atlas-iac/docs/hermes_suite_multipass.md

168 lines
10 KiB
Markdown

# Multi-pass implementation grouping
The suite planner keeps the same HTTPS endpoint, request fields, authentication,
provider permissions, generalized-data approval, and asynchronous job lifecycle.
The required result object remains `result.groups`, with `name`, `description`,
and `members`. Final groups now contain **one to five** cases and names are at
most **64 characters**, unique after whitespace and case normalization.
## Revisions and process
- Configuration: `suite-v6-20260929` (HTTP compatibility identifier).
- Policy: `implementation-five-v1-20260929`.
- Prompt: `implementation-proximity-multipass-v3-20260929`.
- Execution: `suite-multipass-v2-20260929`.
A single server-side job performs:
1. Proposal A on the complete suite, in stable alias order, without a size cap.
2. Proposal B in a reproducible SHA-256 alias order, in a separate fresh CLI
invocation. It sees no proposal A or earlier conversation. Both use exactly
the same implementation-effort system instructions and output schema.
3. Whole-suite reconciliation using both memberships and all original fields.
Success criteria remain primary evidence. There is no majority voting or
transitive closure of pairwise similarity. Natural families are still uncapped.
4. If any reconciled family exceeds five, review every such original family using
all case fields, distinguishing different machinery from cheap variations.
5. One additional bounded audit of those decisions, including descendants that
were split into smaller groups. It may reverse an unjustified split or refine
a broad family, but cannot cross the reconciled family boundaries. There is
no agreement loop.
The final two passes record explicit keep/split rationales and exact short
source-field quotations. The code verifies that evidence belongs to an assigned
alias and is present verbatim in that source field. This checks support provenance,
not the correctness of a model's engineering interpretation. Unknown implementation
details remain concise uncertainty, not invented equipment or procedures.
Each invocation retains six CLI turns for structured output. Model review passes
and CLI turns are separate counters. All calls use the originally selected provider
and pinned model. The model remains `claude-opus-4-8[1m]`, with canonical runtime
identity checked as `claude-opus-4-8`, firstParty, medium effort, reported 1M context
and 64K output ceiling. No new provider fallback or tools are enabled.
## Deterministic task sizing and names
For each remaining coherent large family, `k = ceil(n / 5)` and
`base, remainder = divmod(n, k)`. The first remainder parts have base+1 cases;
the rest have base. Thus 6 becomes 3+3, 7 becomes 4+3, 11 becomes 4+4+3, and
14 becomes 5+5+4. No leftover singleton is created by a five-at-a-time slice.
The semantic review supplies internal variation sets. A deterministic packer uses
those hints and stable complete-record hashes/aliases to fill the calculated
sizes. Equivalent content remains distinct by alias. Hint order, source order,
and member order do not cause gratuitous reshuffling. Different model decisions
can still change membership between fresh jobs; determinism is conditional on the
same natural families and variation hints, not a claim that inference is deterministic.
Natural names must be concise, meaningful, distinct, and at most 56 characters,
leaving room for numbering within the 64-character final limit. Collisions or
invalid names fail validation; the service never truncates them or adds numbers
to unrelated families. Capacity-only parts have the exact same base name:
`Reset recovery (1/3)`, `Reset recovery (2/3)`, `Reset recovery (3/3)`.
Their descriptions identify a work-size division and use only the reviewed common
work applicable to all members. No other part's specific objectives are copied.
Every intermediate partition and final result is checked for exact alias coverage.
Final size, names, balanced part sizes, and campaign/suite ownership are checked
programmatically. Incomplete or invalid work is never returned as completed.
## Review artifact and operational metadata
`GET /suite-planning/v1/jobs/<id>/result` returns an optional top-level
**`review_summary`**, alongside the unchanged **`result`** object. It contains:
- `proposal_disagreements`: affected aliases, differing pair count, and both full
membership partitions in compressed form. Group names/numbers are not compared.
- `reconciled_families`: the first natural partition with concise rationale,
common work, exact source-field support, and uncertainty.
- `large_family_review` and `decision_audit`: explicit decisions for each original
oversized membership set, including semantic-split versus coherent-keep reasons.
- `natural_families`: the final conceptual partition before task sizing.
- `capacity_divisions`: base names, calculated sizes, final part names, common
implementation rationale, and remaining uncertainty.
- `unresolved_uncertainties` and `counts`: separate natural-family/final-task counts
and natural/final singleton counts. Size compliance is not semantic-quality evidence.
Review content lives only in the authorized in-memory result cache, with the same
one-hour TTL and restart loss as results. It is absent from routine logs, status
responses, and SQLite job metadata. It contains source-derived material and exact
quotations, so it is not an export-safe ClickUp artifact. No new hierarchy level
or client request fields are introduced.
Status/result metadata includes `policy_revision`, `execution_revision`,
`prompt_revision`, `model_pass_count`, `natural_family_count`, `final_task_count`,
`singleton_statistics`, aggregate usage/cost, and safe per-invocation `passes`.
Each pass records its stage, actual model, system/schema hashes, ordering hash,
capacity check, remaining budgets, usage, timing, and CLI diagnostics. The top-level
`cli_diagnostics` concerns the last invocation; `turns` is the sum of reported CLI
turns, while `model_pass_count` counts complete model review invocations.
`execution_progress.current_pass` reports the current stage during a running job.
While a CLI invocation runs, progress updates every five seconds with `heartbeat_at`,
`pass_elapsed_seconds`, `job_elapsed_seconds`, `job_remaining_seconds`,
`completed_model_passes`, `maximum_model_passes`, `cli_running`, `cli_output_bytes`,
and `last_cli_activity_seconds_ago`. The heartbeat means the worker is alive; a
change in output bytes means CLI activity. Neither is a percentage complete or
proof that a particular case has been reasoned about. No event bodies or reasoning
text are exposed. Poll the existing status URL every five seconds to display it.
## Shared bounds and failure behavior
The user-approved 1200-second maximum job deadline and USD 10 CLI estimated-cost
guard apply to **all calls combined**, including routing time. These are also the
defaults when omitted. Explicit lower client limits remain effective. This guard
uses Claude CLI estimated-cost accounting; it is not subscription billing. Each invocation receives
only the remaining time and estimated-cost allowance. Available usage/cost is
accumulated after every call; unknown cost accounting stops further hosted calls.
No partial partition substitutes for unfinished passes. The CLI's estimated-cost
ceiling can overshoot within one generation, as before; this is not a guaranteed
subscription billing limit. An observed overrun fails the job rather than allowing
further calls or reporting successful completion.
Request and complete response envelopes stay bounded at 1 MiB. Every expanded pass
is checked before dispatch, including original source, proposals/reviews, system
instructions, and schema. The same conservative byte bound and six possible CLI
outputs are reserved against context. Exact tokenization and output reservation
remain unverified estimates; actual incomplete generation fails explicitly.
Preflight checks initial and minimum reconciliation capacity; it identifies that
actual later-pass capacity will be checked before each invocation. It does not
claim to know model-generated proposal sizes in advance.
The existing 8K-context local route cannot admit the new reconciliation schema and
instructions. Local-only requests fail with capacity errors and never start hosted
jobs. The separate `/local-model` endpoint and its limits are unchanged. No bounds
other than the explicitly approved time/cost guards were increased, and no provider
was switched after an error.
Failures include existing CLI diagnostics and a fixed `review_pass`, completed-pass
count, and available aggregate usage. New explicit codes include `pass_capacity`,
`pass_request_too_large`, `job_time_budget_exhausted`, `job_cost_budget_exhausted`,
`budget_accounting_unavailable`, and bounded review/assignment validation errors.
Raw provider messages, source text, and generated explanations are not error details.
The laptop submits one request with its existing scoped token, then polls the same
job. Retry a timed-out submission using the same key/body; an intentional fresh
attempt needs a new Idempotency-Key. Include policy/prompt/execution revisions in
local cache provenance. No real suite is submitted by deployment checks.
## Verification and rollback
Synthetic unit tests cover the balanced-size examples, semantic subdivisions and
a bounded reversal, differing proposals, duplicate text, name normalization and
collisions, source support, deadline/cost exhaustion, cancellation, privacy,
authorization, idempotency, and 14/75/363-case capacity and sizing. Native CLI
loopback and authenticated LAN/provider acceptance measurements are recorded after
deployment. Mocked outputs alone do not establish model quality or throughput.
Deploy or roll back through Git and Flux. Revert only the multi-pass commits to
restore the preceding single-pass policy while preserving the earlier CLI diagnostic
fix. Wait for active jobs before a worker restart because completed results are
held in memory. No Vault credential or ingress/routing changes are required.
The first live 363-case multi-pass run at the previous 900-second / USD 5
guards stopped after 747.513 seconds during large-family review. Its actual CLI
final subtype was `error_max_budget_usd`; three full passes had completed. The CLI
reported an aggregate estimate of 5.62765, showing the documented within-generation
overshoot. It was not a timeout, turn-limit failure, or accepted partial answer.
The adapter now reports this condition as `job_cost_budget_exhausted`.