atlas-iac/docs/hermes_suite_multipass.md

336 lines
21 KiB
Markdown

# Multi-pass implementation grouping
The suite planner keeps the same HTTPS endpoint, request fields, authentication,
provider permissions, generalized-data approval, and asynchronous job lifecycle.
The required result object remains `result.groups`, with `name`, `description`,
and `members`. Final groups now contain **one to five** cases and names are at
most **64 characters**, unique after whitespace and case normalization.
## Revisions and process
- Configuration: `suite-v6-20260929` (HTTP compatibility identifier).
- Policy: `implementation-five-v1-20260929`.
- Prompt: `implementation-proximity-multipass-v4-20260929`.
- Execution: `suite-multipass-v7-20260929`.
A single server-side job performs:
1. Proposal A on the complete suite, in stable alias order, without a size cap.
2. Proposal B in a reproducible SHA-256 alias order, in a separate fresh CLI
invocation. It sees no proposal A or earlier conversation. Both use exactly
the same implementation-effort system instructions and output schema.
3. Whole-suite reconciliation using both memberships and all original fields.
Success criteria remain primary evidence. There is no majority voting or
transitive closure of pairwise similarity. Natural families are still uncapped.
4. If any reconciled family exceeds five, review every such original family using
all case fields, distinguishing different machinery from cheap variations.
5. One additional bounded audit of those decisions, including descendants that
were split into smaller groups. It may reverse an unjustified split or refine
a broad family, but cannot cross the reconciled family boundaries. There is
no agreement loop.
The final two passes record explicit keep/split rationales and exact short
source-field quotations. The code verifies that evidence belongs to an assigned
alias and is present verbatim in that source field. This checks support provenance,
not the correctness of a model's engineering interpretation. Unknown implementation
details remain concise uncertainty, not invented equipment or procedures.
The internal structured-output schema requires one property for each exact source
alias. Its value selects a natural-family index and, in review passes, a variation
index. The server reconstructs memberships and runs the normal independent checks.
Review decisions also use required keys derived from exact original memberships.
This avoids free-form membership lists silently omitting or repeating aliases.
The public groups schema is unchanged. Invalid keys, indexes, empty groups, and
missing review decisions still fail closed; no missing assignment is fabricated.
Each invocation retains six CLI turns for structured output. Model review passes
and CLI turns are separate counters. All calls use the originally selected provider
and pinned model. The current default is `claude-opus-5-5[1m]`, with canonical runtime
identity checked as `claude-opus-5-5`, firstParty, reported 1M context
and 128K model output ceiling, with requests limited to 64K output tokens.
The server invokes native Claude Code CLI 2.1.285 using
the existing first-party OAuth account, not a separately configured API-key
account. The CLI itself communicates with Anthropic over HTTPS. No new provider
fallback or tools are enabled.
Effort defaults to `high` and increases to `xhigh` for suites with at least 100
cases OR 128 KiB of canonical UTF-8 source JSON (campaign, suite, and all cases).
Preflight selects the effort once for the whole job; every discovery, review,
audit and structured-output repair uses that selection. Source size excludes
transport whitespace, prompts, schemas, and generated proposals. Thresholds are
published in capabilities; selection metadata includes counts, bytes, and triggers.
Existing deadlines and cost guards still apply and never cause a silent effort
reduction. See the [effort policy details](hermes_suite_claude55.md).
Sonnet 5.5 is also verified on the 14-case synthetic fixture and available through
server configuration. See the [matched comparison](hermes_suite_claude55.md).
Earlier acceptance measurements below used Opus 4.8; they do not establish large
suite capacity or quality for the new models.
## Deterministic task sizing and names
For each remaining coherent large family, `k = ceil(n / 5)` and
`base, remainder = divmod(n, k)`. The first remainder parts have base+1 cases;
the rest have base. Thus 6 becomes 3+3, 7 becomes 4+3, 11 becomes 4+4+3, and
14 becomes 5+5+4. No leftover singleton is created by a five-at-a-time slice.
The semantic review supplies internal variation sets. A deterministic packer uses
those hints and stable complete-record hashes/aliases to fill the calculated
sizes. Equivalent content remains distinct by alias. Hint order, source order,
and member order do not cause gratuitous reshuffling. Different model decisions
can still change membership between fresh jobs; determinism is conditional on the
same natural families and variation hints, not a claim that inference is deterministic.
Natural names must be concise, meaningful, distinct, and at most 56 characters,
leaving room for numbering within the 64-character final limit. Collisions or
invalid names fail validation; the service never truncates them or adds numbers
to unrelated families. Capacity-only parts have the exact same base name:
`Reset recovery (1/3)`, `Reset recovery (2/3)`, `Reset recovery (3/3)`.
Their descriptions identify a work-size division and use only the reviewed common
work applicable to all members. No other part's specific objectives are copied.
Every intermediate partition and final result is checked for exact alias coverage.
Final size, names, balanced part sizes, and campaign/suite ownership are checked
programmatically. Incomplete or invalid work is never returned as completed.
## Review artifact and operational metadata
`GET /suite-planning/v1/jobs/<id>/result` returns an optional top-level
**`review_summary`**, alongside the unchanged **`result`** object. It contains:
- `proposal_disagreements`: affected aliases, differing pair count, and both full
membership partitions in compressed form. Group names/numbers are not compared.
- `reconciled_families`: the first natural partition with concise rationale,
common work, exact source-field support, and uncertainty.
- `large_family_review` and `decision_audit`: explicit decisions for each original
oversized membership set, including semantic-split versus coherent-keep reasons.
- `natural_families`: the final conceptual partition before task sizing.
- `capacity_divisions`: base names, calculated sizes, final part names, common
implementation rationale, and remaining uncertainty.
- `unresolved_uncertainties` and `counts`: separate natural-family/final-task counts
and natural/final singleton counts. Size compliance is not semantic-quality evidence.
Review content lives only in the authorized in-memory result cache, with the same
one-hour TTL and restart loss as results. It is absent from routine logs, status
responses, and SQLite job metadata. It contains source-derived material and exact
quotations, so it is not an export-safe ClickUp artifact. No new hierarchy level
or client request fields are introduced.
Status/result metadata includes `policy_revision`, `execution_revision`,
`prompt_revision`, `model_pass_count`, `natural_family_count`, `final_task_count`,
`singleton_statistics`, aggregate usage/cost, and safe per-invocation `passes`.
Each pass records its stage, actual model, system/schema hashes, ordering hash,
capacity check, remaining budgets, usage, timing, and CLI diagnostics. The top-level
`cli_diagnostics` concerns the last invocation; `turns` is the sum of reported CLI
turns, while `model_pass_count` counts complete model review invocations.
`execution_progress.current_pass` reports the current stage during a running job.
While a CLI invocation runs, progress updates every five seconds with `heartbeat_at`,
`pass_elapsed_seconds`, `job_elapsed_seconds`, `job_remaining_seconds`,
`completed_model_passes`, `maximum_model_passes`, `cli_running`, `cli_output_bytes`,
and `last_cli_activity_seconds_ago`. The heartbeat means the worker is alive; a
change in output bytes means CLI activity. Neither is a percentage complete or
proof that a particular case has been reasoned about. No event bodies or reasoning
text are exposed. `cost_used_usd_estimate` counts completed invocations; the
in-flight invocation is not included until its terminal usage is available.
Poll the existing status URL every five seconds to display progress.
To inspect progress from WSL, keep the existing bearer token private and use the
job ID returned by submission. This command only reads status:
```bash
curl -q --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
--connect-timeout 10 --max-time 45 --fail-with-body --silent --show-error \
--config <(printf 'header = "Authorization: Bearer %s"\n' "$SUITE_PLANNING_TOKEN") \
"https://worker.bstein.dev/suite-planning/v1/jobs/$JOB_ID"
```
An existing Python poller can display `envelope["execution_progress"]` after each
status read. No separate pass submissions, new request fields, or client-side
reconciliation are required. Prefer the actual stage and elapsed time over an
invented percentage, because different passes have very different runtimes.
## Shared bounds and failure behavior
The user-approved 1800-second maximum job deadline and USD 30 CLI estimated-cost
guard apply to **all calls combined**, including routing time. These are also the
defaults when omitted. Explicit lower client limits remain effective. This guard
uses Claude CLI estimated-cost accounting; it is not subscription billing. Each invocation receives
only the remaining time and estimated-cost allowance. Available usage/cost is
accumulated after every call; unknown cost accounting stops further hosted calls.
No partial partition substitutes for unfinished passes. The CLI's estimated-cost
ceiling can overshoot within one generation, as before; this is not a guaranteed
subscription billing limit. An observed overrun fails the job rather than allowing
further calls or reporting successful completion.
Request and complete response envelopes stay bounded at 1 MiB. Every expanded pass
is checked before dispatch, including original source, proposals/reviews, system
instructions, and schema. The same conservative byte bound and six possible CLI
outputs are reserved against context. Exact tokenization and output reservation
remain unverified estimates; actual incomplete generation fails explicitly.
Preflight checks initial and minimum reconciliation capacity; it identifies that
actual later-pass capacity will be checked before each invocation. It does not
claim to know model-generated proposal sizes in advance.
The existing 8K-context local route cannot admit the new reconciliation schema and
instructions. Local-only requests fail with capacity errors and never start hosted
jobs. The separate `/local-model` endpoint and its limits are unchanged. No bounds
other than the explicitly approved time/cost guards were increased, and no provider
was switched after an error.
Failures include existing CLI diagnostics and a fixed `review_pass`, completed-pass
count, and available aggregate usage. New explicit codes include `pass_capacity`,
`pass_request_too_large`, `job_time_budget_exhausted`, `job_cost_budget_exhausted`,
`budget_accounting_unavailable`, and bounded review/assignment validation errors.
Raw provider messages, source text, and generated explanations are not error details.
The laptop submits one request with its existing scoped token, then polls the same
job. Retry a timed-out submission using the same key/body; an intentional fresh
attempt needs a new Idempotency-Key. Include policy/prompt/execution revisions in
local cache provenance. No real suite is submitted by deployment checks.
## Verification and rollback
Synthetic unit tests cover the balanced-size examples, semantic subdivisions and
a bounded reversal, differing proposals, duplicate text, name normalization and
collisions, source support, deadline/cost exhaustion, cancellation, privacy,
authorization, idempotency, and 14/75/363-case capacity and sizing. Native CLI
loopback and authenticated LAN/provider acceptance measurements are recorded after
deployment. Mocked outputs alone do not establish model quality or throughput.
Deploy or roll back through Git and Flux. Revert only the multi-pass commits to
restore the preceding single-pass policy while preserving the earlier CLI diagnostic
fix. Wait for active jobs before a worker restart because completed results are
held in memory. No Vault credential or ingress/routing changes are required.
The policy implementation is commit `7af4f4f2`; the approved 20-minute / USD 10
estimate guard and progress reporting are commit `15c05cb7`. Required-alias
structured output is commit `9c327c6f`. To restore the prior
single-pass implementation after active jobs finish, revert these through the normal
Git deployment branch, preserving the earlier six-turn diagnostics fix:
```bash
git revert --no-commit 9c327c6f 15c05cb7 7af4f4f2
git commit -m "hermes: restore prior suite grouping policy"
git push origin HEAD:main
flux reconcile kustomization hermes --namespace flux-system --with-source
```
The first live 363-case multi-pass run at the previous 900-second / USD 5
guards stopped after 747.513 seconds during large-family review. Its actual CLI
final subtype was `error_max_budget_usd`; three full passes had completed. The CLI
reported an aggregate estimate of 5.62765, showing the documented within-generation
overshoot. It was not a timeout, turn-limit failure, or accepted partial answer.
The adapter now reports this condition as `job_cost_budget_exhausted`.
## Synthetic acceptance measurements
The suite inputs are synthetic fixtures, never roster-derived material. Tests ran
through authenticated TLS from `titan-jh` (`192.168.22.8`) to the LAN ingress
`192.168.22.50:443`, with hostname verification, proxy bypass, and no redirects.
This verifies that LAN test path; it does not replace a WSL connectivity test.
| Cases | Request bytes | Wall seconds | Model passes / CLI turns | Natural families | Final tasks | Natural / final singletons | Input / output tokens |
| --- | ---: | ---: | --- | ---: | ---: | --- | --- |
| 14 | 9,847 | 61.805 | 3 / 6 | 9 | 9 | 4 / 4 | 18,216 / 5,794 |
| 75 | 55,984 | 274.620 | 5 / 12 | 9 | 21 | 3 / 3 | 174,942 / 28,950 |
| 38 balanced fixture | 23,871 | 231.337 | 5 / 15 | 4 | 10 | 0 / 0 | 139,789 / 24,681 |
| 363, current required-alias schema | 273,763 | 842.777 | 5 / 12 | 9 | 75 | 3 / 3 | 983,491 / 91,932 |
| 363, old limits | 273,761 | 747.513 | 3 completed; pass 4 failed | Not finalized | No result | Not finalized | 629,141 / 64,424 |
Token totals include all completed CLI turns across the job, not unique source
size. The old-limit 363 result includes reported failed-invocation usage and was
never returned as completed. The earlier 14/75/38 and old-limit 363 CLI cost
estimates were 0.23593, 1.59846, 1.31597, and 5.62765 respectively; these are not subscription bills.
The current 363-case run reported a CLI cost estimate of 7.215755 and a complete
response envelope of 122,115 bytes. Its five passes took 80.962, 98.247, 350.114,
148.165, and 165.211 seconds, with 2/2/4/2/2 CLI turns. Six coherent 60-case
families became twelve five-case tasks each, plus three genuine singletons.
The 135 authorized status samples covered all five stages and recorded 59 changes
in CLI output activity. No compaction or truncation was reported. No terminal rate-limit
error or gateway job retry was reported; structured-output CLI turns are recorded
separately and do not imply provider-network retries.
For every completed run above, natural-family pair precision and recall were
1.0 against the independently defined implementation patterns, exact alias
coverage passed, and final groups stayed pure to those patterns. No incorrect
merges, unnecessary semantic splits, or cross-family description attribution
were observed during review. Equivalent text under separate aliases remained
separate objectives. Natural-family quality was scored before capacity division;
a cap-compliant result alone does not prove semantic correctness.
The 38-case fixture interleaves four reset-related mechanisms with shared generic
subject descriptions and distinct success criteria: RPC/stub observations, pulse
and oscilloscope timing, offline report parsing, and concurrent queue instrumentation.
It yielded natural sizes 6/7/11/14 and final sizes 3+3 / 4+3 / 4+4+3 / 5+5+4.
Thirty-eight authorized status samples showed five stages, changing heartbeats,
and nineteen CLI activity changes without exposing event bodies.
The actual 14/75/38/363 independent proposals agreed on membership. Disagreement
reconciliation, genuine semantic subdivision, reversal of an unjustified split,
and naming-collision rejection are separately tested with controlled mocked
responses; do not describe those as observed live-provider disagreements.
The public 363 fixture has description lengths 48-245 characters (mean 234.5),
preconditions 67-143 (mean 137.2), success criteria 73-135 (mean 131.2), and case
labels 7-15. It interleaves six mechanisms, includes identical distant case text,
and has distinctive beginning/middle/end objectives. These measurements describe
the tested material; longer or more ambiguous real suites may take more work.
An installed native CLI loopback probe captured all complete case fields and full
system instructions in every emitted provider request for 14/75/363 cases, with
64,000 output limits. It performed 3/5/5 passes and deterministic final sizing to
9/21/75 tasks using mock responses. This verifies transport and orchestration,
not live semantic quality. Live successful results reported no compaction or
truncation and passed exact coverage and source-quote support checks. These checks
cannot prove the provider internally attended to every field; exact tokenization
and sufficient output reservation for arbitrary suites remain unverified.
The updated code passes 109 synthetic/mocked regression tests. Kubernetes render,
client dry-run, and Flux diff checks passed before deployment. Policy controls,
case coverage validation, scoped authentication, and provider pinning remain in
place. The invalid-token check returned 401 and same-key submission replay reused
the same job. No automatic real-data rerun was performed.
The second 363-case attempt, under the 1200-second / USD 10 guards, failed after
228.260 seconds in proposal B with `invalid_case_assignments`. Both invocations
returned successful CLI results, but the second partition failed server membership
validation. Its reported total was 281,227 input / 23,796 output tokens and 2.001035
in CLI estimated-cost accounting. The previous validator did not retain counts of
missing, repeated, or unknown assignments, and transient output was cleaned; the
specific mismatch cannot be recovered. No partial partition was accepted. The new
required-alias schema and count-only failure diagnostics address this failure mode.
The updated native CLI loopback test deliberately omitted one required alias from
its first 14-case structured output. The installed CLI rejected that tool payload
and requested a correction, then the complete job passed. This demonstrates actual
schema-repair behavior within the existing six-turn headroom. The updated
14/75/363 transport runs made 4/5/5 mock provider requests across 3/5/5 model passes;
all transmitted the full source and system instructions. No hosted inference was
used for this transport/repair test.
The 14/75 and balanced-38 live measurements preceded the required-alias internal
serialization change. The final serialization was exercised by native loopback at
all three public sizes and by the successful hosted 363-case five-pass job. The
engineering objective, model, effort, cap, and public result shape stayed the same.
The complete suite is supported directly for the tested 273,763-byte request.
This is evidence for this fixture, not a guarantee that every 363-case request
will fit the same time/output allowance. There is no remaining blocker for starting
a new laptop attempt with the approved Claude credential and adequate job limits.
The result still requires the client's independent coverage audit and engineering
review. No real job was run by these checks.
Full synthetic results, review summaries, safe per-pass diagnostics, input/response
sizes, source-field length distributions, failed attempts, and native transport
measurements are in [the acceptance evidence](evidence/hermes_suite_multipass_20260929.json).
Routine planner logs contained only allowed operational fields, the completed
job's SQLite metadata contained no result/review/case content, and temporary CLI
job directories were empty after completion.
Execution revision `suite-multipass-v4-20260929` raises only the shared job defaults
and maximums to 1800 seconds and USD 30 in CLI estimate accounting. Explicit lower
client limits remain effective. The existing request fields are unchanged. No
inference or regression tests were rerun for this limit-only update, as requested;
the acceptance measurements above retain their original revisions and bounds.