atlas-iac/docs/hermes_suite_capacity_20260930.md

6.2 KiB

Suite review capacity fix

The latest failed 363-case job was 3707c784abdb48a3ac625b2f965d1310, created 2026-09-30 at 04:06:17.079260 UTC. Its persisted error was pass_capacity in decision_audit, after four successful model invocations. The previous progress display was stale: large-family review succeeded, then audit admission failed before launching a CLI process.

Established failure

The audit's complete input, system instructions and schema occupied 638,566 UTF-8 bytes. The conservative capacity check added 8,192 harness tokens and reserved six possible 64,000-token CLI outputs: 1,030,758 against 1,000,000. This was a server reservation failure, not measured provider context exhaustion. No tokenizer measurement was available. Audit output reservation was 43,920, within the 64,000 output setting. Transport size was within the 1 MiB bound.

The audit had no subprocess exit/signal, terminal event, provider stop reason, turn count, output usage or timeout: no audit CLI process was launched. Do not attribute the preceding successful review's diagnostics to this rejected audit.

Completed pass Seconds CLI turns Input tokens Output tokens CLI USD estimate
proposal_a 451.994 2 159912 54847 1.736588
proposal_b 460.241 3 372684 57051 2.631756
reconciliation 465.694 2 199517 60783 2.013728
large_family_review 302.696 2 220788 41005 1.703252

All four exited zero, terminal subtype success, structured output at result.structured_output, observed stop reason tool_use, no compaction. Counts aggregate multiple CLI turns; they are not individual context occupancy. Total usage: 952901 input and 213686 output tokens, including 143418 thinking tokens. Runtime validation confirmed firstParty claude-opus-5-5, native CLI 2.1.285, xhigh effort, reported context 1M and model output ceiling 128K. The configured output allowance remained 64K per turn and turn ceiling six.

The job spent 1680.696 seconds and USD 8.085324 estimated, leaving 1919.3 seconds and USD 21.914676. Time and cost were not implicated. At reconciliation, budgeting each remaining pass at the slowest/costliest completed pass plus 30 percent gives 3002.112 seconds and USD 15.139259. With the successful review now measured, the same method projects 2790.307 seconds and USD 13.932204. Keep 3600 seconds and USD 30 for variance; these are CLI estimates, not a claim of subscription charges.

Changes

  • Reserve four to six full 64K outputs according to the complete pass size. The exact failed audit now admits five turns: reservation 966758, headroom 33242. Every invocation retains at least four turns for structured-output repair.
  • If a whole review still exceeds capacity, automatically batch whole already reconciled natural families. Both proposals and reconciliation still consider the complete suite. Every family member retains all supplied source fields.
  • Admit every review batch before launching it; maximum eight batches per review stage, nineteen calls total, under the same single deadline and cost budget. Combine and validate exact global coverage, original family boundaries, names, evidence and final five-case packaging before marking the job complete.
  • Report the attempted stage before capacity checks. Safe failure details now include the capacity reason, input byte/bound counts, context/output limits and available/configured/minimum turns. Per-pass metadata records actual turn allowance, reservation/headroom and optional review_batch information. CLI diagnostics add maximum observed per-message input/output counts when available. Prompt/response bodies and raw provider errors remain excluded.

This is bounded review batching, not unlimited input support. An initial proposal or reconciliation that cannot fit, or one natural family too large to review intact, still fails explicitly. No source text is silently compacted or truncated; no independent arbitrary batches become permanent grouping boundaries.

Verification and limits

181 focused local tests passed. The native pinned CLI passed a loopback mock transport regression at the exact failed 638566-byte audit size, including an injected missing assignment repaired in three CLI turns under the new five-turn ceiling. Complete input and system instructions reached the mock transport; all five stages completed, with 90 final tasks and normal validation. This is a parser, transport and capacity regression, not a hosted grouping-quality benchmark.

A separate hosted 363-case synthetic stress run was stopped at the user's request during large-family review. Its first three passes finished; their cumulative CLI estimate was USD 3.00636. The in-flight call's terminal usage was not obtained. It is not a completed acceptance test. No real case material was rerun. Matching the failed source's total bytes did not establish equal semantic difficulty or field-length distribution, which the historical record did not retain.

Client handoff

Endpoint, authentication, routing restrictions, provider/model, request fields, required result.groups shape and 45-second submission/poll HTTP timeout stay unchanged. Review evidence remains in the authorized result's top-level review_summary; operational metadata is available in job status and passes.

AI_MAX_SECONDS = 3600
AI_MAX_COST_USD = 30
AI_JOB_REVISION = "suite-v10-capacity-20260930-1"

Use a fresh Idempotency-Key for the new real attempt. Reusing the failed key returns the failed job; it does not rerun under the new revision.

  • Configuration: suite-v6-20260929 (unchanged HTTP compatibility identifier).
  • Execution: suite-multipass-v10-20260930.
  • Prompt: implementation-proximity-multipass-v5-20260930.
  • Policy: implementation-five-v1-20260929.
  • Capacity: suite-context-turns-v1-20260930.

The prompt revision adds only the bounded review-batch scope instruction; the success-criteria-first implementation objective remains. The base prompt hash is unchanged; per-pass system hashes identify the actual stage instructions.

Rollback: when no jobs are active, revert the capacity-fix Git commit and reconcile the hermes Flux Kustomization. Do not edit live ConfigMaps or deployments. A rollout clears volatile results, so save wanted completed results first.