113 lines
6.2 KiB
Markdown
113 lines
6.2 KiB
Markdown
# Suite review capacity fix
|
|
|
|
The latest failed 363-case job was `3707c784abdb48a3ac625b2f965d1310`, created
|
|
2026-09-30 at 04:06:17.079260 UTC. Its persisted error was `pass_capacity` in
|
|
`decision_audit`, after four successful model invocations. The previous progress
|
|
display was stale: large-family review succeeded, then audit admission failed
|
|
before launching a CLI process.
|
|
|
|
## Established failure
|
|
|
|
The audit's complete input, system instructions and schema occupied 638,566
|
|
UTF-8 bytes. The conservative capacity check added 8,192 harness tokens and
|
|
reserved six possible 64,000-token CLI outputs: 1,030,758 against 1,000,000.
|
|
This was a server reservation failure, not measured provider context exhaustion.
|
|
No tokenizer measurement was available. Audit output reservation was 43,920,
|
|
within the 64,000 output setting. Transport size was within the 1 MiB bound.
|
|
|
|
The audit had no subprocess exit/signal, terminal event, provider stop reason,
|
|
turn count, output usage or timeout: no audit CLI process was launched. Do not
|
|
attribute the preceding successful review's diagnostics to this rejected audit.
|
|
|
|
| Completed pass | Seconds | CLI turns | Input tokens | Output tokens | CLI USD estimate |
|
|
| --- | ---: | ---: | ---: | ---: | ---: |
|
|
| proposal_a | 451.994 | 2 | 159912 | 54847 | 1.736588 |
|
|
| proposal_b | 460.241 | 3 | 372684 | 57051 | 2.631756 |
|
|
| reconciliation | 465.694 | 2 | 199517 | 60783 | 2.013728 |
|
|
| large_family_review | 302.696 | 2 | 220788 | 41005 | 1.703252 |
|
|
|
|
All four exited zero, terminal subtype `success`, structured output at
|
|
`result.structured_output`, observed stop reason `tool_use`, no compaction.
|
|
Counts aggregate multiple CLI turns; they are not individual context occupancy.
|
|
Total usage: 952901 input and 213686 output tokens, including 143418 thinking
|
|
tokens. Runtime validation confirmed firstParty `claude-opus-5-5`, native CLI
|
|
2.1.285, xhigh effort, reported context 1M and model output ceiling 128K. The
|
|
configured output allowance remained 64K per turn and turn ceiling six.
|
|
|
|
The job spent 1680.696 seconds and USD 8.085324 estimated, leaving 1919.3 seconds
|
|
and USD 21.914676. Time and cost were not implicated. At reconciliation, budgeting
|
|
each remaining pass at the slowest/costliest completed pass plus 30 percent gives
|
|
3002.112 seconds and USD 15.139259. With the successful review now measured, the
|
|
same method projects 2790.307 seconds and USD 13.932204. Keep 3600 seconds and
|
|
USD 30 for variance; these are CLI estimates, not a claim of subscription charges.
|
|
|
|
## Changes
|
|
|
|
- Reserve four to six full 64K outputs according to the complete pass size.
|
|
The exact failed audit now admits five turns: reservation 966758, headroom
|
|
33242. Every invocation retains at least four turns for structured-output repair.
|
|
- If a whole review still exceeds capacity, automatically batch whole already
|
|
reconciled natural families. Both proposals and reconciliation still consider
|
|
the complete suite. Every family member retains all supplied source fields.
|
|
- Admit every review batch before launching it; maximum eight batches per review
|
|
stage, nineteen calls total, under the same single deadline and cost budget.
|
|
Combine and validate exact global coverage, original family boundaries, names,
|
|
evidence and final five-case packaging before marking the job complete.
|
|
- Report the attempted stage before capacity checks. Safe failure details now
|
|
include the capacity reason, input byte/bound counts, context/output limits
|
|
and available/configured/minimum turns. Per-pass metadata records actual turn
|
|
allowance, reservation/headroom and optional `review_batch` information.
|
|
CLI diagnostics add maximum observed per-message input/output counts when
|
|
available. Prompt/response bodies and raw provider errors remain excluded.
|
|
|
|
This is bounded review batching, not unlimited input support. An initial proposal
|
|
or reconciliation that cannot fit, or one natural family too large to review
|
|
intact, still fails explicitly. No source text is silently compacted or truncated;
|
|
no independent arbitrary batches become permanent grouping boundaries.
|
|
|
|
## Verification and limits
|
|
|
|
181 focused local tests passed. The native pinned CLI passed a loopback mock
|
|
transport regression at the exact failed 638566-byte audit size, including an
|
|
injected missing assignment repaired in three CLI turns under the new five-turn
|
|
ceiling. Complete input and system instructions reached the mock transport; all
|
|
five stages completed, with 90 final tasks and normal validation. This is a parser,
|
|
transport and capacity regression, not a hosted grouping-quality benchmark.
|
|
|
|
A separate hosted 363-case synthetic stress run was stopped at the user's request
|
|
during large-family review. Its first three passes finished; their cumulative CLI
|
|
estimate was USD 3.00636. The in-flight call's terminal usage was not obtained.
|
|
It is not a completed acceptance test. No real case material was rerun. Matching
|
|
the failed source's total bytes did not establish equal semantic difficulty or
|
|
field-length distribution, which the historical record did not retain.
|
|
|
|
## Client handoff
|
|
|
|
Endpoint, authentication, routing restrictions, provider/model, request fields,
|
|
required `result.groups` shape and 45-second submission/poll HTTP timeout stay
|
|
unchanged. Review evidence remains in the authorized result's top-level
|
|
`review_summary`; operational metadata is available in job status and `passes`.
|
|
|
|
```python
|
|
AI_MAX_SECONDS = 3600
|
|
AI_MAX_COST_USD = 30
|
|
AI_JOB_REVISION = "suite-v10-capacity-20260930-1"
|
|
```
|
|
|
|
Use a fresh `Idempotency-Key` for the new real attempt. Reusing the failed key
|
|
returns the failed job; it does not rerun under the new revision.
|
|
|
|
- Configuration: `suite-v6-20260929` (unchanged HTTP compatibility identifier).
|
|
- Execution: `suite-multipass-v10-20260930`.
|
|
- Prompt: `implementation-proximity-multipass-v5-20260930`.
|
|
- Policy: `implementation-five-v1-20260929`.
|
|
- Capacity: `suite-context-turns-v1-20260930`.
|
|
|
|
The prompt revision adds only the bounded review-batch scope instruction; the
|
|
success-criteria-first implementation objective remains. The base prompt hash
|
|
is unchanged; per-pass system hashes identify the actual stage instructions.
|
|
|
|
Rollback: when no jobs are active, revert the capacity-fix Git commit and reconcile
|
|
the `hermes` Flux Kustomization. Do not edit live ConfigMaps or deployments. A
|
|
rollout clears volatile results, so save wanted completed results first.
|