# Suite recovery and adaptive analysis ## Latest real-job diagnosis Exact failed job: `90ab61bf285e4c89a41d893d4fee649d`, located by descending persisted creation time, failed status and 363-case count. Its request selected 3600 seconds and an estimated-cost guard of 30. No real source was inspected or rerun. The record is not the earlier decision-audit capacity failure. | Measurement | Retained evidence | | --- | --- | | Error | `incomplete_generation`, `failure_stage=cli_final_result` | | Failed pass | `proposal_b`; one completed model pass | | CLI exit / signal | 1 / null; `nonzero_exit` | | Terminal event | `result`, subtype `success`, **is_error=true** | | Provider / assistant stop | `stop_sequence` / `stop_sequence` | | API error indicator | true; allowlisted category `unknown`; HTTP status unknown | | Structured result | Absent from terminal output, tool input and JSON text candidates | | Turns / ceiling | 1 / 6; no turn-limit or structured-retry-limit outcome | | B input / output usage | CLI reported 0 / 0; does not independently prove no provider work occurred | | B observed model capacity | Unknown; failed terminal event had no modelUsage limits | | B admission | Same 439270-byte input/system/schema size as A; conservative bound 447462; six full outputs reserved: 831462 under 1M | | B configured output / reasoning | 64000 per turn / xhigh; reasoning-token limit unknown | | B output bytes | Last heartbeat observed 54366; exact terminal byte count was not retained | | B elapsed / timeout | Last heartbeat 411.8 s; approximately 412.9 s including boundary overhead / 2791.629775 s watchdog | | Provider timeout | Unknown | | Job elapsed / remaining | 1221.233 s / 2378.8 s | | Total input / output | 319884 / 98177; output includes 82581 thinking tokens | | Cost used / remaining old guard | USD 3.243076 / 26.756924 estimated | | Route | Claude CLI, first-party OAuth; selected `claude-opus-5-5[1m]` | | Runtime model evidence | A confirmed firstParty canonical `claude-opus-5-5`; B initialization matched requested model, but no successful terminal canonical-model verification | | CLI | 2.1.285 | | Old configuration / prompt | `suite-v6-20260929` / `implementation-proximity-multipass-v5-20260930` | | Old execution / policy | `suite-multipass-v10-20260930` / `implementation-five-v1-20260929` | Proposal A completed in 808.317 seconds, three CLI turns, exit zero and valid structured output. B produced an unsuccessful API-error event and exited nonzero. The underlying provider cause is **unknown**: raw CLI text was deliberately deleted. No retained evidence identifies rate limiting, context/output exhaustion, timeout, turn exhaustion or an adapter rejecting valid structured content. The old adapter collapsed this API error into `incomplete_generation` and the workflow had no retry. The new classifier treats an unclassified API error as eligible for a bounded retry, with `provider_failure_transience=unconfirmed`; that is not proof it was transient. New diagnostics retain exact output bytes and per-attempt elapsed time. ## Time and cost The maximum accepted `execution.max_seconds` is **7200**; omission still uses 1800. All profile, grouping, review, retry and repair calls share one absolute deadline. Submit/status/result HTTP timeouts remain 45 seconds on the laptop. The cost guard is **removed**, following the user's subsequent instruction. The CLI receives no `--max-budget-usd`. Legacy positive finite `max_cost_usd` values or null are accepted but never enforced. Capability/status metadata reports `cost_guard_enforced=false` and a null cost ceiling. Measured estimates remain informational; unknown usage/cost does not prevent further work. These estimates are not a statement of subscription billing charges. 7200 seconds is the recommended next limit. Five passes at A's 808.317 seconds would total 4041.585 seconds, leaving 3158.415 seconds for variability/recovery. Hierarchical execution has a different call mix, so this is a planning reference, not a guarantee or a measured hierarchical runtime. No fixed deadline guarantees success through a prolonged provider outage. Later stages retain explicit time floors. Global stages reserve `min(500, max(10, job_seconds / 8))` seconds per remaining global stage; pending profile batches reserve 30 seconds each, and source/review batches 60 seconds. These are admission/watchdog reservations, not guaranteed runtimes. Each attempt receives only time remaining after future reservations. Insufficient remaining reservation returns `insufficient_remaining_pass_budget`; the absolute deadline returns `job_time_budget_exhausted`. No per-pass allowance resets the job clock. ## Recovery and checkpoints At most three attempts per logical pass, each a fresh isolated CLI session with the same provider, model, complete pass source and objective. Retry backoff is 2 then 4 seconds, cancellable. CLI turns start at four and can increase to five then six only when complete-input context admission allows. No compaction, truncation, provider fallback or partial result is accepted. Retriable outcomes include unsuccessful/missing terminal generation, CLI turn or structured-output retry exhaustion, rate limiting, temporary transport failures, nonzero transient exits, and distinct pass timeouts. One additional structured assignment/JSON repair attempt is allowed. Successful final JSON text envelopes are recognized after the normal completion, runtime-model and isolation guards; normal schema/membership validation still applies. Authentication/authorization, model/isolation changes, cancellation, invalid requests, confirmed capacity violations and deterministic evidence/review failures are not blindly retried. Capacity or exhausted generation can select hierarchical analysis instead of repeating the same full-suite request. Final errors identify turn exhaustion, generation retries, output/context limits, provider failure, pass timeout or insufficient remaining budget. Status contains attempts, terminal metadata and retained checkpoint names, never generated content. Every **validated** output is checkpointed in active-job memory. Identity binds the complete request hash, credential owner/scope, routing, provider, canonical model, configuration/prompt/policy/execution revisions, exact pass source, schema, objective and relevant prior-output hashes. Altering any input prevents reuse. Completed predecessors survive recovery of a later pass. A matching checkpoint restores a copy without another model call. Contents never enter SQLite, logs, telemetry or ordinary status; all checkpoints are removed when the job ends or its deadline expires (at most 7200 seconds). **Restart limitation:** checkpoints are not durable. A worker restart marks an in-flight job `interrupted_no_retry`; no automatic provider resubmission occurs. A fresh client attempt after restart regenerates its passes. This deliberately uses the requested active-job recovery option instead of adding disk copies of source-derived outputs. Results retain the existing one-hour volatile retention. ## Adaptive execution Small requests prefer direct independent proposals, whole-suite reconciliation, large-family review and decision audit. Initial hierarchical selection occurs for 250+ cases, 256 KiB+ canonical source, or a p95 source field of 8192+ bytes. Direct capacity failure also selects hierarchical mode before inference. Retriable direct generation exhaustion or confirmed output/context limits can switch modes within the same job and remaining deadline; the transition is recorded. Hierarchical mode: 1. Extract profiles in deterministic batches of at most 24 cases or about 48 KiB. A single larger case remains intact and must pass capacity admission. Each profile preserves target, success-criteria work, setup, stimulus, observations, machinery, variations, uncertainty, exact alias and original source evidence. 2. Run two independent proposals over **every profile in the complete suite**, with an alternative reproducible order for the second. Reconcile globally; profile-processing batches never determine family membership. 3. Reopen all original generalized fields for every candidate family in bounded whole-family source reviews. Verify original evidence, split unsupported compatibility, and reject ownership drift. Only then accept natural families. 4. Perform the existing oversized-family review, decision audit, exact coverage, unique short names and deterministic balanced five-case packaging. Profiles have explicit size bounds; they are not arbitrary silent summarization. No profile can be missing, invented, or silently truncated. At the current 400-case service maximum, bounded profiles plus membership-only proposal context fit one global comparison request. A worst-case 400-alias ASCII regression admitted profile reconciliation at 618260 input/system/schema bytes, with five reserved CLI outputs. Every real expanded request is checked again. The bounded contract avoids needing overlapping-neighborhood reconciliation; no such fallback is claimed. A single full-source family that cannot fit intact still fails explicitly, as do excessive profile/source/review batches. Transport and case ceilings remain 1 MiB and 400, not unlimited input support. Bounds: 32 profile batches, 16 source-review batches, eight batches per semantic review, 64 validated logical calls and 128 model invocations including retries. All share the same deadline. Optional metadata includes `execution_mode`, `strategy_events`, `direct_pass_attempts`, `retry_count`, `checkpointed_passes`, `checkpoint_reused`, profile/source batch counts, `cross_batch_review_count`, `model_pass_count`, `natural_family_count`, `final_task_count` and per-attempt measurements. Content-free source byte distributions now support better future fault fixtures. Engineering rationales remain only in authorized `review_summary`. ## Verification 200 focused local tests passed. They cover isolated retries, turn adaptation, assignment repair, non-retriable errors, preserved predecessor checkpoints, checkpoint scope/expiry/restart semantics, global profile coverage, complete source reopening, no fallback, shared deadlines, reserved future time, and cost estimates exceeding the old guard without stopping work. Manifest render and client dry run passed; Flux diff was reviewed. Hosted native Opus 5.5 comparison used the same synthetic 14 cases, interleaved between byte-array decoder tests and physical pulse/reset timing tests. Duplicate text retained distinct aliases. Hierarchical mode deliberately used four profile batches, then global comparisons, full-source review and both semantic reviews. | Hosted execution | Seconds | Calls | Input / output tokens | CLI estimate | Natural / final groups | | --- | ---: | ---: | --- | ---: | --- | | Direct | 79.009 | 5 | 60130 / 9282 | USD 0.426160 | 2 / 4 | | Hierarchical | 148.752 | 10 | 111617 / 17770 | USD 0.801868 | 2 / 4 | Both had exact alias coverage, zero incorrect merge pairs, zero unnecessary semantic split pairs, no retries and valid balanced 4+3 packages per family. Descriptions kept decoder and hardware-capture objectives separate. The direct answer inferred that fault cases varied only in stimulus; the hierarchical answer more accurately retained uncertainty about unspecified fault patterns. This is a small quality check, not proof of real-suite quality or runtime. The 363-case regression used 410322 source bytes, matching the failed record, with 45 interleaved machinery patterns and distant identical records. Actual historical per-field lengths were not retained, so that distribution cannot be matched or claimed equivalent. Both modes completed with **mocked model answers from the independent fixture map**, producing 45 natural families and 90 final tasks, maximum five members, balanced sizes and unique short names. Direct used five calls (175646 response bytes); hierarchical used 26 (224822 response bytes). Mocked execution was about 6.6 seconds per mode; these are local orchestration measurements, not hosted latency, cost or a model-quality result. No new hosted 363-case benchmark or real-suite rerun was performed. Detailed [synthetic evidence](evidence/hermes_suite_recovery_20260930.json) and [the retained safe failure record](evidence/hermes_suite_failure_90ab61bf_metadata.json) are separate. No source fields or raw provider messages from the real suite are included in either artifact. ## Client handoff and revisions No new request fields. `execution.strategy="whole_suite"` still submits one complete suite and receives one combined `result.groups` result. Authentication, endpoint, credential scopes, pinned Opus 5.5, effort thresholds and routing remain. ```python AI_MAX_SECONDS = 7200 AI_JOB_REVISION = "suite-adaptive-v11-20260930-1" ``` Generate a **new Idempotency-Key**. `AI_MAX_COST_USD` may remain at its existing value if the client sends it; no client edit is needed to disable server cost limits. Do not reuse the failed job's key. - Configuration: `suite-v6-20260929`, HTTP compatibility unchanged. - Execution: `suite-adaptive-v11-20260930`. - Prompt: `implementation-proximity-adaptive-v6-20260930`. - Grouping policy: `implementation-five-v1-20260929`, unchanged. - Strategy: `suite-adaptive-selection-v1-20260930`. - Capacity reservation: `suite-context-turns-v1-20260930`. Rollback while idle: revert the recovery deployment Git commit and reconcile the `hermes` Flux Kustomization. Do not edit live resources. Reverting reinstates the old 3600-second maximum and cost guard. Save any desired volatile results first.