atlas-iac/docs/hermes_suite_recovery_20260930.md

13 KiB

Suite recovery and adaptive analysis

Latest real-job diagnosis

Exact failed job: 90ab61bf285e4c89a41d893d4fee649d, located by descending persisted creation time, failed status and 363-case count. Its request selected 3600 seconds and an estimated-cost guard of 30. No real source was inspected or rerun. The record is not the earlier decision-audit capacity failure.

Measurement Retained evidence
Error incomplete_generation, failure_stage=cli_final_result
Failed pass proposal_b; one completed model pass
CLI exit / signal 1 / null; nonzero_exit
Terminal event result, subtype success, is_error=true
Provider / assistant stop stop_sequence / stop_sequence
API error indicator true; allowlisted category unknown; HTTP status unknown
Structured result Absent from terminal output, tool input and JSON text candidates
Turns / ceiling 1 / 6; no turn-limit or structured-retry-limit outcome
B input / output usage CLI reported 0 / 0; does not independently prove no provider work occurred
B observed model capacity Unknown; failed terminal event had no modelUsage limits
B admission Same 439270-byte input/system/schema size as A; conservative bound 447462; six full outputs reserved: 831462 under 1M
B configured output / reasoning 64000 per turn / xhigh; reasoning-token limit unknown
B output bytes Last heartbeat observed 54366; exact terminal byte count was not retained
B elapsed / timeout Last heartbeat 411.8 s; approximately 412.9 s including boundary overhead / 2791.629775 s watchdog
Provider timeout Unknown
Job elapsed / remaining 1221.233 s / 2378.8 s
Total input / output 319884 / 98177; output includes 82581 thinking tokens
Cost used / remaining old guard USD 3.243076 / 26.756924 estimated
Route Claude CLI, first-party OAuth; selected claude-opus-5-5[1m]
Runtime model evidence A confirmed firstParty canonical claude-opus-5-5; B initialization matched requested model, but no successful terminal canonical-model verification
CLI 2.1.285
Old configuration / prompt suite-v6-20260929 / implementation-proximity-multipass-v5-20260930
Old execution / policy suite-multipass-v10-20260930 / implementation-five-v1-20260929

Proposal A completed in 808.317 seconds, three CLI turns, exit zero and valid structured output. B produced an unsuccessful API-error event and exited nonzero. The underlying provider cause is unknown: raw CLI text was deliberately deleted. No retained evidence identifies rate limiting, context/output exhaustion, timeout, turn exhaustion or an adapter rejecting valid structured content. The old adapter collapsed this API error into incomplete_generation and the workflow had no retry. The new classifier treats an unclassified API error as eligible for a bounded retry, with provider_failure_transience=unconfirmed; that is not proof it was transient. New diagnostics retain exact output bytes and per-attempt elapsed time.

Time and cost

The maximum accepted execution.max_seconds is 7200; omission still uses 1800. All profile, grouping, review, retry and repair calls share one absolute deadline. Submit/status/result HTTP timeouts remain 45 seconds on the laptop.

The cost guard is removed, following the user's subsequent instruction. The CLI receives no --max-budget-usd. Legacy positive finite max_cost_usd values or null are accepted but never enforced. Capability/status metadata reports cost_guard_enforced=false and a null cost ceiling. Measured estimates remain informational; unknown usage/cost does not prevent further work. These estimates are not a statement of subscription billing charges.

7200 seconds is the recommended next limit. Five passes at A's 808.317 seconds would total 4041.585 seconds, leaving 3158.415 seconds for variability/recovery. Hierarchical execution has a different call mix, so this is a planning reference, not a guarantee or a measured hierarchical runtime. No fixed deadline guarantees success through a prolonged provider outage.

Later stages retain explicit time floors. Global stages reserve min(500, max(10, job_seconds / 8)) seconds per remaining global stage; pending profile batches reserve 30 seconds each, and source/review batches 60 seconds. These are admission/watchdog reservations, not guaranteed runtimes. Each attempt receives only time remaining after future reservations. Insufficient remaining reservation returns insufficient_remaining_pass_budget; the absolute deadline returns job_time_budget_exhausted. No per-pass allowance resets the job clock.

Recovery and checkpoints

At most three attempts per logical pass, each a fresh isolated CLI session with the same provider, model, complete pass source and objective. Retry backoff is 2 then 4 seconds, cancellable. CLI turns start at four and can increase to five then six only when complete-input context admission allows. No compaction, truncation, provider fallback or partial result is accepted.

Retriable outcomes include unsuccessful/missing terminal generation, CLI turn or structured-output retry exhaustion, rate limiting, temporary transport failures, nonzero transient exits, and distinct pass timeouts. One additional structured assignment/JSON repair attempt is allowed. Successful final JSON text envelopes are recognized after the normal completion, runtime-model and isolation guards; normal schema/membership validation still applies.

Authentication/authorization, model/isolation changes, cancellation, invalid requests, confirmed capacity violations and deterministic evidence/review failures are not blindly retried. Capacity or exhausted generation can select hierarchical analysis instead of repeating the same full-suite request. Final errors identify turn exhaustion, generation retries, output/context limits, provider failure, pass timeout or insufficient remaining budget. Status contains attempts, terminal metadata and retained checkpoint names, never generated content.

Every validated output is checkpointed in active-job memory. Identity binds the complete request hash, credential owner/scope, routing, provider, canonical model, configuration/prompt/policy/execution revisions, exact pass source, schema, objective and relevant prior-output hashes. Altering any input prevents reuse. Completed predecessors survive recovery of a later pass. A matching checkpoint restores a copy without another model call. Contents never enter SQLite, logs, telemetry or ordinary status; all checkpoints are removed when the job ends or its deadline expires (at most 7200 seconds).

Restart limitation: checkpoints are not durable. A worker restart marks an in-flight job interrupted_no_retry; no automatic provider resubmission occurs. A fresh client attempt after restart regenerates its passes. This deliberately uses the requested active-job recovery option instead of adding disk copies of source-derived outputs. Results retain the existing one-hour volatile retention.

Adaptive execution

Small requests prefer direct independent proposals, whole-suite reconciliation, large-family review and decision audit. Initial hierarchical selection occurs for 250+ cases, 256 KiB+ canonical source, or a p95 source field of 8192+ bytes. Direct capacity failure also selects hierarchical mode before inference. Retriable direct generation exhaustion or confirmed output/context limits can switch modes within the same job and remaining deadline; the transition is recorded.

Hierarchical mode:

  1. Extract profiles in deterministic batches of at most 24 cases or about 48 KiB. A single larger case remains intact and must pass capacity admission. Each profile preserves target, success-criteria work, setup, stimulus, observations, machinery, variations, uncertainty, exact alias and original source evidence.
  2. Run two independent proposals over every profile in the complete suite, with an alternative reproducible order for the second. Reconcile globally; profile-processing batches never determine family membership.
  3. Reopen all original generalized fields for every candidate family in bounded whole-family source reviews. Verify original evidence, split unsupported compatibility, and reject ownership drift. Only then accept natural families.
  4. Perform the existing oversized-family review, decision audit, exact coverage, unique short names and deterministic balanced five-case packaging.

Profiles have explicit size bounds; they are not arbitrary silent summarization. No profile can be missing, invented, or silently truncated. At the current 400-case service maximum, bounded profiles plus membership-only proposal context fit one global comparison request. A worst-case 400-alias ASCII regression admitted profile reconciliation at 618260 input/system/schema bytes, with five reserved CLI outputs. Every real expanded request is checked again. The bounded contract avoids needing overlapping-neighborhood reconciliation; no such fallback is claimed. A single full-source family that cannot fit intact still fails explicitly, as do excessive profile/source/review batches. Transport and case ceilings remain 1 MiB and 400, not unlimited input support.

Bounds: 32 profile batches, 16 source-review batches, eight batches per semantic review, 64 validated logical calls and 128 model invocations including retries. All share the same deadline. Optional metadata includes execution_mode, strategy_events, direct_pass_attempts, retry_count, checkpointed_passes, checkpoint_reused, profile/source batch counts, cross_batch_review_count, model_pass_count, natural_family_count, final_task_count and per-attempt measurements. Content-free source byte distributions now support better future fault fixtures. Engineering rationales remain only in authorized review_summary.

Verification

200 focused local tests passed. They cover isolated retries, turn adaptation, assignment repair, non-retriable errors, preserved predecessor checkpoints, checkpoint scope/expiry/restart semantics, global profile coverage, complete source reopening, no fallback, shared deadlines, reserved future time, and cost estimates exceeding the old guard without stopping work. Manifest render and client dry run passed; Flux diff was reviewed.

Hosted native Opus 5.5 comparison used the same synthetic 14 cases, interleaved between byte-array decoder tests and physical pulse/reset timing tests. Duplicate text retained distinct aliases. Hierarchical mode deliberately used four profile batches, then global comparisons, full-source review and both semantic reviews.

Hosted execution Seconds Calls Input / output tokens CLI estimate Natural / final groups
Direct 79.009 5 60130 / 9282 USD 0.426160 2 / 4
Hierarchical 148.752 10 111617 / 17770 USD 0.801868 2 / 4

Both had exact alias coverage, zero incorrect merge pairs, zero unnecessary semantic split pairs, no retries and valid balanced 4+3 packages per family. Descriptions kept decoder and hardware-capture objectives separate. The direct answer inferred that fault cases varied only in stimulus; the hierarchical answer more accurately retained uncertainty about unspecified fault patterns. This is a small quality check, not proof of real-suite quality or runtime.

The 363-case regression used 410322 source bytes, matching the failed record, with 45 interleaved machinery patterns and distant identical records. Actual historical per-field lengths were not retained, so that distribution cannot be matched or claimed equivalent. Both modes completed with mocked model answers from the independent fixture map, producing 45 natural families and 90 final tasks, maximum five members, balanced sizes and unique short names. Direct used five calls (175646 response bytes); hierarchical used 26 (224822 response bytes). Mocked execution was about 6.6 seconds per mode; these are local orchestration measurements, not hosted latency, cost or a model-quality result. No new hosted 363-case benchmark or real-suite rerun was performed.

Detailed synthetic evidence and the retained safe failure record are separate. No source fields or raw provider messages from the real suite are included in either artifact.

Client handoff and revisions

No new request fields. execution.strategy="whole_suite" still submits one complete suite and receives one combined result.groups result. Authentication, endpoint, credential scopes, pinned Opus 5.5, effort thresholds and routing remain.

AI_MAX_SECONDS = 7200
AI_JOB_REVISION = "suite-adaptive-v11-20260930-1"

Generate a new Idempotency-Key. AI_MAX_COST_USD may remain at its existing value if the client sends it; no client edit is needed to disable server cost limits. Do not reuse the failed job's key.

  • Configuration: suite-v6-20260929, HTTP compatibility unchanged.
  • Execution: suite-adaptive-v11-20260930.
  • Prompt: implementation-proximity-adaptive-v6-20260930.
  • Grouping policy: implementation-five-v1-20260929, unchanged.
  • Strategy: suite-adaptive-selection-v1-20260930.
  • Capacity reservation: suite-context-turns-v1-20260930.

Rollback while idle: revert the recovery deployment Git commit and reconcile the hermes Flux Kustomization. Do not edit live resources. Reverting reinstates the old 3600-second maximum and cost guard. Save any desired volatile results first.