atlas-iac/docs/hermes_suite_recovery_20260930.md

222 lines
13 KiB
Markdown

# Suite recovery and adaptive analysis
## Latest real-job diagnosis
Exact failed job: `90ab61bf285e4c89a41d893d4fee649d`, located by descending
persisted creation time, failed status and 363-case count. Its request selected
3600 seconds and an estimated-cost guard of 30. No real source was inspected or
rerun. The record is not the earlier decision-audit capacity failure.
| Measurement | Retained evidence |
| --- | --- |
| Error | `incomplete_generation`, `failure_stage=cli_final_result` |
| Failed pass | `proposal_b`; one completed model pass |
| CLI exit / signal | 1 / null; `nonzero_exit` |
| Terminal event | `result`, subtype `success`, **is_error=true** |
| Provider / assistant stop | `stop_sequence` / `stop_sequence` |
| API error indicator | true; allowlisted category `unknown`; HTTP status unknown |
| Structured result | Absent from terminal output, tool input and JSON text candidates |
| Turns / ceiling | 1 / 6; no turn-limit or structured-retry-limit outcome |
| B input / output usage | CLI reported 0 / 0; does not independently prove no provider work occurred |
| B observed model capacity | Unknown; failed terminal event had no modelUsage limits |
| B admission | Same 439270-byte input/system/schema size as A; conservative bound 447462; six full outputs reserved: 831462 under 1M |
| B configured output / reasoning | 64000 per turn / xhigh; reasoning-token limit unknown |
| B output bytes | Last heartbeat observed 54366; exact terminal byte count was not retained |
| B elapsed / timeout | Last heartbeat 411.8 s; approximately 412.9 s including boundary overhead / 2791.629775 s watchdog |
| Provider timeout | Unknown |
| Job elapsed / remaining | 1221.233 s / 2378.8 s |
| Total input / output | 319884 / 98177; output includes 82581 thinking tokens |
| Cost used / remaining old guard | USD 3.243076 / 26.756924 estimated |
| Route | Claude CLI, first-party OAuth; selected `claude-opus-5-5[1m]` |
| Runtime model evidence | A confirmed firstParty canonical `claude-opus-5-5`; B initialization matched requested model, but no successful terminal canonical-model verification |
| CLI | 2.1.285 |
| Old configuration / prompt | `suite-v6-20260929` / `implementation-proximity-multipass-v5-20260930` |
| Old execution / policy | `suite-multipass-v10-20260930` / `implementation-five-v1-20260929` |
Proposal A completed in 808.317 seconds, three CLI turns, exit zero and valid
structured output. B produced an unsuccessful API-error event and exited nonzero.
The underlying provider cause is **unknown**: raw CLI text was deliberately deleted.
No retained evidence identifies rate limiting, context/output exhaustion, timeout,
turn exhaustion or an adapter rejecting valid structured content. The old adapter
collapsed this API error into `incomplete_generation` and the workflow had no retry.
The new classifier treats an unclassified API error as eligible for a bounded
retry, with `provider_failure_transience=unconfirmed`; that is not proof it was
transient. New diagnostics retain exact output bytes and per-attempt elapsed time.
## Time and cost
The maximum accepted `execution.max_seconds` is **7200**; omission still uses 1800.
All profile, grouping, review, retry and repair calls share one absolute deadline.
Submit/status/result HTTP timeouts remain 45 seconds on the laptop.
The cost guard is **removed**, following the user's subsequent instruction. The
CLI receives no `--max-budget-usd`. Legacy positive finite `max_cost_usd` values
or null are accepted but never enforced. Capability/status metadata reports
`cost_guard_enforced=false` and a null cost ceiling. Measured estimates remain
informational; unknown usage/cost does not prevent further work. These estimates
are not a statement of subscription billing charges.
7200 seconds is the recommended next limit. Five passes at A's 808.317 seconds
would total 4041.585 seconds, leaving 3158.415 seconds for variability/recovery.
Hierarchical execution has a different call mix, so this is a planning reference,
not a guarantee or a measured hierarchical runtime. No fixed deadline guarantees
success through a prolonged provider outage.
Later stages retain explicit time floors. Global stages reserve
`min(500, max(10, job_seconds / 8))` seconds per remaining global stage; pending
profile batches reserve 30 seconds each, and source/review batches 60 seconds.
These are admission/watchdog reservations, not guaranteed runtimes. Each attempt
receives only time remaining after future reservations. Insufficient remaining
reservation returns `insufficient_remaining_pass_budget`; the absolute deadline
returns `job_time_budget_exhausted`. No per-pass allowance resets the job clock.
## Recovery and checkpoints
At most three attempts per logical pass, each a fresh isolated CLI session with
the same provider, model, complete pass source and objective. Retry backoff is
2 then 4 seconds, cancellable. CLI turns start at four and can increase to five
then six only when complete-input context admission allows. No compaction,
truncation, provider fallback or partial result is accepted.
Retriable outcomes include unsuccessful/missing terminal generation, CLI turn or
structured-output retry exhaustion, rate limiting, temporary transport failures,
nonzero transient exits, and distinct pass timeouts. One additional structured
assignment/JSON repair attempt is allowed. Successful final JSON text envelopes
are recognized after the normal completion, runtime-model and isolation guards;
normal schema/membership validation still applies.
Authentication/authorization, model/isolation changes, cancellation, invalid
requests, confirmed capacity violations and deterministic evidence/review failures
are not blindly retried. Capacity or exhausted generation can select hierarchical
analysis instead of repeating the same full-suite request. Final errors identify
turn exhaustion, generation retries, output/context limits, provider failure,
pass timeout or insufficient remaining budget. Status contains attempts, terminal
metadata and retained checkpoint names, never generated content.
Every **validated** output is checkpointed in active-job memory. Identity binds
the complete request hash, credential owner/scope, routing, provider, canonical
model, configuration/prompt/policy/execution revisions, exact pass source, schema,
objective and relevant prior-output hashes. Altering any input prevents reuse.
Completed predecessors survive recovery of a later pass. A matching checkpoint
restores a copy without another model call. Contents never enter SQLite, logs,
telemetry or ordinary status; all checkpoints are removed when the job ends or
its deadline expires (at most 7200 seconds).
**Restart limitation:** checkpoints are not durable. A worker restart marks an
in-flight job `interrupted_no_retry`; no automatic provider resubmission occurs.
A fresh client attempt after restart regenerates its passes. This deliberately
uses the requested active-job recovery option instead of adding disk copies of
source-derived outputs. Results retain the existing one-hour volatile retention.
## Adaptive execution
Small requests prefer direct independent proposals, whole-suite reconciliation,
large-family review and decision audit. Initial hierarchical selection occurs for
250+ cases, 256 KiB+ canonical source, or a p95 source field of 8192+ bytes. Direct
capacity failure also selects hierarchical mode before inference. Retriable direct
generation exhaustion or confirmed output/context limits can switch modes within
the same job and remaining deadline; the transition is recorded.
Hierarchical mode:
1. Extract profiles in deterministic batches of at most 24 cases or about 48 KiB.
A single larger case remains intact and must pass capacity admission. Each
profile preserves target, success-criteria work, setup, stimulus, observations,
machinery, variations, uncertainty, exact alias and original source evidence.
2. Run two independent proposals over **every profile in the complete suite**,
with an alternative reproducible order for the second. Reconcile globally;
profile-processing batches never determine family membership.
3. Reopen all original generalized fields for every candidate family in bounded
whole-family source reviews. Verify original evidence, split unsupported
compatibility, and reject ownership drift. Only then accept natural families.
4. Perform the existing oversized-family review, decision audit, exact coverage,
unique short names and deterministic balanced five-case packaging.
Profiles have explicit size bounds; they are not arbitrary silent summarization.
No profile can be missing, invented, or silently truncated. At the current
400-case service maximum, bounded profiles plus membership-only proposal context
fit one global comparison request. A worst-case 400-alias ASCII regression admitted
profile reconciliation at 618260 input/system/schema bytes, with five reserved CLI
outputs. Every real expanded request is checked again. The bounded contract avoids
needing overlapping-neighborhood reconciliation; no such fallback is claimed.
A single full-source family that cannot fit intact still fails explicitly, as do
excessive profile/source/review batches. Transport and case ceilings remain 1 MiB
and 400, not unlimited input support.
Bounds: 32 profile batches, 16 source-review batches, eight batches per semantic
review, 64 validated logical calls and 128 model invocations including retries.
All share the same deadline. Optional metadata includes `execution_mode`,
`strategy_events`, `direct_pass_attempts`, `retry_count`, `checkpointed_passes`,
`checkpoint_reused`, profile/source batch counts, `cross_batch_review_count`,
`model_pass_count`, `natural_family_count`, `final_task_count` and per-attempt
measurements. Content-free source byte distributions now support better future
fault fixtures. Engineering rationales remain only in authorized `review_summary`.
## Verification
200 focused local tests passed. They cover isolated retries, turn adaptation,
assignment repair, non-retriable errors, preserved predecessor checkpoints,
checkpoint scope/expiry/restart semantics, global profile coverage, complete source
reopening, no fallback, shared deadlines, reserved future time, and cost estimates
exceeding the old guard without stopping work. Manifest render and client dry run
passed; Flux diff was reviewed.
Hosted native Opus 5.5 comparison used the same synthetic 14 cases, interleaved
between byte-array decoder tests and physical pulse/reset timing tests. Duplicate
text retained distinct aliases. Hierarchical mode deliberately used four profile
batches, then global comparisons, full-source review and both semantic reviews.
| Hosted execution | Seconds | Calls | Input / output tokens | CLI estimate | Natural / final groups |
| --- | ---: | ---: | --- | ---: | --- |
| Direct | 79.009 | 5 | 60130 / 9282 | USD 0.426160 | 2 / 4 |
| Hierarchical | 148.752 | 10 | 111617 / 17770 | USD 0.801868 | 2 / 4 |
Both had exact alias coverage, zero incorrect merge pairs, zero unnecessary
semantic split pairs, no retries and valid balanced 4+3 packages per family.
Descriptions kept decoder and hardware-capture objectives separate. The direct
answer inferred that fault cases varied only in stimulus; the hierarchical answer
more accurately retained uncertainty about unspecified fault patterns. This is a
small quality check, not proof of real-suite quality or runtime.
The 363-case regression used 410322 source bytes, matching the failed record, with
45 interleaved machinery patterns and distant identical records. Actual historical
per-field lengths were not retained, so that distribution cannot be matched or
claimed equivalent. Both modes completed with **mocked model answers from the
independent fixture map**, producing 45 natural families and 90 final tasks,
maximum five members, balanced sizes and unique short names. Direct used five
calls (175646 response bytes); hierarchical used 26 (224822 response bytes).
Mocked execution was about 6.6 seconds per mode; these are local orchestration
measurements, not hosted latency, cost or a model-quality result. No new hosted
363-case benchmark or real-suite rerun was performed.
Detailed [synthetic evidence](evidence/hermes_suite_recovery_20260930.json) and
[the retained safe failure record](evidence/hermes_suite_failure_90ab61bf_metadata.json)
are separate. No source fields or raw provider messages from the real suite are
included in either artifact.
## Client handoff and revisions
No new request fields. `execution.strategy="whole_suite"` still submits one
complete suite and receives one combined `result.groups` result. Authentication,
endpoint, credential scopes, pinned Opus 5.5, effort thresholds and routing remain.
```python
AI_MAX_SECONDS = 7200
AI_JOB_REVISION = "suite-adaptive-v11-20260930-1"
```
Generate a **new Idempotency-Key**. `AI_MAX_COST_USD` may remain at its existing
value if the client sends it; no client edit is needed to disable server cost
limits. Do not reuse the failed job's key.
- Configuration: `suite-v6-20260929`, HTTP compatibility unchanged.
- Execution: `suite-adaptive-v11-20260930`.
- Prompt: `implementation-proximity-adaptive-v6-20260930`.
- Grouping policy: `implementation-five-v1-20260929`, unchanged.
- Strategy: `suite-adaptive-selection-v1-20260930`.
- Capacity reservation: `suite-context-turns-v1-20260930`.
Rollback while idle: revert the recovery deployment Git commit and reconcile the
`hermes` Flux Kustomization. Do not edit live resources. Reverting reinstates the
old 3600-second maximum and cost guard. Save any desired volatile results first.