222 lines
13 KiB
Markdown
222 lines
13 KiB
Markdown
# Suite recovery and adaptive analysis
|
|
|
|
## Latest real-job diagnosis
|
|
|
|
Exact failed job: `90ab61bf285e4c89a41d893d4fee649d`, located by descending
|
|
persisted creation time, failed status and 363-case count. Its request selected
|
|
3600 seconds and an estimated-cost guard of 30. No real source was inspected or
|
|
rerun. The record is not the earlier decision-audit capacity failure.
|
|
|
|
| Measurement | Retained evidence |
|
|
| --- | --- |
|
|
| Error | `incomplete_generation`, `failure_stage=cli_final_result` |
|
|
| Failed pass | `proposal_b`; one completed model pass |
|
|
| CLI exit / signal | 1 / null; `nonzero_exit` |
|
|
| Terminal event | `result`, subtype `success`, **is_error=true** |
|
|
| Provider / assistant stop | `stop_sequence` / `stop_sequence` |
|
|
| API error indicator | true; allowlisted category `unknown`; HTTP status unknown |
|
|
| Structured result | Absent from terminal output, tool input and JSON text candidates |
|
|
| Turns / ceiling | 1 / 6; no turn-limit or structured-retry-limit outcome |
|
|
| B input / output usage | CLI reported 0 / 0; does not independently prove no provider work occurred |
|
|
| B observed model capacity | Unknown; failed terminal event had no modelUsage limits |
|
|
| B admission | Same 439270-byte input/system/schema size as A; conservative bound 447462; six full outputs reserved: 831462 under 1M |
|
|
| B configured output / reasoning | 64000 per turn / xhigh; reasoning-token limit unknown |
|
|
| B output bytes | Last heartbeat observed 54366; exact terminal byte count was not retained |
|
|
| B elapsed / timeout | Last heartbeat 411.8 s; approximately 412.9 s including boundary overhead / 2791.629775 s watchdog |
|
|
| Provider timeout | Unknown |
|
|
| Job elapsed / remaining | 1221.233 s / 2378.8 s |
|
|
| Total input / output | 319884 / 98177; output includes 82581 thinking tokens |
|
|
| Cost used / remaining old guard | USD 3.243076 / 26.756924 estimated |
|
|
| Route | Claude CLI, first-party OAuth; selected `claude-opus-5-5[1m]` |
|
|
| Runtime model evidence | A confirmed firstParty canonical `claude-opus-5-5`; B initialization matched requested model, but no successful terminal canonical-model verification |
|
|
| CLI | 2.1.285 |
|
|
| Old configuration / prompt | `suite-v6-20260929` / `implementation-proximity-multipass-v5-20260930` |
|
|
| Old execution / policy | `suite-multipass-v10-20260930` / `implementation-five-v1-20260929` |
|
|
|
|
Proposal A completed in 808.317 seconds, three CLI turns, exit zero and valid
|
|
structured output. B produced an unsuccessful API-error event and exited nonzero.
|
|
The underlying provider cause is **unknown**: raw CLI text was deliberately deleted.
|
|
No retained evidence identifies rate limiting, context/output exhaustion, timeout,
|
|
turn exhaustion or an adapter rejecting valid structured content. The old adapter
|
|
collapsed this API error into `incomplete_generation` and the workflow had no retry.
|
|
The new classifier treats an unclassified API error as eligible for a bounded
|
|
retry, with `provider_failure_transience=unconfirmed`; that is not proof it was
|
|
transient. New diagnostics retain exact output bytes and per-attempt elapsed time.
|
|
|
|
## Time and cost
|
|
|
|
The maximum accepted `execution.max_seconds` is **7200**; omission still uses 1800.
|
|
All profile, grouping, review, retry and repair calls share one absolute deadline.
|
|
Submit/status/result HTTP timeouts remain 45 seconds on the laptop.
|
|
|
|
The cost guard is **removed**, following the user's subsequent instruction. The
|
|
CLI receives no `--max-budget-usd`. Legacy positive finite `max_cost_usd` values
|
|
or null are accepted but never enforced. Capability/status metadata reports
|
|
`cost_guard_enforced=false` and a null cost ceiling. Measured estimates remain
|
|
informational; unknown usage/cost does not prevent further work. These estimates
|
|
are not a statement of subscription billing charges.
|
|
|
|
7200 seconds is the recommended next limit. Five passes at A's 808.317 seconds
|
|
would total 4041.585 seconds, leaving 3158.415 seconds for variability/recovery.
|
|
Hierarchical execution has a different call mix, so this is a planning reference,
|
|
not a guarantee or a measured hierarchical runtime. No fixed deadline guarantees
|
|
success through a prolonged provider outage.
|
|
|
|
Later stages retain explicit time floors. Global stages reserve
|
|
`min(500, max(10, job_seconds / 8))` seconds per remaining global stage; pending
|
|
profile batches reserve 30 seconds each, and source/review batches 60 seconds.
|
|
These are admission/watchdog reservations, not guaranteed runtimes. Each attempt
|
|
receives only time remaining after future reservations. Insufficient remaining
|
|
reservation returns `insufficient_remaining_pass_budget`; the absolute deadline
|
|
returns `job_time_budget_exhausted`. No per-pass allowance resets the job clock.
|
|
|
|
## Recovery and checkpoints
|
|
|
|
At most three attempts per logical pass, each a fresh isolated CLI session with
|
|
the same provider, model, complete pass source and objective. Retry backoff is
|
|
2 then 4 seconds, cancellable. CLI turns start at four and can increase to five
|
|
then six only when complete-input context admission allows. No compaction,
|
|
truncation, provider fallback or partial result is accepted.
|
|
|
|
Retriable outcomes include unsuccessful/missing terminal generation, CLI turn or
|
|
structured-output retry exhaustion, rate limiting, temporary transport failures,
|
|
nonzero transient exits, and distinct pass timeouts. One additional structured
|
|
assignment/JSON repair attempt is allowed. Successful final JSON text envelopes
|
|
are recognized after the normal completion, runtime-model and isolation guards;
|
|
normal schema/membership validation still applies.
|
|
|
|
Authentication/authorization, model/isolation changes, cancellation, invalid
|
|
requests, confirmed capacity violations and deterministic evidence/review failures
|
|
are not blindly retried. Capacity or exhausted generation can select hierarchical
|
|
analysis instead of repeating the same full-suite request. Final errors identify
|
|
turn exhaustion, generation retries, output/context limits, provider failure,
|
|
pass timeout or insufficient remaining budget. Status contains attempts, terminal
|
|
metadata and retained checkpoint names, never generated content.
|
|
|
|
Every **validated** output is checkpointed in active-job memory. Identity binds
|
|
the complete request hash, credential owner/scope, routing, provider, canonical
|
|
model, configuration/prompt/policy/execution revisions, exact pass source, schema,
|
|
objective and relevant prior-output hashes. Altering any input prevents reuse.
|
|
Completed predecessors survive recovery of a later pass. A matching checkpoint
|
|
restores a copy without another model call. Contents never enter SQLite, logs,
|
|
telemetry or ordinary status; all checkpoints are removed when the job ends or
|
|
its deadline expires (at most 7200 seconds).
|
|
|
|
**Restart limitation:** checkpoints are not durable. A worker restart marks an
|
|
in-flight job `interrupted_no_retry`; no automatic provider resubmission occurs.
|
|
A fresh client attempt after restart regenerates its passes. This deliberately
|
|
uses the requested active-job recovery option instead of adding disk copies of
|
|
source-derived outputs. Results retain the existing one-hour volatile retention.
|
|
|
|
## Adaptive execution
|
|
|
|
Small requests prefer direct independent proposals, whole-suite reconciliation,
|
|
large-family review and decision audit. Initial hierarchical selection occurs for
|
|
250+ cases, 256 KiB+ canonical source, or a p95 source field of 8192+ bytes. Direct
|
|
capacity failure also selects hierarchical mode before inference. Retriable direct
|
|
generation exhaustion or confirmed output/context limits can switch modes within
|
|
the same job and remaining deadline; the transition is recorded.
|
|
|
|
Hierarchical mode:
|
|
|
|
1. Extract profiles in deterministic batches of at most 24 cases or about 48 KiB.
|
|
A single larger case remains intact and must pass capacity admission. Each
|
|
profile preserves target, success-criteria work, setup, stimulus, observations,
|
|
machinery, variations, uncertainty, exact alias and original source evidence.
|
|
2. Run two independent proposals over **every profile in the complete suite**,
|
|
with an alternative reproducible order for the second. Reconcile globally;
|
|
profile-processing batches never determine family membership.
|
|
3. Reopen all original generalized fields for every candidate family in bounded
|
|
whole-family source reviews. Verify original evidence, split unsupported
|
|
compatibility, and reject ownership drift. Only then accept natural families.
|
|
4. Perform the existing oversized-family review, decision audit, exact coverage,
|
|
unique short names and deterministic balanced five-case packaging.
|
|
|
|
Profiles have explicit size bounds; they are not arbitrary silent summarization.
|
|
No profile can be missing, invented, or silently truncated. At the current
|
|
400-case service maximum, bounded profiles plus membership-only proposal context
|
|
fit one global comparison request. A worst-case 400-alias ASCII regression admitted
|
|
profile reconciliation at 618260 input/system/schema bytes, with five reserved CLI
|
|
outputs. Every real expanded request is checked again. The bounded contract avoids
|
|
needing overlapping-neighborhood reconciliation; no such fallback is claimed.
|
|
A single full-source family that cannot fit intact still fails explicitly, as do
|
|
excessive profile/source/review batches. Transport and case ceilings remain 1 MiB
|
|
and 400, not unlimited input support.
|
|
|
|
Bounds: 32 profile batches, 16 source-review batches, eight batches per semantic
|
|
review, 64 validated logical calls and 128 model invocations including retries.
|
|
All share the same deadline. Optional metadata includes `execution_mode`,
|
|
`strategy_events`, `direct_pass_attempts`, `retry_count`, `checkpointed_passes`,
|
|
`checkpoint_reused`, profile/source batch counts, `cross_batch_review_count`,
|
|
`model_pass_count`, `natural_family_count`, `final_task_count` and per-attempt
|
|
measurements. Content-free source byte distributions now support better future
|
|
fault fixtures. Engineering rationales remain only in authorized `review_summary`.
|
|
|
|
## Verification
|
|
|
|
200 focused local tests passed. They cover isolated retries, turn adaptation,
|
|
assignment repair, non-retriable errors, preserved predecessor checkpoints,
|
|
checkpoint scope/expiry/restart semantics, global profile coverage, complete source
|
|
reopening, no fallback, shared deadlines, reserved future time, and cost estimates
|
|
exceeding the old guard without stopping work. Manifest render and client dry run
|
|
passed; Flux diff was reviewed.
|
|
|
|
Hosted native Opus 5.5 comparison used the same synthetic 14 cases, interleaved
|
|
between byte-array decoder tests and physical pulse/reset timing tests. Duplicate
|
|
text retained distinct aliases. Hierarchical mode deliberately used four profile
|
|
batches, then global comparisons, full-source review and both semantic reviews.
|
|
|
|
| Hosted execution | Seconds | Calls | Input / output tokens | CLI estimate | Natural / final groups |
|
|
| --- | ---: | ---: | --- | ---: | --- |
|
|
| Direct | 79.009 | 5 | 60130 / 9282 | USD 0.426160 | 2 / 4 |
|
|
| Hierarchical | 148.752 | 10 | 111617 / 17770 | USD 0.801868 | 2 / 4 |
|
|
|
|
Both had exact alias coverage, zero incorrect merge pairs, zero unnecessary
|
|
semantic split pairs, no retries and valid balanced 4+3 packages per family.
|
|
Descriptions kept decoder and hardware-capture objectives separate. The direct
|
|
answer inferred that fault cases varied only in stimulus; the hierarchical answer
|
|
more accurately retained uncertainty about unspecified fault patterns. This is a
|
|
small quality check, not proof of real-suite quality or runtime.
|
|
|
|
The 363-case regression used 410322 source bytes, matching the failed record, with
|
|
45 interleaved machinery patterns and distant identical records. Actual historical
|
|
per-field lengths were not retained, so that distribution cannot be matched or
|
|
claimed equivalent. Both modes completed with **mocked model answers from the
|
|
independent fixture map**, producing 45 natural families and 90 final tasks,
|
|
maximum five members, balanced sizes and unique short names. Direct used five
|
|
calls (175646 response bytes); hierarchical used 26 (224822 response bytes).
|
|
Mocked execution was about 6.6 seconds per mode; these are local orchestration
|
|
measurements, not hosted latency, cost or a model-quality result. No new hosted
|
|
363-case benchmark or real-suite rerun was performed.
|
|
|
|
Detailed [synthetic evidence](evidence/hermes_suite_recovery_20260930.json) and
|
|
[the retained safe failure record](evidence/hermes_suite_failure_90ab61bf_metadata.json)
|
|
are separate. No source fields or raw provider messages from the real suite are
|
|
included in either artifact.
|
|
|
|
## Client handoff and revisions
|
|
|
|
No new request fields. `execution.strategy="whole_suite"` still submits one
|
|
complete suite and receives one combined `result.groups` result. Authentication,
|
|
endpoint, credential scopes, pinned Opus 5.5, effort thresholds and routing remain.
|
|
|
|
```python
|
|
AI_MAX_SECONDS = 7200
|
|
AI_JOB_REVISION = "suite-adaptive-v11-20260930-1"
|
|
```
|
|
|
|
Generate a **new Idempotency-Key**. `AI_MAX_COST_USD` may remain at its existing
|
|
value if the client sends it; no client edit is needed to disable server cost
|
|
limits. Do not reuse the failed job's key.
|
|
|
|
- Configuration: `suite-v6-20260929`, HTTP compatibility unchanged.
|
|
- Execution: `suite-adaptive-v11-20260930`.
|
|
- Prompt: `implementation-proximity-adaptive-v6-20260930`.
|
|
- Grouping policy: `implementation-five-v1-20260929`, unchanged.
|
|
- Strategy: `suite-adaptive-selection-v1-20260930`.
|
|
- Capacity reservation: `suite-context-turns-v1-20260930`.
|
|
|
|
Rollback while idle: revert the recovery deployment Git commit and reconcile the
|
|
`hermes` Flux Kustomization. Do not edit live resources. Reverting reinstates the
|
|
old 3600-second maximum and cost guard. Save any desired volatile results first.
|