atlas-iac/docs/hermes_suite_failure_diagnosis_20260929.md

170 lines
9.5 KiB
Markdown

# Suite planner failure diagnosis, 2026-09-29
## Established facts and the historical limit of the evidence
The exact failed job is `40f473da1bd8495180a6cd7aff30c801`.
Its saved metadata reports 14 cases and **70.201 seconds**, not the OCR estimate.
It used configuration `suite-v6-20260929`, prompt
`implementation-proximity-v2-20260929`, and the selected Claude route.
The saved error is `incomplete_generation` with empty details.
The historical terminal condition is **unknown**. The old adapter discarded
exit status and final-event diagnostics when raising that error. The per-job
tmpfs directory was then deleted on all exit paths; `/jobs` contained zero
remaining suite directories. Stderr was discarded during execution. The worker
had not restarted. No valid result was stored for this failed job. Cleanup was
intentional for data minimization; the diagnostic loss was an adapter defect.
No request, source text, real generated descriptions, or raw provider error was
read as part of this investigation.
## Exact original code paths
In `suite_backends.parse_claude`, `incomplete_generation` could mean:
1. No recognized initialization event or no terminal `type=result` event.
2. A terminal result with `is_error=true` or `subtype` other than `success`
(except recognized provider authentication/rate-limit errors).
3. A successful-looking final event with `stop_reason=max_tokens` or
`model_context_window_exceeded`.
In `suite_backends.claude_generate`, it also meant a nonzero subprocess exit
**after** parsing an otherwise acceptable structured result. Thus a valid payload
could have existed on that branch, but the historical artifact needed to prove
or validate it no longer exists.
Missing/non-object `result.structured_output` and malformed stream JSON were
`invalid_json_result`, not `incomplete_generation`. Normal schema and exactly-once
membership checks run afterwards in `Jobs.run`, with distinct
`invalid_json_result`, `invalid_case_assignments`, or `duplicate_family_name`
errors. Those validation failures were not relabeled as incomplete generation.
Subprocess deadline, cancellation, output-byte overflow, and worker restart also
had distinct error codes. A CLI-internal execution error could still have entered
path 2, and a signal/nonzero exit could enter paths 1 or 4.
## Historical operational metadata
| Item | Failed job evidence |
| --- | --- |
| CLI exit code / terminating signal | Unknown; discarded |
| Actual final event subtype / error indicator | Unknown; deleted |
| Actual CLI turns / configured ceiling | Unknown / 3 |
| Structured-output retry limit and outcome | Unknown |
| Provider stop reason | Unknown |
| Actual input/output/reasoning token usage | Unknown; saved usage is null |
| Structured output present / location | Unknown; artifact deleted |
| Exact remaining subprocess deadline | Unknown; normally request max_seconds minus routing time, at most 900 seconds |
| Provider-internal execution timeout | Unknown |
| Compaction / truncation | Unknown; both saved null |
| Configured output / reasoning | 64,000 per-call output ceiling; medium effort; separate reasoning-token budget unknown |
| Preflight output_reservation_verified | False; an estimate, not a runtime stop indication |
| Selected versus observed model | Selected `claude-opus-4-8`; failed job's terminal runtime model metadata unavailable |
The same installed Claude Code 2.1.226 binary's loopback request included
`max_tokens=64000` and `x-stainless-timeout: 600`. The latter is an observed SDK
request header, not a measurement or guarantee of a provider-internal timeout,
and it cannot establish the historical failure's cause. Successful jobs report
context 1,000,000 and output ceiling 64,000; those numbers do not establish actual
usage or the stop condition in the failed job.
## Comparison with successful runs
The original 14/75/363 acceptance jobs each reported two CLI turns. The earlier
successful real 14-case job `276720d47aa541038237e284e2058533` reported three turns,
70.2 seconds, and 16,947 input / 5,656 output tokens. The revised-prompt synthetic
job `661aefd2d1f14579a264a34d64bf704b` also reported three turns, 20.544 seconds,
and 11,010 input / 1,833 output tokens. Only their operational metadata was read.
The prompt update expanded the implementation-effort instructions and added prompt
provenance. It did not change the CLI invocation, three-turn ceiling, model,
reasoning effort, output ceiling, parser, or result schema. Successful runs already
reached the configured ceiling before this failure, including one before the new
prompt. A three-turn assumption was therefore fragile; the new prompt is not an
established unique cause.
## Deterministic synthetic reproduction and change
`scripts/ops/hermes_suite_cli_probe.py` runs the installed native binary against a
loopback mock using a fake OAuth value, with no hosted inference. The mock sends
three invalid StructuredOutput tool arguments, then a valid complete synthetic
result. The system prompt and complete input remain in every captured request.
| Ceiling | CLI exit | Final subtype | CLI-reported turns | Actual mock generation calls | Structured output |
| ---: | ---: | --- | ---: | ---: | --- |
| 3 | 1 | error_max_turns | 4 | 3 | Absent |
| 6 | 0 | success | 5 | 4 | Present; schema and exact membership pass |
The CLI's turn counter can include the next/terminal loop iteration. It is
recorded as reported, not interpreted as a count of paid provider calls. This
reproduces a concrete failure class that the old adapter collapsed into the
historical error label; it does **not** prove the historical job took that branch.
Execution revision `claude-diagnostics-turns-v1-20260929`:
- Adds allowlisted failure-stage, exit/signal, final-event, turn/retry-limit,
stop-reason, usage, timing, and structured-output-presence/location diagnostics.
- Detects possible JSON in final text or StructuredOutput assistant tool arguments
without accepting those candidates or retaining their contents.
- Preserves safe usage/diagnostics on schema and membership failures.
- Allows six CLI turns within the existing deadline and estimated-cost ceiling;
preflight reserves six possible outputs against context. No automatic job retry.
- Keeps input/output cleanup and content-free routine logs. No raw CLI errors,
prompts, response descriptions, or credentials are added to persisted metadata.
The configuration revision, prompt revision/hash, model, medium effort, token
ceiling, schema, authentication, provider restrictions, and coverage checks are
unchanged. Historical jobs retain their old metadata. The new execution revision
is additive metadata in capabilities, preflight, and new jobs.
65 focused tests pass, covering turn/structured-retry/budget errors, explicit
provider stops, missing/malformed terminal events, candidate output locations,
nonzero exits with valid output, actual subprocess signals/timeouts/cancellation,
cleanup, privacy canaries, exact membership, and idempotency. Native loopback
reproduction also passes. Kustomize rendering and client dry-run passed; Flux diff
shows only the planner ConfigMap and planner restart annotation.
## Deployed synthetic verification
Flux deployed code commit `2d762776c100c280d491cfbd11119fb57c627ab0`; the new
planner pod is Ready. LAN HTTPS testing ran from `titan-jh` at 192.168.22.8 to
192.168.22.50 with hostname TLS verification, no proxies, and no redirects.
The user's laptop was not used. No active job was present before rollout.
Fresh synthetic job `6ad9067bfbc1441890a1fe49b8f3de70` completed:
- 14 cases; 9,607 serialized request bytes.
- 22.694 server wall seconds, 23.186 client wall seconds; no job retry or rate-limit event.
- CLI exit 0, final `result/success`, `is_error=false`, three reported turns.
- Stop reason `tool_use`; structured object at `result.structured_output`.
- 11,073 input / 1,905 output tokens; 21,970 API milliseconds; USD 0.10299 CLI estimate.
- Canonical model `claude-opus-4-8`, firstParty, reported 1,000,000 context / 64,000 output.
- Six families, three singletons, exact alias coverage, pair precision/recall 1.0,
zero incorrect or missed merge pairs. No observed cross-family objective attribution.
- Same-key replay returned the same job; invalid credentials returned 401.
- Two routine log records contained only approved operational metadata; zero remaining
temporary suite directories after both live and loopback tests.
The native loopback three-versus-six-turn regression also passed on the deployed
binary. These checks establish the change's behavior on synthetic material, not
the deleted historical terminal condition or guaranteed success on future real jobs.
See [the synthetic verification evidence](evidence/hermes_suite_diagnostics_20260929.json).
## Recovery and the laptop's next attempt
The old answer cannot be recovered or validated after cleanup. No assignments
will be fabricated. No real suite has been rerun. The laptop must intentionally
create a **new attempt with a new Idempotency-Key** after reviewing this verification;
reusing the old key correctly returns the same failed job. The input/request
schema and Python parsing contract do not require a change for this fix.
Rollback only this diagnostic/turn-limit commit through Git and Flux. Keep the
prior implementation-proximity prompt commit. Rollouts should occur with no active
jobs; in-memory completed results expire on a worker restart.
Exact rollback for this change (from a clean checkout, after active jobs finish):
```bash
git revert 2d762776c100c280d491cfbd11119fb57c627ab0
git push origin HEAD:main
flux reconcile kustomization hermes --namespace flux-system --with-source
```