atlas-iac/docs/hermes_suite_failure_diagnosis_20260929.md

9.5 KiB

Suite planner failure diagnosis, 2026-09-29

Established facts and the historical limit of the evidence

The exact failed job is 40f473da1bd8495180a6cd7aff30c801. Its saved metadata reports 14 cases and 70.201 seconds, not the OCR estimate. It used configuration suite-v6-20260929, prompt implementation-proximity-v2-20260929, and the selected Claude route. The saved error is incomplete_generation with empty details.

The historical terminal condition is unknown. The old adapter discarded exit status and final-event diagnostics when raising that error. The per-job tmpfs directory was then deleted on all exit paths; /jobs contained zero remaining suite directories. Stderr was discarded during execution. The worker had not restarted. No valid result was stored for this failed job. Cleanup was intentional for data minimization; the diagnostic loss was an adapter defect. No request, source text, real generated descriptions, or raw provider error was read as part of this investigation.

Exact original code paths

In suite_backends.parse_claude, incomplete_generation could mean:

  1. No recognized initialization event or no terminal type=result event.
  2. A terminal result with is_error=true or subtype other than success (except recognized provider authentication/rate-limit errors).
  3. A successful-looking final event with stop_reason=max_tokens or model_context_window_exceeded.

In suite_backends.claude_generate, it also meant a nonzero subprocess exit after parsing an otherwise acceptable structured result. Thus a valid payload could have existed on that branch, but the historical artifact needed to prove or validate it no longer exists.

Missing/non-object result.structured_output and malformed stream JSON were invalid_json_result, not incomplete_generation. Normal schema and exactly-once membership checks run afterwards in Jobs.run, with distinct invalid_json_result, invalid_case_assignments, or duplicate_family_name errors. Those validation failures were not relabeled as incomplete generation. Subprocess deadline, cancellation, output-byte overflow, and worker restart also had distinct error codes. A CLI-internal execution error could still have entered path 2, and a signal/nonzero exit could enter paths 1 or 4.

Historical operational metadata

Item Failed job evidence
CLI exit code / terminating signal Unknown; discarded
Actual final event subtype / error indicator Unknown; deleted
Actual CLI turns / configured ceiling Unknown / 3
Structured-output retry limit and outcome Unknown
Provider stop reason Unknown
Actual input/output/reasoning token usage Unknown; saved usage is null
Structured output present / location Unknown; artifact deleted
Exact remaining subprocess deadline Unknown; normally request max_seconds minus routing time, at most 900 seconds
Provider-internal execution timeout Unknown
Compaction / truncation Unknown; both saved null
Configured output / reasoning 64,000 per-call output ceiling; medium effort; separate reasoning-token budget unknown
Preflight output_reservation_verified False; an estimate, not a runtime stop indication
Selected versus observed model Selected claude-opus-4-8; failed job's terminal runtime model metadata unavailable

The same installed Claude Code 2.1.226 binary's loopback request included max_tokens=64000 and x-stainless-timeout: 600. The latter is an observed SDK request header, not a measurement or guarantee of a provider-internal timeout, and it cannot establish the historical failure's cause. Successful jobs report context 1,000,000 and output ceiling 64,000; those numbers do not establish actual usage or the stop condition in the failed job.

Comparison with successful runs

The original 14/75/363 acceptance jobs each reported two CLI turns. The earlier successful real 14-case job 276720d47aa541038237e284e2058533 reported three turns, 70.2 seconds, and 16,947 input / 5,656 output tokens. The revised-prompt synthetic job 661aefd2d1f14579a264a34d64bf704b also reported three turns, 20.544 seconds, and 11,010 input / 1,833 output tokens. Only their operational metadata was read.

The prompt update expanded the implementation-effort instructions and added prompt provenance. It did not change the CLI invocation, three-turn ceiling, model, reasoning effort, output ceiling, parser, or result schema. Successful runs already reached the configured ceiling before this failure, including one before the new prompt. A three-turn assumption was therefore fragile; the new prompt is not an established unique cause.

Deterministic synthetic reproduction and change

scripts/ops/hermes_suite_cli_probe.py runs the installed native binary against a loopback mock using a fake OAuth value, with no hosted inference. The mock sends three invalid StructuredOutput tool arguments, then a valid complete synthetic result. The system prompt and complete input remain in every captured request.

Ceiling CLI exit Final subtype CLI-reported turns Actual mock generation calls Structured output
3 1 error_max_turns 4 3 Absent
6 0 success 5 4 Present; schema and exact membership pass

The CLI's turn counter can include the next/terminal loop iteration. It is recorded as reported, not interpreted as a count of paid provider calls. This reproduces a concrete failure class that the old adapter collapsed into the historical error label; it does not prove the historical job took that branch.

Execution revision claude-diagnostics-turns-v1-20260929:

  • Adds allowlisted failure-stage, exit/signal, final-event, turn/retry-limit, stop-reason, usage, timing, and structured-output-presence/location diagnostics.
  • Detects possible JSON in final text or StructuredOutput assistant tool arguments without accepting those candidates or retaining their contents.
  • Preserves safe usage/diagnostics on schema and membership failures.
  • Allows six CLI turns within the existing deadline and estimated-cost ceiling; preflight reserves six possible outputs against context. No automatic job retry.
  • Keeps input/output cleanup and content-free routine logs. No raw CLI errors, prompts, response descriptions, or credentials are added to persisted metadata.

The configuration revision, prompt revision/hash, model, medium effort, token ceiling, schema, authentication, provider restrictions, and coverage checks are unchanged. Historical jobs retain their old metadata. The new execution revision is additive metadata in capabilities, preflight, and new jobs.

65 focused tests pass, covering turn/structured-retry/budget errors, explicit provider stops, missing/malformed terminal events, candidate output locations, nonzero exits with valid output, actual subprocess signals/timeouts/cancellation, cleanup, privacy canaries, exact membership, and idempotency. Native loopback reproduction also passes. Kustomize rendering and client dry-run passed; Flux diff shows only the planner ConfigMap and planner restart annotation.

Deployed synthetic verification

Flux deployed code commit 2d762776c100c280d491cfbd11119fb57c627ab0; the new planner pod is Ready. LAN HTTPS testing ran from titan-jh at 192.168.22.8 to 192.168.22.50 with hostname TLS verification, no proxies, and no redirects. The user's laptop was not used. No active job was present before rollout.

Fresh synthetic job 6ad9067bfbc1441890a1fe49b8f3de70 completed:

  • 14 cases; 9,607 serialized request bytes.
  • 22.694 server wall seconds, 23.186 client wall seconds; no job retry or rate-limit event.
  • CLI exit 0, final result/success, is_error=false, three reported turns.
  • Stop reason tool_use; structured object at result.structured_output.
  • 11,073 input / 1,905 output tokens; 21,970 API milliseconds; USD 0.10299 CLI estimate.
  • Canonical model claude-opus-4-8, firstParty, reported 1,000,000 context / 64,000 output.
  • Six families, three singletons, exact alias coverage, pair precision/recall 1.0, zero incorrect or missed merge pairs. No observed cross-family objective attribution.
  • Same-key replay returned the same job; invalid credentials returned 401.
  • Two routine log records contained only approved operational metadata; zero remaining temporary suite directories after both live and loopback tests.

The native loopback three-versus-six-turn regression also passed on the deployed binary. These checks establish the change's behavior on synthetic material, not the deleted historical terminal condition or guaranteed success on future real jobs. See the synthetic verification evidence.

Recovery and the laptop's next attempt

The old answer cannot be recovered or validated after cleanup. No assignments will be fabricated. No real suite has been rerun. The laptop must intentionally create a new attempt with a new Idempotency-Key after reviewing this verification; reusing the old key correctly returns the same failed job. The input/request schema and Python parsing contract do not require a change for this fix.

Rollback only this diagnostic/turn-limit commit through Git and Flux. Keep the prior implementation-proximity prompt commit. Rollouts should occur with no active jobs; in-memory completed results expire on a worker restart.

Exact rollback for this change (from a clean checkout, after active jobs finish):

git revert 2d762776c100c280d491cfbd11119fb57c627ab0
git push origin HEAD:main
flux reconcile kustomization hermes --namespace flux-system --with-source