9.5 KiB
Suite planner failure diagnosis, 2026-09-29
Established facts and the historical limit of the evidence
The exact failed job is 40f473da1bd8495180a6cd7aff30c801.
Its saved metadata reports 14 cases and 70.201 seconds, not the OCR estimate.
It used configuration suite-v6-20260929, prompt
implementation-proximity-v2-20260929, and the selected Claude route.
The saved error is incomplete_generation with empty details.
The historical terminal condition is unknown. The old adapter discarded
exit status and final-event diagnostics when raising that error. The per-job
tmpfs directory was then deleted on all exit paths; /jobs contained zero
remaining suite directories. Stderr was discarded during execution. The worker
had not restarted. No valid result was stored for this failed job. Cleanup was
intentional for data minimization; the diagnostic loss was an adapter defect.
No request, source text, real generated descriptions, or raw provider error was
read as part of this investigation.
Exact original code paths
In suite_backends.parse_claude, incomplete_generation could mean:
- No recognized initialization event or no terminal
type=resultevent. - A terminal result with
is_error=trueorsubtypeother thansuccess(except recognized provider authentication/rate-limit errors). - A successful-looking final event with
stop_reason=max_tokensormodel_context_window_exceeded.
In suite_backends.claude_generate, it also meant a nonzero subprocess exit
after parsing an otherwise acceptable structured result. Thus a valid payload
could have existed on that branch, but the historical artifact needed to prove
or validate it no longer exists.
Missing/non-object result.structured_output and malformed stream JSON were
invalid_json_result, not incomplete_generation. Normal schema and exactly-once
membership checks run afterwards in Jobs.run, with distinct
invalid_json_result, invalid_case_assignments, or duplicate_family_name
errors. Those validation failures were not relabeled as incomplete generation.
Subprocess deadline, cancellation, output-byte overflow, and worker restart also
had distinct error codes. A CLI-internal execution error could still have entered
path 2, and a signal/nonzero exit could enter paths 1 or 4.
Historical operational metadata
| Item | Failed job evidence |
|---|---|
| CLI exit code / terminating signal | Unknown; discarded |
| Actual final event subtype / error indicator | Unknown; deleted |
| Actual CLI turns / configured ceiling | Unknown / 3 |
| Structured-output retry limit and outcome | Unknown |
| Provider stop reason | Unknown |
| Actual input/output/reasoning token usage | Unknown; saved usage is null |
| Structured output present / location | Unknown; artifact deleted |
| Exact remaining subprocess deadline | Unknown; normally request max_seconds minus routing time, at most 900 seconds |
| Provider-internal execution timeout | Unknown |
| Compaction / truncation | Unknown; both saved null |
| Configured output / reasoning | 64,000 per-call output ceiling; medium effort; separate reasoning-token budget unknown |
| Preflight output_reservation_verified | False; an estimate, not a runtime stop indication |
| Selected versus observed model | Selected claude-opus-4-8; failed job's terminal runtime model metadata unavailable |
The same installed Claude Code 2.1.226 binary's loopback request included
max_tokens=64000 and x-stainless-timeout: 600. The latter is an observed SDK
request header, not a measurement or guarantee of a provider-internal timeout,
and it cannot establish the historical failure's cause. Successful jobs report
context 1,000,000 and output ceiling 64,000; those numbers do not establish actual
usage or the stop condition in the failed job.
Comparison with successful runs
The original 14/75/363 acceptance jobs each reported two CLI turns. The earlier
successful real 14-case job 276720d47aa541038237e284e2058533 reported three turns,
70.2 seconds, and 16,947 input / 5,656 output tokens. The revised-prompt synthetic
job 661aefd2d1f14579a264a34d64bf704b also reported three turns, 20.544 seconds,
and 11,010 input / 1,833 output tokens. Only their operational metadata was read.
The prompt update expanded the implementation-effort instructions and added prompt provenance. It did not change the CLI invocation, three-turn ceiling, model, reasoning effort, output ceiling, parser, or result schema. Successful runs already reached the configured ceiling before this failure, including one before the new prompt. A three-turn assumption was therefore fragile; the new prompt is not an established unique cause.
Deterministic synthetic reproduction and change
scripts/ops/hermes_suite_cli_probe.py runs the installed native binary against a
loopback mock using a fake OAuth value, with no hosted inference. The mock sends
three invalid StructuredOutput tool arguments, then a valid complete synthetic
result. The system prompt and complete input remain in every captured request.
| Ceiling | CLI exit | Final subtype | CLI-reported turns | Actual mock generation calls | Structured output |
|---|---|---|---|---|---|
| 3 | 1 | error_max_turns | 4 | 3 | Absent |
| 6 | 0 | success | 5 | 4 | Present; schema and exact membership pass |
The CLI's turn counter can include the next/terminal loop iteration. It is recorded as reported, not interpreted as a count of paid provider calls. This reproduces a concrete failure class that the old adapter collapsed into the historical error label; it does not prove the historical job took that branch.
Execution revision claude-diagnostics-turns-v1-20260929:
- Adds allowlisted failure-stage, exit/signal, final-event, turn/retry-limit, stop-reason, usage, timing, and structured-output-presence/location diagnostics.
- Detects possible JSON in final text or StructuredOutput assistant tool arguments without accepting those candidates or retaining their contents.
- Preserves safe usage/diagnostics on schema and membership failures.
- Allows six CLI turns within the existing deadline and estimated-cost ceiling; preflight reserves six possible outputs against context. No automatic job retry.
- Keeps input/output cleanup and content-free routine logs. No raw CLI errors, prompts, response descriptions, or credentials are added to persisted metadata.
The configuration revision, prompt revision/hash, model, medium effort, token ceiling, schema, authentication, provider restrictions, and coverage checks are unchanged. Historical jobs retain their old metadata. The new execution revision is additive metadata in capabilities, preflight, and new jobs.
65 focused tests pass, covering turn/structured-retry/budget errors, explicit provider stops, missing/malformed terminal events, candidate output locations, nonzero exits with valid output, actual subprocess signals/timeouts/cancellation, cleanup, privacy canaries, exact membership, and idempotency. Native loopback reproduction also passes. Kustomize rendering and client dry-run passed; Flux diff shows only the planner ConfigMap and planner restart annotation.
Deployed synthetic verification
Flux deployed code commit 2d762776c100c280d491cfbd11119fb57c627ab0; the new
planner pod is Ready. LAN HTTPS testing ran from titan-jh at 192.168.22.8 to
192.168.22.50 with hostname TLS verification, no proxies, and no redirects.
The user's laptop was not used. No active job was present before rollout.
Fresh synthetic job 6ad9067bfbc1441890a1fe49b8f3de70 completed:
- 14 cases; 9,607 serialized request bytes.
- 22.694 server wall seconds, 23.186 client wall seconds; no job retry or rate-limit event.
- CLI exit 0, final
result/success,is_error=false, three reported turns. - Stop reason
tool_use; structured object atresult.structured_output. - 11,073 input / 1,905 output tokens; 21,970 API milliseconds; USD 0.10299 CLI estimate.
- Canonical model
claude-opus-4-8, firstParty, reported 1,000,000 context / 64,000 output. - Six families, three singletons, exact alias coverage, pair precision/recall 1.0, zero incorrect or missed merge pairs. No observed cross-family objective attribution.
- Same-key replay returned the same job; invalid credentials returned 401.
- Two routine log records contained only approved operational metadata; zero remaining temporary suite directories after both live and loopback tests.
The native loopback three-versus-six-turn regression also passed on the deployed binary. These checks establish the change's behavior on synthetic material, not the deleted historical terminal condition or guaranteed success on future real jobs. See the synthetic verification evidence.
Recovery and the laptop's next attempt
The old answer cannot be recovered or validated after cleanup. No assignments will be fabricated. No real suite has been rerun. The laptop must intentionally create a new attempt with a new Idempotency-Key after reviewing this verification; reusing the old key correctly returns the same failed job. The input/request schema and Python parsing contract do not require a change for this fix.
Rollback only this diagnostic/turn-limit commit through Git and Flux. Keep the prior implementation-proximity prompt commit. Rollouts should occur with no active jobs; in-memory completed results expire on a worker restart.
Exact rollback for this change (from a clean checkout, after active jobs finish):
git revert 2d762776c100c280d491cfbd11119fb57c627ab0
git push origin HEAD:main
flux reconcile kustomization hermes --namespace flux-system --with-source