170 lines
9.5 KiB
Markdown
170 lines
9.5 KiB
Markdown
# Suite planner failure diagnosis, 2026-09-29
|
|
|
|
## Established facts and the historical limit of the evidence
|
|
|
|
The exact failed job is `40f473da1bd8495180a6cd7aff30c801`.
|
|
Its saved metadata reports 14 cases and **70.201 seconds**, not the OCR estimate.
|
|
It used configuration `suite-v6-20260929`, prompt
|
|
`implementation-proximity-v2-20260929`, and the selected Claude route.
|
|
The saved error is `incomplete_generation` with empty details.
|
|
|
|
The historical terminal condition is **unknown**. The old adapter discarded
|
|
exit status and final-event diagnostics when raising that error. The per-job
|
|
tmpfs directory was then deleted on all exit paths; `/jobs` contained zero
|
|
remaining suite directories. Stderr was discarded during execution. The worker
|
|
had not restarted. No valid result was stored for this failed job. Cleanup was
|
|
intentional for data minimization; the diagnostic loss was an adapter defect.
|
|
No request, source text, real generated descriptions, or raw provider error was
|
|
read as part of this investigation.
|
|
|
|
## Exact original code paths
|
|
|
|
In `suite_backends.parse_claude`, `incomplete_generation` could mean:
|
|
|
|
1. No recognized initialization event or no terminal `type=result` event.
|
|
2. A terminal result with `is_error=true` or `subtype` other than `success`
|
|
(except recognized provider authentication/rate-limit errors).
|
|
3. A successful-looking final event with `stop_reason=max_tokens` or
|
|
`model_context_window_exceeded`.
|
|
|
|
In `suite_backends.claude_generate`, it also meant a nonzero subprocess exit
|
|
**after** parsing an otherwise acceptable structured result. Thus a valid payload
|
|
could have existed on that branch, but the historical artifact needed to prove
|
|
or validate it no longer exists.
|
|
|
|
Missing/non-object `result.structured_output` and malformed stream JSON were
|
|
`invalid_json_result`, not `incomplete_generation`. Normal schema and exactly-once
|
|
membership checks run afterwards in `Jobs.run`, with distinct
|
|
`invalid_json_result`, `invalid_case_assignments`, or `duplicate_family_name`
|
|
errors. Those validation failures were not relabeled as incomplete generation.
|
|
Subprocess deadline, cancellation, output-byte overflow, and worker restart also
|
|
had distinct error codes. A CLI-internal execution error could still have entered
|
|
path 2, and a signal/nonzero exit could enter paths 1 or 4.
|
|
|
|
## Historical operational metadata
|
|
|
|
| Item | Failed job evidence |
|
|
| --- | --- |
|
|
| CLI exit code / terminating signal | Unknown; discarded |
|
|
| Actual final event subtype / error indicator | Unknown; deleted |
|
|
| Actual CLI turns / configured ceiling | Unknown / 3 |
|
|
| Structured-output retry limit and outcome | Unknown |
|
|
| Provider stop reason | Unknown |
|
|
| Actual input/output/reasoning token usage | Unknown; saved usage is null |
|
|
| Structured output present / location | Unknown; artifact deleted |
|
|
| Exact remaining subprocess deadline | Unknown; normally request max_seconds minus routing time, at most 900 seconds |
|
|
| Provider-internal execution timeout | Unknown |
|
|
| Compaction / truncation | Unknown; both saved null |
|
|
| Configured output / reasoning | 64,000 per-call output ceiling; medium effort; separate reasoning-token budget unknown |
|
|
| Preflight output_reservation_verified | False; an estimate, not a runtime stop indication |
|
|
| Selected versus observed model | Selected `claude-opus-4-8`; failed job's terminal runtime model metadata unavailable |
|
|
|
|
The same installed Claude Code 2.1.226 binary's loopback request included
|
|
`max_tokens=64000` and `x-stainless-timeout: 600`. The latter is an observed SDK
|
|
request header, not a measurement or guarantee of a provider-internal timeout,
|
|
and it cannot establish the historical failure's cause. Successful jobs report
|
|
context 1,000,000 and output ceiling 64,000; those numbers do not establish actual
|
|
usage or the stop condition in the failed job.
|
|
|
|
## Comparison with successful runs
|
|
|
|
The original 14/75/363 acceptance jobs each reported two CLI turns. The earlier
|
|
successful real 14-case job `276720d47aa541038237e284e2058533` reported three turns,
|
|
70.2 seconds, and 16,947 input / 5,656 output tokens. The revised-prompt synthetic
|
|
job `661aefd2d1f14579a264a34d64bf704b` also reported three turns, 20.544 seconds,
|
|
and 11,010 input / 1,833 output tokens. Only their operational metadata was read.
|
|
|
|
The prompt update expanded the implementation-effort instructions and added prompt
|
|
provenance. It did not change the CLI invocation, three-turn ceiling, model,
|
|
reasoning effort, output ceiling, parser, or result schema. Successful runs already
|
|
reached the configured ceiling before this failure, including one before the new
|
|
prompt. A three-turn assumption was therefore fragile; the new prompt is not an
|
|
established unique cause.
|
|
|
|
## Deterministic synthetic reproduction and change
|
|
|
|
`scripts/ops/hermes_suite_cli_probe.py` runs the installed native binary against a
|
|
loopback mock using a fake OAuth value, with no hosted inference. The mock sends
|
|
three invalid StructuredOutput tool arguments, then a valid complete synthetic
|
|
result. The system prompt and complete input remain in every captured request.
|
|
|
|
| Ceiling | CLI exit | Final subtype | CLI-reported turns | Actual mock generation calls | Structured output |
|
|
| ---: | ---: | --- | ---: | ---: | --- |
|
|
| 3 | 1 | error_max_turns | 4 | 3 | Absent |
|
|
| 6 | 0 | success | 5 | 4 | Present; schema and exact membership pass |
|
|
|
|
The CLI's turn counter can include the next/terminal loop iteration. It is
|
|
recorded as reported, not interpreted as a count of paid provider calls. This
|
|
reproduces a concrete failure class that the old adapter collapsed into the
|
|
historical error label; it does **not** prove the historical job took that branch.
|
|
|
|
Execution revision `claude-diagnostics-turns-v1-20260929`:
|
|
|
|
- Adds allowlisted failure-stage, exit/signal, final-event, turn/retry-limit,
|
|
stop-reason, usage, timing, and structured-output-presence/location diagnostics.
|
|
- Detects possible JSON in final text or StructuredOutput assistant tool arguments
|
|
without accepting those candidates or retaining their contents.
|
|
- Preserves safe usage/diagnostics on schema and membership failures.
|
|
- Allows six CLI turns within the existing deadline and estimated-cost ceiling;
|
|
preflight reserves six possible outputs against context. No automatic job retry.
|
|
- Keeps input/output cleanup and content-free routine logs. No raw CLI errors,
|
|
prompts, response descriptions, or credentials are added to persisted metadata.
|
|
|
|
The configuration revision, prompt revision/hash, model, medium effort, token
|
|
ceiling, schema, authentication, provider restrictions, and coverage checks are
|
|
unchanged. Historical jobs retain their old metadata. The new execution revision
|
|
is additive metadata in capabilities, preflight, and new jobs.
|
|
|
|
65 focused tests pass, covering turn/structured-retry/budget errors, explicit
|
|
provider stops, missing/malformed terminal events, candidate output locations,
|
|
nonzero exits with valid output, actual subprocess signals/timeouts/cancellation,
|
|
cleanup, privacy canaries, exact membership, and idempotency. Native loopback
|
|
reproduction also passes. Kustomize rendering and client dry-run passed; Flux diff
|
|
shows only the planner ConfigMap and planner restart annotation.
|
|
|
|
## Deployed synthetic verification
|
|
|
|
Flux deployed code commit `2d762776c100c280d491cfbd11119fb57c627ab0`; the new
|
|
planner pod is Ready. LAN HTTPS testing ran from `titan-jh` at 192.168.22.8 to
|
|
192.168.22.50 with hostname TLS verification, no proxies, and no redirects.
|
|
The user's laptop was not used. No active job was present before rollout.
|
|
|
|
Fresh synthetic job `6ad9067bfbc1441890a1fe49b8f3de70` completed:
|
|
|
|
- 14 cases; 9,607 serialized request bytes.
|
|
- 22.694 server wall seconds, 23.186 client wall seconds; no job retry or rate-limit event.
|
|
- CLI exit 0, final `result/success`, `is_error=false`, three reported turns.
|
|
- Stop reason `tool_use`; structured object at `result.structured_output`.
|
|
- 11,073 input / 1,905 output tokens; 21,970 API milliseconds; USD 0.10299 CLI estimate.
|
|
- Canonical model `claude-opus-4-8`, firstParty, reported 1,000,000 context / 64,000 output.
|
|
- Six families, three singletons, exact alias coverage, pair precision/recall 1.0,
|
|
zero incorrect or missed merge pairs. No observed cross-family objective attribution.
|
|
- Same-key replay returned the same job; invalid credentials returned 401.
|
|
- Two routine log records contained only approved operational metadata; zero remaining
|
|
temporary suite directories after both live and loopback tests.
|
|
|
|
The native loopback three-versus-six-turn regression also passed on the deployed
|
|
binary. These checks establish the change's behavior on synthetic material, not
|
|
the deleted historical terminal condition or guaranteed success on future real jobs.
|
|
See [the synthetic verification evidence](evidence/hermes_suite_diagnostics_20260929.json).
|
|
|
|
## Recovery and the laptop's next attempt
|
|
|
|
The old answer cannot be recovered or validated after cleanup. No assignments
|
|
will be fabricated. No real suite has been rerun. The laptop must intentionally
|
|
create a **new attempt with a new Idempotency-Key** after reviewing this verification;
|
|
reusing the old key correctly returns the same failed job. The input/request
|
|
schema and Python parsing contract do not require a change for this fix.
|
|
|
|
Rollback only this diagnostic/turn-limit commit through Git and Flux. Keep the
|
|
prior implementation-proximity prompt commit. Rollouts should occur with no active
|
|
jobs; in-memory completed results expire on a worker restart.
|
|
|
|
Exact rollback for this change (from a clean checkout, after active jobs finish):
|
|
|
|
```bash
|
|
git revert 2d762776c100c280d491cfbd11119fb57c627ab0
|
|
git push origin HEAD:main
|
|
flux reconcile kustomization hermes --namespace flux-system --with-source
|
|
```
|