atlas-iac/docs/hermes_suite_61_failure_20260929.md
jenkins 1bc2f22bcc
Some checks failed
Tests / Declarative: Post Actions failed: 42, skipped: 89, passed: 4218
docs(hermes): record 61-case diagnostic acceptance
2026-09-29 16:48:41 -05:00

89 lines
4.7 KiB
Markdown

# Suite proposal failure diagnosis
The real job `ca6290e71cb14bd1af64864f7c3ea256` failed in `proposal_a`
after 30.264 seconds, with no completed model pass. Only stored operational
metadata was inspected. No real case content was read or resubmitted.
## Established failure
`parse_claude()` rejected the terminal CLI result in its `cli_final_result`
branch because `is_error` was true. The subtype was `success`, which alone
does not establish successful generation. The process exited with code 1.
The final result event and initialization event were both present.
The CLI reported one of six turns, `stop_sequence`, zero input/output/cache
tokens, zero API duration, and no runtime model-usage entry. There was no
structured result, structured tool input, or JSON text candidate. The job
did not reach schema extraction or assignment validation.
There is no evidence of a turn limit, structured-output retry limit, explicit
context/output exhaustion, process signal, cancellation, or subprocess timeout.
The subprocess allowance was 1799.965410697274 seconds. The provider timeout
and reasoning-token limit remain unknown. The configured output cap was
64000 tokens; no observed output-limit entry was returned for this job.
Zero reported usage does not prove that no content reached the provider.
The original CLI/provider cause is unknown. CLI output and temporary files
were deleted by normal cleanup; the old diagnostics did not preserve the
assistant error category. There is no complete answer to recover. The
61-case count is not an established cause.
## Server-only diagnostic correction
Execution revision `suite-multipass-v8-20260929` preserves the prompt revision
`implementation-proximity-multipass-v4-20260929`, compatibility revision
`suite-v6-20260929`, selected model, effort policy, authentication, routing,
timeouts, six-turn limit, and strict output/coverage checks.
It adds `assistant_error_code`, `assistant_api_error_seen`, and
`cli_transport_error` to safe CLI diagnostics. Categories use fixed allowlists;
transport signatures use exact CLI-marked error messages, without retaining
the message. Recognized upstream failures receive specific existing error
codes. Unknown failures remain explicit errors. No retry or fallback is added.
The previous hardcoded `reasoning_effort: medium` diagnostic was incorrect.
The real job selected `high`; v8 reports the actual invocation configuration.
## Validation
- 149 focused tests pass, including terminal error flags, category mapping,
privacy canaries, process exits/signals/timeouts, turn limits, and coverage.
- Kustomize render and client dry-run pass for `services/hermes`.
- Isolated native Claude CLI 2.1.285 tests with a loopback mock returned
`subtype: success`, `is_error: true`, exit 1, one turn, zero usage for HTTP
400/401/429/503 failures. A refused loopback connection produced the same
pattern with no HTTP status and assistant category `server_error`.
- These reproductions establish the adapter/envelope behavior, not the
specific historical upstream failure. They used fake credentials and no
hosted inference.
The deployed live synthetic check completed one 61-case `proposal_a` on
`claude-opus-5-5` at high effort in 28.806 seconds, using three of six turns.
The serialized request was 64502 bytes (64350 bytes of source content;
73903 bytes including pass instructions/schema). The CLI reported 52404 input
and 3627 output tokens across the invocation, including 541 thinking tokens,
and a $0.282156 cost estimate (not a subscription billing statement).
Strict validation accepted all 61 aliases exactly once. All nine natural
families matched the synthetic fixture, with zero false/missed merge pairs.
The terminal event had `is_error: false`, exit 0, and an object at
`result.structured_output`. No compaction event or truncation was detected.
There were no operator retries. This exercises one complete proposal, not the
full multi-pass workflow, and does not establish the original upstream cause.
Flux applied code commit `6178265f70c8a2abbef2b1c78bd591475eeb97a5` and the
planner rolled out successfully. Authenticated in-pod preflight reports v8;
this is not a new laptop/LAN connectivity check. Measured safe evidence is in
`docs/evidence/hermes_suite_61_diagnostic_20260929.json`.
## Client handling and rollback
No client request/schema change is required. Reusing the old idempotency key
returns the old failed job. A user-initiated retry requires a fresh attempt
and idempotency key; the prepared source can be reused. Do not automatically
resubmit the real suite.
Rollback by reverting the v8 diagnostic commit in Git and reconciling the
`hermes` Flux Kustomization. This restores v7 diagnostics; it does not recover
deleted output or change any completed/failed job record.