Some checks failed
Tests / Declarative: Post Actions failed: 42, skipped: 89, passed: 4218
89 lines
4.7 KiB
Markdown
89 lines
4.7 KiB
Markdown
# Suite proposal failure diagnosis
|
|
|
|
The real job `ca6290e71cb14bd1af64864f7c3ea256` failed in `proposal_a`
|
|
after 30.264 seconds, with no completed model pass. Only stored operational
|
|
metadata was inspected. No real case content was read or resubmitted.
|
|
|
|
## Established failure
|
|
|
|
`parse_claude()` rejected the terminal CLI result in its `cli_final_result`
|
|
branch because `is_error` was true. The subtype was `success`, which alone
|
|
does not establish successful generation. The process exited with code 1.
|
|
The final result event and initialization event were both present.
|
|
|
|
The CLI reported one of six turns, `stop_sequence`, zero input/output/cache
|
|
tokens, zero API duration, and no runtime model-usage entry. There was no
|
|
structured result, structured tool input, or JSON text candidate. The job
|
|
did not reach schema extraction or assignment validation.
|
|
|
|
There is no evidence of a turn limit, structured-output retry limit, explicit
|
|
context/output exhaustion, process signal, cancellation, or subprocess timeout.
|
|
The subprocess allowance was 1799.965410697274 seconds. The provider timeout
|
|
and reasoning-token limit remain unknown. The configured output cap was
|
|
64000 tokens; no observed output-limit entry was returned for this job.
|
|
Zero reported usage does not prove that no content reached the provider.
|
|
|
|
The original CLI/provider cause is unknown. CLI output and temporary files
|
|
were deleted by normal cleanup; the old diagnostics did not preserve the
|
|
assistant error category. There is no complete answer to recover. The
|
|
61-case count is not an established cause.
|
|
|
|
## Server-only diagnostic correction
|
|
|
|
Execution revision `suite-multipass-v8-20260929` preserves the prompt revision
|
|
`implementation-proximity-multipass-v4-20260929`, compatibility revision
|
|
`suite-v6-20260929`, selected model, effort policy, authentication, routing,
|
|
timeouts, six-turn limit, and strict output/coverage checks.
|
|
|
|
It adds `assistant_error_code`, `assistant_api_error_seen`, and
|
|
`cli_transport_error` to safe CLI diagnostics. Categories use fixed allowlists;
|
|
transport signatures use exact CLI-marked error messages, without retaining
|
|
the message. Recognized upstream failures receive specific existing error
|
|
codes. Unknown failures remain explicit errors. No retry or fallback is added.
|
|
|
|
The previous hardcoded `reasoning_effort: medium` diagnostic was incorrect.
|
|
The real job selected `high`; v8 reports the actual invocation configuration.
|
|
|
|
## Validation
|
|
|
|
- 149 focused tests pass, including terminal error flags, category mapping,
|
|
privacy canaries, process exits/signals/timeouts, turn limits, and coverage.
|
|
- Kustomize render and client dry-run pass for `services/hermes`.
|
|
- Isolated native Claude CLI 2.1.285 tests with a loopback mock returned
|
|
`subtype: success`, `is_error: true`, exit 1, one turn, zero usage for HTTP
|
|
400/401/429/503 failures. A refused loopback connection produced the same
|
|
pattern with no HTTP status and assistant category `server_error`.
|
|
- These reproductions establish the adapter/envelope behavior, not the
|
|
specific historical upstream failure. They used fake credentials and no
|
|
hosted inference.
|
|
|
|
The deployed live synthetic check completed one 61-case `proposal_a` on
|
|
`claude-opus-5-5` at high effort in 28.806 seconds, using three of six turns.
|
|
The serialized request was 64502 bytes (64350 bytes of source content;
|
|
73903 bytes including pass instructions/schema). The CLI reported 52404 input
|
|
and 3627 output tokens across the invocation, including 541 thinking tokens,
|
|
and a $0.282156 cost estimate (not a subscription billing statement).
|
|
|
|
Strict validation accepted all 61 aliases exactly once. All nine natural
|
|
families matched the synthetic fixture, with zero false/missed merge pairs.
|
|
The terminal event had `is_error: false`, exit 0, and an object at
|
|
`result.structured_output`. No compaction event or truncation was detected.
|
|
There were no operator retries. This exercises one complete proposal, not the
|
|
full multi-pass workflow, and does not establish the original upstream cause.
|
|
|
|
Flux applied code commit `6178265f70c8a2abbef2b1c78bd591475eeb97a5` and the
|
|
planner rolled out successfully. Authenticated in-pod preflight reports v8;
|
|
this is not a new laptop/LAN connectivity check. Measured safe evidence is in
|
|
`docs/evidence/hermes_suite_61_diagnostic_20260929.json`.
|
|
|
|
## Client handling and rollback
|
|
|
|
No client request/schema change is required. Reusing the old idempotency key
|
|
returns the old failed job. A user-initiated retry requires a fresh attempt
|
|
and idempotency key; the prepared source can be reused. Do not automatically
|
|
resubmit the real suite.
|
|
|
|
Rollback by reverting the v8 diagnostic commit in Git and reconciling the
|
|
`hermes` Flux Kustomization. This restores v7 diagnostics; it does not recover
|
|
deleted output or change any completed/failed job record.
|