atlas-iac/docs/hermes_suite_61_failure_20260929.md

3.7 KiB

Suite proposal failure diagnosis

The real job ca6290e71cb14bd1af64864f7c3ea256 failed in proposal_a after 30.264 seconds, with no completed model pass. Only stored operational metadata was inspected. No real case content was read or resubmitted.

Established failure

parse_claude() rejected the terminal CLI result in its cli_final_result branch because is_error was true. The subtype was success, which alone does not establish successful generation. The process exited with code 1. The final result event and initialization event were both present.

The CLI reported one of six turns, stop_sequence, zero input/output/cache tokens, zero API duration, and no runtime model-usage entry. There was no structured result, structured tool input, or JSON text candidate. The job did not reach schema extraction or assignment validation.

There is no evidence of a turn limit, structured-output retry limit, explicit context/output exhaustion, process signal, cancellation, or subprocess timeout. The subprocess allowance was 1799.965410697274 seconds. The provider timeout and reasoning-token limit remain unknown. The configured output cap was 64000 tokens; no observed output-limit entry was returned for this job. Zero reported usage does not prove that no content reached the provider.

The original CLI/provider cause is unknown. CLI output and temporary files were deleted by normal cleanup; the old diagnostics did not preserve the assistant error category. There is no complete answer to recover. The 61-case count is not an established cause.

Server-only diagnostic correction

Execution revision suite-multipass-v8-20260929 preserves the prompt revision implementation-proximity-multipass-v4-20260929, compatibility revision suite-v6-20260929, selected model, effort policy, authentication, routing, timeouts, six-turn limit, and strict output/coverage checks.

It adds assistant_error_code, assistant_api_error_seen, and cli_transport_error to safe CLI diagnostics. Categories use fixed allowlists; transport signatures use exact CLI-marked error messages, without retaining the message. Recognized upstream failures receive specific existing error codes. Unknown failures remain explicit errors. No retry or fallback is added.

The previous hardcoded reasoning_effort: medium diagnostic was incorrect. The real job selected high; v8 reports the actual invocation configuration.

Validation

  • 149 focused tests pass, including terminal error flags, category mapping, privacy canaries, process exits/signals/timeouts, turn limits, and coverage.
  • Kustomize render and client dry-run pass for services/hermes.
  • Isolated native Claude CLI 2.1.285 tests with a loopback mock returned subtype: success, is_error: true, exit 1, one turn, zero usage for HTTP 400/401/429/503 failures. A refused loopback connection produced the same pattern with no HTTP status and assistant category server_error.
  • These reproductions establish the adapter/envelope behavior, not the specific historical upstream failure. They used fake credentials and no hosted inference.

The separately recorded live synthetic check exercises one complete 61-case proposal, not the full multi-pass workflow. It must not be represented as a rerun or recovery of the failed real job.

Client handling and rollback

No client request/schema change is required. Reusing the old idempotency key returns the old failed job. A user-initiated retry requires a fresh attempt and idempotency key; the prepared source can be reused. Do not automatically resubmit the real suite.

Rollback by reverting the v8 diagnostic commit in Git and reconciling the hermes Flux Kustomization. This restores v7 diagnostics; it does not recover deleted output or change any completed/failed job record.