atlas-iac/docs/hermes_suite_claude55.md

8.0 KiB

Claude 5.5 suite comparison

The suite planner now pins native Claude Code 2.1.285. Opus 5.5 and Sonnet 5.5 both returned their exact canonical model identities on the existing first-party OAuth account. The default deployment selects claude-opus-5-5; Sonnet is a server configuration option, not an automatic fallback. The HTTPS contract, credential scopes, implementation objective, five-case task cap, 30-minute deadline, and USD 30 CLI estimated-cost guard are unchanged.

Execution revision suite-multipass-v7-20260929 uses high by default and xhigh for at least 100 cases OR at least 131,072 bytes (128 KiB) of complete suite content. Byte size is the canonical UTF-8 JSON of campaign, suite, and all cases, without transport whitespace, routing/execution fields, prompt text, schemas, or generated proposals. The selected effort is fixed for every pass in that job, including structured-output repair; later review prompts do not change it. There is no classifier or auxiliary inference call to choose effort.

Capabilities expose the thresholds in models.claude.reasoning_policy. Preflight and job metadata expose selection.reasoning plus selection.reasoning_selection with the policy revision, measured source bytes, case count, and triggered thresholds. Per-pass metadata and completed results also contain reasoning. The policy revision is suite-size-effort-v1-20260929. The client does not need a new request field. Explicit smaller job budgets remain effective; the service fails instead of reducing effort or accepting partial work.

The comparison below used medium; its runtime and cost measurements do not measure high or xhigh effort. Both Opus 5.5 and Sonnet 5.5 support low, medium, high, xhigh, and max. High is the next step above medium; xhigh and max spend more reasoning tokens and may take longer, without guaranteeing better grouping. No hosted inference or real roster rerun is needed to apply this setting. The HTTP compatibility and prompt revisions remain unchanged.

Matched 14-case results

Each run used the same 9,851-byte normalized synthetic request, complete source fields, schema, and three-pass workflow (two independent proposals and one whole-suite reconciliation). Each completed in six CLI turns without a retry. No family exceeded five, so these hosted tests did not need the two extra oversized-family review passes.

Measurement Opus 5.5 Sonnet 5.5
Wall seconds 39.256 24.799
Aggregate input tokens 21,936 21,438
Aggregate output tokens 4,243 3,814
CLI estimated USD 0.172604 0.081016
Natural families / final tasks 9 / 9 9 / 9
Singleton families 4 4
Exact alias coverage Yes Yes
Incorrect merge / missed merge pairs 0 / 0 0 / 0

Both produced identical memberships matching the fixture's independent mechanism labels, including distant related cases, identical text under different aliases, and watchdog terminology shared by different test machinery. Review found no description importing another family's objectives. Opus explicitly retained unknown procedure details for the thermal, acoustic, and build singleton fixtures. Sonnet's descriptions were more concise and omitted those uncertainty notes.

Membership quality tied on this small fixture. Sonnet used about 53% less CLI estimated cost and finished about 37% sooner. This is one observation per model, not evidence that their performance or quality is equivalent on a difficult 363-case suite. The amounts are CLI accounting estimates, not subscription charges or proof of a dollar-denominated account bill.

Runtime limits and isolation

Both CLI result envelopes report a 1,000,000-token context and 128,000-token model output ceiling. The service deliberately continues to request at most 64,000 output tokens per invocation. selection.output is that service limit; selection.reported_output and model_usage[*].maxOutputTokens describe the reported model ceiling. The native CLI loopback check captured max_tokens=64000 and complete source/system text on all four requests, including one deliberately invalid structured assignment followed by repair.

The newer CLI loads a built-in instruction plugin even in safe mode unless it is explicitly disabled. The service now disables it, all hooks, bundled skills, provider connectors, plugin synchronization, built-in subagents, and git instructions. Each CLI invocation allowlists only its pinned model and turns off automatic model switching when a request is flagged. Runtime identity, empty plugins/MCP lists, and tool isolation are still checked before accepting output. No provider or model fallback was added. The first tiny Opus readiness probe failed with zero usage and insufficient retained diagnostics; a fresh probe succeeded. Neither measured suite job needed a retry.

The artifact is downloaded from the official release host at startup with a pinned version, 240,327,864-byte size, and SHA-256:

33dad1ec615a2e08cc78b494f05c110e49916de2c79d78ec8799ebf46b233d29

This removes the planner's dependency on the shared agent tools volume. Other Hermes CLIs and the isolated local-model endpoint are unchanged. A cold start now requires the official release download host to be reachable; failed download or checksum verification leaves the worker unavailable instead of using another CLI.

Configuration and rollback

The HTTP compatibility revision remains suite-v6-20260929, the prompt remains implementation-proximity-multipass-v4-20260929, and the execution revision is suite-multipass-v7-20260929. Authenticated capabilities additionally expose claude_model_options and model_selection: server_configuration. No new client request field is accepted or required.

To select Sonnet, change the existing Flux manifest environment entry and its rollout annotation, commit, push, and reconcile Hermes while no job is active:

- name: PLANNING_CLAUDE_MODEL
  value: claude-sonnet-5-5

The allowlisted IDs are claude-opus-4-8, claude-opus-5-5, and claude-sonnet-5-5. Reverting the upgrade commit restores the former CLI staging, model pin, and execution revision. Reconcile with:

flux reconcile kustomization hermes --namespace flux-system --with-source

Do not restart during a job: provider work is never automatically retried, and in-memory results expire on worker restart. The server comparison runs used the actual multi-pass backend inside isolated temporary directories, not laptop connectivity or HTTP submission. Results and concise review are retained in synthetic evidence. The operator probe accepts no input file and constructs only the fixed synthetic 14-case fixture. Focused mocked regression tests passed: 122. Kustomize render and client dry-run passed; Flux diff was limited to the planner Deployment and ConfigMap. The documented combined server-side/client dry-run flags are incompatible with the installed kubectl, so the client dry-run was run without --server-side.

Deployment verification

Flux applied 206c61993d1ff50ed64ab61164b9eee8cd183b45, and the new worker reports CLI 2.1.285. In-cluster application HTTP checks returned authenticated health/capabilities/preflight 200, invalid-token 401, management-path 404, and local-only preflight 422 without selecting an external provider. Preflight pins claude-opus-5-5; capabilities expose all three server configuration choices. These checks used loopback HTTP with the approved client forwarding address and do not establish end-to-end LAN HTTPS connectivity.

The operator host is currently on 192.168.1.64; direct TLS-verified connection to 192.168.22.50:443 timed out before HTTP. No LAN routing, access restrictions, TLS configuration, ingress paths, or client URL were changed to work around that host limitation. The existing client endpoint remains https://worker.bstein.dev/suite-planning.

Official references: model configuration and CLI settings.