# Claude 5.5 suite comparison The suite planner now pins native Claude Code 2.1.285. Opus 5.5 and Sonnet 5.5 both returned their exact canonical model identities on the existing first-party OAuth account. The default deployment selects `claude-opus-5-5`; Sonnet is a server configuration option, not an automatic fallback. The HTTPS contract, credential scopes, implementation objective, five-case task cap, 30-minute deadline, and USD 30 CLI estimated-cost guard are unchanged. Execution revision `suite-multipass-v7-20260929` uses `high` by default and `xhigh` for at least 100 cases OR at least 131,072 bytes (128 KiB) of complete suite content. Byte size is the canonical UTF-8 JSON of `campaign`, `suite`, and all `cases`, without transport whitespace, routing/execution fields, prompt text, schemas, or generated proposals. The selected effort is fixed for every pass in that job, including structured-output repair; later review prompts do not change it. There is no classifier or auxiliary inference call to choose effort. Capabilities expose the thresholds in `models.claude.reasoning_policy`. Preflight and job metadata expose `selection.reasoning` plus `selection.reasoning_selection` with the policy revision, measured source bytes, case count, and triggered thresholds. Per-pass metadata and completed results also contain `reasoning`. The policy revision is `suite-size-effort-v1-20260929`. The client does not need a new request field. Explicit smaller job budgets remain effective; the service fails instead of reducing effort or accepting partial work. The comparison below used `medium`; its runtime and cost measurements do not measure high or xhigh effort. Both Opus 5.5 and Sonnet 5.5 support `low`, `medium`, `high`, `xhigh`, and `max`. High is the next step above medium; xhigh and max spend more reasoning tokens and may take longer, without guaranteeing better grouping. No hosted inference or real roster rerun is needed to apply this setting. The HTTP compatibility and prompt revisions remain unchanged. ## Matched 14-case results Each run used the same 9,851-byte normalized synthetic request, complete source fields, schema, and three-pass workflow (two independent proposals and one whole-suite reconciliation). Each completed in six CLI turns without a retry. No family exceeded five, so these hosted tests did not need the two extra oversized-family review passes. | Measurement | Opus 5.5 | Sonnet 5.5 | | --- | ---: | ---: | | Wall seconds | 39.256 | 24.799 | | Aggregate input tokens | 21,936 | 21,438 | | Aggregate output tokens | 4,243 | 3,814 | | CLI estimated USD | 0.172604 | 0.081016 | | Natural families / final tasks | 9 / 9 | 9 / 9 | | Singleton families | 4 | 4 | | Exact alias coverage | Yes | Yes | | Incorrect merge / missed merge pairs | 0 / 0 | 0 / 0 | Both produced identical memberships matching the fixture's independent mechanism labels, including distant related cases, identical text under different aliases, and watchdog terminology shared by different test machinery. Review found no description importing another family's objectives. Opus explicitly retained unknown procedure details for the thermal, acoustic, and build singleton fixtures. Sonnet's descriptions were more concise and omitted those uncertainty notes. Membership quality tied on this small fixture. Sonnet used about 53% less CLI estimated cost and finished about 37% sooner. This is one observation per model, not evidence that their performance or quality is equivalent on a difficult 363-case suite. The amounts are CLI accounting estimates, not subscription charges or proof of a dollar-denominated account bill. ## Runtime limits and isolation Both CLI result envelopes report a 1,000,000-token context and 128,000-token model output ceiling. The service deliberately continues to request at most 64,000 output tokens per invocation. `selection.output` is that service limit; `selection.reported_output` and `model_usage[*].maxOutputTokens` describe the reported model ceiling. The native CLI loopback check captured `max_tokens=64000` and complete source/system text on all four requests, including one deliberately invalid structured assignment followed by repair. The newer CLI loads a built-in instruction plugin even in safe mode unless it is explicitly disabled. The service now disables it, all hooks, bundled skills, provider connectors, plugin synchronization, built-in subagents, and git instructions. Each CLI invocation allowlists only its pinned model and turns off automatic model switching when a request is flagged. Runtime identity, empty plugins/MCP lists, and tool isolation are still checked before accepting output. No provider or model fallback was added. The first tiny Opus readiness probe failed with zero usage and insufficient retained diagnostics; a fresh probe succeeded. Neither measured suite job needed a retry. The artifact is downloaded from the official release host at startup with a pinned version, 240,327,864-byte size, and SHA-256: ```text 33dad1ec615a2e08cc78b494f05c110e49916de2c79d78ec8799ebf46b233d29 ``` This removes the planner's dependency on the shared agent tools volume. Other Hermes CLIs and the isolated local-model endpoint are unchanged. A cold start now requires the official release download host to be reachable; failed download or checksum verification leaves the worker unavailable instead of using another CLI. ## Configuration and rollback The HTTP compatibility revision remains `suite-v6-20260929`, the prompt remains `implementation-proximity-multipass-v4-20260929`, and the execution revision is `suite-multipass-v7-20260929`. Authenticated capabilities additionally expose `claude_model_options` and `model_selection: server_configuration`. No new client request field is accepted or required. To select Sonnet, change the existing Flux manifest environment entry and its rollout annotation, commit, push, and reconcile Hermes while no job is active: ```yaml - name: PLANNING_CLAUDE_MODEL value: claude-sonnet-5-5 ``` The allowlisted IDs are `claude-opus-4-8`, `claude-opus-5-5`, and `claude-sonnet-5-5`. Reverting the upgrade commit restores the former CLI staging, model pin, and execution revision. Reconcile with: ```bash flux reconcile kustomization hermes --namespace flux-system --with-source ``` Do not restart during a job: provider work is never automatically retried, and in-memory results expire on worker restart. The server comparison runs used the actual multi-pass backend inside isolated temporary directories, not laptop connectivity or HTTP submission. Results and concise review are retained in [synthetic evidence](evidence/hermes_suite_claude55_20260929.json). The operator probe accepts no input file and constructs only the fixed synthetic 14-case fixture. Focused mocked regression tests passed: 122. Kustomize render and client dry-run passed; Flux diff was limited to the planner Deployment and ConfigMap. The documented combined server-side/client dry-run flags are incompatible with the installed kubectl, so the client dry-run was run without `--server-side`. ## Deployment verification Flux applied `206c61993d1ff50ed64ab61164b9eee8cd183b45`, and the new worker reports CLI `2.1.285`. In-cluster application HTTP checks returned authenticated health/capabilities/preflight 200, invalid-token 401, management-path 404, and local-only preflight 422 without selecting an external provider. Preflight pins `claude-opus-5-5`; capabilities expose all three server configuration choices. These checks used loopback HTTP with the approved client forwarding address and do not establish end-to-end LAN HTTPS connectivity. The operator host is currently on `192.168.1.64`; direct TLS-verified connection to `192.168.22.50:443` timed out before HTTP. No LAN routing, access restrictions, TLS configuration, ingress paths, or client URL were changed to work around that host limitation. The existing client endpoint remains `https://worker.bstein.dev/suite-planning`. Official references: [model configuration](https://code.claude.com/docs/en/model-config) and [CLI settings](https://code.claude.com/docs/en/settings-reference).