8.1 KiB
Claude 5.5 suite comparison
The suite planner now pins native Claude Code 2.1.285. Opus 5.5 and Sonnet 5.5
both returned their exact canonical model identities on the existing first-party
OAuth account. The default deployment selects claude-opus-5-5; Sonnet is a
server configuration option, not an automatic fallback. The HTTPS contract,
credential scopes, implementation objective, five-case task cap,
30-minute default deadline, and USD 30 CLI estimated-cost guard are unchanged.
An explicit execution.max_seconds: 3600 now permits a 60-minute total job; see
current budget evidence.
Execution revision suite-multipass-v9-20260929 preserves the effort policy: high by default and
xhigh for at least 100 cases OR at least 131,072 bytes (128 KiB) of complete
suite content. Byte size is the canonical UTF-8 JSON of campaign, suite, and
all cases, without transport whitespace, routing/execution fields, prompt text,
schemas, or generated proposals. The selected effort is fixed for every pass in
that job, including structured-output repair; later review prompts do not change
it. There is no classifier or auxiliary inference call to choose effort.
Capabilities expose the thresholds in models.claude.reasoning_policy.
Preflight and job metadata expose selection.reasoning plus
selection.reasoning_selection with the policy revision, measured source bytes,
case count, and triggered thresholds. Per-pass metadata and completed results
also contain reasoning. The policy revision is suite-size-effort-v1-20260929.
The client does not need a new request field. Explicit smaller job budgets remain
effective; the service fails instead of reducing effort or accepting partial work.
The comparison below used medium; its runtime and cost measurements do not
measure high or xhigh effort. Both Opus 5.5 and
Sonnet 5.5 support low, medium, high, xhigh, and max. High is the next
step above medium; xhigh and max spend more reasoning tokens and may take longer,
without guaranteeing better grouping. No hosted inference or real roster rerun
is needed to apply this setting. The HTTP compatibility and prompt revisions
remain unchanged.
Matched 14-case results
Each run used the same 9,851-byte normalized synthetic request, complete source fields, schema, and three-pass workflow (two independent proposals and one whole-suite reconciliation). Each completed in six CLI turns without a retry. No family exceeded five, so these hosted tests did not need the two extra oversized-family review passes.
| Measurement | Opus 5.5 | Sonnet 5.5 |
|---|---|---|
| Wall seconds | 39.256 | 24.799 |
| Aggregate input tokens | 21,936 | 21,438 |
| Aggregate output tokens | 4,243 | 3,814 |
| CLI estimated USD | 0.172604 | 0.081016 |
| Natural families / final tasks | 9 / 9 | 9 / 9 |
| Singleton families | 4 | 4 |
| Exact alias coverage | Yes | Yes |
| Incorrect merge / missed merge pairs | 0 / 0 | 0 / 0 |
Both produced identical memberships matching the fixture's independent mechanism labels, including distant related cases, identical text under different aliases, and watchdog terminology shared by different test machinery. Review found no description importing another family's objectives. Opus explicitly retained unknown procedure details for the thermal, acoustic, and build singleton fixtures. Sonnet's descriptions were more concise and omitted those uncertainty notes.
Membership quality tied on this small fixture. Sonnet used about 53% less CLI estimated cost and finished about 37% sooner. This is one observation per model, not evidence that their performance or quality is equivalent on a difficult 363-case suite. The amounts are CLI accounting estimates, not subscription charges or proof of a dollar-denominated account bill.
Runtime limits and isolation
Both CLI result envelopes report a 1,000,000-token context and 128,000-token model
output ceiling. The service deliberately continues to request at most 64,000
output tokens per invocation. selection.output is that service limit;
selection.reported_output and model_usage[*].maxOutputTokens describe the
reported model ceiling. The native CLI loopback check captured max_tokens=64000
and complete source/system text on all four requests, including one deliberately
invalid structured assignment followed by repair.
The newer CLI loads a built-in instruction plugin even in safe mode unless it is explicitly disabled. The service now disables it, all hooks, bundled skills, provider connectors, plugin synchronization, built-in subagents, and git instructions. Each CLI invocation allowlists only its pinned model and turns off automatic model switching when a request is flagged. Runtime identity, empty plugins/MCP lists, and tool isolation are still checked before accepting output. No provider or model fallback was added. The first tiny Opus readiness probe failed with zero usage and insufficient retained diagnostics; a fresh probe succeeded. Neither measured suite job needed a retry.
The artifact is downloaded from the official release host at startup with a pinned version, 240,327,864-byte size, and SHA-256:
33dad1ec615a2e08cc78b494f05c110e49916de2c79d78ec8799ebf46b233d29
This removes the planner's dependency on the shared agent tools volume. Other Hermes CLIs and the isolated local-model endpoint are unchanged. A cold start now requires the official release download host to be reachable; failed download or checksum verification leaves the worker unavailable instead of using another CLI.
Configuration and rollback
The HTTP compatibility revision remains suite-v6-20260929, the prompt remains
implementation-proximity-multipass-v4-20260929, and the execution revision is
suite-multipass-v9-20260929. Authenticated capabilities additionally expose
claude_model_options and model_selection: server_configuration. No new client
request field is accepted or required.
To select Sonnet, change the existing Flux manifest environment entry and its rollout annotation, commit, push, and reconcile Hermes while no job is active:
- name: PLANNING_CLAUDE_MODEL
value: claude-sonnet-5-5
The allowlisted IDs are claude-opus-4-8, claude-opus-5-5, and
claude-sonnet-5-5. Reverting the upgrade commit restores the former CLI staging,
model pin, and execution revision. Reconcile with:
flux reconcile kustomization hermes --namespace flux-system --with-source
Do not restart during a job: provider work is never automatically retried, and
in-memory results expire on worker restart. The server comparison runs used the
actual multi-pass backend inside isolated temporary directories, not laptop
connectivity or HTTP submission. Results and concise review are retained in
synthetic evidence. The operator
probe accepts no input file and constructs only the fixed synthetic 14-case
fixture. Focused mocked regression tests passed: 122. Kustomize render and client
dry-run passed; Flux diff was limited to the planner Deployment and ConfigMap.
The documented combined server-side/client dry-run flags are incompatible with
the installed kubectl, so the client dry-run was run without --server-side.
Deployment verification
Flux applied 206c61993d1ff50ed64ab61164b9eee8cd183b45, and the new worker
reports CLI 2.1.285. In-cluster application HTTP checks returned authenticated
health/capabilities/preflight 200, invalid-token 401, management-path 404, and
local-only preflight 422 without selecting an external provider. Preflight pins
claude-opus-5-5; capabilities expose all three server configuration choices.
These checks used loopback HTTP with the approved client forwarding address and
do not establish end-to-end LAN HTTPS connectivity.
The operator host is currently on 192.168.1.64; direct TLS-verified connection
to 192.168.22.50:443 timed out before HTTP. No LAN routing, access restrictions,
TLS configuration, ingress paths, or client URL were changed to work around that
host limitation. The existing client endpoint remains
https://worker.bstein.dev/suite-planning.
Official references: model configuration and CLI settings.