151 lines
8.0 KiB
Markdown
151 lines
8.0 KiB
Markdown
# Claude 5.5 suite comparison
|
|
|
|
The suite planner now pins native Claude Code 2.1.285. Opus 5.5 and Sonnet 5.5
|
|
both returned their exact canonical model identities on the existing first-party
|
|
OAuth account. The default deployment selects `claude-opus-5-5`; Sonnet is a
|
|
server configuration option, not an automatic fallback. The HTTPS contract,
|
|
credential scopes, implementation objective, five-case task cap,
|
|
30-minute deadline, and USD 30 CLI estimated-cost guard are unchanged.
|
|
|
|
Execution revision `suite-multipass-v7-20260929` uses `high` by default and
|
|
`xhigh` for at least 100 cases OR at least 131,072 bytes (128 KiB) of complete
|
|
suite content. Byte size is the canonical UTF-8 JSON of `campaign`, `suite`, and
|
|
all `cases`, without transport whitespace, routing/execution fields, prompt text,
|
|
schemas, or generated proposals. The selected effort is fixed for every pass in
|
|
that job, including structured-output repair; later review prompts do not change
|
|
it. There is no classifier or auxiliary inference call to choose effort.
|
|
|
|
Capabilities expose the thresholds in `models.claude.reasoning_policy`.
|
|
Preflight and job metadata expose `selection.reasoning` plus
|
|
`selection.reasoning_selection` with the policy revision, measured source bytes,
|
|
case count, and triggered thresholds. Per-pass metadata and completed results
|
|
also contain `reasoning`. The policy revision is `suite-size-effort-v1-20260929`.
|
|
The client does not need a new request field. Explicit smaller job budgets remain
|
|
effective; the service fails instead of reducing effort or accepting partial work.
|
|
|
|
The comparison below used `medium`; its runtime and cost measurements do not
|
|
measure high or xhigh effort. Both Opus 5.5 and
|
|
Sonnet 5.5 support `low`, `medium`, `high`, `xhigh`, and `max`. High is the next
|
|
step above medium; xhigh and max spend more reasoning tokens and may take longer,
|
|
without guaranteeing better grouping. No hosted inference or real roster rerun
|
|
is needed to apply this setting. The HTTP compatibility and prompt revisions
|
|
remain unchanged.
|
|
|
|
## Matched 14-case results
|
|
|
|
Each run used the same 9,851-byte normalized synthetic request, complete source
|
|
fields, schema, and three-pass workflow (two independent proposals and one
|
|
whole-suite reconciliation). Each completed in six CLI turns without a retry.
|
|
No family exceeded five, so these hosted tests did not need the two extra
|
|
oversized-family review passes.
|
|
|
|
| Measurement | Opus 5.5 | Sonnet 5.5 |
|
|
| --- | ---: | ---: |
|
|
| Wall seconds | 39.256 | 24.799 |
|
|
| Aggregate input tokens | 21,936 | 21,438 |
|
|
| Aggregate output tokens | 4,243 | 3,814 |
|
|
| CLI estimated USD | 0.172604 | 0.081016 |
|
|
| Natural families / final tasks | 9 / 9 | 9 / 9 |
|
|
| Singleton families | 4 | 4 |
|
|
| Exact alias coverage | Yes | Yes |
|
|
| Incorrect merge / missed merge pairs | 0 / 0 | 0 / 0 |
|
|
|
|
Both produced identical memberships matching the fixture's independent mechanism
|
|
labels, including distant related cases, identical text under different aliases,
|
|
and watchdog terminology shared by different test machinery. Review found no
|
|
description importing another family's objectives. Opus explicitly retained
|
|
unknown procedure details for the thermal, acoustic, and build singleton fixtures.
|
|
Sonnet's descriptions were more concise and omitted those uncertainty notes.
|
|
|
|
Membership quality tied on this small fixture. Sonnet used about 53% less CLI
|
|
estimated cost and finished about 37% sooner. This is one observation per model,
|
|
not evidence that their performance or quality is equivalent on a difficult
|
|
363-case suite. The amounts are CLI accounting estimates, not subscription
|
|
charges or proof of a dollar-denominated account bill.
|
|
|
|
## Runtime limits and isolation
|
|
|
|
Both CLI result envelopes report a 1,000,000-token context and 128,000-token model
|
|
output ceiling. The service deliberately continues to request at most 64,000
|
|
output tokens per invocation. `selection.output` is that service limit;
|
|
`selection.reported_output` and `model_usage[*].maxOutputTokens` describe the
|
|
reported model ceiling. The native CLI loopback check captured `max_tokens=64000`
|
|
and complete source/system text on all four requests, including one deliberately
|
|
invalid structured assignment followed by repair.
|
|
|
|
The newer CLI loads a built-in instruction plugin even in safe mode unless it is
|
|
explicitly disabled. The service now disables it, all hooks, bundled skills,
|
|
provider connectors, plugin synchronization, built-in subagents, and git
|
|
instructions. Each CLI invocation allowlists only its pinned model and turns off
|
|
automatic model switching when a request is flagged. Runtime identity, empty
|
|
plugins/MCP lists, and tool isolation are still checked before accepting output.
|
|
No provider or model fallback was added. The first tiny Opus readiness probe
|
|
failed with zero usage and insufficient retained diagnostics; a fresh probe
|
|
succeeded. Neither measured suite job needed a retry.
|
|
|
|
The artifact is downloaded from the official release host at startup with a
|
|
pinned version, 240,327,864-byte size, and SHA-256:
|
|
|
|
```text
|
|
33dad1ec615a2e08cc78b494f05c110e49916de2c79d78ec8799ebf46b233d29
|
|
```
|
|
|
|
This removes the planner's dependency on the shared agent tools volume. Other
|
|
Hermes CLIs and the isolated local-model endpoint are unchanged. A cold start now
|
|
requires the official release download host to be reachable; failed download or
|
|
checksum verification leaves the worker unavailable instead of using another CLI.
|
|
|
|
## Configuration and rollback
|
|
|
|
The HTTP compatibility revision remains `suite-v6-20260929`, the prompt remains
|
|
`implementation-proximity-multipass-v4-20260929`, and the execution revision is
|
|
`suite-multipass-v7-20260929`. Authenticated capabilities additionally expose
|
|
`claude_model_options` and `model_selection: server_configuration`. No new client
|
|
request field is accepted or required.
|
|
|
|
To select Sonnet, change the existing Flux manifest environment entry and its
|
|
rollout annotation, commit, push, and reconcile Hermes while no job is active:
|
|
|
|
```yaml
|
|
- name: PLANNING_CLAUDE_MODEL
|
|
value: claude-sonnet-5-5
|
|
```
|
|
|
|
The allowlisted IDs are `claude-opus-4-8`, `claude-opus-5-5`, and
|
|
`claude-sonnet-5-5`. Reverting the upgrade commit restores the former CLI staging,
|
|
model pin, and execution revision. Reconcile with:
|
|
|
|
```bash
|
|
flux reconcile kustomization hermes --namespace flux-system --with-source
|
|
```
|
|
|
|
Do not restart during a job: provider work is never automatically retried, and
|
|
in-memory results expire on worker restart. The server comparison runs used the
|
|
actual multi-pass backend inside isolated temporary directories, not laptop
|
|
connectivity or HTTP submission. Results and concise review are retained in
|
|
[synthetic evidence](evidence/hermes_suite_claude55_20260929.json). The operator
|
|
probe accepts no input file and constructs only the fixed synthetic 14-case
|
|
fixture. Focused mocked regression tests passed: 122. Kustomize render and client
|
|
dry-run passed; Flux diff was limited to the planner Deployment and ConfigMap.
|
|
The documented combined server-side/client dry-run flags are incompatible with
|
|
the installed kubectl, so the client dry-run was run without `--server-side`.
|
|
|
|
## Deployment verification
|
|
|
|
Flux applied `206c61993d1ff50ed64ab61164b9eee8cd183b45`, and the new worker
|
|
reports CLI `2.1.285`. In-cluster application HTTP checks returned authenticated
|
|
health/capabilities/preflight 200, invalid-token 401, management-path 404, and
|
|
local-only preflight 422 without selecting an external provider. Preflight pins
|
|
`claude-opus-5-5`; capabilities expose all three server configuration choices.
|
|
These checks used loopback HTTP with the approved client forwarding address and
|
|
do not establish end-to-end LAN HTTPS connectivity.
|
|
|
|
The operator host is currently on `192.168.1.64`; direct TLS-verified connection
|
|
to `192.168.22.50:443` timed out before HTTP. No LAN routing, access restrictions,
|
|
TLS configuration, ingress paths, or client URL were changed to work around that
|
|
host limitation. The existing client endpoint remains
|
|
`https://worker.bstein.dev/suite-planning`.
|
|
|
|
Official references: [model configuration](https://code.claude.com/docs/en/model-config)
|
|
and [CLI settings](https://code.claude.com/docs/en/settings-reference).
|