hermes: upgrade suite planner for Claude 5.5 models

This commit is contained in:
jenkins 2026-09-29 15:50:36 -05:00
parent 3036064bf7
commit 206c61993d
10 changed files with 2433 additions and 29 deletions

File diff suppressed because it is too large Load Diff

View File

@ -0,0 +1,110 @@
# Claude 5.5 suite comparison
The suite planner now pins native Claude Code 2.1.285. Opus 5.5 and Sonnet 5.5
both returned their exact canonical model identities on the existing first-party
OAuth account. The default deployment selects `claude-opus-5-5`; Sonnet is a
server configuration option, not an automatic fallback. The HTTPS contract,
credential scopes, implementation objective, medium effort, five-case task cap,
30-minute deadline, and USD 30 CLI estimated-cost guard are unchanged.
## Matched 14-case results
Each run used the same 9,851-byte normalized synthetic request, complete source
fields, schema, and three-pass workflow (two independent proposals and one
whole-suite reconciliation). Each completed in six CLI turns without a retry.
No family exceeded five, so these hosted tests did not need the two extra
oversized-family review passes.
| Measurement | Opus 5.5 | Sonnet 5.5 |
| --- | ---: | ---: |
| Wall seconds | 39.256 | 24.799 |
| Aggregate input tokens | 21,936 | 21,438 |
| Aggregate output tokens | 4,243 | 3,814 |
| CLI estimated USD | 0.172604 | 0.081016 |
| Natural families / final tasks | 9 / 9 | 9 / 9 |
| Singleton families | 4 | 4 |
| Exact alias coverage | Yes | Yes |
| Incorrect merge / missed merge pairs | 0 / 0 | 0 / 0 |
Both produced identical memberships matching the fixture's independent mechanism
labels, including distant related cases, identical text under different aliases,
and watchdog terminology shared by different test machinery. Review found no
description importing another family's objectives. Opus explicitly retained
unknown procedure details for the thermal, acoustic, and build singleton fixtures.
Sonnet's descriptions were more concise and omitted those uncertainty notes.
Membership quality tied on this small fixture. Sonnet used about 53% less CLI
estimated cost and finished about 37% sooner. This is one observation per model,
not evidence that their performance or quality is equivalent on a difficult
363-case suite. The amounts are CLI accounting estimates, not subscription
charges or proof of a dollar-denominated account bill.
## Runtime limits and isolation
Both CLI result envelopes report a 1,000,000-token context and 128,000-token model
output ceiling. The service deliberately continues to request at most 64,000
output tokens per invocation. `selection.output` is that service limit;
`selection.reported_output` and `model_usage[*].maxOutputTokens` describe the
reported model ceiling. The native CLI loopback check captured `max_tokens=64000`
and complete source/system text on all four requests, including one deliberately
invalid structured assignment followed by repair.
The newer CLI loads a built-in instruction plugin even in safe mode unless it is
explicitly disabled. The service now disables it, all hooks, bundled skills,
provider connectors, plugin synchronization, built-in subagents, and git
instructions. Each CLI invocation allowlists only its pinned model and turns off
automatic model switching when a request is flagged. Runtime identity, empty
plugins/MCP lists, and tool isolation are still checked before accepting output.
No provider or model fallback was added. The first tiny Opus readiness probe
failed with zero usage and insufficient retained diagnostics; a fresh probe
succeeded. Neither measured suite job needed a retry.
The artifact is downloaded from the official release host at startup with a
pinned version, 240,327,864-byte size, and SHA-256:
```text
33dad1ec615a2e08cc78b494f05c110e49916de2c79d78ec8799ebf46b233d29
```
This removes the planner's dependency on the shared agent tools volume. Other
Hermes CLIs and the isolated local-model endpoint are unchanged. A cold start now
requires the official release download host to be reachable; failed download or
checksum verification leaves the worker unavailable instead of using another CLI.
## Configuration and rollback
The HTTP compatibility revision remains `suite-v6-20260929`, the prompt remains
`implementation-proximity-multipass-v4-20260929`, and the execution revision is
`suite-multipass-v5-20260929`. Authenticated capabilities additionally expose
`claude_model_options` and `model_selection: server_configuration`. No new client
request field is accepted or required.
To select Sonnet, change the existing Flux manifest environment entry and its
rollout annotation, commit, push, and reconcile Hermes while no job is active:
```yaml
- name: PLANNING_CLAUDE_MODEL
value: claude-sonnet-5-5
```
The allowlisted IDs are `claude-opus-4-8`, `claude-opus-5-5`, and
`claude-sonnet-5-5`. Reverting the upgrade commit restores the former CLI staging,
model pin, and execution revision. Reconcile with:
```bash
flux reconcile kustomization hermes --namespace flux-system --with-source
```
Do not restart during a job: provider work is never automatically retried, and
in-memory results expire on worker restart. The server comparison runs used the
actual multi-pass backend inside isolated temporary directories, not laptop
connectivity or HTTP submission. Results and concise review are retained in
[synthetic evidence](evidence/hermes_suite_claude55_20260929.json). The operator
probe accepts no input file and constructs only the fixed synthetic 14-case
fixture. Focused mocked regression tests passed: 122. Kustomize render and client
dry-run passed; Flux diff was limited to the planner Deployment and ConfigMap.
The documented combined server-side/client dry-run flags are incompatible with
the installed kubectl, so the client dry-run was run without `--server-side`.
Official references: [model configuration](https://code.claude.com/docs/en/model-config)
and [CLI settings](https://code.claude.com/docs/en/settings-reference).

View File

@ -11,7 +11,7 @@ most **64 characters**, unique after whitespace and case normalization.
- Configuration: `suite-v6-20260929` (HTTP compatibility identifier).
- Policy: `implementation-five-v1-20260929`.
- Prompt: `implementation-proximity-multipass-v4-20260929`.
- Execution: `suite-multipass-v4-20260929`.
- Execution: `suite-multipass-v5-20260929`.
A single server-side job performs:
@ -45,13 +45,19 @@ missing review decisions still fail closed; no missing assignment is fabricated.
Each invocation retains six CLI turns for structured output. Model review passes
and CLI turns are separate counters. All calls use the originally selected provider
and pinned model. The model remains `claude-opus-4-8[1m]`, with canonical runtime
identity checked as `claude-opus-4-8`, firstParty, medium effort, reported 1M context
and 64K output ceiling. The server invokes native Claude Code CLI 2.1.226 using
and pinned model. The current default is `claude-opus-5-5[1m]`, with canonical runtime
identity checked as `claude-opus-5-5`, firstParty, medium effort, reported 1M context
and 128K model output ceiling, with requests limited to 64K output tokens.
The server invokes native Claude Code CLI 2.1.285 using
the existing first-party OAuth account, not a separately configured API-key
account. The CLI itself communicates with Anthropic over HTTPS. No new provider
fallback or tools are enabled.
Sonnet 5.5 is also verified on the 14-case synthetic fixture and available through
server configuration. See the [matched comparison](hermes_suite_claude55.md).
Earlier acceptance measurements below used Opus 4.8; they do not establish large
suite capacity or quality for the new models.
## Deterministic task sizing and names
For each remaining coherent large family, `k = ceil(n / 5)` and

View File

@ -0,0 +1,65 @@
#!/usr/bin/env python3
"""Compare one configured Claude model on the public synthetic 14-case suite.
Run server-side with the planner modules and scoped OAuth mount. This invokes
the same multi-pass backend without creating an HTTP job. Output is synthetic
evidence only; progress contains metadata. Never accepts a roster input file.
"""
import json
import os
import sys
import threading
import time
sys.path.insert(0, os.environ.get("SUITE_PROBE_MODULE_DIR", "/opt/planner"))
import suite_backends
from suite_contract import (EXECUTION_REVISION, MODELS, PROMPT_REVISION, Problem,
encoded, validate_request, validate_result)
from suite_multipass import generate, preflight_workflow
from suite_synthetic import fixture, score
def main():
"""Run one fixed fixture and emit normalized results plus measured metadata."""
model = MODELS["claude"]["model"]
if model not in {"claude-opus-5-5", "claude-sonnet-5-5"}:
raise SystemExit("Configure one of the two approved comparison models")
if os.environ.get("SUITE_PROBE_CLAUDE_BIN"):
suite_backends.CLAUDE_BIN = os.environ["SUITE_PROBE_CLAUDE_BIN"]
request, expected = fixture(14)
request["routing"] = {"allow_external": True, "allowed_external_providers": ["claude"]}
request = validate_request(request, ["claude"])
selected = preflight_workflow(request)
started, last_report = time.monotonic(), 0.0
last_stage = None
def progress(value):
"""Emit stage and elapsed time without model-generated content."""
nonlocal last_report, last_stage
stage = value["current_pass"]
if stage != last_stage or time.monotonic() - last_report >= 20:
print(json.dumps({"model": model, "stage": stage,
"elapsed_seconds": value["job_elapsed_seconds"],
"completed_passes": value["completed_model_passes"]}),
file=sys.stderr, flush=True)
last_report, last_stage = time.monotonic(), stage
report = {"model": model, "execution_revision": EXECUTION_REVISION,
"prompt_revision": PROMPT_REVISION, "selection": selected,
"request_bytes": len(encoded(request)), "case_count": 14,
"automatic_retries": 0, "test_path": "server_side_multipass_backend"}
try:
result, metadata = generate(request, selected, threading.Event(), "192.168.22.8", progress)
validate_result(result, request)
report.update(status="completed", result=result, metadata=metadata,
quality=score(result, expected),
singletons=sum(len(g["members"]) == 1 for g in result["groups"]))
except Problem as exc:
report.update(status="failed", **exc.document())
report["wall_seconds"] = round(time.monotonic() - started, 3)
print(json.dumps(report, sort_keys=True), flush=True)
return 0 if report["status"] == "completed" else 1
if __name__ == "__main__":
raise SystemExit(main())

View File

@ -12,7 +12,7 @@ import re
import subprocess
import threading
from suite_contract import (COST_LIMIT, EXECUTION_REVISION, MAX_BODY, MAX_CASES, MAX_RESULT, MODELS, PROMPT_REVISION,
from suite_contract import (CLAUDE_MODELS, CLAUDE_VERSION, COST_LIMIT, EXECUTION_REVISION, MAX_BODY, MAX_CASES, MAX_RESULT, MODELS, PROMPT_REVISION,
PROMPT_SHA256, REVISION, TIMEOUT, Problem, encoded,
preflight, validate_request)
from suite_jobs import Jobs
@ -124,6 +124,7 @@ class Handler(BaseHTTPRequestHandler):
"prompt_revision": PROMPT_REVISION, "prompt_sha256": PROMPT_SHA256})
if method == "GET" and self.path == "/v1/capabilities":
return self.send(200, {"configuration_revision": REVISION, "models": MODELS,
"claude_model_options": list(CLAUDE_MODELS), "model_selection": "server_configuration",
"execution_revision": EXECUTION_REVISION,
"policy_revision": POLICY_REVISION, "max_final_group_cases": 5,
"max_final_name_characters": 64, "model_passes": {"minimum": 3, "maximum": 5},
@ -214,7 +215,7 @@ def main():
raise SystemExit("Pinned CLI binary mismatch")
version = subprocess.run(["/opt/cli/claude", "--version"], capture_output=True,
text=True, timeout=10, check=True)
if version.stdout.strip() != "2.1.226 (Claude Code)":
if version.stdout.strip() != CLAUDE_VERSION + " (Claude Code)":
raise SystemExit("Pinned CLI version mismatch")
private = ThreadingHTTPServer(("0.0.0.0", 9001), Decision)
threading.Thread(target=private.serve_forever, daemon=True).start()

View File

@ -11,7 +11,7 @@ import time
from urllib.error import HTTPError, URLError
from urllib.request import HTTPRedirectHandler, ProxyHandler, Request, build_opener
from suite_contract import CLAUDE_MAX_TURNS, MODELS, SCHEMA, SYSTEM, Problem, encoded, prompt
from suite_contract import CLAUDE_MAX_TURNS, CLAUDE_MODELS, MODELS, SCHEMA, SYSTEM, Problem, encoded, prompt
from suite_cli_diagnostics import snapshot
SWITCHYARD = "http://hermes-switchyard.hermes.svc.cluster.local:9005/v1/chat/completions"
@ -97,12 +97,19 @@ def local_generate(request, cancel, client_ip, *, invocation=None):
def claude_command(model, max_cost):
"""Use the native pinned binary, never the privileged Hermes shell wrapper."""
if model not in CLAUDE_MODELS:
raise Problem("unsupported_backend", 422)
settings = {"enabledPlugins": {"agents-md@builtin": False,
"cc-plugin-agents-md@builtin": False},
"disableAllHooks": True, "disableBundledSkills": True,
"disableClaudeAiConnectors": True, "syncClaudeAiPlugins": False,
"availableModels": [model], "switchModelsOnFlag": False}
return [CLAUDE_BIN, "-p", "--output-format", "stream-json", "--verbose",
"--no-session-persistence", "--safe-mode", "--tools", "",
"--strict-mcp-config", "--mcp-config", '{"mcpServers":{}}',
"--setting-sources", "", "--disable-slash-commands",
"--setting-sources", "", "--settings", encoded(settings).decode(), "--disable-slash-commands",
"--permission-mode", "dontAsk", "--no-chrome",
"--model", MODELS["claude"]["cli_model"], "--effort", "medium", "--max-budget-usd", str(max_cost),
"--model", model + "[1m]", "--effort", "medium", "--max-budget-usd", str(max_cost),
"--max-turns", str(CLAUDE_MAX_TURNS), "--system-prompt", SYSTEM,
"--json-schema", encoded(SCHEMA).decode()]
@ -117,7 +124,9 @@ def claude_environment(directory, token):
"DISABLE_ERROR_REPORTING", "DISABLE_AUTOUPDATER", "DISABLE_UPDATES",
"DISABLE_PROMPT_CACHING", "CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC",
"CLAUDE_CODE_DISABLE_BACKGROUND_TASKS", "CLAUDE_CODE_DISABLE_TERMINAL_TITLE",
"CLAUDE_CODE_DISABLE_AUTO_MEMORY", "CLAUDE_CODE_SKIP_PROMPT_HISTORY"):
"CLAUDE_CODE_DISABLE_AUTO_MEMORY", "CLAUDE_CODE_SKIP_PROMPT_HISTORY",
"CLAUDE_AGENT_SDK_DISABLE_BUILTIN_AGENTS", "CLAUDE_CODE_DISABLE_GIT_INSTRUCTIONS",
"CLAUDE_CODE_DISABLE_BUNDLED_SKILLS"):
env[key] = "1"
return env
@ -180,7 +189,8 @@ def parse_claude(raw, expected_model, **process_info):
if not isinstance(models, dict) or not models or set(models) - aliases:
fail("model_changed", "runtime_model")
for limits in models.values():
if (not isinstance(limits, dict) or limits.get("contextWindow") != 1000000 or limits.get("maxOutputTokens") != 64000
if (not isinstance(limits, dict) or limits.get("contextWindow") != 1000000
or limits.get("maxOutputTokens") != CLAUDE_MODELS.get(expected_model)
or limits.get("canonicalModel") != expected_model
or limits.get("provider") != "firstParty"):
fail("backend_capabilities_changed", "runtime_model_limits")
@ -196,7 +206,7 @@ def parse_claude(raw, expected_model, **process_info):
fail("incomplete_generation", "process_exit")
# Persist measured metadata only, including on later schema/coverage failure.
safe_models = {name: {"canonicalModel": expected_model, "provider": "firstParty",
"contextWindow": 1000000, "maxOutputTokens": 64000}
"contextWindow": 1000000, "maxOutputTokens": CLAUDE_MODELS[expected_model]}
for name in models}
return result, {"model": expected_model, "compaction": False,
"truncation": False, "compaction_signal": "CLI events and disabled compaction",

View File

@ -30,6 +30,7 @@ def usage_counts(value):
"input_tokens", "output_tokens", "cache_read_input_tokens", "cache_creation_input_tokens")
if key in value}
for key, fields in (("server_tool_use", ("web_search_requests", "web_fetch_requests")),
("output_tokens_details", ("thinking_tokens",)),
("cache_creation", ("ephemeral_1h_input_tokens", "ephemeral_5m_input_tokens"))):
if isinstance(value.get(key), dict):
result[key] = {field: number(value[key][field]) for field in fields if field in value[key]}

View File

@ -3,12 +3,22 @@ from __future__ import annotations
import hashlib
import json
import os
import re
from collections import Counter
REVISION = "suite-v6-20260929"
PROMPT_REVISION = "implementation-proximity-multipass-v4-20260929"
EXECUTION_REVISION = "suite-multipass-v4-20260929"
EXECUTION_REVISION = "suite-multipass-v5-20260929"
CLAUDE_VERSION = "2.1.285"
CLAUDE_MODELS = {
"claude-opus-4-8": 64000,
"claude-opus-5-5": 128000,
"claude-sonnet-5-5": 128000,
}
CLAUDE_MODEL = os.environ.get("PLANNING_CLAUDE_MODEL", "claude-opus-4-8")
if CLAUDE_MODEL not in CLAUDE_MODELS:
raise RuntimeError("Unsupported configured Claude model")
CLAUDE_MAX_TURNS = 6
MAX_BODY = 1 << 20
MAX_RESULT = 1 << 20
@ -23,9 +33,10 @@ MODELS = {
"local": {"model": "qwen2.5:14b-instruct-q4_0", "context": 8192,
"output": 2048, "overhead": 1024, "backend": "ollama-model-gate",
"enabled": True, "reasoning": "none"},
"claude": {"model": "claude-opus-4-8", "context": 1000000,
"cli_model": "claude-opus-4-8[1m]",
"output": 64000, "overhead": 8192, "backend": "claude-code-2.1.226",
"claude": {"model": CLAUDE_MODEL, "context": 1000000,
"cli_model": CLAUDE_MODEL + "[1m]",
"output": 64000, "reported_output": CLAUDE_MODELS[CLAUDE_MODEL],
"overhead": 8192, "backend": "claude-code-" + CLAUDE_VERSION,
"enabled": True, "reasoning": "medium", "max_turns": CLAUDE_MAX_TURNS},
"codex": {"model": "gpt-6-astra", "context": 258400,
"output": None, "overhead": None, "backend": "codex-subscription-broker",

View File

@ -36,7 +36,7 @@ spec:
app: hermes-suite-planner
annotations:
fluentbit.io/exclude: "true"
ai.bstein.dev/config-rev: suite-v6-multipass-cap5-v4-20260929
ai.bstein.dev/config-rev: suite-v6-multipass-cap5-v5-20260929
vault.hashicorp.com/agent-inject: "true"
vault.hashicorp.com/agent-pre-populate-only: "true"
vault.hashicorp.com/agent-init-first: "true"
@ -80,7 +80,7 @@ spec:
automountServiceAccountToken: false
enableServiceLinks: false
terminationGracePeriodSeconds: 15
# The installed amd64 CLI is copied from the existing RWO tools volume.
# The pinned native CLI is amd64; retain the existing worker placement.
nodeSelector:
kubernetes.io/hostname: titan-22
securityContext:
@ -97,18 +97,27 @@ spec:
command: [python, -c]
args:
- |
import hashlib,pathlib,shutil
source=pathlib.Path('/installed/lib/node_modules/@anthropic-ai/claude-code/bin/claude.exe')
assert hashlib.sha256(source.read_bytes()).hexdigest() == '4e9bec1177ce9690e8bd988b710ac24105e70da428dd094c5adcbbe786a55555'
shutil.copyfile(source, '/opt/cli/claude')
pathlib.Path('/opt/cli/claude').chmod(0o555)
import hashlib,pathlib,urllib.request
target=pathlib.Path('/opt/cli/claude')
url='https://downloads.claude.ai/claude-code-releases/2.1.285/linux-x64/claude'
digest=hashlib.sha256()
size=0
with urllib.request.urlopen(url, timeout=120) as source, target.open('wb') as output:
while data:=source.read(1048576):
size+=len(data)
if size > 240327864:
raise RuntimeError('CLI artifact exceeds pinned size')
digest.update(data)
output.write(data)
if size != 240327864 or digest.hexdigest() != '33dad1ec615a2e08cc78b494f05c110e49916de2c79d78ec8799ebf46b233d29':
raise RuntimeError('CLI artifact verification failed')
target.chmod(0o555)
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: [ALL]
volumeMounts:
- {name: installed, mountPath: /installed, subPath: tools, readOnly: true}
- {name: cli, mountPath: /opt/cli}
resources:
requests: {cpu: 25m, memory: 256Mi}
@ -122,7 +131,8 @@ spec:
- {name: PYTHONUNBUFFERED, value: "1"}
# User approved generalized CASE records on this Claude account, 2026-09-29.
- {name: PLANNING_GENERALIZED_CLAUDE_APPROVED, value: "true"}
- {name: PLANNING_CLAUDE_SHA256, value: 4e9bec1177ce9690e8bd988b710ac24105e70da428dd094c5adcbbe786a55555}
- {name: PLANNING_CLAUDE_SHA256, value: 33dad1ec615a2e08cc78b494f05c110e49916de2c79d78ec8799ebf46b233d29}
- {name: PLANNING_CLAUDE_MODEL, value: claude-opus-5-5}
ports:
- {name: http, containerPort: 9000}
- {name: decision, containerPort: 9001}
@ -148,10 +158,6 @@ spec:
- name: scripts
configMap:
name: hermes-suite-planner
- name: installed
persistentVolumeClaim:
claimName: hermes-agent-home
readOnly: true
- name: cli
emptyDir: {medium: Memory, sizeLimit: 512Mi}
- name: jobs

View File

@ -0,0 +1,74 @@
"""Model upgrades preserve exact identity, capacity, and CLI isolation."""
import json
from pathlib import Path
import subprocess
import sys
import pytest
SCRIPTS = Path(__file__).resolve().parents[2] / "services/hermes/scripts"
sys.path.insert(0, str(SCRIPTS))
import suite_backends
from suite_cli_diagnostics import usage_counts
from suite_contract import CLAUDE_MODELS, Problem
@pytest.mark.parametrize("model", CLAUDE_MODELS)
def test_only_configured_model_is_allowed_and_plugins_are_disabled(model):
"""Each invocation pins one model, with fallback and customizations disabled."""
command = suite_backends.claude_command(model, 30)
settings = json.loads(command[command.index("--settings") + 1])
assert command[command.index("--model") + 1] == model + "[1m]"
assert settings["availableModels"] == [model]
assert settings["switchModelsOnFlag"] is False
assert settings["disableAllHooks"] and settings["disableClaudeAiConnectors"]
assert settings["syncClaudeAiPlugins"] is False
assert settings["enabledPlugins"]["cc-plugin-agents-md@builtin"] is False
assert "--fallback-model" not in command
env = suite_backends.claude_environment("/jobs/fresh", "synthetic")
assert env["CLAUDE_CODE_MAX_OUTPUT_TOKENS"] == "64000"
assert env["CLAUDE_AGENT_SDK_DISABLE_BUILTIN_AGENTS"] == "1"
@pytest.mark.parametrize("model", ["claude-opus-5-5", "claude-sonnet-5-5"])
@pytest.mark.parametrize("failure", [None, "substitution", "capacity", "plugin"])
def test_new_runtime_envelope_remains_fail_closed(model, failure):
"""A requested name never substitutes for the actual runtime model identity."""
init = {"type": "system", "subtype": "init", "model": model + "[1m]",
"tools": ["StructuredOutput"], "mcp_servers": [], "plugins": []}
limits = {"contextWindow": 1000000, "maxOutputTokens": 128000,
"canonicalModel": model, "provider": "firstParty"}
result = {"type": "result", "subtype": "success", "is_error": False,
"modelUsage": {model + "[1m]": limits}, "usage": {},
"structured_output": {"groups": []}}
if failure == "substitution":
limits["canonicalModel"] = "claude-opus-4-8"
elif failure == "capacity":
limits["contextWindow"] = 200000
elif failure == "plugin":
init["plugins"] = [{"name": "unapproved"}]
raw = "\n".join(map(json.dumps, (init, result)))
if failure:
with pytest.raises(Problem):
suite_backends.parse_claude(raw, model, exit_code=0)
else:
_, metadata = suite_backends.parse_claude(raw, model, exit_code=0)
assert metadata["model"] == model
assert metadata["model_usage"][model + "[1m]"]["maxOutputTokens"] == 128000
def test_unknown_model_cannot_start_a_cli_process():
"""Operator configuration and command construction both reject unknown IDs."""
with pytest.raises(Problem):
suite_backends.claude_command("unapproved-model", 30)
result = subprocess.run([sys.executable, "-c", "import suite_contract"],
cwd=SCRIPTS, env={"PLANNING_CLAUDE_MODEL": "unapproved-model"},
capture_output=True, text=True)
assert result.returncode != 0
assert "Unsupported configured Claude model" in result.stderr
def test_thinking_usage_is_counted_without_retaining_unrecognized_fields():
"""New CLI usage details retain measurements and drop arbitrary content."""
assert usage_counts({"output_tokens_details": {"thinking_tokens": 123, "content": "private"}}) == {
"output_tokens_details": {"thinking_tokens": 123}}