hermes: upgrade suite planner for Claude 5.5 models
This commit is contained in:
parent
3036064bf7
commit
206c61993d
2120
docs/evidence/hermes_suite_claude55_20260929.json
Normal file
2120
docs/evidence/hermes_suite_claude55_20260929.json
Normal file
File diff suppressed because it is too large
Load Diff
110
docs/hermes_suite_claude55.md
Normal file
110
docs/hermes_suite_claude55.md
Normal file
@ -0,0 +1,110 @@
|
||||
# Claude 5.5 suite comparison
|
||||
|
||||
The suite planner now pins native Claude Code 2.1.285. Opus 5.5 and Sonnet 5.5
|
||||
both returned their exact canonical model identities on the existing first-party
|
||||
OAuth account. The default deployment selects `claude-opus-5-5`; Sonnet is a
|
||||
server configuration option, not an automatic fallback. The HTTPS contract,
|
||||
credential scopes, implementation objective, medium effort, five-case task cap,
|
||||
30-minute deadline, and USD 30 CLI estimated-cost guard are unchanged.
|
||||
|
||||
## Matched 14-case results
|
||||
|
||||
Each run used the same 9,851-byte normalized synthetic request, complete source
|
||||
fields, schema, and three-pass workflow (two independent proposals and one
|
||||
whole-suite reconciliation). Each completed in six CLI turns without a retry.
|
||||
No family exceeded five, so these hosted tests did not need the two extra
|
||||
oversized-family review passes.
|
||||
|
||||
| Measurement | Opus 5.5 | Sonnet 5.5 |
|
||||
| --- | ---: | ---: |
|
||||
| Wall seconds | 39.256 | 24.799 |
|
||||
| Aggregate input tokens | 21,936 | 21,438 |
|
||||
| Aggregate output tokens | 4,243 | 3,814 |
|
||||
| CLI estimated USD | 0.172604 | 0.081016 |
|
||||
| Natural families / final tasks | 9 / 9 | 9 / 9 |
|
||||
| Singleton families | 4 | 4 |
|
||||
| Exact alias coverage | Yes | Yes |
|
||||
| Incorrect merge / missed merge pairs | 0 / 0 | 0 / 0 |
|
||||
|
||||
Both produced identical memberships matching the fixture's independent mechanism
|
||||
labels, including distant related cases, identical text under different aliases,
|
||||
and watchdog terminology shared by different test machinery. Review found no
|
||||
description importing another family's objectives. Opus explicitly retained
|
||||
unknown procedure details for the thermal, acoustic, and build singleton fixtures.
|
||||
Sonnet's descriptions were more concise and omitted those uncertainty notes.
|
||||
|
||||
Membership quality tied on this small fixture. Sonnet used about 53% less CLI
|
||||
estimated cost and finished about 37% sooner. This is one observation per model,
|
||||
not evidence that their performance or quality is equivalent on a difficult
|
||||
363-case suite. The amounts are CLI accounting estimates, not subscription
|
||||
charges or proof of a dollar-denominated account bill.
|
||||
|
||||
## Runtime limits and isolation
|
||||
|
||||
Both CLI result envelopes report a 1,000,000-token context and 128,000-token model
|
||||
output ceiling. The service deliberately continues to request at most 64,000
|
||||
output tokens per invocation. `selection.output` is that service limit;
|
||||
`selection.reported_output` and `model_usage[*].maxOutputTokens` describe the
|
||||
reported model ceiling. The native CLI loopback check captured `max_tokens=64000`
|
||||
and complete source/system text on all four requests, including one deliberately
|
||||
invalid structured assignment followed by repair.
|
||||
|
||||
The newer CLI loads a built-in instruction plugin even in safe mode unless it is
|
||||
explicitly disabled. The service now disables it, all hooks, bundled skills,
|
||||
provider connectors, plugin synchronization, built-in subagents, and git
|
||||
instructions. Each CLI invocation allowlists only its pinned model and turns off
|
||||
automatic model switching when a request is flagged. Runtime identity, empty
|
||||
plugins/MCP lists, and tool isolation are still checked before accepting output.
|
||||
No provider or model fallback was added. The first tiny Opus readiness probe
|
||||
failed with zero usage and insufficient retained diagnostics; a fresh probe
|
||||
succeeded. Neither measured suite job needed a retry.
|
||||
|
||||
The artifact is downloaded from the official release host at startup with a
|
||||
pinned version, 240,327,864-byte size, and SHA-256:
|
||||
|
||||
```text
|
||||
33dad1ec615a2e08cc78b494f05c110e49916de2c79d78ec8799ebf46b233d29
|
||||
```
|
||||
|
||||
This removes the planner's dependency on the shared agent tools volume. Other
|
||||
Hermes CLIs and the isolated local-model endpoint are unchanged. A cold start now
|
||||
requires the official release download host to be reachable; failed download or
|
||||
checksum verification leaves the worker unavailable instead of using another CLI.
|
||||
|
||||
## Configuration and rollback
|
||||
|
||||
The HTTP compatibility revision remains `suite-v6-20260929`, the prompt remains
|
||||
`implementation-proximity-multipass-v4-20260929`, and the execution revision is
|
||||
`suite-multipass-v5-20260929`. Authenticated capabilities additionally expose
|
||||
`claude_model_options` and `model_selection: server_configuration`. No new client
|
||||
request field is accepted or required.
|
||||
|
||||
To select Sonnet, change the existing Flux manifest environment entry and its
|
||||
rollout annotation, commit, push, and reconcile Hermes while no job is active:
|
||||
|
||||
```yaml
|
||||
- name: PLANNING_CLAUDE_MODEL
|
||||
value: claude-sonnet-5-5
|
||||
```
|
||||
|
||||
The allowlisted IDs are `claude-opus-4-8`, `claude-opus-5-5`, and
|
||||
`claude-sonnet-5-5`. Reverting the upgrade commit restores the former CLI staging,
|
||||
model pin, and execution revision. Reconcile with:
|
||||
|
||||
```bash
|
||||
flux reconcile kustomization hermes --namespace flux-system --with-source
|
||||
```
|
||||
|
||||
Do not restart during a job: provider work is never automatically retried, and
|
||||
in-memory results expire on worker restart. The server comparison runs used the
|
||||
actual multi-pass backend inside isolated temporary directories, not laptop
|
||||
connectivity or HTTP submission. Results and concise review are retained in
|
||||
[synthetic evidence](evidence/hermes_suite_claude55_20260929.json). The operator
|
||||
probe accepts no input file and constructs only the fixed synthetic 14-case
|
||||
fixture. Focused mocked regression tests passed: 122. Kustomize render and client
|
||||
dry-run passed; Flux diff was limited to the planner Deployment and ConfigMap.
|
||||
The documented combined server-side/client dry-run flags are incompatible with
|
||||
the installed kubectl, so the client dry-run was run without `--server-side`.
|
||||
|
||||
Official references: [model configuration](https://code.claude.com/docs/en/model-config)
|
||||
and [CLI settings](https://code.claude.com/docs/en/settings-reference).
|
||||
@ -11,7 +11,7 @@ most **64 characters**, unique after whitespace and case normalization.
|
||||
- Configuration: `suite-v6-20260929` (HTTP compatibility identifier).
|
||||
- Policy: `implementation-five-v1-20260929`.
|
||||
- Prompt: `implementation-proximity-multipass-v4-20260929`.
|
||||
- Execution: `suite-multipass-v4-20260929`.
|
||||
- Execution: `suite-multipass-v5-20260929`.
|
||||
|
||||
A single server-side job performs:
|
||||
|
||||
@ -45,13 +45,19 @@ missing review decisions still fail closed; no missing assignment is fabricated.
|
||||
|
||||
Each invocation retains six CLI turns for structured output. Model review passes
|
||||
and CLI turns are separate counters. All calls use the originally selected provider
|
||||
and pinned model. The model remains `claude-opus-4-8[1m]`, with canonical runtime
|
||||
identity checked as `claude-opus-4-8`, firstParty, medium effort, reported 1M context
|
||||
and 64K output ceiling. The server invokes native Claude Code CLI 2.1.226 using
|
||||
and pinned model. The current default is `claude-opus-5-5[1m]`, with canonical runtime
|
||||
identity checked as `claude-opus-5-5`, firstParty, medium effort, reported 1M context
|
||||
and 128K model output ceiling, with requests limited to 64K output tokens.
|
||||
The server invokes native Claude Code CLI 2.1.285 using
|
||||
the existing first-party OAuth account, not a separately configured API-key
|
||||
account. The CLI itself communicates with Anthropic over HTTPS. No new provider
|
||||
fallback or tools are enabled.
|
||||
|
||||
Sonnet 5.5 is also verified on the 14-case synthetic fixture and available through
|
||||
server configuration. See the [matched comparison](hermes_suite_claude55.md).
|
||||
Earlier acceptance measurements below used Opus 4.8; they do not establish large
|
||||
suite capacity or quality for the new models.
|
||||
|
||||
## Deterministic task sizing and names
|
||||
|
||||
For each remaining coherent large family, `k = ceil(n / 5)` and
|
||||
|
||||
65
scripts/ops/hermes_suite_model_compare.py
Normal file
65
scripts/ops/hermes_suite_model_compare.py
Normal file
@ -0,0 +1,65 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Compare one configured Claude model on the public synthetic 14-case suite.
|
||||
|
||||
Run server-side with the planner modules and scoped OAuth mount. This invokes
|
||||
the same multi-pass backend without creating an HTTP job. Output is synthetic
|
||||
evidence only; progress contains metadata. Never accepts a roster input file.
|
||||
"""
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import threading
|
||||
import time
|
||||
|
||||
sys.path.insert(0, os.environ.get("SUITE_PROBE_MODULE_DIR", "/opt/planner"))
|
||||
import suite_backends
|
||||
from suite_contract import (EXECUTION_REVISION, MODELS, PROMPT_REVISION, Problem,
|
||||
encoded, validate_request, validate_result)
|
||||
from suite_multipass import generate, preflight_workflow
|
||||
from suite_synthetic import fixture, score
|
||||
|
||||
|
||||
def main():
|
||||
"""Run one fixed fixture and emit normalized results plus measured metadata."""
|
||||
model = MODELS["claude"]["model"]
|
||||
if model not in {"claude-opus-5-5", "claude-sonnet-5-5"}:
|
||||
raise SystemExit("Configure one of the two approved comparison models")
|
||||
if os.environ.get("SUITE_PROBE_CLAUDE_BIN"):
|
||||
suite_backends.CLAUDE_BIN = os.environ["SUITE_PROBE_CLAUDE_BIN"]
|
||||
request, expected = fixture(14)
|
||||
request["routing"] = {"allow_external": True, "allowed_external_providers": ["claude"]}
|
||||
request = validate_request(request, ["claude"])
|
||||
selected = preflight_workflow(request)
|
||||
started, last_report = time.monotonic(), 0.0
|
||||
last_stage = None
|
||||
|
||||
def progress(value):
|
||||
"""Emit stage and elapsed time without model-generated content."""
|
||||
nonlocal last_report, last_stage
|
||||
stage = value["current_pass"]
|
||||
if stage != last_stage or time.monotonic() - last_report >= 20:
|
||||
print(json.dumps({"model": model, "stage": stage,
|
||||
"elapsed_seconds": value["job_elapsed_seconds"],
|
||||
"completed_passes": value["completed_model_passes"]}),
|
||||
file=sys.stderr, flush=True)
|
||||
last_report, last_stage = time.monotonic(), stage
|
||||
|
||||
report = {"model": model, "execution_revision": EXECUTION_REVISION,
|
||||
"prompt_revision": PROMPT_REVISION, "selection": selected,
|
||||
"request_bytes": len(encoded(request)), "case_count": 14,
|
||||
"automatic_retries": 0, "test_path": "server_side_multipass_backend"}
|
||||
try:
|
||||
result, metadata = generate(request, selected, threading.Event(), "192.168.22.8", progress)
|
||||
validate_result(result, request)
|
||||
report.update(status="completed", result=result, metadata=metadata,
|
||||
quality=score(result, expected),
|
||||
singletons=sum(len(g["members"]) == 1 for g in result["groups"]))
|
||||
except Problem as exc:
|
||||
report.update(status="failed", **exc.document())
|
||||
report["wall_seconds"] = round(time.monotonic() - started, 3)
|
||||
print(json.dumps(report, sort_keys=True), flush=True)
|
||||
return 0 if report["status"] == "completed" else 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@ -12,7 +12,7 @@ import re
|
||||
import subprocess
|
||||
import threading
|
||||
|
||||
from suite_contract import (COST_LIMIT, EXECUTION_REVISION, MAX_BODY, MAX_CASES, MAX_RESULT, MODELS, PROMPT_REVISION,
|
||||
from suite_contract import (CLAUDE_MODELS, CLAUDE_VERSION, COST_LIMIT, EXECUTION_REVISION, MAX_BODY, MAX_CASES, MAX_RESULT, MODELS, PROMPT_REVISION,
|
||||
PROMPT_SHA256, REVISION, TIMEOUT, Problem, encoded,
|
||||
preflight, validate_request)
|
||||
from suite_jobs import Jobs
|
||||
@ -124,6 +124,7 @@ class Handler(BaseHTTPRequestHandler):
|
||||
"prompt_revision": PROMPT_REVISION, "prompt_sha256": PROMPT_SHA256})
|
||||
if method == "GET" and self.path == "/v1/capabilities":
|
||||
return self.send(200, {"configuration_revision": REVISION, "models": MODELS,
|
||||
"claude_model_options": list(CLAUDE_MODELS), "model_selection": "server_configuration",
|
||||
"execution_revision": EXECUTION_REVISION,
|
||||
"policy_revision": POLICY_REVISION, "max_final_group_cases": 5,
|
||||
"max_final_name_characters": 64, "model_passes": {"minimum": 3, "maximum": 5},
|
||||
@ -214,7 +215,7 @@ def main():
|
||||
raise SystemExit("Pinned CLI binary mismatch")
|
||||
version = subprocess.run(["/opt/cli/claude", "--version"], capture_output=True,
|
||||
text=True, timeout=10, check=True)
|
||||
if version.stdout.strip() != "2.1.226 (Claude Code)":
|
||||
if version.stdout.strip() != CLAUDE_VERSION + " (Claude Code)":
|
||||
raise SystemExit("Pinned CLI version mismatch")
|
||||
private = ThreadingHTTPServer(("0.0.0.0", 9001), Decision)
|
||||
threading.Thread(target=private.serve_forever, daemon=True).start()
|
||||
|
||||
@ -11,7 +11,7 @@ import time
|
||||
from urllib.error import HTTPError, URLError
|
||||
from urllib.request import HTTPRedirectHandler, ProxyHandler, Request, build_opener
|
||||
|
||||
from suite_contract import CLAUDE_MAX_TURNS, MODELS, SCHEMA, SYSTEM, Problem, encoded, prompt
|
||||
from suite_contract import CLAUDE_MAX_TURNS, CLAUDE_MODELS, MODELS, SCHEMA, SYSTEM, Problem, encoded, prompt
|
||||
from suite_cli_diagnostics import snapshot
|
||||
|
||||
SWITCHYARD = "http://hermes-switchyard.hermes.svc.cluster.local:9005/v1/chat/completions"
|
||||
@ -97,12 +97,19 @@ def local_generate(request, cancel, client_ip, *, invocation=None):
|
||||
|
||||
def claude_command(model, max_cost):
|
||||
"""Use the native pinned binary, never the privileged Hermes shell wrapper."""
|
||||
if model not in CLAUDE_MODELS:
|
||||
raise Problem("unsupported_backend", 422)
|
||||
settings = {"enabledPlugins": {"agents-md@builtin": False,
|
||||
"cc-plugin-agents-md@builtin": False},
|
||||
"disableAllHooks": True, "disableBundledSkills": True,
|
||||
"disableClaudeAiConnectors": True, "syncClaudeAiPlugins": False,
|
||||
"availableModels": [model], "switchModelsOnFlag": False}
|
||||
return [CLAUDE_BIN, "-p", "--output-format", "stream-json", "--verbose",
|
||||
"--no-session-persistence", "--safe-mode", "--tools", "",
|
||||
"--strict-mcp-config", "--mcp-config", '{"mcpServers":{}}',
|
||||
"--setting-sources", "", "--disable-slash-commands",
|
||||
"--setting-sources", "", "--settings", encoded(settings).decode(), "--disable-slash-commands",
|
||||
"--permission-mode", "dontAsk", "--no-chrome",
|
||||
"--model", MODELS["claude"]["cli_model"], "--effort", "medium", "--max-budget-usd", str(max_cost),
|
||||
"--model", model + "[1m]", "--effort", "medium", "--max-budget-usd", str(max_cost),
|
||||
"--max-turns", str(CLAUDE_MAX_TURNS), "--system-prompt", SYSTEM,
|
||||
"--json-schema", encoded(SCHEMA).decode()]
|
||||
|
||||
@ -117,7 +124,9 @@ def claude_environment(directory, token):
|
||||
"DISABLE_ERROR_REPORTING", "DISABLE_AUTOUPDATER", "DISABLE_UPDATES",
|
||||
"DISABLE_PROMPT_CACHING", "CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC",
|
||||
"CLAUDE_CODE_DISABLE_BACKGROUND_TASKS", "CLAUDE_CODE_DISABLE_TERMINAL_TITLE",
|
||||
"CLAUDE_CODE_DISABLE_AUTO_MEMORY", "CLAUDE_CODE_SKIP_PROMPT_HISTORY"):
|
||||
"CLAUDE_CODE_DISABLE_AUTO_MEMORY", "CLAUDE_CODE_SKIP_PROMPT_HISTORY",
|
||||
"CLAUDE_AGENT_SDK_DISABLE_BUILTIN_AGENTS", "CLAUDE_CODE_DISABLE_GIT_INSTRUCTIONS",
|
||||
"CLAUDE_CODE_DISABLE_BUNDLED_SKILLS"):
|
||||
env[key] = "1"
|
||||
return env
|
||||
|
||||
@ -180,7 +189,8 @@ def parse_claude(raw, expected_model, **process_info):
|
||||
if not isinstance(models, dict) or not models or set(models) - aliases:
|
||||
fail("model_changed", "runtime_model")
|
||||
for limits in models.values():
|
||||
if (not isinstance(limits, dict) or limits.get("contextWindow") != 1000000 or limits.get("maxOutputTokens") != 64000
|
||||
if (not isinstance(limits, dict) or limits.get("contextWindow") != 1000000
|
||||
or limits.get("maxOutputTokens") != CLAUDE_MODELS.get(expected_model)
|
||||
or limits.get("canonicalModel") != expected_model
|
||||
or limits.get("provider") != "firstParty"):
|
||||
fail("backend_capabilities_changed", "runtime_model_limits")
|
||||
@ -196,7 +206,7 @@ def parse_claude(raw, expected_model, **process_info):
|
||||
fail("incomplete_generation", "process_exit")
|
||||
# Persist measured metadata only, including on later schema/coverage failure.
|
||||
safe_models = {name: {"canonicalModel": expected_model, "provider": "firstParty",
|
||||
"contextWindow": 1000000, "maxOutputTokens": 64000}
|
||||
"contextWindow": 1000000, "maxOutputTokens": CLAUDE_MODELS[expected_model]}
|
||||
for name in models}
|
||||
return result, {"model": expected_model, "compaction": False,
|
||||
"truncation": False, "compaction_signal": "CLI events and disabled compaction",
|
||||
|
||||
@ -30,6 +30,7 @@ def usage_counts(value):
|
||||
"input_tokens", "output_tokens", "cache_read_input_tokens", "cache_creation_input_tokens")
|
||||
if key in value}
|
||||
for key, fields in (("server_tool_use", ("web_search_requests", "web_fetch_requests")),
|
||||
("output_tokens_details", ("thinking_tokens",)),
|
||||
("cache_creation", ("ephemeral_1h_input_tokens", "ephemeral_5m_input_tokens"))):
|
||||
if isinstance(value.get(key), dict):
|
||||
result[key] = {field: number(value[key][field]) for field in fields if field in value[key]}
|
||||
|
||||
@ -3,12 +3,22 @@ from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
from collections import Counter
|
||||
|
||||
REVISION = "suite-v6-20260929"
|
||||
PROMPT_REVISION = "implementation-proximity-multipass-v4-20260929"
|
||||
EXECUTION_REVISION = "suite-multipass-v4-20260929"
|
||||
EXECUTION_REVISION = "suite-multipass-v5-20260929"
|
||||
CLAUDE_VERSION = "2.1.285"
|
||||
CLAUDE_MODELS = {
|
||||
"claude-opus-4-8": 64000,
|
||||
"claude-opus-5-5": 128000,
|
||||
"claude-sonnet-5-5": 128000,
|
||||
}
|
||||
CLAUDE_MODEL = os.environ.get("PLANNING_CLAUDE_MODEL", "claude-opus-4-8")
|
||||
if CLAUDE_MODEL not in CLAUDE_MODELS:
|
||||
raise RuntimeError("Unsupported configured Claude model")
|
||||
CLAUDE_MAX_TURNS = 6
|
||||
MAX_BODY = 1 << 20
|
||||
MAX_RESULT = 1 << 20
|
||||
@ -23,9 +33,10 @@ MODELS = {
|
||||
"local": {"model": "qwen2.5:14b-instruct-q4_0", "context": 8192,
|
||||
"output": 2048, "overhead": 1024, "backend": "ollama-model-gate",
|
||||
"enabled": True, "reasoning": "none"},
|
||||
"claude": {"model": "claude-opus-4-8", "context": 1000000,
|
||||
"cli_model": "claude-opus-4-8[1m]",
|
||||
"output": 64000, "overhead": 8192, "backend": "claude-code-2.1.226",
|
||||
"claude": {"model": CLAUDE_MODEL, "context": 1000000,
|
||||
"cli_model": CLAUDE_MODEL + "[1m]",
|
||||
"output": 64000, "reported_output": CLAUDE_MODELS[CLAUDE_MODEL],
|
||||
"overhead": 8192, "backend": "claude-code-" + CLAUDE_VERSION,
|
||||
"enabled": True, "reasoning": "medium", "max_turns": CLAUDE_MAX_TURNS},
|
||||
"codex": {"model": "gpt-6-astra", "context": 258400,
|
||||
"output": None, "overhead": None, "backend": "codex-subscription-broker",
|
||||
|
||||
@ -36,7 +36,7 @@ spec:
|
||||
app: hermes-suite-planner
|
||||
annotations:
|
||||
fluentbit.io/exclude: "true"
|
||||
ai.bstein.dev/config-rev: suite-v6-multipass-cap5-v4-20260929
|
||||
ai.bstein.dev/config-rev: suite-v6-multipass-cap5-v5-20260929
|
||||
vault.hashicorp.com/agent-inject: "true"
|
||||
vault.hashicorp.com/agent-pre-populate-only: "true"
|
||||
vault.hashicorp.com/agent-init-first: "true"
|
||||
@ -80,7 +80,7 @@ spec:
|
||||
automountServiceAccountToken: false
|
||||
enableServiceLinks: false
|
||||
terminationGracePeriodSeconds: 15
|
||||
# The installed amd64 CLI is copied from the existing RWO tools volume.
|
||||
# The pinned native CLI is amd64; retain the existing worker placement.
|
||||
nodeSelector:
|
||||
kubernetes.io/hostname: titan-22
|
||||
securityContext:
|
||||
@ -97,18 +97,27 @@ spec:
|
||||
command: [python, -c]
|
||||
args:
|
||||
- |
|
||||
import hashlib,pathlib,shutil
|
||||
source=pathlib.Path('/installed/lib/node_modules/@anthropic-ai/claude-code/bin/claude.exe')
|
||||
assert hashlib.sha256(source.read_bytes()).hexdigest() == '4e9bec1177ce9690e8bd988b710ac24105e70da428dd094c5adcbbe786a55555'
|
||||
shutil.copyfile(source, '/opt/cli/claude')
|
||||
pathlib.Path('/opt/cli/claude').chmod(0o555)
|
||||
import hashlib,pathlib,urllib.request
|
||||
target=pathlib.Path('/opt/cli/claude')
|
||||
url='https://downloads.claude.ai/claude-code-releases/2.1.285/linux-x64/claude'
|
||||
digest=hashlib.sha256()
|
||||
size=0
|
||||
with urllib.request.urlopen(url, timeout=120) as source, target.open('wb') as output:
|
||||
while data:=source.read(1048576):
|
||||
size+=len(data)
|
||||
if size > 240327864:
|
||||
raise RuntimeError('CLI artifact exceeds pinned size')
|
||||
digest.update(data)
|
||||
output.write(data)
|
||||
if size != 240327864 or digest.hexdigest() != '33dad1ec615a2e08cc78b494f05c110e49916de2c79d78ec8799ebf46b233d29':
|
||||
raise RuntimeError('CLI artifact verification failed')
|
||||
target.chmod(0o555)
|
||||
securityContext:
|
||||
allowPrivilegeEscalation: false
|
||||
readOnlyRootFilesystem: true
|
||||
capabilities:
|
||||
drop: [ALL]
|
||||
volumeMounts:
|
||||
- {name: installed, mountPath: /installed, subPath: tools, readOnly: true}
|
||||
- {name: cli, mountPath: /opt/cli}
|
||||
resources:
|
||||
requests: {cpu: 25m, memory: 256Mi}
|
||||
@ -122,7 +131,8 @@ spec:
|
||||
- {name: PYTHONUNBUFFERED, value: "1"}
|
||||
# User approved generalized CASE records on this Claude account, 2026-09-29.
|
||||
- {name: PLANNING_GENERALIZED_CLAUDE_APPROVED, value: "true"}
|
||||
- {name: PLANNING_CLAUDE_SHA256, value: 4e9bec1177ce9690e8bd988b710ac24105e70da428dd094c5adcbbe786a55555}
|
||||
- {name: PLANNING_CLAUDE_SHA256, value: 33dad1ec615a2e08cc78b494f05c110e49916de2c79d78ec8799ebf46b233d29}
|
||||
- {name: PLANNING_CLAUDE_MODEL, value: claude-opus-5-5}
|
||||
ports:
|
||||
- {name: http, containerPort: 9000}
|
||||
- {name: decision, containerPort: 9001}
|
||||
@ -148,10 +158,6 @@ spec:
|
||||
- name: scripts
|
||||
configMap:
|
||||
name: hermes-suite-planner
|
||||
- name: installed
|
||||
persistentVolumeClaim:
|
||||
claimName: hermes-agent-home
|
||||
readOnly: true
|
||||
- name: cli
|
||||
emptyDir: {medium: Memory, sizeLimit: 512Mi}
|
||||
- name: jobs
|
||||
|
||||
74
testing/tests/test_suite_claude_models.py
Normal file
74
testing/tests/test_suite_claude_models.py
Normal file
@ -0,0 +1,74 @@
|
||||
"""Model upgrades preserve exact identity, capacity, and CLI isolation."""
|
||||
import json
|
||||
from pathlib import Path
|
||||
import subprocess
|
||||
import sys
|
||||
|
||||
import pytest
|
||||
|
||||
SCRIPTS = Path(__file__).resolve().parents[2] / "services/hermes/scripts"
|
||||
sys.path.insert(0, str(SCRIPTS))
|
||||
import suite_backends
|
||||
from suite_cli_diagnostics import usage_counts
|
||||
from suite_contract import CLAUDE_MODELS, Problem
|
||||
|
||||
|
||||
@pytest.mark.parametrize("model", CLAUDE_MODELS)
|
||||
def test_only_configured_model_is_allowed_and_plugins_are_disabled(model):
|
||||
"""Each invocation pins one model, with fallback and customizations disabled."""
|
||||
command = suite_backends.claude_command(model, 30)
|
||||
settings = json.loads(command[command.index("--settings") + 1])
|
||||
assert command[command.index("--model") + 1] == model + "[1m]"
|
||||
assert settings["availableModels"] == [model]
|
||||
assert settings["switchModelsOnFlag"] is False
|
||||
assert settings["disableAllHooks"] and settings["disableClaudeAiConnectors"]
|
||||
assert settings["syncClaudeAiPlugins"] is False
|
||||
assert settings["enabledPlugins"]["cc-plugin-agents-md@builtin"] is False
|
||||
assert "--fallback-model" not in command
|
||||
env = suite_backends.claude_environment("/jobs/fresh", "synthetic")
|
||||
assert env["CLAUDE_CODE_MAX_OUTPUT_TOKENS"] == "64000"
|
||||
assert env["CLAUDE_AGENT_SDK_DISABLE_BUILTIN_AGENTS"] == "1"
|
||||
|
||||
|
||||
@pytest.mark.parametrize("model", ["claude-opus-5-5", "claude-sonnet-5-5"])
|
||||
@pytest.mark.parametrize("failure", [None, "substitution", "capacity", "plugin"])
|
||||
def test_new_runtime_envelope_remains_fail_closed(model, failure):
|
||||
"""A requested name never substitutes for the actual runtime model identity."""
|
||||
init = {"type": "system", "subtype": "init", "model": model + "[1m]",
|
||||
"tools": ["StructuredOutput"], "mcp_servers": [], "plugins": []}
|
||||
limits = {"contextWindow": 1000000, "maxOutputTokens": 128000,
|
||||
"canonicalModel": model, "provider": "firstParty"}
|
||||
result = {"type": "result", "subtype": "success", "is_error": False,
|
||||
"modelUsage": {model + "[1m]": limits}, "usage": {},
|
||||
"structured_output": {"groups": []}}
|
||||
if failure == "substitution":
|
||||
limits["canonicalModel"] = "claude-opus-4-8"
|
||||
elif failure == "capacity":
|
||||
limits["contextWindow"] = 200000
|
||||
elif failure == "plugin":
|
||||
init["plugins"] = [{"name": "unapproved"}]
|
||||
raw = "\n".join(map(json.dumps, (init, result)))
|
||||
if failure:
|
||||
with pytest.raises(Problem):
|
||||
suite_backends.parse_claude(raw, model, exit_code=0)
|
||||
else:
|
||||
_, metadata = suite_backends.parse_claude(raw, model, exit_code=0)
|
||||
assert metadata["model"] == model
|
||||
assert metadata["model_usage"][model + "[1m]"]["maxOutputTokens"] == 128000
|
||||
|
||||
|
||||
def test_unknown_model_cannot_start_a_cli_process():
|
||||
"""Operator configuration and command construction both reject unknown IDs."""
|
||||
with pytest.raises(Problem):
|
||||
suite_backends.claude_command("unapproved-model", 30)
|
||||
result = subprocess.run([sys.executable, "-c", "import suite_contract"],
|
||||
cwd=SCRIPTS, env={"PLANNING_CLAUDE_MODEL": "unapproved-model"},
|
||||
capture_output=True, text=True)
|
||||
assert result.returncode != 0
|
||||
assert "Unsupported configured Claude model" in result.stderr
|
||||
|
||||
|
||||
def test_thinking_usage_is_counted_without_retaining_unrecognized_fields():
|
||||
"""New CLI usage details retain measurements and drop arbitrary content."""
|
||||
assert usage_counts({"output_tokens_details": {"thinking_tokens": 123, "content": "private"}}) == {
|
||||
"output_tokens_details": {"thinking_tokens": 123}}
|
||||
Loading…
x
Reference in New Issue
Block a user