# Suite planning API This optional API groups one complete campaign/suite into implementation families. The existing `/local-model` API, GPU allocation, serving model, and limits are unchanged. Deployment uses Flux; the application and workbook stay on the laptop. ## Current multi-pass policy New jobs use [the multi-pass implementation policy](hermes_suite_multipass.md): two independent full-suite proposals, reconciliation, up to two bounded semantic reviews, and a programmatic five-case final cap. Names are at most 64 characters. The endpoint and request fields are unchanged. Optional top-level `review_summary` is returned only with the authorized result. Configuration remains `suite-v6-20260929`; policy, prompt, and execution revisions identify the changed grouping behavior. The prior single-pass acceptance results below are historical and do not measure the new workflow. The new process shares the user-approved 1800-second deadline and USD 30 CLI estimate guard across all invocations. This is not subscription billing. Progress updates every five seconds during a CLI pass. The separate local-model endpoint is unchanged; the 8K local route cannot admit the new multi-pass reconciliation schema. ## Earlier CLI failure diagnostics update (historical) Execution revision `claude-diagnostics-turns-v1-20260929` adds content-free `cli_diagnostics` and failure-stage details, including exit/signal, final subtype, reported turns, observed stop reason, token counts, and structured-output location. Unknown measurements remain null. Source, generated descriptions, raw provider errors, and transcripts are not retained. The normal schema and exact membership checks still gate completion. Failed validation now preserves operational usage. The CLI turn ceiling is six, with the same model, medium effort, 64,000 per-call output ceiling, maximum 900-second job deadline, and USD 5 CLI estimated-cost cap. Preflight reserves all six possible outputs against context. This is headroom within one fresh job, not an automatic job retry or a fallback. The historical three-turn ceiling was exhausted by a deterministic synthetic structured-output repair test; six turns completed that same test. This does not establish the cause of a historical failure whose terminal event was deleted. The HTTP configuration revision and implementation-proximity prompt revision/hash are unchanged. `execution_revision` records the new behavior in capability, preflight, and job metadata. Use a new idempotency key only when intentionally starting a new attempt; the failed historical job is never automatically rerun. See [the diagnosis and verification record](hermes_suite_failure_diagnosis_20260929.md). Rollback this update by reverting its Git commit and reconciling the Hermes Flux Kustomization; do not revert the earlier prompt commit. Wait for no active jobs before rollout because results are held in memory. ## Implementation-proximity prompt update New jobs use prompt revision `implementation-proximity-v2-20260929`. The HTTP configuration revision remains `suite-v6-20260929`; request and result schemas, models, reasoning effort, token limits, authentication, and routing policy are unchanged. `prompt_revision` and `prompt_sha256` are additive provenance in health, capabilities, preflight selection, and persisted job/status/result metadata. Historical jobs keep their original metadata; they are not relabeled on replay. The objective is the additional test-development effort after implementing a representative case. Success criteria provide the primary evidence of actions, observations, measurements, and assertions. Description supplies behavior and operation; preconditions and other material constraints supply setup/state; case type is supporting context. There is no keyword or word-count weighting. The prompt distinguishes inexpensive input/assertion variations from new test machinery, checks entire-family coherence, reviews singletons, preserves uncertainty in generalized values, and confines names/descriptions to assigned test objectives. The same system prompt is supplied to the actual grouping invocation and remains applicable during structured-output turns or format repair. There is no separate repair model. Cases enter a fresh isolated session without any previous grouping as a target answer. The expanded prompt is included in the existing conservative preflight accounting; no context or output setting has been increased. For a fresh comparison, keep the same case request but use a NEW `Idempotency-Key`. Reusing the old key returns the old job instead of invoking the revised prompt. Include `prompt_revision` or `prompt_sha256` in the laptop's result-cache identity. The synthetic check fixture is `testing/fixtures/suite_implementation_proximity.json`; only its `request` object is sent, never its independent expected-family mapping. No real roster job is launched by this deployment or its acceptance checks. Verified on the deployed endpoint: the 14-case adversarial synthetic suite completed in 21.059 seconds with six implementation families, exact alias coverage, and zero incorrect or missed merge pairs against its independent expected map. Identical broad descriptions did not merge different machinery. Nominal/negative-result and extra-assertion variants shared their existing observation code; an underspecified case retained its uncertainty. Every observed native CLI request in the loopback structured-output transport check contained the complete revised system prompt. The focused unit suite passes 36 tests. These checks establish the synthetic behavior, not the quality of a real-suite rerun. See the [prompt-update acceptance evidence](evidence/hermes_suite_prompt_v2_20260929.json). ## Acceptance results, 2026-09-29 The final runs use `suite-v6-20260929`, deployed commit `299b9ed1`, including case-type fields. These supersede the preliminary runs without that field. All tests used synthetic content. HTTPS tests ran from `titan-jh`, source `192.168.22.8`, directly on its LAN interface to `192.168.22.50:443`, with hostname certificate verification and proxies disabled. The user's WSL route has not been tested; run the health command below from the laptop before uploading anything. | Cases | Actual request bytes | CLI input/output tokens | Client wall seconds | Families / singletons | Exact coverage | Pattern match | | --- | ---: | --- | ---: | --- | --- | --- | | 14 | 9,847 | 4,185 / 1,250 | 16.947 | 9 / 4 | Yes | Exact | | 75 | 55,984 | 19,145 / 2,336 | 25.755 | 9 / 3 | Yes | Exact | | 363 | 273,761 | 89,753 / 8,568 | 80.071 | 9 / 3 | Yes | Exact | Request sizes are the UTF-8 compact JSON bytes actually submitted, excluding HTTP headers. Token counts are CLI-reported aggregate usage across each fresh job's two structured-output turns; they are not an exact standalone input-token count. All three final jobs completed on their first attempt with no observed retries or rate-limit errors. Two CLI turns are part of producing the structured result, not automatic job retries. CLI cost estimates were USD 0.052175, 0.154125, and 0.662965; these are estimates, not verified subscription charges. Remaining account quota is not available through this endpoint and may be consumed by other account users. The actual terminal `modelUsage` reports `claude-opus-4-8[1m]`, with `canonicalModel=claude-opus-4-8`, `provider=firstParty`, context 1,000,000 and maximum output 64,000. The invocation requests `claude-opus-4-8[1m]`; these results also independently verify the returned model metadata, rather than trusting the request flag or the internal `atlas/planning/claude` route label. Provider physical hardware and inference region are unknown. This verifies the tested suite sizes, not the entire advertised million-token context. The 363-case fixture has descriptions of 48-245 bytes (median 234), preconditions 67-143 (median 136), success criteria 73-135 (median 132), and case types 7-15 bytes. Most cases have multi-sentence fields plus operating conditions, verification method, and a generic reference. Six machinery patterns are interleaved throughout the input. Watchdog terminology occurs in three different machinery patterns. A distant pair has identical complete case text but different aliases. Distinctive thermal, acoustic, and build cases occur at the beginning, middle, and end. Against the independent synthetic implementation map, all sizes had pair precision and recall 1.0, zero incorrect merge pairs, and zero missed merge pairs. Manual review of all 27 family descriptions found no objectives attributed to another family. Four singletons in the 14-case fixture and three in the larger fixtures are expected. These are clear synthetic patterns, not evidence of quality on the user's more ambiguous real suite. No real case quality is claimed. Evidence for complete input handling goes beyond IDs: - Unit checks compare every serialized case and field with the original fixture. - The deployed native CLI was also run against a loopback mock provider, with the exact production command and full 14/75/363 inputs. The captured requests contained the complete source string, byte for byte, including every field. The largest first provider request was 287,914 bytes. This transport check used no hosted model. - Live jobs returned the expected distinct beginning/middle/end patterns and reunited interleaved cases. Initialization allowed only `StructuredOutput`, with no MCP, plugins, or file-reading tools. Each job used fresh HOME/config and stdin. - No CLI compaction event or incomplete-output signal was observed. Compaction was disabled, and the gateway performs no summarization or selective file reading. Those observations support complete transmission and the tested grouping behavior. They cannot inspect provider internals or prove that a model attended equally to every word. The response's false compaction/truncation values describe gateway/CLI observations, not a provider-side attestation. An exact tokenizer is unavailable; preflight's byte bound and output reservation are documented below. The local-only two-case parser example also completed in 3.535 seconds on the RTX backend with exact coverage and no hosted destination. An earlier local result incorrectly appended descriptions to aliases and was rejected; the new adapter constrains the local output schema to supplied aliases. The existing local-model endpoint itself is unchanged. Verification also covered authentication failures, permission escalation, unknown providers, synthetic-versus-operational scope, local-only capacity rejection, management-path rejection, cross-client ownership, idempotent replay, oversized HTTP 413, and content-free routine worker logs. Mocked backend failures verified that no local failure launches Claude. No shared backend was disrupted. The focused unit suite passes 35 tests. Service renders, client dry-runs, and Flux diffs passed. The installed kubectl rejects the repository's combined `--server-side --dry-run=client` syntax; the supported client dry-run was used instead. Full evidence and successful normalized envelopes: - [Acceptance measurements and policy checks](evidence/hermes_suite_acceptance_20260929.json) - [Actual complete 14-case Claude response](evidence/hermes_suite_success_14.json) - [Actual complete two-case local response](evidence/hermes_suite_local_success.json) - [All final synthetic family results](evidence/hermes_suite_results_synthetic.json) Ready: synthetic client testing, direct complete-suite execution at all three tested sizes, and the separately approved Claude credential. Remaining client work: WSL connectivity, explicit client-envelope adaptation, and a controlled real pilot after reviewing this handoff. Codex capacity verification and provider/account retention or training controls remain unresolved. The user's account-specific approval is recorded; broader organizational approval or release classification has not been established by this infrastructure work. ## Connection and credentials ``` LAN_HOST: worker.bstein.dev LAN_IP: 192.168.22.50 PORT: 443 BASE_URL: https://worker.bstein.dev/suite-planning HEALTH_URL: https://worker.bstein.dev/suite-planning/healthz CAPABILITIES_URL: https://worker.bstein.dev/suite-planning/v1/capabilities PREFLIGHT_URL: https://worker.bstein.dev/suite-planning/v1/preflight SUBMIT_URL: https://worker.bstein.dev/suite-planning/v1/jobs AUTHENTICATION_HEADER: Authorization: Bearer VAULT_PATH: kv/atlas/hermes/suite-planning-api LOCAL_ONLY_FIELD: token SYNTHETIC_EXTERNAL_FIELD: synthetic_token APPROVED_OPERATIONAL_FIELD: operational_token INITIAL_CONCURRENCY: 1; busy submissions return 429; no waiting queue JOB_TIMEOUT: up to 1800 seconds HTTP_TIMEOUT: client 45 seconds; submit/status do not wait for inference TLS: existing worker.bstein.dev certificate; normal trusted CA verification ``` Retrieve the appropriate field privately from [Vault](https://vault.bstein.dev/ui/vault/secrets/kv/show/atlas/hermes/suite-planning-api). Neither token grants provider, Vault, Kubernetes, shell, or ClickUp access. The old local-model token remains at `kv/atlas/hermes/model-gate-lan-api`. Tokens are distinct; an old local-model token does not authenticate this API. The credentials have different server-enforced scopes: | Vault field | Local inference | Hosted inference | | --- | --- | --- | | `token` | Allowed within local capacity | Prohibited | | `synthetic_token` | Allowed within local capacity | Exact API fixtures only; full content hash checked | | `operational_token` | Allowed within local capacity | Generalized CASE records, Claude only, under the explicit account approval | `synthetic_token` has not been broadened. Changing a fixture or sending real generalized records with that token returns `external_data_not_approved` even with `allow_external=true`. The operational credential is separate, grants only `claude`, and requires the server's `PLANNING_GENERALIZED_CLAUDE_APPROVED=true`. The user explicitly approved the existing Hermes first-party Claude OAuth account for generalized CASE records on 2026-09-29. No real records were used in acceptance. Original identifiers, uncensored text, and alias mappings remain on the laptop. The server cannot determine whether arbitrary prose is adequately generalized; that remains the application's and organization's responsibility. Account-specific training and retention settings remain unverified. A request flag cannot grant credential permission or override data scope. ## Request contract Submit with `POST https://worker.bstein.dev/suite-planning/v1/jobs`. The [request JSON Schema](contracts/suite_planning_request.schema.json) describes the accepted wire object. The [result JSON Schema](contracts/suite_planning_result.schema.json) describes `result`, not the whole job envelope. Additional server checks enforce UTF-8 byte limits, unique aliases, ownership, permissions, and exact result coverage. ``` { "campaign": "SYNTHETIC", "suite": "PARSER", "cases": [ {"alias": "CASE-1", "description": "Parse a valid configuration and assert accepted fields."}, {"alias": "CASE-2", "description": "Parse an invalid configuration and assert diagnostic fields."} ], "routing": { "allow_external": false, "allowed_external_providers": [] }, "execution": { "strategy": "whole_suite", "max_seconds": 1800, "max_cost_usd": 30 } } ``` `campaign`, `suite`, and nonempty `cases` are required. Each case requires a unique `CASE-*` alias and a description. Optional case fields are `success_criteria`, `preconditions`, `operating_condition`, `case_type`, `verification_method`, `target`, `swci`, `verifies`, `functional_area`, `functional_group`, and `functional_group_name`. Values are strings or null. Optional per-case campaign/suite values must exactly match the job's ownership. Unknown fields are rejected. `[reference]` is ordinary source text, never an alias. Identical descriptions with different aliases remain distinct cases. Missing routing policy means local-only. External provider names are `claude` and `codex`. The requested list must be contained in the credential's permissions; unknown providers, malformed booleans, duplicate names, and permission escalation are rejected before routing. `allow_external=false` cannot include providers. `whole_suite` is the only strategy. No batching, summarization, or reconciliation is silently substituted. `max_seconds` is 10-1800. The Claude CLI guard is at most USD 30 in its estimated usage accounting; this is not a verified subscription billing ceiling and can overshoot within a single provider call. The gateway deterministically selects a fitting permitted backend, preferring local capacity. It asks the fixed Switchyard route to confirm that binding, then performs one inference attempt. Switchyard receives only `select`, never case text. Its new routes each have exactly one target, zero client retries, and no classifier: `atlas/planning/local`, `atlas/planning/claude`, `atlas/planning/codex`. These are decision routes; clients use the HTTPS job API to obtain inference results. Existing general Hermes routes retain their old behavior and are not this API. ## Job lifecycle and normalized response - `GET /healthz`: authenticated service readiness. - `GET /v1/capabilities`: permissions, models, configuration revision, and limits. - `POST /v1/preflight`: validate the identical proposed job without generation. - `POST /v1/jobs`: require `Idempotency-Key`, return 202 and a job ID. - `GET /v1/jobs/`: status and metadata. - `GET /v1/jobs//result`: status plus the normalized result when complete. - `DELETE /v1/jobs/`: request cancellation. - `GET /v1/synthetic/14`, `/75`, `/363`: immutable synthetic fixtures. Retry a submission with the SAME key and SAME body. It returns the original job, not a second provider attempt. A changed body with that key returns 409. Keys are scoped to the authenticated credential and retained seven days. The maximum is 10,000 metadata records; excess submissions fail rather than evicting live keys. No provider retry or fallback occurs after an error. An operator may intentionally start a new attempt with a new key after diagnosing the previous failure. Results remain in memory for at most one hour (cleanup every 30 seconds), with a 128-result bound. Durable local SQLite holds only operational metadata and hashes, not prompts or results. Restarts mark unfinished jobs `interrupted_no_retry` and never resubmit them. Previously completed results are lost on restart; retrieval returns 410. A cancellation holds the slot until the underlying attempt ends. Claude cancellation kills the isolated process group. Local generation cannot be cancelled at the existing API; its eventual result is discarded. Provider-side cancellation and exact billing after disconnection are not guaranteed. Example completed envelope (illustrative synthetic values): ``` { "job_id": "0123456789abcdef0123456789abcdef", "status": "completed", "configuration_revision": "suite-v6-20260929", "routing": {"allow_external": true, "allowed_external_providers": ["claude"]}, "selection": {"provider": "claude", "backend": "claude-code-2.1.226", "model": "claude-opus-4-8"}, "model": "claude-opus-4-8", "attempted_destinations": ["switchyard:atlas/planning/claude", "claude:claude-opus-4-8"], "compaction": false, "truncation": false, "usage": {"input_tokens": 1000, "output_tokens": 100}, "wall_seconds": 12.3, "result": { "groups": [{"name": "Configuration parser fixtures", "description": "Parameterize configuration text and assert parsed fields or diagnostics.", "members": ["CASE-1", "CASE-2"]}] } } ``` The real envelope also includes input bytes, input hash, reservation method, model usage, timing, and estimated cost where available. Unknown values are null. Failed jobs retain status and fixed error codes but no provider stderr. Status remains HTTP 200 for an existing failed job; inspect `status` and `error.code`. Invalid submissions use 400/401/403/409/413/415/422/429 as appropriate. Application errors have the shape `{"error":{"code":"provider_forbidden","details":{}}}`. Errors raised by Traefik, including its 413 body rejection, can be plain text. Check HTTP status before parsing JSON. With a large upload, a client can see a connection reset while still writing after ingress rejects the request; check the advertised byte cap locally before upload. `Expect: 100-continue` produced an explicit HTTP 413 in the LAN curl acceptance check. Failures distinguish routing, backend availability, provider authentication, timeouts, invalid JSON, incomplete generation, changed model/capabilities, compaction, and invalid assignments. Failed output is never a completed result. The result is a JSON object. The underlying Claude stream/result envelope and Ollama's JSON-in-a-string representation are normalized on the server. Every alias must occur exactly once. Duplicate, missing, invented, or foreign aliases reject the entire result. Names must be unique within the returned suite. Coverage validation is independent of the model and does not prove semantic quality. ## Mapping the prepared client job The prepared `job.json` is not the wire request. Unknown top-level and case fields are rejected; this service does not silently unpack the client envelope. | Prepared client field | Service mapping | | --- | --- | | `job_id` | Use a stable `Idempotency-Key` header, 8-128 ASCII letters/digits/underscore/hyphen. The service returns its own `job_id`; retain both locally. Hash an incompatible client ID once rather than generating a new retry key. | | `client_spec_version` | Keep locally and validate compatibility with `/v1/capabilities`; not a submitted field. | | `task` | This endpoint implements only whole-suite implementation-family planning. Validate that task locally; not a submitted field. | | `routing` | Send only `allow_external` and `allowed_external_providers`. | | `execution` | Send only `strategy`, `max_seconds`, `max_cost_usd`. | | `instructions` | Not supported as custom instructions. The server uses the documented fixed implementation-family prompt. If custom instructions are required, this contract does not yet support that request; do not silently discard them. | | `input` | Lift its one complete suite into top-level `campaign`, `suite`, and `cases`. Keep these identity strings generalized if needed. | | `output_schema` | Not accepted. The service pins groups with `name`, `description`, `members`; compare the client schema and explicitly adapt or reject incompatible expectations. | | case `case_id` | Rename to `alias`; must already be the unique CASE alias, never the original identifier. | | case `name` | No separate wire field. Preserve it explicitly in `description`, for example `Name: ...\nDescription: ...`, if the model needs it. Do not silently lose meaningful name content. | | case `description` | `description`, optionally combined with the name as above. | | case `preconditions`, `case_type`, `success_criteria` | Same-named fields, strings or null. | The application must select one exact campaign/suite before constructing this object. It must retain the complete original case records and alias mapping locally. Group `members` contain aliases only. Server-generated family names and descriptions do not become automatically permitted ClickUp export content. Only small client functions are needed: build the wire projection, validate the fixed task/schema contract, preflight, submit with a stable key, poll, save the result, and independently audit aliases. No importer replacement is needed. ## Installed route capabilities | Property | Codex | Claude | Existing local API | | --- | --- | --- | --- | | Installed client | Codex CLI 0.154.0 | Claude Code 2.1.226, pinned binary SHA-256 | Ollama 0.13.5 | | Existing broker mode | Direct subscription Responses transport, not `codex exec` | Native CLI print mode behind a wrapper | Native `/api/generate` | | New planner backend | Disabled pending effective output-budget verification | Fresh native CLI process, wrapper bypass avoided | Unchanged model-gate LAN listener | | Auth | ChatGPT Pro claim verified locally | Existing first-party OAuth setup token; associated credential metadata says Max 20x | Scoped local bearer | | Visible catalog | `gpt-6-astra`, `gpt-5.6-sol`, `gpt-5.6-terra`, `gpt-5.6-luna`, `gpt-5.5` | Fable 5, Opus 5 with 1M option, Sonnet 5, Haiku 4.5 | Pinned Qwen 2.5 14B Q4 | | Exact selected ID | `gpt-6-astra` reserved, not enabled | `claude-opus-4-8` | `qwen2.5:14b-instruct-q4_0` | | Context evidence | Local account model cache 272,000, 95% effective = 258,400 | Actual CLI response: 1,000,000 | Verified serving configuration: 8,192 | | Output control | Existing broker removes public token-limit fields; effective maximum unverified | Actual CLI reports 64,000; environment pins that ceiling | 2,048 | | New planner concurrency | None | One across the whole planning service | Shares the existing serialized GPU backend | | Hardware | Provider hosted; broker on titan-22 | Provider hosted; CLI on titan-22 | RTX 3080 10GB on titan-24 | The Claude live catalog resolves `claude-fable-5[1m]` to `claude-fable-5`, `default`/`opus[1m]` to `claude-opus-5[1m]`, `sonnet` to `claude-sonnet-5`, and `haiku` to `claude-haiku-4-5-20251001`. Catalog availability is not proof that every model has been exercised. Fable 5 suite requests reported `claude-opus-4-8` at runtime and were rejected. The initial working route therefore explicitly pins the observed Opus 4.8 model; the advertised Fable name is not verified for full-suite execution. Only the tiny Opus/Fable probes and recorded planner acceptance jobs were executed. No Claude agent wrote infrastructure code. Subscription quotas and account retention settings are not verified by model discovery. The legacy Codex and Claude wrapper scripts enable permission bypass; the Claude wrapper also enables automatic compaction and shared settings. Neither is used by the new worker. Codex supports `exec --ephemeral --output-schema --json` and stdin, but that separate CLI execution mode has not been approved as a whole-suite backend. The direct Codex broker uses `store=false`, streaming Responses, and a 900-second read timeout. That flag does not establish provider Zero Data Retention. Claude invocation, with the schema supplied by the server: ``` /opt/cli/claude -p --output-format stream-json --verbose \ --no-session-persistence --safe-mode --tools '' \ --strict-mcp-config --mcp-config '{"mcpServers":{}}' \ --setting-sources '' --disable-slash-commands --permission-mode dontAsk \ --no-chrome --model 'claude-opus-4-8[1m]' --effort medium \ --max-budget-usd 5 --max-turns 6 \ --system-prompt '' --json-schema '' ``` The complete suite is an isolated input file connected to stdin, never an argv string or interpolated shell command. The process receives only a short allowlisted environment, its OAuth token, and a fresh temporary HOME/config directory. `DISABLE_COMPACT`, `DISABLE_AUTO_COMPACT`, `DISABLE_TELEMETRY`, `DISABLE_ERROR_REPORTING`, `DISABLE_PROMPT_CACHING`, `CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC`, `CLAUDE_CODE_DISABLE_AUTO_MEMORY`, and `CLAUDE_CODE_SKIP_PROMPT_HISTORY` are enabled; retries and updates are disabled. The actual initialization event must report only `StructuredOutput`, no MCP servers or plugins. The worker rejects compaction events, unexpected models, incomplete output, changed limits, and missing terminal results. These are client controls, not provider ZDR. Account-specific training/retention settings remain unverified. See [Claude data usage](https://code.claude.com/docs/en/data-usage), [Claude CLI flags](https://code.claude.com/docs/en/cli-reference), [Claude environment controls](https://code.claude.com/docs/en/env-vars), and [Codex noninteractive behavior](https://learn.chatgpt.com/docs/non-interactive-mode). ## Capacity and transport The new API accepts at most 1 MiB and 400 cases per request, with at most 32 KiB per field. The existing local endpoint still accepts only 128 KiB and its original context/output limits. Claude subprocess stdout is bounded to 4 MiB; normalized results to 1 MiB. Final names are at most 64 characters, descriptions at most 240, and groups at most five cases. Preflight includes the entire serialized suite, fixed instructions, schema, reserved harness overhead, and output space. There is no accurate account-specific tokenizer. The input check uses UTF-8 byte count plus 8,192 reserved harness tokens, with the 64,000 output ceiling reserved for each of at most six turns against context. This is a conservative bound, not a measured token count. The output reservation is an explicit estimate allowing a family per case and 8,192 reasoning tokens; it is not a verified bound on arbitrary generated wording. Medium adaptive reasoning can consume output budget. Incomplete output fails explicitly; the server never drops cases to turn it into success. Small local requests reserve 1,024 overhead tokens and 2,048 output tokens within 8,192 context. The realistic 14/75/363 fixtures exceed this conservative local whole-suite budget and must use the approved hosted path or fail preflight. All CLI invocations share at most 1800 seconds per suite job; cancellation and timeouts terminate its process group. Async HTTP calls finish promptly, with a 30-second body-read timeout, 40-second ingress response-header timeout, and recommended 45-second client timeout. Switchyard's ten-minute internal request limit carries only a tiny immediate routing decision, not the long inference job or its large body. If a real complete suite cannot fit, the current API returns capacity/unsupported. A future strategy would extract profiles with preserved aliases, generate overlapping implementation candidates across all batches, compare candidates across batch boundaries, then reconcile and audit the entire suite. Independent batch grouping followed by concatenation is not supported. Transport chunking would not add context. ## WSL examples Supply only the scoped credential privately. Do not add `-k` or `-L`. ``` read -rsp 'Suite API token: ' SUITE_PLANNING_TOKEN export SUITE_PLANNING_TOKEN curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \ --connect-timeout 10 --max-time 45 --fail-with-body \ -H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \ https://worker.bstein.dev/suite-planning/healthz ``` Local-only synthetic grouping (any scoped credential): ``` cat > local-synthetic.json <<'JSON' {"campaign":"SYNTHETIC","suite":"PARSER","cases":[{"alias":"CASE-1","description":"Parse valid configuration text and check returned fields.","preconditions":"An isolated parser fixture is reset before each case.","case_type":"nominal","success_criteria":"Returned fields match the supplied configuration values."},{"alias":"CASE-2","description":"Parse invalid configuration text and check returned diagnostics.","preconditions":"An isolated parser fixture is reset before each case.","case_type":"fault injection","success_criteria":"The malformed input produces the specified diagnostic fields."}],"routing":{"allow_external":false,"allowed_external_providers":[]},"execution":{"strategy":"whole_suite"}} JSON curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \ --connect-timeout 10 --max-time 45 --fail-with-body \ -H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \ -H 'Content-Type: application/json' -H 'Idempotency-Key: local-parser-pilot-001' \ --data-binary @local-synthetic.json \ https://worker.bstein.dev/suite-planning/v1/jobs ``` Externally allowed, complete 363-case synthetic grouping (use `synthetic_token`): ``` curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \ --connect-timeout 10 --max-time 45 --fail-with-body \ -H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \ https://worker.bstein.dev/suite-planning/v1/synthetic/363 > synthetic-suite.json python3 - <<'PY' import json with open('synthetic-suite.json') as stream: request = json.load(stream) request['routing'] = {'allow_external': True, 'allowed_external_providers': ['claude']} request['execution'] = {'strategy': 'whole_suite', 'max_seconds': 1800, 'max_cost_usd': 30} with open('synthetic-request.json', 'w') as stream: json.dump(request, stream) PY curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \ --connect-timeout 10 --max-time 45 --fail-with-body \ -H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \ -H 'Content-Type: application/json' --data-binary @synthetic-request.json \ https://worker.bstein.dev/suite-planning/v1/preflight curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \ --connect-timeout 10 --max-time 45 --fail-with-body \ -H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \ -H 'Content-Type: application/json' -H 'Idempotency-Key: synthetic-363-pilot-001' \ --data-binary @synthetic-request.json \ https://worker.bstein.dev/suite-planning/v1/jobs > submitted-job.json JOB_ID=$(python3 -c 'import json; print(json.load(open("submitted-job.json"))["job_id"])') curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \ --connect-timeout 10 --max-time 45 --fail-with-body \ -H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \ "https://worker.bstein.dev/suite-planning/v1/jobs/$JOB_ID/result" ``` Repeat the last GET until `status` is `completed`, `failed`, or `cancelled`. Save the result locally before retention expires. Keep the original key for network retries. Use a new key only when intentionally starting another provider job. Capability, status-only, and cancellation requests: ``` curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \ --connect-timeout 10 --max-time 45 --fail-with-body \ -H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \ https://worker.bstein.dev/suite-planning/v1/capabilities curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \ --connect-timeout 10 --max-time 45 --fail-with-body \ -H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \ "https://worker.bstein.dev/suite-planning/v1/jobs/$JOB_ID" curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \ --connect-timeout 10 --max-time 45 --fail-with-body -X DELETE \ -H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \ "https://worker.bstein.dev/suite-planning/v1/jobs/$JOB_ID" ``` For an approved generalized job, obtain `operational_token`, construct the same wire object locally, and request only `allowed_external_providers:["claude"]`. The credential name does not go in the HTTP request; only its secret value is used as the bearer. Neither a provider credential nor a Vault token belongs in this header. The standalone `scripts/ops/hermes_suite_probe.py` uses only Python urllib and the same API credential. Its HTTPS handler implements the equivalent of curl `--resolve`, retaining certificate verification and bypassing proxies. It retrieves synthetic fixtures, submits, checks idempotency, polls, and validates exact coverage. It does not integrate with or replace the importer. Proposed application additions are `get_capabilities()`, `preflight_suite()`, `submit_suite(idempotency_key)`, `get_job_result()`, and `cancel_job()`. Preserve the application's local selection, alias mapping, ownership and coverage audits, and export policy. ## Deployment, logging, and rollback The planner runs on titan-22 alongside the existing RWO CLI tools volume. Only an init container reads its tools subdirectory and copies the hash-pinned native binary. The serving container has no agent home, cluster token, repository, ClickUp credentials, provider API keys, or arbitrary command endpoint. Its requests are 100m CPU/512 MiB and limits are 2 CPU/2 GiB. No shared GPU workload is changed. Input/output/config/debug storage is per-job tmpfs, removed on success, failure, timeout, and cancellation. Pod termination also clears it. The observed CLI writes configuration and backup files despite no session persistence; these share that temporary directory. The durable metadata PVC uses local-path and is outside the Longhorn backup path. Ordinary worker logs contain job ID, selected provider/model, status and duration only. Provider stderr is discarded. The pod is excluded from Fluent Bit; no prompt is sent to the existing Switchyard routing logs. TLS ingress does not log bodies. Central archival deletion is not claimed or required because the new path does not send content there; account-side retention remains separate. Rollback through Git/Flux: To revoke only generalized-data permission, set `PLANNING_GENERALIZED_CLAUDE_APPROVED=false` and reconcile `hermes`. Revoke/rotate `operational_token` through Vault if the credential itself must be invalidated, then restart through a Flux-tracked deployment revision so the pre-populated credential mount refreshes. Keep both older token fields intact. 1. Remove the three suite-planner resource entries and its ConfigMap generator from `services/hermes/kustomization.yaml`; reconcile `hermes` to remove the endpoint. 2. Remove `suite_decision`, the three suite targets/routes, and the corresponding Switchyard revision update. Reconcile; keep unrelated routes unchanged. 3. Remove the suite Vault seed/bootstrap resources and the suite role/policy additions if no longer needed. Remove the created Vault role/policy and secret through the normal Vault administrative workflow; removing a completed Job does not revoke them. 4. Deleting the metadata PVC removes retry protection, so do it only after all jobs are terminal and no client can submit. No provider job is replayed automatically. The existing `/local-model` endpoint and token do not require rollback changes.