atlas-iac/docs/hermes_suite_planning.md

37 KiB

Suite planning API

This optional API groups one complete campaign/suite into implementation families. The existing /local-model API, GPU allocation, serving model, and limits are unchanged. Deployment uses Flux; the application and workbook stay on the laptop.

Current multi-pass policy

New jobs use the multi-pass implementation policy: two independent full-suite proposals, reconciliation, up to two bounded semantic reviews, and a programmatic five-case final cap. Names are at most 64 characters. The endpoint and request fields are unchanged. Optional top-level review_summary is returned only with the authorized result. Configuration remains suite-v6-20260929; policy, prompt, and execution revisions identify the changed grouping behavior. The prior single-pass acceptance results below are historical and do not measure the new workflow. The new process shares the user-approved 1200-second deadline and USD 10 CLI estimate guard across all invocations. This is not subscription billing. Progress updates every five seconds during a CLI pass. The separate local-model endpoint is unchanged; the 8K local route cannot admit the new multi-pass reconciliation schema.

Earlier CLI failure diagnostics update (historical)

Execution revision claude-diagnostics-turns-v1-20260929 adds content-free cli_diagnostics and failure-stage details, including exit/signal, final subtype, reported turns, observed stop reason, token counts, and structured-output location. Unknown measurements remain null. Source, generated descriptions, raw provider errors, and transcripts are not retained. The normal schema and exact membership checks still gate completion. Failed validation now preserves operational usage.

The CLI turn ceiling is six, with the same model, medium effort, 64,000 per-call output ceiling, maximum 900-second job deadline, and USD 5 CLI estimated-cost cap. Preflight reserves all six possible outputs against context. This is headroom within one fresh job, not an automatic job retry or a fallback. The historical three-turn ceiling was exhausted by a deterministic synthetic structured-output repair test; six turns completed that same test. This does not establish the cause of a historical failure whose terminal event was deleted.

The HTTP configuration revision and implementation-proximity prompt revision/hash are unchanged. execution_revision records the new behavior in capability, preflight, and job metadata. Use a new idempotency key only when intentionally starting a new attempt; the failed historical job is never automatically rerun. See the diagnosis and verification record. Rollback this update by reverting its Git commit and reconciling the Hermes Flux Kustomization; do not revert the earlier prompt commit. Wait for no active jobs before rollout because results are held in memory.

Implementation-proximity prompt update

New jobs use prompt revision implementation-proximity-v2-20260929. The HTTP configuration revision remains suite-v6-20260929; request and result schemas, models, reasoning effort, token limits, authentication, and routing policy are unchanged. prompt_revision and prompt_sha256 are additive provenance in health, capabilities, preflight selection, and persisted job/status/result metadata. Historical jobs keep their original metadata; they are not relabeled on replay.

The objective is the additional test-development effort after implementing a representative case. Success criteria provide the primary evidence of actions, observations, measurements, and assertions. Description supplies behavior and operation; preconditions and other material constraints supply setup/state; case type is supporting context. There is no keyword or word-count weighting. The prompt distinguishes inexpensive input/assertion variations from new test machinery, checks entire-family coherence, reviews singletons, preserves uncertainty in generalized values, and confines names/descriptions to assigned test objectives.

The same system prompt is supplied to the actual grouping invocation and remains applicable during structured-output turns or format repair. There is no separate repair model. Cases enter a fresh isolated session without any previous grouping as a target answer. The expanded prompt is included in the existing conservative preflight accounting; no context or output setting has been increased.

For a fresh comparison, keep the same case request but use a NEW Idempotency-Key. Reusing the old key returns the old job instead of invoking the revised prompt. Include prompt_revision or prompt_sha256 in the laptop's result-cache identity. The synthetic check fixture is testing/fixtures/suite_implementation_proximity.json; only its request object is sent, never its independent expected-family mapping. No real roster job is launched by this deployment or its acceptance checks.

Verified on the deployed endpoint: the 14-case adversarial synthetic suite completed in 21.059 seconds with six implementation families, exact alias coverage, and zero incorrect or missed merge pairs against its independent expected map. Identical broad descriptions did not merge different machinery. Nominal/negative-result and extra-assertion variants shared their existing observation code; an underspecified case retained its uncertainty. Every observed native CLI request in the loopback structured-output transport check contained the complete revised system prompt. The focused unit suite passes 36 tests. These checks establish the synthetic behavior, not the quality of a real-suite rerun. See the prompt-update acceptance evidence.

Acceptance results, 2026-09-29

The final runs use suite-v6-20260929, deployed commit 299b9ed1, including case-type fields. These supersede the preliminary runs without that field. All tests used synthetic content. HTTPS tests ran from titan-jh, source 192.168.22.8, directly on its LAN interface to 192.168.22.50:443, with hostname certificate verification and proxies disabled. The user's WSL route has not been tested; run the health command below from the laptop before uploading anything.

Cases Actual request bytes CLI input/output tokens Client wall seconds Families / singletons Exact coverage Pattern match
14 9,847 4,185 / 1,250 16.947 9 / 4 Yes Exact
75 55,984 19,145 / 2,336 25.755 9 / 3 Yes Exact
363 273,761 89,753 / 8,568 80.071 9 / 3 Yes Exact

Request sizes are the UTF-8 compact JSON bytes actually submitted, excluding HTTP headers. Token counts are CLI-reported aggregate usage across each fresh job's two structured-output turns; they are not an exact standalone input-token count. All three final jobs completed on their first attempt with no observed retries or rate-limit errors. Two CLI turns are part of producing the structured result, not automatic job retries. CLI cost estimates were USD 0.052175, 0.154125, and 0.662965; these are estimates, not verified subscription charges. Remaining account quota is not available through this endpoint and may be consumed by other account users.

The actual terminal modelUsage reports claude-opus-4-8[1m], with canonicalModel=claude-opus-4-8, provider=firstParty, context 1,000,000 and maximum output 64,000. The invocation requests claude-opus-4-8[1m]; these results also independently verify the returned model metadata, rather than trusting the request flag or the internal atlas/planning/claude route label. Provider physical hardware and inference region are unknown. This verifies the tested suite sizes, not the entire advertised million-token context.

The 363-case fixture has descriptions of 48-245 bytes (median 234), preconditions 67-143 (median 136), success criteria 73-135 (median 132), and case types 7-15 bytes. Most cases have multi-sentence fields plus operating conditions, verification method, and a generic reference. Six machinery patterns are interleaved throughout the input. Watchdog terminology occurs in three different machinery patterns. A distant pair has identical complete case text but different aliases. Distinctive thermal, acoustic, and build cases occur at the beginning, middle, and end.

Against the independent synthetic implementation map, all sizes had pair precision and recall 1.0, zero incorrect merge pairs, and zero missed merge pairs. Manual review of all 27 family descriptions found no objectives attributed to another family. Four singletons in the 14-case fixture and three in the larger fixtures are expected. These are clear synthetic patterns, not evidence of quality on the user's more ambiguous real suite. No real case quality is claimed.

Evidence for complete input handling goes beyond IDs:

  • Unit checks compare every serialized case and field with the original fixture.
  • The deployed native CLI was also run against a loopback mock provider, with the exact production command and full 14/75/363 inputs. The captured requests contained the complete source string, byte for byte, including every field. The largest first provider request was 287,914 bytes. This transport check used no hosted model.
  • Live jobs returned the expected distinct beginning/middle/end patterns and reunited interleaved cases. Initialization allowed only StructuredOutput, with no MCP, plugins, or file-reading tools. Each job used fresh HOME/config and stdin.
  • No CLI compaction event or incomplete-output signal was observed. Compaction was disabled, and the gateway performs no summarization or selective file reading.

Those observations support complete transmission and the tested grouping behavior. They cannot inspect provider internals or prove that a model attended equally to every word. The response's false compaction/truncation values describe gateway/CLI observations, not a provider-side attestation. An exact tokenizer is unavailable; preflight's byte bound and output reservation are documented below.

The local-only two-case parser example also completed in 3.535 seconds on the RTX backend with exact coverage and no hosted destination. An earlier local result incorrectly appended descriptions to aliases and was rejected; the new adapter constrains the local output schema to supplied aliases. The existing local-model endpoint itself is unchanged.

Verification also covered authentication failures, permission escalation, unknown providers, synthetic-versus-operational scope, local-only capacity rejection, management-path rejection, cross-client ownership, idempotent replay, oversized HTTP 413, and content-free routine worker logs. Mocked backend failures verified that no local failure launches Claude. No shared backend was disrupted. The focused unit suite passes 35 tests. Service renders, client dry-runs, and Flux diffs passed. The installed kubectl rejects the repository's combined --server-side --dry-run=client syntax; the supported client dry-run was used instead.

Full evidence and successful normalized envelopes:

Ready: synthetic client testing, direct complete-suite execution at all three tested sizes, and the separately approved Claude credential. Remaining client work: WSL connectivity, explicit client-envelope adaptation, and a controlled real pilot after reviewing this handoff. Codex capacity verification and provider/account retention or training controls remain unresolved. The user's account-specific approval is recorded; broader organizational approval or release classification has not been established by this infrastructure work.

Connection and credentials

LAN_HOST: worker.bstein.dev
LAN_IP: 192.168.22.50
PORT: 443
BASE_URL: https://worker.bstein.dev/suite-planning
HEALTH_URL: https://worker.bstein.dev/suite-planning/healthz
CAPABILITIES_URL: https://worker.bstein.dev/suite-planning/v1/capabilities
PREFLIGHT_URL: https://worker.bstein.dev/suite-planning/v1/preflight
SUBMIT_URL: https://worker.bstein.dev/suite-planning/v1/jobs
AUTHENTICATION_HEADER: Authorization: Bearer <scoped token>
VAULT_PATH: kv/atlas/hermes/suite-planning-api
LOCAL_ONLY_FIELD: token
SYNTHETIC_EXTERNAL_FIELD: synthetic_token
APPROVED_OPERATIONAL_FIELD: operational_token
INITIAL_CONCURRENCY: 1; busy submissions return 429; no waiting queue
JOB_TIMEOUT: up to 1200 seconds
HTTP_TIMEOUT: client 45 seconds; submit/status do not wait for inference
TLS: existing worker.bstein.dev certificate; normal trusted CA verification

Retrieve the appropriate field privately from Vault. Neither token grants provider, Vault, Kubernetes, shell, or ClickUp access. The old local-model token remains at kv/atlas/hermes/model-gate-lan-api. Tokens are distinct; an old local-model token does not authenticate this API.

The credentials have different server-enforced scopes:

Vault field Local inference Hosted inference
token Allowed within local capacity Prohibited
synthetic_token Allowed within local capacity Exact API fixtures only; full content hash checked
operational_token Allowed within local capacity Generalized CASE records, Claude only, under the explicit account approval

synthetic_token has not been broadened. Changing a fixture or sending real generalized records with that token returns external_data_not_approved even with allow_external=true. The operational credential is separate, grants only claude, and requires the server's PLANNING_GENERALIZED_CLAUDE_APPROVED=true. The user explicitly approved the existing Hermes first-party Claude OAuth account for generalized CASE records on 2026-09-29. No real records were used in acceptance. Original identifiers, uncensored text, and alias mappings remain on the laptop. The server cannot determine whether arbitrary prose is adequately generalized; that remains the application's and organization's responsibility. Account-specific training and retention settings remain unverified. A request flag cannot grant credential permission or override data scope.

Request contract

Submit with POST https://worker.bstein.dev/suite-planning/v1/jobs. The request JSON Schema describes the accepted wire object. The result JSON Schema describes result, not the whole job envelope. Additional server checks enforce UTF-8 byte limits, unique aliases, ownership, permissions, and exact result coverage.

{
  "campaign": "SYNTHETIC",
  "suite": "PARSER",
  "cases": [
    {"alias": "CASE-1", "description": "Parse a valid configuration and assert accepted fields."},
    {"alias": "CASE-2", "description": "Parse an invalid configuration and assert diagnostic fields."}
  ],
  "routing": {
    "allow_external": false,
    "allowed_external_providers": []
  },
  "execution": {
    "strategy": "whole_suite",
    "max_seconds": 1200,
    "max_cost_usd": 10
  }
}

campaign, suite, and nonempty cases are required. Each case requires a unique CASE-* alias and a description. Optional case fields are success_criteria, preconditions, operating_condition, case_type, verification_method, target, swci, verifies, functional_area, functional_group, and functional_group_name. Values are strings or null. Optional per-case campaign/suite values must exactly match the job's ownership. Unknown fields are rejected. [reference] is ordinary source text, never an alias. Identical descriptions with different aliases remain distinct cases.

Missing routing policy means local-only. External provider names are claude and codex. The requested list must be contained in the credential's permissions; unknown providers, malformed booleans, duplicate names, and permission escalation are rejected before routing. allow_external=false cannot include providers. whole_suite is the only strategy. No batching, summarization, or reconciliation is silently substituted. max_seconds is 10-1200. The Claude CLI guard is at most USD 10 in its estimated usage accounting; this is not a verified subscription billing ceiling and can overshoot within a single provider call.

The gateway deterministically selects a fitting permitted backend, preferring local capacity. It asks the fixed Switchyard route to confirm that binding, then performs one inference attempt. Switchyard receives only select, never case text. Its new routes each have exactly one target, zero client retries, and no classifier: atlas/planning/local, atlas/planning/claude, atlas/planning/codex. These are decision routes; clients use the HTTPS job API to obtain inference results. Existing general Hermes routes retain their old behavior and are not this API.

Job lifecycle and normalized response

  • GET /healthz: authenticated service readiness.
  • GET /v1/capabilities: permissions, models, configuration revision, and limits.
  • POST /v1/preflight: validate the identical proposed job without generation.
  • POST /v1/jobs: require Idempotency-Key, return 202 and a job ID.
  • GET /v1/jobs/<id>: status and metadata.
  • GET /v1/jobs/<id>/result: status plus the normalized result when complete.
  • DELETE /v1/jobs/<id>: request cancellation.
  • GET /v1/synthetic/14, /75, /363: immutable synthetic fixtures.

Retry a submission with the SAME key and SAME body. It returns the original job, not a second provider attempt. A changed body with that key returns 409. Keys are scoped to the authenticated credential and retained seven days. The maximum is 10,000 metadata records; excess submissions fail rather than evicting live keys. No provider retry or fallback occurs after an error. An operator may intentionally start a new attempt with a new key after diagnosing the previous failure.

Results remain in memory for at most one hour (cleanup every 30 seconds), with a 128-result bound. Durable local SQLite holds only operational metadata and hashes, not prompts or results. Restarts mark unfinished jobs interrupted_no_retry and never resubmit them. Previously completed results are lost on restart; retrieval returns 410. A cancellation holds the slot until the underlying attempt ends. Claude cancellation kills the isolated process group. Local generation cannot be cancelled at the existing API; its eventual result is discarded. Provider-side cancellation and exact billing after disconnection are not guaranteed.

Example completed envelope (illustrative synthetic values):

{
  "job_id": "0123456789abcdef0123456789abcdef",
  "status": "completed",
  "configuration_revision": "suite-v6-20260929",
  "routing": {"allow_external": true, "allowed_external_providers": ["claude"]},
  "selection": {"provider": "claude", "backend": "claude-code-2.1.226", "model": "claude-opus-4-8"},
  "model": "claude-opus-4-8",
  "attempted_destinations": ["switchyard:atlas/planning/claude", "claude:claude-opus-4-8"],
  "compaction": false,
  "truncation": false,
  "usage": {"input_tokens": 1000, "output_tokens": 100},
  "wall_seconds": 12.3,
  "result": {
    "groups": [{"name": "Configuration parser fixtures", "description": "Parameterize configuration text and assert parsed fields or diagnostics.", "members": ["CASE-1", "CASE-2"]}]
  }
}

The real envelope also includes input bytes, input hash, reservation method, model usage, timing, and estimated cost where available. Unknown values are null. Failed jobs retain status and fixed error codes but no provider stderr. Status remains HTTP 200 for an existing failed job; inspect status and error.code. Invalid submissions use 400/401/403/409/413/415/422/429 as appropriate. Application errors have the shape {"error":{"code":"provider_forbidden","details":{}}}. Errors raised by Traefik, including its 413 body rejection, can be plain text. Check HTTP status before parsing JSON. With a large upload, a client can see a connection reset while still writing after ingress rejects the request; check the advertised byte cap locally before upload. Expect: 100-continue produced an explicit HTTP 413 in the LAN curl acceptance check. Failures distinguish routing, backend availability, provider authentication, timeouts, invalid JSON, incomplete generation, changed model/capabilities, compaction, and invalid assignments. Failed output is never a completed result.

The result is a JSON object. The underlying Claude stream/result envelope and Ollama's JSON-in-a-string representation are normalized on the server. Every alias must occur exactly once. Duplicate, missing, invented, or foreign aliases reject the entire result. Names must be unique within the returned suite. Coverage validation is independent of the model and does not prove semantic quality.

Mapping the prepared client job

The prepared job.json is not the wire request. Unknown top-level and case fields are rejected; this service does not silently unpack the client envelope.

Prepared client field Service mapping
job_id Use a stable Idempotency-Key header, 8-128 ASCII letters/digits/underscore/hyphen. The service returns its own job_id; retain both locally. Hash an incompatible client ID once rather than generating a new retry key.
client_spec_version Keep locally and validate compatibility with /v1/capabilities; not a submitted field.
task This endpoint implements only whole-suite implementation-family planning. Validate that task locally; not a submitted field.
routing Send only allow_external and allowed_external_providers.
execution Send only strategy, max_seconds, max_cost_usd.
instructions Not supported as custom instructions. The server uses the documented fixed implementation-family prompt. If custom instructions are required, this contract does not yet support that request; do not silently discard them.
input Lift its one complete suite into top-level campaign, suite, and cases. Keep these identity strings generalized if needed.
output_schema Not accepted. The service pins groups with name, description, members; compare the client schema and explicitly adapt or reject incompatible expectations.
case case_id Rename to alias; must already be the unique CASE alias, never the original identifier.
case name No separate wire field. Preserve it explicitly in description, for example Name: ...\nDescription: ..., if the model needs it. Do not silently lose meaningful name content.
case description description, optionally combined with the name as above.
case preconditions, case_type, success_criteria Same-named fields, strings or null.

The application must select one exact campaign/suite before constructing this object. It must retain the complete original case records and alias mapping locally. Group members contain aliases only. Server-generated family names and descriptions do not become automatically permitted ClickUp export content.

Only small client functions are needed: build the wire projection, validate the fixed task/schema contract, preflight, submit with a stable key, poll, save the result, and independently audit aliases. No importer replacement is needed.

Installed route capabilities

Property Codex Claude Existing local API
Installed client Codex CLI 0.154.0 Claude Code 2.1.226, pinned binary SHA-256 Ollama 0.13.5
Existing broker mode Direct subscription Responses transport, not codex exec Native CLI print mode behind a wrapper Native /api/generate
New planner backend Disabled pending effective output-budget verification Fresh native CLI process, wrapper bypass avoided Unchanged model-gate LAN listener
Auth ChatGPT Pro claim verified locally Existing first-party OAuth setup token; associated credential metadata says Max 20x Scoped local bearer
Visible catalog gpt-6-astra, gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna, gpt-5.5 Fable 5, Opus 5 with 1M option, Sonnet 5, Haiku 4.5 Pinned Qwen 2.5 14B Q4
Exact selected ID gpt-6-astra reserved, not enabled claude-opus-4-8 qwen2.5:14b-instruct-q4_0
Context evidence Local account model cache 272,000, 95% effective = 258,400 Actual CLI response: 1,000,000 Verified serving configuration: 8,192
Output control Existing broker removes public token-limit fields; effective maximum unverified Actual CLI reports 64,000; environment pins that ceiling 2,048
New planner concurrency None One across the whole planning service Shares the existing serialized GPU backend
Hardware Provider hosted; broker on titan-22 Provider hosted; CLI on titan-22 RTX 3080 10GB on titan-24

The Claude live catalog resolves claude-fable-5[1m] to claude-fable-5, default/opus[1m] to claude-opus-5[1m], sonnet to claude-sonnet-5, and haiku to claude-haiku-4-5-20251001. Catalog availability is not proof that every model has been exercised. Fable 5 suite requests reported claude-opus-4-8 at runtime and were rejected. The initial working route therefore explicitly pins the observed Opus 4.8 model; the advertised Fable name is not verified for full-suite execution. Only the tiny Opus/Fable probes and recorded planner acceptance jobs were executed. No Claude agent wrote infrastructure code. Subscription quotas and account retention settings are not verified by model discovery.

The legacy Codex and Claude wrapper scripts enable permission bypass; the Claude wrapper also enables automatic compaction and shared settings. Neither is used by the new worker. Codex supports exec --ephemeral --output-schema --json and stdin, but that separate CLI execution mode has not been approved as a whole-suite backend. The direct Codex broker uses store=false, streaming Responses, and a 900-second read timeout. That flag does not establish provider Zero Data Retention.

Claude invocation, with the schema supplied by the server:

/opt/cli/claude -p --output-format stream-json --verbose \
  --no-session-persistence --safe-mode --tools '' \
  --strict-mcp-config --mcp-config '{"mcpServers":{}}' \
  --setting-sources '' --disable-slash-commands --permission-mode dontAsk \
  --no-chrome --model 'claude-opus-4-8[1m]' --effort medium \
  --max-budget-usd 5 --max-turns 6 \
  --system-prompt '<fixed server prompt>' --json-schema '<fixed schema>'

The complete suite is an isolated input file connected to stdin, never an argv string or interpolated shell command. The process receives only a short allowlisted environment, its OAuth token, and a fresh temporary HOME/config directory. DISABLE_COMPACT, DISABLE_AUTO_COMPACT, DISABLE_TELEMETRY, DISABLE_ERROR_REPORTING, DISABLE_PROMPT_CACHING, CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC, CLAUDE_CODE_DISABLE_AUTO_MEMORY, and CLAUDE_CODE_SKIP_PROMPT_HISTORY are enabled; retries and updates are disabled. The actual initialization event must report only StructuredOutput, no MCP servers or plugins. The worker rejects compaction events, unexpected models, incomplete output, changed limits, and missing terminal results.

These are client controls, not provider ZDR. Account-specific training/retention settings remain unverified. See Claude data usage, Claude CLI flags, Claude environment controls, and Codex noninteractive behavior.

Capacity and transport

The new API accepts at most 1 MiB and 400 cases per request, with at most 32 KiB per field. The existing local endpoint still accepts only 128 KiB and its original context/output limits. Claude subprocess stdout is bounded to 4 MiB; normalized results to 1 MiB. Final names are at most 64 characters, descriptions at most 240, and groups at most five cases.

Preflight includes the entire serialized suite, fixed instructions, schema, reserved harness overhead, and output space. There is no accurate account-specific tokenizer. The input check uses UTF-8 byte count plus 8,192 reserved harness tokens, with the 64,000 output ceiling reserved for each of at most six turns against context. This is a conservative bound, not a measured token count. The output reservation is an explicit estimate allowing a family per case and 8,192 reasoning tokens; it is not a verified bound on arbitrary generated wording. Medium adaptive reasoning can consume output budget. Incomplete output fails explicitly; the server never drops cases to turn it into success.

Small local requests reserve 1,024 overhead tokens and 2,048 output tokens within 8,192 context. The realistic 14/75/363 fixtures exceed this conservative local whole-suite budget and must use the approved hosted path or fail preflight.

All CLI invocations share at most 1200 seconds per suite job; cancellation and timeouts terminate its process group. Async HTTP calls finish promptly, with a 30-second body-read timeout, 40-second ingress response-header timeout, and recommended 45-second client timeout. Switchyard's ten-minute internal request limit carries only a tiny immediate routing decision, not the long inference job or its large body.

If a real complete suite cannot fit, the current API returns capacity/unsupported. A future strategy would extract profiles with preserved aliases, generate overlapping implementation candidates across all batches, compare candidates across batch boundaries, then reconcile and audit the entire suite. Independent batch grouping followed by concatenation is not supported. Transport chunking would not add context.

WSL examples

Supply only the scoped credential privately. Do not add -k or -L.

read -rsp 'Suite API token: ' SUITE_PLANNING_TOKEN
export SUITE_PLANNING_TOKEN
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
  --connect-timeout 10 --max-time 45 --fail-with-body \
  -H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \
  https://worker.bstein.dev/suite-planning/healthz

Local-only synthetic grouping (any scoped credential):

cat > local-synthetic.json <<'JSON'
{"campaign":"SYNTHETIC","suite":"PARSER","cases":[{"alias":"CASE-1","description":"Parse valid configuration text and check returned fields.","preconditions":"An isolated parser fixture is reset before each case.","case_type":"nominal","success_criteria":"Returned fields match the supplied configuration values."},{"alias":"CASE-2","description":"Parse invalid configuration text and check returned diagnostics.","preconditions":"An isolated parser fixture is reset before each case.","case_type":"fault injection","success_criteria":"The malformed input produces the specified diagnostic fields."}],"routing":{"allow_external":false,"allowed_external_providers":[]},"execution":{"strategy":"whole_suite"}}
JSON
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
  --connect-timeout 10 --max-time 45 --fail-with-body \
  -H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \
  -H 'Content-Type: application/json' -H 'Idempotency-Key: local-parser-pilot-001' \
  --data-binary @local-synthetic.json \
  https://worker.bstein.dev/suite-planning/v1/jobs

Externally allowed, complete 363-case synthetic grouping (use synthetic_token):

curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
  --connect-timeout 10 --max-time 45 --fail-with-body \
  -H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \
  https://worker.bstein.dev/suite-planning/v1/synthetic/363 > synthetic-suite.json
python3 - <<'PY'
import json
with open('synthetic-suite.json') as stream:
    request = json.load(stream)
request['routing'] = {'allow_external': True, 'allowed_external_providers': ['claude']}
request['execution'] = {'strategy': 'whole_suite', 'max_seconds': 1200, 'max_cost_usd': 10}
with open('synthetic-request.json', 'w') as stream:
    json.dump(request, stream)
PY
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
  --connect-timeout 10 --max-time 45 --fail-with-body \
  -H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \
  -H 'Content-Type: application/json' --data-binary @synthetic-request.json \
  https://worker.bstein.dev/suite-planning/v1/preflight
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
  --connect-timeout 10 --max-time 45 --fail-with-body \
  -H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \
  -H 'Content-Type: application/json' -H 'Idempotency-Key: synthetic-363-pilot-001' \
  --data-binary @synthetic-request.json \
  https://worker.bstein.dev/suite-planning/v1/jobs > submitted-job.json
JOB_ID=$(python3 -c 'import json; print(json.load(open("submitted-job.json"))["job_id"])')
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
  --connect-timeout 10 --max-time 45 --fail-with-body \
  -H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \
  "https://worker.bstein.dev/suite-planning/v1/jobs/$JOB_ID/result"

Repeat the last GET until status is completed, failed, or cancelled. Save the result locally before retention expires. Keep the original key for network retries. Use a new key only when intentionally starting another provider job.

Capability, status-only, and cancellation requests:

curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
  --connect-timeout 10 --max-time 45 --fail-with-body \
  -H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \
  https://worker.bstein.dev/suite-planning/v1/capabilities
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
  --connect-timeout 10 --max-time 45 --fail-with-body \
  -H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \
  "https://worker.bstein.dev/suite-planning/v1/jobs/$JOB_ID"
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
  --connect-timeout 10 --max-time 45 --fail-with-body -X DELETE \
  -H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \
  "https://worker.bstein.dev/suite-planning/v1/jobs/$JOB_ID"

For an approved generalized job, obtain operational_token, construct the same wire object locally, and request only allowed_external_providers:["claude"]. The credential name does not go in the HTTP request; only its secret value is used as the bearer. Neither a provider credential nor a Vault token belongs in this header.

The standalone scripts/ops/hermes_suite_probe.py uses only Python urllib and the same API credential. Its HTTPS handler implements the equivalent of curl --resolve, retaining certificate verification and bypassing proxies. It retrieves synthetic fixtures, submits, checks idempotency, polls, and validates exact coverage. It does not integrate with or replace the importer. Proposed application additions are get_capabilities(), preflight_suite(), submit_suite(idempotency_key), get_job_result(), and cancel_job(). Preserve the application's local selection, alias mapping, ownership and coverage audits, and export policy.

Deployment, logging, and rollback

The planner runs on titan-22 alongside the existing RWO CLI tools volume. Only an init container reads its tools subdirectory and copies the hash-pinned native binary. The serving container has no agent home, cluster token, repository, ClickUp credentials, provider API keys, or arbitrary command endpoint. Its requests are 100m CPU/512 MiB and limits are 2 CPU/2 GiB. No shared GPU workload is changed.

Input/output/config/debug storage is per-job tmpfs, removed on success, failure, timeout, and cancellation. Pod termination also clears it. The observed CLI writes configuration and backup files despite no session persistence; these share that temporary directory. The durable metadata PVC uses local-path and is outside the Longhorn backup path. Ordinary worker logs contain job ID, selected provider/model, status and duration only. Provider stderr is discarded. The pod is excluded from Fluent Bit; no prompt is sent to the existing Switchyard routing logs. TLS ingress does not log bodies. Central archival deletion is not claimed or required because the new path does not send content there; account-side retention remains separate.

Rollback through Git/Flux:

To revoke only generalized-data permission, set PLANNING_GENERALIZED_CLAUDE_APPROVED=false and reconcile hermes. Revoke/rotate operational_token through Vault if the credential itself must be invalidated, then restart through a Flux-tracked deployment revision so the pre-populated credential mount refreshes. Keep both older token fields intact.

  1. Remove the three suite-planner resource entries and its ConfigMap generator from services/hermes/kustomization.yaml; reconcile hermes to remove the endpoint.
  2. Remove suite_decision, the three suite targets/routes, and the corresponding Switchyard revision update. Reconcile; keep unrelated routes unchanged.
  3. Remove the suite Vault seed/bootstrap resources and the suite role/policy additions if no longer needed. Remove the created Vault role/policy and secret through the normal Vault administrative workflow; removing a completed Job does not revoke them.
  4. Deleting the metadata PVC removes retry protection, so do it only after all jobs are terminal and no client can submit. No provider job is replayed automatically.

The existing /local-model endpoint and token do not require rollback changes.