613 lines
37 KiB
Markdown
613 lines
37 KiB
Markdown
# Suite planning API
|
|
|
|
This optional API groups one complete campaign/suite into implementation families.
|
|
The existing `/local-model` API, GPU allocation, serving model, and limits are unchanged.
|
|
Deployment uses Flux; the application and workbook stay on the laptop.
|
|
|
|
## Current multi-pass policy
|
|
|
|
New jobs use [the multi-pass implementation policy](hermes_suite_multipass.md):
|
|
two independent full-suite proposals, reconciliation, up to two bounded semantic
|
|
reviews, and a programmatic five-case final cap. Names are at most 64 characters.
|
|
The endpoint and request fields are unchanged. Optional top-level `review_summary`
|
|
is returned only with the authorized result. Configuration remains `suite-v6-20260929`;
|
|
policy, prompt, and execution revisions identify the changed grouping behavior.
|
|
The prior single-pass acceptance results below are historical and do not measure
|
|
the new workflow. The new process shares the user-approved 1800-second deadline
|
|
and USD 30 CLI estimate guard across all invocations. This is not subscription
|
|
billing. Progress updates every five seconds during a CLI pass. The separate local-model endpoint is unchanged;
|
|
the 8K local route cannot admit the new multi-pass reconciliation schema.
|
|
|
|
## Earlier CLI failure diagnostics update (historical)
|
|
|
|
Execution revision `claude-diagnostics-turns-v1-20260929` adds content-free
|
|
`cli_diagnostics` and failure-stage details, including exit/signal, final subtype,
|
|
reported turns, observed stop reason, token counts, and structured-output location.
|
|
Unknown measurements remain null. Source, generated descriptions, raw provider
|
|
errors, and transcripts are not retained. The normal schema and exact membership
|
|
checks still gate completion. Failed validation now preserves operational usage.
|
|
|
|
The CLI turn ceiling is six, with the same model, medium effort, 64,000 per-call
|
|
output ceiling, maximum 900-second job deadline, and USD 5 CLI estimated-cost cap.
|
|
Preflight reserves all six possible outputs against context. This is headroom
|
|
within one fresh job, not an automatic job retry or a fallback. The historical
|
|
three-turn ceiling was exhausted by a deterministic synthetic structured-output
|
|
repair test; six turns completed that same test. This does not establish the
|
|
cause of a historical failure whose terminal event was deleted.
|
|
|
|
The HTTP configuration revision and implementation-proximity prompt revision/hash
|
|
are unchanged. `execution_revision` records the new behavior in capability,
|
|
preflight, and job metadata. Use a new idempotency key only when intentionally
|
|
starting a new attempt; the failed historical job is never automatically rerun.
|
|
See [the diagnosis and verification record](hermes_suite_failure_diagnosis_20260929.md).
|
|
Rollback this update by reverting its Git commit and reconciling the Hermes Flux
|
|
Kustomization; do not revert the earlier prompt commit. Wait for no active jobs
|
|
before rollout because results are held in memory.
|
|
|
|
## Implementation-proximity prompt update
|
|
|
|
New jobs use prompt revision `implementation-proximity-v2-20260929`. The HTTP
|
|
configuration revision remains `suite-v6-20260929`; request and result schemas,
|
|
models, reasoning effort, token limits, authentication, and routing policy are
|
|
unchanged. `prompt_revision` and `prompt_sha256` are additive provenance in health,
|
|
capabilities, preflight selection, and persisted job/status/result metadata.
|
|
Historical jobs keep their original metadata; they are not relabeled on replay.
|
|
|
|
The objective is the additional test-development effort after implementing a
|
|
representative case. Success criteria provide the primary evidence of actions,
|
|
observations, measurements, and assertions. Description supplies behavior and
|
|
operation; preconditions and other material constraints supply setup/state;
|
|
case type is supporting context. There is no keyword or word-count weighting.
|
|
The prompt distinguishes inexpensive input/assertion variations from new test
|
|
machinery, checks entire-family coherence, reviews singletons, preserves uncertainty
|
|
in generalized values, and confines names/descriptions to assigned test objectives.
|
|
|
|
The same system prompt is supplied to the actual grouping invocation and remains
|
|
applicable during structured-output turns or format repair. There is no separate
|
|
repair model. Cases enter a fresh isolated session without any previous grouping
|
|
as a target answer. The expanded prompt is included in the existing conservative
|
|
preflight accounting; no context or output setting has been increased.
|
|
|
|
For a fresh comparison, keep the same case request but use a NEW `Idempotency-Key`.
|
|
Reusing the old key returns the old job instead of invoking the revised prompt.
|
|
Include `prompt_revision` or `prompt_sha256` in the laptop's result-cache identity.
|
|
The synthetic check fixture is
|
|
`testing/fixtures/suite_implementation_proximity.json`; only its `request` object is
|
|
sent, never its independent expected-family mapping. No real roster job is launched
|
|
by this deployment or its acceptance checks.
|
|
|
|
Verified on the deployed endpoint: the 14-case adversarial synthetic suite completed
|
|
in 21.059 seconds with six implementation families, exact alias coverage, and zero
|
|
incorrect or missed merge pairs against its independent expected map. Identical
|
|
broad descriptions did not merge different machinery. Nominal/negative-result and
|
|
extra-assertion variants shared their existing observation code; an underspecified
|
|
case retained its uncertainty. Every observed native CLI request in the loopback
|
|
structured-output transport check contained the complete revised system prompt.
|
|
The focused unit suite passes 36 tests. These checks establish the synthetic behavior,
|
|
not the quality of a real-suite rerun. See the
|
|
[prompt-update acceptance evidence](evidence/hermes_suite_prompt_v2_20260929.json).
|
|
|
|
## Acceptance results, 2026-09-29
|
|
|
|
The final runs use `suite-v6-20260929`, deployed commit `299b9ed1`, including
|
|
case-type fields. These supersede the preliminary runs without that field.
|
|
All tests used synthetic content. HTTPS tests ran from `titan-jh`, source
|
|
`192.168.22.8`, directly on its LAN interface to `192.168.22.50:443`, with hostname
|
|
certificate verification and proxies disabled. The user's WSL route has not been
|
|
tested; run the health command below from the laptop before uploading anything.
|
|
|
|
| Cases | Actual request bytes | CLI input/output tokens | Client wall seconds | Families / singletons | Exact coverage | Pattern match |
|
|
| --- | ---: | --- | ---: | --- | --- | --- |
|
|
| 14 | 9,847 | 4,185 / 1,250 | 16.947 | 9 / 4 | Yes | Exact |
|
|
| 75 | 55,984 | 19,145 / 2,336 | 25.755 | 9 / 3 | Yes | Exact |
|
|
| 363 | 273,761 | 89,753 / 8,568 | 80.071 | 9 / 3 | Yes | Exact |
|
|
|
|
Request sizes are the UTF-8 compact JSON bytes actually submitted, excluding HTTP
|
|
headers. Token counts are CLI-reported aggregate usage across each fresh job's
|
|
two structured-output turns; they are not an exact standalone input-token count.
|
|
All three final jobs completed on their first attempt with no observed retries or
|
|
rate-limit errors. Two CLI turns are part of producing the structured result, not
|
|
automatic job retries. CLI cost estimates were USD 0.052175, 0.154125, and 0.662965;
|
|
these are estimates, not verified subscription charges. Remaining account quota
|
|
is not available through this endpoint and may be consumed by other account users.
|
|
|
|
The actual terminal `modelUsage` reports `claude-opus-4-8[1m]`, with
|
|
`canonicalModel=claude-opus-4-8`, `provider=firstParty`, context 1,000,000 and maximum
|
|
output 64,000. The invocation requests `claude-opus-4-8[1m]`; these results also
|
|
independently verify the returned model metadata, rather than trusting the request
|
|
flag or the internal `atlas/planning/claude` route label. Provider physical hardware
|
|
and inference region are unknown. This verifies the tested suite sizes, not the
|
|
entire advertised million-token context.
|
|
|
|
The 363-case fixture has descriptions of 48-245 bytes (median 234), preconditions
|
|
67-143 (median 136), success criteria 73-135 (median 132), and case types 7-15 bytes.
|
|
Most cases have multi-sentence fields plus operating conditions, verification
|
|
method, and a generic reference. Six machinery patterns are interleaved throughout
|
|
the input. Watchdog terminology occurs in three different machinery patterns. A
|
|
distant pair has identical complete case text but different aliases. Distinctive
|
|
thermal, acoustic, and build cases occur at the beginning, middle, and end.
|
|
|
|
Against the independent synthetic implementation map, all sizes had pair precision
|
|
and recall 1.0, zero incorrect merge pairs, and zero missed merge pairs. Manual
|
|
review of all 27 family descriptions found no objectives attributed to another
|
|
family. Four singletons in the 14-case fixture and three in the larger fixtures
|
|
are expected. These are clear synthetic patterns, not evidence of quality on the
|
|
user's more ambiguous real suite. No real case quality is claimed.
|
|
|
|
Evidence for complete input handling goes beyond IDs:
|
|
|
|
- Unit checks compare every serialized case and field with the original fixture.
|
|
- The deployed native CLI was also run against a loopback mock provider, with the
|
|
exact production command and full 14/75/363 inputs. The captured requests contained
|
|
the complete source string, byte for byte, including every field. The largest
|
|
first provider request was 287,914 bytes. This transport check used no hosted model.
|
|
- Live jobs returned the expected distinct beginning/middle/end patterns and
|
|
reunited interleaved cases. Initialization allowed only `StructuredOutput`, with
|
|
no MCP, plugins, or file-reading tools. Each job used fresh HOME/config and stdin.
|
|
- No CLI compaction event or incomplete-output signal was observed. Compaction was
|
|
disabled, and the gateway performs no summarization or selective file reading.
|
|
|
|
Those observations support complete transmission and the tested grouping behavior.
|
|
They cannot inspect provider internals or prove that a model attended equally to
|
|
every word. The response's false compaction/truncation values describe gateway/CLI
|
|
observations, not a provider-side attestation. An exact tokenizer is unavailable;
|
|
preflight's byte bound and output reservation are documented below.
|
|
|
|
The local-only two-case parser example also completed in 3.535 seconds on the RTX
|
|
backend with exact coverage and no hosted destination. An earlier local result
|
|
incorrectly appended descriptions to aliases and was rejected; the new adapter
|
|
constrains the local output schema to supplied aliases. The existing local-model
|
|
endpoint itself is unchanged.
|
|
|
|
Verification also covered authentication failures, permission escalation, unknown
|
|
providers, synthetic-versus-operational scope, local-only capacity rejection,
|
|
management-path rejection, cross-client ownership, idempotent replay, oversized
|
|
HTTP 413, and content-free routine worker logs. Mocked backend failures verified
|
|
that no local failure launches Claude. No shared backend was disrupted. The focused
|
|
unit suite passes 35 tests. Service renders, client dry-runs, and Flux diffs passed.
|
|
The installed kubectl rejects the repository's combined `--server-side --dry-run=client`
|
|
syntax; the supported client dry-run was used instead.
|
|
|
|
Full evidence and successful normalized envelopes:
|
|
|
|
- [Acceptance measurements and policy checks](evidence/hermes_suite_acceptance_20260929.json)
|
|
- [Actual complete 14-case Claude response](evidence/hermes_suite_success_14.json)
|
|
- [Actual complete two-case local response](evidence/hermes_suite_local_success.json)
|
|
- [All final synthetic family results](evidence/hermes_suite_results_synthetic.json)
|
|
|
|
Ready: synthetic client testing, direct complete-suite execution at all three
|
|
tested sizes, and the separately approved Claude credential. Remaining client work:
|
|
WSL connectivity, explicit client-envelope adaptation, and a controlled real pilot
|
|
after reviewing this handoff. Codex capacity verification and provider/account
|
|
retention or training controls remain unresolved. The user's account-specific
|
|
approval is recorded; broader organizational approval or release classification
|
|
has not been established by this infrastructure work.
|
|
|
|
## Connection and credentials
|
|
|
|
```
|
|
LAN_HOST: worker.bstein.dev
|
|
LAN_IP: 192.168.22.50
|
|
PORT: 443
|
|
BASE_URL: https://worker.bstein.dev/suite-planning
|
|
HEALTH_URL: https://worker.bstein.dev/suite-planning/healthz
|
|
CAPABILITIES_URL: https://worker.bstein.dev/suite-planning/v1/capabilities
|
|
PREFLIGHT_URL: https://worker.bstein.dev/suite-planning/v1/preflight
|
|
SUBMIT_URL: https://worker.bstein.dev/suite-planning/v1/jobs
|
|
AUTHENTICATION_HEADER: Authorization: Bearer <scoped token>
|
|
VAULT_PATH: kv/atlas/hermes/suite-planning-api
|
|
LOCAL_ONLY_FIELD: token
|
|
SYNTHETIC_EXTERNAL_FIELD: synthetic_token
|
|
APPROVED_OPERATIONAL_FIELD: operational_token
|
|
INITIAL_CONCURRENCY: 1; busy submissions return 429; no waiting queue
|
|
JOB_TIMEOUT: up to 1800 seconds
|
|
HTTP_TIMEOUT: client 45 seconds; submit/status do not wait for inference
|
|
TLS: existing worker.bstein.dev certificate; normal trusted CA verification
|
|
```
|
|
|
|
Retrieve the appropriate field privately from
|
|
[Vault](https://vault.bstein.dev/ui/vault/secrets/kv/show/atlas/hermes/suite-planning-api).
|
|
Neither token grants provider, Vault, Kubernetes, shell, or ClickUp access.
|
|
The old local-model token remains at `kv/atlas/hermes/model-gate-lan-api`.
|
|
Tokens are distinct; an old local-model token does not authenticate this API.
|
|
|
|
The credentials have different server-enforced scopes:
|
|
|
|
| Vault field | Local inference | Hosted inference |
|
|
| --- | --- | --- |
|
|
| `token` | Allowed within local capacity | Prohibited |
|
|
| `synthetic_token` | Allowed within local capacity | Exact API fixtures only; full content hash checked |
|
|
| `operational_token` | Allowed within local capacity | Generalized CASE records, Claude only, under the explicit account approval |
|
|
|
|
`synthetic_token` has not been broadened. Changing a fixture or sending real
|
|
generalized records with that token returns `external_data_not_approved` even
|
|
with `allow_external=true`. The operational credential is separate, grants only
|
|
`claude`, and requires the server's `PLANNING_GENERALIZED_CLAUDE_APPROVED=true`.
|
|
The user explicitly approved the existing Hermes first-party Claude OAuth account
|
|
for generalized CASE records on 2026-09-29. No real records were used in acceptance.
|
|
Original identifiers, uncensored text, and alias mappings remain on the laptop.
|
|
The server cannot determine whether arbitrary prose is adequately generalized;
|
|
that remains the application's and organization's responsibility. Account-specific
|
|
training and retention settings remain unverified. A request flag cannot grant
|
|
credential permission or override data scope.
|
|
|
|
## Request contract
|
|
|
|
Submit with `POST https://worker.bstein.dev/suite-planning/v1/jobs`.
|
|
The [request JSON Schema](contracts/suite_planning_request.schema.json) describes
|
|
the accepted wire object. The [result JSON Schema](contracts/suite_planning_result.schema.json)
|
|
describes `result`, not the whole job envelope. Additional server checks enforce
|
|
UTF-8 byte limits, unique aliases, ownership, permissions, and exact result coverage.
|
|
|
|
```
|
|
{
|
|
"campaign": "SYNTHETIC",
|
|
"suite": "PARSER",
|
|
"cases": [
|
|
{"alias": "CASE-1", "description": "Parse a valid configuration and assert accepted fields."},
|
|
{"alias": "CASE-2", "description": "Parse an invalid configuration and assert diagnostic fields."}
|
|
],
|
|
"routing": {
|
|
"allow_external": false,
|
|
"allowed_external_providers": []
|
|
},
|
|
"execution": {
|
|
"strategy": "whole_suite",
|
|
"max_seconds": 1800,
|
|
"max_cost_usd": 30
|
|
}
|
|
}
|
|
```
|
|
|
|
`campaign`, `suite`, and nonempty `cases` are required. Each case requires a
|
|
unique `CASE-*` alias and a description. Optional case fields are
|
|
`success_criteria`, `preconditions`, `operating_condition`, `case_type`,
|
|
`verification_method`, `target`, `swci`, `verifies`, `functional_area`,
|
|
`functional_group`, and `functional_group_name`. Values are strings or null.
|
|
Optional per-case campaign/suite values must exactly match the job's ownership.
|
|
Unknown fields are rejected. `[reference]` is ordinary source text, never an alias.
|
|
Identical descriptions with different aliases remain distinct cases.
|
|
|
|
Missing routing policy means local-only. External provider names are `claude` and
|
|
`codex`. The requested list must be contained in the credential's permissions;
|
|
unknown providers, malformed booleans, duplicate names, and permission escalation
|
|
are rejected before routing. `allow_external=false` cannot include providers.
|
|
`whole_suite` is the only strategy. No batching, summarization, or reconciliation
|
|
is silently substituted. `max_seconds` is 10-1800. The Claude CLI guard is at most
|
|
USD 30 in its estimated usage accounting; this is not a verified subscription
|
|
billing ceiling and can overshoot within a single provider call.
|
|
|
|
The gateway deterministically selects a fitting permitted backend, preferring
|
|
local capacity. It asks the fixed Switchyard route to confirm that binding, then
|
|
performs one inference attempt. Switchyard receives only `select`, never case text.
|
|
Its new routes each have exactly one target, zero client retries, and no classifier:
|
|
`atlas/planning/local`, `atlas/planning/claude`, `atlas/planning/codex`.
|
|
These are decision routes; clients use the HTTPS job API to obtain inference results.
|
|
Existing general Hermes routes retain their old behavior and are not this API.
|
|
|
|
## Job lifecycle and normalized response
|
|
|
|
- `GET /healthz`: authenticated service readiness.
|
|
- `GET /v1/capabilities`: permissions, models, configuration revision, and limits.
|
|
- `POST /v1/preflight`: validate the identical proposed job without generation.
|
|
- `POST /v1/jobs`: require `Idempotency-Key`, return 202 and a job ID.
|
|
- `GET /v1/jobs/<id>`: status and metadata.
|
|
- `GET /v1/jobs/<id>/result`: status plus the normalized result when complete.
|
|
- `DELETE /v1/jobs/<id>`: request cancellation.
|
|
- `GET /v1/synthetic/14`, `/75`, `/363`: immutable synthetic fixtures.
|
|
|
|
Retry a submission with the SAME key and SAME body. It returns the original job,
|
|
not a second provider attempt. A changed body with that key returns 409. Keys are
|
|
scoped to the authenticated credential and retained seven days. The maximum is
|
|
10,000 metadata records; excess submissions fail rather than evicting live keys.
|
|
No provider retry or fallback occurs after an error. An operator may intentionally
|
|
start a new attempt with a new key after diagnosing the previous failure.
|
|
|
|
Results remain in memory for at most one hour (cleanup every 30 seconds), with a
|
|
128-result bound. Durable local SQLite holds only operational metadata and hashes,
|
|
not prompts or results. Restarts mark unfinished jobs `interrupted_no_retry` and
|
|
never resubmit them. Previously completed results are lost on restart; retrieval
|
|
returns 410. A cancellation holds the slot until the underlying attempt ends.
|
|
Claude cancellation kills the isolated process group. Local generation cannot be
|
|
cancelled at the existing API; its eventual result is discarded. Provider-side
|
|
cancellation and exact billing after disconnection are not guaranteed.
|
|
|
|
Example completed envelope (illustrative synthetic values):
|
|
|
|
```
|
|
{
|
|
"job_id": "0123456789abcdef0123456789abcdef",
|
|
"status": "completed",
|
|
"configuration_revision": "suite-v6-20260929",
|
|
"routing": {"allow_external": true, "allowed_external_providers": ["claude"]},
|
|
"selection": {"provider": "claude", "backend": "claude-code-2.1.226", "model": "claude-opus-4-8"},
|
|
"model": "claude-opus-4-8",
|
|
"attempted_destinations": ["switchyard:atlas/planning/claude", "claude:claude-opus-4-8"],
|
|
"compaction": false,
|
|
"truncation": false,
|
|
"usage": {"input_tokens": 1000, "output_tokens": 100},
|
|
"wall_seconds": 12.3,
|
|
"result": {
|
|
"groups": [{"name": "Configuration parser fixtures", "description": "Parameterize configuration text and assert parsed fields or diagnostics.", "members": ["CASE-1", "CASE-2"]}]
|
|
}
|
|
}
|
|
```
|
|
|
|
The real envelope also includes input bytes, input hash, reservation method,
|
|
model usage, timing, and estimated cost where available. Unknown values are null.
|
|
Failed jobs retain status and fixed error codes but no provider stderr. Status
|
|
remains HTTP 200 for an existing failed job; inspect `status` and `error.code`.
|
|
Invalid submissions use 400/401/403/409/413/415/422/429 as appropriate.
|
|
Application errors have the shape `{"error":{"code":"provider_forbidden","details":{}}}`.
|
|
Errors raised by Traefik, including its 413 body rejection, can be plain text.
|
|
Check HTTP status before parsing JSON. With a large upload, a client can see a
|
|
connection reset while still writing after ingress rejects the request; check
|
|
the advertised byte cap locally before upload. `Expect: 100-continue` produced
|
|
an explicit HTTP 413 in the LAN curl acceptance check.
|
|
Failures distinguish routing, backend availability, provider authentication,
|
|
timeouts, invalid JSON, incomplete generation, changed model/capabilities,
|
|
compaction, and invalid assignments. Failed output is never a completed result.
|
|
|
|
The result is a JSON object. The underlying Claude stream/result envelope and
|
|
Ollama's JSON-in-a-string representation are normalized on the server.
|
|
Every alias must occur exactly once. Duplicate, missing, invented, or foreign
|
|
aliases reject the entire result. Names must be unique within the returned suite.
|
|
Coverage validation is independent of the model and does not prove semantic quality.
|
|
|
|
## Mapping the prepared client job
|
|
|
|
The prepared `job.json` is not the wire request. Unknown top-level and case fields
|
|
are rejected; this service does not silently unpack the client envelope.
|
|
|
|
| Prepared client field | Service mapping |
|
|
| --- | --- |
|
|
| `job_id` | Use a stable `Idempotency-Key` header, 8-128 ASCII letters/digits/underscore/hyphen. The service returns its own `job_id`; retain both locally. Hash an incompatible client ID once rather than generating a new retry key. |
|
|
| `client_spec_version` | Keep locally and validate compatibility with `/v1/capabilities`; not a submitted field. |
|
|
| `task` | This endpoint implements only whole-suite implementation-family planning. Validate that task locally; not a submitted field. |
|
|
| `routing` | Send only `allow_external` and `allowed_external_providers`. |
|
|
| `execution` | Send only `strategy`, `max_seconds`, `max_cost_usd`. |
|
|
| `instructions` | Not supported as custom instructions. The server uses the documented fixed implementation-family prompt. If custom instructions are required, this contract does not yet support that request; do not silently discard them. |
|
|
| `input` | Lift its one complete suite into top-level `campaign`, `suite`, and `cases`. Keep these identity strings generalized if needed. |
|
|
| `output_schema` | Not accepted. The service pins groups with `name`, `description`, `members`; compare the client schema and explicitly adapt or reject incompatible expectations. |
|
|
| case `case_id` | Rename to `alias`; must already be the unique CASE alias, never the original identifier. |
|
|
| case `name` | No separate wire field. Preserve it explicitly in `description`, for example `Name: ...\nDescription: ...`, if the model needs it. Do not silently lose meaningful name content. |
|
|
| case `description` | `description`, optionally combined with the name as above. |
|
|
| case `preconditions`, `case_type`, `success_criteria` | Same-named fields, strings or null. |
|
|
|
|
The application must select one exact campaign/suite before constructing this
|
|
object. It must retain the complete original case records and alias mapping
|
|
locally. Group `members` contain aliases only. Server-generated family names and
|
|
descriptions do not become automatically permitted ClickUp export content.
|
|
|
|
Only small client functions are needed: build the wire projection, validate the
|
|
fixed task/schema contract, preflight, submit with a stable key, poll, save the
|
|
result, and independently audit aliases. No importer replacement is needed.
|
|
|
|
## Installed route capabilities
|
|
|
|
| Property | Codex | Claude | Existing local API |
|
|
| --- | --- | --- | --- |
|
|
| Installed client | Codex CLI 0.154.0 | Claude Code 2.1.226, pinned binary SHA-256 | Ollama 0.13.5 |
|
|
| Existing broker mode | Direct subscription Responses transport, not `codex exec` | Native CLI print mode behind a wrapper | Native `/api/generate` |
|
|
| New planner backend | Disabled pending effective output-budget verification | Fresh native CLI process, wrapper bypass avoided | Unchanged model-gate LAN listener |
|
|
| Auth | ChatGPT Pro claim verified locally | Existing first-party OAuth setup token; associated credential metadata says Max 20x | Scoped local bearer |
|
|
| Visible catalog | `gpt-6-astra`, `gpt-5.6-sol`, `gpt-5.6-terra`, `gpt-5.6-luna`, `gpt-5.5` | Fable 5, Opus 5 with 1M option, Sonnet 5, Haiku 4.5 | Pinned Qwen 2.5 14B Q4 |
|
|
| Exact selected ID | `gpt-6-astra` reserved, not enabled | `claude-opus-4-8` | `qwen2.5:14b-instruct-q4_0` |
|
|
| Context evidence | Local account model cache 272,000, 95% effective = 258,400 | Actual CLI response: 1,000,000 | Verified serving configuration: 8,192 |
|
|
| Output control | Existing broker removes public token-limit fields; effective maximum unverified | Actual CLI reports 64,000; environment pins that ceiling | 2,048 |
|
|
| New planner concurrency | None | One across the whole planning service | Shares the existing serialized GPU backend |
|
|
| Hardware | Provider hosted; broker on titan-22 | Provider hosted; CLI on titan-22 | RTX 3080 10GB on titan-24 |
|
|
|
|
The Claude live catalog resolves `claude-fable-5[1m]` to `claude-fable-5`,
|
|
`default`/`opus[1m]` to `claude-opus-5[1m]`, `sonnet` to `claude-sonnet-5`,
|
|
and `haiku` to `claude-haiku-4-5-20251001`. Catalog availability is not proof
|
|
that every model has been exercised. Fable 5 suite requests reported `claude-opus-4-8` at runtime and were rejected.
|
|
The initial working route therefore explicitly pins the observed Opus 4.8 model;
|
|
the advertised Fable name is not verified for full-suite execution.
|
|
Only the tiny Opus/Fable probes and recorded
|
|
planner acceptance jobs were executed. No Claude agent wrote infrastructure code.
|
|
Subscription quotas and account retention settings are not verified by model discovery.
|
|
|
|
The legacy Codex and Claude wrapper scripts enable permission bypass; the Claude
|
|
wrapper also enables automatic compaction and shared settings. Neither is used by
|
|
the new worker. Codex supports `exec --ephemeral --output-schema --json` and stdin,
|
|
but that separate CLI execution mode has not been approved as a whole-suite backend.
|
|
The direct Codex broker uses `store=false`, streaming Responses, and a 900-second
|
|
read timeout. That flag does not establish provider Zero Data Retention.
|
|
|
|
Claude invocation, with the schema supplied by the server:
|
|
|
|
```
|
|
/opt/cli/claude -p --output-format stream-json --verbose \
|
|
--no-session-persistence --safe-mode --tools '' \
|
|
--strict-mcp-config --mcp-config '{"mcpServers":{}}' \
|
|
--setting-sources '' --disable-slash-commands --permission-mode dontAsk \
|
|
--no-chrome --model 'claude-opus-4-8[1m]' --effort medium \
|
|
--max-budget-usd 5 --max-turns 6 \
|
|
--system-prompt '<fixed server prompt>' --json-schema '<fixed schema>'
|
|
```
|
|
|
|
The complete suite is an isolated input file connected to stdin, never an argv
|
|
string or interpolated shell command. The process receives only a short allowlisted
|
|
environment, its OAuth token, and a fresh temporary HOME/config directory.
|
|
`DISABLE_COMPACT`, `DISABLE_AUTO_COMPACT`, `DISABLE_TELEMETRY`,
|
|
`DISABLE_ERROR_REPORTING`, `DISABLE_PROMPT_CACHING`,
|
|
`CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC`, `CLAUDE_CODE_DISABLE_AUTO_MEMORY`,
|
|
and `CLAUDE_CODE_SKIP_PROMPT_HISTORY` are enabled; retries and updates are disabled.
|
|
The actual initialization event must report only `StructuredOutput`, no MCP servers
|
|
or plugins. The worker rejects compaction events, unexpected models, incomplete
|
|
output, changed limits, and missing terminal results.
|
|
|
|
These are client controls, not provider ZDR. Account-specific training/retention
|
|
settings remain unverified. See [Claude data usage](https://code.claude.com/docs/en/data-usage),
|
|
[Claude CLI flags](https://code.claude.com/docs/en/cli-reference),
|
|
[Claude environment controls](https://code.claude.com/docs/en/env-vars), and
|
|
[Codex noninteractive behavior](https://learn.chatgpt.com/docs/non-interactive-mode).
|
|
|
|
## Capacity and transport
|
|
|
|
The new API accepts at most 1 MiB and 400 cases per request, with at most 32 KiB
|
|
per field. The existing local endpoint still accepts only 128 KiB and its original
|
|
context/output limits. Claude subprocess stdout is bounded to 4 MiB; normalized
|
|
results to 1 MiB. Final names are at most 64 characters, descriptions at most 240, and groups at most five cases.
|
|
|
|
Preflight includes the entire serialized suite, fixed instructions, schema, reserved
|
|
harness overhead, and output space. There is no accurate account-specific tokenizer.
|
|
The input check uses UTF-8 byte count plus 8,192 reserved harness tokens, with the
|
|
64,000 output ceiling reserved for each of at most six turns against context. This is a conservative bound,
|
|
not a measured token count. The output reservation is an explicit estimate allowing
|
|
a family per case and 8,192 reasoning tokens; it is not a verified bound on arbitrary
|
|
generated wording. Medium adaptive reasoning can consume output budget. Incomplete
|
|
output fails explicitly; the server never drops cases to turn it into success.
|
|
|
|
Small local requests reserve 1,024 overhead tokens and 2,048 output tokens within
|
|
8,192 context. The realistic 14/75/363 fixtures exceed this conservative local
|
|
whole-suite budget and must use the approved hosted path or fail preflight.
|
|
|
|
All CLI invocations share at most 1800 seconds per suite job; cancellation and timeouts terminate its
|
|
process group. Async HTTP calls finish promptly, with a 30-second body-read timeout,
|
|
40-second ingress response-header timeout, and recommended 45-second client timeout.
|
|
Switchyard's ten-minute internal request limit carries only a tiny immediate routing
|
|
decision, not the long inference job or its large body.
|
|
|
|
If a real complete suite cannot fit, the current API returns capacity/unsupported.
|
|
A future strategy would extract profiles with preserved aliases, generate overlapping
|
|
implementation candidates across all batches, compare candidates across batch
|
|
boundaries, then reconcile and audit the entire suite. Independent batch grouping
|
|
followed by concatenation is not supported. Transport chunking would not add context.
|
|
|
|
## WSL examples
|
|
|
|
Supply only the scoped credential privately. Do not add `-k` or `-L`.
|
|
|
|
```
|
|
read -rsp 'Suite API token: ' SUITE_PLANNING_TOKEN
|
|
export SUITE_PLANNING_TOKEN
|
|
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
|
|
--connect-timeout 10 --max-time 45 --fail-with-body \
|
|
-H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \
|
|
https://worker.bstein.dev/suite-planning/healthz
|
|
```
|
|
|
|
Local-only synthetic grouping (any scoped credential):
|
|
|
|
```
|
|
cat > local-synthetic.json <<'JSON'
|
|
{"campaign":"SYNTHETIC","suite":"PARSER","cases":[{"alias":"CASE-1","description":"Parse valid configuration text and check returned fields.","preconditions":"An isolated parser fixture is reset before each case.","case_type":"nominal","success_criteria":"Returned fields match the supplied configuration values."},{"alias":"CASE-2","description":"Parse invalid configuration text and check returned diagnostics.","preconditions":"An isolated parser fixture is reset before each case.","case_type":"fault injection","success_criteria":"The malformed input produces the specified diagnostic fields."}],"routing":{"allow_external":false,"allowed_external_providers":[]},"execution":{"strategy":"whole_suite"}}
|
|
JSON
|
|
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
|
|
--connect-timeout 10 --max-time 45 --fail-with-body \
|
|
-H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \
|
|
-H 'Content-Type: application/json' -H 'Idempotency-Key: local-parser-pilot-001' \
|
|
--data-binary @local-synthetic.json \
|
|
https://worker.bstein.dev/suite-planning/v1/jobs
|
|
```
|
|
|
|
Externally allowed, complete 363-case synthetic grouping (use `synthetic_token`):
|
|
|
|
```
|
|
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
|
|
--connect-timeout 10 --max-time 45 --fail-with-body \
|
|
-H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \
|
|
https://worker.bstein.dev/suite-planning/v1/synthetic/363 > synthetic-suite.json
|
|
python3 - <<'PY'
|
|
import json
|
|
with open('synthetic-suite.json') as stream:
|
|
request = json.load(stream)
|
|
request['routing'] = {'allow_external': True, 'allowed_external_providers': ['claude']}
|
|
request['execution'] = {'strategy': 'whole_suite', 'max_seconds': 1800, 'max_cost_usd': 30}
|
|
with open('synthetic-request.json', 'w') as stream:
|
|
json.dump(request, stream)
|
|
PY
|
|
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
|
|
--connect-timeout 10 --max-time 45 --fail-with-body \
|
|
-H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \
|
|
-H 'Content-Type: application/json' --data-binary @synthetic-request.json \
|
|
https://worker.bstein.dev/suite-planning/v1/preflight
|
|
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
|
|
--connect-timeout 10 --max-time 45 --fail-with-body \
|
|
-H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \
|
|
-H 'Content-Type: application/json' -H 'Idempotency-Key: synthetic-363-pilot-001' \
|
|
--data-binary @synthetic-request.json \
|
|
https://worker.bstein.dev/suite-planning/v1/jobs > submitted-job.json
|
|
JOB_ID=$(python3 -c 'import json; print(json.load(open("submitted-job.json"))["job_id"])')
|
|
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
|
|
--connect-timeout 10 --max-time 45 --fail-with-body \
|
|
-H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \
|
|
"https://worker.bstein.dev/suite-planning/v1/jobs/$JOB_ID/result"
|
|
```
|
|
|
|
Repeat the last GET until `status` is `completed`, `failed`, or `cancelled`.
|
|
Save the result locally before retention expires. Keep the original key for network
|
|
retries. Use a new key only when intentionally starting another provider job.
|
|
|
|
Capability, status-only, and cancellation requests:
|
|
|
|
```
|
|
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
|
|
--connect-timeout 10 --max-time 45 --fail-with-body \
|
|
-H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \
|
|
https://worker.bstein.dev/suite-planning/v1/capabilities
|
|
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
|
|
--connect-timeout 10 --max-time 45 --fail-with-body \
|
|
-H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \
|
|
"https://worker.bstein.dev/suite-planning/v1/jobs/$JOB_ID"
|
|
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
|
|
--connect-timeout 10 --max-time 45 --fail-with-body -X DELETE \
|
|
-H "Authorization: Bearer $SUITE_PLANNING_TOKEN" \
|
|
"https://worker.bstein.dev/suite-planning/v1/jobs/$JOB_ID"
|
|
```
|
|
|
|
For an approved generalized job, obtain `operational_token`, construct the same
|
|
wire object locally, and request only `allowed_external_providers:["claude"]`.
|
|
The credential name does not go in the HTTP request; only its secret value is
|
|
used as the bearer. Neither a provider credential nor a Vault token belongs in
|
|
this header.
|
|
|
|
The standalone `scripts/ops/hermes_suite_probe.py` uses only Python urllib and
|
|
the same API credential. Its HTTPS handler implements the equivalent of curl
|
|
`--resolve`, retaining certificate verification and bypassing proxies. It retrieves
|
|
synthetic fixtures, submits, checks idempotency, polls, and validates exact coverage.
|
|
It does not integrate with or replace the importer. Proposed application additions
|
|
are `get_capabilities()`, `preflight_suite()`, `submit_suite(idempotency_key)`,
|
|
`get_job_result()`, and `cancel_job()`. Preserve the application's local selection,
|
|
alias mapping, ownership and coverage audits, and export policy.
|
|
|
|
## Deployment, logging, and rollback
|
|
|
|
The planner runs on titan-22 alongside the existing RWO CLI tools volume. Only an
|
|
init container reads its tools subdirectory and copies the hash-pinned native
|
|
binary. The serving container has no agent home, cluster token, repository, ClickUp
|
|
credentials, provider API keys, or arbitrary command endpoint. Its requests are
|
|
100m CPU/512 MiB and limits are 2 CPU/2 GiB. No shared GPU workload is changed.
|
|
|
|
Input/output/config/debug storage is per-job tmpfs, removed on success, failure,
|
|
timeout, and cancellation. Pod termination also clears it. The observed CLI writes
|
|
configuration and backup files despite no session persistence; these share that
|
|
temporary directory. The durable metadata PVC uses local-path and is outside the
|
|
Longhorn backup path. Ordinary worker logs contain job ID, selected provider/model,
|
|
status and duration only. Provider stderr is discarded. The pod is excluded from
|
|
Fluent Bit; no prompt is sent to the existing Switchyard routing logs. TLS ingress
|
|
does not log bodies. Central archival deletion is not claimed or required because
|
|
the new path does not send content there; account-side retention remains separate.
|
|
|
|
Rollback through Git/Flux:
|
|
|
|
To revoke only generalized-data permission, set
|
|
`PLANNING_GENERALIZED_CLAUDE_APPROVED=false` and reconcile `hermes`.
|
|
Revoke/rotate `operational_token` through Vault if the credential itself must be
|
|
invalidated, then restart through a Flux-tracked deployment revision so the
|
|
pre-populated credential mount refreshes. Keep both older token fields intact.
|
|
|
|
1. Remove the three suite-planner resource entries and its ConfigMap generator from
|
|
`services/hermes/kustomization.yaml`; reconcile `hermes` to remove the endpoint.
|
|
2. Remove `suite_decision`, the three suite targets/routes, and the corresponding
|
|
Switchyard revision update. Reconcile; keep unrelated routes unchanged.
|
|
3. Remove the suite Vault seed/bootstrap resources and the suite role/policy additions
|
|
if no longer needed. Remove the created Vault role/policy and secret through the
|
|
normal Vault administrative workflow; removing a completed Job does not revoke them.
|
|
4. Deleting the metadata PVC removes retry protection, so do it only after all jobs
|
|
are terminal and no client can submit. No provider job is replayed automatically.
|
|
|
|
The existing `/local-model` endpoint and token do not require rollback changes.
|