atlas-iac/docs/hermes_suite_job_budget.md

137 lines
7.1 KiB
Markdown

# Suite job time and cost budget
Current limits and recovery behavior supersede the historical details below. See
[the adaptive execution handoff](hermes_suite_recovery_20260930.md): 7200 seconds,
no estimated-cost cutoff, bounded pass recovery and optional hierarchical execution.
New explicitly requested jobs may use `execution.max_seconds: 3600`. Omitting
the field still selects 1800 seconds; valid values are integers from 10 to 3600.
The USD 30 CLI estimated-cost maximum/default is unchanged. These figures are
CLI accounting guards, not proof of subscription billing charges.
```json
{
"execution": {
"strategy": "whole_suite",
"max_seconds": 3600,
"max_cost_usd": 30
}
}
```
This is the execution section of the existing request, not a complete request.
Use a fresh job attempt/idempotency key when intentionally resubmitting a failed
job. No real suite was rerun during this change.
See the [latest real-job capacity diagnosis](hermes_suite_capacity_20260930.md).
The budget verification below is historical; the time and cost limits remain unchanged.
## Runtime and revisions
- HTTP compatibility configuration: `suite-v6-20260929`, unchanged.
- Execution: `suite-multipass-v10-20260930`.
- Prompt: `implementation-proximity-multipass-v5-20260930`.
- Policy: `implementation-five-v1-20260929`, unchanged.
- Model: `claude-opus-5-5`, CLI setting `claude-opus-5-5[1m]`.
- Native Claude Code: 2.1.285, existing pinned binary and first-party OAuth.
Opus 5.5 was intentionally selected in the earlier user-authorized model
upgrade, not by this budget change. Completed real-job CLI results verified
Opus 5.5, firstParty, 1M context and 128K reported model output ceiling. The
service still requests at most 64K output per model turn, with at most six CLI
turns per invocation. Aggregated invocation output can therefore exceed 64K.
The source default now agrees with the existing deployment pin. A job never
changes model or falls back; mismatched runtime identity is rejected.
Effort remains high by default, xhigh for 100+ cases or 128 KiB+ canonical
source content. All passes keep the same selected effort. This is not changed
to make a slow job fit.
## One shared deadline
Request validation and the published JSON schema allow 3600 only when explicitly
requested. Capabilities expose `max_seconds: 3600` and
`default_max_seconds: 1800`. Preflight returns the normalized `execution` object,
and job metadata records it. No new client request field is required.
One absolute monotonic deadline starts before routing, passes through the
orchestrator, and reaches each CLI watchdog. Each call receives only the
remaining time and cost. Status uses that same deadline, including a clamped
zero remaining time at expiry. No pass independently receives another hour.
Shared deadline expiry reports `job_time_budget_exhausted`, including when the
watchdog stops an active CLI process. A distinct provider timeout keeps its
provider error classification; a standalone subprocess watchdog can still
report `timeout`. Output and coverage validation remain strict.
The API remains asynchronous at `https://worker.bstein.dev/suite-planning`.
Client submission/polling stays at 45 seconds. The 30-second HTTP body-read and
40-second ingress response-header timeouts, ingress, authentication, routing,
concurrency, and USD 30 guard are unchanged.
## Failed 363-case job: observed usage
Safe metadata for `9af0db7c71b04882a3a2c0449e83b74c` shows:
| Completed pass | Seconds | Input tokens | Output tokens | Thinking tokens | Estimated USD |
| --- | ---: | ---: | ---: | ---: | ---: |
| proposal_a | 505.343 | 369479 | 57879 | 41501 | 2.635496 |
| proposal_b | 385.938 | 364164 | 47941 | 34819 | 2.415476 |
| reconciliation | 847.944 | 398332 | 108603 | 76973 | 3.765388 |
| Recorded total | 1739.225 | 1131975 | 214423 | 153293 | 8.816360 |
`execution_progress.cost_used_usd_estimate` was **8.81636**, against 30.0.
Cache counters and provider web-tool counts were zero. Each completed invocation
used three CLI turns on Opus 5.5 at xhigh effort. These token counts aggregate
multiple turns/passes; they are not one context-window measurement.
`large_family_review` received 60.72140585165471 seconds. The job failed at
1800.064 wall seconds when the old watchdog terminated it, reporting exit 143,
termination reason `timeout`, and no terminal CLI event. Its token usage and
cost are unknown, so 8.81636 is the recorded completed-pass cost, not an exact
total including interrupted work. The decision audit never started.
If each remaining review takes as long and costs as much as reconciliation,
a fresh full job would take approximately 3435 seconds (57.3 minutes) and
USD 16.35 in CLI accounting. This is a scenario estimate, not a verified upper
bound. It supports the requested 3600-second allowance and retaining USD 30;
there is no evidence requiring a higher cost guard. Content difficulty and
structured-output repair can still exhaust either guard.
## Verification and rollback
Regression checks include explicit opt-in/default/schema bounds, an accelerated
five-pass job lasting 2510 simulated seconds under one 3600-second deadline,
declining allocations after routing, shared-watchdog classification and zero
remaining time, distinct provider timeouts, and existing coverage/isolation tests.
No actual hour-long wait is represented as tested by the accelerated clock.
The pre-deployment benchmark uses an isolated loopback async API and a temporary
database with only the public synthetic 75-case fixture. It exercises the actual
pinned hosted CLI/account, preflight, idempotent submission and 45-second HTTP
polling. Its metadata is recorded separately in the accompanying evidence. It
does not establish laptop connectivity or guarantee the real 363-case runtime.
The synthetic job completed all five passes in **201.953 server wall seconds**
(202.407 seconds including client preflight/polling), with 223758 input tokens,
27713 output tokens including 2823 thinking tokens, and USD **1.449292** estimated
cost. The request was 55986 serialized bytes. Runtime identity was verified as
Opus 5.5, firstParty, at high effort. There were no automatic job retries.
All 75 aliases appeared exactly once. The natural partition matched all nine
synthetic mechanisms with zero false/missed merge pairs; sizing produced 21
final tasks with three singletons. This membership score is distinct from a
comprehensive semantic review of generated descriptions.
The API accepted an explicit 3600-second preflight/submission and rejected 3601.
Per-pass remaining-time allocations were 3599.966, 3584.405, 3555.589, 3487.413,
and 3442.319 seconds; the final watchdog received the last allocation.
Idempotent replay returned the same job. HTTP polling retained 45 seconds.
All 163 regression tests passed; Kustomize render and client dry-run passed.
Flux diff affects only the planner script ConfigMap and rollout annotation.
Full safe measurements: [budget evidence](evidence/hermes_suite_budget_20260929.json).
Rollback by reverting the v9 implementation commit through Git and reconciling
the Hermes Flux Kustomization while no job is active. This restores the 1800-second
maximum and old timeout classification. Do not manually edit the deployment.