137 lines
7.1 KiB
Markdown
137 lines
7.1 KiB
Markdown
# Suite job time and cost budget
|
|
|
|
Current limits and recovery behavior supersede the historical details below. See
|
|
[the adaptive execution handoff](hermes_suite_recovery_20260930.md): 7200 seconds,
|
|
no estimated-cost cutoff, bounded pass recovery and optional hierarchical execution.
|
|
|
|
New explicitly requested jobs may use `execution.max_seconds: 3600`. Omitting
|
|
the field still selects 1800 seconds; valid values are integers from 10 to 3600.
|
|
The USD 30 CLI estimated-cost maximum/default is unchanged. These figures are
|
|
CLI accounting guards, not proof of subscription billing charges.
|
|
|
|
```json
|
|
{
|
|
"execution": {
|
|
"strategy": "whole_suite",
|
|
"max_seconds": 3600,
|
|
"max_cost_usd": 30
|
|
}
|
|
}
|
|
```
|
|
|
|
This is the execution section of the existing request, not a complete request.
|
|
Use a fresh job attempt/idempotency key when intentionally resubmitting a failed
|
|
job. No real suite was rerun during this change.
|
|
|
|
See the [latest real-job capacity diagnosis](hermes_suite_capacity_20260930.md).
|
|
The budget verification below is historical; the time and cost limits remain unchanged.
|
|
|
|
## Runtime and revisions
|
|
|
|
- HTTP compatibility configuration: `suite-v6-20260929`, unchanged.
|
|
- Execution: `suite-multipass-v10-20260930`.
|
|
- Prompt: `implementation-proximity-multipass-v5-20260930`.
|
|
- Policy: `implementation-five-v1-20260929`, unchanged.
|
|
- Model: `claude-opus-5-5`, CLI setting `claude-opus-5-5[1m]`.
|
|
- Native Claude Code: 2.1.285, existing pinned binary and first-party OAuth.
|
|
|
|
Opus 5.5 was intentionally selected in the earlier user-authorized model
|
|
upgrade, not by this budget change. Completed real-job CLI results verified
|
|
Opus 5.5, firstParty, 1M context and 128K reported model output ceiling. The
|
|
service still requests at most 64K output per model turn, with at most six CLI
|
|
turns per invocation. Aggregated invocation output can therefore exceed 64K.
|
|
The source default now agrees with the existing deployment pin. A job never
|
|
changes model or falls back; mismatched runtime identity is rejected.
|
|
|
|
Effort remains high by default, xhigh for 100+ cases or 128 KiB+ canonical
|
|
source content. All passes keep the same selected effort. This is not changed
|
|
to make a slow job fit.
|
|
|
|
## One shared deadline
|
|
|
|
Request validation and the published JSON schema allow 3600 only when explicitly
|
|
requested. Capabilities expose `max_seconds: 3600` and
|
|
`default_max_seconds: 1800`. Preflight returns the normalized `execution` object,
|
|
and job metadata records it. No new client request field is required.
|
|
|
|
One absolute monotonic deadline starts before routing, passes through the
|
|
orchestrator, and reaches each CLI watchdog. Each call receives only the
|
|
remaining time and cost. Status uses that same deadline, including a clamped
|
|
zero remaining time at expiry. No pass independently receives another hour.
|
|
|
|
Shared deadline expiry reports `job_time_budget_exhausted`, including when the
|
|
watchdog stops an active CLI process. A distinct provider timeout keeps its
|
|
provider error classification; a standalone subprocess watchdog can still
|
|
report `timeout`. Output and coverage validation remain strict.
|
|
|
|
The API remains asynchronous at `https://worker.bstein.dev/suite-planning`.
|
|
Client submission/polling stays at 45 seconds. The 30-second HTTP body-read and
|
|
40-second ingress response-header timeouts, ingress, authentication, routing,
|
|
concurrency, and USD 30 guard are unchanged.
|
|
|
|
## Failed 363-case job: observed usage
|
|
|
|
Safe metadata for `9af0db7c71b04882a3a2c0449e83b74c` shows:
|
|
|
|
| Completed pass | Seconds | Input tokens | Output tokens | Thinking tokens | Estimated USD |
|
|
| --- | ---: | ---: | ---: | ---: | ---: |
|
|
| proposal_a | 505.343 | 369479 | 57879 | 41501 | 2.635496 |
|
|
| proposal_b | 385.938 | 364164 | 47941 | 34819 | 2.415476 |
|
|
| reconciliation | 847.944 | 398332 | 108603 | 76973 | 3.765388 |
|
|
| Recorded total | 1739.225 | 1131975 | 214423 | 153293 | 8.816360 |
|
|
|
|
`execution_progress.cost_used_usd_estimate` was **8.81636**, against 30.0.
|
|
Cache counters and provider web-tool counts were zero. Each completed invocation
|
|
used three CLI turns on Opus 5.5 at xhigh effort. These token counts aggregate
|
|
multiple turns/passes; they are not one context-window measurement.
|
|
|
|
`large_family_review` received 60.72140585165471 seconds. The job failed at
|
|
1800.064 wall seconds when the old watchdog terminated it, reporting exit 143,
|
|
termination reason `timeout`, and no terminal CLI event. Its token usage and
|
|
cost are unknown, so 8.81636 is the recorded completed-pass cost, not an exact
|
|
total including interrupted work. The decision audit never started.
|
|
|
|
If each remaining review takes as long and costs as much as reconciliation,
|
|
a fresh full job would take approximately 3435 seconds (57.3 minutes) and
|
|
USD 16.35 in CLI accounting. This is a scenario estimate, not a verified upper
|
|
bound. It supports the requested 3600-second allowance and retaining USD 30;
|
|
there is no evidence requiring a higher cost guard. Content difficulty and
|
|
structured-output repair can still exhaust either guard.
|
|
|
|
## Verification and rollback
|
|
|
|
Regression checks include explicit opt-in/default/schema bounds, an accelerated
|
|
five-pass job lasting 2510 simulated seconds under one 3600-second deadline,
|
|
declining allocations after routing, shared-watchdog classification and zero
|
|
remaining time, distinct provider timeouts, and existing coverage/isolation tests.
|
|
No actual hour-long wait is represented as tested by the accelerated clock.
|
|
|
|
The pre-deployment benchmark uses an isolated loopback async API and a temporary
|
|
database with only the public synthetic 75-case fixture. It exercises the actual
|
|
pinned hosted CLI/account, preflight, idempotent submission and 45-second HTTP
|
|
polling. Its metadata is recorded separately in the accompanying evidence. It
|
|
does not establish laptop connectivity or guarantee the real 363-case runtime.
|
|
|
|
The synthetic job completed all five passes in **201.953 server wall seconds**
|
|
(202.407 seconds including client preflight/polling), with 223758 input tokens,
|
|
27713 output tokens including 2823 thinking tokens, and USD **1.449292** estimated
|
|
cost. The request was 55986 serialized bytes. Runtime identity was verified as
|
|
Opus 5.5, firstParty, at high effort. There were no automatic job retries.
|
|
|
|
All 75 aliases appeared exactly once. The natural partition matched all nine
|
|
synthetic mechanisms with zero false/missed merge pairs; sizing produced 21
|
|
final tasks with three singletons. This membership score is distinct from a
|
|
comprehensive semantic review of generated descriptions.
|
|
|
|
The API accepted an explicit 3600-second preflight/submission and rejected 3601.
|
|
Per-pass remaining-time allocations were 3599.966, 3584.405, 3555.589, 3487.413,
|
|
and 3442.319 seconds; the final watchdog received the last allocation.
|
|
Idempotent replay returned the same job. HTTP polling retained 45 seconds.
|
|
All 163 regression tests passed; Kustomize render and client dry-run passed.
|
|
Flux diff affects only the planner script ConfigMap and rollout annotation.
|
|
Full safe measurements: [budget evidence](evidence/hermes_suite_budget_20260929.json).
|
|
|
|
Rollback by reverting the v9 implementation commit through Git and reconciling
|
|
the Hermes Flux Kustomization while no job is active. This restores the 1800-second
|
|
maximum and old timeout classification. Do not manually edit the deployment.
|