atlas-iac/docs/hermes_suite_job_budget.md

7.1 KiB

Suite job time and cost budget

Current limits and recovery behavior supersede the historical details below. See the adaptive execution handoff: 7200 seconds, no estimated-cost cutoff, bounded pass recovery and optional hierarchical execution.

New explicitly requested jobs may use execution.max_seconds: 3600. Omitting the field still selects 1800 seconds; valid values are integers from 10 to 3600. The USD 30 CLI estimated-cost maximum/default is unchanged. These figures are CLI accounting guards, not proof of subscription billing charges.

{
  "execution": {
    "strategy": "whole_suite",
    "max_seconds": 3600,
    "max_cost_usd": 30
  }
}

This is the execution section of the existing request, not a complete request. Use a fresh job attempt/idempotency key when intentionally resubmitting a failed job. No real suite was rerun during this change.

See the latest real-job capacity diagnosis. The budget verification below is historical; the time and cost limits remain unchanged.

Runtime and revisions

  • HTTP compatibility configuration: suite-v6-20260929, unchanged.
  • Execution: suite-multipass-v10-20260930.
  • Prompt: implementation-proximity-multipass-v5-20260930.
  • Policy: implementation-five-v1-20260929, unchanged.
  • Model: claude-opus-5-5, CLI setting claude-opus-5-5[1m].
  • Native Claude Code: 2.1.285, existing pinned binary and first-party OAuth.

Opus 5.5 was intentionally selected in the earlier user-authorized model upgrade, not by this budget change. Completed real-job CLI results verified Opus 5.5, firstParty, 1M context and 128K reported model output ceiling. The service still requests at most 64K output per model turn, with at most six CLI turns per invocation. Aggregated invocation output can therefore exceed 64K. The source default now agrees with the existing deployment pin. A job never changes model or falls back; mismatched runtime identity is rejected.

Effort remains high by default, xhigh for 100+ cases or 128 KiB+ canonical source content. All passes keep the same selected effort. This is not changed to make a slow job fit.

One shared deadline

Request validation and the published JSON schema allow 3600 only when explicitly requested. Capabilities expose max_seconds: 3600 and default_max_seconds: 1800. Preflight returns the normalized execution object, and job metadata records it. No new client request field is required.

One absolute monotonic deadline starts before routing, passes through the orchestrator, and reaches each CLI watchdog. Each call receives only the remaining time and cost. Status uses that same deadline, including a clamped zero remaining time at expiry. No pass independently receives another hour.

Shared deadline expiry reports job_time_budget_exhausted, including when the watchdog stops an active CLI process. A distinct provider timeout keeps its provider error classification; a standalone subprocess watchdog can still report timeout. Output and coverage validation remain strict.

The API remains asynchronous at https://worker.bstein.dev/suite-planning. Client submission/polling stays at 45 seconds. The 30-second HTTP body-read and 40-second ingress response-header timeouts, ingress, authentication, routing, concurrency, and USD 30 guard are unchanged.

Failed 363-case job: observed usage

Safe metadata for 9af0db7c71b04882a3a2c0449e83b74c shows:

Completed pass Seconds Input tokens Output tokens Thinking tokens Estimated USD
proposal_a 505.343 369479 57879 41501 2.635496
proposal_b 385.938 364164 47941 34819 2.415476
reconciliation 847.944 398332 108603 76973 3.765388
Recorded total 1739.225 1131975 214423 153293 8.816360

execution_progress.cost_used_usd_estimate was 8.81636, against 30.0. Cache counters and provider web-tool counts were zero. Each completed invocation used three CLI turns on Opus 5.5 at xhigh effort. These token counts aggregate multiple turns/passes; they are not one context-window measurement.

large_family_review received 60.72140585165471 seconds. The job failed at 1800.064 wall seconds when the old watchdog terminated it, reporting exit 143, termination reason timeout, and no terminal CLI event. Its token usage and cost are unknown, so 8.81636 is the recorded completed-pass cost, not an exact total including interrupted work. The decision audit never started.

If each remaining review takes as long and costs as much as reconciliation, a fresh full job would take approximately 3435 seconds (57.3 minutes) and USD 16.35 in CLI accounting. This is a scenario estimate, not a verified upper bound. It supports the requested 3600-second allowance and retaining USD 30; there is no evidence requiring a higher cost guard. Content difficulty and structured-output repair can still exhaust either guard.

Verification and rollback

Regression checks include explicit opt-in/default/schema bounds, an accelerated five-pass job lasting 2510 simulated seconds under one 3600-second deadline, declining allocations after routing, shared-watchdog classification and zero remaining time, distinct provider timeouts, and existing coverage/isolation tests. No actual hour-long wait is represented as tested by the accelerated clock.

The pre-deployment benchmark uses an isolated loopback async API and a temporary database with only the public synthetic 75-case fixture. It exercises the actual pinned hosted CLI/account, preflight, idempotent submission and 45-second HTTP polling. Its metadata is recorded separately in the accompanying evidence. It does not establish laptop connectivity or guarantee the real 363-case runtime.

The synthetic job completed all five passes in 201.953 server wall seconds (202.407 seconds including client preflight/polling), with 223758 input tokens, 27713 output tokens including 2823 thinking tokens, and USD 1.449292 estimated cost. The request was 55986 serialized bytes. Runtime identity was verified as Opus 5.5, firstParty, at high effort. There were no automatic job retries.

All 75 aliases appeared exactly once. The natural partition matched all nine synthetic mechanisms with zero false/missed merge pairs; sizing produced 21 final tasks with three singletons. This membership score is distinct from a comprehensive semantic review of generated descriptions.

The API accepted an explicit 3600-second preflight/submission and rejected 3601. Per-pass remaining-time allocations were 3599.966, 3584.405, 3555.589, 3487.413, and 3442.319 seconds; the final watchdog received the last allocation. Idempotent replay returned the same job. HTTP polling retained 45 seconds. All 163 regression tests passed; Kustomize render and client dry-run passed. Flux diff affects only the planner script ConfigMap and rollout annotation. Full safe measurements: budget evidence.

Rollback by reverting the v9 implementation commit through Git and reconciling the Hermes Flux Kustomization while no job is active. This restores the 1800-second maximum and old timeout classification. Do not manually edit the deployment.