atlas-iac/docs/hermes_suite_multipass.md

17 KiB

Multi-pass implementation grouping

The suite planner keeps the same HTTPS endpoint, request fields, authentication, provider permissions, generalized-data approval, and asynchronous job lifecycle. The required result object remains result.groups, with name, description, and members. Final groups now contain one to five cases and names are at most 64 characters, unique after whitespace and case normalization.

Revisions and process

  • Configuration: suite-v6-20260929 (HTTP compatibility identifier).
  • Policy: implementation-five-v1-20260929.
  • Prompt: implementation-proximity-multipass-v4-20260929.
  • Execution: suite-multipass-v3-20260929.

A single server-side job performs:

  1. Proposal A on the complete suite, in stable alias order, without a size cap.
  2. Proposal B in a reproducible SHA-256 alias order, in a separate fresh CLI invocation. It sees no proposal A or earlier conversation. Both use exactly the same implementation-effort system instructions and output schema.
  3. Whole-suite reconciliation using both memberships and all original fields. Success criteria remain primary evidence. There is no majority voting or transitive closure of pairwise similarity. Natural families are still uncapped.
  4. If any reconciled family exceeds five, review every such original family using all case fields, distinguishing different machinery from cheap variations.
  5. One additional bounded audit of those decisions, including descendants that were split into smaller groups. It may reverse an unjustified split or refine a broad family, but cannot cross the reconciled family boundaries. There is no agreement loop.

The final two passes record explicit keep/split rationales and exact short source-field quotations. The code verifies that evidence belongs to an assigned alias and is present verbatim in that source field. This checks support provenance, not the correctness of a model's engineering interpretation. Unknown implementation details remain concise uncertainty, not invented equipment or procedures.

The internal structured-output schema requires one property for each exact source alias. Its value selects a natural-family index and, in review passes, a variation index. The server reconstructs memberships and runs the normal independent checks. Review decisions also use required keys derived from exact original memberships. This avoids free-form membership lists silently omitting or repeating aliases. The public groups schema is unchanged. Invalid keys, indexes, empty groups, and missing review decisions still fail closed; no missing assignment is fabricated.

Each invocation retains six CLI turns for structured output. Model review passes and CLI turns are separate counters. All calls use the originally selected provider and pinned model. The model remains claude-opus-4-8[1m], with canonical runtime identity checked as claude-opus-4-8, firstParty, medium effort, reported 1M context and 64K output ceiling. The server invokes native Claude Code CLI 2.1.226 using the existing first-party OAuth account, not a separately configured API-key account. The CLI itself communicates with Anthropic over HTTPS. No new provider fallback or tools are enabled.

Deterministic task sizing and names

For each remaining coherent large family, k = ceil(n / 5) and base, remainder = divmod(n, k). The first remainder parts have base+1 cases; the rest have base. Thus 6 becomes 3+3, 7 becomes 4+3, 11 becomes 4+4+3, and 14 becomes 5+5+4. No leftover singleton is created by a five-at-a-time slice.

The semantic review supplies internal variation sets. A deterministic packer uses those hints and stable complete-record hashes/aliases to fill the calculated sizes. Equivalent content remains distinct by alias. Hint order, source order, and member order do not cause gratuitous reshuffling. Different model decisions can still change membership between fresh jobs; determinism is conditional on the same natural families and variation hints, not a claim that inference is deterministic.

Natural names must be concise, meaningful, distinct, and at most 56 characters, leaving room for numbering within the 64-character final limit. Collisions or invalid names fail validation; the service never truncates them or adds numbers to unrelated families. Capacity-only parts have the exact same base name: Reset recovery (1/3), Reset recovery (2/3), Reset recovery (3/3). Their descriptions identify a work-size division and use only the reviewed common work applicable to all members. No other part's specific objectives are copied.

Every intermediate partition and final result is checked for exact alias coverage. Final size, names, balanced part sizes, and campaign/suite ownership are checked programmatically. Incomplete or invalid work is never returned as completed.

Review artifact and operational metadata

GET /suite-planning/v1/jobs/<id>/result returns an optional top-level review_summary, alongside the unchanged result object. It contains:

  • proposal_disagreements: affected aliases, differing pair count, and both full membership partitions in compressed form. Group names/numbers are not compared.
  • reconciled_families: the first natural partition with concise rationale, common work, exact source-field support, and uncertainty.
  • large_family_review and decision_audit: explicit decisions for each original oversized membership set, including semantic-split versus coherent-keep reasons.
  • natural_families: the final conceptual partition before task sizing.
  • capacity_divisions: base names, calculated sizes, final part names, common implementation rationale, and remaining uncertainty.
  • unresolved_uncertainties and counts: separate natural-family/final-task counts and natural/final singleton counts. Size compliance is not semantic-quality evidence.

Review content lives only in the authorized in-memory result cache, with the same one-hour TTL and restart loss as results. It is absent from routine logs, status responses, and SQLite job metadata. It contains source-derived material and exact quotations, so it is not an export-safe ClickUp artifact. No new hierarchy level or client request fields are introduced.

Status/result metadata includes policy_revision, execution_revision, prompt_revision, model_pass_count, natural_family_count, final_task_count, singleton_statistics, aggregate usage/cost, and safe per-invocation passes. Each pass records its stage, actual model, system/schema hashes, ordering hash, capacity check, remaining budgets, usage, timing, and CLI diagnostics. The top-level cli_diagnostics concerns the last invocation; turns is the sum of reported CLI turns, while model_pass_count counts complete model review invocations. execution_progress.current_pass reports the current stage during a running job. While a CLI invocation runs, progress updates every five seconds with heartbeat_at, pass_elapsed_seconds, job_elapsed_seconds, job_remaining_seconds, completed_model_passes, maximum_model_passes, cli_running, cli_output_bytes, and last_cli_activity_seconds_ago. The heartbeat means the worker is alive; a change in output bytes means CLI activity. Neither is a percentage complete or proof that a particular case has been reasoned about. No event bodies or reasoning text are exposed. cost_used_usd_estimate counts completed invocations; the in-flight invocation is not included until its terminal usage is available. Poll the existing status URL every five seconds to display progress.

To inspect progress from WSL, keep the existing bearer token private and use the job ID returned by submission. This command only reads status:

curl -q --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
  --connect-timeout 10 --max-time 45 --fail-with-body --silent --show-error \
  --config <(printf 'header = "Authorization: Bearer %s"\n' "$SUITE_PLANNING_TOKEN") \
  "https://worker.bstein.dev/suite-planning/v1/jobs/$JOB_ID"

An existing Python poller can display envelope["execution_progress"] after each status read. No separate pass submissions, new request fields, or client-side reconciliation are required. Prefer the actual stage and elapsed time over an invented percentage, because different passes have very different runtimes.

Shared bounds and failure behavior

The user-approved 1200-second maximum job deadline and USD 10 CLI estimated-cost guard apply to all calls combined, including routing time. These are also the defaults when omitted. Explicit lower client limits remain effective. This guard uses Claude CLI estimated-cost accounting; it is not subscription billing. Each invocation receives only the remaining time and estimated-cost allowance. Available usage/cost is accumulated after every call; unknown cost accounting stops further hosted calls. No partial partition substitutes for unfinished passes. The CLI's estimated-cost ceiling can overshoot within one generation, as before; this is not a guaranteed subscription billing limit. An observed overrun fails the job rather than allowing further calls or reporting successful completion.

Request and complete response envelopes stay bounded at 1 MiB. Every expanded pass is checked before dispatch, including original source, proposals/reviews, system instructions, and schema. The same conservative byte bound and six possible CLI outputs are reserved against context. Exact tokenization and output reservation remain unverified estimates; actual incomplete generation fails explicitly. Preflight checks initial and minimum reconciliation capacity; it identifies that actual later-pass capacity will be checked before each invocation. It does not claim to know model-generated proposal sizes in advance.

The existing 8K-context local route cannot admit the new reconciliation schema and instructions. Local-only requests fail with capacity errors and never start hosted jobs. The separate /local-model endpoint and its limits are unchanged. No bounds other than the explicitly approved time/cost guards were increased, and no provider was switched after an error.

Failures include existing CLI diagnostics and a fixed review_pass, completed-pass count, and available aggregate usage. New explicit codes include pass_capacity, pass_request_too_large, job_time_budget_exhausted, job_cost_budget_exhausted, budget_accounting_unavailable, and bounded review/assignment validation errors. Raw provider messages, source text, and generated explanations are not error details.

The laptop submits one request with its existing scoped token, then polls the same job. Retry a timed-out submission using the same key/body; an intentional fresh attempt needs a new Idempotency-Key. Include policy/prompt/execution revisions in local cache provenance. No real suite is submitted by deployment checks.

Verification and rollback

Synthetic unit tests cover the balanced-size examples, semantic subdivisions and a bounded reversal, differing proposals, duplicate text, name normalization and collisions, source support, deadline/cost exhaustion, cancellation, privacy, authorization, idempotency, and 14/75/363-case capacity and sizing. Native CLI loopback and authenticated LAN/provider acceptance measurements are recorded after deployment. Mocked outputs alone do not establish model quality or throughput.

Deploy or roll back through Git and Flux. Revert only the multi-pass commits to restore the preceding single-pass policy while preserving the earlier CLI diagnostic fix. Wait for active jobs before a worker restart because completed results are held in memory. No Vault credential or ingress/routing changes are required.

The policy implementation is commit 7af4f4f2; the approved 20-minute / USD 10 estimate guard and progress reporting are commit 15c05cb7. To restore the prior single-pass implementation after active jobs finish, revert both through the normal Git deployment branch, preserving the earlier six-turn diagnostics fix:

git revert --no-commit 15c05cb7 7af4f4f2
git commit -m "hermes: restore prior suite grouping policy"
git push origin HEAD:main
flux reconcile kustomization hermes --namespace flux-system --with-source

The first live 363-case multi-pass run at the previous 900-second / USD 5 guards stopped after 747.513 seconds during large-family review. Its actual CLI final subtype was error_max_budget_usd; three full passes had completed. The CLI reported an aggregate estimate of 5.62765, showing the documented within-generation overshoot. It was not a timeout, turn-limit failure, or accepted partial answer. The adapter now reports this condition as job_cost_budget_exhausted.

Synthetic acceptance measurements

The suite inputs are synthetic fixtures, never roster-derived material. Tests ran through authenticated TLS from titan-jh (192.168.22.8) to the LAN ingress 192.168.22.50:443, with hostname verification, proxy bypass, and no redirects. This verifies that LAN test path; it does not replace a WSL connectivity test.

Cases Request bytes Wall seconds Model passes / CLI turns Natural families Final tasks Natural / final singletons Input / output tokens
14 9,847 61.805 3 / 6 9 9 4 / 4 18,216 / 5,794
75 55,984 274.620 5 / 12 9 21 3 / 3 174,942 / 28,950
38 balanced fixture 23,871 231.337 5 / 15 4 10 0 / 0 139,789 / 24,681
363, old limits 273,761 747.513 3 completed; pass 4 failed Not finalized No result Not finalized 629,141 / 64,424

Token totals include all completed CLI turns across the job, not unique source size. The old-limit 363 result includes reported failed-invocation usage and was never returned as completed. CLI estimated costs were 0.23593, 1.59846, 1.31597, and 5.62765 respectively; these are not subscription bills.

For every completed run above, natural-family pair precision and recall were 1.0 against the independently defined implementation patterns, exact alias coverage passed, and final groups stayed pure to those patterns. No incorrect merges, unnecessary semantic splits, or cross-family description attribution were observed during review. Equivalent text under separate aliases remained separate objectives. Natural-family quality was scored before capacity division; a cap-compliant result alone does not prove semantic correctness.

The 38-case fixture interleaves four reset-related mechanisms with shared generic subject descriptions and distinct success criteria: RPC/stub observations, pulse and oscilloscope timing, offline report parsing, and concurrent queue instrumentation. It yielded natural sizes 6/7/11/14 and final sizes 3+3 / 4+3 / 4+4+3 / 5+5+4. Thirty-eight authorized status samples showed five stages, changing heartbeats, and nineteen CLI activity changes without exposing event bodies.

The actual 14/75/38 independent proposals agreed on membership. Disagreement reconciliation, genuine semantic subdivision, reversal of an unjustified split, and naming-collision rejection are separately tested with controlled mocked responses; do not describe those as observed live-provider disagreements.

The public 363 fixture has description lengths 48-245 characters (mean 234.5), preconditions 67-143 (mean 137.2), success criteria 73-135 (mean 131.2), and case labels 7-15. It interleaves six mechanisms, includes identical distant case text, and has distinctive beginning/middle/end objectives. These measurements describe the tested material; longer or more ambiguous real suites may take more work.

An installed native CLI loopback probe captured all complete case fields and full system instructions in every emitted provider request for 14/75/363 cases, with 64,000 output limits. It performed 3/5/5 passes and deterministic final sizing to 9/21/75 tasks using mock responses. This verifies transport and orchestration, not live semantic quality. Live successful results reported no compaction or truncation and passed exact coverage and source-quote support checks. These checks cannot prove the provider internally attended to every field; exact tokenization and sufficient output reservation for arbitrary suites remain unverified.

The updated code passes 109 synthetic/mocked regression tests. Kubernetes render, client dry-run, and Flux diff checks passed before deployment. Policy controls, case coverage validation, scoped authentication, and provider pinning remain in place. The invalid-token check returned 401 and same-key submission replay reused the same job. No automatic real-data rerun was performed.

The second 363-case attempt, under the 1200-second / USD 10 guards, failed after 228.260 seconds in proposal B with invalid_case_assignments. Both invocations returned successful CLI results, but the second partition failed server membership validation. Its reported total was 281,227 input / 23,796 output tokens and 2.001035 in CLI estimated-cost accounting. The previous validator did not retain counts of missing, repeated, or unknown assignments, and transient output was cleaned; the specific mismatch cannot be recovered. No partial partition was accepted. The new required-alias schema and count-only failure diagnostics address this failure mode.