336 lines
21 KiB
Markdown
336 lines
21 KiB
Markdown
# Multi-pass implementation grouping
|
|
|
|
The suite planner keeps the same HTTPS endpoint, request fields, authentication,
|
|
provider permissions, generalized-data approval, and asynchronous job lifecycle.
|
|
The required result object remains `result.groups`, with `name`, `description`,
|
|
and `members`. Final groups now contain **one to five** cases and names are at
|
|
most **64 characters**, unique after whitespace and case normalization.
|
|
|
|
## Revisions and process
|
|
|
|
- Configuration: `suite-v6-20260929` (HTTP compatibility identifier).
|
|
- Policy: `implementation-five-v1-20260929`.
|
|
- Prompt: `implementation-proximity-multipass-v4-20260929`.
|
|
- Execution: `suite-multipass-v7-20260929`.
|
|
|
|
A single server-side job performs:
|
|
|
|
1. Proposal A on the complete suite, in stable alias order, without a size cap.
|
|
2. Proposal B in a reproducible SHA-256 alias order, in a separate fresh CLI
|
|
invocation. It sees no proposal A or earlier conversation. Both use exactly
|
|
the same implementation-effort system instructions and output schema.
|
|
3. Whole-suite reconciliation using both memberships and all original fields.
|
|
Success criteria remain primary evidence. There is no majority voting or
|
|
transitive closure of pairwise similarity. Natural families are still uncapped.
|
|
4. If any reconciled family exceeds five, review every such original family using
|
|
all case fields, distinguishing different machinery from cheap variations.
|
|
5. One additional bounded audit of those decisions, including descendants that
|
|
were split into smaller groups. It may reverse an unjustified split or refine
|
|
a broad family, but cannot cross the reconciled family boundaries. There is
|
|
no agreement loop.
|
|
|
|
The final two passes record explicit keep/split rationales and exact short
|
|
source-field quotations. The code verifies that evidence belongs to an assigned
|
|
alias and is present verbatim in that source field. This checks support provenance,
|
|
not the correctness of a model's engineering interpretation. Unknown implementation
|
|
details remain concise uncertainty, not invented equipment or procedures.
|
|
|
|
The internal structured-output schema requires one property for each exact source
|
|
alias. Its value selects a natural-family index and, in review passes, a variation
|
|
index. The server reconstructs memberships and runs the normal independent checks.
|
|
Review decisions also use required keys derived from exact original memberships.
|
|
This avoids free-form membership lists silently omitting or repeating aliases.
|
|
The public groups schema is unchanged. Invalid keys, indexes, empty groups, and
|
|
missing review decisions still fail closed; no missing assignment is fabricated.
|
|
|
|
Each invocation retains six CLI turns for structured output. Model review passes
|
|
and CLI turns are separate counters. All calls use the originally selected provider
|
|
and pinned model. The current default is `claude-opus-5-5[1m]`, with canonical runtime
|
|
identity checked as `claude-opus-5-5`, firstParty, reported 1M context
|
|
and 128K model output ceiling, with requests limited to 64K output tokens.
|
|
The server invokes native Claude Code CLI 2.1.285 using
|
|
the existing first-party OAuth account, not a separately configured API-key
|
|
account. The CLI itself communicates with Anthropic over HTTPS. No new provider
|
|
fallback or tools are enabled.
|
|
|
|
Effort defaults to `high` and increases to `xhigh` for suites with at least 100
|
|
cases OR 128 KiB of canonical UTF-8 source JSON (campaign, suite, and all cases).
|
|
Preflight selects the effort once for the whole job; every discovery, review,
|
|
audit and structured-output repair uses that selection. Source size excludes
|
|
transport whitespace, prompts, schemas, and generated proposals. Thresholds are
|
|
published in capabilities; selection metadata includes counts, bytes, and triggers.
|
|
Existing deadlines and cost guards still apply and never cause a silent effort
|
|
reduction. See the [effort policy details](hermes_suite_claude55.md).
|
|
|
|
Sonnet 5.5 is also verified on the 14-case synthetic fixture and available through
|
|
server configuration. See the [matched comparison](hermes_suite_claude55.md).
|
|
Earlier acceptance measurements below used Opus 4.8; they do not establish large
|
|
suite capacity or quality for the new models.
|
|
|
|
## Deterministic task sizing and names
|
|
|
|
For each remaining coherent large family, `k = ceil(n / 5)` and
|
|
`base, remainder = divmod(n, k)`. The first remainder parts have base+1 cases;
|
|
the rest have base. Thus 6 becomes 3+3, 7 becomes 4+3, 11 becomes 4+4+3, and
|
|
14 becomes 5+5+4. No leftover singleton is created by a five-at-a-time slice.
|
|
|
|
The semantic review supplies internal variation sets. A deterministic packer uses
|
|
those hints and stable complete-record hashes/aliases to fill the calculated
|
|
sizes. Equivalent content remains distinct by alias. Hint order, source order,
|
|
and member order do not cause gratuitous reshuffling. Different model decisions
|
|
can still change membership between fresh jobs; determinism is conditional on the
|
|
same natural families and variation hints, not a claim that inference is deterministic.
|
|
|
|
Natural names must be concise, meaningful, distinct, and at most 56 characters,
|
|
leaving room for numbering within the 64-character final limit. Collisions or
|
|
invalid names fail validation; the service never truncates them or adds numbers
|
|
to unrelated families. Capacity-only parts have the exact same base name:
|
|
`Reset recovery (1/3)`, `Reset recovery (2/3)`, `Reset recovery (3/3)`.
|
|
Their descriptions identify a work-size division and use only the reviewed common
|
|
work applicable to all members. No other part's specific objectives are copied.
|
|
|
|
Every intermediate partition and final result is checked for exact alias coverage.
|
|
Final size, names, balanced part sizes, and campaign/suite ownership are checked
|
|
programmatically. Incomplete or invalid work is never returned as completed.
|
|
|
|
## Review artifact and operational metadata
|
|
|
|
`GET /suite-planning/v1/jobs/<id>/result` returns an optional top-level
|
|
**`review_summary`**, alongside the unchanged **`result`** object. It contains:
|
|
|
|
- `proposal_disagreements`: affected aliases, differing pair count, and both full
|
|
membership partitions in compressed form. Group names/numbers are not compared.
|
|
- `reconciled_families`: the first natural partition with concise rationale,
|
|
common work, exact source-field support, and uncertainty.
|
|
- `large_family_review` and `decision_audit`: explicit decisions for each original
|
|
oversized membership set, including semantic-split versus coherent-keep reasons.
|
|
- `natural_families`: the final conceptual partition before task sizing.
|
|
- `capacity_divisions`: base names, calculated sizes, final part names, common
|
|
implementation rationale, and remaining uncertainty.
|
|
- `unresolved_uncertainties` and `counts`: separate natural-family/final-task counts
|
|
and natural/final singleton counts. Size compliance is not semantic-quality evidence.
|
|
|
|
Review content lives only in the authorized in-memory result cache, with the same
|
|
one-hour TTL and restart loss as results. It is absent from routine logs, status
|
|
responses, and SQLite job metadata. It contains source-derived material and exact
|
|
quotations, so it is not an export-safe ClickUp artifact. No new hierarchy level
|
|
or client request fields are introduced.
|
|
|
|
Status/result metadata includes `policy_revision`, `execution_revision`,
|
|
`prompt_revision`, `model_pass_count`, `natural_family_count`, `final_task_count`,
|
|
`singleton_statistics`, aggregate usage/cost, and safe per-invocation `passes`.
|
|
Each pass records its stage, actual model, system/schema hashes, ordering hash,
|
|
capacity check, remaining budgets, usage, timing, and CLI diagnostics. The top-level
|
|
`cli_diagnostics` concerns the last invocation; `turns` is the sum of reported CLI
|
|
turns, while `model_pass_count` counts complete model review invocations.
|
|
`execution_progress.current_pass` reports the current stage during a running job.
|
|
While a CLI invocation runs, progress updates every five seconds with `heartbeat_at`,
|
|
`pass_elapsed_seconds`, `job_elapsed_seconds`, `job_remaining_seconds`,
|
|
`completed_model_passes`, `maximum_model_passes`, `cli_running`, `cli_output_bytes`,
|
|
and `last_cli_activity_seconds_ago`. The heartbeat means the worker is alive; a
|
|
change in output bytes means CLI activity. Neither is a percentage complete or
|
|
proof that a particular case has been reasoned about. No event bodies or reasoning
|
|
text are exposed. `cost_used_usd_estimate` counts completed invocations; the
|
|
in-flight invocation is not included until its terminal usage is available.
|
|
Poll the existing status URL every five seconds to display progress.
|
|
|
|
To inspect progress from WSL, keep the existing bearer token private and use the
|
|
job ID returned by submission. This command only reads status:
|
|
|
|
```bash
|
|
curl -q --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
|
|
--connect-timeout 10 --max-time 45 --fail-with-body --silent --show-error \
|
|
--config <(printf 'header = "Authorization: Bearer %s"\n' "$SUITE_PLANNING_TOKEN") \
|
|
"https://worker.bstein.dev/suite-planning/v1/jobs/$JOB_ID"
|
|
```
|
|
|
|
An existing Python poller can display `envelope["execution_progress"]` after each
|
|
status read. No separate pass submissions, new request fields, or client-side
|
|
reconciliation are required. Prefer the actual stage and elapsed time over an
|
|
invented percentage, because different passes have very different runtimes.
|
|
|
|
## Shared bounds and failure behavior
|
|
|
|
The user-approved 1800-second maximum job deadline and USD 30 CLI estimated-cost
|
|
guard apply to **all calls combined**, including routing time. These are also the
|
|
defaults when omitted. Explicit lower client limits remain effective. This guard
|
|
uses Claude CLI estimated-cost accounting; it is not subscription billing. Each invocation receives
|
|
only the remaining time and estimated-cost allowance. Available usage/cost is
|
|
accumulated after every call; unknown cost accounting stops further hosted calls.
|
|
No partial partition substitutes for unfinished passes. The CLI's estimated-cost
|
|
ceiling can overshoot within one generation, as before; this is not a guaranteed
|
|
subscription billing limit. An observed overrun fails the job rather than allowing
|
|
further calls or reporting successful completion.
|
|
|
|
Request and complete response envelopes stay bounded at 1 MiB. Every expanded pass
|
|
is checked before dispatch, including original source, proposals/reviews, system
|
|
instructions, and schema. The same conservative byte bound and six possible CLI
|
|
outputs are reserved against context. Exact tokenization and output reservation
|
|
remain unverified estimates; actual incomplete generation fails explicitly.
|
|
Preflight checks initial and minimum reconciliation capacity; it identifies that
|
|
actual later-pass capacity will be checked before each invocation. It does not
|
|
claim to know model-generated proposal sizes in advance.
|
|
|
|
The existing 8K-context local route cannot admit the new reconciliation schema and
|
|
instructions. Local-only requests fail with capacity errors and never start hosted
|
|
jobs. The separate `/local-model` endpoint and its limits are unchanged. No bounds
|
|
other than the explicitly approved time/cost guards were increased, and no provider
|
|
was switched after an error.
|
|
|
|
Failures include existing CLI diagnostics and a fixed `review_pass`, completed-pass
|
|
count, and available aggregate usage. New explicit codes include `pass_capacity`,
|
|
`pass_request_too_large`, `job_time_budget_exhausted`, `job_cost_budget_exhausted`,
|
|
`budget_accounting_unavailable`, and bounded review/assignment validation errors.
|
|
Raw provider messages, source text, and generated explanations are not error details.
|
|
|
|
The laptop submits one request with its existing scoped token, then polls the same
|
|
job. Retry a timed-out submission using the same key/body; an intentional fresh
|
|
attempt needs a new Idempotency-Key. Include policy/prompt/execution revisions in
|
|
local cache provenance. No real suite is submitted by deployment checks.
|
|
|
|
## Verification and rollback
|
|
|
|
Synthetic unit tests cover the balanced-size examples, semantic subdivisions and
|
|
a bounded reversal, differing proposals, duplicate text, name normalization and
|
|
collisions, source support, deadline/cost exhaustion, cancellation, privacy,
|
|
authorization, idempotency, and 14/75/363-case capacity and sizing. Native CLI
|
|
loopback and authenticated LAN/provider acceptance measurements are recorded after
|
|
deployment. Mocked outputs alone do not establish model quality or throughput.
|
|
|
|
Deploy or roll back through Git and Flux. Revert only the multi-pass commits to
|
|
restore the preceding single-pass policy while preserving the earlier CLI diagnostic
|
|
fix. Wait for active jobs before a worker restart because completed results are
|
|
held in memory. No Vault credential or ingress/routing changes are required.
|
|
|
|
The policy implementation is commit `7af4f4f2`; the approved 20-minute / USD 10
|
|
estimate guard and progress reporting are commit `15c05cb7`. Required-alias
|
|
structured output is commit `9c327c6f`. To restore the prior
|
|
single-pass implementation after active jobs finish, revert these through the normal
|
|
Git deployment branch, preserving the earlier six-turn diagnostics fix:
|
|
|
|
```bash
|
|
git revert --no-commit 9c327c6f 15c05cb7 7af4f4f2
|
|
git commit -m "hermes: restore prior suite grouping policy"
|
|
git push origin HEAD:main
|
|
flux reconcile kustomization hermes --namespace flux-system --with-source
|
|
```
|
|
|
|
The first live 363-case multi-pass run at the previous 900-second / USD 5
|
|
guards stopped after 747.513 seconds during large-family review. Its actual CLI
|
|
final subtype was `error_max_budget_usd`; three full passes had completed. The CLI
|
|
reported an aggregate estimate of 5.62765, showing the documented within-generation
|
|
overshoot. It was not a timeout, turn-limit failure, or accepted partial answer.
|
|
The adapter now reports this condition as `job_cost_budget_exhausted`.
|
|
|
|
## Synthetic acceptance measurements
|
|
|
|
The suite inputs are synthetic fixtures, never roster-derived material. Tests ran
|
|
through authenticated TLS from `titan-jh` (`192.168.22.8`) to the LAN ingress
|
|
`192.168.22.50:443`, with hostname verification, proxy bypass, and no redirects.
|
|
This verifies that LAN test path; it does not replace a WSL connectivity test.
|
|
|
|
| Cases | Request bytes | Wall seconds | Model passes / CLI turns | Natural families | Final tasks | Natural / final singletons | Input / output tokens |
|
|
| --- | ---: | ---: | --- | ---: | ---: | --- | --- |
|
|
| 14 | 9,847 | 61.805 | 3 / 6 | 9 | 9 | 4 / 4 | 18,216 / 5,794 |
|
|
| 75 | 55,984 | 274.620 | 5 / 12 | 9 | 21 | 3 / 3 | 174,942 / 28,950 |
|
|
| 38 balanced fixture | 23,871 | 231.337 | 5 / 15 | 4 | 10 | 0 / 0 | 139,789 / 24,681 |
|
|
| 363, current required-alias schema | 273,763 | 842.777 | 5 / 12 | 9 | 75 | 3 / 3 | 983,491 / 91,932 |
|
|
| 363, old limits | 273,761 | 747.513 | 3 completed; pass 4 failed | Not finalized | No result | Not finalized | 629,141 / 64,424 |
|
|
|
|
Token totals include all completed CLI turns across the job, not unique source
|
|
size. The old-limit 363 result includes reported failed-invocation usage and was
|
|
never returned as completed. The earlier 14/75/38 and old-limit 363 CLI cost
|
|
estimates were 0.23593, 1.59846, 1.31597, and 5.62765 respectively; these are not subscription bills.
|
|
|
|
The current 363-case run reported a CLI cost estimate of 7.215755 and a complete
|
|
response envelope of 122,115 bytes. Its five passes took 80.962, 98.247, 350.114,
|
|
148.165, and 165.211 seconds, with 2/2/4/2/2 CLI turns. Six coherent 60-case
|
|
families became twelve five-case tasks each, plus three genuine singletons.
|
|
The 135 authorized status samples covered all five stages and recorded 59 changes
|
|
in CLI output activity. No compaction or truncation was reported. No terminal rate-limit
|
|
error or gateway job retry was reported; structured-output CLI turns are recorded
|
|
separately and do not imply provider-network retries.
|
|
|
|
For every completed run above, natural-family pair precision and recall were
|
|
1.0 against the independently defined implementation patterns, exact alias
|
|
coverage passed, and final groups stayed pure to those patterns. No incorrect
|
|
merges, unnecessary semantic splits, or cross-family description attribution
|
|
were observed during review. Equivalent text under separate aliases remained
|
|
separate objectives. Natural-family quality was scored before capacity division;
|
|
a cap-compliant result alone does not prove semantic correctness.
|
|
|
|
The 38-case fixture interleaves four reset-related mechanisms with shared generic
|
|
subject descriptions and distinct success criteria: RPC/stub observations, pulse
|
|
and oscilloscope timing, offline report parsing, and concurrent queue instrumentation.
|
|
It yielded natural sizes 6/7/11/14 and final sizes 3+3 / 4+3 / 4+4+3 / 5+5+4.
|
|
Thirty-eight authorized status samples showed five stages, changing heartbeats,
|
|
and nineteen CLI activity changes without exposing event bodies.
|
|
|
|
The actual 14/75/38/363 independent proposals agreed on membership. Disagreement
|
|
reconciliation, genuine semantic subdivision, reversal of an unjustified split,
|
|
and naming-collision rejection are separately tested with controlled mocked
|
|
responses; do not describe those as observed live-provider disagreements.
|
|
|
|
The public 363 fixture has description lengths 48-245 characters (mean 234.5),
|
|
preconditions 67-143 (mean 137.2), success criteria 73-135 (mean 131.2), and case
|
|
labels 7-15. It interleaves six mechanisms, includes identical distant case text,
|
|
and has distinctive beginning/middle/end objectives. These measurements describe
|
|
the tested material; longer or more ambiguous real suites may take more work.
|
|
|
|
An installed native CLI loopback probe captured all complete case fields and full
|
|
system instructions in every emitted provider request for 14/75/363 cases, with
|
|
64,000 output limits. It performed 3/5/5 passes and deterministic final sizing to
|
|
9/21/75 tasks using mock responses. This verifies transport and orchestration,
|
|
not live semantic quality. Live successful results reported no compaction or
|
|
truncation and passed exact coverage and source-quote support checks. These checks
|
|
cannot prove the provider internally attended to every field; exact tokenization
|
|
and sufficient output reservation for arbitrary suites remain unverified.
|
|
|
|
The updated code passes 109 synthetic/mocked regression tests. Kubernetes render,
|
|
client dry-run, and Flux diff checks passed before deployment. Policy controls,
|
|
case coverage validation, scoped authentication, and provider pinning remain in
|
|
place. The invalid-token check returned 401 and same-key submission replay reused
|
|
the same job. No automatic real-data rerun was performed.
|
|
|
|
The second 363-case attempt, under the 1200-second / USD 10 guards, failed after
|
|
228.260 seconds in proposal B with `invalid_case_assignments`. Both invocations
|
|
returned successful CLI results, but the second partition failed server membership
|
|
validation. Its reported total was 281,227 input / 23,796 output tokens and 2.001035
|
|
in CLI estimated-cost accounting. The previous validator did not retain counts of
|
|
missing, repeated, or unknown assignments, and transient output was cleaned; the
|
|
specific mismatch cannot be recovered. No partial partition was accepted. The new
|
|
required-alias schema and count-only failure diagnostics address this failure mode.
|
|
|
|
The updated native CLI loopback test deliberately omitted one required alias from
|
|
its first 14-case structured output. The installed CLI rejected that tool payload
|
|
and requested a correction, then the complete job passed. This demonstrates actual
|
|
schema-repair behavior within the existing six-turn headroom. The updated
|
|
14/75/363 transport runs made 4/5/5 mock provider requests across 3/5/5 model passes;
|
|
all transmitted the full source and system instructions. No hosted inference was
|
|
used for this transport/repair test.
|
|
|
|
The 14/75 and balanced-38 live measurements preceded the required-alias internal
|
|
serialization change. The final serialization was exercised by native loopback at
|
|
all three public sizes and by the successful hosted 363-case five-pass job. The
|
|
engineering objective, model, effort, cap, and public result shape stayed the same.
|
|
|
|
The complete suite is supported directly for the tested 273,763-byte request.
|
|
This is evidence for this fixture, not a guarantee that every 363-case request
|
|
will fit the same time/output allowance. There is no remaining blocker for starting
|
|
a new laptop attempt with the approved Claude credential and adequate job limits.
|
|
The result still requires the client's independent coverage audit and engineering
|
|
review. No real job was run by these checks.
|
|
|
|
Full synthetic results, review summaries, safe per-pass diagnostics, input/response
|
|
sizes, source-field length distributions, failed attempts, and native transport
|
|
measurements are in [the acceptance evidence](evidence/hermes_suite_multipass_20260929.json).
|
|
Routine planner logs contained only allowed operational fields, the completed
|
|
job's SQLite metadata contained no result/review/case content, and temporary CLI
|
|
job directories were empty after completion.
|
|
|
|
Execution revision `suite-multipass-v4-20260929` raises only the shared job defaults
|
|
and maximums to 1800 seconds and USD 30 in CLI estimate accounting. Explicit lower
|
|
client limits remain effective. The existing request fields are unchanged. No
|
|
inference or regression tests were rerun for this limit-only update, as requested;
|
|
the acceptance measurements above retain their original revisions and bounds.
|