277 lines
17 KiB
Markdown
277 lines
17 KiB
Markdown
# Multi-pass implementation grouping
|
|
|
|
The suite planner keeps the same HTTPS endpoint, request fields, authentication,
|
|
provider permissions, generalized-data approval, and asynchronous job lifecycle.
|
|
The required result object remains `result.groups`, with `name`, `description`,
|
|
and `members`. Final groups now contain **one to five** cases and names are at
|
|
most **64 characters**, unique after whitespace and case normalization.
|
|
|
|
## Revisions and process
|
|
|
|
- Configuration: `suite-v6-20260929` (HTTP compatibility identifier).
|
|
- Policy: `implementation-five-v1-20260929`.
|
|
- Prompt: `implementation-proximity-multipass-v4-20260929`.
|
|
- Execution: `suite-multipass-v3-20260929`.
|
|
|
|
A single server-side job performs:
|
|
|
|
1. Proposal A on the complete suite, in stable alias order, without a size cap.
|
|
2. Proposal B in a reproducible SHA-256 alias order, in a separate fresh CLI
|
|
invocation. It sees no proposal A or earlier conversation. Both use exactly
|
|
the same implementation-effort system instructions and output schema.
|
|
3. Whole-suite reconciliation using both memberships and all original fields.
|
|
Success criteria remain primary evidence. There is no majority voting or
|
|
transitive closure of pairwise similarity. Natural families are still uncapped.
|
|
4. If any reconciled family exceeds five, review every such original family using
|
|
all case fields, distinguishing different machinery from cheap variations.
|
|
5. One additional bounded audit of those decisions, including descendants that
|
|
were split into smaller groups. It may reverse an unjustified split or refine
|
|
a broad family, but cannot cross the reconciled family boundaries. There is
|
|
no agreement loop.
|
|
|
|
The final two passes record explicit keep/split rationales and exact short
|
|
source-field quotations. The code verifies that evidence belongs to an assigned
|
|
alias and is present verbatim in that source field. This checks support provenance,
|
|
not the correctness of a model's engineering interpretation. Unknown implementation
|
|
details remain concise uncertainty, not invented equipment or procedures.
|
|
|
|
The internal structured-output schema requires one property for each exact source
|
|
alias. Its value selects a natural-family index and, in review passes, a variation
|
|
index. The server reconstructs memberships and runs the normal independent checks.
|
|
Review decisions also use required keys derived from exact original memberships.
|
|
This avoids free-form membership lists silently omitting or repeating aliases.
|
|
The public groups schema is unchanged. Invalid keys, indexes, empty groups, and
|
|
missing review decisions still fail closed; no missing assignment is fabricated.
|
|
|
|
Each invocation retains six CLI turns for structured output. Model review passes
|
|
and CLI turns are separate counters. All calls use the originally selected provider
|
|
and pinned model. The model remains `claude-opus-4-8[1m]`, with canonical runtime
|
|
identity checked as `claude-opus-4-8`, firstParty, medium effort, reported 1M context
|
|
and 64K output ceiling. The server invokes native Claude Code CLI 2.1.226 using
|
|
the existing first-party OAuth account, not a separately configured API-key
|
|
account. The CLI itself communicates with Anthropic over HTTPS. No new provider
|
|
fallback or tools are enabled.
|
|
|
|
## Deterministic task sizing and names
|
|
|
|
For each remaining coherent large family, `k = ceil(n / 5)` and
|
|
`base, remainder = divmod(n, k)`. The first remainder parts have base+1 cases;
|
|
the rest have base. Thus 6 becomes 3+3, 7 becomes 4+3, 11 becomes 4+4+3, and
|
|
14 becomes 5+5+4. No leftover singleton is created by a five-at-a-time slice.
|
|
|
|
The semantic review supplies internal variation sets. A deterministic packer uses
|
|
those hints and stable complete-record hashes/aliases to fill the calculated
|
|
sizes. Equivalent content remains distinct by alias. Hint order, source order,
|
|
and member order do not cause gratuitous reshuffling. Different model decisions
|
|
can still change membership between fresh jobs; determinism is conditional on the
|
|
same natural families and variation hints, not a claim that inference is deterministic.
|
|
|
|
Natural names must be concise, meaningful, distinct, and at most 56 characters,
|
|
leaving room for numbering within the 64-character final limit. Collisions or
|
|
invalid names fail validation; the service never truncates them or adds numbers
|
|
to unrelated families. Capacity-only parts have the exact same base name:
|
|
`Reset recovery (1/3)`, `Reset recovery (2/3)`, `Reset recovery (3/3)`.
|
|
Their descriptions identify a work-size division and use only the reviewed common
|
|
work applicable to all members. No other part's specific objectives are copied.
|
|
|
|
Every intermediate partition and final result is checked for exact alias coverage.
|
|
Final size, names, balanced part sizes, and campaign/suite ownership are checked
|
|
programmatically. Incomplete or invalid work is never returned as completed.
|
|
|
|
## Review artifact and operational metadata
|
|
|
|
`GET /suite-planning/v1/jobs/<id>/result` returns an optional top-level
|
|
**`review_summary`**, alongside the unchanged **`result`** object. It contains:
|
|
|
|
- `proposal_disagreements`: affected aliases, differing pair count, and both full
|
|
membership partitions in compressed form. Group names/numbers are not compared.
|
|
- `reconciled_families`: the first natural partition with concise rationale,
|
|
common work, exact source-field support, and uncertainty.
|
|
- `large_family_review` and `decision_audit`: explicit decisions for each original
|
|
oversized membership set, including semantic-split versus coherent-keep reasons.
|
|
- `natural_families`: the final conceptual partition before task sizing.
|
|
- `capacity_divisions`: base names, calculated sizes, final part names, common
|
|
implementation rationale, and remaining uncertainty.
|
|
- `unresolved_uncertainties` and `counts`: separate natural-family/final-task counts
|
|
and natural/final singleton counts. Size compliance is not semantic-quality evidence.
|
|
|
|
Review content lives only in the authorized in-memory result cache, with the same
|
|
one-hour TTL and restart loss as results. It is absent from routine logs, status
|
|
responses, and SQLite job metadata. It contains source-derived material and exact
|
|
quotations, so it is not an export-safe ClickUp artifact. No new hierarchy level
|
|
or client request fields are introduced.
|
|
|
|
Status/result metadata includes `policy_revision`, `execution_revision`,
|
|
`prompt_revision`, `model_pass_count`, `natural_family_count`, `final_task_count`,
|
|
`singleton_statistics`, aggregate usage/cost, and safe per-invocation `passes`.
|
|
Each pass records its stage, actual model, system/schema hashes, ordering hash,
|
|
capacity check, remaining budgets, usage, timing, and CLI diagnostics. The top-level
|
|
`cli_diagnostics` concerns the last invocation; `turns` is the sum of reported CLI
|
|
turns, while `model_pass_count` counts complete model review invocations.
|
|
`execution_progress.current_pass` reports the current stage during a running job.
|
|
While a CLI invocation runs, progress updates every five seconds with `heartbeat_at`,
|
|
`pass_elapsed_seconds`, `job_elapsed_seconds`, `job_remaining_seconds`,
|
|
`completed_model_passes`, `maximum_model_passes`, `cli_running`, `cli_output_bytes`,
|
|
and `last_cli_activity_seconds_ago`. The heartbeat means the worker is alive; a
|
|
change in output bytes means CLI activity. Neither is a percentage complete or
|
|
proof that a particular case has been reasoned about. No event bodies or reasoning
|
|
text are exposed. `cost_used_usd_estimate` counts completed invocations; the
|
|
in-flight invocation is not included until its terminal usage is available.
|
|
Poll the existing status URL every five seconds to display progress.
|
|
|
|
To inspect progress from WSL, keep the existing bearer token private and use the
|
|
job ID returned by submission. This command only reads status:
|
|
|
|
```bash
|
|
curl -q --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
|
|
--connect-timeout 10 --max-time 45 --fail-with-body --silent --show-error \
|
|
--config <(printf 'header = "Authorization: Bearer %s"\n' "$SUITE_PLANNING_TOKEN") \
|
|
"https://worker.bstein.dev/suite-planning/v1/jobs/$JOB_ID"
|
|
```
|
|
|
|
An existing Python poller can display `envelope["execution_progress"]` after each
|
|
status read. No separate pass submissions, new request fields, or client-side
|
|
reconciliation are required. Prefer the actual stage and elapsed time over an
|
|
invented percentage, because different passes have very different runtimes.
|
|
|
|
## Shared bounds and failure behavior
|
|
|
|
The user-approved 1200-second maximum job deadline and USD 10 CLI estimated-cost
|
|
guard apply to **all calls combined**, including routing time. These are also the
|
|
defaults when omitted. Explicit lower client limits remain effective. This guard
|
|
uses Claude CLI estimated-cost accounting; it is not subscription billing. Each invocation receives
|
|
only the remaining time and estimated-cost allowance. Available usage/cost is
|
|
accumulated after every call; unknown cost accounting stops further hosted calls.
|
|
No partial partition substitutes for unfinished passes. The CLI's estimated-cost
|
|
ceiling can overshoot within one generation, as before; this is not a guaranteed
|
|
subscription billing limit. An observed overrun fails the job rather than allowing
|
|
further calls or reporting successful completion.
|
|
|
|
Request and complete response envelopes stay bounded at 1 MiB. Every expanded pass
|
|
is checked before dispatch, including original source, proposals/reviews, system
|
|
instructions, and schema. The same conservative byte bound and six possible CLI
|
|
outputs are reserved against context. Exact tokenization and output reservation
|
|
remain unverified estimates; actual incomplete generation fails explicitly.
|
|
Preflight checks initial and minimum reconciliation capacity; it identifies that
|
|
actual later-pass capacity will be checked before each invocation. It does not
|
|
claim to know model-generated proposal sizes in advance.
|
|
|
|
The existing 8K-context local route cannot admit the new reconciliation schema and
|
|
instructions. Local-only requests fail with capacity errors and never start hosted
|
|
jobs. The separate `/local-model` endpoint and its limits are unchanged. No bounds
|
|
other than the explicitly approved time/cost guards were increased, and no provider
|
|
was switched after an error.
|
|
|
|
Failures include existing CLI diagnostics and a fixed `review_pass`, completed-pass
|
|
count, and available aggregate usage. New explicit codes include `pass_capacity`,
|
|
`pass_request_too_large`, `job_time_budget_exhausted`, `job_cost_budget_exhausted`,
|
|
`budget_accounting_unavailable`, and bounded review/assignment validation errors.
|
|
Raw provider messages, source text, and generated explanations are not error details.
|
|
|
|
The laptop submits one request with its existing scoped token, then polls the same
|
|
job. Retry a timed-out submission using the same key/body; an intentional fresh
|
|
attempt needs a new Idempotency-Key. Include policy/prompt/execution revisions in
|
|
local cache provenance. No real suite is submitted by deployment checks.
|
|
|
|
## Verification and rollback
|
|
|
|
Synthetic unit tests cover the balanced-size examples, semantic subdivisions and
|
|
a bounded reversal, differing proposals, duplicate text, name normalization and
|
|
collisions, source support, deadline/cost exhaustion, cancellation, privacy,
|
|
authorization, idempotency, and 14/75/363-case capacity and sizing. Native CLI
|
|
loopback and authenticated LAN/provider acceptance measurements are recorded after
|
|
deployment. Mocked outputs alone do not establish model quality or throughput.
|
|
|
|
Deploy or roll back through Git and Flux. Revert only the multi-pass commits to
|
|
restore the preceding single-pass policy while preserving the earlier CLI diagnostic
|
|
fix. Wait for active jobs before a worker restart because completed results are
|
|
held in memory. No Vault credential or ingress/routing changes are required.
|
|
|
|
The policy implementation is commit `7af4f4f2`; the approved 20-minute / USD 10
|
|
estimate guard and progress reporting are commit `15c05cb7`. To restore the prior
|
|
single-pass implementation after active jobs finish, revert both through the normal
|
|
Git deployment branch, preserving the earlier six-turn diagnostics fix:
|
|
|
|
```bash
|
|
git revert --no-commit 15c05cb7 7af4f4f2
|
|
git commit -m "hermes: restore prior suite grouping policy"
|
|
git push origin HEAD:main
|
|
flux reconcile kustomization hermes --namespace flux-system --with-source
|
|
```
|
|
|
|
The first live 363-case multi-pass run at the previous 900-second / USD 5
|
|
guards stopped after 747.513 seconds during large-family review. Its actual CLI
|
|
final subtype was `error_max_budget_usd`; three full passes had completed. The CLI
|
|
reported an aggregate estimate of 5.62765, showing the documented within-generation
|
|
overshoot. It was not a timeout, turn-limit failure, or accepted partial answer.
|
|
The adapter now reports this condition as `job_cost_budget_exhausted`.
|
|
|
|
## Synthetic acceptance measurements
|
|
|
|
The suite inputs are synthetic fixtures, never roster-derived material. Tests ran
|
|
through authenticated TLS from `titan-jh` (`192.168.22.8`) to the LAN ingress
|
|
`192.168.22.50:443`, with hostname verification, proxy bypass, and no redirects.
|
|
This verifies that LAN test path; it does not replace a WSL connectivity test.
|
|
|
|
| Cases | Request bytes | Wall seconds | Model passes / CLI turns | Natural families | Final tasks | Natural / final singletons | Input / output tokens |
|
|
| --- | ---: | ---: | --- | ---: | ---: | --- | --- |
|
|
| 14 | 9,847 | 61.805 | 3 / 6 | 9 | 9 | 4 / 4 | 18,216 / 5,794 |
|
|
| 75 | 55,984 | 274.620 | 5 / 12 | 9 | 21 | 3 / 3 | 174,942 / 28,950 |
|
|
| 38 balanced fixture | 23,871 | 231.337 | 5 / 15 | 4 | 10 | 0 / 0 | 139,789 / 24,681 |
|
|
| 363, old limits | 273,761 | 747.513 | 3 completed; pass 4 failed | Not finalized | No result | Not finalized | 629,141 / 64,424 |
|
|
|
|
Token totals include all completed CLI turns across the job, not unique source
|
|
size. The old-limit 363 result includes reported failed-invocation usage and was
|
|
never returned as completed. CLI estimated costs were 0.23593, 1.59846, 1.31597,
|
|
and 5.62765 respectively; these are not subscription bills.
|
|
|
|
For every completed run above, natural-family pair precision and recall were
|
|
1.0 against the independently defined implementation patterns, exact alias
|
|
coverage passed, and final groups stayed pure to those patterns. No incorrect
|
|
merges, unnecessary semantic splits, or cross-family description attribution
|
|
were observed during review. Equivalent text under separate aliases remained
|
|
separate objectives. Natural-family quality was scored before capacity division;
|
|
a cap-compliant result alone does not prove semantic correctness.
|
|
|
|
The 38-case fixture interleaves four reset-related mechanisms with shared generic
|
|
subject descriptions and distinct success criteria: RPC/stub observations, pulse
|
|
and oscilloscope timing, offline report parsing, and concurrent queue instrumentation.
|
|
It yielded natural sizes 6/7/11/14 and final sizes 3+3 / 4+3 / 4+4+3 / 5+5+4.
|
|
Thirty-eight authorized status samples showed five stages, changing heartbeats,
|
|
and nineteen CLI activity changes without exposing event bodies.
|
|
|
|
The actual 14/75/38 independent proposals agreed on membership. Disagreement
|
|
reconciliation, genuine semantic subdivision, reversal of an unjustified split,
|
|
and naming-collision rejection are separately tested with controlled mocked
|
|
responses; do not describe those as observed live-provider disagreements.
|
|
|
|
The public 363 fixture has description lengths 48-245 characters (mean 234.5),
|
|
preconditions 67-143 (mean 137.2), success criteria 73-135 (mean 131.2), and case
|
|
labels 7-15. It interleaves six mechanisms, includes identical distant case text,
|
|
and has distinctive beginning/middle/end objectives. These measurements describe
|
|
the tested material; longer or more ambiguous real suites may take more work.
|
|
|
|
An installed native CLI loopback probe captured all complete case fields and full
|
|
system instructions in every emitted provider request for 14/75/363 cases, with
|
|
64,000 output limits. It performed 3/5/5 passes and deterministic final sizing to
|
|
9/21/75 tasks using mock responses. This verifies transport and orchestration,
|
|
not live semantic quality. Live successful results reported no compaction or
|
|
truncation and passed exact coverage and source-quote support checks. These checks
|
|
cannot prove the provider internally attended to every field; exact tokenization
|
|
and sufficient output reservation for arbitrary suites remain unverified.
|
|
|
|
The updated code passes 109 synthetic/mocked regression tests. Kubernetes render,
|
|
client dry-run, and Flux diff checks passed before deployment. Policy controls,
|
|
case coverage validation, scoped authentication, and provider pinning remain in
|
|
place. The invalid-token check returned 401 and same-key submission replay reused
|
|
the same job. No automatic real-data rerun was performed.
|
|
|
|
The second 363-case attempt, under the 1200-second / USD 10 guards, failed after
|
|
228.260 seconds in proposal B with `invalid_case_assignments`. Both invocations
|
|
returned successful CLI results, but the second partition failed server membership
|
|
validation. Its reported total was 281,227 input / 23,796 output tokens and 2.001035
|
|
in CLI estimated-cost accounting. The previous validator did not retain counts of
|
|
missing, repeated, or unknown assignments, and transient output was cleaned; the
|
|
specific mismatch cannot be recovered. No partial partition was accepted. The new
|
|
required-alias schema and count-only failure diagnostics address this failure mode.
|