168 lines
10 KiB
Markdown
168 lines
10 KiB
Markdown
# Multi-pass implementation grouping
|
|
|
|
The suite planner keeps the same HTTPS endpoint, request fields, authentication,
|
|
provider permissions, generalized-data approval, and asynchronous job lifecycle.
|
|
The required result object remains `result.groups`, with `name`, `description`,
|
|
and `members`. Final groups now contain **one to five** cases and names are at
|
|
most **64 characters**, unique after whitespace and case normalization.
|
|
|
|
## Revisions and process
|
|
|
|
- Configuration: `suite-v6-20260929` (HTTP compatibility identifier).
|
|
- Policy: `implementation-five-v1-20260929`.
|
|
- Prompt: `implementation-proximity-multipass-v3-20260929`.
|
|
- Execution: `suite-multipass-v2-20260929`.
|
|
|
|
A single server-side job performs:
|
|
|
|
1. Proposal A on the complete suite, in stable alias order, without a size cap.
|
|
2. Proposal B in a reproducible SHA-256 alias order, in a separate fresh CLI
|
|
invocation. It sees no proposal A or earlier conversation. Both use exactly
|
|
the same implementation-effort system instructions and output schema.
|
|
3. Whole-suite reconciliation using both memberships and all original fields.
|
|
Success criteria remain primary evidence. There is no majority voting or
|
|
transitive closure of pairwise similarity. Natural families are still uncapped.
|
|
4. If any reconciled family exceeds five, review every such original family using
|
|
all case fields, distinguishing different machinery from cheap variations.
|
|
5. One additional bounded audit of those decisions, including descendants that
|
|
were split into smaller groups. It may reverse an unjustified split or refine
|
|
a broad family, but cannot cross the reconciled family boundaries. There is
|
|
no agreement loop.
|
|
|
|
The final two passes record explicit keep/split rationales and exact short
|
|
source-field quotations. The code verifies that evidence belongs to an assigned
|
|
alias and is present verbatim in that source field. This checks support provenance,
|
|
not the correctness of a model's engineering interpretation. Unknown implementation
|
|
details remain concise uncertainty, not invented equipment or procedures.
|
|
|
|
Each invocation retains six CLI turns for structured output. Model review passes
|
|
and CLI turns are separate counters. All calls use the originally selected provider
|
|
and pinned model. The model remains `claude-opus-4-8[1m]`, with canonical runtime
|
|
identity checked as `claude-opus-4-8`, firstParty, medium effort, reported 1M context
|
|
and 64K output ceiling. No new provider fallback or tools are enabled.
|
|
|
|
## Deterministic task sizing and names
|
|
|
|
For each remaining coherent large family, `k = ceil(n / 5)` and
|
|
`base, remainder = divmod(n, k)`. The first remainder parts have base+1 cases;
|
|
the rest have base. Thus 6 becomes 3+3, 7 becomes 4+3, 11 becomes 4+4+3, and
|
|
14 becomes 5+5+4. No leftover singleton is created by a five-at-a-time slice.
|
|
|
|
The semantic review supplies internal variation sets. A deterministic packer uses
|
|
those hints and stable complete-record hashes/aliases to fill the calculated
|
|
sizes. Equivalent content remains distinct by alias. Hint order, source order,
|
|
and member order do not cause gratuitous reshuffling. Different model decisions
|
|
can still change membership between fresh jobs; determinism is conditional on the
|
|
same natural families and variation hints, not a claim that inference is deterministic.
|
|
|
|
Natural names must be concise, meaningful, distinct, and at most 56 characters,
|
|
leaving room for numbering within the 64-character final limit. Collisions or
|
|
invalid names fail validation; the service never truncates them or adds numbers
|
|
to unrelated families. Capacity-only parts have the exact same base name:
|
|
`Reset recovery (1/3)`, `Reset recovery (2/3)`, `Reset recovery (3/3)`.
|
|
Their descriptions identify a work-size division and use only the reviewed common
|
|
work applicable to all members. No other part's specific objectives are copied.
|
|
|
|
Every intermediate partition and final result is checked for exact alias coverage.
|
|
Final size, names, balanced part sizes, and campaign/suite ownership are checked
|
|
programmatically. Incomplete or invalid work is never returned as completed.
|
|
|
|
## Review artifact and operational metadata
|
|
|
|
`GET /suite-planning/v1/jobs/<id>/result` returns an optional top-level
|
|
**`review_summary`**, alongside the unchanged **`result`** object. It contains:
|
|
|
|
- `proposal_disagreements`: affected aliases, differing pair count, and both full
|
|
membership partitions in compressed form. Group names/numbers are not compared.
|
|
- `reconciled_families`: the first natural partition with concise rationale,
|
|
common work, exact source-field support, and uncertainty.
|
|
- `large_family_review` and `decision_audit`: explicit decisions for each original
|
|
oversized membership set, including semantic-split versus coherent-keep reasons.
|
|
- `natural_families`: the final conceptual partition before task sizing.
|
|
- `capacity_divisions`: base names, calculated sizes, final part names, common
|
|
implementation rationale, and remaining uncertainty.
|
|
- `unresolved_uncertainties` and `counts`: separate natural-family/final-task counts
|
|
and natural/final singleton counts. Size compliance is not semantic-quality evidence.
|
|
|
|
Review content lives only in the authorized in-memory result cache, with the same
|
|
one-hour TTL and restart loss as results. It is absent from routine logs, status
|
|
responses, and SQLite job metadata. It contains source-derived material and exact
|
|
quotations, so it is not an export-safe ClickUp artifact. No new hierarchy level
|
|
or client request fields are introduced.
|
|
|
|
Status/result metadata includes `policy_revision`, `execution_revision`,
|
|
`prompt_revision`, `model_pass_count`, `natural_family_count`, `final_task_count`,
|
|
`singleton_statistics`, aggregate usage/cost, and safe per-invocation `passes`.
|
|
Each pass records its stage, actual model, system/schema hashes, ordering hash,
|
|
capacity check, remaining budgets, usage, timing, and CLI diagnostics. The top-level
|
|
`cli_diagnostics` concerns the last invocation; `turns` is the sum of reported CLI
|
|
turns, while `model_pass_count` counts complete model review invocations.
|
|
`execution_progress.current_pass` reports the current stage during a running job.
|
|
While a CLI invocation runs, progress updates every five seconds with `heartbeat_at`,
|
|
`pass_elapsed_seconds`, `job_elapsed_seconds`, `job_remaining_seconds`,
|
|
`completed_model_passes`, `maximum_model_passes`, `cli_running`, `cli_output_bytes`,
|
|
and `last_cli_activity_seconds_ago`. The heartbeat means the worker is alive; a
|
|
change in output bytes means CLI activity. Neither is a percentage complete or
|
|
proof that a particular case has been reasoned about. No event bodies or reasoning
|
|
text are exposed. Poll the existing status URL every five seconds to display it.
|
|
|
|
## Shared bounds and failure behavior
|
|
|
|
The user-approved 1200-second maximum job deadline and USD 10 CLI estimated-cost
|
|
guard apply to **all calls combined**, including routing time. These are also the
|
|
defaults when omitted. Explicit lower client limits remain effective. This guard
|
|
uses Claude CLI estimated-cost accounting; it is not subscription billing. Each invocation receives
|
|
only the remaining time and estimated-cost allowance. Available usage/cost is
|
|
accumulated after every call; unknown cost accounting stops further hosted calls.
|
|
No partial partition substitutes for unfinished passes. The CLI's estimated-cost
|
|
ceiling can overshoot within one generation, as before; this is not a guaranteed
|
|
subscription billing limit. An observed overrun fails the job rather than allowing
|
|
further calls or reporting successful completion.
|
|
|
|
Request and complete response envelopes stay bounded at 1 MiB. Every expanded pass
|
|
is checked before dispatch, including original source, proposals/reviews, system
|
|
instructions, and schema. The same conservative byte bound and six possible CLI
|
|
outputs are reserved against context. Exact tokenization and output reservation
|
|
remain unverified estimates; actual incomplete generation fails explicitly.
|
|
Preflight checks initial and minimum reconciliation capacity; it identifies that
|
|
actual later-pass capacity will be checked before each invocation. It does not
|
|
claim to know model-generated proposal sizes in advance.
|
|
|
|
The existing 8K-context local route cannot admit the new reconciliation schema and
|
|
instructions. Local-only requests fail with capacity errors and never start hosted
|
|
jobs. The separate `/local-model` endpoint and its limits are unchanged. No bounds
|
|
other than the explicitly approved time/cost guards were increased, and no provider
|
|
was switched after an error.
|
|
|
|
Failures include existing CLI diagnostics and a fixed `review_pass`, completed-pass
|
|
count, and available aggregate usage. New explicit codes include `pass_capacity`,
|
|
`pass_request_too_large`, `job_time_budget_exhausted`, `job_cost_budget_exhausted`,
|
|
`budget_accounting_unavailable`, and bounded review/assignment validation errors.
|
|
Raw provider messages, source text, and generated explanations are not error details.
|
|
|
|
The laptop submits one request with its existing scoped token, then polls the same
|
|
job. Retry a timed-out submission using the same key/body; an intentional fresh
|
|
attempt needs a new Idempotency-Key. Include policy/prompt/execution revisions in
|
|
local cache provenance. No real suite is submitted by deployment checks.
|
|
|
|
## Verification and rollback
|
|
|
|
Synthetic unit tests cover the balanced-size examples, semantic subdivisions and
|
|
a bounded reversal, differing proposals, duplicate text, name normalization and
|
|
collisions, source support, deadline/cost exhaustion, cancellation, privacy,
|
|
authorization, idempotency, and 14/75/363-case capacity and sizing. Native CLI
|
|
loopback and authenticated LAN/provider acceptance measurements are recorded after
|
|
deployment. Mocked outputs alone do not establish model quality or throughput.
|
|
|
|
Deploy or roll back through Git and Flux. Revert only the multi-pass commits to
|
|
restore the preceding single-pass policy while preserving the earlier CLI diagnostic
|
|
fix. Wait for active jobs before a worker restart because completed results are
|
|
held in memory. No Vault credential or ingress/routing changes are required.
|
|
|
|
The first live 363-case multi-pass run at the previous 900-second / USD 5
|
|
guards stopped after 747.513 seconds during large-family review. Its actual CLI
|
|
final subtype was `error_max_budget_usd`; three full passes had completed. The CLI
|
|
reported an aggregate estimate of 5.62765, showing the documented within-generation
|
|
overshoot. It was not a timeout, turn-limit failure, or accepted partial answer.
|
|
The adapter now reports this condition as `job_cost_budget_exhausted`.
|