10 KiB
Multi-pass implementation grouping
The suite planner keeps the same HTTPS endpoint, request fields, authentication,
provider permissions, generalized-data approval, and asynchronous job lifecycle.
The required result object remains result.groups, with name, description,
and members. Final groups now contain one to five cases and names are at
most 64 characters, unique after whitespace and case normalization.
Revisions and process
- Configuration:
suite-v6-20260929(HTTP compatibility identifier). - Policy:
implementation-five-v1-20260929. - Prompt:
implementation-proximity-multipass-v3-20260929. - Execution:
suite-multipass-v2-20260929.
A single server-side job performs:
- Proposal A on the complete suite, in stable alias order, without a size cap.
- Proposal B in a reproducible SHA-256 alias order, in a separate fresh CLI invocation. It sees no proposal A or earlier conversation. Both use exactly the same implementation-effort system instructions and output schema.
- Whole-suite reconciliation using both memberships and all original fields. Success criteria remain primary evidence. There is no majority voting or transitive closure of pairwise similarity. Natural families are still uncapped.
- If any reconciled family exceeds five, review every such original family using all case fields, distinguishing different machinery from cheap variations.
- One additional bounded audit of those decisions, including descendants that were split into smaller groups. It may reverse an unjustified split or refine a broad family, but cannot cross the reconciled family boundaries. There is no agreement loop.
The final two passes record explicit keep/split rationales and exact short source-field quotations. The code verifies that evidence belongs to an assigned alias and is present verbatim in that source field. This checks support provenance, not the correctness of a model's engineering interpretation. Unknown implementation details remain concise uncertainty, not invented equipment or procedures.
Each invocation retains six CLI turns for structured output. Model review passes
and CLI turns are separate counters. All calls use the originally selected provider
and pinned model. The model remains claude-opus-4-8[1m], with canonical runtime
identity checked as claude-opus-4-8, firstParty, medium effort, reported 1M context
and 64K output ceiling. No new provider fallback or tools are enabled.
Deterministic task sizing and names
For each remaining coherent large family, k = ceil(n / 5) and
base, remainder = divmod(n, k). The first remainder parts have base+1 cases;
the rest have base. Thus 6 becomes 3+3, 7 becomes 4+3, 11 becomes 4+4+3, and
14 becomes 5+5+4. No leftover singleton is created by a five-at-a-time slice.
The semantic review supplies internal variation sets. A deterministic packer uses those hints and stable complete-record hashes/aliases to fill the calculated sizes. Equivalent content remains distinct by alias. Hint order, source order, and member order do not cause gratuitous reshuffling. Different model decisions can still change membership between fresh jobs; determinism is conditional on the same natural families and variation hints, not a claim that inference is deterministic.
Natural names must be concise, meaningful, distinct, and at most 56 characters,
leaving room for numbering within the 64-character final limit. Collisions or
invalid names fail validation; the service never truncates them or adds numbers
to unrelated families. Capacity-only parts have the exact same base name:
Reset recovery (1/3), Reset recovery (2/3), Reset recovery (3/3).
Their descriptions identify a work-size division and use only the reviewed common
work applicable to all members. No other part's specific objectives are copied.
Every intermediate partition and final result is checked for exact alias coverage. Final size, names, balanced part sizes, and campaign/suite ownership are checked programmatically. Incomplete or invalid work is never returned as completed.
Review artifact and operational metadata
GET /suite-planning/v1/jobs/<id>/result returns an optional top-level
review_summary, alongside the unchanged result object. It contains:
proposal_disagreements: affected aliases, differing pair count, and both full membership partitions in compressed form. Group names/numbers are not compared.reconciled_families: the first natural partition with concise rationale, common work, exact source-field support, and uncertainty.large_family_reviewanddecision_audit: explicit decisions for each original oversized membership set, including semantic-split versus coherent-keep reasons.natural_families: the final conceptual partition before task sizing.capacity_divisions: base names, calculated sizes, final part names, common implementation rationale, and remaining uncertainty.unresolved_uncertaintiesandcounts: separate natural-family/final-task counts and natural/final singleton counts. Size compliance is not semantic-quality evidence.
Review content lives only in the authorized in-memory result cache, with the same one-hour TTL and restart loss as results. It is absent from routine logs, status responses, and SQLite job metadata. It contains source-derived material and exact quotations, so it is not an export-safe ClickUp artifact. No new hierarchy level or client request fields are introduced.
Status/result metadata includes policy_revision, execution_revision,
prompt_revision, model_pass_count, natural_family_count, final_task_count,
singleton_statistics, aggregate usage/cost, and safe per-invocation passes.
Each pass records its stage, actual model, system/schema hashes, ordering hash,
capacity check, remaining budgets, usage, timing, and CLI diagnostics. The top-level
cli_diagnostics concerns the last invocation; turns is the sum of reported CLI
turns, while model_pass_count counts complete model review invocations.
execution_progress.current_pass reports the current stage during a running job.
While a CLI invocation runs, progress updates every five seconds with heartbeat_at,
pass_elapsed_seconds, job_elapsed_seconds, job_remaining_seconds,
completed_model_passes, maximum_model_passes, cli_running, cli_output_bytes,
and last_cli_activity_seconds_ago. The heartbeat means the worker is alive; a
change in output bytes means CLI activity. Neither is a percentage complete or
proof that a particular case has been reasoned about. No event bodies or reasoning
text are exposed. Poll the existing status URL every five seconds to display it.
Shared bounds and failure behavior
The user-approved 1200-second maximum job deadline and USD 10 CLI estimated-cost guard apply to all calls combined, including routing time. These are also the defaults when omitted. Explicit lower client limits remain effective. This guard uses Claude CLI estimated-cost accounting; it is not subscription billing. Each invocation receives only the remaining time and estimated-cost allowance. Available usage/cost is accumulated after every call; unknown cost accounting stops further hosted calls. No partial partition substitutes for unfinished passes. The CLI's estimated-cost ceiling can overshoot within one generation, as before; this is not a guaranteed subscription billing limit. An observed overrun fails the job rather than allowing further calls or reporting successful completion.
Request and complete response envelopes stay bounded at 1 MiB. Every expanded pass is checked before dispatch, including original source, proposals/reviews, system instructions, and schema. The same conservative byte bound and six possible CLI outputs are reserved against context. Exact tokenization and output reservation remain unverified estimates; actual incomplete generation fails explicitly. Preflight checks initial and minimum reconciliation capacity; it identifies that actual later-pass capacity will be checked before each invocation. It does not claim to know model-generated proposal sizes in advance.
The existing 8K-context local route cannot admit the new reconciliation schema and
instructions. Local-only requests fail with capacity errors and never start hosted
jobs. The separate /local-model endpoint and its limits are unchanged. No bounds
other than the explicitly approved time/cost guards were increased, and no provider
was switched after an error.
Failures include existing CLI diagnostics and a fixed review_pass, completed-pass
count, and available aggregate usage. New explicit codes include pass_capacity,
pass_request_too_large, job_time_budget_exhausted, job_cost_budget_exhausted,
budget_accounting_unavailable, and bounded review/assignment validation errors.
Raw provider messages, source text, and generated explanations are not error details.
The laptop submits one request with its existing scoped token, then polls the same job. Retry a timed-out submission using the same key/body; an intentional fresh attempt needs a new Idempotency-Key. Include policy/prompt/execution revisions in local cache provenance. No real suite is submitted by deployment checks.
Verification and rollback
Synthetic unit tests cover the balanced-size examples, semantic subdivisions and a bounded reversal, differing proposals, duplicate text, name normalization and collisions, source support, deadline/cost exhaustion, cancellation, privacy, authorization, idempotency, and 14/75/363-case capacity and sizing. Native CLI loopback and authenticated LAN/provider acceptance measurements are recorded after deployment. Mocked outputs alone do not establish model quality or throughput.
Deploy or roll back through Git and Flux. Revert only the multi-pass commits to restore the preceding single-pass policy while preserving the earlier CLI diagnostic fix. Wait for active jobs before a worker restart because completed results are held in memory. No Vault credential or ingress/routing changes are required.
The first live 363-case multi-pass run at the previous 900-second / USD 5
guards stopped after 747.513 seconds during large-family review. Its actual CLI
final subtype was error_max_budget_usd; three full passes had completed. The CLI
reported an aggregate estimate of 5.62765, showing the documented within-generation
overshoot. It was not a timeout, turn-limit failure, or accepted partial answer.
The adapter now reports this condition as job_cost_budget_exhausted.