hermes: require one structured assignment per suite case

This commit is contained in:
jenkins 2026-09-29 14:10:03 -05:00
parent 15c05cb7fb
commit 9c327c6fd0
10 changed files with 342 additions and 27 deletions

View File

@ -10,8 +10,8 @@ most **64 characters**, unique after whitespace and case normalization.
- Configuration: `suite-v6-20260929` (HTTP compatibility identifier).
- Policy: `implementation-five-v1-20260929`.
- Prompt: `implementation-proximity-multipass-v3-20260929`.
- Execution: `suite-multipass-v2-20260929`.
- Prompt: `implementation-proximity-multipass-v4-20260929`.
- Execution: `suite-multipass-v3-20260929`.
A single server-side job performs:
@ -35,11 +35,22 @@ alias and is present verbatim in that source field. This checks support provenan
not the correctness of a model's engineering interpretation. Unknown implementation
details remain concise uncertainty, not invented equipment or procedures.
The internal structured-output schema requires one property for each exact source
alias. Its value selects a natural-family index and, in review passes, a variation
index. The server reconstructs memberships and runs the normal independent checks.
Review decisions also use required keys derived from exact original memberships.
This avoids free-form membership lists silently omitting or repeating aliases.
The public groups schema is unchanged. Invalid keys, indexes, empty groups, and
missing review decisions still fail closed; no missing assignment is fabricated.
Each invocation retains six CLI turns for structured output. Model review passes
and CLI turns are separate counters. All calls use the originally selected provider
and pinned model. The model remains `claude-opus-4-8[1m]`, with canonical runtime
identity checked as `claude-opus-4-8`, firstParty, medium effort, reported 1M context
and 64K output ceiling. No new provider fallback or tools are enabled.
and 64K output ceiling. The server invokes native Claude Code CLI 2.1.226 using
the existing first-party OAuth account, not a separately configured API-key
account. The CLI itself communicates with Anthropic over HTTPS. No new provider
fallback or tools are enabled.
## Deterministic task sizing and names
@ -104,7 +115,24 @@ While a CLI invocation runs, progress updates every five seconds with `heartbeat
and `last_cli_activity_seconds_ago`. The heartbeat means the worker is alive; a
change in output bytes means CLI activity. Neither is a percentage complete or
proof that a particular case has been reasoned about. No event bodies or reasoning
text are exposed. Poll the existing status URL every five seconds to display it.
text are exposed. `cost_used_usd_estimate` counts completed invocations; the
in-flight invocation is not included until its terminal usage is available.
Poll the existing status URL every five seconds to display progress.
To inspect progress from WSL, keep the existing bearer token private and use the
job ID returned by submission. This command only reads status:
```bash
curl -q --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
--connect-timeout 10 --max-time 45 --fail-with-body --silent --show-error \
--config <(printf 'header = "Authorization: Bearer %s"\n' "$SUITE_PLANNING_TOKEN") \
"https://worker.bstein.dev/suite-planning/v1/jobs/$JOB_ID"
```
An existing Python poller can display `envelope["execution_progress"]` after each
status read. No separate pass submissions, new request fields, or client-side
reconciliation are required. Prefer the actual stage and elapsed time over an
invented percentage, because different passes have very different runtimes.
## Shared bounds and failure behavior
@ -159,9 +187,90 @@ restore the preceding single-pass policy while preserving the earlier CLI diagno
fix. Wait for active jobs before a worker restart because completed results are
held in memory. No Vault credential or ingress/routing changes are required.
The policy implementation is commit `7af4f4f2`; the approved 20-minute / USD 10
estimate guard and progress reporting are commit `15c05cb7`. To restore the prior
single-pass implementation after active jobs finish, revert both through the normal
Git deployment branch, preserving the earlier six-turn diagnostics fix:
```bash
git revert --no-commit 15c05cb7 7af4f4f2
git commit -m "hermes: restore prior suite grouping policy"
git push origin HEAD:main
flux reconcile kustomization hermes --namespace flux-system --with-source
```
The first live 363-case multi-pass run at the previous 900-second / USD 5
guards stopped after 747.513 seconds during large-family review. Its actual CLI
final subtype was `error_max_budget_usd`; three full passes had completed. The CLI
reported an aggregate estimate of 5.62765, showing the documented within-generation
overshoot. It was not a timeout, turn-limit failure, or accepted partial answer.
The adapter now reports this condition as `job_cost_budget_exhausted`.
## Synthetic acceptance measurements
The suite inputs are synthetic fixtures, never roster-derived material. Tests ran
through authenticated TLS from `titan-jh` (`192.168.22.8`) to the LAN ingress
`192.168.22.50:443`, with hostname verification, proxy bypass, and no redirects.
This verifies that LAN test path; it does not replace a WSL connectivity test.
| Cases | Request bytes | Wall seconds | Model passes / CLI turns | Natural families | Final tasks | Natural / final singletons | Input / output tokens |
| --- | ---: | ---: | --- | ---: | ---: | --- | --- |
| 14 | 9,847 | 61.805 | 3 / 6 | 9 | 9 | 4 / 4 | 18,216 / 5,794 |
| 75 | 55,984 | 274.620 | 5 / 12 | 9 | 21 | 3 / 3 | 174,942 / 28,950 |
| 38 balanced fixture | 23,871 | 231.337 | 5 / 15 | 4 | 10 | 0 / 0 | 139,789 / 24,681 |
| 363, old limits | 273,761 | 747.513 | 3 completed; pass 4 failed | Not finalized | No result | Not finalized | 629,141 / 64,424 |
Token totals include all completed CLI turns across the job, not unique source
size. The old-limit 363 result includes reported failed-invocation usage and was
never returned as completed. CLI estimated costs were 0.23593, 1.59846, 1.31597,
and 5.62765 respectively; these are not subscription bills.
For every completed run above, natural-family pair precision and recall were
1.0 against the independently defined implementation patterns, exact alias
coverage passed, and final groups stayed pure to those patterns. No incorrect
merges, unnecessary semantic splits, or cross-family description attribution
were observed during review. Equivalent text under separate aliases remained
separate objectives. Natural-family quality was scored before capacity division;
a cap-compliant result alone does not prove semantic correctness.
The 38-case fixture interleaves four reset-related mechanisms with shared generic
subject descriptions and distinct success criteria: RPC/stub observations, pulse
and oscilloscope timing, offline report parsing, and concurrent queue instrumentation.
It yielded natural sizes 6/7/11/14 and final sizes 3+3 / 4+3 / 4+4+3 / 5+5+4.
Thirty-eight authorized status samples showed five stages, changing heartbeats,
and nineteen CLI activity changes without exposing event bodies.
The actual 14/75/38 independent proposals agreed on membership. Disagreement
reconciliation, genuine semantic subdivision, reversal of an unjustified split,
and naming-collision rejection are separately tested with controlled mocked
responses; do not describe those as observed live-provider disagreements.
The public 363 fixture has description lengths 48-245 characters (mean 234.5),
preconditions 67-143 (mean 137.2), success criteria 73-135 (mean 131.2), and case
labels 7-15. It interleaves six mechanisms, includes identical distant case text,
and has distinctive beginning/middle/end objectives. These measurements describe
the tested material; longer or more ambiguous real suites may take more work.
An installed native CLI loopback probe captured all complete case fields and full
system instructions in every emitted provider request for 14/75/363 cases, with
64,000 output limits. It performed 3/5/5 passes and deterministic final sizing to
9/21/75 tasks using mock responses. This verifies transport and orchestration,
not live semantic quality. Live successful results reported no compaction or
truncation and passed exact coverage and source-quote support checks. These checks
cannot prove the provider internally attended to every field; exact tokenization
and sufficient output reservation for arbitrary suites remain unverified.
The updated code passes 109 synthetic/mocked regression tests. Kubernetes render,
client dry-run, and Flux diff checks passed before deployment. Policy controls,
case coverage validation, scoped authentication, and provider pinning remain in
place. The invalid-token check returned 401 and same-key submission replay reused
the same job. No automatic real-data rerun was performed.
The second 363-case attempt, under the 1200-second / USD 10 guards, failed after
228.260 seconds in proposal B with `invalid_case_assignments`. Both invocations
returned successful CLI results, but the second partition failed server membership
validation. Its reported total was 281,227 input / 23,796 output tokens and 2.001035
in CLI estimated-cost accounting. The previous validator did not retain counts of
missing, repeated, or unknown assignments, and transient output was cleaned; the
specific mismatch cannot be recovered. No partial partition was accepted. The new
required-alias schema and count-only failure diagnostics address this failure mode.

View File

@ -13,11 +13,12 @@ The endpoint and request fields are unchanged. Optional top-level `review_summar
is returned only with the authorized result. Configuration remains `suite-v6-20260929`;
policy, prompt, and execution revisions identify the changed grouping behavior.
The prior single-pass acceptance results below are historical and do not measure
the new workflow. The new process shares the existing 900-second / USD 5 estimated
job limits across all invocations. The separate local-model endpoint is unchanged;
the new workflow. The new process shares the user-approved 1200-second deadline
and USD 10 CLI estimate guard across all invocations. This is not subscription
billing. Progress updates every five seconds during a CLI pass. The separate local-model endpoint is unchanged;
the 8K local route cannot admit the new multi-pass reconciliation schema.
## CLI failure diagnostics update
## Earlier CLI failure diagnostics update (historical)
Execution revision `claude-diagnostics-turns-v1-20260929` adds content-free
`cli_diagnostics` and failure-stage details, including exit/signal, final subtype,
@ -199,7 +200,7 @@ LOCAL_ONLY_FIELD: token
SYNTHETIC_EXTERNAL_FIELD: synthetic_token
APPROVED_OPERATIONAL_FIELD: operational_token
INITIAL_CONCURRENCY: 1; busy submissions return 429; no waiting queue
JOB_TIMEOUT: up to 900 seconds
JOB_TIMEOUT: up to 1200 seconds
HTTP_TIMEOUT: client 45 seconds; submit/status do not wait for inference
TLS: existing worker.bstein.dev certificate; normal trusted CA verification
```
@ -252,8 +253,8 @@ UTF-8 byte limits, unique aliases, ownership, permissions, and exact result cove
},
"execution": {
"strategy": "whole_suite",
"max_seconds": 900,
"max_cost_usd": 5
"max_seconds": 1200,
"max_cost_usd": 10
}
}
```
@ -272,8 +273,8 @@ Missing routing policy means local-only. External provider names are `claude` an
unknown providers, malformed booleans, duplicate names, and permission escalation
are rejected before routing. `allow_external=false` cannot include providers.
`whole_suite` is the only strategy. No batching, summarization, or reconciliation
is silently substituted. `max_seconds` is 10-900. The Claude CLI budget is at most
USD 5 in its estimated usage accounting; this is not a verified subscription
is silently substituted. `max_seconds` is 10-1200. The Claude CLI guard is at most
USD 10 in its estimated usage accounting; this is not a verified subscription
billing ceiling and can overshoot within a single provider call.
The gateway deterministically selects a fitting permitted backend, preferring
@ -448,7 +449,7 @@ settings remain unverified. See [Claude data usage](https://code.claude.com/docs
The new API accepts at most 1 MiB and 400 cases per request, with at most 32 KiB
per field. The existing local endpoint still accepts only 128 KiB and its original
context/output limits. Claude subprocess stdout is bounded to 4 MiB; normalized
results to 1 MiB. Names are at most 80 characters, descriptions at most 240.
results to 1 MiB. Final names are at most 64 characters, descriptions at most 240, and groups at most five cases.
Preflight includes the entire serialized suite, fixed instructions, schema, reserved
harness overhead, and output space. There is no accurate account-specific tokenizer.
@ -463,7 +464,7 @@ Small local requests reserve 1,024 overhead tokens and 2,048 output tokens withi
8,192 context. The realistic 14/75/363 fixtures exceed this conservative local
whole-suite budget and must use the approved hosted path or fail preflight.
CLI execution lasts at most 900 seconds; cancellation and timeouts terminate its
All CLI invocations share at most 1200 seconds per suite job; cancellation and timeouts terminate its
process group. Async HTTP calls finish promptly, with a 30-second body-read timeout,
40-second ingress response-header timeout, and recommended 45-second client timeout.
Switchyard's ten-minute internal request limit carries only a tiny immediate routing
@ -514,7 +515,7 @@ import json
with open('synthetic-suite.json') as stream:
request = json.load(stream)
request['routing'] = {'allow_external': True, 'allowed_external_providers': ['claude']}
request['execution'] = {'strategy': 'whole_suite', 'max_seconds': 900, 'max_cost_usd': 5}
request['execution'] = {'strategy': 'whole_suite', 'max_seconds': 1200, 'max_cost_usd': 10}
with open('synthetic-request.json', 'w') as stream:
json.dump(request, stream)
PY

View File

@ -5,11 +5,12 @@ Execute inside the planner using Python stdin. No hosted inference is performed.
This checks full-field transmission and fresh CLI sessions, not grouping quality.
"""
import json
import os
from http.server import BaseHTTPRequestHandler, HTTPServer
import sys
import threading
sys.path.insert(0, '/opt/planner')
sys.path.insert(0, os.environ.get('SUITE_PROBE_MODULE_DIR', '/opt/planner'))
import suite_backends
from suite_contract import validate_request, validate_result
from suite_multipass import generate, preflight_workflow
@ -53,7 +54,17 @@ def response_value():
value['decisions'] = [{'source_members': g['members'], 'decision': 'keep',
'rationale': 'Mock transport check of a coherent implementation.',
'evidence': g['evidence']} for g in groups if len(g['members']) > 5]
return value
assignments, wire_groups = {}, []
for index, group in enumerate(groups):
wire_groups.append({k:v for k,v in group.items() if k not in {'members','variation_sets'}})
for alias in group['members']:
assignments[alias] = {'family':index,'variation':0} if 'variation_sets' in group else index
wire = {'groups':wire_groups,'assignments':assignments}
if 'decisions' in value:
keys = {frozenset(v):k for k,v in active['review_keys'].items()}
wire['decisions'] = {keys[frozenset(d['source_members'])]:
{k:v for k,v in d.items() if k!='source_members'} for d in value['decisions']}
return wire
class Provider(BaseHTTPRequestHandler):
@ -72,6 +83,11 @@ class Provider(BaseHTTPRequestHandler):
'complete_system': system, 'max_tokens': body.get('max_tokens')})
assert complete and system
value = response_value()
# Force one schema rejection; the actual CLI must repair it within turns.
omitted = len(json.loads(active['input'])['suite']['cases']) == 14 and len(seen) == 1
if omitted:
del value['assignments'][next(iter(value['assignments']))]
seen[-1]['missing_assignment_injected'] = omitted
block = {'type': 'tool_use', 'id': 'mock-output', 'name': 'StructuredOutput', 'input': {}}
message = {'id': 'mock', 'type': 'message', 'role': 'assistant', 'model': 'claude-opus-4-8',
'content': [], 'stop_reason': None, 'stop_sequence': None,

View File

@ -61,6 +61,7 @@ configMapGenerator:
- suite_jobs.py=scripts/suite_jobs.py
- suite_backends.py=scripts/suite_backends.py
- suite_cli_diagnostics.py=scripts/suite_cli_diagnostics.py
- suite_assignments.py=scripts/suite_assignments.py
- suite_policy.py=scripts/suite_policy.py
- suite_sizing.py=scripts/suite_sizing.py
- suite_multipass.py=scripts/suite_multipass.py

View File

@ -0,0 +1,108 @@
"""Require one model assignment per source alias, then rebuild public partitions."""
from __future__ import annotations
import copy
from suite_contract import Problem, digest
WIRE_INSTRUCTIONS = (
"OUTPUT REPRESENTATION: Define your natural families in the groups array. "
"Do not list members there. The assignments object has one REQUIRED key for "
"EVERY exact input alias. Its value identifies the ZERO-BASED groups array index. "
"In review passes each assignment instead has family (that index) and variation "
"(a nonnegative integer naming a related variation set within that family). "
"Identical content should use the same family and variation. Every group must "
"be used. Do not omit or rename aliases. This changes only serialization, not "
"the natural-family objective or the independent assessment. For large-family "
"decisions use the required review_keys from review_material: those keys identify "
"exact original memberships, not family names. Never invent decision keys."
)
def decision_keys(context):
"""Index original oversized memberships independently of names or proposal order."""
original = context.get("original_partition", context.get("natural_partition", {}))
return {"F-" + digest(sorted(g["members"]))[:20]: sorted(g["members"])
for g in original.get("groups", []) if len(g["members"]) > 5}
def wire_schema(schema, request, context):
"""Replace freely enumerated memberships with required, exactly keyed assignments."""
wire = copy.deepcopy(schema)
group = wire["properties"]["groups"]["items"]
rich = "variation_sets" in group["properties"]
for field in ("members", "variation_sets"):
group["properties"].pop(field, None)
if field in group["required"]:
group["required"].remove(field)
aliases = sorted(c["alias"] for c in request["cases"])
index = {"type": "integer", "minimum": 0, "maximum": len(aliases) - 1}
assignment = {"type": "object", "additionalProperties": False,
"required": ["family", "variation"],
"properties": {"family": index, "variation": index}} if rich else index
wire["required"].append("assignments")
wire["properties"]["assignments"] = {
"type": "object", "additionalProperties": False, "required": aliases,
"properties": {alias: assignment for alias in aliases}}
if "decisions" in wire["properties"]:
item = wire["properties"]["decisions"]["items"]
del item["properties"]["source_members"]
item["required"].remove("source_members")
keys = decision_keys(context)
wire["properties"]["decisions"] = {
"type": "object", "additionalProperties": False, "required": sorted(keys),
"properties": {key: item for key in sorted(keys)}}
return wire
def decode(value, call, request):
"""Reconstruct groups only after strict alias keys and valid indexes are checked."""
schema = call["schema"]
if not isinstance(value, dict) or set(value) != set(schema["required"]):
raise Problem("invalid_json_result", 502, failure_stage="assignment_envelope")
aliases = {c["alias"] for c in request["cases"]}
assignments, groups = value["assignments"], copy.deepcopy(value["groups"])
if not isinstance(assignments, dict):
raise Problem("invalid_case_assignments", 502, failure_stage="assignment_keys")
if set(assignments) != aliases:
raise Problem("invalid_case_assignments", 502, failure_stage="assignment_keys",
missing_case_count=len(aliases - set(assignments)),
unknown_case_count=len(set(assignments) - aliases))
if not isinstance(groups, list) or not 1 <= len(groups) <= len(aliases):
raise Problem("invalid_case_assignments", 502, failure_stage="assignment_groups")
required = set(schema["properties"]["groups"]["items"]["required"])
rich = "common_work" in required
for group in groups:
if not isinstance(group, dict) or set(group) != required:
raise Problem("invalid_json_result", 502, failure_stage="assignment_group_fields")
group["members"] = []
variations = [{} for _ in groups]
for alias in sorted(aliases):
assignment = assignments[alias]
if rich:
if not isinstance(assignment, dict) or set(assignment) != {"family", "variation"}:
raise Problem("invalid_case_assignments", 502, failure_stage="assignment_value")
family, variation = assignment["family"], assignment["variation"]
if type(variation) is not int or not 0 <= variation < len(aliases):
raise Problem("invalid_variation_assignments", 502)
else:
family, variation = assignment, 0
if type(family) is not int or not 0 <= family < len(groups):
raise Problem("invalid_case_assignments", 502, failure_stage="assignment_index")
groups[family]["members"].append(alias)
variations[family].setdefault(variation, []).append(alias)
if any(not g["members"] for g in groups):
raise Problem("invalid_case_assignments", 502, failure_stage="empty_assigned_group")
if rich:
for group, sets in zip(groups, variations):
group["variation_sets"] = [sets[index] for index in sorted(sets)]
result = {"groups": groups}
if "decisions" in value:
decisions, keys = value["decisions"], call["review_keys"]
if not isinstance(decisions, dict) or set(decisions) != set(keys):
raise Problem("incomplete_large_family_review", 502, failure_stage="decision_keys")
if any(not isinstance(item, dict) or set(item) != {"decision", "rationale", "evidence"}
for item in decisions.values()):
raise Problem("invalid_review_output", 502)
result["decisions"] = [{**decisions[key], "source_members": keys[key]} for key in sorted(keys)]
return result

View File

@ -7,8 +7,8 @@ import re
from collections import Counter
REVISION = "suite-v6-20260929"
PROMPT_REVISION = "implementation-proximity-multipass-v3-20260929"
EXECUTION_REVISION = "suite-multipass-v2-20260929"
PROMPT_REVISION = "implementation-proximity-multipass-v4-20260929"
EXECUTION_REVISION = "suite-multipass-v3-20260929"
CLAUDE_MAX_TURNS = 6
MAX_BODY = 1 << 20
MAX_RESULT = 1 << 20
@ -209,8 +209,12 @@ def validate_partition(result, request, *, name_limit=64, group_limit=None, uniq
if group_limit is not None and len(members) > group_limit:
raise Problem("group_size_limit", 502)
found.extend(members)
if Counter(found) != Counter(c["alias"] for c in request["cases"]):
raise Problem("invalid_case_assignments", 502)
counts = Counter(found)
expected = {c["alias"] for c in request["cases"]}
if counts != Counter(expected):
raise Problem("invalid_case_assignments", 502, missing_case_count=len(expected - counts.keys()),
unknown_case_count=len(counts.keys() - expected),
duplicate_assignment_count=sum(count - 1 for count in counts.values()))
if len(encoded(result)) > MAX_RESULT:
raise Problem("response_too_large", 502)
return result

View File

@ -6,6 +6,7 @@ import math
import time
import suite_backends
from suite_assignments import decode
from suite_contract import (EXECUTION_REVISION, MAX_BODY, MAX_RESULT, MODELS, PROMPT_REVISION,
PROMPT_SHA256, REVISION, Problem, digest, encoded, validate_partition)
from suite_policy import (BASE_NAME_LIMIT, MAX_GROUP, POLICY_REVISION, invocation,
@ -164,7 +165,7 @@ class Workflow:
self.spent += cost
self.checkpoint()
report({"cli_running": False})
return value
return decode(value, call, source)
def execute(self):
"""Perform independent discovery, reconciliation, and bounded semantic reviews."""

View File

@ -6,6 +6,7 @@ import re
from collections import Counter
from suite_contract import FIELDS, Problem, SCHEMA, SYSTEM, encoded, validate_partition
from suite_assignments import WIRE_INSTRUCTIONS, decision_keys, wire_schema
POLICY_REVISION = "implementation-five-v1-20260929"
MAX_GROUP = 5
@ -109,8 +110,14 @@ def invocation(stage, request, context=None):
if stage == "decision_audit":
instruction += "\n" + REVIEW + "\n" + AUDIT
source = {k: request[k] for k in ("campaign", "suite", "cases")}
return {"stage": stage, "system": SYSTEM + "\n\nPASS INSTRUCTIONS:\n" + instruction,
"schema": schema, "input": encoded({"suite": source, "review_material": context or {}}).decode()}
context = dict(context or {})
keys = decision_keys(context) if stage in STAGES[3:] else {}
if keys:
context["review_keys"] = keys
return {"stage": stage,
"schema": wire_schema(schema, request, context), "review_keys": keys,
"input": encoded({"suite": source, "review_material": context}).decode(),
"system": SYSTEM + "\n\nPASS INSTRUCTIONS:\n" + instruction + "\n\n" + WIRE_INSTRUCTIONS}
def projection(value):

View File

@ -36,7 +36,7 @@ spec:
app: hermes-suite-planner
annotations:
fluentbit.io/exclude: "true"
ai.bstein.dev/config-rev: suite-v6-multipass-cap5-v2-20260929
ai.bstein.dev/config-rev: suite-v6-multipass-cap5-v3-20260929
vault.hashicorp.com/agent-inject: "true"
vault.hashicorp.com/agent-pre-populate-only: "true"
vault.hashicorp.com/agent-init-first: "true"

View File

@ -13,6 +13,7 @@ import suite_multipass as workflow
from suite_contract import MODELS, Problem, encoded, preflight, validate_partition, validate_request, validate_result
from suite_jobs import Jobs
from suite_policy import invocation, validate_natural
from suite_assignments import decode, decision_keys
from suite_sizing import balanced_sizes, cap_families, disagreements
from suite_synthetic import fixture
@ -49,6 +50,22 @@ def public(value):
return {'groups': [{k: g[k] for k in ('name', 'description', 'members')} for g in value['groups']]}
def wire_value(value, call):
"""Encode fixed test decisions using the model's required per-alias representation."""
assignments, groups = {}, []
for index, group in enumerate(value['groups']):
groups.append({k:v for k,v in group.items() if k not in {'members','variation_sets'}})
variation = {a:i for i,part in enumerate(group.get('variation_sets',[])) for a in part}
for alias in group['members']:
assignments[alias] = {'family':index,'variation':variation[alias]} if variation else index
result = {'groups':groups,'assignments':assignments}
if 'decisions' in value:
key_by_members = {frozenset(v):k for k,v in call['review_keys'].items()}
result['decisions'] = {key_by_members[frozenset(d['source_members'])]:
{k:v for k,v in d.items() if k!='source_members'} for d in value['decisions']}
return result
def review(value, original):
"""Attach one explicit keep/split decision for each original oversized family."""
value = copy.deepcopy(value)
@ -80,7 +97,7 @@ def install_backend(monkeypatch, source, reconciled, *, a=None, b=None, reviewed
if progress:
progress({'cli_running': True, 'cli_output_bytes': 120, 'last_cli_activity_seconds_ago': 0})
cost = costs[len(calls)-1] if costs else 0.1
return copy.deepcopy(answers[invocation['stage']]), {
return wire_value(copy.deepcopy(answers[invocation['stage']]), invocation), {
'model': MODELS['claude']['model'], 'usage': {'input_tokens': 100, 'output_tokens': 40},
'cost_usd_estimate': cost, 'turns': 2, 'duration_api_ms': 10,
'compaction': False, 'truncation': False, 'cli_diagnostics': {'exit_code': 0}}
@ -117,7 +134,8 @@ def test_independent_orders_and_content_are_reproducible(monkeypatch):
assert calls[0][1]['system'] == calls[1][1]['system']
assert a['suite']['cases'] != b['suite']['cases']
assert b['suite']['cases'] == workflow.ordered_request(source, True)['cases']
assert all('maxItems' not in c[1]['schema']['properties']['groups']['items']['properties']['members'] for c in calls)
assert all('members' not in c[1]['schema']['properties']['groups']['items']['properties'] for c in calls)
assert all(set(c[1]['schema']['properties']['assignments']['required']) == {x['alias'] for x in source['cases']} for c in calls)
assert 'proposal_a' in json.loads(calls[2][1]['input'])['review_material']
@ -157,6 +175,56 @@ def test_cli_estimate_guard_default_and_client_override():
validate_request(source, ['claude'])
@pytest.mark.parametrize('size', [14,75,363])
def test_every_alias_is_required_in_internal_wire_schema(size):
source = request(size)
aliases = [c['alias'] for c in source['cases']]
value = natural(source, [aliases])
for stage in ('proposal_a','reconciliation','large_family_review','decision_audit'):
context = {'original_partition':value} if stage=='decision_audit' else {'natural_partition':value}
call = invocation(stage, source, None if stage=='proposal_a' else context)
data = public(value) if stage=='proposal_a' else review(value,value) if stage in ('large_family_review','decision_audit') else value
wire = wire_value(data, call)
schema = call['schema']['properties']['assignments']
assert set(schema['required']) == set(schema['properties']) == set(aliases)
assert schema['additionalProperties'] is False
assert decode(wire,call,source) == data
@pytest.mark.parametrize('mode,stage', [
('missing','assignment_keys'),('unknown','assignment_keys'),('wrong_index','assignment_index'),
('bool_index','assignment_index'),('unused_family','empty_assigned_group'),
('array_instead','assignment_keys'),('extra_group_field','assignment_group_fields')])
def test_invalid_wire_assignments_fail_without_partial_output_or_content(mode,stage):
source = request(14)
aliases = [c['alias'] for c in source['cases']]
call = invocation('proposal_a',source)
wire = wire_value(public(natural(source,[aliases])),call)
if mode=='missing': del wire['assignments'][aliases[0]]
elif mode=='unknown': wire['assignments'][CANARY]=0
elif mode=='wrong_index': wire['assignments'][aliases[0]]=1
elif mode=='bool_index': wire['assignments'][aliases[0]]=True
elif mode=='unused_family': wire['groups'].append(dict(wire['groups'][0]))
elif mode=='array_instead': wire['assignments']=[]
else: wire['groups'][0]['members']=aliases
with pytest.raises(Problem) as raised: decode(wire,call,source)
assert raised.value.details['failure_stage']==stage
assert CANARY not in json.dumps(raised.value.document())
def test_review_wire_keys_reference_exact_memberships_and_cannot_omit_a_decision():
source=request(14)
aliases=[c['alias'] for c in source['cases']]
natural_value=natural(source,[aliases[:7],aliases[7:]])
context={'original_partition':natural_value}
call=invocation('decision_audit',source,context)
keys=decision_keys(context)
assert len(keys)==2 and set(call['schema']['properties']['decisions']['required'])==set(keys)
wire=wire_value(review(natural_value,natural_value),call)
del wire['decisions'][next(iter(keys))]
with pytest.raises(Problem,match='incomplete_large_family_review'): decode(wire,call,source)
def test_real_semantic_subdivisions_precede_work_sizing(monkeypatch):
source = request(11)
aliases = [c['alias'] for c in source['cases']]