hermes: require one structured assignment per suite case
This commit is contained in:
parent
15c05cb7fb
commit
9c327c6fd0
@ -10,8 +10,8 @@ most **64 characters**, unique after whitespace and case normalization.
|
||||
|
||||
- Configuration: `suite-v6-20260929` (HTTP compatibility identifier).
|
||||
- Policy: `implementation-five-v1-20260929`.
|
||||
- Prompt: `implementation-proximity-multipass-v3-20260929`.
|
||||
- Execution: `suite-multipass-v2-20260929`.
|
||||
- Prompt: `implementation-proximity-multipass-v4-20260929`.
|
||||
- Execution: `suite-multipass-v3-20260929`.
|
||||
|
||||
A single server-side job performs:
|
||||
|
||||
@ -35,11 +35,22 @@ alias and is present verbatim in that source field. This checks support provenan
|
||||
not the correctness of a model's engineering interpretation. Unknown implementation
|
||||
details remain concise uncertainty, not invented equipment or procedures.
|
||||
|
||||
The internal structured-output schema requires one property for each exact source
|
||||
alias. Its value selects a natural-family index and, in review passes, a variation
|
||||
index. The server reconstructs memberships and runs the normal independent checks.
|
||||
Review decisions also use required keys derived from exact original memberships.
|
||||
This avoids free-form membership lists silently omitting or repeating aliases.
|
||||
The public groups schema is unchanged. Invalid keys, indexes, empty groups, and
|
||||
missing review decisions still fail closed; no missing assignment is fabricated.
|
||||
|
||||
Each invocation retains six CLI turns for structured output. Model review passes
|
||||
and CLI turns are separate counters. All calls use the originally selected provider
|
||||
and pinned model. The model remains `claude-opus-4-8[1m]`, with canonical runtime
|
||||
identity checked as `claude-opus-4-8`, firstParty, medium effort, reported 1M context
|
||||
and 64K output ceiling. No new provider fallback or tools are enabled.
|
||||
and 64K output ceiling. The server invokes native Claude Code CLI 2.1.226 using
|
||||
the existing first-party OAuth account, not a separately configured API-key
|
||||
account. The CLI itself communicates with Anthropic over HTTPS. No new provider
|
||||
fallback or tools are enabled.
|
||||
|
||||
## Deterministic task sizing and names
|
||||
|
||||
@ -104,7 +115,24 @@ While a CLI invocation runs, progress updates every five seconds with `heartbeat
|
||||
and `last_cli_activity_seconds_ago`. The heartbeat means the worker is alive; a
|
||||
change in output bytes means CLI activity. Neither is a percentage complete or
|
||||
proof that a particular case has been reasoned about. No event bodies or reasoning
|
||||
text are exposed. Poll the existing status URL every five seconds to display it.
|
||||
text are exposed. `cost_used_usd_estimate` counts completed invocations; the
|
||||
in-flight invocation is not included until its terminal usage is available.
|
||||
Poll the existing status URL every five seconds to display progress.
|
||||
|
||||
To inspect progress from WSL, keep the existing bearer token private and use the
|
||||
job ID returned by submission. This command only reads status:
|
||||
|
||||
```bash
|
||||
curl -q --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
|
||||
--connect-timeout 10 --max-time 45 --fail-with-body --silent --show-error \
|
||||
--config <(printf 'header = "Authorization: Bearer %s"\n' "$SUITE_PLANNING_TOKEN") \
|
||||
"https://worker.bstein.dev/suite-planning/v1/jobs/$JOB_ID"
|
||||
```
|
||||
|
||||
An existing Python poller can display `envelope["execution_progress"]` after each
|
||||
status read. No separate pass submissions, new request fields, or client-side
|
||||
reconciliation are required. Prefer the actual stage and elapsed time over an
|
||||
invented percentage, because different passes have very different runtimes.
|
||||
|
||||
## Shared bounds and failure behavior
|
||||
|
||||
@ -159,9 +187,90 @@ restore the preceding single-pass policy while preserving the earlier CLI diagno
|
||||
fix. Wait for active jobs before a worker restart because completed results are
|
||||
held in memory. No Vault credential or ingress/routing changes are required.
|
||||
|
||||
The policy implementation is commit `7af4f4f2`; the approved 20-minute / USD 10
|
||||
estimate guard and progress reporting are commit `15c05cb7`. To restore the prior
|
||||
single-pass implementation after active jobs finish, revert both through the normal
|
||||
Git deployment branch, preserving the earlier six-turn diagnostics fix:
|
||||
|
||||
```bash
|
||||
git revert --no-commit 15c05cb7 7af4f4f2
|
||||
git commit -m "hermes: restore prior suite grouping policy"
|
||||
git push origin HEAD:main
|
||||
flux reconcile kustomization hermes --namespace flux-system --with-source
|
||||
```
|
||||
|
||||
The first live 363-case multi-pass run at the previous 900-second / USD 5
|
||||
guards stopped after 747.513 seconds during large-family review. Its actual CLI
|
||||
final subtype was `error_max_budget_usd`; three full passes had completed. The CLI
|
||||
reported an aggregate estimate of 5.62765, showing the documented within-generation
|
||||
overshoot. It was not a timeout, turn-limit failure, or accepted partial answer.
|
||||
The adapter now reports this condition as `job_cost_budget_exhausted`.
|
||||
|
||||
## Synthetic acceptance measurements
|
||||
|
||||
The suite inputs are synthetic fixtures, never roster-derived material. Tests ran
|
||||
through authenticated TLS from `titan-jh` (`192.168.22.8`) to the LAN ingress
|
||||
`192.168.22.50:443`, with hostname verification, proxy bypass, and no redirects.
|
||||
This verifies that LAN test path; it does not replace a WSL connectivity test.
|
||||
|
||||
| Cases | Request bytes | Wall seconds | Model passes / CLI turns | Natural families | Final tasks | Natural / final singletons | Input / output tokens |
|
||||
| --- | ---: | ---: | --- | ---: | ---: | --- | --- |
|
||||
| 14 | 9,847 | 61.805 | 3 / 6 | 9 | 9 | 4 / 4 | 18,216 / 5,794 |
|
||||
| 75 | 55,984 | 274.620 | 5 / 12 | 9 | 21 | 3 / 3 | 174,942 / 28,950 |
|
||||
| 38 balanced fixture | 23,871 | 231.337 | 5 / 15 | 4 | 10 | 0 / 0 | 139,789 / 24,681 |
|
||||
| 363, old limits | 273,761 | 747.513 | 3 completed; pass 4 failed | Not finalized | No result | Not finalized | 629,141 / 64,424 |
|
||||
|
||||
Token totals include all completed CLI turns across the job, not unique source
|
||||
size. The old-limit 363 result includes reported failed-invocation usage and was
|
||||
never returned as completed. CLI estimated costs were 0.23593, 1.59846, 1.31597,
|
||||
and 5.62765 respectively; these are not subscription bills.
|
||||
|
||||
For every completed run above, natural-family pair precision and recall were
|
||||
1.0 against the independently defined implementation patterns, exact alias
|
||||
coverage passed, and final groups stayed pure to those patterns. No incorrect
|
||||
merges, unnecessary semantic splits, or cross-family description attribution
|
||||
were observed during review. Equivalent text under separate aliases remained
|
||||
separate objectives. Natural-family quality was scored before capacity division;
|
||||
a cap-compliant result alone does not prove semantic correctness.
|
||||
|
||||
The 38-case fixture interleaves four reset-related mechanisms with shared generic
|
||||
subject descriptions and distinct success criteria: RPC/stub observations, pulse
|
||||
and oscilloscope timing, offline report parsing, and concurrent queue instrumentation.
|
||||
It yielded natural sizes 6/7/11/14 and final sizes 3+3 / 4+3 / 4+4+3 / 5+5+4.
|
||||
Thirty-eight authorized status samples showed five stages, changing heartbeats,
|
||||
and nineteen CLI activity changes without exposing event bodies.
|
||||
|
||||
The actual 14/75/38 independent proposals agreed on membership. Disagreement
|
||||
reconciliation, genuine semantic subdivision, reversal of an unjustified split,
|
||||
and naming-collision rejection are separately tested with controlled mocked
|
||||
responses; do not describe those as observed live-provider disagreements.
|
||||
|
||||
The public 363 fixture has description lengths 48-245 characters (mean 234.5),
|
||||
preconditions 67-143 (mean 137.2), success criteria 73-135 (mean 131.2), and case
|
||||
labels 7-15. It interleaves six mechanisms, includes identical distant case text,
|
||||
and has distinctive beginning/middle/end objectives. These measurements describe
|
||||
the tested material; longer or more ambiguous real suites may take more work.
|
||||
|
||||
An installed native CLI loopback probe captured all complete case fields and full
|
||||
system instructions in every emitted provider request for 14/75/363 cases, with
|
||||
64,000 output limits. It performed 3/5/5 passes and deterministic final sizing to
|
||||
9/21/75 tasks using mock responses. This verifies transport and orchestration,
|
||||
not live semantic quality. Live successful results reported no compaction or
|
||||
truncation and passed exact coverage and source-quote support checks. These checks
|
||||
cannot prove the provider internally attended to every field; exact tokenization
|
||||
and sufficient output reservation for arbitrary suites remain unverified.
|
||||
|
||||
The updated code passes 109 synthetic/mocked regression tests. Kubernetes render,
|
||||
client dry-run, and Flux diff checks passed before deployment. Policy controls,
|
||||
case coverage validation, scoped authentication, and provider pinning remain in
|
||||
place. The invalid-token check returned 401 and same-key submission replay reused
|
||||
the same job. No automatic real-data rerun was performed.
|
||||
|
||||
The second 363-case attempt, under the 1200-second / USD 10 guards, failed after
|
||||
228.260 seconds in proposal B with `invalid_case_assignments`. Both invocations
|
||||
returned successful CLI results, but the second partition failed server membership
|
||||
validation. Its reported total was 281,227 input / 23,796 output tokens and 2.001035
|
||||
in CLI estimated-cost accounting. The previous validator did not retain counts of
|
||||
missing, repeated, or unknown assignments, and transient output was cleaned; the
|
||||
specific mismatch cannot be recovered. No partial partition was accepted. The new
|
||||
required-alias schema and count-only failure diagnostics address this failure mode.
|
||||
|
||||
@ -13,11 +13,12 @@ The endpoint and request fields are unchanged. Optional top-level `review_summar
|
||||
is returned only with the authorized result. Configuration remains `suite-v6-20260929`;
|
||||
policy, prompt, and execution revisions identify the changed grouping behavior.
|
||||
The prior single-pass acceptance results below are historical and do not measure
|
||||
the new workflow. The new process shares the existing 900-second / USD 5 estimated
|
||||
job limits across all invocations. The separate local-model endpoint is unchanged;
|
||||
the new workflow. The new process shares the user-approved 1200-second deadline
|
||||
and USD 10 CLI estimate guard across all invocations. This is not subscription
|
||||
billing. Progress updates every five seconds during a CLI pass. The separate local-model endpoint is unchanged;
|
||||
the 8K local route cannot admit the new multi-pass reconciliation schema.
|
||||
|
||||
## CLI failure diagnostics update
|
||||
## Earlier CLI failure diagnostics update (historical)
|
||||
|
||||
Execution revision `claude-diagnostics-turns-v1-20260929` adds content-free
|
||||
`cli_diagnostics` and failure-stage details, including exit/signal, final subtype,
|
||||
@ -199,7 +200,7 @@ LOCAL_ONLY_FIELD: token
|
||||
SYNTHETIC_EXTERNAL_FIELD: synthetic_token
|
||||
APPROVED_OPERATIONAL_FIELD: operational_token
|
||||
INITIAL_CONCURRENCY: 1; busy submissions return 429; no waiting queue
|
||||
JOB_TIMEOUT: up to 900 seconds
|
||||
JOB_TIMEOUT: up to 1200 seconds
|
||||
HTTP_TIMEOUT: client 45 seconds; submit/status do not wait for inference
|
||||
TLS: existing worker.bstein.dev certificate; normal trusted CA verification
|
||||
```
|
||||
@ -252,8 +253,8 @@ UTF-8 byte limits, unique aliases, ownership, permissions, and exact result cove
|
||||
},
|
||||
"execution": {
|
||||
"strategy": "whole_suite",
|
||||
"max_seconds": 900,
|
||||
"max_cost_usd": 5
|
||||
"max_seconds": 1200,
|
||||
"max_cost_usd": 10
|
||||
}
|
||||
}
|
||||
```
|
||||
@ -272,8 +273,8 @@ Missing routing policy means local-only. External provider names are `claude` an
|
||||
unknown providers, malformed booleans, duplicate names, and permission escalation
|
||||
are rejected before routing. `allow_external=false` cannot include providers.
|
||||
`whole_suite` is the only strategy. No batching, summarization, or reconciliation
|
||||
is silently substituted. `max_seconds` is 10-900. The Claude CLI budget is at most
|
||||
USD 5 in its estimated usage accounting; this is not a verified subscription
|
||||
is silently substituted. `max_seconds` is 10-1200. The Claude CLI guard is at most
|
||||
USD 10 in its estimated usage accounting; this is not a verified subscription
|
||||
billing ceiling and can overshoot within a single provider call.
|
||||
|
||||
The gateway deterministically selects a fitting permitted backend, preferring
|
||||
@ -448,7 +449,7 @@ settings remain unverified. See [Claude data usage](https://code.claude.com/docs
|
||||
The new API accepts at most 1 MiB and 400 cases per request, with at most 32 KiB
|
||||
per field. The existing local endpoint still accepts only 128 KiB and its original
|
||||
context/output limits. Claude subprocess stdout is bounded to 4 MiB; normalized
|
||||
results to 1 MiB. Names are at most 80 characters, descriptions at most 240.
|
||||
results to 1 MiB. Final names are at most 64 characters, descriptions at most 240, and groups at most five cases.
|
||||
|
||||
Preflight includes the entire serialized suite, fixed instructions, schema, reserved
|
||||
harness overhead, and output space. There is no accurate account-specific tokenizer.
|
||||
@ -463,7 +464,7 @@ Small local requests reserve 1,024 overhead tokens and 2,048 output tokens withi
|
||||
8,192 context. The realistic 14/75/363 fixtures exceed this conservative local
|
||||
whole-suite budget and must use the approved hosted path or fail preflight.
|
||||
|
||||
CLI execution lasts at most 900 seconds; cancellation and timeouts terminate its
|
||||
All CLI invocations share at most 1200 seconds per suite job; cancellation and timeouts terminate its
|
||||
process group. Async HTTP calls finish promptly, with a 30-second body-read timeout,
|
||||
40-second ingress response-header timeout, and recommended 45-second client timeout.
|
||||
Switchyard's ten-minute internal request limit carries only a tiny immediate routing
|
||||
@ -514,7 +515,7 @@ import json
|
||||
with open('synthetic-suite.json') as stream:
|
||||
request = json.load(stream)
|
||||
request['routing'] = {'allow_external': True, 'allowed_external_providers': ['claude']}
|
||||
request['execution'] = {'strategy': 'whole_suite', 'max_seconds': 900, 'max_cost_usd': 5}
|
||||
request['execution'] = {'strategy': 'whole_suite', 'max_seconds': 1200, 'max_cost_usd': 10}
|
||||
with open('synthetic-request.json', 'w') as stream:
|
||||
json.dump(request, stream)
|
||||
PY
|
||||
|
||||
@ -5,11 +5,12 @@ Execute inside the planner using Python stdin. No hosted inference is performed.
|
||||
This checks full-field transmission and fresh CLI sessions, not grouping quality.
|
||||
"""
|
||||
import json
|
||||
import os
|
||||
from http.server import BaseHTTPRequestHandler, HTTPServer
|
||||
import sys
|
||||
import threading
|
||||
|
||||
sys.path.insert(0, '/opt/planner')
|
||||
sys.path.insert(0, os.environ.get('SUITE_PROBE_MODULE_DIR', '/opt/planner'))
|
||||
import suite_backends
|
||||
from suite_contract import validate_request, validate_result
|
||||
from suite_multipass import generate, preflight_workflow
|
||||
@ -53,7 +54,17 @@ def response_value():
|
||||
value['decisions'] = [{'source_members': g['members'], 'decision': 'keep',
|
||||
'rationale': 'Mock transport check of a coherent implementation.',
|
||||
'evidence': g['evidence']} for g in groups if len(g['members']) > 5]
|
||||
return value
|
||||
assignments, wire_groups = {}, []
|
||||
for index, group in enumerate(groups):
|
||||
wire_groups.append({k:v for k,v in group.items() if k not in {'members','variation_sets'}})
|
||||
for alias in group['members']:
|
||||
assignments[alias] = {'family':index,'variation':0} if 'variation_sets' in group else index
|
||||
wire = {'groups':wire_groups,'assignments':assignments}
|
||||
if 'decisions' in value:
|
||||
keys = {frozenset(v):k for k,v in active['review_keys'].items()}
|
||||
wire['decisions'] = {keys[frozenset(d['source_members'])]:
|
||||
{k:v for k,v in d.items() if k!='source_members'} for d in value['decisions']}
|
||||
return wire
|
||||
|
||||
|
||||
class Provider(BaseHTTPRequestHandler):
|
||||
@ -72,6 +83,11 @@ class Provider(BaseHTTPRequestHandler):
|
||||
'complete_system': system, 'max_tokens': body.get('max_tokens')})
|
||||
assert complete and system
|
||||
value = response_value()
|
||||
# Force one schema rejection; the actual CLI must repair it within turns.
|
||||
omitted = len(json.loads(active['input'])['suite']['cases']) == 14 and len(seen) == 1
|
||||
if omitted:
|
||||
del value['assignments'][next(iter(value['assignments']))]
|
||||
seen[-1]['missing_assignment_injected'] = omitted
|
||||
block = {'type': 'tool_use', 'id': 'mock-output', 'name': 'StructuredOutput', 'input': {}}
|
||||
message = {'id': 'mock', 'type': 'message', 'role': 'assistant', 'model': 'claude-opus-4-8',
|
||||
'content': [], 'stop_reason': None, 'stop_sequence': None,
|
||||
|
||||
@ -61,6 +61,7 @@ configMapGenerator:
|
||||
- suite_jobs.py=scripts/suite_jobs.py
|
||||
- suite_backends.py=scripts/suite_backends.py
|
||||
- suite_cli_diagnostics.py=scripts/suite_cli_diagnostics.py
|
||||
- suite_assignments.py=scripts/suite_assignments.py
|
||||
- suite_policy.py=scripts/suite_policy.py
|
||||
- suite_sizing.py=scripts/suite_sizing.py
|
||||
- suite_multipass.py=scripts/suite_multipass.py
|
||||
|
||||
108
services/hermes/scripts/suite_assignments.py
Normal file
108
services/hermes/scripts/suite_assignments.py
Normal file
@ -0,0 +1,108 @@
|
||||
"""Require one model assignment per source alias, then rebuild public partitions."""
|
||||
from __future__ import annotations
|
||||
|
||||
import copy
|
||||
|
||||
from suite_contract import Problem, digest
|
||||
|
||||
WIRE_INSTRUCTIONS = (
|
||||
"OUTPUT REPRESENTATION: Define your natural families in the groups array. "
|
||||
"Do not list members there. The assignments object has one REQUIRED key for "
|
||||
"EVERY exact input alias. Its value identifies the ZERO-BASED groups array index. "
|
||||
"In review passes each assignment instead has family (that index) and variation "
|
||||
"(a nonnegative integer naming a related variation set within that family). "
|
||||
"Identical content should use the same family and variation. Every group must "
|
||||
"be used. Do not omit or rename aliases. This changes only serialization, not "
|
||||
"the natural-family objective or the independent assessment. For large-family "
|
||||
"decisions use the required review_keys from review_material: those keys identify "
|
||||
"exact original memberships, not family names. Never invent decision keys."
|
||||
)
|
||||
|
||||
|
||||
def decision_keys(context):
|
||||
"""Index original oversized memberships independently of names or proposal order."""
|
||||
original = context.get("original_partition", context.get("natural_partition", {}))
|
||||
return {"F-" + digest(sorted(g["members"]))[:20]: sorted(g["members"])
|
||||
for g in original.get("groups", []) if len(g["members"]) > 5}
|
||||
|
||||
|
||||
def wire_schema(schema, request, context):
|
||||
"""Replace freely enumerated memberships with required, exactly keyed assignments."""
|
||||
wire = copy.deepcopy(schema)
|
||||
group = wire["properties"]["groups"]["items"]
|
||||
rich = "variation_sets" in group["properties"]
|
||||
for field in ("members", "variation_sets"):
|
||||
group["properties"].pop(field, None)
|
||||
if field in group["required"]:
|
||||
group["required"].remove(field)
|
||||
aliases = sorted(c["alias"] for c in request["cases"])
|
||||
index = {"type": "integer", "minimum": 0, "maximum": len(aliases) - 1}
|
||||
assignment = {"type": "object", "additionalProperties": False,
|
||||
"required": ["family", "variation"],
|
||||
"properties": {"family": index, "variation": index}} if rich else index
|
||||
wire["required"].append("assignments")
|
||||
wire["properties"]["assignments"] = {
|
||||
"type": "object", "additionalProperties": False, "required": aliases,
|
||||
"properties": {alias: assignment for alias in aliases}}
|
||||
if "decisions" in wire["properties"]:
|
||||
item = wire["properties"]["decisions"]["items"]
|
||||
del item["properties"]["source_members"]
|
||||
item["required"].remove("source_members")
|
||||
keys = decision_keys(context)
|
||||
wire["properties"]["decisions"] = {
|
||||
"type": "object", "additionalProperties": False, "required": sorted(keys),
|
||||
"properties": {key: item for key in sorted(keys)}}
|
||||
return wire
|
||||
|
||||
|
||||
def decode(value, call, request):
|
||||
"""Reconstruct groups only after strict alias keys and valid indexes are checked."""
|
||||
schema = call["schema"]
|
||||
if not isinstance(value, dict) or set(value) != set(schema["required"]):
|
||||
raise Problem("invalid_json_result", 502, failure_stage="assignment_envelope")
|
||||
aliases = {c["alias"] for c in request["cases"]}
|
||||
assignments, groups = value["assignments"], copy.deepcopy(value["groups"])
|
||||
if not isinstance(assignments, dict):
|
||||
raise Problem("invalid_case_assignments", 502, failure_stage="assignment_keys")
|
||||
if set(assignments) != aliases:
|
||||
raise Problem("invalid_case_assignments", 502, failure_stage="assignment_keys",
|
||||
missing_case_count=len(aliases - set(assignments)),
|
||||
unknown_case_count=len(set(assignments) - aliases))
|
||||
if not isinstance(groups, list) or not 1 <= len(groups) <= len(aliases):
|
||||
raise Problem("invalid_case_assignments", 502, failure_stage="assignment_groups")
|
||||
required = set(schema["properties"]["groups"]["items"]["required"])
|
||||
rich = "common_work" in required
|
||||
for group in groups:
|
||||
if not isinstance(group, dict) or set(group) != required:
|
||||
raise Problem("invalid_json_result", 502, failure_stage="assignment_group_fields")
|
||||
group["members"] = []
|
||||
variations = [{} for _ in groups]
|
||||
for alias in sorted(aliases):
|
||||
assignment = assignments[alias]
|
||||
if rich:
|
||||
if not isinstance(assignment, dict) or set(assignment) != {"family", "variation"}:
|
||||
raise Problem("invalid_case_assignments", 502, failure_stage="assignment_value")
|
||||
family, variation = assignment["family"], assignment["variation"]
|
||||
if type(variation) is not int or not 0 <= variation < len(aliases):
|
||||
raise Problem("invalid_variation_assignments", 502)
|
||||
else:
|
||||
family, variation = assignment, 0
|
||||
if type(family) is not int or not 0 <= family < len(groups):
|
||||
raise Problem("invalid_case_assignments", 502, failure_stage="assignment_index")
|
||||
groups[family]["members"].append(alias)
|
||||
variations[family].setdefault(variation, []).append(alias)
|
||||
if any(not g["members"] for g in groups):
|
||||
raise Problem("invalid_case_assignments", 502, failure_stage="empty_assigned_group")
|
||||
if rich:
|
||||
for group, sets in zip(groups, variations):
|
||||
group["variation_sets"] = [sets[index] for index in sorted(sets)]
|
||||
result = {"groups": groups}
|
||||
if "decisions" in value:
|
||||
decisions, keys = value["decisions"], call["review_keys"]
|
||||
if not isinstance(decisions, dict) or set(decisions) != set(keys):
|
||||
raise Problem("incomplete_large_family_review", 502, failure_stage="decision_keys")
|
||||
if any(not isinstance(item, dict) or set(item) != {"decision", "rationale", "evidence"}
|
||||
for item in decisions.values()):
|
||||
raise Problem("invalid_review_output", 502)
|
||||
result["decisions"] = [{**decisions[key], "source_members": keys[key]} for key in sorted(keys)]
|
||||
return result
|
||||
@ -7,8 +7,8 @@ import re
|
||||
from collections import Counter
|
||||
|
||||
REVISION = "suite-v6-20260929"
|
||||
PROMPT_REVISION = "implementation-proximity-multipass-v3-20260929"
|
||||
EXECUTION_REVISION = "suite-multipass-v2-20260929"
|
||||
PROMPT_REVISION = "implementation-proximity-multipass-v4-20260929"
|
||||
EXECUTION_REVISION = "suite-multipass-v3-20260929"
|
||||
CLAUDE_MAX_TURNS = 6
|
||||
MAX_BODY = 1 << 20
|
||||
MAX_RESULT = 1 << 20
|
||||
@ -209,8 +209,12 @@ def validate_partition(result, request, *, name_limit=64, group_limit=None, uniq
|
||||
if group_limit is not None and len(members) > group_limit:
|
||||
raise Problem("group_size_limit", 502)
|
||||
found.extend(members)
|
||||
if Counter(found) != Counter(c["alias"] for c in request["cases"]):
|
||||
raise Problem("invalid_case_assignments", 502)
|
||||
counts = Counter(found)
|
||||
expected = {c["alias"] for c in request["cases"]}
|
||||
if counts != Counter(expected):
|
||||
raise Problem("invalid_case_assignments", 502, missing_case_count=len(expected - counts.keys()),
|
||||
unknown_case_count=len(counts.keys() - expected),
|
||||
duplicate_assignment_count=sum(count - 1 for count in counts.values()))
|
||||
if len(encoded(result)) > MAX_RESULT:
|
||||
raise Problem("response_too_large", 502)
|
||||
return result
|
||||
|
||||
@ -6,6 +6,7 @@ import math
|
||||
import time
|
||||
|
||||
import suite_backends
|
||||
from suite_assignments import decode
|
||||
from suite_contract import (EXECUTION_REVISION, MAX_BODY, MAX_RESULT, MODELS, PROMPT_REVISION,
|
||||
PROMPT_SHA256, REVISION, Problem, digest, encoded, validate_partition)
|
||||
from suite_policy import (BASE_NAME_LIMIT, MAX_GROUP, POLICY_REVISION, invocation,
|
||||
@ -164,7 +165,7 @@ class Workflow:
|
||||
self.spent += cost
|
||||
self.checkpoint()
|
||||
report({"cli_running": False})
|
||||
return value
|
||||
return decode(value, call, source)
|
||||
|
||||
def execute(self):
|
||||
"""Perform independent discovery, reconciliation, and bounded semantic reviews."""
|
||||
|
||||
@ -6,6 +6,7 @@ import re
|
||||
from collections import Counter
|
||||
|
||||
from suite_contract import FIELDS, Problem, SCHEMA, SYSTEM, encoded, validate_partition
|
||||
from suite_assignments import WIRE_INSTRUCTIONS, decision_keys, wire_schema
|
||||
|
||||
POLICY_REVISION = "implementation-five-v1-20260929"
|
||||
MAX_GROUP = 5
|
||||
@ -109,8 +110,14 @@ def invocation(stage, request, context=None):
|
||||
if stage == "decision_audit":
|
||||
instruction += "\n" + REVIEW + "\n" + AUDIT
|
||||
source = {k: request[k] for k in ("campaign", "suite", "cases")}
|
||||
return {"stage": stage, "system": SYSTEM + "\n\nPASS INSTRUCTIONS:\n" + instruction,
|
||||
"schema": schema, "input": encoded({"suite": source, "review_material": context or {}}).decode()}
|
||||
context = dict(context or {})
|
||||
keys = decision_keys(context) if stage in STAGES[3:] else {}
|
||||
if keys:
|
||||
context["review_keys"] = keys
|
||||
return {"stage": stage,
|
||||
"schema": wire_schema(schema, request, context), "review_keys": keys,
|
||||
"input": encoded({"suite": source, "review_material": context}).decode(),
|
||||
"system": SYSTEM + "\n\nPASS INSTRUCTIONS:\n" + instruction + "\n\n" + WIRE_INSTRUCTIONS}
|
||||
|
||||
|
||||
def projection(value):
|
||||
|
||||
@ -36,7 +36,7 @@ spec:
|
||||
app: hermes-suite-planner
|
||||
annotations:
|
||||
fluentbit.io/exclude: "true"
|
||||
ai.bstein.dev/config-rev: suite-v6-multipass-cap5-v2-20260929
|
||||
ai.bstein.dev/config-rev: suite-v6-multipass-cap5-v3-20260929
|
||||
vault.hashicorp.com/agent-inject: "true"
|
||||
vault.hashicorp.com/agent-pre-populate-only: "true"
|
||||
vault.hashicorp.com/agent-init-first: "true"
|
||||
|
||||
@ -13,6 +13,7 @@ import suite_multipass as workflow
|
||||
from suite_contract import MODELS, Problem, encoded, preflight, validate_partition, validate_request, validate_result
|
||||
from suite_jobs import Jobs
|
||||
from suite_policy import invocation, validate_natural
|
||||
from suite_assignments import decode, decision_keys
|
||||
from suite_sizing import balanced_sizes, cap_families, disagreements
|
||||
from suite_synthetic import fixture
|
||||
|
||||
@ -49,6 +50,22 @@ def public(value):
|
||||
return {'groups': [{k: g[k] for k in ('name', 'description', 'members')} for g in value['groups']]}
|
||||
|
||||
|
||||
def wire_value(value, call):
|
||||
"""Encode fixed test decisions using the model's required per-alias representation."""
|
||||
assignments, groups = {}, []
|
||||
for index, group in enumerate(value['groups']):
|
||||
groups.append({k:v for k,v in group.items() if k not in {'members','variation_sets'}})
|
||||
variation = {a:i for i,part in enumerate(group.get('variation_sets',[])) for a in part}
|
||||
for alias in group['members']:
|
||||
assignments[alias] = {'family':index,'variation':variation[alias]} if variation else index
|
||||
result = {'groups':groups,'assignments':assignments}
|
||||
if 'decisions' in value:
|
||||
key_by_members = {frozenset(v):k for k,v in call['review_keys'].items()}
|
||||
result['decisions'] = {key_by_members[frozenset(d['source_members'])]:
|
||||
{k:v for k,v in d.items() if k!='source_members'} for d in value['decisions']}
|
||||
return result
|
||||
|
||||
|
||||
def review(value, original):
|
||||
"""Attach one explicit keep/split decision for each original oversized family."""
|
||||
value = copy.deepcopy(value)
|
||||
@ -80,7 +97,7 @@ def install_backend(monkeypatch, source, reconciled, *, a=None, b=None, reviewed
|
||||
if progress:
|
||||
progress({'cli_running': True, 'cli_output_bytes': 120, 'last_cli_activity_seconds_ago': 0})
|
||||
cost = costs[len(calls)-1] if costs else 0.1
|
||||
return copy.deepcopy(answers[invocation['stage']]), {
|
||||
return wire_value(copy.deepcopy(answers[invocation['stage']]), invocation), {
|
||||
'model': MODELS['claude']['model'], 'usage': {'input_tokens': 100, 'output_tokens': 40},
|
||||
'cost_usd_estimate': cost, 'turns': 2, 'duration_api_ms': 10,
|
||||
'compaction': False, 'truncation': False, 'cli_diagnostics': {'exit_code': 0}}
|
||||
@ -117,7 +134,8 @@ def test_independent_orders_and_content_are_reproducible(monkeypatch):
|
||||
assert calls[0][1]['system'] == calls[1][1]['system']
|
||||
assert a['suite']['cases'] != b['suite']['cases']
|
||||
assert b['suite']['cases'] == workflow.ordered_request(source, True)['cases']
|
||||
assert all('maxItems' not in c[1]['schema']['properties']['groups']['items']['properties']['members'] for c in calls)
|
||||
assert all('members' not in c[1]['schema']['properties']['groups']['items']['properties'] for c in calls)
|
||||
assert all(set(c[1]['schema']['properties']['assignments']['required']) == {x['alias'] for x in source['cases']} for c in calls)
|
||||
assert 'proposal_a' in json.loads(calls[2][1]['input'])['review_material']
|
||||
|
||||
|
||||
@ -157,6 +175,56 @@ def test_cli_estimate_guard_default_and_client_override():
|
||||
validate_request(source, ['claude'])
|
||||
|
||||
|
||||
@pytest.mark.parametrize('size', [14,75,363])
|
||||
def test_every_alias_is_required_in_internal_wire_schema(size):
|
||||
source = request(size)
|
||||
aliases = [c['alias'] for c in source['cases']]
|
||||
value = natural(source, [aliases])
|
||||
for stage in ('proposal_a','reconciliation','large_family_review','decision_audit'):
|
||||
context = {'original_partition':value} if stage=='decision_audit' else {'natural_partition':value}
|
||||
call = invocation(stage, source, None if stage=='proposal_a' else context)
|
||||
data = public(value) if stage=='proposal_a' else review(value,value) if stage in ('large_family_review','decision_audit') else value
|
||||
wire = wire_value(data, call)
|
||||
schema = call['schema']['properties']['assignments']
|
||||
assert set(schema['required']) == set(schema['properties']) == set(aliases)
|
||||
assert schema['additionalProperties'] is False
|
||||
assert decode(wire,call,source) == data
|
||||
|
||||
|
||||
@pytest.mark.parametrize('mode,stage', [
|
||||
('missing','assignment_keys'),('unknown','assignment_keys'),('wrong_index','assignment_index'),
|
||||
('bool_index','assignment_index'),('unused_family','empty_assigned_group'),
|
||||
('array_instead','assignment_keys'),('extra_group_field','assignment_group_fields')])
|
||||
def test_invalid_wire_assignments_fail_without_partial_output_or_content(mode,stage):
|
||||
source = request(14)
|
||||
aliases = [c['alias'] for c in source['cases']]
|
||||
call = invocation('proposal_a',source)
|
||||
wire = wire_value(public(natural(source,[aliases])),call)
|
||||
if mode=='missing': del wire['assignments'][aliases[0]]
|
||||
elif mode=='unknown': wire['assignments'][CANARY]=0
|
||||
elif mode=='wrong_index': wire['assignments'][aliases[0]]=1
|
||||
elif mode=='bool_index': wire['assignments'][aliases[0]]=True
|
||||
elif mode=='unused_family': wire['groups'].append(dict(wire['groups'][0]))
|
||||
elif mode=='array_instead': wire['assignments']=[]
|
||||
else: wire['groups'][0]['members']=aliases
|
||||
with pytest.raises(Problem) as raised: decode(wire,call,source)
|
||||
assert raised.value.details['failure_stage']==stage
|
||||
assert CANARY not in json.dumps(raised.value.document())
|
||||
|
||||
|
||||
def test_review_wire_keys_reference_exact_memberships_and_cannot_omit_a_decision():
|
||||
source=request(14)
|
||||
aliases=[c['alias'] for c in source['cases']]
|
||||
natural_value=natural(source,[aliases[:7],aliases[7:]])
|
||||
context={'original_partition':natural_value}
|
||||
call=invocation('decision_audit',source,context)
|
||||
keys=decision_keys(context)
|
||||
assert len(keys)==2 and set(call['schema']['properties']['decisions']['required'])==set(keys)
|
||||
wire=wire_value(review(natural_value,natural_value),call)
|
||||
del wire['decisions'][next(iter(keys))]
|
||||
with pytest.raises(Problem,match='incomplete_large_family_review'): decode(wire,call,source)
|
||||
|
||||
|
||||
def test_real_semantic_subdivisions_precede_work_sizing(monkeypatch):
|
||||
source = request(11)
|
||||
aliases = [c['alias'] for c in source['cases']]
|
||||
|
||||
Loading…
x
Reference in New Issue
Block a user