Fail-closed Hermes full-handoff acceptance harness #19

Manually merged
atlas-reconciler merged 0 commits from feature/hermes-full-handoff-acceptance into main 2026-08-18 21:28:00 +00:00

Repairs the independent review blockers on the same branch and PR. Previous reviewed head: c808baff40a12a90f43c2e5e39b94d139f88ea29. Repaired head: 02e559d7ac80cde6b2f2b87a4cba32bcb79d8971.

Mandatory P1: the zero-evidence fail-open is closed

evaluate_names_absent returned PASS when its step exited 0 with no output. Five mandatory checks — identity.no-provider-api-key-in-pod, identity.no-provider-api-key-in-manifest, forge.worker-holds-no-git-credential-variable, pool.coordinator-retains-sole-state-ownership, access.no-cluster-admin-binding-for-agent — could therefore report a pass on no evidence at all, converting a NO_GO into a GO on exactly the checks that assert credentials, a cluster-admin binding, and shared coordinator state are absent.

Both name rules now resolve their step through a single guard in _line_step, so zero observations are NOT_RUN. Regressions pin all five real catalog specs (not synthetic look-alikes) across four axes each: silent, whitespace-only, healthy-evidence PASS, and forbidden-name FAIL. A further regression drives the whole 63-check catalog with every other check held at PASS and asserts the verdict is NO_GO, with an all-PASS control asserting GO.

Both reachable silence paths are reproduced against the real tools:

  • Shell pipeline — a POSIX pipeline reports the status of its last stage, so the harness's own frozen env_names template with a missing upstream binary exits 0 with 0 bytes. Executed through /bin/sh; fully hermetic.
  • kubectl jsonpath — a drifted parent field, a drifted leaf field, and an unmatched container filter each exit 0 with no usable evidence. The exact output bytes are pinned hermetically; a second test confirms the live rc=0 behaviour against a read-only kubectl get and skips when no cluster is reachable.

The pool.coordinator-retains-sole-state-ownership projection now emits one <volume>=<claim> line per template volume, so a volume without a PVC still counts as an observation and only real drift produces silence. Verified against a live StatefulSet.

Closed hardening and evidence defects

  • Repository pin boundary — Gitea paths are pinned to atlas/titan-iac on an exact segment boundary, so atlas/titan-iac-evil and atlas/titan-iacx are refused.
  • Dot segments — any relative segment is rejected, including percent-encoded (%2e%2e) and double-encoded forms, before the request is built.
  • Impersonation--as/--as-group/--as-uid are refused for every argv, in every mode, from every vantage. The inner command of kubectl exec is re-checked under the same rule instead of being granted an unconditional exemption, and validate_catalog no longer guards only the operator vantage. No audited self-probe needs impersonation, so nothing is left enabled by convention.
  • Unattestable binariesflux and helm are removed from the allowlist; they had no pinned release digest and no catalog step used them. Flux/Helm status is read through kubectl. A regression asserts ALLOWED_BINARIES, EXPECTED_PATHS, EXPECTED_SHA256, and POD_COMMAND_PATHS have identical key sets, so an allowlisted binary the executor would refuse to attest cannot reappear.
  • Inert flags removed--concurrency was validated and stored but never used (run_catalog is sequential), and --expect-telegram-sessions was store_true with default=True so it could never be False. Both are gone along with the dead concurrency bound; Telegram continuity remains mandatory.
  • Armed create-uncertainty — the pull index is read page by page to its last page instead of one fixed limit=50, and the create response is treated as an authoritative source for the pull number. Every number either source names is registered and closed, so an unusable index can no longer orphan an open draft. When anything is uncertain the evidence carries residue_ref and an exact manual_cleanup command list. The previously verified path where a non-zero create is corroborated by an exact index match still passes.
  • Provenance for unrecorded steps — bulk-evidence steps keep executable_path and executable_sha256, so withholding bytes no longer withholds binary attestation.
  • Runbook format command — the documented ruff format --check scope now names the files this change owns and passes verbatim (33 files, exit 0). The previous testing/quality_*.py glob pulled in the unmodified, not-format-clean testing/quality_gate.py.
  • No repo-wide contract broadening — the hygiene.legacy_line_exceptions mechanism is reverted; testing/quality_hygiene.py is byte-identical to main and the contract change here is purely additive (74 insertions, 0 deletions). See the honest gate status below.

hermes_handoff_arming.py is split out of hermes_handoff_ephemeral.py so both stay under the 500-line cap.

Exact-head validation at 02e559d7

  • Canonical local quality gate: 754 tests pass; docs, smell, unit, and coverage all ok.
  • hygiene is failed, and it is red on main too. Exactly four pre-existing test files exceed 500 lines: scripts/tests/test_dashboards_render_atlas.py (516), testing/tests/test_hermes_chat_quality.py (2242), testing/tests/test_hermes_cli_lanes.py (1971), testing/tests/test_hermes_coordinator.py (510). All four are untouched here and byte-identical at merge-base 30259b52 and at current origin/main; a pristine export of each reports the same four issues and nothing else. Grandfathering or splitting them is a repo-wide contract decision that belongs to the canonical contract change in PR #14/#15, per the repair instruction not to add broad legacy LOC exceptions here.
  • Per-file coverage: all 16 handoff modules at >=95% line and branch (lowest 97.27% line, 95.27% branch); new hermes_handoff_arming.py at 100%/100%.
  • Deterministic mutation gate: 13/13 killed, including new mutants that restore the zero-evidence pass, restore the bare-prefix repository pin, stop re-checking a kubectl exec inner command, and drop attestation from an unrecorded step.
  • Ruff 0.16.3 check and format --check over every changed file: PASS. compileall: PASS. git diff --check origin/main...HEAD: PASS.
  • LOC: every file added or edited here is under 500 (largest: hermes_handoff_evaluators.py at 497).
  • Hermeticity: the full handoff suite passes under env -i with no PYTHONPATH/HOME/KUBECONFIG (425 passed, 1 live-cluster test skipped).
  • Adversarial re-runs: complete default-mode inventory is 63 checks / 15 groups / 75 steps with validate_catalog reporting 0 problems; 66 unique argv contain zero mutating commands (the only create tokens are side-effect-free kubectl auth can-i reviews), zero impersonation arguments, zero credential paths, and GET-only Gitea calls. A driven full run through main() executes 81 commands with zero mutations and zero POST/PATCH/DELETE. With no release inputs supplied, the run fails closed with 0 spawned commands.
  • Kustomize render plus client-only dry-run: PASS for services/hermes (113 docs), services/bstein-dev-home (35), clusters/atlas/flux-system (100).
  • detect-secrets 1.5.0 over every changed file: every hit is a pinned SHA-256 digest, a Git commit SHA, or a synthetic redaction fixture. Both new files scan clean.
  • Current main d8f2d818b9a552ea6c2d7fe86554be829bd5ffff: merge-tree clean. Recomputed over all 28 pairwise combinations, #19 is clean with main, #17, #18, and #20, and conflicts with #12, #14, #15, and #16. #18+#19 became clean because the duplicated hygiene mechanism is gone; #12+#19 is new because #12 advanced to ed43dbaa. The runbook names the exact overlapping files per pair and requires recomputation before integration.

Human review is required. This PR remains open and draft. Nothing was merged, force-pushed, deployed, reconciled, published, approved, or live-mutated, and no credential was read or created.

Repairs the independent review blockers on the same branch and PR. Previous reviewed head: `c808baff40a12a90f43c2e5e39b94d139f88ea29`. Repaired head: `02e559d7ac80cde6b2f2b87a4cba32bcb79d8971`. ### Mandatory P1: the zero-evidence fail-open is closed `evaluate_names_absent` returned `PASS` when its step exited `0` with no output. Five **mandatory** checks — `identity.no-provider-api-key-in-pod`, `identity.no-provider-api-key-in-manifest`, `forge.worker-holds-no-git-credential-variable`, `pool.coordinator-retains-sole-state-ownership`, `access.no-cluster-admin-binding-for-agent` — could therefore report a pass on no evidence at all, converting a `NO_GO` into a `GO` on exactly the checks that assert credentials, a cluster-admin binding, and shared coordinator state are *absent*. Both name rules now resolve their step through a single guard in `_line_step`, so zero observations are `NOT_RUN`. Regressions pin all five **real catalog specs** (not synthetic look-alikes) across four axes each: silent, whitespace-only, healthy-evidence `PASS`, and forbidden-name `FAIL`. A further regression drives the whole 63-check catalog with every other check held at `PASS` and asserts the verdict is `NO_GO`, with an all-`PASS` control asserting `GO`. Both reachable silence paths are reproduced against the real tools: - **Shell pipeline** — a POSIX pipeline reports the status of its *last* stage, so the harness's own frozen `env_names` template with a missing upstream binary exits `0` with 0 bytes. Executed through `/bin/sh`; fully hermetic. - **kubectl jsonpath** — a drifted parent field, a drifted leaf field, and an unmatched container filter each exit `0` with no usable evidence. The exact output bytes are pinned hermetically; a second test confirms the live `rc=0` behaviour against a read-only `kubectl get` and skips when no cluster is reachable. The `pool.coordinator-retains-sole-state-ownership` projection now emits one `<volume>=<claim>` line per template volume, so a volume without a PVC still counts as an observation and only real drift produces silence. Verified against a live StatefulSet. ### Closed hardening and evidence defects - **Repository pin boundary** — Gitea paths are pinned to `atlas/titan-iac` on an exact segment boundary, so `atlas/titan-iac-evil` and `atlas/titan-iacx` are refused. - **Dot segments** — any relative segment is rejected, including percent-encoded (`%2e%2e`) and double-encoded forms, before the request is built. - **Impersonation** — `--as`/`--as-group`/`--as-uid` are refused for every argv, in every mode, from every vantage. The inner command of `kubectl exec` is re-checked under the same rule instead of being granted an unconditional exemption, and `validate_catalog` no longer guards only the operator vantage. No audited self-probe needs impersonation, so nothing is left enabled by convention. - **Unattestable binaries** — `flux` and `helm` are removed from the allowlist; they had no pinned release digest and no catalog step used them. Flux/Helm status is read through `kubectl`. A regression asserts `ALLOWED_BINARIES`, `EXPECTED_PATHS`, `EXPECTED_SHA256`, and `POD_COMMAND_PATHS` have identical key sets, so an allowlisted binary the executor would refuse to attest cannot reappear. - **Inert flags removed** — `--concurrency` was validated and stored but never used (`run_catalog` is sequential), and `--expect-telegram-sessions` was `store_true` with `default=True` so it could never be `False`. Both are gone along with the dead concurrency bound; Telegram continuity remains mandatory. - **Armed create-uncertainty** — the pull index is read page by page to its last page instead of one fixed `limit=50`, and the create response is treated as an authoritative source for the pull number. Every number either source names is registered and closed, so an unusable index can no longer orphan an open draft. When anything is uncertain the evidence carries `residue_ref` and an exact `manual_cleanup` command list. The previously verified path where a non-zero create is corroborated by an exact index match still passes. - **Provenance for unrecorded steps** — bulk-evidence steps keep `executable_path` and `executable_sha256`, so withholding bytes no longer withholds binary attestation. - **Runbook format command** — the documented `ruff format --check` scope now names the files this change owns and passes verbatim (33 files, exit 0). The previous `testing/quality_*.py` glob pulled in the unmodified, not-format-clean `testing/quality_gate.py`. - **No repo-wide contract broadening** — the `hygiene.legacy_line_exceptions` mechanism is reverted; `testing/quality_hygiene.py` is byte-identical to `main` and the contract change here is **purely additive** (74 insertions, 0 deletions). See the honest gate status below. `hermes_handoff_arming.py` is split out of `hermes_handoff_ephemeral.py` so both stay under the 500-line cap. ### Exact-head validation at `02e559d7` - Canonical local quality gate: **754 tests pass**; `docs`, `smell`, `unit`, and `coverage` all `ok`. - **`hygiene` is `failed`, and it is red on `main` too.** Exactly four pre-existing test files exceed 500 lines: `scripts/tests/test_dashboards_render_atlas.py` (516), `testing/tests/test_hermes_chat_quality.py` (2242), `testing/tests/test_hermes_cli_lanes.py` (1971), `testing/tests/test_hermes_coordinator.py` (510). All four are untouched here and byte-identical at merge-base `30259b52` and at current `origin/main`; a pristine export of each reports the same four issues and nothing else. Grandfathering or splitting them is a repo-wide contract decision that belongs to the canonical contract change in PR #14/#15, per the repair instruction not to add broad legacy LOC exceptions here. - Per-file coverage: all **16** handoff modules at >=95% line **and** branch (lowest 97.27% line, 95.27% branch); new `hermes_handoff_arming.py` at 100%/100%. - Deterministic mutation gate: **13/13 killed**, including new mutants that restore the zero-evidence pass, restore the bare-prefix repository pin, stop re-checking a `kubectl exec` inner command, and drop attestation from an unrecorded step. - Ruff 0.16.3 `check` and `format --check` over every changed file: PASS. `compileall`: PASS. `git diff --check origin/main...HEAD`: PASS. - LOC: every file added or edited here is under 500 (largest: `hermes_handoff_evaluators.py` at 497). - Hermeticity: the full handoff suite passes under `env -i` with no `PYTHONPATH`/`HOME`/`KUBECONFIG` (425 passed, 1 live-cluster test skipped). - Adversarial re-runs: complete default-mode inventory is 63 checks / 15 groups / 75 steps with `validate_catalog` reporting 0 problems; 66 unique argv contain **zero** mutating commands (the only `create` tokens are side-effect-free `kubectl auth can-i` reviews), zero impersonation arguments, zero credential paths, and `GET`-only Gitea calls. A driven full run through `main()` executes 81 commands with zero mutations and zero `POST`/`PATCH`/`DELETE`. With no release inputs supplied, the run fails closed with **0** spawned commands. - Kustomize render plus client-only dry-run: PASS for `services/hermes` (113 docs), `services/bstein-dev-home` (35), `clusters/atlas/flux-system` (100). - `detect-secrets` 1.5.0 over every changed file: every hit is a pinned SHA-256 digest, a Git commit SHA, or a synthetic redaction fixture. Both new files scan clean. - Current `main` `d8f2d818b9a552ea6c2d7fe86554be829bd5ffff`: merge-tree **clean**. Recomputed over all 28 pairwise combinations, #19 is clean with `main`, #17, #18, and #20, and conflicts with #12, #14, #15, and #16. #18+#19 became clean because the duplicated hygiene mechanism is gone; #12+#19 is new because #12 advanced to `ed43dbaa`. The runbook names the exact overlapping files per pair and requires recomputation before integration. Human review is required. This PR remains open and draft. Nothing was merged, force-pushed, deployed, reconciled, published, approved, or live-mutated, and no credential was read or created.
hermes-automation added 1 commit 2026-08-17 10:14:47 +00:00
Decides whether the Hermes platform handoff is fit to release, and refuses
to round an absence of evidence up to a pass.

The harness is read-only by default and classifies 71 checks PASS / FAIL /
NOT_RUN / NOT_APPLICABLE. Any mandatory FAIL or NOT_RUN is NO_GO, and so is a
harness-level problem: an unreachable vantage, a catalog entry whose evidence
no longer exists, an expired deadline, or an evaluator that raised.

Evidence comes from two vantages that cannot cover for each other: an external
read-only operator kubeconfig, and the Hermes agent probing itself from inside
its own pod. Before any check runs, the harness asks each vantage who it is and
stops if they are the same principal, because dual-vantage evidence from one
identity is a restatement rather than a corroboration. `--as` is rejected for
every operator-side command and reachable only as the inner command of a
`kubectl exec`, so impersonation can never stand in for a real self-probe. A
deny check needs a live refused request, not only an authorization review.

Two safety properties are structural rather than conventional, enforced where
an argv becomes a subprocess: the default mode mutates nothing (mutating verbs
require a server dry run; there is deliberately no live TokenRequest probe,
because a successful one would mint a real credential), and no probe can pull a
credential value into a report (no vault/sops/curl, secrets readable only with
-o name, environment probes list names, shell only through frozen reviewed
templates). Captures are bounded before they are screened, and the rendered
report is re-screened before it is written.

Mutation lives behind a separate arming flag with an exact confirmation phrase,
a caller-supplied unique ref, a preflight that refuses a protected push target
before any network call, and a cleanup whose verification is itself mandatory.
A default run reports those four checks NOT_RUN.

The catalog is declarative so a reviewer reads what is asserted rather than how
it is plumbed, and so structural properties can be proven over every entry
before a run. Catalog drift surfaces as NOT_RUN, which stops the release.

docs/hermes_full_handoff_acceptance.md carries the merge order for PRs #14-#18
on top of the merged #13 baseline, the image build and Flux rollout, the
rollback point for each step, the go/no-go checklist, and the limits that are
asserted rather than exercised.

Validation: 295 handoff tests pass with 100% line coverage on all 15 new
modules; the full unit suite is 647 passed with two failures that reproduce
unchanged on origin/main; Ruff, py_compile, kustomize render, and a diff
credential screen are clean; a live read-only run against Atlas returns NO_GO
for the pre-merge cluster with no unscreened fields in the report.
hermes-automation added 2 commits 2026-08-17 13:13:58 +00:00
hermes-automation added 2 commits 2026-08-17 16:27:35 +00:00
evaluate_names_absent returned PASS when its step exited 0 with no output,
so five mandatory checks - the ones asserting that provider API keys, forge
credentials, a cluster-admin binding, and shared coordinator state are
absent - could report a pass on no evidence and turn a NO_GO into a GO.
Both name rules now resolve their step through one guard in _line_step, so
zero observations are NOT_RUN. Regressions pin all five real catalog specs
plus both reachable silence paths: a POSIX pipeline whose status comes from
its last stage, and a drifted kubectl -o jsonpath. The pool claim projection
emits one <volume>=<claim> line per template volume so a volume without a
PVC still counts as an observation rather than reading as drift.

Also closes the review's reachable hardening and evidence defects:

- pin Gitea paths to atlas/titan-iac on an exact segment boundary and
  reject relative segments, including percent-encoded ones
- forbid impersonation structurally in every mode and vantage; the inner
  command of kubectl exec is re-checked rather than exempted, and
  validate_catalog no longer guards only the operator vantage
- drop flux and helm from the binary allowlist; they had no pinned release
  digest, so no allowlisted binary can now be admitted that the executor
  would refuse to attest
- remove the inert --concurrency and --expect-telegram-sessions flags and
  the dead concurrency bound; Telegram continuity stays mandatory
- read the ephemeral pull index page by page, treat the create response as
  an authoritative source for the pull number, close every number either
  source names, and surface residue_ref plus exact manual_cleanup commands
  when creation is uncertain
- keep executable_path and executable_sha256 on unrecorded bulk-evidence
  steps so withholding bytes never withholds binary attestation
- revert the repo-wide hygiene legacy-exception mechanism; the contract
  change here is purely additive and the four pre-existing over-cap files
  are left to the canonical contract change in PR #14/#15
- correct the runbook ruff format scope so the documented command passes

Split hermes_handoff_arming.py out of hermes_handoff_ephemeral.py to keep
both modules under the 500-line cap. All 16 handoff modules hold at least
95% line and branch coverage; the mutation gate is 13/13.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Refresh #12 and #18 to their live Gitea heads, record #19's replaced head
exactly, and replace the pairwise conflict set with a freshly computed one
over all 28 combinations. #18+#19 is now clean because #19 no longer
duplicates the repo-wide hygiene legacy-exception mechanism that #18 also
carries; #12+#19 is new because #12 advanced. Name the exact overlapping
files per conflicting pair and how to resolve each additively.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
bstein added 25 commits 2026-08-18 09:33:42 +00:00
The local goal judge scored every worker report against "did the reviewed
implementation reach a shippable state". A read-only reviewer that returned a
completed BLOCK verdict with findings was therefore resumed turn after turn with
an instruction to repair code it was forbidden to touch (observed live on
t_dbdcd739), burning subscription capacity and risking an unbounded loop.

Completion is now judged against the action the card assigned:

* Cards declare their role explicitly with Hermes-Task-Role / Hermes-Expected-
  Output metadata. Pre-contract cards fall back to a narrow inference that needs
  a read-only scope, a requested verdict, no requested mutation deliverable, and
  a report that changed no files.
* Role resolution reads only the card itself. Prior attempts, parent results,
  cross-task history and comments appended to the worker context can no longer
  reassign the role.
* Review, audit and diagnostic cards finalize deterministically on a truthful
  SHIP or BLOCK verdict with evidence, and fail closed on a missing,
  unrecognized or self-contradictory verdict, on a BLOCK without findings, and
  on a verdict without evidence. Every rejection reason carries the read-only
  guard, so a resumed review is never told to edit the reviewed implementation.
* Implementation cards keep the fail-closed model judge unchanged, including the
  unfinished-work heuristic and judge-unavailable rejection.
* All judge reasons are bounded, single-line and secret-redacted before they
  reach Kanban metadata, comments and continuation prompts.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The distributed worker pool stacks on exactly two open pull requests, in this
order: PR 14 supplies the broker-only SCM boundary the mediators route through
and the hermes-scm-boundary-v2 ConfigMap they mount, and PR 15 supplies the
cli_lane_* decomposition -- including canonical_run_id and the eligibility
predicate on claim_ready -- that the coordinator depends on. Neither can be
dropped without breaking a fixed P0 boundary, so both are carried here as
prerequisites and this branch must not merge before them.

PR 16 (agent image release lane) and PR 19 (full-handoff acceptance harness)
are NOT prerequisites and are deliberately absent, so reviewing this branch no
longer means approving them.

PR 14 and PR 15 conflict with each other in nine paths. Each is resolved to the
resolution already reviewed on this branch at 4d4cf1bd.
Three fenced worker Pods claim Hermes Kanban runs through a coordinator that
owns every state transition, with per-ordinal HMAC authority, a mediated
broker-only SCM path, and durable per-ordinal workspaces.

Content is the reviewed head of PR #18 (689bcb6e) with PR 16's and PR 19's
contributions removed: they were merged in only to validate co-existence and are
not prerequisites, so this branch no longer carries them as ancestors. Only PR 14
and PR 15 remain, because the broker boundary and the cli_lane_* decomposition
are load-bearing for two of the fixed P0 boundaries.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Independent review t_5975c06a blocked this branch on a P1: a Kanban write that
failed while a lease expired left a `lease_failed` row that was invisible to
every pass, immortal to garbage collection, and fatal to the coordinator. It
poisoned `reconcile()` forever with a conflicting-duplicate primary key,
produced a spurious capability `block_task` from `dispatch()`, and -- because
startup maintenance ran unguarded before the port bound, against a store on a
PVC -- crash-looped the coordinator with no automatic recovery.

`lease_failed` is now a retryable state that every maintenance pass drains, and
a row only reaches a terminal state on authoritative evidence about its exact
Kanban run, so nothing is collected before its outcome is known and nothing is
silently dropped. Each row, task, and board is processed in isolation, and a
coordinator-side fault is never converted into a Kanban mutation. Startup runs
through the same guarded cycle as the steady-state loop.

The wire protocol and the durable store are now separate modules, and the
maintenance passes moved out of the coordinator, so each file stays under the
managed line ceiling with room for the recovery logic.

Also closes three consequential handoff risks the same review raised:

* mediator-N pinned itself hard to worker-N while sharing a ReadWriteOnce
  claim, so a drain or preemption that moved only the lower-priority worker
  deadlocked the ordinal on Multi-Attach until an operator deleted a Pod. The
  shared workspace is now ReadWriteMany (as the hermes-chat tenant workspaces
  already are on the same class), colocation is a preference, and the mediator
  shares the worker's preemption priority, so each Pod reschedules on its own.
* the broker permits only branch creation, so a retry that added commits could
  never submit and the run's work was discarded with the failure. Submission
  now targets a fresh attempt- or content-scoped ref in the same reviewed
  namespace -- never an update -- and is idempotent under replay. A refused
  submission downgrades the result and says why instead of unwinding the run.
* the provider CLIs were reinstalled into an emptyDir on every Pod start inside
  the 10m Flux health window for the whole hermes app. They now install once
  per pinned version onto a durable volume, re-verified against the real
  binaries and time-bounded, and the best-effort pool no longer gates the
  health of the app its dependents wait on.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Document what an operator needs to know about the repairs: why the ordinal
workspace claim is ReadWriteMany and why colocation is only a preference, that
`lease_failed` is retryable rather than terminal and what a persistent deferred
park means, why a retry publishes an attempt-scoped ref and what to expect from
the extra drafts, and why the best-effort pool is deliberately absent from the
hermes Kustomization health checks.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Repairs the blockers from the independent review of the previous head. The
role-aware verdict contract was correct but reachable only through one call
site and only on cards written with real newlines, so most real review cards
never used it.

* The role-blind call is no longer a weaker classifier that can reject before
  the role-aware judge runs. Without a card there is no defensible
  role-dependent judgement, so `unfinished_result_reason()` applies only the
  card-independent checks. That closes the short-circuit at every call site,
  including the one PR15 moves to `cli_lane_execution`, and it makes a
  journalled terminal record accepted under one version of these semantics
  re-validate under any other instead of being quarantined into a re-dispatch
  of an already-accepted task.
* The lane resolves the role from the card and passes it, and records it in
  the Kanban metadata. The verdict contract binds only where the lane can buy
  another turn: in single-shot mode a rejection discards the worker's real
  result, so review cards keep the relaxation without gaining any rejection
  single-shot mode did not already have.
* Card scope expands the literal \n escapes the board stores in one-line
  bodies, so explicit `Hermes-Task-Role` / `Hermes-Expected-Output` directives
  are honoured on the 6 of 78 live cards that carry no real newline, and the
  read-only, verdict and mutation heuristics stop being cut apart by them.
* Card scope now ends at the first non-card H2 and at the runner's controller
  evidence, which is emitted under its own heading. Goal-controller rejection
  history can no longer sit inside the card, and an upstream heading rename
  fails closed instead of admitting history into role resolution.
* The inference recognises the SHIP/BLOCK-shaped deliverables real cards
  actually use: 21 of 78 live cards resolve to review, up from 10, with no
  implementation card misclassified. Cards asking for a findings list rather
  than a verdict deliberately stay on the model judge, since the verdict is
  the review contract's only gate.
* A declared review role no longer outranks mutation evidence: a report that
  changed files falls back to the implementation regime.
* Judge reasons go through the agent runtime's canonical redactor, extended
  for the two shapes it deliberately passes through and this lane handles -
  `scheme://user:secret@host` and a credential named in prose - while a 40-hex
  commit SHA survives as evidence.

Regressions cover the recovered t_dbdcd739 incident, a verbatim snapshot of
every live card on three boards with a hand-labelled expected role, the
upgrade and single-shot properties against the previous gate, the upstream
context-heading contract, and the end-to-end `execute_claim` shape that used
to burn every goal turn.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Add kanban routing keys (provider_quota_min_remaining_percent: 15,
provider_capacity_cooldown_seconds: 300, provider_auth_cooldown_seconds:
3600), a lane-metrics port/Service on 9011 with service-annotation
scraping, monitoring ingress for the new port, and the lane's quota
metrics URL env. Based on PR #15 (fix/hermes-result-decomposition-
reliability); stacked because this work rides on the decomposed lane.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Based on PR #15 (fix/hermes-result-decomposition-reliability); stacked
on the decomposed cli_lane modules.

- cli_lane_quota: soft-exclude a provider from NEW cli-auto work below
  the remaining-quota threshold (both-below prefers more remaining;
  fetch failure fails open with a metric).
- cli_lane_health: lane now writes provider health (G7) with classified
  failure reasons splitting the capacity conflation (quota/auth/
  rate-limit/transport) and cooldown hysteresis; re-admission only on
  full cooldown expiry, passed quota reset, or fresh success (G4).
- cli_lane_routing: capacity-limited health now excludes a provider
  (G3); cooldown/reset-aware re-admission.
- cli_lane_failover: explicit cli-codex-*/cli-claude-* assignees fail
  closed as transient instead of switching providers (G5); fallback
  depth stays bounded at two hosted providers (G1) with effort
  preserved; Switchyard outages block transient, not capability (G9).
- cli_lane_metrics: route-decision/fallback counters, quota and
  soft-exclusion gauges, pod-local scrape server (G6).
- cli_lane_provider: worker env drops ANTHROPIC_API_KEY, CLAUDE_API_KEY,
  OPENAI_API_KEY, API_SERVER_KEY so no metered path exists (G10).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Deterministic coverage for the quota-aware lane: threshold boundaries
(14.9/15/15.1), both-below preference, fetch-failure fail-open, cooldown
elapsed-vs-not hysteresis, quota-reset recovery (never for auth),
explicit fail-closed in both directions, bounded double-failure block,
failure-reason classification, metrics emission, and worker env key
stripping. Based on PR #15 (fix/hermes-result-decomposition-reliability).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
CAPACITY_PATTERN (the gate that sets result.capacity_failure) lacked the
bare unauthorized/forbidden/401/403 signals that classify_capacity_failure
already recognizes, so an auth blip surfacing only as "403 Forbidden"
was blocked as capability instead of transient: no auto failover, no
health cooldown. Add 401|403|unauthorized|forbidden to the gate so it
matches the classifier; reason classification still distinguishes auth
from quota/rate-limit/transport.

Tests: an auto card failing with only "403 Forbidden" now fails over to
the other provider, records an auth cooldown (authenticated:false), and
classifies the fallback reason as auth in metrics; a manual card with the
same failure still fails closed as transient. Router-outage tests moved to
test_hermes_cli_router_outage.py to keep both files <500 LOC.

Based on PR #15 (fix/hermes-result-decomposition-reliability).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Reconcile two independent test/gate reorganizations:
- Gate/semgrep/mailu: keep main's #16 dual-metric implementation.
- quality_contract.json: union #16 image-builder + #14 scm/node + #15 cli_lane.
- agent-deployment.yaml: keep #14 gitea removal + #16 image-build-token + #15 probe.
- Test splits: main's chat/coordinator/agent organization is authoritative;
  drop #15's redundant competing splits and #14's stale cli-lane duplicates;
  keep #15's cli-lane decomposition suite and port the execution-safety test.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
# Conflicts:
#	services/hermes/networkpolicy.yaml
#	testing/quality_contract.json
#	testing/tests/test_hermes_agent_security.py
# Conflicts:
#	services/hermes/scripts/cli_lane_runner.py
# Conflicts:
#	scripts/tests/test_dashboards_render_atlas_drilldowns.py
#	scripts/tests/test_dashboards_render_jobs.py
#	services/hermes/scm-common/scripts/scm_broker.py
#	services/hermes/scripts/cli_lane_dispatch.py
#	services/hermes/scripts/cli_lane_execution.py
#	testing/quality_contract.json
#	testing/tests/test_hermes_agent_access.py
#	testing/tests/test_hermes_agent_security.py
#	testing/tests/test_hermes_chat_config.py
#	testing/tests/test_hermes_chat_images.py
#	testing/tests/test_hermes_chat_provider_auth.py
#	testing/tests/test_hermes_chat_quality.py
#	testing/tests/test_hermes_chat_support.py
#	testing/tests/test_hermes_chat_voice.py
#	testing/tests/test_hermes_cli_finalization_edges.py
#	testing/tests/test_hermes_cli_foundation_coverage.py
#	testing/tests/test_hermes_cli_lanes_configuration.py
#	testing/tests/test_hermes_cli_recovery_edges.py
#	testing/tests/test_hermes_cli_retention_edges.py
#	testing/tests/test_hermes_coordinator.py
#	testing/tests/test_hermes_coordinator_boards.py
#	testing/tests/test_hermes_coordinator_support.py
Post-merge fixes after rebasing the distributed worker pool onto the
review-goal-semantics train tip:

- Pool SCM submission tests target the train's relocated receive-pack
  scanner: FEATURE_REF_RE now lives in receive_pack_scan, bodies are
  built via the shared _receive_command helper (valid pack), and the
  update-rejection assertion matches the train's message.
- Take the train's canonical test_hermes_scm_broker,
  test_hermes_cli_dispatch_runtime and test_hermes_cli_execution_edges,
  which exercise the train's broker/dispatch/execution behavior.
- Runtime staging tests patch os.fchown alongside os.chown so the
  UID-10000 _write_secret path passes under a non-root gate runner
  (production ownership behavior unchanged).
- Split the jenkins build-evidence contracts out of
  test_hermes_runtime_access into test_hermes_runtime_evidence to keep
  both files under the 500-LOC hygiene ceiling after the merge.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
# Conflicts:
#	ci/scripts/semgrep_report.py
#	testing/quality_contract.json
#	testing/quality_coverage.py
#	testing/tests/conftest.py
#	testing/tests/test_hermes_auto_router.py
#	testing/tests/test_quality_contract.py
#	testing/tests/test_quality_coverage_helpers.py
#	testing/tests/test_semgrep_report.py
atlas-reconciler changed title from WIP: Fail-closed Hermes full-handoff acceptance harness to Fail-closed Hermes full-handoff acceptance harness 2026-08-18 21:26:52 +00:00
atlas-reconciler manually merged commit 1c4ed3af93 into main 2026-08-18 21:28:00 +00:00
Sign in to join this conversation.
No Reviewers
No Label
2 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: titan/atlas-iac#19
No description provided.