Repairs the blockers from the independent review of the previous head. The
role-aware verdict contract was correct but reachable only through one call
site and only on cards written with real newlines, so most real review cards
never used it.
* The role-blind call is no longer a weaker classifier that can reject before
the role-aware judge runs. Without a card there is no defensible
role-dependent judgement, so `unfinished_result_reason()` applies only the
card-independent checks. That closes the short-circuit at every call site,
including the one PR15 moves to `cli_lane_execution`, and it makes a
journalled terminal record accepted under one version of these semantics
re-validate under any other instead of being quarantined into a re-dispatch
of an already-accepted task.
* The lane resolves the role from the card and passes it, and records it in
the Kanban metadata. The verdict contract binds only where the lane can buy
another turn: in single-shot mode a rejection discards the worker's real
result, so review cards keep the relaxation without gaining any rejection
single-shot mode did not already have.
* Card scope expands the literal \n escapes the board stores in one-line
bodies, so explicit `Hermes-Task-Role` / `Hermes-Expected-Output` directives
are honoured on the 6 of 78 live cards that carry no real newline, and the
read-only, verdict and mutation heuristics stop being cut apart by them.
* Card scope now ends at the first non-card H2 and at the runner's controller
evidence, which is emitted under its own heading. Goal-controller rejection
history can no longer sit inside the card, and an upstream heading rename
fails closed instead of admitting history into role resolution.
* The inference recognises the SHIP/BLOCK-shaped deliverables real cards
actually use: 21 of 78 live cards resolve to review, up from 10, with no
implementation card misclassified. Cards asking for a findings list rather
than a verdict deliberately stay on the model judge, since the verdict is
the review contract's only gate.
* A declared review role no longer outranks mutation evidence: a report that
changed files falls back to the implementation regime.
* Judge reasons go through the agent runtime's canonical redactor, extended
for the two shapes it deliberately passes through and this lane handles -
`scheme://user:secret@host` and a credential named in prose - while a 40-hex
commit SHA survives as evidence.
Regressions cover the recovered t_dbdcd739 incident, a verbatim snapshot of
every live card on three boards with a hand-labelled expected role, the
upgrade and single-shot properties against the previous gate, the upstream
context-heading contract, and the end-to-end `execute_claim` shape that used
to burn every goal turn.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The local goal judge scored every worker report against "did the reviewed
implementation reach a shippable state". A read-only reviewer that returned a
completed BLOCK verdict with findings was therefore resumed turn after turn with
an instruction to repair code it was forbidden to touch (observed live on
t_dbdcd739), burning subscription capacity and risking an unbounded loop.
Completion is now judged against the action the card assigned:
* Cards declare their role explicitly with Hermes-Task-Role / Hermes-Expected-
Output metadata. Pre-contract cards fall back to a narrow inference that needs
a read-only scope, a requested verdict, no requested mutation deliverable, and
a report that changed no files.
* Role resolution reads only the card itself. Prior attempts, parent results,
cross-task history and comments appended to the worker context can no longer
reassign the role.
* Review, audit and diagnostic cards finalize deterministically on a truthful
SHIP or BLOCK verdict with evidence, and fail closed on a missing,
unrecognized or self-contradictory verdict, on a BLOCK without findings, and
on a verdict without evidence. Every rejection reason carries the read-only
guard, so a resumed review is never told to edit the reviewed implementation.
* Implementation cards keep the fail-closed model judge unchanged, including the
unfinished-work heuristic and judge-unavailable rejection.
* All judge reasons are bounded, single-line and secret-redacted before they
reach Kanban metadata, comments and continuation prompts.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The sticky-block gate added in the previous commit classifies a task from
the `created` event payload that upstream `create_task` writes. That
producer is code we do not own, so trusting it silently was the gap: if
upstream renamed the key, dropped it, or stopped deriving it from
`initial_status`, the image would still build and ship a consumer that
mis-classifies every task it reads.
Anchor the producer contract at build time, before the regression suite
runs, with three assert-only preconditions: the `initial_status="blocked"`
park resolves `task_status` to `"blocked"`, every non-park creation
resolves it to something else, and the `created` event carries that same
variable under `"status"`. None of them rewrite the producer.
Textual anchors cannot see dataflow, so add the runtime net the reviewer
asked for. The suite now drives the real API: create + claim an ordinary
task, trip the circuit breaker once at failure_limit=1 so it parks with a
`gave_up` event (leaving its own `created` event as the most recent
create/block/unblock row), then recompute at failure_limit=2 and require
promotion to ready. That case is red under an unconditional-true created
predicate and red under producer drift that labels every created event
blocked, while the explicit block/unblock, dependency-promotion and
circuit-breaker-at-current-limit cases stay green. Non-blocked and
malformed created payloads are pinned as controls, and the gate now
rejects non-dict payloads rather than trusting `.get`.
Also make the live placement correction durable: titan-04 is cordoned
after repeated kernel undervoltage and kubelet failure and titan-19 was
probe/Longhorn unstable under worker load, so both join the hard NotIn
list; titan-05 is healthy but sits at 3592m/3600m requested CPU, so the
main hermes container gives back 50m (350m -> 300m) to schedule there.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three narrowly scoped Hermes reliability fixes backed by live evidence
from the Cassandra/titan-iac proof run.
Worker concurrency. Three simultaneous direct CLI workers on the 4-core
hermes-agent node drove load to ~45 and made the hermes and oauth2-proxy
containers fail their probes, leaving the pod 8/10 Ready; two workers
stayed at 10/10. Cap HERMES_CLI_LANE_CONCURRENCY at 2 and lower the
cli-lane-runner CPU limit from 3 to 2 so the dashboard and auth sidecars
keep a guaranteed share of the node. Requests are unchanged: the pod
still asks for 745m total, so placement does not move.
Service links. Kubernetes injects a service-link variable pair for every
service in the namespace, and hermes-claude-broker produces
HERMES_CLAUDE_BROKER_PORT=tcp://10.43.31.76:9006 — a value the broker
parses as an int. That contaminated worker and test environments even
though the deployment already addresses every service by DNS name. Set
enableServiceLinks: false on the hermes-agent pod spec.
Blocked-task scheduling. create_task(initial_status="blocked") records a
created event carrying status=blocked but never a blocked event, while
_has_sticky_block() only inspects blocked/unblocked events. recompute_ready()
considers blocked tasks, so an explicitly parked task with no incomplete
parent auto-promoted on the next dispatcher cycle. Teach _has_sticky_block()
to also recognize a created event whose payload status is blocked, which
covers tasks created before this image patch without adding a persisted
field. Dependency-driven promotion and the circuit-breaker failure-limit
guard are untouched; unblock_task() still releases either kind of block.
hermes-kanban-blocked-regression.py runs against the real upstream
kanban_db API during the image build, so the build fails if any of these
semantics regress.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>