The repair action no longer creates a Kubernetes Job and polls it. Ariadne
patches the fixture ConfigMap directly through its own k8s client, which
removes roughly 40 seconds of pod scheduling from the loop, drops the two
failure modes that Job introduced (volume attach and node selection), and
turns an opaque pod log into an Ariadne event.
- execute_repair returns {action, target, succeeded, error} and issues one
merge patch; no Job, no polling, no injectable clock, never raises
- the patch writes the same terminal value every time, so idempotency needs
no duplicate guard; one action per incident is still enforced upstream
- orchestrator records {repair, target}; the state machine, rebuild trigger
and failure path are unchanged
- retires the unused repair-image setting
Ariadne's service account now needs get+patch on that one ConfigMap by name
instead of Job create; the batch/jobs grant and the hermes-demo-repair
service account can be retired.
480 pass in the hermes suite.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Second entry in the action registry, proving it is a real extension point.
No cluster mutation: the action is one Jenkins rebuild.
- hermes_infra_signals: reviewable marker set across DNS/connectivity,
image pull, upstream 5xx and agent-channel loss; Ariadne independently
confirms a marker in the evidence before any retry, and records which
marker justified it. "no space left on device" is deliberately excluded
because a retry lands on the same full volume.
- decision: classification -> action registry (action_classifications),
falling back to the previous single-classification behavior
- repair: retry_build posts to /build for unparameterized real jobs and
buildWithParameters for the fixture demo job
- orchestrator: retry path records requested/accepted/executed and moves
the incident to awaiting_rebuild so the existing success path resolves
it; one action per incident still enforced, so a retry cannot loop
- events layer split out of the orchestrator to stay under the LOC cap
Motivated by real failures tonight: pip DNS resolution and a Gitea 443
connect timeout, plus live incident metis/271 (SCM checkout timeout).
30 new tests; 368 pass in the hermes suite.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Longhorn RWO attach latency and per-node engine availability made the
PVC fixture unreliable for a fast repeatable demo. The fixture is now a
ConfigMap; the repair Job runs bitnami/kubectl under the least-privilege
hermes-demo-repair SA and patches state=healthy.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Non-worker nodes (titan-2x lane) are not Longhorn-ready; volume attach
fails there. Same constraint as the demo test-runner.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>