7 Commits

Author SHA1 Message Date
codex
136caa8477 feat(hermes): search a service's own namespace for log evidence
All checks were successful
Tests / Declarative: Post Actions passed: 1200
Log evidence covered the demo namespace plus jenkins. That is right for the
common case, since CI failures happen in Jenkins agent pods, but it means a
build failure that correlates with the service itself being unhealthy carries
no trace of the service at all.

ARIADNE_HERMES_JOB_NAMESPACES maps a job to its own namespace, which is added
alongside jenkins rather than replacing it. Unmapped jobs are unchanged, and a
namespace already in the list is never duplicated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 12:32:59 -03:00
codex
281f06dccd feat(hermes): triage the branches inside multibranch folders
All checks were successful
Tests / Declarative: Post Actions passed: 1192
A multibranch project is a folder, not a job: its /api/json carries no
lastBuild at all, only a jobs array with one child per branch. Detection read
lastBuild, so hermes-code-demo-branches was returned as skipped on every tick
since it was created and no branch could ever be triaged.

Expand a folder into its branch jobs and process each. Jenkins addresses them
as parent/job/branch, which the existing fetch already builds correctly, and
branch names stay URL-encoded because re-encoding them yields a 404.

Expansion is capped by ARIADNE_HERMES_MAX_BRANCHES (default 5): a repository
with many active branches would otherwise multiply the watched job count
without limit and could open one incident per failing branch in a single tick.
A name containing a slash is refused outright, since it would escape the
parent and address an unrelated job.

The Jenkins transport moves to its own module so the orchestrator reads as
decision logic, and jobs[name] rides along on the existing request rather than
costing a second call per job per tick.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 11:01:57 -03:00
codex
1b29b56f50 feat(hermes): escalate builds that never reach a terminal result
Detection only ever considered builds that finished, so a build wedged on a
network call was invisible: no incident, no issue, no alert, while it held one
of the five Jenkins agent slots. Observed on metis build 272, which sat on
apt-get update for 75 minutes and starved the triage demo of an agent with
nothing anywhere saying so.

A build still running past ARIADNE_HERMES_HUNG_BUILD_MINUTES (default 45) now
files an issue and lands human_required. No model is consulted: the console is
still being written, so a root cause would be invented. Escalates once per
build, and a cap of zero disables the check.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 05:13:11 -03:00
codex
b2fcad4116 refactor(hermes-triage): repair the fixture in-process, not via a spawned Job
The repair action no longer creates a Kubernetes Job and polls it. Ariadne
patches the fixture ConfigMap directly through its own k8s client, which
removes roughly 40 seconds of pod scheduling from the loop, drops the two
failure modes that Job introduced (volume attach and node selection), and
turns an opaque pod log into an Ariadne event.

- execute_repair returns {action, target, succeeded, error} and issues one
  merge patch; no Job, no polling, no injectable clock, never raises
- the patch writes the same terminal value every time, so idempotency needs
  no duplicate guard; one action per incident is still enforced upstream
- orchestrator records {repair, target}; the state machine, rebuild trigger
  and failure path are unchanged
- retires the unused repair-image setting

Ariadne's service account now needs get+patch on that one ConfigMap by name
instead of Job create; the batch/jobs grant and the hermes-demo-repair
service account can be retired.

480 pass in the hermes suite.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 22:29:57 -03:00
codex
950d014707 feat(hermes-triage): real services get both an issue and a patch proposal
The code path was gated on a single job id, so homegrown services could
never produce a pull request. Escalated incidents on mapped repositories
now additionally attempt a bounded patch proposal, and the filed issue
links it.

- hermes_code_flow.propose_for_incident: additive entry point invoked only
  from the escalation branch, so an auto-remediated failure never also
  gets a patch and a failed Hermes run never spends tokens on one
- three gates before any HTTP call: code enabled, not the legacy demo job,
  and the job has a repo mapping; unmapped jobs make zero calls
- never raises: a failed proposal cannot change the incident outcome or
  break the tick
- issue body links the proposal when a pull request was opened
- propose_code_fix added to the bounded action-label set

The legacy demo-job short-circuit is untouched and all of its tests pass
unchanged.

8 new tests; 480 pass in the hermes suite.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 21:29:54 -03:00
codex
6da560810f feat(hermes-triage): Gitea issues for escalations, patch proposals for real repos
Two capabilities that make triage useful outside the demo surface.

Issues: when triage concludes a human is needed, file an issue in the
failing service's own repository carrying classification, confidence, the
facts with their sources, the inferences and a Jenkins link, plus a footer
stating Hermes has no write access and nothing was changed. Opt-in per job
via a repo map, deduplicated by job+classification so a repeatedly failing
job yields one issue per kind of failure rather than one per build, and
capped per tick. Disabled by default.

Real-repo patches: candidate files are selected from the console failure
regions (Python, Rust and JS/TS reference patterns), filtered to each
repo's allowed prefixes and suffixes, ranked earliest-failure-first with
source preferred over test files, and fetched whole - never truncated,
because a patch anchor must match exactly. Per-job owner/repo/base-branch
resolution; the patch is validated against the file the model actually
chose, and an unlisted path is rejected.

Legacy single-repo demo behaviour is preserved unchanged.

131 new tests; 472 pass in the hermes suite.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 21:04:32 -03:00
codex
8c65b7bd60 feat(hermes-triage): retry_transient_infra action for connectivity failures
Second entry in the action registry, proving it is a real extension point.
No cluster mutation: the action is one Jenkins rebuild.

- hermes_infra_signals: reviewable marker set across DNS/connectivity,
  image pull, upstream 5xx and agent-channel loss; Ariadne independently
  confirms a marker in the evidence before any retry, and records which
  marker justified it. "no space left on device" is deliberately excluded
  because a retry lands on the same full volume.
- decision: classification -> action registry (action_classifications),
  falling back to the previous single-classification behavior
- repair: retry_build posts to /build for unparameterized real jobs and
  buildWithParameters for the fixture demo job
- orchestrator: retry path records requested/accepted/executed and moves
  the incident to awaiting_rebuild so the existing success path resolves
  it; one action per incident still enforced, so a retry cannot loop
- events layer split out of the orchestrator to stay under the LOC cap

Motivated by real failures tonight: pip DNS resolution and a Gitea 443
connect timeout, plus live incident metis/271 (SCM checkout timeout).

30 new tests; 368 pass in the hermes suite.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 20:42:48 -03:00