From 70d8f5f35fbe55b88806ff07908c11c477380c2e Mon Sep 17 00:00:00 2001 From: jenkins Date: Wed, 5 Aug 2026 18:51:52 -0300 Subject: [PATCH] docs(hermes-triage): investigative-triage demonstration record + measured runbook timings Co-Authored-By: Claude Fable 5 --- knowledge/hermes-investigative-triage-demo.md | 74 +++++++++++++++++++ knowledge/runbooks/hermes-triage-demo.md | 26 ++++--- .../hermes-investigative-triage-demo.md | 74 +++++++++++++++++++ 3 files changed, 163 insertions(+), 11 deletions(-) create mode 100644 knowledge/hermes-investigative-triage-demo.md create mode 100644 services/comms/knowledge/hermes-investigative-triage-demo.md diff --git a/knowledge/hermes-investigative-triage-demo.md b/knowledge/hermes-investigative-triage-demo.md new file mode 100644 index 000000000..f3dfe91b7 --- /dev/null +++ b/knowledge/hermes-investigative-triage-demo.md @@ -0,0 +1,74 @@ +# Hermes Investigative Triage Demonstration + +- Date: 2026-08-05 +- Run ID: run_d2f61902c00044948a3912a94e895ab9 +- Duration: 128 seconds, status completed +- Usage: {"input_tokens": 200961, "output_tokens": 4370, "total_tokens": 205331} +- Mode: supervised, read-only; approvals `smart` with mutation deny-list; zero mutations, zero denied approvals +- Tool audit (from the run's SSE event stream): 8 skill loads (triage-titan-test-failures pack: orchestrator, jenkins-retained-evidence, platform-quality-metrics, kubernetes-readonly-failure-classifier, flux-git-change-correlation, ...), 4 read-only terminal batches (Ariadne internal API sweep, kubectl logs/jobs/configmap in hermes-triage-demo, VictoriaMetrics queries incl. last_over_time fallback, Flux/Gitea revision correlation) + +Unlike the automated loop (where Ariadne supplies a sanitized bundle), this run +was given ONLY an incident ID and gathered all evidence itself — the operator's +triage path executed by the agent. Notable behaviors: honest dead-end reporting +(Jenkins 403 -> pivot to Kubernetes-retained logs; Ariadne audit endpoints 404), +instant-query lookback miss -> switched to last_over_time, causal chain +reconstruction across five sources, and an unprompted, legitimate hardening +suggestion (add a build label to ariadne_hermes_triage_action_total). + +## Prompt + +```text +Use $triage-titan-test-failures. +Investigate Jenkins incident hermes-triage-demo/10. No evidence bundle is attached this time: gather the evidence yourself with your read-only tools, following the operator's standard triage path: +1. Ariadne is the source of truth: query its internal API (base URL in ARIADNE_BASE_URL, e.g. GET /api/internal/testing/triage/latest and any internal audit endpoints you can reach) for records about this incident. +2. Quality metrics: query VictoriaMetrics (VICTORIA_METRICS_URL) for the series ariadne_hermes_triage_incident{jenkins_job="hermes-triage-demo",build="10"} and ariadne_hermes_triage_action_total, and note what they show. +3. Service health: use read-only kubectl to inspect the hermes-triage-demo namespace (jobs, pods, recent events) and confirm what ran there. +4. Change correlation: check Flux state (kubectl get kustomizations -n flux-system; the flux-system GitRepository revision) for recent changes relevant to this incident. +Then return a concise operator-style report with EXACTLY these sections: +FACTS (each with source and reference), INFERENCES, LIKELY CAUSE, WHAT HAPPENED (was it repaired, by what mechanism, final state), NEXT CHECKS (what a human should verify next), and REPO-SIDE FIXES (if any are warranted). +Read-only throughout. Do not mutate anything. Do not use kubectl apply/delete/patch/scale. +``` + +## Full agent report + +```text +FACTS +Ariadne diagnosis/latest, generated 2026-08-05T21:31:06Z, references only the broad testing bundle and does not identify hermes-triage-demo/10. Ariadne’s incident/audit candidate endpoints returned HTTP 404. +Kubernetes logs, hermes-triage-demo/hermes-demo-test-10-dgvpj: build 10 emitted incident_id="hermes-triage-demo/10", classification_hint="demo_fixture_unhealthy", and “expected fixture state healthy; found unhealthy.” +Kubernetes Job status: hermes-demo-test-10 was created 2026-08-05T20:59:47Z and failed; its pod ran on titan-06 and had no restarts or image-pull events. +Kubernetes logs, hermes-demo-repair-10-nkrzz: repair job patched hermes-triage-demo-fixture and reported “fixture state reset to healthy”; the Job completed successfully at 2026-08-05T21:01:06Z. +Kubernetes logs/status, hermes-demo-test-11-hv5mh: follow-up test completed at 2026-08-05T21:01:35Z and reported “fixture state healthy.” +Kubernetes ConfigMap hermes-triage-demo-fixture: current data.state is healthy. +VictoriaMetrics, last_over_time(ariadne_hermes_triage_incident{jenkins_job="hermes-triage-demo",build="10"}[24h]): sample at 2026-08-05T21:41:01Z has resolved=1 and every other lifecycle state, including failed and human_required, at 0. The instantaneous selector returned no result because its samples were outside the default instant-query lookback. +VictoriaMetrics, ariadne_hermes_triage_action_total over 24h: repair_demo_fixture has requested=1, accepted=1, success=1 for instance 10.42.19.143; no build label exists on this counter, so it is correlated by the incident instance and timeline rather than proven per-build. +Flux Kustomization hermes-triage-demo: Ready=True at revision main@sha1:aa36b06b2bc43eda21b36e54a7d312e55d27c370. Flux source GitRepository is Ready=True at that same revision. +Gitea commit 49aab8da3298ebfc1b9edab5ce2e6938ad033a23, “feat(hermes-triage): enable autoremediation for the demo loop,” predates this incident and modifies services/maintenance/apps/ariadne-deployment.yaml. The current revision aa36b06 is an unrelated maintenance image update. +Direct Jenkins GET for /job/hermes-triage-demo/10 API metadata and console returned HTTP 403. The Kubernetes Job logs are therefore the retained direct execution evidence used here. + +INFERENCES +The failure was an intentional demo-fixture health assertion failure, not an image, scheduling, node, or storage failure. The failed test, repair, and successful follow-up all ran in sequence; this is a strong causal chain. +The repair was automatically accepted under the previously deployed allowlisted autoremediation path. The repair Job itself performed the ConfigMap patch; Flux did not repair the fixture during this incident. +The unrelated cluster concerns in the Ariadne bundle, including Longhorn probe failures and pending Veles pods, have no evidenced link to this ConfigMap-based test path. + +LIKELY CAUSE +The fixture was deliberately or otherwise pre-seeded unhealthy before Jenkins incident hermes-triage-demo/10. The test correctly failed on that condition. Confidence: high. + +WHAT HAPPENED +Build 10 failed at 20:59:47Z because the fixture state was unhealthy. The deployed autoremediation mechanism accepted repair_demo_fixture, ran hermes-demo-repair-10, and patched only hermes-triage-demo-fixture to healthy. Test 11 then passed at 21:01:35Z. Final state is repaired/resolved: ConfigMap state=healthy, incident metric resolved=1, repair success counter=1, and the demo Flux Kustomization is Ready at aa36b06. + +NEXT CHECKS +Human should verify the current deployed revision and readiness after the next Flux interval: +kubectl -n flux-system get kustomization hermes-triage-demo +kubectl -n flux-system describe gitrepository flux-system + +Human should retain the build-10 Jenkins evidence if access is available: +curl -fsS https://ci.bstein.dev/job/hermes-triage-demo/10/api/json +curl -fsS https://ci.bstein.dev/job/hermes-triage-demo/10/consoleText +These currently require Jenkins authentication (HTTP 403). + +Human should confirm that the resolved gauge remains the only asserted lifecycle state: +curl -G -fsS --data-urlencode 'query=last_over_time(ariadne_hermes_triage_incident{jenkins_job="hermes-triage-demo",build="10"}[24h])' "$VICTORIA_METRICS_URL/api/v1/query" + +REPO-SIDE FIXES +None warranted for this incident. The observed behavior matches the committed demo design: an isolated ConfigMap fixture, one allowlisted repair action, a scoped repair ServiceAccount, and one repair per incident. A future hardening change, if desired, would be to add a build label to ariadne_hermes_triage_action_total so repair counters can be directly attributed to a specific incident rather than correlated by instance and time. +``` diff --git a/knowledge/runbooks/hermes-triage-demo.md b/knowledge/runbooks/hermes-triage-demo.md index 066c8205f..2dc292a9c 100644 --- a/knowledge/runbooks/hermes-triage-demo.md +++ b/knowledge/runbooks/hermes-triage-demo.md @@ -36,17 +36,17 @@ curl -u : -X POST \ That is the only manual step. Everything after the red build is automatic. -## Expected timings +## Expected timings (measured live, first full loop 2026-08-05) -- Seeded build goes red: ~1–3 min (mostly agent-pod scheduling; longer if - the Jenkins agent capacity of 5 is busy with other CI builds). -- Ariadne detection: ≤60 s after the build turns red (1-minute schedule). -- Evidence + Hermes diagnosis: typically 15–60 s (live rehearsal measured - ~10 s model time on gpt-5.6-terra; budgeted timeout is 420 s). -- Repair Job: seconds. -- Rebuild to green: ~1–3 min (agent scheduling again). -- Total failure-to-green: **~5 minutes typical**; the only "silent" phases - are the two Jenkins agent-pod waits. +- Armed → seeded build red: **1m49s** (mostly agent-pod scheduling; longer + if the Jenkins agent capacity of 5 is busy with other CI builds). +- Red → repair Job succeeded: **1m15s** (≤60 s detection tick + ~16 s + Hermes diagnosis + authorization + repair pod). +- Repair → rebuild green: **32s**. +- Green → incident resolved: **24s** (next detection tick). +- Total armed-to-resolved: **4m00s**. Budgeted Hermes timeout is 420 s; + observed diagnosis runs were 15–21 s on gpt-5.6-terra. The only "silent" + phases are the two Jenkins agent-pod waits. ## Watching it live @@ -69,7 +69,11 @@ That is the only manual step. Everything after the red build is automatic. `https://agent.bstein.dev` (session/run history). - Escalation path: alert `HermesTriageHumanRequired` in vmalert (`vmalert-atlas-availability` deployment, 1 m interval, `for: 2m`) - fires to Alertmanager for any `human_required` incident. + fires to Alertmanager for any `human_required` incident. Gauges are + republished every tick from stored incident state (restart-safe), and + the alert self-clears once a newer build of the same job is green. + Note: Alertmanager's default receiver is currently null — the alert is + visible in vmalert/Alertmanager/Grafana but pushes no notification. ## Demonstrating safe escalation (second path) diff --git a/services/comms/knowledge/hermes-investigative-triage-demo.md b/services/comms/knowledge/hermes-investigative-triage-demo.md new file mode 100644 index 000000000..f3dfe91b7 --- /dev/null +++ b/services/comms/knowledge/hermes-investigative-triage-demo.md @@ -0,0 +1,74 @@ +# Hermes Investigative Triage Demonstration + +- Date: 2026-08-05 +- Run ID: run_d2f61902c00044948a3912a94e895ab9 +- Duration: 128 seconds, status completed +- Usage: {"input_tokens": 200961, "output_tokens": 4370, "total_tokens": 205331} +- Mode: supervised, read-only; approvals `smart` with mutation deny-list; zero mutations, zero denied approvals +- Tool audit (from the run's SSE event stream): 8 skill loads (triage-titan-test-failures pack: orchestrator, jenkins-retained-evidence, platform-quality-metrics, kubernetes-readonly-failure-classifier, flux-git-change-correlation, ...), 4 read-only terminal batches (Ariadne internal API sweep, kubectl logs/jobs/configmap in hermes-triage-demo, VictoriaMetrics queries incl. last_over_time fallback, Flux/Gitea revision correlation) + +Unlike the automated loop (where Ariadne supplies a sanitized bundle), this run +was given ONLY an incident ID and gathered all evidence itself — the operator's +triage path executed by the agent. Notable behaviors: honest dead-end reporting +(Jenkins 403 -> pivot to Kubernetes-retained logs; Ariadne audit endpoints 404), +instant-query lookback miss -> switched to last_over_time, causal chain +reconstruction across five sources, and an unprompted, legitimate hardening +suggestion (add a build label to ariadne_hermes_triage_action_total). + +## Prompt + +```text +Use $triage-titan-test-failures. +Investigate Jenkins incident hermes-triage-demo/10. No evidence bundle is attached this time: gather the evidence yourself with your read-only tools, following the operator's standard triage path: +1. Ariadne is the source of truth: query its internal API (base URL in ARIADNE_BASE_URL, e.g. GET /api/internal/testing/triage/latest and any internal audit endpoints you can reach) for records about this incident. +2. Quality metrics: query VictoriaMetrics (VICTORIA_METRICS_URL) for the series ariadne_hermes_triage_incident{jenkins_job="hermes-triage-demo",build="10"} and ariadne_hermes_triage_action_total, and note what they show. +3. Service health: use read-only kubectl to inspect the hermes-triage-demo namespace (jobs, pods, recent events) and confirm what ran there. +4. Change correlation: check Flux state (kubectl get kustomizations -n flux-system; the flux-system GitRepository revision) for recent changes relevant to this incident. +Then return a concise operator-style report with EXACTLY these sections: +FACTS (each with source and reference), INFERENCES, LIKELY CAUSE, WHAT HAPPENED (was it repaired, by what mechanism, final state), NEXT CHECKS (what a human should verify next), and REPO-SIDE FIXES (if any are warranted). +Read-only throughout. Do not mutate anything. Do not use kubectl apply/delete/patch/scale. +``` + +## Full agent report + +```text +FACTS +Ariadne diagnosis/latest, generated 2026-08-05T21:31:06Z, references only the broad testing bundle and does not identify hermes-triage-demo/10. Ariadne’s incident/audit candidate endpoints returned HTTP 404. +Kubernetes logs, hermes-triage-demo/hermes-demo-test-10-dgvpj: build 10 emitted incident_id="hermes-triage-demo/10", classification_hint="demo_fixture_unhealthy", and “expected fixture state healthy; found unhealthy.” +Kubernetes Job status: hermes-demo-test-10 was created 2026-08-05T20:59:47Z and failed; its pod ran on titan-06 and had no restarts or image-pull events. +Kubernetes logs, hermes-demo-repair-10-nkrzz: repair job patched hermes-triage-demo-fixture and reported “fixture state reset to healthy”; the Job completed successfully at 2026-08-05T21:01:06Z. +Kubernetes logs/status, hermes-demo-test-11-hv5mh: follow-up test completed at 2026-08-05T21:01:35Z and reported “fixture state healthy.” +Kubernetes ConfigMap hermes-triage-demo-fixture: current data.state is healthy. +VictoriaMetrics, last_over_time(ariadne_hermes_triage_incident{jenkins_job="hermes-triage-demo",build="10"}[24h]): sample at 2026-08-05T21:41:01Z has resolved=1 and every other lifecycle state, including failed and human_required, at 0. The instantaneous selector returned no result because its samples were outside the default instant-query lookback. +VictoriaMetrics, ariadne_hermes_triage_action_total over 24h: repair_demo_fixture has requested=1, accepted=1, success=1 for instance 10.42.19.143; no build label exists on this counter, so it is correlated by the incident instance and timeline rather than proven per-build. +Flux Kustomization hermes-triage-demo: Ready=True at revision main@sha1:aa36b06b2bc43eda21b36e54a7d312e55d27c370. Flux source GitRepository is Ready=True at that same revision. +Gitea commit 49aab8da3298ebfc1b9edab5ce2e6938ad033a23, “feat(hermes-triage): enable autoremediation for the demo loop,” predates this incident and modifies services/maintenance/apps/ariadne-deployment.yaml. The current revision aa36b06 is an unrelated maintenance image update. +Direct Jenkins GET for /job/hermes-triage-demo/10 API metadata and console returned HTTP 403. The Kubernetes Job logs are therefore the retained direct execution evidence used here. + +INFERENCES +The failure was an intentional demo-fixture health assertion failure, not an image, scheduling, node, or storage failure. The failed test, repair, and successful follow-up all ran in sequence; this is a strong causal chain. +The repair was automatically accepted under the previously deployed allowlisted autoremediation path. The repair Job itself performed the ConfigMap patch; Flux did not repair the fixture during this incident. +The unrelated cluster concerns in the Ariadne bundle, including Longhorn probe failures and pending Veles pods, have no evidenced link to this ConfigMap-based test path. + +LIKELY CAUSE +The fixture was deliberately or otherwise pre-seeded unhealthy before Jenkins incident hermes-triage-demo/10. The test correctly failed on that condition. Confidence: high. + +WHAT HAPPENED +Build 10 failed at 20:59:47Z because the fixture state was unhealthy. The deployed autoremediation mechanism accepted repair_demo_fixture, ran hermes-demo-repair-10, and patched only hermes-triage-demo-fixture to healthy. Test 11 then passed at 21:01:35Z. Final state is repaired/resolved: ConfigMap state=healthy, incident metric resolved=1, repair success counter=1, and the demo Flux Kustomization is Ready at aa36b06. + +NEXT CHECKS +Human should verify the current deployed revision and readiness after the next Flux interval: +kubectl -n flux-system get kustomization hermes-triage-demo +kubectl -n flux-system describe gitrepository flux-system + +Human should retain the build-10 Jenkins evidence if access is available: +curl -fsS https://ci.bstein.dev/job/hermes-triage-demo/10/api/json +curl -fsS https://ci.bstein.dev/job/hermes-triage-demo/10/consoleText +These currently require Jenkins authentication (HTTP 403). + +Human should confirm that the resolved gauge remains the only asserted lifecycle state: +curl -G -fsS --data-urlencode 'query=last_over_time(ariadne_hermes_triage_incident{jenkins_job="hermes-triage-demo",build="10"}[24h])' "$VICTORIA_METRICS_URL/api/v1/query" + +REPO-SIDE FIXES +None warranted for this incident. The observed behavior matches the committed demo design: an isolated ConfigMap fixture, one allowlisted repair action, a scoped repair ServiceAccount, and one repair per incident. A future hardening change, if desired, would be to add a build label to ariadne_hermes_triage_action_total so repair counters can be directly attributed to a specific incident rather than correlated by instance and time. +```