4825 Commits

Author SHA1 Message Date
jenkins
aab389499c fix(ariadne): let the pod finish booting before liveness judges it
The earlier probe fix addressed slow /health responses under load, but the
restarts continued with a different signature: connection refused rather than
timeout, meaning the app was not listening yet. Ariadne runs migrations and
builds its cron schedule before binding, which can outlast what liveness
allows from initialDelaySeconds, so the kubelet kept restarting a pod that
was merely still starting.

Add a startupProbe granting up to five minutes to come up, after which
liveness takes over unchanged. This is the case startupProbe exists for.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 06:16:09 -03:00
flux-bot
c23234fffd chore(maintenance): automated image update 2026-08-06 09:11:55 +00:00
jenkins
0bd47ef6a4 feat(jenkins): install junit and pipeline-stage-view
Without the junit plugin the pipeline's junit step throws NoSuchMethodError,
jenkins.failed_tests is always empty, and auto-triage has only raw console
text to reason from. That was the root of two separate diagnosis failures:
the enforced failure being crowded out of the evidence budget, and the
patcher being unable to locate the defective source file.

Pinned to junit 1369.v15da_00283f06, the newest release that runs on core
2.528.3 - the current 1418 requires 2.533. scm-api moves 724 -> 728 because
the workflow-cps these pull in requires it; verified the whole 30-plugin
dependency closure needs no core newer than 2.528.3, and rehearsed the exact
jenkins-plugin-cli install in a throwaway pod before committing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 06:08:17 -03:00
flux-bot
c7f670a53c chore(maintenance): automated image update 2026-08-06 08:55:38 +00:00
flux-bot
1611844e69 chore(maintenance): automated image update 2026-08-06 08:22:34 +00:00
flux-bot
7248d54b34 chore(cassandra): automated image update 2026-08-06 08:21:47 +00:00
jenkins
699d1a5689 fix(ariadne): stop the kubelet killing a healthy triage pod
The liveness probe used the default timeoutSeconds of 1. The auto-triage tick
runs every minute and spends most of it waiting on Jenkins, OpenSearch, Gitea
and Hermes, so against a 500m CPU limit /health occasionally answers in over a
second. Three of those and the container is killed, dropping triage ticks for
the length of a restart. Observed 11 times in 139 minutes, with the pod
sitting 1/2 Ready and restarting repeatedly.

Give both probes a 5s timeout and let liveness tolerate five failures, so a
busy tick is no longer mistaken for a hung process.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 05:20:35 -03:00
flux-bot
ff8da346d2 chore(cassandra): automated image update 2026-08-06 08:17:41 +00:00
jenkins
20c5b326c5 docs(runbook): Claude primary, hung-build escalation, narrowed alerting
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 05:15:42 -03:00
flux-bot
4d4d4c481c chore(cassandra): automated image update 2026-08-06 08:15:03 +00:00
flux-bot
8de1a87ffe chore(cassandra): automated image update 2026-08-06 08:14:29 +00:00
jenkins
c2c8a81c7f feat(hermes): switch the primary model to Claude Opus 5
The Codex weekly limit is close, so make anthropic/claude-opus-5 primary and
demote openai-codex to first fallback with the local gpt-oss:20b behind it.

The credential is a Claude subscription OAuth token, not an API key. The
anthropic provider resolves ANTHROPIC_API_KEY, then ANTHROPIC_TOKEN, then
CLAUDE_CODE_OAUTH_TOKEN, so the OAuth token must arrive under the last name
to be treated correctly. Marked optional so Hermes still starts and falls
back if the Secret is absent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 05:00:09 -03:00
jenkins
4d3982ac96 docs: add cluster architecture diagrams 2026-08-06 04:42:12 -03:00
jenkins
901c106a92 docs: record the first autonomous source repair on a real service repo
ariadne PR #3 from incident ariadne/409, chosen path ariadne/utils/errors.py -
a file that appears nowhere in the build console and was reached only by
following the failing test's imports.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 04:28:54 -03:00
flux-bot
d52498273c chore(cassandra): automated image update 2026-08-06 06:50:03 +00:00
flux-bot
f19272f7a2 chore(cassandra): automated image update 2026-08-06 06:46:05 +00:00
flux-bot
f6d5a2a677 chore(cassandra): automated image update 2026-08-06 06:42:06 +00:00
flux-bot
6f4f5f8afa chore(cassandra): automated image update 2026-08-06 06:40:06 +00:00
flux-bot
426b6db1e7 chore(maintenance): automated image update 2026-08-06 04:47:42 +00:00
jenkins
a1b79e89b7 docs: state what each demo proves and record the two selection defects
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 01:39:55 -03:00
flux-bot
384de8652a chore(maintenance): automated image update 2026-08-06 04:33:02 +00:00
flux-bot
f44c8d26a4 chore(cassandra): automated image update 2026-08-06 04:27:41 +00:00
flux-bot
7a5db4cf5a chore(cassandra): automated image update 2026-08-06 04:24:11 +00:00
flux-bot
d0f9773368 chore(cassandra): automated image update 2026-08-06 04:19:49 +00:00
flux-bot
8c4f90a39c chore(cassandra): automated image update 2026-08-06 04:19:31 +00:00
flux-bot
62034931e2 chore(maintenance): automated image update 2026-08-06 04:10:10 +00:00
jenkins
2b0fdd1a03 feat(monitoring): make Hermes triage email mean something
Every human_required escalation already files an issue in the failing
service's own repository, and mailing on each one made the inbox the loudest
and least informative output of the system. Replace the blanket alert with
two narrow ones: a repair that ran and failed, which is the only case where
the automation acted and left things no better, and an escalation still
untouched after six hours.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 00:26:34 -03:00
flux-bot
2a75b6174a chore(maintenance): automated image update 2026-08-06 03:11:10 +00:00
flux-bot
a08815fb3f chore(maintenance): automated image update 2026-08-06 03:09:11 +00:00
flux-bot
2a297c962e chore(maintenance): automated image update 2026-08-06 03:07:10 +00:00
flux-bot
4ff140ac45 chore(maintenance): automated image update 2026-08-06 03:02:09 +00:00
flux-bot
ecd4d8c48d chore(bstein-dev-home): automated image update 2026-08-06 02:56:09 +00:00
flux-bot
c06cc8cbe0 chore(bstein-dev-home): automated image update 2026-08-06 02:52:08 +00:00
jenkins
88f05cfe37 docs: record the in-process repair timings and the classification-bleed fix
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 23:36:58 -03:00
jenkins
3bd8d249cc docs(runbook): record the faster in-process repair timings
Red-build to green-rebuild is 1m04s on Ariadne 0.1.0-402, down from 4m00s
armed-to-resolved when the repair spawned its own Kubernetes Job. Reframe the
timings around the red build, since the arming leg depends on the agent pool.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 23:34:09 -03:00
flux-bot
2a3b7c3727 chore(maintenance): automated image update 2026-08-06 02:25:00 +00:00
jenkins
2ce849c6d7 docs(runbook): warn that a full agent pool stalls the demo
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 23:20:30 -03:00
jenkins
ea206e98fc feat(demo): preflight the agent-pool cap and open code-demo PRs
Two conditions silently break a rehearsal. A saturated Kubernetes agent pool
leaves the demo build queued reporting that all nodes are offline, and an open
hermes-repair PR makes the duplicate guard refuse a new proposal. Report both
so the operator sees them before starting rather than mid-demo.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 23:19:37 -03:00
jenkins
97dfb1a844 feat(ariadne): allowlist the transient-infra retry action
The classification->action registry already mapped transient_infra_failure to
retry_transient_infra, but the action allowlist held only repair_demo_fixture,
so that route always died at the action_not_allowlisted gate. Add the action
and state the registry explicitly rather than relying on the code default.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 23:17:32 -03:00
flux-bot
56f5e5a123 chore(maintenance): automated image update 2026-08-06 02:10:50 +00:00
jenkins
af4d9db64c fix(vault-injector): run two replicas so restarts cannot skip injection
The webhook is failurePolicy: Ignore, so with one replica any pod created
during an injector restart is admitted without its Vault agent sidecar and
then crash-loops forever on a missing /vault/secrets file, with nothing to
indicate injection was skipped. Hit twice while rolling ariadne.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 22:56:49 -03:00
flux-bot
5df98f65f7 chore(maintenance): automated image update 2026-08-06 01:49:38 +00:00
jenkins
33085879b7 chore(hermes-triage-demo): retire repair Job service account and role
The fixture repair now runs in-process inside Ariadne, so the
hermes-demo-repair service account and its Role/RoleBinding have no
remaining user.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 22:31:11 -03:00
jenkins
f4d8b2c01d feat(hermes-triage-demo): least-privilege ConfigMap patch for in-process repair
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 22:25:52 -03:00
jenkins
1894078b62 fix(scripts): dashboard renderer wrote outside the repo after the layout move
ROOT still used parents[1], which resolved to scripts/ once the renderer
moved into scripts/render/. Every --build run wrote a phantom
scripts/services/monitoring tree and silently left the real dashboards
untouched. Points at the repo root again and removes the stray tree.

Also adds Hermes triage panels to the Atlas Testing dashboard: open
escalations awaiting a human, automated actions succeeded, Hermes
diagnosis latency, actions by result, and incident state by job.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 21:41:02 -03:00
flux-bot
b929e67692 chore(maintenance): automated image update 2026-08-06 00:39:46 +00:00
jenkins
84ff647898 fix(monitoring): Alertmanager must HELO with a FQDN for Mailu
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 21:33:11 -03:00
jenkins
664e436575 feat(hermes-triage): real-repo patch proposals + Alertmanager email escalation
- Ariadne: per-repo code config for metis, lesavka, soteria,
  bstein-dev-home and ariadne, each with its own base branch, source path
  prefixes and file suffixes so a proposal can only touch that repo's
  source tree
- Alertmanager: the only receiver was an empty "default", so every alert
  fired into a void. HermesTriageHumanRequired now routes to an email
  receiver via Mailu's in-cluster local-domain relay, with resolved
  notices; scoped to service=hermes-triage so nothing else mails yet.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 21:30:34 -03:00
flux-bot
72fc09783b chore(maintenance): automated image update 2026-08-06 00:15:26 +00:00
jenkins
46bd4daf88 docs(hermes-triage): delivery record + enable Gitea issue filing config
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 21:01:50 -03:00