Second entry in the action registry, proving it is a real extension point. No cluster mutation: the action is one Jenkins rebuild. - hermes_infra_signals: reviewable marker set across DNS/connectivity, image pull, upstream 5xx and agent-channel loss; Ariadne independently confirms a marker in the evidence before any retry, and records which marker justified it. "no space left on device" is deliberately excluded because a retry lands on the same full volume. - decision: classification -> action registry (action_classifications), falling back to the previous single-classification behavior - repair: retry_build posts to /build for unparameterized real jobs and buildWithParameters for the fixture demo job - orchestrator: retry path records requested/accepted/executed and moves the incident to awaiting_rebuild so the existing success path resolves it; one action per incident still enforced, so a retry cannot loop - events layer split out of the orchestrator to stay under the LOC cap Motivated by real failures tonight: pip DNS resolution and a Gitea 443 connect timeout, plus live incident metis/271 (SCM checkout timeout). 30 new tests; 368 pass in the hermes suite. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
ariadne
Ariadne is the Atlas admin and account automation service.
It sits behind the portal and handles the jobs that are annoying or risky to do by hand: approving access, syncing account state, rotating service passwords, cleaning stale Kubernetes work, checking platform health, and keeping a few service integrations lined up.
How it works
Ariadne is a FastAPI service with a small scheduler. It talks to Keycloak, Vault, Mailu, Nextcloud, Wger, Firefly, Jenkins, Metis, Kubernetes, and a few Atlas-specific services through focused adapters under ariadne/services/.
The API is split between admin routes, account self-service routes, internal event hooks, and Prometheus metrics. Background jobs store run history in the Ariadne database so failures can be inspected later instead of vanishing into logs.
The following are notes for future Brad.
Bring-up dependencies
Ariadne needs:
- Kubernetes API, service DNS, and Ariadne's service account/RBAC
- the Ariadne database, plus the portal database if portal/account sync is enabled
- Vault or the Kubernetes secrets that Vault normally feeds it
- Keycloak/OIDC, because auth and profile sync assume it exists
- ingress/proxy plumbing if humans are going to use it through the portal
- the services for whatever jobs are enabled: Mailu, Nextcloud, Vaultwarden, Wger, Firefly, Jenkins, Metis, OpenSearch, and the comms/game-mode pieces
It can start before every integration is perfect, but the matching scheduled jobs will fail or no-op until their service is actually alive. In a total bring-up, wait for storage, Flux, Postgres, Vault, Keycloak, and ingress first. Afterwards Ariadne becomes useful glue.
Useful routes:
GET /healthGET /metricsGET /api/admin/cluster/statePOST /api/admin/access/requests/{username}/approvePOST /api/account/mailu/rotatePOST /api/account/wger/resetPOST /api/account/firefly/resetPOST /events
Development
python -m pytest
ruff check .
Most runtime behavior is configured through environment variables in ariadne/settings.py. Service-specific logic is in the small adapter modules; ariadne/app.py is focused on request flow and task orchestration.