codex ca2324ab95
All checks were successful
Tests / Declarative: Post Actions passed: 1251
feat(hermes): reclaim exhausted workspace storage, and stop hung builds
Two real actions, both for failures that interrupted this project while it was
being built, and both backed by code Ariadne already had.

reclaim_workspace_storage closes a gap left open deliberately: 'no space left
on device' was excluded from the transient retry because a rebuild lands on
the same full volume and either fails identically or hides a capacity problem.
Reclaiming first makes the retry meaningful. The reclaim is the existing
scheduled cleanup, with its own deletion budget, so no new capability is
granted - it is only reachable from triage now. Its signature requires both a
storage marker and a workspace hint, because a full disk elsewhere in the
cluster is a different failure that reclaiming Jenkins workspaces would not
address.

Hung-build detection previously filed an issue and left the build running,
holding one of five Jenkins agent slots and starving every other job - the
actual harm. Ariadne now stops it as well, which is reversible: the job can
simply be built again.

The action executors move to their own module. They are the only code in
triage that changes anything outside Ariadne and should be reviewable as one
unit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 20:55:35 -03:00
2026-06-19 21:27:06 +00:00

ariadne

Ariadne is the Atlas admin and account automation service.

It sits behind the portal and handles the jobs that are annoying or risky to do by hand: approving access, syncing account state, rotating service passwords, cleaning stale Kubernetes work, checking platform health, and keeping a few service integrations lined up.

How it works

Ariadne is a FastAPI service with a small scheduler. It talks to Keycloak, Vault, Mailu, Nextcloud, Wger, Firefly, Jenkins, Metis, Kubernetes, and a few Atlas-specific services through focused adapters under ariadne/services/.

The API is split between admin routes, account self-service routes, internal event hooks, and Prometheus metrics. Background jobs store run history in the Ariadne database so failures can be inspected later instead of vanishing into logs.

The following are notes for future Brad.

Bring-up dependencies

Ariadne needs:

  • Kubernetes API, service DNS, and Ariadne's service account/RBAC
  • the Ariadne database, plus the portal database if portal/account sync is enabled
  • Vault or the Kubernetes secrets that Vault normally feeds it
  • Keycloak/OIDC, because auth and profile sync assume it exists
  • ingress/proxy plumbing if humans are going to use it through the portal
  • the services for whatever jobs are enabled: Mailu, Nextcloud, Vaultwarden, Wger, Firefly, Jenkins, Metis, OpenSearch, and the comms/game-mode pieces

It can start before every integration is perfect, but the matching scheduled jobs will fail or no-op until their service is actually alive. In a total bring-up, wait for storage, Flux, Postgres, Vault, Keycloak, and ingress first. Afterwards Ariadne becomes useful glue.

Useful routes:

  • GET /health
  • GET /metrics
  • GET /api/admin/cluster/state
  • POST /api/admin/access/requests/{username}/approve
  • POST /api/account/mailu/rotate
  • POST /api/account/wger/reset
  • POST /api/account/firefly/reset
  • POST /events

Development

python -m pytest
ruff check .

Most runtime behavior is configured through environment variables in ariadne/settings.py. Service-specific logic is in the small adapter modules; ariadne/app.py is focused on request flow and task orchestration.

Description
atlas cluster job management tool with reporting for prometheus
Readme 4.4 MiB
Languages
Python 100%