atlas-iac/docs/cluster-audit-20261002/RECOVERY_TOOL_PLAN.md

13 KiB

Make recovery dependable: existing-tool implementation plan

Proposal for review, 2026-10-02. No implementation, recovery actions or failure drills have been authorized or performed by this document.

Scope update: the user's latest direction is standard Kubernetes/component configuration first. The matrix and coordination ideas below are candidate capabilities, not a required integration project. Apply the decision gate in INTEGRATED_PLAN.md: use native behavior, correct configuration and remove unnecessary overlap first; change a homegrown program only for a demonstrated remaining gap. The 4-8 day tool-work estimate is conditional and is not an approved baseline work package.

The existing homegrown tools are the basis of this work. Each must have a clear responsibility, usable recovery prerequisites and evidence that its recovery succeeds. The objective is to repair gaps in the existing system and its coordination, without introducing another overlapping general-purpose recovery controller.

Responsibility and acceptance matrix

Tool Intended responsibility in this plan Evidence already inspected Proposed work Required proof after approval
Soteria Backup coverage, freshness, recovery-point selection and isolated volume-restore workflow Longhorn driver settings, policy schedule/exclusions, error and bucket telemetry; existing PVC restore-drill notes Diagnose backend errors and mounted-volume skips; calculate eligible coverage and achievable schedule; report success only for completed recoverable backups; track approved local-only protection separately; integrate external database backup status A fresh backup and a restore into a separate target, with application/data checks. A missed/failed backup becomes visibly unhealthy rather than a successful scan being mistaken for successful protection
Ananke Power-event coordination and ordered infrastructure shutdown/startup/recovery Actual coordinator on titan-db and peer on titan-24; UPS configuration, startup checks, node inventory and selected datastore/recovery code Align inventory with actual node roles; verify external-PostgreSQL recovery handling; base shutdown timing on measured behavior; coordinate maintenance and bounded node repair; require service/storage validation before declaring recovery complete Simulated dependency/peer failures stop safely; later controlled recovery follows the real dependency order, respects deadlines and has one acting owner
Metis Rebuild or replace a node whose OS/runtime medium is no longer trustworthy Node/image/flash-host configuration, runner/deployment/RBAC and mounted inventory/data dependencies Verify the effective inventory path, supported hardware images and flash-host identities; protect non-target disks and existing replicas; define backed-up, operator-authorized rebuild and post-join validation; keep a usable independent recovery package A disposable test device or spare node is rebuilt and rejoins correctly; target identity and image verification prevent an unintended disk write; Longhorn data disks are preserved
Ariadne Bounded CI/application incident handling within its existing remit Deployment configuration and existing Hermes triage/remediation restrictions Recognize shared node/storage/DB incidents as dependency failures; defer repeated rebuilds while that dependency is broken; retain bounded retries and action allowlists; publish the owning infrastructure incident A simulated shared failure produces one escalated dependency incident rather than repeated build/restart churn; a recoverable CI fault still follows its existing bounded path
Hermes recovery integration Diagnosis and existing scoped repair assistance Existing integration/action restrictions, not an exhaustive source audit Preserve current permissions and action limits; attach safe diagnostic evidence and escalation reasons; do not make model output sufficient authority to restore databases, flash disks or change cluster ownership Invalid/unapproved actions remain rejected; unavailable model assistance does not prevent deterministic recovery or reading the operator runbook
Node helpers Narrow host-specific detection or repair Titan-22 link helper and several restart/cleanup DaemonSets Keep useful bounded repairs; give each an explicit incident owner, cooldown and stop condition; retire unconditional or overlapping restart/cleanup behavior after replacement is verified The same outage cannot trigger concurrent restarts from multiple owners; repeated failure stops with an actionable condition
Flux and Helm Restore approved steady-state configuration after infrastructure is usable Rendered/live differences, creation-only Kustomization annotations, absent Helm drift detection Resolve desired-state differences before changing ownership; represent temporary recovery exceptions explicitly and expire them; validate runtime health after convergence Reconciliation produces intended live state without undoing a legitimate recovery operation or silently retaining an old workaround

These are proposed responsibilities and acceptance tests. The matrix does not claim the current programs already expose all the required APIs, locks, integrations or restore guarantees. Soteria's existing restore checklist is at services/maintenance/NOTES.md; it should be validated and extended rather than duplicated.

1. Establish recovery prerequisites through the existing toolchain

External Kubernetes database

The current datastore is PostgreSQL on titan-db. It needs a database-native backup and its associated K3s recovery material; an etcd-only procedure does not cover it.

Proposed ownership:

  • A versioned host-side PostgreSQL backup/verification job performs scheduled backups independently of Kubernetes. Maintain it with the out-of-cluster host tooling used by Ananke. This is proposed integration, not a verified existing Ananke backup feature.
  • Soteria consumes safe completion/freshness/restore-verification metadata so this database appears in the same recovery inventory as application volumes. Do not require an in-cluster Soteria process to be running in order to recover the Kubernetes datastore.
  • Ananke checks datastore availability and reports the specific external-database blocker during startup. An unavailable API alone must not trigger a destructive database restore.
  • The approved recovery package includes credentials/keys needed to recover without depending on a functioning in-cluster Vault or registry. Keep protected copies outside the failure domain, with tightly controlled access; never put secret values in the report or Git.

First acceptance gate: restore a new database backup into an isolated target, verify the required database/recovery material, record recovery duration and prove the original is untouched. Full replacement of the live datastore remains a separately reviewed cutover.

Application volumes and local job state

  • Use Soteria's actual configured Longhorn path for eligible volumes, after diagnosing the observed backend errors and skip decisions.
  • Keep database-consistent backups for transactional applications in addition to whatever volume snapshots are appropriate.
  • Establish an explicit approved-local backup/recovery method for local-path suite job state. Do not enable B2 export of local-only records as a side effect of closing a backup gap.
  • Record backup age, target, completion result, last restore verification and recovery owner for every service. An intentionally reconstructible cache is an explicit exclusion, not a missing entry.

First acceptance gate: successful isolated restores of a representative Longhorn volume, an application database and synthetic local job/checkpoint state. Complete per-service restore coverage follows in the continuity wave; three sample restores do not prove all services recover.

2. Make the tools cooperate during one incident

If a demonstrated failure still requires multiple custom tools after simpler configuration/ownership fixes, consider the following operational requirements. They do not mandate a shared protocol or new service:

  • Each incident identifies the affected node/service, current owner, safe evidence, attempted actions, remaining retries, cooldown and blocking dependency.
  • Only one owner performs a disruptive action on the same target at a time. Reuse existing locks/coordination where suitable, and test gaps before designing extensions.
  • Full-cluster recovery coordination must still work when the Kubernetes API is unavailable; a Kubernetes Lease cannot be the only protection for out-of-cluster actions.
  • Metis rebuild activity excludes Ananke's automatic return-to-service for that node. Ananke recovery excludes discretionary descheduling and conflicting node restarts. Ariadne defers application retries while an infrastructure dependency is unresolved.
  • Action eligibility is distinct from detection. Suspecting a dead node is not sufficient authority to detach a possibly live writer or overwrite a disk.
  • Completion requires recovery checks appropriate to the failure: API/SSH/runtime reachability, a test pod, storage mount/read-write checks on a disposable target, and service-level readiness. A node becoming Ready is only one check.
  • If automated repair cannot proceed safely or exhausts retries, it leaves a clear blocked incident and the next operator action. It must not silently start a new incident to reset its retry allowance.

This is a design requirement for the existing programs, not an assertion that a shared lock or incident protocol is already implemented.

3. Example: Titan-22 loses connectivity again

  1. Ananke and monitoring identify host/link loss. The narrow link helper may perform its already-approved bounded repair; other restart owners defer while that action is active.
  2. Kubernetes handles normal workload reconciliation where compatible capacity and storage allow it. Ananke reports which required services remain unavailable. Spare capacity is supplied by the capacity wave, not by assuming recovery software can create it.
  3. Any storage repair checks whether the old writer is truly stopped before reattachment. Soteria provides a verified recovery point if storage contents must be restored; it does not automatically overwrite live data because a link went down.
  4. Ariadne avoids treating resulting CI failures as independent application defects and repeatedly restarting builds.
  5. If the node's OS/runtime medium requires replacement, an authorized Metis rebuild uses the correct hardware image and preserves unrelated storage. A bad cable or power supply still needs physical repair first.
  6. Ananke performs the return-to-service checks, Flux converges approved configuration, and service checks determine whether the incident can close.

The final operator view should say, for example: "Node link recovered; storage validation pending; affected services X and Y; next check Z." It should not require the owner to inspect four separate repair loops to learn why a service is still down.

4. Prove recovery without risking current services

Implementation should progress through:

  1. Read-only source/runtime review: confirm exact deployed versions, current responsibilities, guards, inventory and effective settings. Full program audits remain outstanding beyond the paths already inspected.
  2. Unit/integration tests with simulated API, node, storage, backup and peer failures. Verify action counts, coordination, cancellation, retry exhaustion and safe escalation.
  3. Isolated backup/restore tests and a Metis test on disposable media or a spare node. Do not execute flashing tests against an active node or restore over a production PVC.
  4. A separately scheduled, limited live recovery drill with a known rollback, spare capacity and current verified backups. Test one failure domain at a time.
  5. Observe stability and require an operator walkthrough: identify the problem, locate the acting tool, understand why it stopped, and perform the documented approved next step.

An isolated restore, a synthetic recovery simulation and a real failover drill demonstrate different things. The completion record must identify which was performed and retain safe evidence.

Deliverables and effort

  • A concise ownership/dependency record for mechanisms actually retained after the configuration-first review.
  • Current verified recovery copies for the control-plane database and first application recovery targets.
  • Small reviewed changes to existing tools only for the gaps proven to need them, including tests for the failures actually observed.
  • Existing status/metrics and focused operator procedures sufficient to explain the remaining recovery paths. A common UI is optional, not a baseline requirement.
  • A documented recovery path usable while the cluster, SSO, registry or primary monitoring is down.

The original assessment's 1-2 day first wave covers the urgent backup/recovery prerequisites. Completing the broader coordinated-tool work is a provisional 4-8 engineering days, spread across the recovery and management-simplification waves, plus scheduled drills and observation. That overlaps the original assessment's automation and recovery estimates; it is not all additional work. Revise the estimate after confirming which safeguards already exist in each deployed program.