157 lines
22 KiB
Markdown
157 lines
22 KiB
Markdown
# Integrated infrastructure and recovery plan
|
|
|
|
Proposal for review, 2026-10-02. Updated to the user's configuration-first direction. It authorizes no cluster changes.
|
|
|
|
## Governing rule: standard configuration first
|
|
|
|
**Owner-operability requirement, added 2026-10-03:** the user must be able to understand and operate the cluster from atlas-iac, the actual tools and their documented configuration without this assistant, conversation history or a model-backed management decision. This is a completion criterion, not an optional documentation task. The audit observations remain dated October 2; this update is not a new live-health assessment.
|
|
|
|
The first implementation should use ordinary Kubernetes and component configuration to address the observed problems. Homegrown recovery changes are conditional: they need a demonstrated remaining requirement that the native controllers, application configuration or normal host administration cannot meet. This rule supersedes broader tool-integration proposals below and in earlier planning documents.
|
|
|
|
For each finding, use this decision order:
|
|
|
|
1. Establish the actual cause and check whether hardware repair or a supported configuration change addresses it.
|
|
2. Use the normal owner: Kubernetes workload controllers/scheduler, kubelet, Flux/Helm, Longhorn, or the application's own supported behavior.
|
|
3. Correct or retire overlapping custom behavior through reviewed changes when it is unnecessary or interferes with that owner. Do not assume more coordination code is needed just because two helpers currently overlap.
|
|
4. Verify service health and the relevant failure/recovery behavior.
|
|
5. Only if a gap remains, specify the smallest change to an existing tool, the evidence requiring it, its scope and its acceptance test. Prefer a configuration correction or bug fix over extending the program's responsibility.
|
|
|
|
The baseline does NOT require a shared incident protocol, new cross-tool locking service, unified recovery API/UI, inventory generator, or general expansion of Ananke/Ariadne/Hermes. Existing useful guards remain until a safe replacement/removal is reviewed. Scope changes are proposals; nothing is being disabled during investigation.
|
|
|
|
## Keep operation understandable without AI
|
|
|
|
- Use the top-level README as the entry point: show the Flux reference chain, the principal folders and safe first checks. Do not require the owner to read the complete historical audit to operate a service.
|
|
- Keep each service's authoritative settings discoverable through its Flux path and `kustomization.yaml`. Document meaningful exceptions, generated-file sources and any out-of-cluster settings. Avoid duplicate hand-maintained explanations of the same setting.
|
|
- Record what runs automatically, its scope, trigger, normal owner and how to stop or undo it. Remove obsolete helpers after validating their replacement instead of indefinitely accumulating recovery layers.
|
|
- Prefer small, explicit manifests and ordinary component features. Introduce a wrapper, controller or generator only when it removes demonstrable complexity and remains easier to inspect than the configuration it replaces.
|
|
- Core startup, backup execution, restoration and routine maintenance must work without model calls. Optional AI triage may assist but must not hold required state or be the sole way to select or execute a recovery procedure.
|
|
- Keep operational documentation beside the maintained code: root README for navigation, service `NOTES.md` for exceptions, tool usage/help for commands, and a short recovery guide for the few cross-service procedures. Historical audit evidence stays separate.
|
|
- Require handoff evidence for each changed component: the owner can find the setting, explain what will happen, inspect health, carry out the documented approved action and locate rollback. Test this without assistance from an AI session.
|
|
|
|
The root README has received a documentation-only navigation/diagnostic update. Detailed recovery commands remain unvalidated until the corresponding implementation and recovery tests are approved and completed. Do not label the existing cluster AI-independent or easy to recover merely because the target is documented.
|
|
|
|
## Minimal first implementation
|
|
|
|
| First action | Normal mechanism | Custom-tool work only if needed |
|
|
|---|---|---|
|
|
| Protect the external PostgreSQL datastore and important application data | Database-native backup tools, appropriate host scheduler, existing storage backup mechanism and isolated restore checks | Repair Soteria configuration or proven defects where it is the selected backup mechanism. Do not require Soteria/Ananke integration before obtaining a valid backup |
|
|
| Repair unreliable power/network/runtime media | Physical repair and supported OS/runtime configuration | Metis may perform a planned rebuild; a rebuild is not an automatic response to any node outage |
|
|
| Correct resource pressure | Accurate requests/limits, suitable node pools, namespace/job concurrency controls and compatible spare capacity | No custom scheduler; correct application/CI configuration before considering admission extensions |
|
|
| Correct placement and service health | Remove unjustified hard pins, supported replica placement, readiness/startup probes, safe rollout strategies and PDBs where useful | No new general restart controller; first eliminate conflicting helpers |
|
|
| Recover the known storage/application incidents | Component-supported diagnostics and recovery, preserved backups, correct CSI/DB/JVM settings | A targeted tool fix only if a repeatable defect remains after configuration/repair |
|
|
| Make Git authoritative and nodes maintainable | Resolve desired-state differences, staged Flux/Helm ownership/drift settings, supported versioned host configuration | No new configuration framework or broad tool integration |
|
|
| Verify availability, backup age and remaining failures | Existing metrics, component health checks and focused operator runbooks | Add only the missing signal, not a new monitoring/control platform |
|
|
|
|
Power/UPS shutdown sequencing, rebuilding a damaged host and restoring lost/corrupted data remain legitimate work outside ordinary pod reconciliation. Ananke, Metis and Soteria can serve those roles where their current implementations fit. Ariadne stays focused on its CI/application remit; Hermes is optional diagnostic assistance, not a prerequisite for cluster stability.
|
|
|
|
## Operating model
|
|
|
|
Basic infrastructure configuration should prevent normal workloads from exhausting the cluster and allow Kubernetes to handle ordinary container replacement and rescheduling. Ananke, Soteria, Metis and Ariadne/Hermes should handle the failure classes that need additional coordination, data recovery or operator assistance.
|
|
|
|
The roles are complementary:
|
|
|
|
| Layer | Normal responsibility | Owner |
|
|
|---|---|---|
|
|
| Physical hosts | Reliable power, network and runtime media; supported host configuration | Hardware maintenance plus versioned node profiles; Metis for authorized rebuilding |
|
|
| Cluster foundation | API/datastore, DNS, ingress, storage, secrets and adequate compatible spare capacity | K3s/Kubernetes, Longhorn and approved Flux-managed configuration |
|
|
| Applications | Accurate resource reservations, readiness, supported replication, persistent state and bounded background jobs | Native workload controllers plus per-service configuration |
|
|
| Exceptional recovery | Power/startup coordination, bounded host repair, data restoration and scoped application remediation | Ananke, Soteria, Metis, Ariadne/Hermes and narrowly scoped helpers |
|
|
| Operator understanding | Actual service health, recovery owner, last action, blocked dependency and next step | Existing monitoring/recovery interfaces backed by a maintained service inventory |
|
|
|
|
Routine Kubernetes reconciliation continues normally. Custom tools should not each implement a competing scheduler, garbage collector or pod restart loop. Disruptive recovery should become an exception rather than a normal requirement for keeping services online.
|
|
|
|
## Pair each finding with prevention and recovery
|
|
|
|
The R-numbers refer to [ASSESSMENT.md](ASSESSMENT.md). Actions below are proposed, not verified existing features.
|
|
|
|
| Finding | Basic infrastructure/configuration change | Existing tool's role | Proof of success |
|
|
|---|---|---|---|
|
|
| R1: External Kubernetes database recovery gap | Scheduled database-native backups on an independent host path, protected K3s recovery material, independent copies; supported OS and later standby design | Ananke checks the actual PostgreSQL dependency; Soteria displays backup/restore verification metadata without becoming a prerequisite for restoring Kubernetes | Isolated restore of a current backup; recovery material available with the cluster unavailable; measured recovery duration |
|
|
| R2: Power, link and runtime-media failures | Repair power/cabling/NIC problems; replace inadequate runtime media; keep suspect nodes out of ordinary placement | Ananke coordinates quarantine and return checks; narrow helpers attempt bounded repairs; Metis rebuilds only when required and authorized | Stable host under representative load; runtime and storage canaries pass; no renewed voltage/link/I/O fault during observation |
|
|
| R3: Insufficient compatible headroom | Accurate requests; limited CI/agent/scan concurrency; compatible node pools with capacity for a worker failure | Ariadne respects job admission and infrastructure incidents; Ananke checks recovery capacity before planned maintenance | Online services remain healthy during approved peak jobs and loss of one eligible worker; excess batch work queues visibly |
|
|
| R4: Shared PostgreSQL resource/probe gap | Explicit measured memory/CPU envelope, tuned database limits, startup/readiness checks and database-consistent backups | Soteria records protection; Ariadne avoids retrying dependent applications indefinitely; Ananke checks DB readiness before dependent startup | No repeat OOM under representative load; readiness reflects a usable database; backup restores successfully |
|
|
| R5: Persistent volume/engine faults | Resolve attachment/engine ownership with data preservation; supported CSI/storage settings, storage reservations and bounded rebuild activity | Ananke coordinates the incident without racing active writers; Soteria supplies a verified restore point when necessary | The two affected consumers mount usable storage and remain healthy; no repeating salvage loop; recovery copy preserved until verification |
|
|
| R6: Backup coverage/freshness gap | Explicit eligible data set, local-only exceptions, achievable schedules/retention and correct driver behavior | Fix Soteria's actual backend/skip failure paths and validate its existing restore flow | Every required service has current protection or an explicit reconstructible-data policy; missed backups and stale metadata are unhealthy |
|
|
| R7: Git/live divergence | Reconcile intended values before replacing creation-only handling; enable appropriate drift checks gradually | Flux owns normal desired state; Ananke/tool exceptions are scoped, visible and expire through an agreed procedure | Git and live state agree on the changed components; recovery does not fight reconciliation; parked services are explicit |
|
|
| R8: Overlapping maintenance | Remove redundant cleanup and unconditional restart behavior after replacements are proven; use native controls for ordinary operation | Existing tools share target ownership and bounded action rules; destructive operations have explicit authority | One incident cannot trigger competing restarts, evictions or rebuilds; exhausted recovery stops and reports why |
|
|
| R9: Power-recovery/inventory drift | One reconciled node/service dependency inventory; measured shutdown timing; independent bootstrap assets | Ananke uses actual roles and dependencies; Metis builds the same approved node profiles | Simulated failure tests and later a controlled power/startup drill follow the dependency order and preserve data |
|
|
| R10: Service failure-domain weaknesses | Spread supported replicas, preserve singleton semantics, remove unnecessary hard pins, provide compatible spare capacity and durable job state | Ananke coordinates failures requiring host action; Soteria protects state; Ariadne handles only its scoped application/CI recovery | Every service has a tested continuity or recovery target; a protected front end is not declared healthy while its backend is unavailable |
|
|
| R11: Unavailable/parked services | Correct OpenSearch placement and heap/request mismatch; diagnose Firefly; explicitly budget and schedule parked GPU capabilities | Ariadne records dependency causes and bounded actions; Metis is relevant only if a required host must be rebuilt | OpenSearch and Firefly pass functional health checks; each parked capability has an approved activation/capacity plan |
|
|
| R12: Unsupported/inconsistent baseline | Tested hardware-specific OS/K3s/storage/runtime versions; staged upgrades and versioned node configuration | Metis uses approved images; Ananke manages safe node return and maintenance order | One cohort upgraded at a time with compatibility, health and recovery checks; deployed versions are visible |
|
|
| R13: DNS/TLS/bootstrap/observability issues | Normalize resolver configuration; one certificate owner; protected bootstrap image set; bounded Job/log retention; missing telemetry shown as unknown | Ananke tests dependencies; Soteria distinguishes backup success from scan success; existing monitoring exposes actionable incidents | TLS/DNS and image bootstrap checks pass; stale metrics cannot appear healthy; a whole-cluster monitoring failure remains visible through an approved independent path |
|
|
|
|
This covers a treatment path for every finding. It does not establish that configuration alone can repair faulty hardware, that currently reserved hardware may be reassigned without approval, or that every application supports uninterrupted failover.
|
|
|
|
## Implementation order: configuration first, targeted tool fixes when justified
|
|
|
|
### 1. Protect data and establish change ownership
|
|
|
|
First, reconcile the live node/service inventory and the few ownership settings required for the components being touched. Do not remove all IfNotPresent annotations at once or perform a cluster-wide reconciliation reset.
|
|
|
|
Create current external-PostgreSQL recovery copies using a versioned host-side job, then prove restoration in isolation. Repair the first Soteria backup failures and prove representative application recovery. Preserve recovery copies before manipulating stuck storage. Prepare a protected recovery package usable without cluster DNS, Vault UI, Harbor or SSO.
|
|
|
|
At the same time, identify any existing automation that could interfere with the specific repair. Prefer removing the overlap or correcting configuration over adding coordination code. Full tool refactoring is not a prerequisite for closing the urgent backup gap.
|
|
|
|
**Gate:** approved backups and recovery prerequisites exist for the next repair; the acting owner and rollback are known. The initial few restores do not count as complete per-service restore coverage.
|
|
|
|
### 2. Repair hosts and make ordinary workloads fit
|
|
|
|
Handle physical power/link/media issues, beginning with confirmed symptoms and dependencies. Restore reliable worker capacity, or explicitly arrange a compatible alternative. Correct application PostgreSQL resources and health checks. Recover OpenSearch and the stalled volumes in small steps. Diagnose Firefly independently.
|
|
|
|
Set measured requests and concurrency for CI, agents, scans and heavy jobs. Keep front-end services and recovery capacity protected while excess batch work waits. Do not silently disable applications to satisfy an availability target.
|
|
|
|
Review Ananke/Metis settings only where the host repair actually depends on them. Bound CI retry and concurrency settings if they amplify the incident. Shared incident ingestion by Ariadne is a possible later extension, not a baseline dependency.
|
|
|
|
**Gate:** normal approved load no longer causes recurring saturation/OOM; repaired nodes stay healthy and affected services work. Quarantined or retired nodes have no required workload stranded on them.
|
|
|
|
### 3. Simplify steady-state management
|
|
|
|
Make Git match the approved live design before changing reconciliation ownership. Use one reviewed node profile per hardware class. Define eligible node pools and hardware exceptions explicitly. Replace overlapping restart/cleanup behavior with native controls and a small number of bounded helpers.
|
|
|
|
Document any necessary tool exception to Flux. A temporary cordon, paused batch workload or recovery state needs an owner, purpose and a checked exit condition. Keep indefinite operator quarantine distinct from an expiring automatic exception. Do not build a new exception controller unless an observed operating requirement justifies it.
|
|
|
|
**Gate:** the next reconciliation, helper restart or routine node reboot does not reintroduce an old workaround or initiate unrelated repair work. The operator can identify the owner of each setting and action.
|
|
|
|
### 4. Establish continuity for every service
|
|
|
|
Demonstrate failure capacity within each compatible pool. Spread stateless replicas and shared dependencies where supported. For stateful services, validate storage, database, session and writer behavior before changing replica counts. Give singletons a measured recovery target and long-running jobs durable progress/recovery semantics.
|
|
|
|
Extend verified backup coverage service by service using the selected existing mechanisms. Native workload controllers and application readiness should handle ordinary dependency recovery. Add Ananke/Ariadne changes only for demonstrated remaining gaps. Keep local-only inference/job data on authorized local infrastructure.
|
|
|
|
**Gate:** every service in the catalog has a tested failure behavior, recovery owner and current data-protection status. If existing hardware cannot meet a target, report the required capacity or explicit scheduling tradeoff before claiming completion.
|
|
|
|
### 5. Validate the combined system and hand it over
|
|
|
|
Upgrade supported software cohorts once backups and recovery prerequisites are proven. Test one failure domain at a time in an approved window. Begin with simulations/disposable targets and proceed to controlled live drills; no spontaneous outage injection.
|
|
|
|
Verify the full path: detection, native rescheduling where appropriate, tool coordination, restored application health, incident closeout and operator explanation. Observe the result for 7-14 days. A fresh configuration and a single passing test are not proof that intermittent faults are resolved.
|
|
|
|
**Gate:** the user can identify a representative fault and follow its documented next step without an AI assistant. Required operations work with AI-assisted management unavailable. A cluster/API failure still has an independent recovery path. The repository and installed tools contain the needed instructions and configuration references.
|
|
|
|
## Keep the configuration small and understandable
|
|
|
|
Use a maintained service catalog to record a small set of operational facts for every service: dependencies, eligible node pool, measured resource budget, persistent-data location and backup policy, health check, recovery owner and expected recovery behavior. Reconcile existing Ananke, Metis and Flux inventories against it. Decide later whether a generator is worthwhile; do not introduce another configuration framework merely to connect the plans.
|
|
|
|
The standard service pattern should use:
|
|
|
|
- Explicit resources where defaults are unsuitable, especially databases, JVMs and model workers.
|
|
- Readiness that reflects ability to serve requests; startup checks for slow initialization; liveness only for conditions a restart can actually correct.
|
|
- Preferred placement with compatible alternatives where possible; hard placement only for real hardware/data constraints.
|
|
- Supported replica spreading and disruption protection, plus actual spare capacity. Neither a PDB nor replicated storage creates application-level HA by itself.
|
|
- Bounded background concurrency, storage/log retention and one normal garbage-collection mechanism.
|
|
- Backup freshness and restore verification that match the service's data policy.
|
|
|
|
Avoid turning Ananke into the scheduler, Metis into an automatic response to every unreachable node, Soteria into a required bootstrap dependency, or Hermes into the authority for destructive recovery. These tools can remain smaller and easier to understand when the foundation handles ordinary operation reliably.
|
|
|
|
## Concrete example: loss of a general-purpose worker
|
|
|
|
With sufficient compatible spare capacity and tested configuration, surviving application replicas keep serving and Kubernetes schedules replacements where their storage and placement permit. Bounded batch workloads leave room for that recovery. Ananke tracks the node incident and coordinates any exceptional host repair; Ariadne defers dependent builds. A volume is restored only if needed, from Soteria's verified recovery point. Metis becomes relevant if the node must be rebuilt.
|
|
|
|
The operator sees affected services, recovery progress, the acting owner and a specific next step. If there is no safe capacity or a possibly live storage writer, the system reports that blocker instead of concealing it behind repeated restarts.
|
|
|
|
## Scope and evidence
|
|
|
|
This integrated plan uses the observations in [ASSESSMENT.md](ASSESSMENT.md), the tool requirements in [RECOVERY_TOOL_PLAN.md](RECOVERY_TOOL_PLAN.md), and the per-service coverage in [SERVICE_PLAN.md](SERVICE_PLAN.md). It adds no claim that an uninspected tool feature already exists. Deployed-version source review, root-cause checks and acceptance evidence remain required before the corresponding implementation is considered complete.
|
|
|
|
The broad 2-4 week estimate from the original assessment remains a provisional envelope for the larger reliability effort, not a commitment to spend that time integrating tools. Re-estimate after the basic fixes: unnecessary custom-tool work should be dropped. Physical access, usable spare capacity, application HA choices and remaining unknowns control duration. No implementation has begun.
|