260 lines
37 KiB
Markdown

# Atlas cluster reliability assessment
Investigation only. Prepared 2026-10-02. Observations collected approximately 06:52-07:30 UTC (01:52-02:30 CDT).
No workloads, node settings, manifests, credentials, storage, or reconciliation state were changed. No recovery script, inference job, restore, or destructive test was run. Local audit files and a detached inspection worktree were created. The pre-existing OpenSearch manifest edit was preserved and is not counted as deployed.
## Assessment
The cluster's recurring instability has several interacting causes. It is not explained by Titan-22 alone. The general-purpose Pi pool is nearly full in scheduler reservations, several hosts have physical or I/O problems, important applications inherit unsuitable resource defaults, and some storage failures persist behind otherwise Running pods. Recovery and configuration authority are fragmented across Flux, Helm, host services, and multiple maintenance controllers.
The first priority should be recoverability: the actual Kubernetes datastore is a single external PostgreSQL server, with no active standby or WAL archiving and no recent backup found. The second is restoring dependable capacity and resolving confirmed application/storage faults. The third is simplifying ownership and measuring service health rather than counting Running pods.
All services remain in scope. Foundational services go first because every other application depends on them, not because the other applications are disposable. Maintaining service during a node failure requires both compatible spare capacity and application-specific state handling. Some current singleton/GPU services cannot offer uninterrupted failover simply by increasing replicas.
## Coverage and limits of this investigation
| Area | Examined |
|---|---|
| Git configuration | Deployed `main` revision `d114ce71e8495d5025a90ee6b5e19c5d911f58b9`; 794 YAML files parsed; all 58 live Flux paths rendered successfully |
| Live workloads | 21 Kubernetes nodes; 130 Deployments, 13 StatefulSets, 34 DaemonSets; 31 CronJobs and 1,775 retained Job objects |
| Networking and policy | 182 Services and EndpointSlices; 42 Ingresses; 47 NetworkPolicies; 38 Certificates; ingress, MetalLB, DNS, webhooks, quotas and placement settings |
| Storage | 98 PVCs, 140 PVs, 134 Longhorn volumes, 358 replica objects, storage-node disks, settings, recurring jobs and backup metadata |
| Host health | All 19 reachable Kubernetes nodes sampled; SSH failure verified for two offline nodes; external `titan-db` database and Ananke services checked |
| History | Available seven-day node/working-set/OOM metrics; 24-hour restart counters; current events and selected host kernel/runtime diagnostics |
| Reachability | 39 distinct ingress hostnames checked from `titan-db`, using LAN IPs, TLS verification, no redirects and no credentials |
The workload inventory is [workloads.csv](workloads.csv). It records every Deployment/StatefulSet/DaemonSet, placement, readiness, images, storage and resources. [namespaces.csv](namespaces.csv) and [SERVICE_PLAN.md](SERVICE_PLAN.md) cover the service families. [node-capacity.csv](node-capacity.csv) contains reservation calculations.
This is broad configuration and operational inspection, not a claim that every application workflow or every line of embedded application code was tested. No authenticated application transactions, mail delivery, physical power tests, disk self-tests, restore drills, deliberate node failures or full cold-start tests were performed. Switch/router configuration, cables, power supplies, battery condition under load and backups outside the inspected locations remain unverified. Metrics can have blind spots during monitoring outages. Runtime configuration of individual applications may require a second, targeted inspection during implementation.
## Current condition
- 19/21 nodes report Ready. Titan-05 has been unreachable to Kubernetes since August 21; Titan-06 since September 21. Titan-04 is Ready but cordoned and showing current undervoltage events.
- The first snapshot contained 635 pods. Of 41 nonterminal pods without Ready=true, 35 were on the two offline nodes. These are not 41 independent application incidents.
- Firefly and the quality Pushgateway deployment were unavailable in both snapshots. OpenSearch remained Pending. The Pushgateway pod has been stuck since September 28; OpenSearch since August 21.
- A Hermes tenant workspace is stuck in a Longhorn recovery loop despite consumer pods reporting Running.
- At the later check, 57/58 Flux Kustomizations reported Ready. The exception was a suspended migration Kustomization carrying an old error; its current path renders. All 21 HelmReleases reported Ready, which does not establish runtime availability.
- All 39 tested LAN hostnames completed TLS. Grafana health, Gitea health and Keycloak discovery returned 200. Registry `/v2/` and unauthenticated suite-planning capabilities returned the expected 401. Three root paths returned 404 and many returned login redirects. These are transport checks, not authenticated functional tests. Gitea health and some login paths took roughly 6 seconds in this sample.
- Direct LAN connections from the audit workstation timed out; the same checks succeeded from the LAN host. The workstation's route must not be mistaken for a cluster-wide ingress failure. No laptop connectivity test was performed.
## Findings and proposed work
### R1. Protect the real Kubernetes database before disruptive work - critical
All three K3s servers use PostgreSQL at `192.168.22.10:5432/k3s` on `titan-db`. They do not currently use embedded etcd as their datastore. The PostgreSQL server is primary, has no replication clients or replication slots, has `archive_mode=off`, and reports zero archived WAL files. Its k3s database is approximately 606 MB.
No database-backup schedule was found in the inspected systemd timers, cron locations or root/postgres crontabs. The discovered dump, role dump and server-token backup files under `/var/backups/k3s` are dated **2025-08-29**. Their restore validity was not tested. This establishes an unverified-current-backup gap, not proof that no other copy exists anywhere.
The host runs Ubuntu 24.10 and PostgreSQL 16.9. Ubuntu 24.10 reached end of life on July 10, 2025. [Ubuntu release notice](https://lists.ubuntu.com/archives/ubuntu-announce/2025-July/000314.html)
**Proposal:** create and verify an application-consistent PostgreSQL backup plus the associated K3s server token, then keep encrypted independent copies with age alerts. Test recovery into an isolated environment. K3s documents that external-database backups are the administrator's responsibility and that the server token is needed for restoration. [K3s backup guidance](https://docs.k3s.io/datastore/backup-restore)
Next, replace the unsupported host OS through a staged migration to a supported baseline. Choose either a supported PostgreSQL standby/failover arrangement or a separately planned migration to embedded etcd; do not combine datastore replatforming with the first emergency repairs. Retaining PostgreSQL initially is the smaller change. A standby alone does not provide safe automatic failover without fencing and client-endpoint handling.
**Estimate:** 4-8 engineering hours for backup scheduling and an isolated restore proof; 1-2 days for a staged host/database migration; 2-4 additional days if database failover is implemented. Physical access and destination capacity may add waiting time.
### R2. Repair power, connectivity and runtime-media faults - critical/high
Titan-04 and Titan-11 emitted undervoltage kernel messages during this inspection. Titan-07 and Titan-08 also have historical undervoltage/throttling flags. A historical flag alone does not establish current undervoltage; the fresh kernel messages on 04/11 do.
Titan-14 repeatedly showed about 77% I/O pressure and about 60% full I/O pressure. Its containerd, image and kubelet paths are mounted from a 58.6 GB USB device identified as `USB DISK`, not from its root SD card. A two-second sample showed that device busy for approximately the entire interval, with about 24 ms average write completion time. A USB-storage task was blocked in D state. This supports a runtime-storage bottleneck; it does not yet prove failed flash hardware.
Titan-22 is currently reachable and its overlay is present. The prior USB-network loss remains a physical/link reliability concern; bounded overlay repair cannot repair a defective NIC, power connection or cable. The Sept 30 `sd*` I/O errors examined here correspond to Longhorn iSCSI devices, so they are not evidence that its local NVMe failed.
**Proposal:** inspect power supplies, USB power budget, cables, switch ports and link counters; replace confirmed weak components. Move write-heavy runtime data off inadequate USB flash/SD media onto suitable SSDs, using a drained and backed-up migration. Maintain quarantine until a repaired node passes stability checks. Do not just uncordon Titan-04 to gain apparent capacity.
**Estimate:** see the node table below. Physical diagnosis of 05/06 cannot be completed remotely while they remain inaccessible.
### R3. Correct reservations and create compatible failure capacity - high
At the snapshot, Pi workers 07/11/12 reserved 99%, 99% and 97% of allocatable CPU respectively. Node-07 reserved about 93% of memory; node-11 about 99%. A Jenkins agent was unschedulable for insufficient CPU/memory plus placement constraints. Titan-12's CPU pressure persisted around 70% through repeated samples.
Spare cluster-wide capacity is misleading: Titan-23 has large spare capacity but is intentionally reserved for simulations; Titan-22 is discouraged for general placement; Titan-24 is reserved for heavy work. Many services require ARM/Pi nodes. Their image architecture and placement must be checked before moving them.
**Proposal:** right-size each service from actual memory peaks and CPU demand, then calculate N+1 capacity within each compatible pool, including DaemonSets, sidecars, volume topology and a failed worker. Reserve headroom for rebuilds and short spikes. Set separate bounded concurrency/resource budgets for CI, AI agents, scans and simulations so an application request cannot consume the platform's recovery capacity. Jobs can queue without making the application front ends unavailable.
The preferred policy remains Pi5 then Pi4 for suitable workloads. If repaired Pi capacity cannot meet the measured N+1 requirement, explicitly approve either a limited general-purpose reservation on existing x86 capacity or additional reliable general-purpose workers. No x86 reassignment is implicit in this report. Hardware sizing needs measured demand and an architecture/placement check; idle RAM alone is not a capacity plan.
**Estimate:** 1-2 days for measurement-based reservations and queue limits; 2-4 days for staged placement changes and failure-capacity verification across service pools. Additional hardware procurement is separate.
### R4. Shared application PostgreSQL has an unsuitable default envelope - high
`postgres/postgres-0` on Titan-07 is separate from the external K3s database. Its main container inherits a 50m CPU/96 MiB memory request and 500m CPU/512 MiB memory limit. The source StatefulSet defines no explicit resources and no readiness, startup or liveness probe for PostgreSQL. Its seven-day maximum working set was about 511.75 MiB. Available metrics record four OOM events for that container; these need not be four container restarts because an OOM can kill a child process.
This database supports many services, while Vault and Keycloak also reside on Titan-07. A nominally Running database pod can therefore conceal a correlated application failure.
**Proposal:** give PostgreSQL explicit measured resources, tune connections/memory to that budget, add a startup probe and meaningful readiness, and avoid an aggressive liveness probe that restarts it during transient storage slowness. Add database-native backups and spread its dependent control services. Build a standby/failover design only with the required capacity and tested storage semantics.
**Estimate:** 4-8 hours for the immediate resource/probe/backup configuration and validation, then separate HA work if approved. Avoid choosing a new memory limit solely by doubling the current one.
### R5. Resolve two persistent Longhorn failures without losing the recovery path - high
1. `hermes/workspace-hermes-chat-tenant-2`, volume `pvc-02d99a30-3757-4a01-abfb-5b30aef793d6`: volume state detached; share manager not becoming ready; an old engine still runs on Titan-23 with an unknown errored replica entry while three desired replicas are stopped. The volume/share manager points toward Titan-14. Repeated salvage, engine-delete and mount-related events are ongoing.
2. The quality Pushgateway volume, `pvc-3a2c0caf-ccdb-4870-ac07-12c99eb72f99`: the pod remains ContainerCreating, while the engine reports Running on Titan-08 but has no replica-mode map and CSI reports the volume not attached.
**Proposal:** map actual mounts, engine ownership, replica revision/health and attachment tickets; preserve verified recovery copies before controlled detach/reattach or engine repair. Stop blind salvage/restart loops for these incidents through a scoped maintenance procedure. Do not force-delete the last replica or assume an `actualSize=0` field means there is no data.
134 Longhorn volumes include 62 detached volumes and 63 with robustness `unknown`. Many are retained/unused volumes; they must not all be called corrupt. There are 43 Released PVs. Classify ownership and retention before deleting anything. The pending Cassandra cache claim uses `WaitForFirstConsumer`, so Pending alone is not evidence of storage failure.
**Estimate:** 4-12 hours for the two incident investigations/recoveries if a usable replica exists. Recovery duration and data integrity remain uncertain until replica inspection. A restore from an older backup, if needed, is a separate user decision.
### R6. Backup configuration exists, but current recoverability is not demonstrated - critical/high
The Longhorn backup target is reachable, but only two BackupVolume records were present, with last backups in June/July 2026. Soteria reported a successful bucket scan with zero new objects in 24 hours; the newest observed object was in July. Its counters included 81 backend errors and many `live_rwo_mount` skips. These are cumulative counters, not a claim about an hourly error rate.
Soteria is configured for the Longhorn driver, a 24-hour age goal, and one policy backup every 1,800 seconds. That has a theoretical maximum of 48 new policy-triggered backups a day before errors and runtime. There are 98 claims, including exclusions and caches; the actual eligible set must be calculated. Mounted RWO skip behavior also needs to match the chosen backup driver.
**Proposal:** audit the driver/API failure path and eligible-PVC inventory, establish per-service backup and restore objectives, and verify fresh completed backups and isolated restores. Use database-consistent backups for PostgreSQL; storage replication is not a backup. Keep authorized local-only suite records/checkpoints on approved local backup infrastructure, not automatically in B2. Excluding local-path from Soteria currently leaves those workloads needing a separate recovery method.
**Estimate:** 1-3 days to repair backup coverage and prove a representative restore set, followed by per-service restore coverage. No destructive restore testing is authorized by this assessment.
### R7. Flux cannot currently be treated as a complete source of truth - high
46 of 58 live Flux Kustomizations have the `kustomize.toolkit.fluxcd.io/ssa: IfNotPresent` annotation. This causes their parent to create these objects but skip subsequent updates to existing definitions. Their child workload reconciliation still operates; it is the Kustomization definitions themselves that are not being kept in sync. [Flux apply policy](https://fluxcd.io/flux/components/kustomize/kustomizations/)
Rendered Git and live definitions differ in 22 fields: 18 suspension settings and four Harbor/Vault health-check, wait or timeout settings. The diff is preserved in [flux-spec-diff.json](flux-spec-diff.json). Field ownership includes historical `kubectl-patch` changes. This explains how a green Flux status can coexist with configuration that differs from the repository.
**Proposal:** first decide and commit the desired state of each divergent field; then remove creation-only handling in small reviewed groups. Do not remove all these annotations at once: Git currently contains suspension settings that would affect active services. Separate bootstrap state from steady-state desired state and define one owner for temporary recovery exceptions with an expiry and audit trail.
All 21 HelmReleases lack configured drift detection. The Weave GitOps release reports Ready but has no backing application pod/deployment and its service has no endpoints. Review that optional service's actual desired state, then add drift checks selectively after capturing legitimate runtime-managed fields. A green Helm release can mean its last install succeeded, not that all objects still exist.
**Estimate:** 1-2 days for ownership cleanup, staged convergence and safeguards. Recovery-controller interactions must be tested before broad enforcement.
### R8. Too many independent mechanisms can restart, evict or clean up workloads - high
Confirmed examples include a DaemonSet that unconditionally restarts `k3s-agent` when its container starts; multiple image-pruning/sweeping mechanisms plus kubelet garbage collection; a per-minute node-placement CronJob; a descheduler every 20 minutes; Ananke startup/recovery actions; Ariadne/Hermes CI recovery; and Flux/Helm rollout behavior.
The descheduler has useful limits (two evictions per node and namespace, node-fit checks, PVC protection), and Hermes CI recovery has an action allowlist. Those protections should be preserved. There is no established common disruption budget across all the independent actors. A restart-count threshold of 12 also needs a time window rather than treating a long-lived historical count as a current incident.
**Proposal:** replace permanent one-shot restart helpers with versioned, idempotent host configuration. Retain kubelet image GC as the normal mechanism and one bounded emergency cleaner, protecting bootstrap images. Establish one maintenance lock, node ownership/cordon annotations, cooldowns, action limits and circuit breakers. Recovery should stop and expose one actionable incident after bounded failure, rather than create an indefinite loop. Pause discretionary movement during node/storage recovery through the approved workflow.
**Estimate:** 2-4 days for consolidation and regression tests; do not disable everything at once or remove known recovery protections without replacements.
### R9. Ananke inventory and the power-recovery procedure need reconciliation - high
Both out-of-cluster Ananke instances were inspected. Titan-db is the coordinator; Titan-24 is a peer forwarding shutdown to Titan-db with local fallback disabled. That is evidence of intended coordination, not two proven competing leaders. Both UPSs currently report online and 100% charge.
The configuration still includes absent nodes 09/10 (explicitly ignored for startup availability), omits Titan-23 from the general worker inventory, and assigns the Harbor bootstrap label to Titan-09 although Harbor is actually pinned to Titan-11. Some differences may be intentional, but they need one maintained inventory and explanation. The configuration mentions etcd restore/snapshot behavior; inspected source has an external-datastore guard and the live K3s unit uses its recognized syntax. No evidence was found that an etcd restore was incorrectly run.
Reported UPS runtimes were approximately 900 seconds for Titan-db's UPS and 1,325 seconds for Titan-24's. The configured default shutdown budget is 1,380 seconds, with a 420-second emergency path and a runtime safety factor. Verify the actual early-trigger/deadline behavior against measured shutdown time; simply comparing these constants does not prove the current trigger is wrong. The bootstrap unit on Titan-db had a restart counter of 35, requiring a bounded failure history and a clear latest-success indicator.
**Proposal:** version and expose host daemon revisions, reconcile the inventory, document startup dependencies and external-PG recovery, and validate shutdown timing with a non-destructive simulation followed by a separately scheduled controlled drill. Gate self-updates during maintenance or power instability. Keep emergency recovery instructions usable without Grafana, Vault UI, or an operational cluster.
**Estimate:** 1-3 days for inventory, tests and operating instructions; a controlled power drill requires a separate maintenance window.
### R10. Improve failure-domain design for every service - high/medium
117 Deployments/StatefulSets request exactly one replica. Some correctly require a singleton; some stateless front ends could run redundantly. Vault, Keycloak and PostgreSQL currently share Titan-07. Harbor, Grafana and Alertmanager share undervoltage-affected Titan-11. Both Vault injector replicas are on Titan-12; its webhook is failurePolicy=Ignore, so failure may produce missing injection rather than block every admission.
There are 27 PDBs, mostly Longhorn-managed. Application protection is sparse. A PDB does not survive a node outage for a singleton or prevent direct pod deletion; replicas, spare capacity, placement and application state still matter. [Kubernetes disruption guidance](https://kubernetes.io/docs/concepts/workloads/pods/disruptions/)
**Proposal:** spread the shared foundation first, then give every service either redundant instances or a documented, tested restart/failover path with a recovery target. Stateful applications need compatible multi-writer/session/database designs before replicas are raised. Preserve stable identifiers, storage ownership and credential boundaries. The suite-planning API needs durable local job state and interrupted-job recovery even if its server is initially singleton. Single-GPU inference requires an approved second local backend or an explicit queued/degraded mode to survive that GPU's loss.
### R11. Restore unavailable applications and distinguish intentionally parked services
OpenSearch is hard-pinned to offline Titan-05. It also requests 768 MiB while its configured JVM heap is 2 GiB; scheduling based on that request understates the intended footprint. An uncommitted local edit removes the pin, but it is not deployed. Move it only after choosing capacity and validating its existing volume and memory configuration. Firefly needs a targeted readiness/dependency diagnosis; its cause was not established by this audit and must not be assumed to be the database OOMs.
Several GPU/auxiliary services are deliberately scaled to zero, including Wolf, the batch Ollama deployment, local image inference, P2Pool and a Sui test wallet. They are not healthy available services merely because desired replicas is zero. Keep them in the catalog with an explicit parked reason and capacity/activation plan. The user's previous GPU reassignment explains some parking; do not reverse it silently.
**Estimate:** OpenSearch 4-8 hours if its volume is usable; Firefly 2-6 hours for diagnosis and a scoped fix; restoring parked capabilities depends on agreed GPU/compute allocation.
### R12. Standardize the supported software baseline after recovery safeguards
Fourteen nodes run K3s v1.33.3 and seven run v1.31.5, with two containerd generations and multiple OS/kernel families. These Kubernetes minor branches are past their upstream end-of-life dates as of this audit. K3s vendor extended support was not established. [Kubernetes release history](https://v1-33.docs.kubernetes.io/releases/patch-releases/)
**Proposal:** maintain a tested compatibility matrix for K3s, Longhorn, CSI, cert-manager, Traefik, GPU runtime and hardware-specific Jetson kernels. Upgrade one supported step and one failure domain at a time, following component compatibility guidance at implementation time. Do not treat a general Ubuntu upgrade as a safe Jetson GPU upgrade. Move host configuration out of scattered always-running privileged repair pods where practicable.
**Estimate:** 2-5 days of staged engineering work plus soak periods, after storage/data recovery is proven. Hardware-specific upgrades may take longer.
### R13. Clean up DNS, certificates, image bootstrap and health reporting
- Titan-24 generated 159 retained sandbox-creation warning events associated with external DNS failure while resolving Docker Hub's pause image. Host DNS resolved successfully during the later check. Cause and duration remain uncertain; this is a transient dependency failure, not established permanent DNS misconfiguration.
- Titan-23 emits repeated nameserver-limit warnings with duplicated public resolvers. Normalize its resolver configuration and validate LAN plus external names.
- Two Harbor Certificate objects compete for the same TLS secret. The live registry certificate worked, but one Certificate remains IncorrectCertificate. Keep one declarative owner.
- Longhorn and other essential images come from Harbor, whose storage and database are themselves cluster dependencies. Protect a minimal bootstrap image set and document recovery order so image GC cannot strand a cold start.
- MetalLB, K3s ServiceLB for mail and legacy service objects coexist. Document which system owns each VIP/port before retiring anything; no active same-VIP conflict was demonstrated.
- Quality Pushgateway is unavailable, and some dashboards use zero fallbacks. Missing/stale telemetry must be shown as unknown, not healthy zero. Grafana alerting and Alertmanager have separate routes; Alertmanager's default receiver is intentionally silent except specific Hermes routes. Grafana has its own email policy. End-to-end delivery was not tested.
- Alerts depend on in-cluster mail and identity/monitoring. Add a narrowly scoped independent reachability/heartbeat signal so a total cluster/mail failure is visible. No messages or new external monitoring integrations were sent/created during this audit.
- Most retained Job objects are old comms jobs (1,396 successful/failed objects). Apply explicit retention after preserving required evidence. This is cleanup and usability work, not an established cause of database overload.
**Estimate:** 1-2 days for DNS/certificate/telemetry/retention corrections and ownership documentation; bootstrap recovery validation is part of the coordinated recovery work.
## Node repair and maintenance estimate
Estimates are hands-on engineering time, not guaranteed elapsed time. No physical repair is established without inspection. Rebuilds, soak tests and obtaining parts add elapsed time.
| Node | Observed condition / role | Proposed treatment | Ballpark |
|---|---|---|---|
| titan-db | External K3s PG16; unsupported Ubuntu24.10; no active replication or current scheduled backup found | Backup/restore proof, supported-host migration, evaluate standby | 4-8h backup first; 1-2d migration |
| titan-0a | Ready control plane; SSD; current API healthy | Preserve external-PG recovery path, standardize host settings and staged upgrade | 2-4h/node plus soak |
| titan-0b | Ready control plane; SSD; current leader/controller lease healthy | Same staged control-plane work | 2-4h plus soak |
| titan-0c | Ready control plane; SSD; sampled load spike but no established persistent fault | Same; watch sustained demand | 2-4h plus soak |
| titan-04 | Ready but cordoned since Sep14; fresh undervoltage | Physical power/cable/USB inspection, then stability gate before return | 1-3h physical; 2-4h validation |
| titan-05 | Offline since Aug21; SSH refused | Check power, address, SSH/runtime and boot media; recover or cleanly retire role | 2-6h diagnosis; 4-8h rebuild if needed |
| titan-06 | Offline since Sep21; no route to host | Physical link/power first; recover or replace | 2-6h diagnosis; 4-8h rebuild if needed |
| titan-07 | ~99% CPU requests, ~93% memory requests; high CPU pressure; PG/Vault/SSO together; historical undervoltage | Move competing work, right-size shared services, inspect power history | 4-8h staged changes |
| titan-08 | ~88% CPU requests; 16 Ready transitions in available 7d history; historical undervoltage; stuck Pushgateway engine | Correlate reboot/runtime/power history; fix volume independently | 3-6h diagnosis, storage work separate |
| titan-11 | Fresh undervoltage; ~99% CPU/memory requests; Harbor/Grafana/alerts | Power repair and reduce dependency concentration | 2-4h physical; 4-8h placement |
| titan-12 | Sustained ~70% CPU pressure; 77 containers; ~228 scheduled exec health checks/minute | Right-size/move workloads; reduce expensive probe process launches after measurement | 4-8h |
| titan-13 | Storage Pi, many replicas; no current kernel fault established | Reserve storage CPU/RAM/network; baseline disks and rebuild performance | 2-4h |
| titan-14 | Sustained severe I/O pressure; USB runtime device saturated | Inspect/replace runtime medium, migrate safely, disentangle stalled mounts | 4-8h plus copy/drain time |
| titan-15 | Storage Pi; CPU pressure observed; many replicas | Storage-only reservation and performance baseline | 2-4h |
| titan-17 | Storage Pi; large replica population; no established hardware fault | Storage baseline and capacity reserve | 2-4h |
| titan-18 | CPU pressure and recent I/O contention; CI/scans; root currently not full | Budget scans/agents and runtime I/O; retain controlled disk policy | 3-6h |
| titan-19 | Storage Pi; no current hardware fault established | Storage baseline and version alignment | 2-4h |
| titan-20 | Jetson local inference; very little MemAvailable in sample; unified-memory demand | Account for GPU/shared memory and measured concurrency; preserve assigned workload | 3-6h |
| titan-21 | Jetson under memory/swap pressure; Java OOM kernel records; Data Prepper restarts | Attribute OOM to exact cgroup, tune JVM and competing workloads | 3-6h |
| titan-22 | Reachable now; previous USB NIC loss; overlay workaround active | NIC/cable/power/switch validation; keep bounded recovery, not dependence on it | 2-4h physical; 2-4h validation |
| titan-23 | Large spare capacity but simulation-reserved; duplicate DNS; absent from general Ananke worker list | Document reservation/ownership, normalize DNS; no reassignment without approval | 2-4h |
| titan-24 | Currently healthy; recent external DNS sandbox errors; suite jobs and recovery peer | Protect long jobs, bootstrap image cache and resolver reliability; document restart ownership | 2-4h |
Do not add every row to the project estimate: baseline, placement and automation work overlap. Disk/UPS/network diagnostics should be extended where symptoms persist; the brief samples are not performance certification.
## Implementation waves for approval
**Latest scope direction:** start with standard Kubernetes/component configuration and physical repairs. Homegrown-tool changes require a demonstrated remaining gap. Do not make broad tool integration, a new coordination service or a new dashboard a prerequisite for stabilization.
**Owner-operability direction, 2026-10-03:** operation and recovery must be understandable from the repository and tools without an AI assistant or conversation history. Prefer removing unnecessary mechanisms over adding management layers. Each implementation handoff must include a short, tested explanation of normal operation, diagnosis and rollback. This documentation update does not refresh the October 2 live observations.
The combined prevention, recovery ownership and acceptance sequence is in [INTEGRATED_PLAN.md](INTEGRATED_PLAN.md). It pairs every finding with basic infrastructure changes and identifies conditional use of existing tools.
The tool-by-tool implementation and acceptance plan for "Make recovery dependable" is in [RECOVERY_TOOL_PLAN.md](RECOVERY_TOOL_PLAN.md). It assigns work to Soteria, Ananke, Metis, Ariadne/Hermes, node helpers and Flux, and distinguishes existing behavior from proposed integration.
| Wave | Work and acceptance gate | Estimate |
|---|---|---|
| 1: Make recovery possible | Current external-PG + token backup; independent restore proof; preserve recovery copies before Longhorn changes; explicit incident list | 1-2 days |
| 2: Stop recurring resource failures | Power/link/runtime-media repairs; PG resources; OpenSearch and two stuck volumes; Firefly diagnosis; CI/AI admission budgets | 2-4 days, overlapping physical work |
| 3: Make desired state predictable | Resolve IfNotPresent/live-Git differences; use native owners; retire unnecessary restart/GC overlap; concise node/service catalog | 2-4 days; re-estimate after basic fixes |
| 4: Give every service a continuity plan | Measured N+1 placement; redundant stateless foundations; stateful backup/failover targets; local job persistence; restore parked services within capacity | 1-3 weeks depending on HA choices |
| 5: Supported baseline and handoff | Staged supported upgrades; restore/failover/cold-start tests; alerts and one-page operator procedures | 3-7 working days plus soak |
Plan on **several days for substantial stabilization, roughly 2-4 weeks for a maintainable baseline**, and longer if new hardware or application HA redesign is needed. Require **at least 7-14 days of observed stability** before calling the recurring-failure problem resolved. These are planning estimates, not promises of uninterrupted availability or full completion within an elapsed window.
Each implementation change should have: a specific owner and dependency, a small Flux-tracked diff (or versioned host configuration), a maintenance impact, preconditions, a validation check, and a rollback/restore path. Reverting Git is not a substitute for reversing a database schema or storage migration. Test those recovery paths before cutover.
## What a maintainable cluster should look like
1. One inventory identifies each service, URL, owner, dependencies, node pool, storage, current backup, restore procedure and expected recovery time. Parked services are explicit.
2. Git describes the actual steady state. Recovery exceptions are scoped, visible and expiring. Host configuration is versioned alongside the supported hardware matrix.
3. One read-only health entry point distinguishes: node/link failure, capacity/scheduling failure, storage failure, application failure, and configuration drift. It links the relevant runbook and evidence.
4. Automated actions have ownership, a maintenance lock, a bounded retry count and a circuit breaker. An unexplained permanent restart loop is a fault, not a recovery strategy.
5. A daily health review verifies real availability, backup age and alert delivery. A monthly maintenance window handles one compatible upgrade group and a restore sample. No manual recurring cleanup is required for normal operation.
6. The operator can diagnose a routine outage and perform the documented approved recovery without knowing Kubernetes internals or searching this conversation.
## Completion criteria
- Every retained node is healthy or deliberately quarantined/retired with no required service depending on it; no ongoing undervoltage, unbounded recovery loops or unexplained Ready flaps.
- No sustained CPU/I/O/memory saturation or recurring resource OOM for normal approved workloads. Compatible N+1 spare capacity is demonstrated, not inferred from cluster totals.
- Every required service has a real health check and a defined node-failure behavior. Stateless services that promise continuity survive a controlled worker loss; singleton services have an agreed measured recovery target.
- All required data has a current authorized backup, an owner and a tested restore path. Local-only data remains local-only.
- Flux/Helm status reflects intended live state; old suspended migration errors and parked services are reported separately from incidents.
- Telemetry missingness is explicit. The owner receives actionable alerts even if primary cluster monitoring/mail is down, using an approved independent channel.
- The user can follow [OPERATOR_GUIDE.md](OPERATOR_GUIDE.md) and the eventual service runbooks to locate and recover a representative fault. The current guide contains diagnosis only; implementation-specific recovery commands should be added after those procedures are validated.
## Remaining unknowns requiring targeted follow-up
Physical causes on 05/06; exact power-supply/cable faults on 04/11; Titan-08 flap causes; recoverable contents of the two stalled volumes; Soteria backend-error mechanism; Firefly readiness cause; exact Java OOM cgroup on Titan-21; whether independent current backups exist elsewhere; switch/router/UPS load-test behavior; authenticated user workflows and mail delivery; per-service HA compatibility; restore and cold-start duration. These are not established root causes merely because nearby symptoms exist.
Raw operational snapshots are held in the private local audit directory `/tmp/atlas-cluster-audit-20261002`; durable curated inventories and evidence are alongside this report. No real suite inputs, generated suite outputs, credentials or application records were collected for the report.