| 3. Recovery coverage | In progress; cloud uploads blocked | External datastore restore and second LAN copy verified; 19 application DB dumps completed; scheduled application backup completed at 08:37 UTC; representative Gitea restore passed. Backblaze storage-cap failure established | Resolve the account storage cap within an agreed budget, complete eligible PVC coverage and Soteria restore, resolve Hermes tenant-2 safely, and provide approved local-only recovery copies |
| 4. Capacity and placement | Partial; CI waiting for compatible capacity | CI agent cap reduced to one; shared default, Metis, IaC and Data Prepper agent templates now exclude storage/quarantined workers. Shared PostgreSQL/OpenSearch resource corrections deployed; Vault injectors spread. Longhorn uses least-effort balancing and one rebuild per node | Restore compatible healthy worker capacity; inspect other projects' inline agent templates. Account for the measured 2-3.5 GiB storage-process footprint and calculate spare capacity for one worker loss |
| 5. Supported software | Planned; backups established | Installed K3s cohorts identified as 1.31.5 and 1.33.3; titan-db Ubuntu 24.10 is beyond support | Select a compatible supported baseline; stage OS/K3s/storage upgrades with rollback. Plan collation/index repair |
| 6. Maintainability and sustained health | In progress | Operator guide and rollback record published; Ananke stale-event loop fixed; updater failure reporting corrected; overlapping restart/eviction behavior removed | Complete the service recovery checklist, validate remaining automation boundaries, arrange controlled recovery checks, then observe 7-14 days |
| Nodes | Established evidence | Remaining action |
| --- | --- | --- |
| 12/13/19 | atlas/root passwords restored; root SSH keys preserved; password-backed sudo works after legacy grant retirement | Observe stability; use the corrected recovery image for future rebuilds |
| 20/21 | atlas sudo already worked; root passwords now match Vault | Continue workload memory diagnosis; Jetson-21 recorded a Data Prepper cgroup OOM |
| 04/11 | Fresh undervoltage messages continued during this audit, despite sufficient disk space | Onboard EXT5V samples were 4.82132 V (04) and 4.75968 V (11). Check delivered power/cables/peripherals; do not clear quarantine based on one sample |
| 14 | Runtime USB flash; 39-45% iowait in the first sample, quiet later, no new kernel I/O error in the two-hour sample; stale RAM-log hooks corrected | Preserve data; investigate intermittent storage stalls and validate sustained operation before return; hardware failure is not proven by I/O pressure alone |
| 18 | Quarantine has reduced current I/O pressure; its 50 MiB RAM-log filesystem is still active | Keep quarantined until sustained storage checks pass; do not apply the inactive-RAM-log repair script to it |
| 05 | Answers LAN ping at .31; expected SSH, kubelet and common recovery ports refuse connections | Console/user-space service check; this is not evidence that the machine is powered off |
| 09/10 | Answer LAN ping and SSH on port 22; recorded 2277 port is closed; host keys differ from saved keys | Confirm reinstall/image identity before supplying passwords; no host-key checks bypassed |
| 06/16 | No LAN ping response; neighbor resolution incomplete; SSH unavailable | Power, link and boot-media checks needed |
The 12 unlocked root passwords match Vault. Nine machines retain locked direct
root accounts (0a/0b/0c/04/07/08/11/db/jh); their administrator account plus sudo
still provides full root control. Root-account lock state is recorded explicitly
in Vault instead of implying that every root_password field is a usable login.
| Native external PostgreSQL backup and second LAN copy | A 54,083,817-byte custom dump restored into a disposable PostgreSQL 16 instance in 11 seconds with 22,350 kine rows; no TCP listener. Both timers enabled; checksum and matching-token checks pass | Disable the two timers; retain all recovery bundles |
| Retired unconditional k3s-agent restart DaemonSet | Flux preview removed only that DaemonSet; reconciliation completed and the DaemonSet is absent | Revert f04c84ee, understanding that returning the DaemonSet restarts agents |
| Recovered metrics Pushgateway | Existing volume attached healthy on titan-19; replacement pod Ready. No volume deletion or data replacement | Revert the focused monitoring placement commits after the original host is repaired |
| Removed eviction for historical restart counts | Descheduler manifest validates and Flux applied the change | Revert fa9251ae |
| Restored GitOps UI Deployment | Helm drift correction recreated weave-gitops; Deployment 1/1 Ready | Revert 6c9398ea to disable ongoing drift correction; this does not remove recovered resources |
| Aligned Flux definition ownership | Removed creation-only policy and adopted already-active service state; no service paths or source refs changed | Revert the focused Flux commits; review suspension fields before doing so |
| Backup freshness alerts | Both datastore-copy timestamps are scraped; native Grafana provisioning reload returned HTTP 200 | Revert the focused monitoring change; backups continue independently |
| Quarantined stalled runtime media | Titan-14 and Titan-18 are SchedulingDisabled via `infrastructure/core/node-maintenance.yaml`; no forced storage detach | After repair, change `unschedulable` to false in Git and verify before removing the prune-disabled Node declaration |
| Bounded CI concurrency | Jenkins controller Ready with runtime cap one (8a6d3db1); inspected pipeline placements exclude storage workers | Revert the focused change and reconcile Jenkins; allow a controller restart and account for renewed storage load |
| Spread Vault injectors | Two healthy replicas on Titan-08 and Titan-22; native anti-affinity and minAvailable=1 PDB | Revert 5f3f3184 |
| Protected application databases | Nineteen logical dumps and globals completed with checksums on Titan-0b; isolated Gitea restore passed in six seconds with 111 tables | Keep copies; stop the daily timer if needed |
| Application PostgreSQL resource/probe repair | Pinned current PostgreSQL 15.18 image, 1 GiB request / 2 GiB limit, native startup/readiness; 2/2 Ready on Titan-17 and healthy attached volume | Revert 103e2064 only after checking destination capacity; another database restart is required |
| Recovered Firefly and OpenSearch | Firefly Ready on Titan-08; OpenSearch Ready with its existing healthy volume on Titan-22. Pi capacity could not fit OpenSearch, so the documented CPU fallback was used; no GPU reassignment | Review capacity before placement rollback; preserve existing volumes |
| Protected local-only data policy | Hermes PVCs excluded from Soteria's cloud policy through Flux | Do not broaden this policy without reviewing data authorization |
| Stopped stale credential-event restart loop | Ananke c73b8fe installed on Titan-db after its complete quality gate; helper generation 796 remained unchanged after normal one-minute checks resumed. Peer 53ad03a no longer runs periodic cluster repairs | Native installer retains `/usr/local/lib/ananke/rollback/ananke.previous`; rolling back restores the defective behavior, so retain a bounded repair interval if needed |
| Disabled premature CI cancellation | Ariadne's elapsed-only cutoff is disabled with value zero (d47e7b8e), verified in the running application | Reverting resumes age-only cancellations; require progress-aware detection before re-enabling |
| Scoped Soteria permissions | Atlas 46596536 applied; intended state updates and credential reads allowed, unrelated secret reads/listing and Hermes job creation denied; policy/usage fingerprints unchanged | Revert only after reviewing why broader access is required; retain state Secrets |
| Recovered Monerod mount | Snapshot df99cdac, clean detach b119076c, restart 0984164f; existing ext4 filesystem mounted on Titan-19, LMDB opened and RPC initialized | Preserve the snapshot and PVC. Returning to Titan-08 requires storage-path validation; do not replay the temporary scale-zero commit as a normal rollback |
| Bounded Longhorn background I/O | b0be99b7 applied; live settings are least-effort balancing and one concurrent rebuild per node | Revert b0be99b7 and reconcile Longhorn; the versioned settings Job restores best-effort and two rebuilds per node |
| Problem | Evidence / implication | Next action |
| --- | --- | --- |
| titan-05 / titan-06 offline | Nineteen of twenty-one nodes Ready at the start of implementation | Physical power/network check; recover or replace; no remote software claim of repair |
| Pi power instability | Fresh undervoltage observed on titan-04 and titan-11 during audit; titan-04 remains cordoned | Check supplies, cables and USB power demand; keep quarantine until a stability check passes |
| titan-14 runtime I/O | USB flash runtime disk at sustained saturation during audit | Replace/migrate to suitable SSD after data preservation and a controlled drain |
| titan-18 storage stalls | I/O some pressure near 99%, load near 39; Longhorn engine image fails to become ready | Reduce CI load, inspect storage/kernel health, move runtime off inadequate media if confirmed |
| titan-08 stale iSCSI session | Pushgateway engine could not log out an obsolete target; replica data remained usable on another host | Repair during a controlled storage maintenance window; do not mass-restart instance managers holding healthy volumes |
| Hermes tenant-2 shared workspace | Existing faulted/detached volume, repeated recovery attempts | Preserve replica/recovery evidence and perform component-supported repair; no source content in routine logs |
| Backup coverage and restore proof | Old Soteria coverage is insufficient; repairing its scheduler does not instantly create all backups | Confirm eligible data, completed backup objects, achievable schedule and representative restores service by service |
| Database collation drift | Native dumps reported stored collation 2.36 versus runtime 2.41 for several databases | Plan index rebuilds with compatible locale settings before refreshing version metadata; do not merely suppress warnings |
| Supported software baseline | Ubuntu 24.10 on titan-db and older Kubernetes cohorts are out of support | Backed-up, staged host/K3s/Longhorn upgrades, one compatible cohort at a time |
| Failure capacity | Pi pool is heavily reserved; unused x86 capacity has explicit simulation/GPU roles | Recalculate compatible N+1 capacity after repairs; do not silently take reserved GPUs or simulation capacity |
## Completion standard
Do not call the entire cluster fixed based on one green snapshot. Open storage
faults, physical power/media problems, unsupported host software and per-service
restore coverage remain visible until addressed. Validate application behavior,
backup freshness, actual restore results, and 7-14 days without the recurring
failures. No automatic real-suite inference jobs or outage drills are part of