5.7 KiB
Cluster stabilization - implementation record
This record begins on 2026-10-03 UTC, following authorization to implement the configuration-first plan. The October 2 audit remains a historical baseline, not a current health claim. Use Cluster operations for the ordinary operator path.
Verified changes
| Change | Evidence | Rollback |
|---|---|---|
| Native external PostgreSQL backup and second LAN copy | A 54,083,817-byte custom dump restored into a disposable PostgreSQL 16 instance in 11 seconds with 22,350 kine rows; no TCP listener. Both timers enabled; checksum and matching-token checks pass | Disable the two timers; retain all recovery bundles |
| Retired unconditional k3s-agent restart DaemonSet | Flux preview removed only that DaemonSet; reconciliation completed and the DaemonSet is absent | Revert f04c84ee, understanding that returning the DaemonSet restarts agents |
| Recovered metrics Pushgateway | Existing volume attached healthy on titan-19; replacement pod Ready. No volume deletion or data replacement | Revert the focused monitoring placement commits after the original host is repaired |
| Removed eviction for historical restart counts | Descheduler manifest validates and Flux applied the change | Revert fa9251ae |
| Restored GitOps UI Deployment | Helm drift correction recreated weave-gitops; Deployment 1/1 Ready | Revert 6c9398ea to disable ongoing drift correction; this does not remove recovered resources |
| Aligned Flux definition ownership | Removed creation-only policy and adopted already-active service state; no service paths or source refs changed | Revert the focused Flux commits; review suspension fields before doing so |
| Protected local-only data policy | Hermes PVCs excluded from Soteria's cloud policy through Flux | Do not broaden this policy without reviewing data authorization |
Host backup implementation is in commit 80aff498 and
its operational notes.
The restore check validates database/schema restoration without global-role/ACL
replay or starting K3s. A complete control-plane disaster drill remains outstanding.
Work being completed
- OpenSearch: preserve the existing disk and 2 GiB heap; request 3 GiB memory
instead of 768 MiB, remove the offline titan-05 pin, and use the chart's actual
nodeAffinitysetting. Pi5 then Pi4 are preferred; titan-22 is the explicit last-resort CPU destination. The original Helm operation must converge before the corrected generation can be verified. - Shared application PostgreSQL: explicit 1 GiB request / 2 GiB limit, startup and readiness checks, bounded exporter resources. Deployment waits for the pre-change database recovery copies to finish.
- Soteria: source fix 7018d4c in the Soteria repository allows live RWO Longhorn snapshots while preserving the restic mount guard. Manual Longhorn requests also enforce exclusions. All Go package tests pass; normal image publication and a representative backup verification remain required.
- Jenkins: a two-agent concurrency cap is prepared. Its existing configuration hash triggers a controller rollout, so apply it with ongoing builds accounted for rather than pretending it is a harmless live-only reload.
Confirmed problems needing further work
| Problem | Evidence / implication | Next action |
|---|---|---|
| titan-05 / titan-06 offline | Nineteen of twenty-one nodes Ready at the start of implementation | Physical power/network check; recover or replace; no remote software claim of repair |
| Pi power instability | Fresh undervoltage observed on titan-04 and titan-11 during audit; titan-04 remains cordoned | Check supplies, cables and USB power demand; keep quarantine until a stability check passes |
| titan-14 runtime I/O | USB flash runtime disk at sustained saturation during audit | Replace/migrate to suitable SSD after data preservation and a controlled drain |
| titan-18 storage stalls | I/O some pressure near 99%, load near 39; Longhorn engine image fails to become ready | Reduce CI load, inspect storage/kernel health, move runtime off inadequate media if confirmed |
| titan-08 stale iSCSI session | Pushgateway engine could not log out an obsolete target; replica data remained usable on another host | Repair during a controlled storage maintenance window; do not mass-restart instance managers holding healthy volumes |
| Hermes tenant-2 shared workspace | Existing faulted/detached volume, repeated recovery attempts | Preserve replica/recovery evidence and perform component-supported repair; no source content in routine logs |
| Backup coverage and restore proof | Old Soteria coverage is insufficient; repairing its scheduler does not instantly create all backups | Confirm eligible data, completed backup objects, achievable schedule and representative restores service by service |
| Supported software baseline | Ubuntu 24.10 on titan-db and older Kubernetes cohorts are out of support | Backed-up, staged host/K3s/Longhorn upgrades, one compatible cohort at a time |
| Failure capacity | Pi pool is heavily reserved; unused x86 capacity has explicit simulation/GPU roles | Recalculate compatible N+1 capacity after repairs; do not silently take reserved GPUs or simulation capacity |
Completion standard
Do not call the entire cluster fixed based on one green snapshot. Open storage faults, physical power/media problems, unsupported host software and per-service restore coverage remain visible until addressed. Validate application behavior, backup freshness, actual restore results, and 7-14 days without the recurring failures. No automatic real-suite inference jobs or outage drills are part of this maintenance work.