atlas-iac/docs/CLUSTER_STABILIZATION.md

5.7 KiB

Cluster stabilization - implementation record

This record begins on 2026-10-03 UTC, following authorization to implement the configuration-first plan. The October 2 audit remains a historical baseline, not a current health claim. Use Cluster operations for the ordinary operator path.

Verified changes

Change Evidence Rollback
Native external PostgreSQL backup and second LAN copy A 54,083,817-byte custom dump restored into a disposable PostgreSQL 16 instance in 11 seconds with 22,350 kine rows; no TCP listener. Both timers enabled; checksum and matching-token checks pass Disable the two timers; retain all recovery bundles
Retired unconditional k3s-agent restart DaemonSet Flux preview removed only that DaemonSet; reconciliation completed and the DaemonSet is absent Revert f04c84ee, understanding that returning the DaemonSet restarts agents
Recovered metrics Pushgateway Existing volume attached healthy on titan-19; replacement pod Ready. No volume deletion or data replacement Revert the focused monitoring placement commits after the original host is repaired
Removed eviction for historical restart counts Descheduler manifest validates and Flux applied the change Revert fa9251ae
Restored GitOps UI Deployment Helm drift correction recreated weave-gitops; Deployment 1/1 Ready Revert 6c9398ea to disable ongoing drift correction; this does not remove recovered resources
Aligned Flux definition ownership Removed creation-only policy and adopted already-active service state; no service paths or source refs changed Revert the focused Flux commits; review suspension fields before doing so
Protected local-only data policy Hermes PVCs excluded from Soteria's cloud policy through Flux Do not broaden this policy without reviewing data authorization

Host backup implementation is in commit 80aff498 and its operational notes. The restore check validates database/schema restoration without global-role/ACL replay or starting K3s. A complete control-plane disaster drill remains outstanding.

Work being completed

  • OpenSearch: preserve the existing disk and 2 GiB heap; request 3 GiB memory instead of 768 MiB, remove the offline titan-05 pin, and use the chart's actual nodeAffinity setting. Pi5 then Pi4 are preferred; titan-22 is the explicit last-resort CPU destination. The original Helm operation must converge before the corrected generation can be verified.
  • Shared application PostgreSQL: explicit 1 GiB request / 2 GiB limit, startup and readiness checks, bounded exporter resources. Deployment waits for the pre-change database recovery copies to finish.
  • Soteria: source fix 7018d4c in the Soteria repository allows live RWO Longhorn snapshots while preserving the restic mount guard. Manual Longhorn requests also enforce exclusions. All Go package tests pass; normal image publication and a representative backup verification remain required.
  • Jenkins: a two-agent concurrency cap is prepared. Its existing configuration hash triggers a controller rollout, so apply it with ongoing builds accounted for rather than pretending it is a harmless live-only reload.

Confirmed problems needing further work

Problem Evidence / implication Next action
titan-05 / titan-06 offline Nineteen of twenty-one nodes Ready at the start of implementation Physical power/network check; recover or replace; no remote software claim of repair
Pi power instability Fresh undervoltage observed on titan-04 and titan-11 during audit; titan-04 remains cordoned Check supplies, cables and USB power demand; keep quarantine until a stability check passes
titan-14 runtime I/O USB flash runtime disk at sustained saturation during audit Replace/migrate to suitable SSD after data preservation and a controlled drain
titan-18 storage stalls I/O some pressure near 99%, load near 39; Longhorn engine image fails to become ready Reduce CI load, inspect storage/kernel health, move runtime off inadequate media if confirmed
titan-08 stale iSCSI session Pushgateway engine could not log out an obsolete target; replica data remained usable on another host Repair during a controlled storage maintenance window; do not mass-restart instance managers holding healthy volumes
Hermes tenant-2 shared workspace Existing faulted/detached volume, repeated recovery attempts Preserve replica/recovery evidence and perform component-supported repair; no source content in routine logs
Backup coverage and restore proof Old Soteria coverage is insufficient; repairing its scheduler does not instantly create all backups Confirm eligible data, completed backup objects, achievable schedule and representative restores service by service
Supported software baseline Ubuntu 24.10 on titan-db and older Kubernetes cohorts are out of support Backed-up, staged host/K3s/Longhorn upgrades, one compatible cohort at a time
Failure capacity Pi pool is heavily reserved; unused x86 capacity has explicit simulation/GPU roles Recalculate compatible N+1 capacity after repairs; do not silently take reserved GPUs or simulation capacity

Completion standard

Do not call the entire cluster fixed based on one green snapshot. Open storage faults, physical power/media problems, unsupported host software and per-service restore coverage remain visible until addressed. Validate application behavior, backup freshness, actual restore results, and 7-14 days without the recurring failures. No automatic real-suite inference jobs or outage drills are part of this maintenance work.