atlas-iac/docs/CLUSTER_STABILIZATION.md

65 lines
5.7 KiB
Markdown
Raw Normal View History

# Cluster stabilization - implementation record
This record begins on 2026-10-03 UTC, following authorization to implement the
[configuration-first plan](cluster-audit-20261002/INTEGRATED_PLAN.md).
The October 2 audit remains a historical baseline, not a current health claim.
Use [Cluster operations](CLUSTER_OPERATIONS.md) for the ordinary operator path.
## Verified changes
| Change | Evidence | Rollback |
| --- | --- | --- |
| Native external PostgreSQL backup and second LAN copy | A 54,083,817-byte custom dump restored into a disposable PostgreSQL 16 instance in 11 seconds with 22,350 kine rows; no TCP listener. Both timers enabled; checksum and matching-token checks pass | Disable the two timers; retain all recovery bundles |
| Retired unconditional k3s-agent restart DaemonSet | Flux preview removed only that DaemonSet; reconciliation completed and the DaemonSet is absent | Revert f04c84ee, understanding that returning the DaemonSet restarts agents |
| Recovered metrics Pushgateway | Existing volume attached healthy on titan-19; replacement pod Ready. No volume deletion or data replacement | Revert the focused monitoring placement commits after the original host is repaired |
| Removed eviction for historical restart counts | Descheduler manifest validates and Flux applied the change | Revert fa9251ae |
| Restored GitOps UI Deployment | Helm drift correction recreated weave-gitops; Deployment 1/1 Ready | Revert 6c9398ea to disable ongoing drift correction; this does not remove recovered resources |
| Aligned Flux definition ownership | Removed creation-only policy and adopted already-active service state; no service paths or source refs changed | Revert the focused Flux commits; review suspension fields before doing so |
| Protected local-only data policy | Hermes PVCs excluded from Soteria's cloud policy through Flux | Do not broaden this policy without reviewing data authorization |
Host backup implementation is in commit 80aff498 and
[its operational notes](../infrastructure/host-backup/NOTES.md).
The restore check validates database/schema restoration without global-role/ACL
replay or starting K3s. A complete control-plane disaster drill remains outstanding.
## Work being completed
* OpenSearch: preserve the existing disk and 2 GiB heap; request 3 GiB memory
instead of 768 MiB, remove the offline titan-05 pin, and use the chart's actual
`nodeAffinity` setting. Pi5 then Pi4 are preferred; titan-22 is the explicit
last-resort CPU destination. The original Helm operation must converge before
the corrected generation can be verified.
* Shared application PostgreSQL: explicit 1 GiB request / 2 GiB limit, startup
and readiness checks, bounded exporter resources. Deployment waits for the
pre-change database recovery copies to finish.
* Soteria: source fix 7018d4c in the Soteria repository allows live RWO Longhorn
snapshots while preserving the restic mount guard. Manual Longhorn requests
also enforce exclusions. All Go package tests pass; normal image publication
and a representative backup verification remain required.
* Jenkins: a two-agent concurrency cap is prepared. Its existing configuration
hash triggers a controller rollout, so apply it with ongoing builds accounted
for rather than pretending it is a harmless live-only reload.
## Confirmed problems needing further work
| Problem | Evidence / implication | Next action |
| --- | --- | --- |
| titan-05 / titan-06 offline | Nineteen of twenty-one nodes Ready at the start of implementation | Physical power/network check; recover or replace; no remote software claim of repair |
| Pi power instability | Fresh undervoltage observed on titan-04 and titan-11 during audit; titan-04 remains cordoned | Check supplies, cables and USB power demand; keep quarantine until a stability check passes |
| titan-14 runtime I/O | USB flash runtime disk at sustained saturation during audit | Replace/migrate to suitable SSD after data preservation and a controlled drain |
| titan-18 storage stalls | I/O some pressure near 99%, load near 39; Longhorn engine image fails to become ready | Reduce CI load, inspect storage/kernel health, move runtime off inadequate media if confirmed |
| titan-08 stale iSCSI session | Pushgateway engine could not log out an obsolete target; replica data remained usable on another host | Repair during a controlled storage maintenance window; do not mass-restart instance managers holding healthy volumes |
| Hermes tenant-2 shared workspace | Existing faulted/detached volume, repeated recovery attempts | Preserve replica/recovery evidence and perform component-supported repair; no source content in routine logs |
| Backup coverage and restore proof | Old Soteria coverage is insufficient; repairing its scheduler does not instantly create all backups | Confirm eligible data, completed backup objects, achievable schedule and representative restores service by service |
| Supported software baseline | Ubuntu 24.10 on titan-db and older Kubernetes cohorts are out of support | Backed-up, staged host/K3s/Longhorn upgrades, one compatible cohort at a time |
| Failure capacity | Pi pool is heavily reserved; unused x86 capacity has explicit simulation/GPU roles | Recalculate compatible N+1 capacity after repairs; do not silently take reserved GPUs or simulation capacity |
## Completion standard
Do not call the entire cluster fixed based on one green snapshot. Open storage
faults, physical power/media problems, unsupported host software and per-service
restore coverage remain visible until addressed. Validate application behavior,
backup freshness, actual restore results, and 7-14 days without the recurring
failures. No automatic real-suite inference jobs or outage drills are part of
this maintenance work.