atlas-iac/docs/CLUSTER_STABILIZATION.md

94 lines
8.5 KiB
Markdown

# Cluster stabilization - implementation record
This record begins on 2026-10-03 UTC, following authorization to implement the
[configuration-first plan](cluster-audit-20261002/INTEGRATED_PLAN.md).
The October 2 audit remains a historical baseline, not a current health claim.
Use [Cluster operations](CLUSTER_OPERATIONS.md) for the ordinary operator path.
## Verified changes
| Change | Evidence | Rollback |
| --- | --- | --- |
| Native external PostgreSQL backup and second LAN copy | A 54,083,817-byte custom dump restored into a disposable PostgreSQL 16 instance in 11 seconds with 22,350 kine rows; no TCP listener. Both timers enabled; checksum and matching-token checks pass | Disable the two timers; retain all recovery bundles |
| Retired unconditional k3s-agent restart DaemonSet | Flux preview removed only that DaemonSet; reconciliation completed and the DaemonSet is absent | Revert f04c84ee, understanding that returning the DaemonSet restarts agents |
| Recovered metrics Pushgateway | Existing volume attached healthy on titan-19; replacement pod Ready. No volume deletion or data replacement | Revert the focused monitoring placement commits after the original host is repaired |
| Removed eviction for historical restart counts | Descheduler manifest validates and Flux applied the change | Revert fa9251ae |
| Restored GitOps UI Deployment | Helm drift correction recreated weave-gitops; Deployment 1/1 Ready | Revert 6c9398ea to disable ongoing drift correction; this does not remove recovered resources |
| Aligned Flux definition ownership | Removed creation-only policy and adopted already-active service state; no service paths or source refs changed | Revert the focused Flux commits; review suspension fields before doing so |
| Backup freshness alerts | Both datastore-copy timestamps are scraped; native Grafana provisioning reload returned HTTP 200 | Revert the focused monitoring change; backups continue independently |
| Quarantined stalled runtime media | Titan-14 and Titan-18 are SchedulingDisabled via `infrastructure/core/node-maintenance.yaml`; no forced storage detach | After repair, change `unschedulable` to false in Git and verify before removing the prune-disabled Node declaration |
| Bounded CI concurrency | Jenkins controller Ready after applying the two-agent cap; existing agents reconnected | Revert 440244f3 and reconcile Jenkins; allow a controller restart |
| Spread Vault injectors | Two healthy replicas on Titan-08 and Titan-22; native anti-affinity and minAvailable=1 PDB | Revert 5f3f3184 |
| Protected application databases | Nineteen logical dumps and globals completed with checksums on Titan-0b; isolated Gitea restore passed in six seconds with 111 tables | Keep copies; stop the daily timer if needed |
| Application PostgreSQL resource/probe repair | Pinned current PostgreSQL 15.18 image, 1 GiB request / 2 GiB limit, native startup/readiness; 2/2 Ready on Titan-17 and healthy attached volume | Revert 103e2064 only after checking destination capacity; another database restart is required |
| Protected local-only data policy | Hermes PVCs excluded from Soteria's cloud policy through Flux | Do not broaden this policy without reviewing data authorization |
Host backup implementation is in commit 80aff498 and
[its operational notes](../infrastructure/host-backup/NOTES.md).
The restore check validates database/schema restoration without global-role/ACL
replay or starting K3s. A complete control-plane disaster drill remains outstanding.
## Work being completed
* OpenSearch: preserve the existing disk and 2 GiB heap; request 3 GiB memory
instead of 768 MiB, remove the offline titan-05 pin, and use the chart's actual
`nodeAffinity` setting. Pi5 then Pi4 are preferred; titan-22 is the explicit
last-resort CPU destination. The original Helm operation must converge before
the corrected generation can be verified.
* Soteria: source fix 7018d4c in the Soteria repository allows live RWO Longhorn
snapshots while preserving the restic mount guard. Manual Longhorn requests
also enforce exclusions. All Go package tests pass; normal image publication
and a representative backup verification remain required.
* Native application-backup schedule and its freshness alert are being installed
after the initial successful copy and restore check.
* Soteria CI: the coverage report was generated after Sonar analysis, causing a
false zero-coverage gate despite 96.4% measured test coverage. Commit f167d7d
orders tests before analysis and awaits the Sonar gate; Jenkins declarative
validation passed. Release build remains pending; do not claim the runtime
backup defect is repaired until the new image and a real eligible backup pass.
## Confirmed problems needing further work
| Problem | Evidence / implication | Next action |
| --- | --- | --- |
| titan-05 / titan-06 offline | Nineteen of twenty-one nodes Ready at the start of implementation | Physical power/network check; recover or replace; no remote software claim of repair |
| Pi power instability | Fresh undervoltage observed on titan-04 and titan-11 during audit; titan-04 remains cordoned | Check supplies, cables and USB power demand; keep quarantine until a stability check passes |
| titan-14 runtime I/O | USB flash runtime disk at sustained saturation during audit | Replace/migrate to suitable SSD after data preservation and a controlled drain |
| titan-18 storage stalls | I/O some pressure near 99%, load near 39; Longhorn engine image fails to become ready | Reduce CI load, inspect storage/kernel health, move runtime off inadequate media if confirmed |
| titan-08 stale iSCSI session | Pushgateway engine could not log out an obsolete target; replica data remained usable on another host | Repair during a controlled storage maintenance window; do not mass-restart instance managers holding healthy volumes |
| Hermes tenant-2 shared workspace | Existing faulted/detached volume, repeated recovery attempts | Preserve replica/recovery evidence and perform component-supported repair; no source content in routine logs |
| Backup coverage and restore proof | Old Soteria coverage is insufficient; repairing its scheduler does not instantly create all backups | Confirm eligible data, completed backup objects, achievable schedule and representative restores service by service |
| Database collation drift | Native dumps reported stored collation 2.36 versus runtime 2.41 for several databases | Plan index rebuilds with compatible locale settings before refreshing version metadata; do not merely suppress warnings |
| Supported software baseline | Ubuntu 24.10 on titan-db and older Kubernetes cohorts are out of support | Backed-up, staged host/K3s/Longhorn upgrades, one compatible cohort at a time |
| Failure capacity | Pi pool is heavily reserved; unused x86 capacity has explicit simulation/GPU roles | Recalculate compatible N+1 capacity after repairs; do not silently take reserved GPUs or simulation capacity |
## Completion standard
Do not call the entire cluster fixed based on one green snapshot. Open storage
faults, physical power/media problems, unsupported host software and per-service
restore coverage remain visible until addressed. Validate application behavior,
backup freshness, actual restore results, and 7-14 days without the recurring
failures. No automatic real-suite inference jobs or outage drills are part of
this maintenance work.
## Operational lessons from this maintenance
The application database move took approximately nine minutes because both
pinned images had to be pulled on the destination. Its volume attached correctly;
PostgreSQL and its exporter became Ready afterward. Preload required images on
an eligible destination before another planned database move. The larger database
dumps also completed much more efficiently through native PostgreSQL than through
`kubectl exec` output streaming. The native recovery script records this path.
At 06:07 UTC, both Titan-04 and Titan-11 emitted fresh undervoltage messages.
Their instantaneous Pi 5 input samples were 4.9245 V and 4.89904 V; these are not
measurements of the transient minimum. Both had `get_throttled=0x50000` between
events. The fault remains active even when that individual sample reports only
historical bits. The user is checking power/boot-media issues physically.
Titan-18 remains a blocker: Firefly and OpenSearch pods are waiting for graceful
termination on its stalled runtime. Do not force detach their mounted volumes
while the old host may still have writers. Its replacement/repair is distinct
from Kubernetes scheduling; a cordon alone cannot repair hung I/O.