atlas-iac/docs/CLUSTER_STABILIZATION.md
2026-10-03 02:51:51 -05:00

12 KiB

Cluster stabilization - implementation record

This record begins on 2026-10-03 UTC, following authorization to implement the configuration-first plan. The October 2 audit remains a historical baseline, not a current health claim. Use Cluster operations for the ordinary operator path.

Verified changes

Change Evidence Rollback
Native external PostgreSQL backup and second LAN copy A 54,083,817-byte custom dump restored into a disposable PostgreSQL 16 instance in 11 seconds with 22,350 kine rows; no TCP listener. Both timers enabled; checksum and matching-token checks pass Disable the two timers; retain all recovery bundles
Retired unconditional k3s-agent restart DaemonSet Flux preview removed only that DaemonSet; reconciliation completed and the DaemonSet is absent Revert f04c84ee, understanding that returning the DaemonSet restarts agents
Recovered metrics Pushgateway Existing volume attached healthy on titan-19; replacement pod Ready. No volume deletion or data replacement Revert the focused monitoring placement commits after the original host is repaired
Removed eviction for historical restart counts Descheduler manifest validates and Flux applied the change Revert fa9251ae
Restored GitOps UI Deployment Helm drift correction recreated weave-gitops; Deployment 1/1 Ready Revert 6c9398ea to disable ongoing drift correction; this does not remove recovered resources
Aligned Flux definition ownership Removed creation-only policy and adopted already-active service state; no service paths or source refs changed Revert the focused Flux commits; review suspension fields before doing so
Backup freshness alerts Both datastore-copy timestamps are scraped; native Grafana provisioning reload returned HTTP 200 Revert the focused monitoring change; backups continue independently
Quarantined stalled runtime media Titan-14 and Titan-18 are SchedulingDisabled via infrastructure/core/node-maintenance.yaml; no forced storage detach After repair, change unschedulable to false in Git and verify before removing the prune-disabled Node declaration
Bounded CI concurrency Jenkins controller Ready after applying the two-agent cap; existing agents reconnected Revert 440244f3 and reconcile Jenkins; allow a controller restart
Spread Vault injectors Two healthy replicas on Titan-08 and Titan-22; native anti-affinity and minAvailable=1 PDB Revert 5f3f3184
Protected application databases Nineteen logical dumps and globals completed with checksums on Titan-0b; isolated Gitea restore passed in six seconds with 111 tables Keep copies; stop the daily timer if needed
Application PostgreSQL resource/probe repair Pinned current PostgreSQL 15.18 image, 1 GiB request / 2 GiB limit, native startup/readiness; 2/2 Ready on Titan-17 and healthy attached volume Revert 103e2064 only after checking destination capacity; another database restart is required
Recovered Firefly and OpenSearch Firefly Ready on Titan-08; OpenSearch Ready with its existing healthy volume on Titan-22. Pi capacity could not fit OpenSearch, so the documented CPU fallback was used; no GPU reassignment Review capacity before placement rollback; preserve existing volumes
Protected local-only data policy Hermes PVCs excluded from Soteria's cloud policy through Flux Do not broaden this policy without reviewing data authorization
Stopped stale credential-event restart loop Ananke c73b8fe installed on Titan-db after its complete quality gate; helper generation 796 remained unchanged after normal one-minute checks resumed. Peer 53ad03a no longer runs periodic cluster repairs Native installer retains /usr/local/lib/ananke/rollback/ananke.previous; rolling back restores the defective behavior, so retain a bounded repair interval if needed
Corrected premature CI cancellation Ariadne's configured elapsed-time cutoff is now 120 minutes, verified in the running Deployment Revert 0180ba66; this returns the 45-minute cutoff that aborted an active Pi build

Host backup implementation is in commit 80aff498 and its operational notes. The restore check validates database/schema restoration without global-role/ACL replay or starting K3s. A complete control-plane disaster drill remains outstanding.

Work being completed

  • Soteria: source fix 7018d4c in the Soteria repository allows live RWO Longhorn snapshots while preserving the restic mount guard. Manual Longhorn requests also enforce exclusions. All Go package tests pass; normal image publication and a representative backup verification remain required.
  • Native application backups now run daily from Titan-0b. The timer is enabled, the verified initial bundle timestamp is scraped, and the 36-hour freshness alert is provisioned. The first scheduled run remains to be observed.
  • Soteria CI: the coverage report was generated after Sonar analysis, causing a false zero-coverage gate despite 96.4% measured test coverage. Commit f167d7d orders tests before analysis and awaits the Sonar gate; Jenkins declarative validation passed. Release build 1292 failed during Git checkout with a 404; current credential/repository checks and build 1293's source checkout succeed. The release retry remains queued; do not claim the runtime backup defect is repaired until the new image and a real eligible backup pass.

Confirmed problems needing further work

Problem Evidence / implication Next action
titan-05 / titan-06 offline Nineteen of twenty-one nodes Ready at the start of implementation Physical power/network check; recover or replace; no remote software claim of repair
Pi power instability Fresh undervoltage observed on titan-04 and titan-11 during audit; titan-04 remains cordoned Check supplies, cables and USB power demand; keep quarantine until a stability check passes
titan-14 runtime I/O USB flash runtime disk at sustained saturation during audit Replace/migrate to suitable SSD after data preservation and a controlled drain
titan-18 storage stalls I/O some pressure near 99%, load near 39; Longhorn engine image fails to become ready Reduce CI load, inspect storage/kernel health, move runtime off inadequate media if confirmed
titan-08 stale iSCSI session Pushgateway engine could not log out an obsolete target; replica data remained usable on another host Repair during a controlled storage maintenance window; do not mass-restart instance managers holding healthy volumes
Hermes tenant-2 shared workspace Existing faulted/detached volume, repeated recovery attempts Preserve replica/recovery evidence and perform component-supported repair; no source content in routine logs
Backup coverage and restore proof Old Soteria coverage is insufficient; repairing its scheduler does not instantly create all backups Confirm eligible data, completed backup objects, achievable schedule and representative restores service by service
Database collation drift Native dumps reported stored collation 2.36 versus runtime 2.41 for several databases Plan index rebuilds with compatible locale settings before refreshing version metadata; do not merely suppress warnings
Supported software baseline Ubuntu 24.10 on titan-db and older Kubernetes cohorts are out of support Backed-up, staged host/K3s/Longhorn upgrades, one compatible cohort at a time
Failure capacity Pi pool is heavily reserved; unused x86 capacity has explicit simulation/GPU roles Recalculate compatible N+1 capacity after repairs; do not silently take reserved GPUs or simulation capacity

Completion standard

Do not call the entire cluster fixed based on one green snapshot. Open storage faults, physical power/media problems, unsupported host software and per-service restore coverage remain visible until addressed. Validate application behavior, backup freshness, actual restore results, and 7-14 days without the recurring failures. No automatic real-suite inference jobs or outage drills are part of this maintenance work.

Operational lessons from this maintenance

The application database move took approximately nine minutes because both pinned images had to be pulled on the destination. Its volume attached correctly; PostgreSQL and its exporter became Ready afterward. Preload required images on an eligible destination before another planned database move. The larger database dumps also completed much more efficiently through native PostgreSQL than through kubectl exec output streaming. The native recovery script records this path.

At 06:07 UTC, both Titan-04 and Titan-11 emitted fresh undervoltage messages. Their instantaneous Pi 5 input samples were 4.9245 V and 4.89904 V; these are not measurements of the transient minimum. Both had get_throttled=0x50000 between events. The fault remains active even when that individual sample reports only historical bits. The user is checking power/boot-media issues physically.

Firefly recovered on Titan-08 at approximately 06:52 UTC after its old pod terminated and its volume moved normally. OpenSearch recovered on Titan-22 around 07:03 UTC after the corrected placement was applied. Commit 2f043d10 temporarily allowed the old offline-node rollback to finish. Normal rollback readiness was restored after the corrected release became Ready. Titan-18 remains cordoned; rescheduling does not repair its runtime media.

Ananke was confirmed restarting jenkins-vault-sync approximately every minute from retained credential-warning events, including after the original failure cleared. The Ananke fixes require a current matching pod UID, active image-pull failure and warning age below ten minutes, with a persisted 30-minute cooldown per helper. Only the coordinator runs periodic cluster repair; the peer retains UPS monitoring and forwarding. Coordinator revision c73b8fe passed the native host's complete quality gate, including the per-file 95% coverage check, and installed successfully. Its temporary one-hour interval was restored to the original 60 seconds at 07:41 UTC. The peer runs 53ad03a, which includes the credential guard and coordinator-only repair. UPS services remain active; the coordinator reports line power. The normal checks no longer roll the healthy credential helper.

The updater previously installed revisions even after failing verification. Its default now keeps the running binary on failure. The host had a known Go auto-toolchain coverage failure; the quality script now pins its resolved Go version for child commands. A cancelled datastore preflight also now returns cancellation before probing a live database. All runtime changes remain gated. Updater commit 53ad03a also preserves the installer's nonzero exit code when fallback is disabled. Three isolated shell regressions passed for strict failure, strict success and explicitly enabled legacy fallback. The coordinator's installed updater script matches that commit; the peer's normal update passed the full quality gate and completed successfully at approximately 07:55 UTC.

At 07:21 UTC, Ariadne aborted Soteria build 1290 solely for exceeding its default 45-minute elapsed-time cap; Go compilation was still active. Atlas commit 0180ba66 raises that deterministic cutoff to 120 minutes. This is a time-budget backstop, not proof that a build is hung. The next Soteria release was queued with PUBLISH_IMAGES=true; its runtime fix remains unverified until publication.

At approximately 07:19 UTC, five-minute I/O pressure waiting fractions were 0.78 on Titan-19, 0.77 on Titan-08, 0.50 on Titan-17 and 0.48 on Titan-13. Titan-08's SSH query timed out and Outline's cold image pull was still pending. Outline subsequently became Ready on Titan-11. These are additional signs that runtime/storage capacity must be addressed; reducing restart churn alone does not establish healthy media.

At 07:40 UTC, all 57 active Flux Kustomizations were Ready. The intentionally suspended bstein-dev-home-migrations Kustomization retains its historical ArtifactFailed condition. This snapshot does not close the hardware, storage, backup coverage or supported-version work above.