138 lines
12 KiB
Markdown
138 lines
12 KiB
Markdown
# Cluster stabilization - implementation record
|
|
|
|
This record begins on 2026-10-03 UTC, following authorization to implement the
|
|
[configuration-first plan](cluster-audit-20261002/INTEGRATED_PLAN.md).
|
|
The October 2 audit remains a historical baseline, not a current health claim.
|
|
Use [Cluster operations](CLUSTER_OPERATIONS.md) for the ordinary operator path.
|
|
|
|
## Verified changes
|
|
|
|
| Change | Evidence | Rollback |
|
|
| --- | --- | --- |
|
|
| Native external PostgreSQL backup and second LAN copy | A 54,083,817-byte custom dump restored into a disposable PostgreSQL 16 instance in 11 seconds with 22,350 kine rows; no TCP listener. Both timers enabled; checksum and matching-token checks pass | Disable the two timers; retain all recovery bundles |
|
|
| Retired unconditional k3s-agent restart DaemonSet | Flux preview removed only that DaemonSet; reconciliation completed and the DaemonSet is absent | Revert f04c84ee, understanding that returning the DaemonSet restarts agents |
|
|
| Recovered metrics Pushgateway | Existing volume attached healthy on titan-19; replacement pod Ready. No volume deletion or data replacement | Revert the focused monitoring placement commits after the original host is repaired |
|
|
| Removed eviction for historical restart counts | Descheduler manifest validates and Flux applied the change | Revert fa9251ae |
|
|
| Restored GitOps UI Deployment | Helm drift correction recreated weave-gitops; Deployment 1/1 Ready | Revert 6c9398ea to disable ongoing drift correction; this does not remove recovered resources |
|
|
| Aligned Flux definition ownership | Removed creation-only policy and adopted already-active service state; no service paths or source refs changed | Revert the focused Flux commits; review suspension fields before doing so |
|
|
| Backup freshness alerts | Both datastore-copy timestamps are scraped; native Grafana provisioning reload returned HTTP 200 | Revert the focused monitoring change; backups continue independently |
|
|
| Quarantined stalled runtime media | Titan-14 and Titan-18 are SchedulingDisabled via `infrastructure/core/node-maintenance.yaml`; no forced storage detach | After repair, change `unschedulable` to false in Git and verify before removing the prune-disabled Node declaration |
|
|
| Bounded CI concurrency | Jenkins controller Ready after applying the two-agent cap; existing agents reconnected | Revert 440244f3 and reconcile Jenkins; allow a controller restart |
|
|
| Spread Vault injectors | Two healthy replicas on Titan-08 and Titan-22; native anti-affinity and minAvailable=1 PDB | Revert 5f3f3184 |
|
|
| Protected application databases | Nineteen logical dumps and globals completed with checksums on Titan-0b; isolated Gitea restore passed in six seconds with 111 tables | Keep copies; stop the daily timer if needed |
|
|
| Application PostgreSQL resource/probe repair | Pinned current PostgreSQL 15.18 image, 1 GiB request / 2 GiB limit, native startup/readiness; 2/2 Ready on Titan-17 and healthy attached volume | Revert 103e2064 only after checking destination capacity; another database restart is required |
|
|
| Recovered Firefly and OpenSearch | Firefly Ready on Titan-08; OpenSearch Ready with its existing healthy volume on Titan-22. Pi capacity could not fit OpenSearch, so the documented CPU fallback was used; no GPU reassignment | Review capacity before placement rollback; preserve existing volumes |
|
|
| Protected local-only data policy | Hermes PVCs excluded from Soteria's cloud policy through Flux | Do not broaden this policy without reviewing data authorization |
|
|
| Stopped stale credential-event restart loop | Ananke c73b8fe installed on Titan-db after its complete quality gate; helper generation 796 remained unchanged after normal one-minute checks resumed. Peer ec7c2ee no longer runs periodic cluster repairs | Native installer retains `/usr/local/lib/ananke/rollback/ananke.previous`; rolling back restores the defective behavior, so retain a bounded repair interval if needed |
|
|
| Corrected premature CI cancellation | Ariadne's configured elapsed-time cutoff is now 120 minutes, verified in the running Deployment | Revert 0180ba66; this returns the 45-minute cutoff that aborted an active Pi build |
|
|
|
|
Host backup implementation is in commit 80aff498 and
|
|
[its operational notes](../infrastructure/host-backup/NOTES.md).
|
|
The restore check validates database/schema restoration without global-role/ACL
|
|
replay or starting K3s. A complete control-plane disaster drill remains outstanding.
|
|
|
|
## Work being completed
|
|
|
|
* Soteria: source fix 7018d4c in the Soteria repository allows live RWO Longhorn
|
|
snapshots while preserving the restic mount guard. Manual Longhorn requests
|
|
also enforce exclusions. All Go package tests pass; normal image publication
|
|
and a representative backup verification remain required.
|
|
* Native application backups now run daily from Titan-0b. The timer is enabled,
|
|
the verified initial bundle timestamp is scraped, and the 36-hour freshness
|
|
alert is provisioned. The first scheduled run remains to be observed.
|
|
* Soteria CI: the coverage report was generated after Sonar analysis, causing a
|
|
false zero-coverage gate despite 96.4% measured test coverage. Commit f167d7d
|
|
orders tests before analysis and awaits the Sonar gate; Jenkins declarative
|
|
validation passed. Release build 1292 failed during Git checkout with a 404;
|
|
current credential/repository checks and build 1293's source checkout succeed.
|
|
The release retry remains queued; do not claim the runtime
|
|
backup defect is repaired until the new image and a real eligible backup pass.
|
|
|
|
|
|
## Confirmed problems needing further work
|
|
|
|
| Problem | Evidence / implication | Next action |
|
|
| --- | --- | --- |
|
|
| titan-05 / titan-06 offline | Nineteen of twenty-one nodes Ready at the start of implementation | Physical power/network check; recover or replace; no remote software claim of repair |
|
|
| Pi power instability | Fresh undervoltage observed on titan-04 and titan-11 during audit; titan-04 remains cordoned | Check supplies, cables and USB power demand; keep quarantine until a stability check passes |
|
|
| titan-14 runtime I/O | USB flash runtime disk at sustained saturation during audit | Replace/migrate to suitable SSD after data preservation and a controlled drain |
|
|
| titan-18 storage stalls | I/O some pressure near 99%, load near 39; Longhorn engine image fails to become ready | Reduce CI load, inspect storage/kernel health, move runtime off inadequate media if confirmed |
|
|
| titan-08 stale iSCSI session | Pushgateway engine could not log out an obsolete target; replica data remained usable on another host | Repair during a controlled storage maintenance window; do not mass-restart instance managers holding healthy volumes |
|
|
| Hermes tenant-2 shared workspace | Existing faulted/detached volume, repeated recovery attempts | Preserve replica/recovery evidence and perform component-supported repair; no source content in routine logs |
|
|
| Backup coverage and restore proof | Old Soteria coverage is insufficient; repairing its scheduler does not instantly create all backups | Confirm eligible data, completed backup objects, achievable schedule and representative restores service by service |
|
|
| Database collation drift | Native dumps reported stored collation 2.36 versus runtime 2.41 for several databases | Plan index rebuilds with compatible locale settings before refreshing version metadata; do not merely suppress warnings |
|
|
| Supported software baseline | Ubuntu 24.10 on titan-db and older Kubernetes cohorts are out of support | Backed-up, staged host/K3s/Longhorn upgrades, one compatible cohort at a time |
|
|
| Failure capacity | Pi pool is heavily reserved; unused x86 capacity has explicit simulation/GPU roles | Recalculate compatible N+1 capacity after repairs; do not silently take reserved GPUs or simulation capacity |
|
|
|
|
## Completion standard
|
|
|
|
Do not call the entire cluster fixed based on one green snapshot. Open storage
|
|
faults, physical power/media problems, unsupported host software and per-service
|
|
restore coverage remain visible until addressed. Validate application behavior,
|
|
backup freshness, actual restore results, and 7-14 days without the recurring
|
|
failures. No automatic real-suite inference jobs or outage drills are part of
|
|
this maintenance work.
|
|
|
|
## Operational lessons from this maintenance
|
|
|
|
The application database move took approximately nine minutes because both
|
|
pinned images had to be pulled on the destination. Its volume attached correctly;
|
|
PostgreSQL and its exporter became Ready afterward. Preload required images on
|
|
an eligible destination before another planned database move. The larger database
|
|
dumps also completed much more efficiently through native PostgreSQL than through
|
|
`kubectl exec` output streaming. The native recovery script records this path.
|
|
|
|
At 06:07 UTC, both Titan-04 and Titan-11 emitted fresh undervoltage messages.
|
|
Their instantaneous Pi 5 input samples were 4.9245 V and 4.89904 V; these are not
|
|
measurements of the transient minimum. Both had `get_throttled=0x50000` between
|
|
events. The fault remains active even when that individual sample reports only
|
|
historical bits. The user is checking power/boot-media issues physically.
|
|
|
|
Firefly recovered on Titan-08 at approximately 06:52 UTC after its old pod
|
|
terminated and its volume moved normally. OpenSearch recovered on Titan-22 around 07:03 UTC after the corrected placement
|
|
was applied. Commit 2f043d10 temporarily allowed the old offline-node rollback
|
|
to finish. Normal rollback readiness was restored after the corrected release
|
|
became Ready.
|
|
Titan-18 remains cordoned; rescheduling does not repair its runtime media.
|
|
|
|
Ananke was confirmed restarting jenkins-vault-sync approximately every minute
|
|
from retained credential-warning events, including after the original failure
|
|
cleared. The Ananke fixes require
|
|
a current matching pod UID, active image-pull failure and warning age below ten
|
|
minutes, with a persisted 30-minute cooldown per helper. Only the coordinator
|
|
runs periodic cluster repair; the peer retains UPS monitoring and forwarding.
|
|
Coordinator revision c73b8fe passed the native host's complete quality gate,
|
|
including the per-file 95% coverage check, and installed successfully. Its
|
|
temporary one-hour interval was restored to the original 60 seconds at 07:41 UTC.
|
|
The peer runs ec7c2ee, which includes the credential guard and coordinator-only
|
|
repair. UPS services remain active; the coordinator reports line power. The
|
|
normal checks no longer roll the healthy credential helper.
|
|
|
|
The updater previously installed revisions even after failing verification.
|
|
Its default now keeps the running binary on failure. The host had a known Go
|
|
auto-toolchain coverage failure; the quality script now pins its resolved Go
|
|
version for child commands. A cancelled datastore preflight also now returns
|
|
cancellation before probing a live database. All runtime changes remain gated.
|
|
Updater commit 53ad03a also preserves the installer's nonzero exit code when
|
|
fallback is disabled. Three isolated shell regressions passed for strict failure,
|
|
strict success and explicitly enabled legacy fallback. The coordinator's installed
|
|
updater script matches that commit; the peer's normal update is in progress.
|
|
|
|
At 07:21 UTC, Ariadne aborted Soteria build 1290 solely for exceeding its default
|
|
45-minute elapsed-time cap; Go compilation was still active. Atlas commit
|
|
0180ba66 raises that deterministic cutoff to 120 minutes. This is a time-budget
|
|
backstop, not proof that a build is hung. The next Soteria release was queued
|
|
with PUBLISH_IMAGES=true; its runtime fix remains unverified until publication.
|
|
|
|
At approximately 07:19 UTC, five-minute I/O pressure waiting fractions were
|
|
0.78 on Titan-19, 0.77 on Titan-08, 0.50 on Titan-17 and 0.48 on Titan-13.
|
|
Titan-08's SSH query timed out and Outline's cold image pull was still pending.
|
|
Outline subsequently became Ready on Titan-11. These are additional signs that
|
|
runtime/storage capacity must be addressed; reducing restart churn alone does
|
|
not establish healthy media.
|
|
|
|
At 07:40 UTC, all 57 active Flux Kustomizations were Ready. The intentionally
|
|
suspended bstein-dev-home-migrations Kustomization retains its historical
|
|
ArtifactFailed condition. This snapshot does not close the hardware, storage,
|
|
backup coverage or supported-version work above.
|