94 lines
8.5 KiB
Markdown
94 lines
8.5 KiB
Markdown
# Cluster stabilization - implementation record
|
|
|
|
This record begins on 2026-10-03 UTC, following authorization to implement the
|
|
[configuration-first plan](cluster-audit-20261002/INTEGRATED_PLAN.md).
|
|
The October 2 audit remains a historical baseline, not a current health claim.
|
|
Use [Cluster operations](CLUSTER_OPERATIONS.md) for the ordinary operator path.
|
|
|
|
## Verified changes
|
|
|
|
| Change | Evidence | Rollback |
|
|
| --- | --- | --- |
|
|
| Native external PostgreSQL backup and second LAN copy | A 54,083,817-byte custom dump restored into a disposable PostgreSQL 16 instance in 11 seconds with 22,350 kine rows; no TCP listener. Both timers enabled; checksum and matching-token checks pass | Disable the two timers; retain all recovery bundles |
|
|
| Retired unconditional k3s-agent restart DaemonSet | Flux preview removed only that DaemonSet; reconciliation completed and the DaemonSet is absent | Revert f04c84ee, understanding that returning the DaemonSet restarts agents |
|
|
| Recovered metrics Pushgateway | Existing volume attached healthy on titan-19; replacement pod Ready. No volume deletion or data replacement | Revert the focused monitoring placement commits after the original host is repaired |
|
|
| Removed eviction for historical restart counts | Descheduler manifest validates and Flux applied the change | Revert fa9251ae |
|
|
| Restored GitOps UI Deployment | Helm drift correction recreated weave-gitops; Deployment 1/1 Ready | Revert 6c9398ea to disable ongoing drift correction; this does not remove recovered resources |
|
|
| Aligned Flux definition ownership | Removed creation-only policy and adopted already-active service state; no service paths or source refs changed | Revert the focused Flux commits; review suspension fields before doing so |
|
|
| Backup freshness alerts | Both datastore-copy timestamps are scraped; native Grafana provisioning reload returned HTTP 200 | Revert the focused monitoring change; backups continue independently |
|
|
| Quarantined stalled runtime media | Titan-14 and Titan-18 are SchedulingDisabled via `infrastructure/core/node-maintenance.yaml`; no forced storage detach | After repair, change `unschedulable` to false in Git and verify before removing the prune-disabled Node declaration |
|
|
| Bounded CI concurrency | Jenkins controller Ready after applying the two-agent cap; existing agents reconnected | Revert 440244f3 and reconcile Jenkins; allow a controller restart |
|
|
| Spread Vault injectors | Two healthy replicas on Titan-08 and Titan-22; native anti-affinity and minAvailable=1 PDB | Revert 5f3f3184 |
|
|
| Protected application databases | Nineteen logical dumps and globals completed with checksums on Titan-0b; isolated Gitea restore passed in six seconds with 111 tables | Keep copies; stop the daily timer if needed |
|
|
| Application PostgreSQL resource/probe repair | Pinned current PostgreSQL 15.18 image, 1 GiB request / 2 GiB limit, native startup/readiness; 2/2 Ready on Titan-17 and healthy attached volume | Revert 103e2064 only after checking destination capacity; another database restart is required |
|
|
| Protected local-only data policy | Hermes PVCs excluded from Soteria's cloud policy through Flux | Do not broaden this policy without reviewing data authorization |
|
|
|
|
Host backup implementation is in commit 80aff498 and
|
|
[its operational notes](../infrastructure/host-backup/NOTES.md).
|
|
The restore check validates database/schema restoration without global-role/ACL
|
|
replay or starting K3s. A complete control-plane disaster drill remains outstanding.
|
|
|
|
## Work being completed
|
|
|
|
* OpenSearch: preserve the existing disk and 2 GiB heap; request 3 GiB memory
|
|
instead of 768 MiB, remove the offline titan-05 pin, and use the chart's actual
|
|
`nodeAffinity` setting. Pi5 then Pi4 are preferred; titan-22 is the explicit
|
|
last-resort CPU destination. The original Helm operation must converge before
|
|
the corrected generation can be verified.
|
|
* Soteria: source fix 7018d4c in the Soteria repository allows live RWO Longhorn
|
|
snapshots while preserving the restic mount guard. Manual Longhorn requests
|
|
also enforce exclusions. All Go package tests pass; normal image publication
|
|
and a representative backup verification remain required.
|
|
* Native application-backup schedule and its freshness alert are being installed
|
|
after the initial successful copy and restore check.
|
|
* Soteria CI: the coverage report was generated after Sonar analysis, causing a
|
|
false zero-coverage gate despite 96.4% measured test coverage. Commit f167d7d
|
|
orders tests before analysis and awaits the Sonar gate; Jenkins declarative
|
|
validation passed. Release build remains pending; do not claim the runtime
|
|
backup defect is repaired until the new image and a real eligible backup pass.
|
|
|
|
|
|
## Confirmed problems needing further work
|
|
|
|
| Problem | Evidence / implication | Next action |
|
|
| --- | --- | --- |
|
|
| titan-05 / titan-06 offline | Nineteen of twenty-one nodes Ready at the start of implementation | Physical power/network check; recover or replace; no remote software claim of repair |
|
|
| Pi power instability | Fresh undervoltage observed on titan-04 and titan-11 during audit; titan-04 remains cordoned | Check supplies, cables and USB power demand; keep quarantine until a stability check passes |
|
|
| titan-14 runtime I/O | USB flash runtime disk at sustained saturation during audit | Replace/migrate to suitable SSD after data preservation and a controlled drain |
|
|
| titan-18 storage stalls | I/O some pressure near 99%, load near 39; Longhorn engine image fails to become ready | Reduce CI load, inspect storage/kernel health, move runtime off inadequate media if confirmed |
|
|
| titan-08 stale iSCSI session | Pushgateway engine could not log out an obsolete target; replica data remained usable on another host | Repair during a controlled storage maintenance window; do not mass-restart instance managers holding healthy volumes |
|
|
| Hermes tenant-2 shared workspace | Existing faulted/detached volume, repeated recovery attempts | Preserve replica/recovery evidence and perform component-supported repair; no source content in routine logs |
|
|
| Backup coverage and restore proof | Old Soteria coverage is insufficient; repairing its scheduler does not instantly create all backups | Confirm eligible data, completed backup objects, achievable schedule and representative restores service by service |
|
|
| Database collation drift | Native dumps reported stored collation 2.36 versus runtime 2.41 for several databases | Plan index rebuilds with compatible locale settings before refreshing version metadata; do not merely suppress warnings |
|
|
| Supported software baseline | Ubuntu 24.10 on titan-db and older Kubernetes cohorts are out of support | Backed-up, staged host/K3s/Longhorn upgrades, one compatible cohort at a time |
|
|
| Failure capacity | Pi pool is heavily reserved; unused x86 capacity has explicit simulation/GPU roles | Recalculate compatible N+1 capacity after repairs; do not silently take reserved GPUs or simulation capacity |
|
|
|
|
## Completion standard
|
|
|
|
Do not call the entire cluster fixed based on one green snapshot. Open storage
|
|
faults, physical power/media problems, unsupported host software and per-service
|
|
restore coverage remain visible until addressed. Validate application behavior,
|
|
backup freshness, actual restore results, and 7-14 days without the recurring
|
|
failures. No automatic real-suite inference jobs or outage drills are part of
|
|
this maintenance work.
|
|
|
|
## Operational lessons from this maintenance
|
|
|
|
The application database move took approximately nine minutes because both
|
|
pinned images had to be pulled on the destination. Its volume attached correctly;
|
|
PostgreSQL and its exporter became Ready afterward. Preload required images on
|
|
an eligible destination before another planned database move. The larger database
|
|
dumps also completed much more efficiently through native PostgreSQL than through
|
|
`kubectl exec` output streaming. The native recovery script records this path.
|
|
|
|
At 06:07 UTC, both Titan-04 and Titan-11 emitted fresh undervoltage messages.
|
|
Their instantaneous Pi 5 input samples were 4.9245 V and 4.89904 V; these are not
|
|
measurements of the transient minimum. Both had `get_throttled=0x50000` between
|
|
events. The fault remains active even when that individual sample reports only
|
|
historical bits. The user is checking power/boot-media issues physically.
|
|
|
|
Titan-18 remains a blocker: Firefly and OpenSearch pods are waiting for graceful
|
|
termination on its stalled runtime. Do not force detach their mounted volumes
|
|
while the old host may still have writers. Its replacement/repair is distinct
|
|
from Kubernetes scheduling; a cordon alone cannot repair hung I/O.
|