65 lines
5.7 KiB
Markdown
65 lines
5.7 KiB
Markdown
# Cluster stabilization - implementation record
|
|
|
|
This record begins on 2026-10-03 UTC, following authorization to implement the
|
|
[configuration-first plan](cluster-audit-20261002/INTEGRATED_PLAN.md).
|
|
The October 2 audit remains a historical baseline, not a current health claim.
|
|
Use [Cluster operations](CLUSTER_OPERATIONS.md) for the ordinary operator path.
|
|
|
|
## Verified changes
|
|
|
|
| Change | Evidence | Rollback |
|
|
| --- | --- | --- |
|
|
| Native external PostgreSQL backup and second LAN copy | A 54,083,817-byte custom dump restored into a disposable PostgreSQL 16 instance in 11 seconds with 22,350 kine rows; no TCP listener. Both timers enabled; checksum and matching-token checks pass | Disable the two timers; retain all recovery bundles |
|
|
| Retired unconditional k3s-agent restart DaemonSet | Flux preview removed only that DaemonSet; reconciliation completed and the DaemonSet is absent | Revert f04c84ee, understanding that returning the DaemonSet restarts agents |
|
|
| Recovered metrics Pushgateway | Existing volume attached healthy on titan-19; replacement pod Ready. No volume deletion or data replacement | Revert the focused monitoring placement commits after the original host is repaired |
|
|
| Removed eviction for historical restart counts | Descheduler manifest validates and Flux applied the change | Revert fa9251ae |
|
|
| Restored GitOps UI Deployment | Helm drift correction recreated weave-gitops; Deployment 1/1 Ready | Revert 6c9398ea to disable ongoing drift correction; this does not remove recovered resources |
|
|
| Aligned Flux definition ownership | Removed creation-only policy and adopted already-active service state; no service paths or source refs changed | Revert the focused Flux commits; review suspension fields before doing so |
|
|
| Protected local-only data policy | Hermes PVCs excluded from Soteria's cloud policy through Flux | Do not broaden this policy without reviewing data authorization |
|
|
|
|
Host backup implementation is in commit 80aff498 and
|
|
[its operational notes](../infrastructure/host-backup/NOTES.md).
|
|
The restore check validates database/schema restoration without global-role/ACL
|
|
replay or starting K3s. A complete control-plane disaster drill remains outstanding.
|
|
|
|
## Work being completed
|
|
|
|
* OpenSearch: preserve the existing disk and 2 GiB heap; request 3 GiB memory
|
|
instead of 768 MiB, remove the offline titan-05 pin, and use the chart's actual
|
|
`nodeAffinity` setting. Pi5 then Pi4 are preferred; titan-22 is the explicit
|
|
last-resort CPU destination. The original Helm operation must converge before
|
|
the corrected generation can be verified.
|
|
* Shared application PostgreSQL: explicit 1 GiB request / 2 GiB limit, startup
|
|
and readiness checks, bounded exporter resources. Deployment waits for the
|
|
pre-change database recovery copies to finish.
|
|
* Soteria: source fix 7018d4c in the Soteria repository allows live RWO Longhorn
|
|
snapshots while preserving the restic mount guard. Manual Longhorn requests
|
|
also enforce exclusions. All Go package tests pass; normal image publication
|
|
and a representative backup verification remain required.
|
|
* Jenkins: a two-agent concurrency cap is prepared. Its existing configuration
|
|
hash triggers a controller rollout, so apply it with ongoing builds accounted
|
|
for rather than pretending it is a harmless live-only reload.
|
|
|
|
## Confirmed problems needing further work
|
|
|
|
| Problem | Evidence / implication | Next action |
|
|
| --- | --- | --- |
|
|
| titan-05 / titan-06 offline | Nineteen of twenty-one nodes Ready at the start of implementation | Physical power/network check; recover or replace; no remote software claim of repair |
|
|
| Pi power instability | Fresh undervoltage observed on titan-04 and titan-11 during audit; titan-04 remains cordoned | Check supplies, cables and USB power demand; keep quarantine until a stability check passes |
|
|
| titan-14 runtime I/O | USB flash runtime disk at sustained saturation during audit | Replace/migrate to suitable SSD after data preservation and a controlled drain |
|
|
| titan-18 storage stalls | I/O some pressure near 99%, load near 39; Longhorn engine image fails to become ready | Reduce CI load, inspect storage/kernel health, move runtime off inadequate media if confirmed |
|
|
| titan-08 stale iSCSI session | Pushgateway engine could not log out an obsolete target; replica data remained usable on another host | Repair during a controlled storage maintenance window; do not mass-restart instance managers holding healthy volumes |
|
|
| Hermes tenant-2 shared workspace | Existing faulted/detached volume, repeated recovery attempts | Preserve replica/recovery evidence and perform component-supported repair; no source content in routine logs |
|
|
| Backup coverage and restore proof | Old Soteria coverage is insufficient; repairing its scheduler does not instantly create all backups | Confirm eligible data, completed backup objects, achievable schedule and representative restores service by service |
|
|
| Supported software baseline | Ubuntu 24.10 on titan-db and older Kubernetes cohorts are out of support | Backed-up, staged host/K3s/Longhorn upgrades, one compatible cohort at a time |
|
|
| Failure capacity | Pi pool is heavily reserved; unused x86 capacity has explicit simulation/GPU roles | Recalculate compatible N+1 capacity after repairs; do not silently take reserved GPUs or simulation capacity |
|
|
|
|
## Completion standard
|
|
|
|
Do not call the entire cluster fixed based on one green snapshot. Open storage
|
|
faults, physical power/media problems, unsupported host software and per-service
|
|
restore coverage remain visible until addressed. Validate application behavior,
|
|
backup freshness, actual restore results, and 7-14 days without the recurring
|
|
failures. No automatic real-suite inference jobs or outage drills are part of
|
|
this maintenance work.
|