diff --git a/docs/CLUSTER_STABILIZATION.md b/docs/CLUSTER_STABILIZATION.md index 76cc92ce..413d47e7 100644 --- a/docs/CLUSTER_STABILIZATION.md +++ b/docs/CLUSTER_STABILIZATION.md @@ -7,29 +7,34 @@ Use [Cluster operations](CLUSTER_OPERATIONS.md) for the ordinary operator path. ## Six-priority progress tracker -Started against the user's approved list on 2026-10-03. Updated: 20:45 UTC. +Started against the user's approved list on 2026-10-03. Updated: 21:22 UTC. "Verified" means the stated check passed; it does not imply a completed soak or failure drill. Physical repairs and disruptive recovery drills remain separate from routine software work. | Priority | Status | Completed evidence | Next action / completion gate | | --- | --- | --- | --- | -| 1. Immediate software repairs | In progress | Soteria test builds 1293/1294 passed; release 1295 failed its supply-chain gate. Monerod's LMDB open fails with I/O error despite healthy Longhorn replicas | Correct the release findings, publish and validate a backup; repair Monerod's storage path without deleting its database | +| 1. Immediate software repairs | In progress | Monerod remounted its existing database on Titan-19, initialized RPC and became 2/2 Ready with zero restarts. Soteria's release permissions/build defects are corrected; release 1297 is running | Verify responsive Monerod RPC and sustained sync; publish Soteria and validate an eligible backup and restore | | 2. Reliable nodes | Partial; physical checks needed | Titan-04/14/18 excluded from new placement; Titan-05/06 remain offline. Fresh Titan-04/11 undervoltage was observed | Power/media checks, preservation or replacement of suspect media, then networking/storage/load validation before return. Establish the actual condition of 09/10/16 | | 3. Recovery coverage | In progress | External datastore restore and second LAN copy verified; 19 application DB dumps completed; scheduled application backup completed at 08:37 UTC; representative Gitea restore passed | Complete eligible PVC coverage, verify Soteria restore, resolve Hermes tenant-2 safely, and provide approved local-only recovery copies | -| 4. Capacity and placement | Partial | CI agent cap reduced to two; shared PostgreSQL/OpenSearch resource corrections deployed; Vault injectors spread | Measure peak demand and I/O, correct remaining requests/placement, and calculate compatible spare capacity for one worker loss | +| 4. Capacity and placement | Partial | CI agent cap reduced to two; shared PostgreSQL/OpenSearch resource corrections deployed; Vault injectors spread. Longhorn now uses least-effort balancing and one rebuild per node | Account for the measured 2-3.5 GiB storage-process footprint, correct remaining requests/placement, and calculate compatible spare capacity for one worker loss | | 5. Supported software | Planned; backups established | Installed K3s cohorts identified as 1.31.5 and 1.33.3; titan-db Ubuntu 24.10 is beyond support | Select a compatible supported baseline; stage OS/K3s/storage upgrades with rollback. Plan collation/index repair | | 6. Maintainability and sustained health | In progress | Operator guide and rollback record published; Ananke stale-event loop fixed; updater failure reporting corrected; overlapping restart/eviction behavior removed | Complete the service recovery checklist, validate remaining automation boundaries, arrange controlled recovery checks, then observe 7-14 days | ### Current work and acceptance checks -- [ ] Resolve Soteria release gate findings without disabling the gate. +- [x] Correct Soteria release gate findings without disabling the gate; verify the local configuration scan. +- [x] Replace Soteria's daemonless Docker build with the existing pinned Kaniko builder; validate the pipeline. +- [x] Deploy and verify scoped Soteria state/credential permissions without altering stored policy or usage. - [ ] Publish the corrected Soteria image through the normal pipeline and Flux. - [ ] Verify an eligible backup and an isolated restore; retain Hermes exclusions. -- [ ] Diagnose Monerod's host/attachment/filesystem error before selecting recovery. +- [x] Diagnose Monerod's host/attachment/filesystem error before selecting recovery. +- [x] Preserve a pre-recovery snapshot and remount the existing Monerod volume through Flux. - [ ] Recover Monerod with its existing data and verify application RPC health. +- [x] Deploy and verify bounded Longhorn balancing/rebuild settings. - [x] Verify the first scheduled native application PostgreSQL backup completed. -- [ ] Recheck health and update this tracker after each deployed repair. +- [x] Recheck controller health: 57/57 active Flux Kustomizations Ready at 21:18 UTC. +- [ ] Complete application checks and the 7-14 day observation period. ## Verified changes @@ -51,6 +56,9 @@ from routine software work. | Protected local-only data policy | Hermes PVCs excluded from Soteria's cloud policy through Flux | Do not broaden this policy without reviewing data authorization | | Stopped stale credential-event restart loop | Ananke c73b8fe installed on Titan-db after its complete quality gate; helper generation 796 remained unchanged after normal one-minute checks resumed. Peer 53ad03a no longer runs periodic cluster repairs | Native installer retains `/usr/local/lib/ananke/rollback/ananke.previous`; rolling back restores the defective behavior, so retain a bounded repair interval if needed | | Corrected premature CI cancellation | Ariadne's configured elapsed-time cutoff is now 120 minutes, verified in the running Deployment | Revert 0180ba66; this returns the 45-minute cutoff that aborted an active Pi build | +| Scoped Soteria permissions | Atlas 46596536 applied; intended state updates and credential reads allowed, unrelated secret reads/listing and Hermes job creation denied; policy/usage fingerprints unchanged | Revert only after reviewing why broader access is required; retain state Secrets | +| Recovered Monerod mount | Snapshot df99cdac, clean detach b119076c, restart 0984164f; existing ext4 filesystem mounted on Titan-19, LMDB opened and RPC initialized | Preserve the snapshot and PVC. Returning to Titan-08 requires storage-path validation; do not replay the temporary scale-zero commit as a normal rollback | +| Bounded Longhorn background I/O | b0be99b7 applied; live settings are least-effort balancing and one concurrent rebuild per node | Revert b0be99b7 and reconcile Longhorn; the versioned settings Job restores best-effort and two rebuilds per node | Host backup implementation is in commit 80aff498 and [its operational notes](../infrastructure/host-backup/NOTES.md). @@ -73,9 +81,59 @@ replay or starting K3s. A complete control-plane disaster drill remains outstand validation passed. Release build 1292 failed during Git checkout with a 404; current credential/repository checks and build 1293's source checkout succeed. Builds 1293/1294 subsequently passed. Release retry 1295 failed the supply-chain - gate; its findings are under investigation. Do not claim the runtime + gate on broad ClusterRole secret access, not a dependency vulnerability. Source + 3c0c868 removes those permissions and the expired exception; Atlas 46596536 + deploys namespaced, resource-named state access. A verified Trivy 0.70.0 local + configuration scan reports zero HIGH/CRITICAL findings using its embedded + checks. Jenkins remains the publication gate. Source 596b2df also replaces a + Docker client with no daemon using the cluster's existing digest-pinned Kaniko + builder, without privileged access. Declarative pipeline validation passed. + Redundant timer build 1296 was stopped before any stage/agent began; release + 1297 runs the complete gate with PUBLISH_IMAGES=true. Do not claim the runtime backup defect is repaired until the new image and a real eligible backup pass. +### October 3 storage recovery and capacity evidence + +Titan-08's kernel recorded I/O failure, an aborted ext4 journal and a read-only +remount of Monerod's device at 11:31 UTC. Longhorn still reported two healthy +replicas on Titan-15/17. Replica health therefore did not establish filesystem +health. A retained local snapshot, `monerod-before-recovery-20261003`, became +ready at 20:52 UTC. It is a recovery point, not an independent backup. + +Monerod was stopped through its Flux-tracked Deployment; CSI completed unstage +at 21:04 UTC and the volume detached normally. Restoring one replica with +Titan-08 excluded moved it to Titan-19. An initial attachment lookup failure +resolved through the normal controller retry. The existing ext4 filesystem +mounted read/write, LMDB opened, RPC initialized at 21:16:58 UTC and the pod +became 2/2 Ready without a restart. The first ten-second RPC check timed out; +application responsiveness and steady synchronization remain explicit checks. +No PVC deletion, forced detach, database replacement or manual filesystem repair +was performed. Retain the snapshot until recovery verification is complete. + +The 21:12 UTC capacity snapshot showed Titan-07/11 CPU requests at 3.53/3.56 of +3.60 allocatable cores; Titan-11 memory requests were 6662 of 6785 MiB. Titan-17/19 +working memory was 5994/6102 of approximately 6656 MiB. Longhorn instance-manager +pods have CPU requests but no memory requests in this installed configuration. +Their observed 24-hour memory peaks were 3525 MiB on Titan-17, 3176 on Titan-19, +2914 on Titan-13 and 2594 on Titan-15. Those bytes cannot be treated as spare +application capacity merely because Kubernetes has not reserved them. + +At 21:16 UTC Titan-19 reported high I/O and memory pressure while the blockchain +opened and CI ran. The native Longhorn settings were changed from best-effort +to least-effort replica balancing and from two to one concurrent rebuild per +node, verified live at 21:18 UTC. This reduces elective replica movement and +concurrent recovery I/O; it does not repair slow media, disable required replica +recovery, or change desired replica counts. Multiple rebuilds may take longer. +The existing 600-second replica replenishment delay is unchanged. Per-volume +overrides remain unchanged (28 least-effort, three disabled at inspection). +Setting semantics were checked against the installed +[Longhorn 1.8.2 source](https://github.com/longhorn/longhorn-manager/blob/v1.8.2/types/setting.go). + +All 57 active Flux Kustomizations were Ready again at 21:18 UTC, including +Hermes after a transient Longhorn webhook timeout cleared. That controller +snapshot is not a declaration of sufficient failure capacity or completed +hardware repair. + ## Confirmed problems needing further work