diff --git a/docs/CLUSTER_OPERATIONS.md b/docs/CLUSTER_OPERATIONS.md index 4d4d30f9..65561cbe 100644 --- a/docs/CLUSTER_OPERATIONS.md +++ b/docs/CLUSTER_OPERATIONS.md @@ -176,6 +176,14 @@ control plane boot. Live Longhorn snapshots and database-native dumps offer different consistency guarantees. A recent backup indicator is not a restore test. Keep database restore checks and service recovery drills explicit. +For cloud backups, inspect completed backup timestamps, not just the backup +target's Available status or a successful bucket listing. The October 3 audit +found a readable Backblaze bucket rejecting uploads with `storage cap exceeded`. +Check Backblaze Caps & Alerts and agree the storage budget before changing it. +Do not delete old recovery points merely to make uploads work. Distinguish +visible-object bytes from billed storage, including retained versions. The +current diagnosis and counts are in the implementation record. + The Hermes namespace is excluded from cloud backup policy. Its workspaces can contain material restricted to local infrastructure. Do not remove that exclusion to make a coverage dashboard greener. diff --git a/docs/CLUSTER_STABILIZATION.md b/docs/CLUSTER_STABILIZATION.md index 413d47e7..2c225b62 100644 --- a/docs/CLUSTER_STABILIZATION.md +++ b/docs/CLUSTER_STABILIZATION.md @@ -7,16 +7,16 @@ Use [Cluster operations](CLUSTER_OPERATIONS.md) for the ordinary operator path. ## Six-priority progress tracker -Started against the user's approved list on 2026-10-03. Updated: 21:22 UTC. +Started against the user's approved list on 2026-10-03. Updated: 21:32 UTC. "Verified" means the stated check passed; it does not imply a completed soak or failure drill. Physical repairs and disruptive recovery drills remain separate from routine software work. | Priority | Status | Completed evidence | Next action / completion gate | | --- | --- | --- | --- | -| 1. Immediate software repairs | In progress | Monerod remounted its existing database on Titan-19, initialized RPC and became 2/2 Ready with zero restarts. Soteria's release permissions/build defects are corrected; release 1297 is running | Verify responsive Monerod RPC and sustained sync; publish Soteria and validate an eligible backup and restore | +| 1. Immediate software repairs | In progress | Monerod remounted its existing database on Titan-19, became 2/2 Ready with zero restarts, returned RPC OK and advanced 20 blocks. Soteria's release permissions/build defects are corrected; release 1297 is running | Observe Monerod latency/sync; publish Soteria and validate an eligible backup and restore | | 2. Reliable nodes | Partial; physical checks needed | Titan-04/14/18 excluded from new placement; Titan-05/06 remain offline. Fresh Titan-04/11 undervoltage was observed | Power/media checks, preservation or replacement of suspect media, then networking/storage/load validation before return. Establish the actual condition of 09/10/16 | -| 3. Recovery coverage | In progress | External datastore restore and second LAN copy verified; 19 application DB dumps completed; scheduled application backup completed at 08:37 UTC; representative Gitea restore passed | Complete eligible PVC coverage, verify Soteria restore, resolve Hermes tenant-2 safely, and provide approved local-only recovery copies | +| 3. Recovery coverage | In progress; cloud uploads blocked | External datastore restore and second LAN copy verified; 19 application DB dumps completed; scheduled application backup completed at 08:37 UTC; representative Gitea restore passed. Backblaze storage-cap failure established | Resolve the account storage cap within an agreed budget, complete eligible PVC coverage and Soteria restore, resolve Hermes tenant-2 safely, and provide approved local-only recovery copies | | 4. Capacity and placement | Partial | CI agent cap reduced to two; shared PostgreSQL/OpenSearch resource corrections deployed; Vault injectors spread. Longhorn now uses least-effort balancing and one rebuild per node | Account for the measured 2-3.5 GiB storage-process footprint, correct remaining requests/placement, and calculate compatible spare capacity for one worker loss | | 5. Supported software | Planned; backups established | Installed K3s cohorts identified as 1.31.5 and 1.33.3; titan-db Ubuntu 24.10 is beyond support | Select a compatible supported baseline; stage OS/K3s/storage upgrades with rollback. Plan collation/index repair | | 6. Maintainability and sustained health | In progress | Operator guide and rollback record published; Ananke stale-event loop fixed; updater failure reporting corrected; overlapping restart/eviction behavior removed | Complete the service recovery checklist, validate remaining automation boundaries, arrange controlled recovery checks, then observe 7-14 days | @@ -30,11 +30,12 @@ from routine software work. - [ ] Verify an eligible backup and an isolated restore; retain Hermes exclusions. - [x] Diagnose Monerod's host/attachment/filesystem error before selecting recovery. - [x] Preserve a pre-recovery snapshot and remount the existing Monerod volume through Flux. -- [ ] Recover Monerod with its existing data and verify application RPC health. +- [x] Recover Monerod with its existing data: get_version returned OK; height advanced from 3775922 to 3775942. Broader get_info requests remain slow. - [x] Deploy and verify bounded Longhorn balancing/rebuild settings. - [x] Verify the first scheduled native application PostgreSQL backup completed. - [x] Recheck controller health: 57/57 active Flux Kustomizations Ready at 21:18 UTC. - [ ] Complete application checks and the 7-14 day observation period. +- [ ] Resolve Backblaze's storage cap before accepting cloud backup coverage; preserve existing copies. ## Verified changes @@ -109,6 +110,9 @@ became 2/2 Ready without a restart. The first ten-second RPC check timed out; application responsiveness and steady synchronization remain explicit checks. No PVC deletion, forced detach, database replacement or manual filesystem repair was performed. Retain the snapshot until recovery verification is complete. +Subsequent `get_version` calls returned status OK with the existing height +3775922, then 3775942 against target 3776212. Synchronization is making progress; +the broader `get_info` call still timed out at 45 seconds during recovery. The 21:12 UTC capacity snapshot showed Titan-07/11 CPU requests at 3.53/3.56 of 3.60 allocatable cores; Titan-11 memory requests were 6662 of 6785 MiB. Titan-17/19 @@ -134,6 +138,40 @@ Hermes after a transient Longhorn webhook timeout cleared. That controller snapshot is not a declaration of sufficient failure capacity or completed hardware repair. +### Cloud backup blocker established + +At approximately 21:29 UTC the Longhorn catalog contained 980 backup objects: +952 Error, 14 Completed and 14 without a reported state. The completed entries +were from June/July; 899 retained errors explicitly reported `storage cap +exceeded`. Recent failed attempts on October 3 reported the same condition. +This is not evidence that rotating an expired key will repair uploads. + +An authenticated Backblaze authorization check succeeded. The synchronized +Longhorn/Soteria key is non-expiring and has writeFiles/readFiles/listFiles +permissions. Its scope also includes account/bucket/key management; replace +that broad shared scope with dedicated bucket-scoped credentials in a separate +controlled rotation, without revoking other consumers blindly. No credentials +were printed or committed. + +Soteria's cached metadata scan at 21:27:15 UTC reported 64,692 visible objects +and 963,207,320,600 bytes in `atlas-soteria`, no new objects in 24 hours, and a +last object modification of July 15. This is a visible-object inventory, not +verified billable storage including historical versions or other buckets. +Backblaze's exact account cap and approved monthly budget are not known from +the supported API checks. The user has been asked to review Caps & Alerts and +provide the budget. No spending cap was raised, old recovery copy removed, or +local-only material uploaded. A readable bucket and an Available backup target +do not establish successful writes. + +Hermes tenant-2 remains unresolved: its volume is detached with a share manager +waiting on Titan-14, while an older engine reports running on Titan-23 with an +ERR replica reference. Three replica objects remain on Titan-13/15/17. These are +controller observations, not evidence of empty or lost data. Both consumers +report Running despite the unavailable share. No workspace content was read, +consumer stopped, volume forcibly detached or replica removed during inspection. +Preserving an approved local recovery copy and repairing that specific engine/ +share ownership remains an explicit recovery task. + ## Confirmed problems needing further work