docs: record verified backup storage cap blocker
Some checks failed
Tests / Declarative: Post Actions failed: 45, skipped: 89, passed: 4270

This commit is contained in:
jenkins 2026-10-03 16:32:20 -05:00
parent 5d50d2f8c8
commit b41dab9886
2 changed files with 50 additions and 4 deletions

View File

@ -176,6 +176,14 @@ control plane boot. Live Longhorn snapshots and database-native dumps offer
different consistency guarantees. A recent backup indicator is not a restore
test. Keep database restore checks and service recovery drills explicit.
For cloud backups, inspect completed backup timestamps, not just the backup
target's Available status or a successful bucket listing. The October 3 audit
found a readable Backblaze bucket rejecting uploads with `storage cap exceeded`.
Check Backblaze Caps & Alerts and agree the storage budget before changing it.
Do not delete old recovery points merely to make uploads work. Distinguish
visible-object bytes from billed storage, including retained versions. The
current diagnosis and counts are in the implementation record.
The Hermes namespace is excluded from cloud backup policy. Its workspaces can
contain material restricted to local infrastructure. Do not remove that exclusion
to make a coverage dashboard greener.

View File

@ -7,16 +7,16 @@ Use [Cluster operations](CLUSTER_OPERATIONS.md) for the ordinary operator path.
## Six-priority progress tracker
Started against the user's approved list on 2026-10-03. Updated: 21:22 UTC.
Started against the user's approved list on 2026-10-03. Updated: 21:32 UTC.
"Verified" means the stated check passed; it does not imply a completed soak or
failure drill. Physical repairs and disruptive recovery drills remain separate
from routine software work.
| Priority | Status | Completed evidence | Next action / completion gate |
| --- | --- | --- | --- |
| 1. Immediate software repairs | In progress | Monerod remounted its existing database on Titan-19, initialized RPC and became 2/2 Ready with zero restarts. Soteria's release permissions/build defects are corrected; release 1297 is running | Verify responsive Monerod RPC and sustained sync; publish Soteria and validate an eligible backup and restore |
| 1. Immediate software repairs | In progress | Monerod remounted its existing database on Titan-19, became 2/2 Ready with zero restarts, returned RPC OK and advanced 20 blocks. Soteria's release permissions/build defects are corrected; release 1297 is running | Observe Monerod latency/sync; publish Soteria and validate an eligible backup and restore |
| 2. Reliable nodes | Partial; physical checks needed | Titan-04/14/18 excluded from new placement; Titan-05/06 remain offline. Fresh Titan-04/11 undervoltage was observed | Power/media checks, preservation or replacement of suspect media, then networking/storage/load validation before return. Establish the actual condition of 09/10/16 |
| 3. Recovery coverage | In progress | External datastore restore and second LAN copy verified; 19 application DB dumps completed; scheduled application backup completed at 08:37 UTC; representative Gitea restore passed | Complete eligible PVC coverage, verify Soteria restore, resolve Hermes tenant-2 safely, and provide approved local-only recovery copies |
| 3. Recovery coverage | In progress; cloud uploads blocked | External datastore restore and second LAN copy verified; 19 application DB dumps completed; scheduled application backup completed at 08:37 UTC; representative Gitea restore passed. Backblaze storage-cap failure established | Resolve the account storage cap within an agreed budget, complete eligible PVC coverage and Soteria restore, resolve Hermes tenant-2 safely, and provide approved local-only recovery copies |
| 4. Capacity and placement | Partial | CI agent cap reduced to two; shared PostgreSQL/OpenSearch resource corrections deployed; Vault injectors spread. Longhorn now uses least-effort balancing and one rebuild per node | Account for the measured 2-3.5 GiB storage-process footprint, correct remaining requests/placement, and calculate compatible spare capacity for one worker loss |
| 5. Supported software | Planned; backups established | Installed K3s cohorts identified as 1.31.5 and 1.33.3; titan-db Ubuntu 24.10 is beyond support | Select a compatible supported baseline; stage OS/K3s/storage upgrades with rollback. Plan collation/index repair |
| 6. Maintainability and sustained health | In progress | Operator guide and rollback record published; Ananke stale-event loop fixed; updater failure reporting corrected; overlapping restart/eviction behavior removed | Complete the service recovery checklist, validate remaining automation boundaries, arrange controlled recovery checks, then observe 7-14 days |
@ -30,11 +30,12 @@ from routine software work.
- [ ] Verify an eligible backup and an isolated restore; retain Hermes exclusions.
- [x] Diagnose Monerod's host/attachment/filesystem error before selecting recovery.
- [x] Preserve a pre-recovery snapshot and remount the existing Monerod volume through Flux.
- [ ] Recover Monerod with its existing data and verify application RPC health.
- [x] Recover Monerod with its existing data: get_version returned OK; height advanced from 3775922 to 3775942. Broader get_info requests remain slow.
- [x] Deploy and verify bounded Longhorn balancing/rebuild settings.
- [x] Verify the first scheduled native application PostgreSQL backup completed.
- [x] Recheck controller health: 57/57 active Flux Kustomizations Ready at 21:18 UTC.
- [ ] Complete application checks and the 7-14 day observation period.
- [ ] Resolve Backblaze's storage cap before accepting cloud backup coverage; preserve existing copies.
## Verified changes
@ -109,6 +110,9 @@ became 2/2 Ready without a restart. The first ten-second RPC check timed out;
application responsiveness and steady synchronization remain explicit checks.
No PVC deletion, forced detach, database replacement or manual filesystem repair
was performed. Retain the snapshot until recovery verification is complete.
Subsequent `get_version` calls returned status OK with the existing height
3775922, then 3775942 against target 3776212. Synchronization is making progress;
the broader `get_info` call still timed out at 45 seconds during recovery.
The 21:12 UTC capacity snapshot showed Titan-07/11 CPU requests at 3.53/3.56 of
3.60 allocatable cores; Titan-11 memory requests were 6662 of 6785 MiB. Titan-17/19
@ -134,6 +138,40 @@ Hermes after a transient Longhorn webhook timeout cleared. That controller
snapshot is not a declaration of sufficient failure capacity or completed
hardware repair.
### Cloud backup blocker established
At approximately 21:29 UTC the Longhorn catalog contained 980 backup objects:
952 Error, 14 Completed and 14 without a reported state. The completed entries
were from June/July; 899 retained errors explicitly reported `storage cap
exceeded`. Recent failed attempts on October 3 reported the same condition.
This is not evidence that rotating an expired key will repair uploads.
An authenticated Backblaze authorization check succeeded. The synchronized
Longhorn/Soteria key is non-expiring and has writeFiles/readFiles/listFiles
permissions. Its scope also includes account/bucket/key management; replace
that broad shared scope with dedicated bucket-scoped credentials in a separate
controlled rotation, without revoking other consumers blindly. No credentials
were printed or committed.
Soteria's cached metadata scan at 21:27:15 UTC reported 64,692 visible objects
and 963,207,320,600 bytes in `atlas-soteria`, no new objects in 24 hours, and a
last object modification of July 15. This is a visible-object inventory, not
verified billable storage including historical versions or other buckets.
Backblaze's exact account cap and approved monthly budget are not known from
the supported API checks. The user has been asked to review Caps & Alerts and
provide the budget. No spending cap was raised, old recovery copy removed, or
local-only material uploaded. A readable bucket and an Available backup target
do not establish successful writes.
Hermes tenant-2 remains unresolved: its volume is detached with a share manager
waiting on Titan-14, while an older engine reports running on Titan-23 with an
ERR replica reference. Three replica objects remain on Titan-13/15/17. These are
controller observations, not evidence of empty or lost data. Both consumers
report Running despite the unavailable share. No workspace content was read,
consumer stopped, volume forcibly detached or replica removed during inspection.
Preserving an approved local recovery copy and repairing that specific engine/
share ownership remains an explicit recovery task.
## Confirmed problems needing further work