docs: record verified backup storage cap blocker
Some checks failed
Tests / Declarative: Post Actions failed: 45, skipped: 89, passed: 4270
Some checks failed
Tests / Declarative: Post Actions failed: 45, skipped: 89, passed: 4270
This commit is contained in:
parent
5d50d2f8c8
commit
b41dab9886
@ -176,6 +176,14 @@ control plane boot. Live Longhorn snapshots and database-native dumps offer
|
||||
different consistency guarantees. A recent backup indicator is not a restore
|
||||
test. Keep database restore checks and service recovery drills explicit.
|
||||
|
||||
For cloud backups, inspect completed backup timestamps, not just the backup
|
||||
target's Available status or a successful bucket listing. The October 3 audit
|
||||
found a readable Backblaze bucket rejecting uploads with `storage cap exceeded`.
|
||||
Check Backblaze Caps & Alerts and agree the storage budget before changing it.
|
||||
Do not delete old recovery points merely to make uploads work. Distinguish
|
||||
visible-object bytes from billed storage, including retained versions. The
|
||||
current diagnosis and counts are in the implementation record.
|
||||
|
||||
The Hermes namespace is excluded from cloud backup policy. Its workspaces can
|
||||
contain material restricted to local infrastructure. Do not remove that exclusion
|
||||
to make a coverage dashboard greener.
|
||||
|
||||
@ -7,16 +7,16 @@ Use [Cluster operations](CLUSTER_OPERATIONS.md) for the ordinary operator path.
|
||||
|
||||
## Six-priority progress tracker
|
||||
|
||||
Started against the user's approved list on 2026-10-03. Updated: 21:22 UTC.
|
||||
Started against the user's approved list on 2026-10-03. Updated: 21:32 UTC.
|
||||
"Verified" means the stated check passed; it does not imply a completed soak or
|
||||
failure drill. Physical repairs and disruptive recovery drills remain separate
|
||||
from routine software work.
|
||||
|
||||
| Priority | Status | Completed evidence | Next action / completion gate |
|
||||
| --- | --- | --- | --- |
|
||||
| 1. Immediate software repairs | In progress | Monerod remounted its existing database on Titan-19, initialized RPC and became 2/2 Ready with zero restarts. Soteria's release permissions/build defects are corrected; release 1297 is running | Verify responsive Monerod RPC and sustained sync; publish Soteria and validate an eligible backup and restore |
|
||||
| 1. Immediate software repairs | In progress | Monerod remounted its existing database on Titan-19, became 2/2 Ready with zero restarts, returned RPC OK and advanced 20 blocks. Soteria's release permissions/build defects are corrected; release 1297 is running | Observe Monerod latency/sync; publish Soteria and validate an eligible backup and restore |
|
||||
| 2. Reliable nodes | Partial; physical checks needed | Titan-04/14/18 excluded from new placement; Titan-05/06 remain offline. Fresh Titan-04/11 undervoltage was observed | Power/media checks, preservation or replacement of suspect media, then networking/storage/load validation before return. Establish the actual condition of 09/10/16 |
|
||||
| 3. Recovery coverage | In progress | External datastore restore and second LAN copy verified; 19 application DB dumps completed; scheduled application backup completed at 08:37 UTC; representative Gitea restore passed | Complete eligible PVC coverage, verify Soteria restore, resolve Hermes tenant-2 safely, and provide approved local-only recovery copies |
|
||||
| 3. Recovery coverage | In progress; cloud uploads blocked | External datastore restore and second LAN copy verified; 19 application DB dumps completed; scheduled application backup completed at 08:37 UTC; representative Gitea restore passed. Backblaze storage-cap failure established | Resolve the account storage cap within an agreed budget, complete eligible PVC coverage and Soteria restore, resolve Hermes tenant-2 safely, and provide approved local-only recovery copies |
|
||||
| 4. Capacity and placement | Partial | CI agent cap reduced to two; shared PostgreSQL/OpenSearch resource corrections deployed; Vault injectors spread. Longhorn now uses least-effort balancing and one rebuild per node | Account for the measured 2-3.5 GiB storage-process footprint, correct remaining requests/placement, and calculate compatible spare capacity for one worker loss |
|
||||
| 5. Supported software | Planned; backups established | Installed K3s cohorts identified as 1.31.5 and 1.33.3; titan-db Ubuntu 24.10 is beyond support | Select a compatible supported baseline; stage OS/K3s/storage upgrades with rollback. Plan collation/index repair |
|
||||
| 6. Maintainability and sustained health | In progress | Operator guide and rollback record published; Ananke stale-event loop fixed; updater failure reporting corrected; overlapping restart/eviction behavior removed | Complete the service recovery checklist, validate remaining automation boundaries, arrange controlled recovery checks, then observe 7-14 days |
|
||||
@ -30,11 +30,12 @@ from routine software work.
|
||||
- [ ] Verify an eligible backup and an isolated restore; retain Hermes exclusions.
|
||||
- [x] Diagnose Monerod's host/attachment/filesystem error before selecting recovery.
|
||||
- [x] Preserve a pre-recovery snapshot and remount the existing Monerod volume through Flux.
|
||||
- [ ] Recover Monerod with its existing data and verify application RPC health.
|
||||
- [x] Recover Monerod with its existing data: get_version returned OK; height advanced from 3775922 to 3775942. Broader get_info requests remain slow.
|
||||
- [x] Deploy and verify bounded Longhorn balancing/rebuild settings.
|
||||
- [x] Verify the first scheduled native application PostgreSQL backup completed.
|
||||
- [x] Recheck controller health: 57/57 active Flux Kustomizations Ready at 21:18 UTC.
|
||||
- [ ] Complete application checks and the 7-14 day observation period.
|
||||
- [ ] Resolve Backblaze's storage cap before accepting cloud backup coverage; preserve existing copies.
|
||||
|
||||
## Verified changes
|
||||
|
||||
@ -109,6 +110,9 @@ became 2/2 Ready without a restart. The first ten-second RPC check timed out;
|
||||
application responsiveness and steady synchronization remain explicit checks.
|
||||
No PVC deletion, forced detach, database replacement or manual filesystem repair
|
||||
was performed. Retain the snapshot until recovery verification is complete.
|
||||
Subsequent `get_version` calls returned status OK with the existing height
|
||||
3775922, then 3775942 against target 3776212. Synchronization is making progress;
|
||||
the broader `get_info` call still timed out at 45 seconds during recovery.
|
||||
|
||||
The 21:12 UTC capacity snapshot showed Titan-07/11 CPU requests at 3.53/3.56 of
|
||||
3.60 allocatable cores; Titan-11 memory requests were 6662 of 6785 MiB. Titan-17/19
|
||||
@ -134,6 +138,40 @@ Hermes after a transient Longhorn webhook timeout cleared. That controller
|
||||
snapshot is not a declaration of sufficient failure capacity or completed
|
||||
hardware repair.
|
||||
|
||||
### Cloud backup blocker established
|
||||
|
||||
At approximately 21:29 UTC the Longhorn catalog contained 980 backup objects:
|
||||
952 Error, 14 Completed and 14 without a reported state. The completed entries
|
||||
were from June/July; 899 retained errors explicitly reported `storage cap
|
||||
exceeded`. Recent failed attempts on October 3 reported the same condition.
|
||||
This is not evidence that rotating an expired key will repair uploads.
|
||||
|
||||
An authenticated Backblaze authorization check succeeded. The synchronized
|
||||
Longhorn/Soteria key is non-expiring and has writeFiles/readFiles/listFiles
|
||||
permissions. Its scope also includes account/bucket/key management; replace
|
||||
that broad shared scope with dedicated bucket-scoped credentials in a separate
|
||||
controlled rotation, without revoking other consumers blindly. No credentials
|
||||
were printed or committed.
|
||||
|
||||
Soteria's cached metadata scan at 21:27:15 UTC reported 64,692 visible objects
|
||||
and 963,207,320,600 bytes in `atlas-soteria`, no new objects in 24 hours, and a
|
||||
last object modification of July 15. This is a visible-object inventory, not
|
||||
verified billable storage including historical versions or other buckets.
|
||||
Backblaze's exact account cap and approved monthly budget are not known from
|
||||
the supported API checks. The user has been asked to review Caps & Alerts and
|
||||
provide the budget. No spending cap was raised, old recovery copy removed, or
|
||||
local-only material uploaded. A readable bucket and an Available backup target
|
||||
do not establish successful writes.
|
||||
|
||||
Hermes tenant-2 remains unresolved: its volume is detached with a share manager
|
||||
waiting on Titan-14, while an older engine reports running on Titan-23 with an
|
||||
ERR replica reference. Three replica objects remain on Titan-13/15/17. These are
|
||||
controller observations, not evidence of empty or lost data. Both consumers
|
||||
report Running despite the unavailable share. No workspace content was read,
|
||||
consumer stopped, volume forcibly detached or replica removed during inspection.
|
||||
Preserving an approved local recovery copy and repairing that specific engine/
|
||||
share ownership remains an explicit recovery task.
|
||||
|
||||
|
||||
## Confirmed problems needing further work
|
||||
|
||||
|
||||
Loading…
x
Reference in New Issue
Block a user