docs: record storage recovery and remaining health checks

This commit is contained in:
jenkins 2026-10-03 16:20:11 -05:00
parent b0be99b739
commit 0ad4ec866b

View File

@ -7,29 +7,34 @@ Use [Cluster operations](CLUSTER_OPERATIONS.md) for the ordinary operator path.
## Six-priority progress tracker
Started against the user's approved list on 2026-10-03. Updated: 20:45 UTC.
Started against the user's approved list on 2026-10-03. Updated: 21:22 UTC.
"Verified" means the stated check passed; it does not imply a completed soak or
failure drill. Physical repairs and disruptive recovery drills remain separate
from routine software work.
| Priority | Status | Completed evidence | Next action / completion gate |
| --- | --- | --- | --- |
| 1. Immediate software repairs | In progress | Soteria test builds 1293/1294 passed; release 1295 failed its supply-chain gate. Monerod's LMDB open fails with I/O error despite healthy Longhorn replicas | Correct the release findings, publish and validate a backup; repair Monerod's storage path without deleting its database |
| 1. Immediate software repairs | In progress | Monerod remounted its existing database on Titan-19, initialized RPC and became 2/2 Ready with zero restarts. Soteria's release permissions/build defects are corrected; release 1297 is running | Verify responsive Monerod RPC and sustained sync; publish Soteria and validate an eligible backup and restore |
| 2. Reliable nodes | Partial; physical checks needed | Titan-04/14/18 excluded from new placement; Titan-05/06 remain offline. Fresh Titan-04/11 undervoltage was observed | Power/media checks, preservation or replacement of suspect media, then networking/storage/load validation before return. Establish the actual condition of 09/10/16 |
| 3. Recovery coverage | In progress | External datastore restore and second LAN copy verified; 19 application DB dumps completed; scheduled application backup completed at 08:37 UTC; representative Gitea restore passed | Complete eligible PVC coverage, verify Soteria restore, resolve Hermes tenant-2 safely, and provide approved local-only recovery copies |
| 4. Capacity and placement | Partial | CI agent cap reduced to two; shared PostgreSQL/OpenSearch resource corrections deployed; Vault injectors spread | Measure peak demand and I/O, correct remaining requests/placement, and calculate compatible spare capacity for one worker loss |
| 4. Capacity and placement | Partial | CI agent cap reduced to two; shared PostgreSQL/OpenSearch resource corrections deployed; Vault injectors spread. Longhorn now uses least-effort balancing and one rebuild per node | Account for the measured 2-3.5 GiB storage-process footprint, correct remaining requests/placement, and calculate compatible spare capacity for one worker loss |
| 5. Supported software | Planned; backups established | Installed K3s cohorts identified as 1.31.5 and 1.33.3; titan-db Ubuntu 24.10 is beyond support | Select a compatible supported baseline; stage OS/K3s/storage upgrades with rollback. Plan collation/index repair |
| 6. Maintainability and sustained health | In progress | Operator guide and rollback record published; Ananke stale-event loop fixed; updater failure reporting corrected; overlapping restart/eviction behavior removed | Complete the service recovery checklist, validate remaining automation boundaries, arrange controlled recovery checks, then observe 7-14 days |
### Current work and acceptance checks
- [ ] Resolve Soteria release gate findings without disabling the gate.
- [x] Correct Soteria release gate findings without disabling the gate; verify the local configuration scan.
- [x] Replace Soteria's daemonless Docker build with the existing pinned Kaniko builder; validate the pipeline.
- [x] Deploy and verify scoped Soteria state/credential permissions without altering stored policy or usage.
- [ ] Publish the corrected Soteria image through the normal pipeline and Flux.
- [ ] Verify an eligible backup and an isolated restore; retain Hermes exclusions.
- [ ] Diagnose Monerod's host/attachment/filesystem error before selecting recovery.
- [x] Diagnose Monerod's host/attachment/filesystem error before selecting recovery.
- [x] Preserve a pre-recovery snapshot and remount the existing Monerod volume through Flux.
- [ ] Recover Monerod with its existing data and verify application RPC health.
- [x] Deploy and verify bounded Longhorn balancing/rebuild settings.
- [x] Verify the first scheduled native application PostgreSQL backup completed.
- [ ] Recheck health and update this tracker after each deployed repair.
- [x] Recheck controller health: 57/57 active Flux Kustomizations Ready at 21:18 UTC.
- [ ] Complete application checks and the 7-14 day observation period.
## Verified changes
@ -51,6 +56,9 @@ from routine software work.
| Protected local-only data policy | Hermes PVCs excluded from Soteria's cloud policy through Flux | Do not broaden this policy without reviewing data authorization |
| Stopped stale credential-event restart loop | Ananke c73b8fe installed on Titan-db after its complete quality gate; helper generation 796 remained unchanged after normal one-minute checks resumed. Peer 53ad03a no longer runs periodic cluster repairs | Native installer retains `/usr/local/lib/ananke/rollback/ananke.previous`; rolling back restores the defective behavior, so retain a bounded repair interval if needed |
| Corrected premature CI cancellation | Ariadne's configured elapsed-time cutoff is now 120 minutes, verified in the running Deployment | Revert 0180ba66; this returns the 45-minute cutoff that aborted an active Pi build |
| Scoped Soteria permissions | Atlas 46596536 applied; intended state updates and credential reads allowed, unrelated secret reads/listing and Hermes job creation denied; policy/usage fingerprints unchanged | Revert only after reviewing why broader access is required; retain state Secrets |
| Recovered Monerod mount | Snapshot df99cdac, clean detach b119076c, restart 0984164f; existing ext4 filesystem mounted on Titan-19, LMDB opened and RPC initialized | Preserve the snapshot and PVC. Returning to Titan-08 requires storage-path validation; do not replay the temporary scale-zero commit as a normal rollback |
| Bounded Longhorn background I/O | b0be99b7 applied; live settings are least-effort balancing and one concurrent rebuild per node | Revert b0be99b7 and reconcile Longhorn; the versioned settings Job restores best-effort and two rebuilds per node |
Host backup implementation is in commit 80aff498 and
[its operational notes](../infrastructure/host-backup/NOTES.md).
@ -73,9 +81,59 @@ replay or starting K3s. A complete control-plane disaster drill remains outstand
validation passed. Release build 1292 failed during Git checkout with a 404;
current credential/repository checks and build 1293's source checkout succeed.
Builds 1293/1294 subsequently passed. Release retry 1295 failed the supply-chain
gate; its findings are under investigation. Do not claim the runtime
gate on broad ClusterRole secret access, not a dependency vulnerability. Source
3c0c868 removes those permissions and the expired exception; Atlas 46596536
deploys namespaced, resource-named state access. A verified Trivy 0.70.0 local
configuration scan reports zero HIGH/CRITICAL findings using its embedded
checks. Jenkins remains the publication gate. Source 596b2df also replaces a
Docker client with no daemon using the cluster's existing digest-pinned Kaniko
builder, without privileged access. Declarative pipeline validation passed.
Redundant timer build 1296 was stopped before any stage/agent began; release
1297 runs the complete gate with PUBLISH_IMAGES=true. Do not claim the runtime
backup defect is repaired until the new image and a real eligible backup pass.
### October 3 storage recovery and capacity evidence
Titan-08's kernel recorded I/O failure, an aborted ext4 journal and a read-only
remount of Monerod's device at 11:31 UTC. Longhorn still reported two healthy
replicas on Titan-15/17. Replica health therefore did not establish filesystem
health. A retained local snapshot, `monerod-before-recovery-20261003`, became
ready at 20:52 UTC. It is a recovery point, not an independent backup.
Monerod was stopped through its Flux-tracked Deployment; CSI completed unstage
at 21:04 UTC and the volume detached normally. Restoring one replica with
Titan-08 excluded moved it to Titan-19. An initial attachment lookup failure
resolved through the normal controller retry. The existing ext4 filesystem
mounted read/write, LMDB opened, RPC initialized at 21:16:58 UTC and the pod
became 2/2 Ready without a restart. The first ten-second RPC check timed out;
application responsiveness and steady synchronization remain explicit checks.
No PVC deletion, forced detach, database replacement or manual filesystem repair
was performed. Retain the snapshot until recovery verification is complete.
The 21:12 UTC capacity snapshot showed Titan-07/11 CPU requests at 3.53/3.56 of
3.60 allocatable cores; Titan-11 memory requests were 6662 of 6785 MiB. Titan-17/19
working memory was 5994/6102 of approximately 6656 MiB. Longhorn instance-manager
pods have CPU requests but no memory requests in this installed configuration.
Their observed 24-hour memory peaks were 3525 MiB on Titan-17, 3176 on Titan-19,
2914 on Titan-13 and 2594 on Titan-15. Those bytes cannot be treated as spare
application capacity merely because Kubernetes has not reserved them.
At 21:16 UTC Titan-19 reported high I/O and memory pressure while the blockchain
opened and CI ran. The native Longhorn settings were changed from best-effort
to least-effort replica balancing and from two to one concurrent rebuild per
node, verified live at 21:18 UTC. This reduces elective replica movement and
concurrent recovery I/O; it does not repair slow media, disable required replica
recovery, or change desired replica counts. Multiple rebuilds may take longer.
The existing 600-second replica replenishment delay is unchanged. Per-volume
overrides remain unchanged (28 least-effort, three disabled at inspection).
Setting semantics were checked against the installed
[Longhorn 1.8.2 source](https://github.com/longhorn/longhorn-manager/blob/v1.8.2/types/setting.go).
All 57 active Flux Kustomizations were Ready again at 21:18 UTC, including
Hermes after a transient Longhorn webhook timeout cleared. That controller
snapshot is not a declaration of sufficient failure capacity or completed
hardware repair.
## Confirmed problems needing further work