docs: record storage recovery and remaining health checks
This commit is contained in:
parent
b0be99b739
commit
0ad4ec866b
@ -7,29 +7,34 @@ Use [Cluster operations](CLUSTER_OPERATIONS.md) for the ordinary operator path.
|
||||
|
||||
## Six-priority progress tracker
|
||||
|
||||
Started against the user's approved list on 2026-10-03. Updated: 20:45 UTC.
|
||||
Started against the user's approved list on 2026-10-03. Updated: 21:22 UTC.
|
||||
"Verified" means the stated check passed; it does not imply a completed soak or
|
||||
failure drill. Physical repairs and disruptive recovery drills remain separate
|
||||
from routine software work.
|
||||
|
||||
| Priority | Status | Completed evidence | Next action / completion gate |
|
||||
| --- | --- | --- | --- |
|
||||
| 1. Immediate software repairs | In progress | Soteria test builds 1293/1294 passed; release 1295 failed its supply-chain gate. Monerod's LMDB open fails with I/O error despite healthy Longhorn replicas | Correct the release findings, publish and validate a backup; repair Monerod's storage path without deleting its database |
|
||||
| 1. Immediate software repairs | In progress | Monerod remounted its existing database on Titan-19, initialized RPC and became 2/2 Ready with zero restarts. Soteria's release permissions/build defects are corrected; release 1297 is running | Verify responsive Monerod RPC and sustained sync; publish Soteria and validate an eligible backup and restore |
|
||||
| 2. Reliable nodes | Partial; physical checks needed | Titan-04/14/18 excluded from new placement; Titan-05/06 remain offline. Fresh Titan-04/11 undervoltage was observed | Power/media checks, preservation or replacement of suspect media, then networking/storage/load validation before return. Establish the actual condition of 09/10/16 |
|
||||
| 3. Recovery coverage | In progress | External datastore restore and second LAN copy verified; 19 application DB dumps completed; scheduled application backup completed at 08:37 UTC; representative Gitea restore passed | Complete eligible PVC coverage, verify Soteria restore, resolve Hermes tenant-2 safely, and provide approved local-only recovery copies |
|
||||
| 4. Capacity and placement | Partial | CI agent cap reduced to two; shared PostgreSQL/OpenSearch resource corrections deployed; Vault injectors spread | Measure peak demand and I/O, correct remaining requests/placement, and calculate compatible spare capacity for one worker loss |
|
||||
| 4. Capacity and placement | Partial | CI agent cap reduced to two; shared PostgreSQL/OpenSearch resource corrections deployed; Vault injectors spread. Longhorn now uses least-effort balancing and one rebuild per node | Account for the measured 2-3.5 GiB storage-process footprint, correct remaining requests/placement, and calculate compatible spare capacity for one worker loss |
|
||||
| 5. Supported software | Planned; backups established | Installed K3s cohorts identified as 1.31.5 and 1.33.3; titan-db Ubuntu 24.10 is beyond support | Select a compatible supported baseline; stage OS/K3s/storage upgrades with rollback. Plan collation/index repair |
|
||||
| 6. Maintainability and sustained health | In progress | Operator guide and rollback record published; Ananke stale-event loop fixed; updater failure reporting corrected; overlapping restart/eviction behavior removed | Complete the service recovery checklist, validate remaining automation boundaries, arrange controlled recovery checks, then observe 7-14 days |
|
||||
|
||||
### Current work and acceptance checks
|
||||
|
||||
- [ ] Resolve Soteria release gate findings without disabling the gate.
|
||||
- [x] Correct Soteria release gate findings without disabling the gate; verify the local configuration scan.
|
||||
- [x] Replace Soteria's daemonless Docker build with the existing pinned Kaniko builder; validate the pipeline.
|
||||
- [x] Deploy and verify scoped Soteria state/credential permissions without altering stored policy or usage.
|
||||
- [ ] Publish the corrected Soteria image through the normal pipeline and Flux.
|
||||
- [ ] Verify an eligible backup and an isolated restore; retain Hermes exclusions.
|
||||
- [ ] Diagnose Monerod's host/attachment/filesystem error before selecting recovery.
|
||||
- [x] Diagnose Monerod's host/attachment/filesystem error before selecting recovery.
|
||||
- [x] Preserve a pre-recovery snapshot and remount the existing Monerod volume through Flux.
|
||||
- [ ] Recover Monerod with its existing data and verify application RPC health.
|
||||
- [x] Deploy and verify bounded Longhorn balancing/rebuild settings.
|
||||
- [x] Verify the first scheduled native application PostgreSQL backup completed.
|
||||
- [ ] Recheck health and update this tracker after each deployed repair.
|
||||
- [x] Recheck controller health: 57/57 active Flux Kustomizations Ready at 21:18 UTC.
|
||||
- [ ] Complete application checks and the 7-14 day observation period.
|
||||
|
||||
## Verified changes
|
||||
|
||||
@ -51,6 +56,9 @@ from routine software work.
|
||||
| Protected local-only data policy | Hermes PVCs excluded from Soteria's cloud policy through Flux | Do not broaden this policy without reviewing data authorization |
|
||||
| Stopped stale credential-event restart loop | Ananke c73b8fe installed on Titan-db after its complete quality gate; helper generation 796 remained unchanged after normal one-minute checks resumed. Peer 53ad03a no longer runs periodic cluster repairs | Native installer retains `/usr/local/lib/ananke/rollback/ananke.previous`; rolling back restores the defective behavior, so retain a bounded repair interval if needed |
|
||||
| Corrected premature CI cancellation | Ariadne's configured elapsed-time cutoff is now 120 minutes, verified in the running Deployment | Revert 0180ba66; this returns the 45-minute cutoff that aborted an active Pi build |
|
||||
| Scoped Soteria permissions | Atlas 46596536 applied; intended state updates and credential reads allowed, unrelated secret reads/listing and Hermes job creation denied; policy/usage fingerprints unchanged | Revert only after reviewing why broader access is required; retain state Secrets |
|
||||
| Recovered Monerod mount | Snapshot df99cdac, clean detach b119076c, restart 0984164f; existing ext4 filesystem mounted on Titan-19, LMDB opened and RPC initialized | Preserve the snapshot and PVC. Returning to Titan-08 requires storage-path validation; do not replay the temporary scale-zero commit as a normal rollback |
|
||||
| Bounded Longhorn background I/O | b0be99b7 applied; live settings are least-effort balancing and one concurrent rebuild per node | Revert b0be99b7 and reconcile Longhorn; the versioned settings Job restores best-effort and two rebuilds per node |
|
||||
|
||||
Host backup implementation is in commit 80aff498 and
|
||||
[its operational notes](../infrastructure/host-backup/NOTES.md).
|
||||
@ -73,9 +81,59 @@ replay or starting K3s. A complete control-plane disaster drill remains outstand
|
||||
validation passed. Release build 1292 failed during Git checkout with a 404;
|
||||
current credential/repository checks and build 1293's source checkout succeed.
|
||||
Builds 1293/1294 subsequently passed. Release retry 1295 failed the supply-chain
|
||||
gate; its findings are under investigation. Do not claim the runtime
|
||||
gate on broad ClusterRole secret access, not a dependency vulnerability. Source
|
||||
3c0c868 removes those permissions and the expired exception; Atlas 46596536
|
||||
deploys namespaced, resource-named state access. A verified Trivy 0.70.0 local
|
||||
configuration scan reports zero HIGH/CRITICAL findings using its embedded
|
||||
checks. Jenkins remains the publication gate. Source 596b2df also replaces a
|
||||
Docker client with no daemon using the cluster's existing digest-pinned Kaniko
|
||||
builder, without privileged access. Declarative pipeline validation passed.
|
||||
Redundant timer build 1296 was stopped before any stage/agent began; release
|
||||
1297 runs the complete gate with PUBLISH_IMAGES=true. Do not claim the runtime
|
||||
backup defect is repaired until the new image and a real eligible backup pass.
|
||||
|
||||
### October 3 storage recovery and capacity evidence
|
||||
|
||||
Titan-08's kernel recorded I/O failure, an aborted ext4 journal and a read-only
|
||||
remount of Monerod's device at 11:31 UTC. Longhorn still reported two healthy
|
||||
replicas on Titan-15/17. Replica health therefore did not establish filesystem
|
||||
health. A retained local snapshot, `monerod-before-recovery-20261003`, became
|
||||
ready at 20:52 UTC. It is a recovery point, not an independent backup.
|
||||
|
||||
Monerod was stopped through its Flux-tracked Deployment; CSI completed unstage
|
||||
at 21:04 UTC and the volume detached normally. Restoring one replica with
|
||||
Titan-08 excluded moved it to Titan-19. An initial attachment lookup failure
|
||||
resolved through the normal controller retry. The existing ext4 filesystem
|
||||
mounted read/write, LMDB opened, RPC initialized at 21:16:58 UTC and the pod
|
||||
became 2/2 Ready without a restart. The first ten-second RPC check timed out;
|
||||
application responsiveness and steady synchronization remain explicit checks.
|
||||
No PVC deletion, forced detach, database replacement or manual filesystem repair
|
||||
was performed. Retain the snapshot until recovery verification is complete.
|
||||
|
||||
The 21:12 UTC capacity snapshot showed Titan-07/11 CPU requests at 3.53/3.56 of
|
||||
3.60 allocatable cores; Titan-11 memory requests were 6662 of 6785 MiB. Titan-17/19
|
||||
working memory was 5994/6102 of approximately 6656 MiB. Longhorn instance-manager
|
||||
pods have CPU requests but no memory requests in this installed configuration.
|
||||
Their observed 24-hour memory peaks were 3525 MiB on Titan-17, 3176 on Titan-19,
|
||||
2914 on Titan-13 and 2594 on Titan-15. Those bytes cannot be treated as spare
|
||||
application capacity merely because Kubernetes has not reserved them.
|
||||
|
||||
At 21:16 UTC Titan-19 reported high I/O and memory pressure while the blockchain
|
||||
opened and CI ran. The native Longhorn settings were changed from best-effort
|
||||
to least-effort replica balancing and from two to one concurrent rebuild per
|
||||
node, verified live at 21:18 UTC. This reduces elective replica movement and
|
||||
concurrent recovery I/O; it does not repair slow media, disable required replica
|
||||
recovery, or change desired replica counts. Multiple rebuilds may take longer.
|
||||
The existing 600-second replica replenishment delay is unchanged. Per-volume
|
||||
overrides remain unchanged (28 least-effort, three disabled at inspection).
|
||||
Setting semantics were checked against the installed
|
||||
[Longhorn 1.8.2 source](https://github.com/longhorn/longhorn-manager/blob/v1.8.2/types/setting.go).
|
||||
|
||||
All 57 active Flux Kustomizations were Ready again at 21:18 UTC, including
|
||||
Hermes after a transient Longhorn webhook timeout cleared. That controller
|
||||
snapshot is not a declaration of sufficient failure capacity or completed
|
||||
hardware repair.
|
||||
|
||||
|
||||
## Confirmed problems needing further work
|
||||
|
||||
|
||||
Loading…
x
Reference in New Issue
Block a user