386 lines
32 KiB
Markdown
386 lines
32 KiB
Markdown
# Cluster stabilization - implementation record
|
|
|
|
This record begins on 2026-10-03 UTC, following authorization to implement the
|
|
[configuration-first plan](cluster-audit-20261002/INTEGRATED_PLAN.md).
|
|
The October 2 audit remains a historical baseline, not a current health claim.
|
|
Use [Cluster operations](CLUSTER_OPERATIONS.md) for the ordinary operator path.
|
|
|
|
## Six-priority progress tracker
|
|
|
|
Started against the user's approved list on 2026-10-03. Updated: 2026-10-04 06:20 UTC.
|
|
"Verified" means the stated check passed; it does not imply a completed soak or
|
|
failure drill. Physical repairs and disruptive recovery drills remain separate
|
|
from routine software work.
|
|
|
|
| Priority | Status | Completed evidence | Next action / completion gate |
|
|
| --- | --- | --- | --- |
|
|
| 1. Immediate software repairs | In progress; Soteria publication blocked | Monerod remains 2/2 Ready; its daemon has zero restarts, status sidecar one. Soteria release 1297 passed enforced gates but failed Go VCS stamping inside the image build; runtime remains 0.1.0-120 | Repair image build metadata, publish Soteria and validate an eligible backup and restore; continue Monerod latency/sync checks |
|
|
| 2. Reliable nodes | Software repairs applied; hardware and identity checks remain | Full root administration verified on 21 reachable machines; Vault verification metadata saved. Repaired credentials on 12/13/19/20/21. Titan-04/11/14/18 excluded from new placement. Titan-19 log rotation repaired | Verify changed SSH identities on 09/10; recover 05/06/16; address fresh power faults and runtime-media stalls. Complete Metis image publication and physical replacement-image validation |
|
|
| 3. Recovery coverage | In progress; cloud uploads blocked | External datastore restore and second LAN copy verified; 19 application DB dumps completed; scheduled application backup completed at 08:37 UTC; representative Gitea restore passed. Backblaze storage-cap failure established | Resolve the account storage cap within an agreed budget, complete eligible PVC coverage and Soteria restore, resolve Hermes tenant-2 safely, and provide approved local-only recovery copies |
|
|
| 4. Capacity and placement | Partial | CI agent cap reduced to two; shared PostgreSQL/OpenSearch resource corrections deployed; Vault injectors spread. Longhorn now uses least-effort balancing and one rebuild per node | Account for the measured 2-3.5 GiB storage-process footprint, correct remaining requests/placement, and calculate compatible spare capacity for one worker loss |
|
|
| 5. Supported software | Planned; backups established | Installed K3s cohorts identified as 1.31.5 and 1.33.3; titan-db Ubuntu 24.10 is beyond support | Select a compatible supported baseline; stage OS/K3s/storage upgrades with rollback. Plan collation/index repair |
|
|
| 6. Maintainability and sustained health | In progress | Operator guide and rollback record published; Ananke stale-event loop fixed; updater failure reporting corrected; overlapping restart/eviction behavior removed | Complete the service recovery checklist, validate remaining automation boundaries, arrange controlled recovery checks, then observe 7-14 days |
|
|
|
|
### Current work and acceptance checks
|
|
|
|
- [x] Correct Soteria release gate findings without disabling the gate; verify the local configuration scan.
|
|
- [x] Replace Soteria's daemonless Docker build with the existing pinned Kaniko builder; validate the pipeline.
|
|
- [x] Deploy and verify scoped Soteria state/credential permissions without altering stored policy or usage.
|
|
- [ ] Publish the corrected Soteria image through the normal pipeline and Flux.
|
|
- [ ] Verify an eligible backup and an isolated restore; retain Hermes exclusions.
|
|
- [x] Diagnose Monerod's host/attachment/filesystem error before selecting recovery.
|
|
- [x] Preserve a pre-recovery snapshot and remount the existing Monerod volume through Flux.
|
|
- [x] Recover Monerod with its existing data: get_version returned OK; height advanced from 3775922 to 3775942. Broader get_info requests remain slow.
|
|
- [x] Deploy and verify bounded Longhorn balancing/rebuild settings.
|
|
- [x] Verify the first scheduled native application PostgreSQL backup completed.
|
|
- [x] Recheck controller health: 57/57 active Flux Kustomizations and 19/21 nodes Ready at 2026-10-04 04:38 UTC; Titan-05/06 remain offline.
|
|
- [ ] Complete application checks and the 7-14 day observation period.
|
|
- [ ] Resolve Backblaze's storage cap before accepting cloud backup coverage; preserve existing copies.
|
|
|
|
## Node access and software repair - October 4
|
|
|
|
The user prioritized node administration and software repairs before hardware
|
|
replacement. No node was reflashed or rebooted during this work.
|
|
|
|
- [x] Verify full administrative execution on all 21 reachable machines,
|
|
including the datastore and bastion; save the verification timestamp, admin
|
|
username, access method and root-password state in Vault custom metadata.
|
|
- [x] Repair Titan-12/13/19: SSH keys worked, but atlas was locked and the root
|
|
password differed from Vault. Both accounts now match their existing Vault
|
|
passwords. Fresh independent password-backed sudo reached UID 0.
|
|
- [x] Repair the existing unlocked root passwords on Titan-20/21 to match Vault.
|
|
- [x] Remove the obsolete Metis passwordless command grants on 12/13/19 after
|
|
verifying password-backed sudo; repeat the sudo verification afterward.
|
|
- [x] Fix the Vault bootstrap writer: preserve all fields, use compare-and-set,
|
|
skip unchanged records, keep temporary files private, remove content-bearing
|
|
errors, and retain the completed Job so Flux does not recreate it after TTL.
|
|
Live Job -5 changed only Titan-23's incorrect .23 address to .24; 25 records
|
|
were unchanged. Atlas commit 1bd8ae31.
|
|
- [x] Exclude Titan-11 from new placement through Flux (c2f41899); existing pods
|
|
continue running. Its fresh voltage faults remain a hardware issue.
|
|
- [x] Repair Titan-19's logrotate post-hook failure: armbian-ramlog was inactive
|
|
but ENABLED=true, copying between obsolete RAM-log locations although /var/log
|
|
already points to external ext4 storage. The guarded versioned script disabled
|
|
those hooks; the native logrotate service completed with result success/status 0.
|
|
- [x] Apply the same guarded repair on Titan-14 after confirming its RAM-log unit
|
|
was already masked/inactive and logs use external ext4 storage. Logrotate then
|
|
completed with result success/status 0. The later short I/O sample was quiet;
|
|
this does not resolve its earlier intermittent storage stalls.
|
|
- [ ] Migrate the active 50 MiB RAM-log filesystems on Titan-12/18 in a controlled
|
|
node maintenance cycle. The obsolete-hook script intentionally refuses them.
|
|
Their native stop action lazily unmounts /var/log while processes may still
|
|
hold files open; do not treat changing ENABLED as a complete live migration.
|
|
Preserve logs and arrange workload evacuation/restart before removing the mount.
|
|
- [x] Correct Ariadne's memory reservation/limit from 128/512 MiB to 512/1024 MiB
|
|
after repeated cgroup OOM kills. The replacement is 2/2 Ready with zero restarts.
|
|
This supplies headroom; it does not establish that its memory growth is bounded.
|
|
- [x] Publish Metis source fixes 1c110c9 and 6213159: native first-boot unit,
|
|
per-injection pending marker, serialized setup, propagated password failures,
|
|
removal of default passwordless grants, cleanup of applied bootstrap passwords,
|
|
and correct file permissions/native systemd links in both image-writing paths.
|
|
- [ ] Complete the normal Metis image publication and Flux rollout. Build 400 was
|
|
still active at the latest check; do not claim the running 0.1.0-399 contains
|
|
these source changes. The next replacement image still needs a physical boot
|
|
check before these changes are considered proven on hardware.
|
|
- [x] Enable Ananke's existing password-backed host-action path. Atlas 886dbeee
|
|
adds 13 atlas-password mappings to the existing Vault CSI synchronizer; all
|
|
matched Vault, with native two-minute rotation enabled. No root passwords are
|
|
copied into that synchronizer. Source Ananke 41fd207 sets the namespace/name
|
|
template and corrects Titan-24's administrator to tethys. Both live host configs
|
|
were patched without changing unrelated settings and passed Ananke validation.
|
|
- [x] Verify read-only privileged actions using Ananke's own coordinator SSH
|
|
identity on all 13 affected workers. Both Ananke services were restarted
|
|
separately and returned active/running with zero automatic restarts. Jenkins
|
|
credential-helper generation remained 796.
|
|
- [x] Restore the peer's missing SSH public key on 04/07/08/11/12/13/19/20/21.
|
|
The key was derived on the authenticated Titan-24 host and matched the existing
|
|
Vault recovery public key. New entries restrict their source to 192.168.22.26;
|
|
existing administrator keys and SSH policies were preserved. The peer's own
|
|
identity then passed password-backed sudo on all 13 workers, plus its existing
|
|
privileged path on 0a/0b/0c/db/22. These are actual remote systemctl version
|
|
checks, not just successful key installation.
|
|
- [ ] Resolve the remaining network/physical findings below.
|
|
|
|
### Current node distinctions
|
|
|
|
| Nodes | Established evidence | Remaining action |
|
|
| --- | --- | --- |
|
|
| 12/13/19 | atlas/root passwords restored; root SSH keys preserved; password-backed sudo works after legacy grant retirement | Observe stability; use the corrected recovery image for future rebuilds |
|
|
| 20/21 | atlas sudo already worked; root passwords now match Vault | Continue workload memory diagnosis; Jetson-21 recorded a Data Prepper cgroup OOM |
|
|
| 04/11 | Fresh undervoltage messages continued during this audit, despite sufficient disk space | Onboard EXT5V samples were 4.82132 V (04) and 4.75968 V (11). Check delivered power/cables/peripherals; do not clear quarantine based on one sample |
|
|
| 14 | Runtime USB flash; 39-45% iowait in the first sample, quiet later, no new kernel I/O error in the two-hour sample; stale RAM-log hooks corrected | Preserve data; investigate intermittent storage stalls and validate sustained operation before return; hardware failure is not proven by I/O pressure alone |
|
|
| 18 | Quarantine has reduced current I/O pressure; its 50 MiB RAM-log filesystem is still active | Keep quarantined until sustained storage checks pass; do not apply the inactive-RAM-log repair script to it |
|
|
| 05 | Answers LAN ping at .31; expected SSH, kubelet and common recovery ports refuse connections | Console/user-space service check; this is not evidence that the machine is powered off |
|
|
| 09/10 | Answer LAN ping and SSH on port 22; recorded 2277 port is closed; host keys differ from saved keys | Confirm reinstall/image identity before supplying passwords; no host-key checks bypassed |
|
|
| 06/16 | No LAN ping response; neighbor resolution incomplete; SSH unavailable | Power, link and boot-media checks needed |
|
|
|
|
The 12 unlocked root passwords match Vault. Nine machines retain locked direct
|
|
root accounts (0a/0b/0c/04/07/08/11/db/jh); their administrator account plus sudo
|
|
still provides full root control. Root-account lock state is recorded explicitly
|
|
in Vault instead of implying that every root_password field is a usable login.
|
|
|
|
Validation: 19 Atlas regression tests passed (credential targeting, secret-safe
|
|
output, legacy-grant preservation, Vault no-op/CAS behavior, preservation/idempotency
|
|
of the Ananke config patch, and restricted peer-key enrollment); relevant Kustomize
|
|
builds, client dry-runs and focused Flux diffs passed. Metis's complete Go tests
|
|
and separate docs/LOC/vet/per-file coverage gate passed. Tests include a real
|
|
throwaway ext4 image for systemd link injection and mocked password-setup failure,
|
|
retry and stale image markers. No live node was burned to test recovery. All 57 active Flux Kustomizations
|
|
remained Ready after these changes; 19/21 Kubernetes nodes are Ready.
|
|
|
|
Ananke still depends on an available Kubernetes API to retrieve its synchronized
|
|
worker sudo credential. This is its existing implementation, not a new offline
|
|
credential cache. Control-plane administration remains available through its
|
|
existing path; use the dated manual SSH/Vault access record for intervention
|
|
outside the automated recovery flow. Do not claim a full power-loss drill from
|
|
read-only connection tests.
|
|
|
|
Rollback: Kubernetes changes use their focused Git commit and Flux. Titan-14/19's
|
|
previous RAM-log config is root-only under
|
|
`/var/lib/atlas-maintenance/ramlog-before-20261004/`; restoring it reintroduces the
|
|
obsolete hooks. Retired sudo grants are root-only under
|
|
`/var/lib/atlas-maintenance/legacy-sudo-20261004/`; do not restore passwordless
|
|
access merely to avoid using the now-working password. Password repairs use the
|
|
already stored Vault credentials; no secret values were changed or committed.
|
|
Native Ananke config backups are root-only under
|
|
`/var/lib/atlas-maintenance/ananke-access-before-20261004/` on both hosts. Reverting
|
|
its lookup settings would remove automated password-backed access; retain the
|
|
working administrator passwords. Its source profiles and the versioned Atlas
|
|
configuration helper document the change.
|
|
Peer key enrollment uses `scripts/install_ananke_peer_key.py`. Rollback removes
|
|
only the newly appended `ananke-tethys-recovery` line after checking another
|
|
administrator session still works; do not replace the whole authorized_keys file.
|
|
|
|
## Verified changes
|
|
|
|
| Change | Evidence | Rollback |
|
|
| --- | --- | --- |
|
|
| Native external PostgreSQL backup and second LAN copy | A 54,083,817-byte custom dump restored into a disposable PostgreSQL 16 instance in 11 seconds with 22,350 kine rows; no TCP listener. Both timers enabled; checksum and matching-token checks pass | Disable the two timers; retain all recovery bundles |
|
|
| Retired unconditional k3s-agent restart DaemonSet | Flux preview removed only that DaemonSet; reconciliation completed and the DaemonSet is absent | Revert f04c84ee, understanding that returning the DaemonSet restarts agents |
|
|
| Recovered metrics Pushgateway | Existing volume attached healthy on titan-19; replacement pod Ready. No volume deletion or data replacement | Revert the focused monitoring placement commits after the original host is repaired |
|
|
| Removed eviction for historical restart counts | Descheduler manifest validates and Flux applied the change | Revert fa9251ae |
|
|
| Restored GitOps UI Deployment | Helm drift correction recreated weave-gitops; Deployment 1/1 Ready | Revert 6c9398ea to disable ongoing drift correction; this does not remove recovered resources |
|
|
| Aligned Flux definition ownership | Removed creation-only policy and adopted already-active service state; no service paths or source refs changed | Revert the focused Flux commits; review suspension fields before doing so |
|
|
| Backup freshness alerts | Both datastore-copy timestamps are scraped; native Grafana provisioning reload returned HTTP 200 | Revert the focused monitoring change; backups continue independently |
|
|
| Quarantined stalled runtime media | Titan-14 and Titan-18 are SchedulingDisabled via `infrastructure/core/node-maintenance.yaml`; no forced storage detach | After repair, change `unschedulable` to false in Git and verify before removing the prune-disabled Node declaration |
|
|
| Bounded CI concurrency | Jenkins controller Ready after applying the two-agent cap; existing agents reconnected | Revert 440244f3 and reconcile Jenkins; allow a controller restart |
|
|
| Spread Vault injectors | Two healthy replicas on Titan-08 and Titan-22; native anti-affinity and minAvailable=1 PDB | Revert 5f3f3184 |
|
|
| Protected application databases | Nineteen logical dumps and globals completed with checksums on Titan-0b; isolated Gitea restore passed in six seconds with 111 tables | Keep copies; stop the daily timer if needed |
|
|
| Application PostgreSQL resource/probe repair | Pinned current PostgreSQL 15.18 image, 1 GiB request / 2 GiB limit, native startup/readiness; 2/2 Ready on Titan-17 and healthy attached volume | Revert 103e2064 only after checking destination capacity; another database restart is required |
|
|
| Recovered Firefly and OpenSearch | Firefly Ready on Titan-08; OpenSearch Ready with its existing healthy volume on Titan-22. Pi capacity could not fit OpenSearch, so the documented CPU fallback was used; no GPU reassignment | Review capacity before placement rollback; preserve existing volumes |
|
|
| Protected local-only data policy | Hermes PVCs excluded from Soteria's cloud policy through Flux | Do not broaden this policy without reviewing data authorization |
|
|
| Stopped stale credential-event restart loop | Ananke c73b8fe installed on Titan-db after its complete quality gate; helper generation 796 remained unchanged after normal one-minute checks resumed. Peer 53ad03a no longer runs periodic cluster repairs | Native installer retains `/usr/local/lib/ananke/rollback/ananke.previous`; rolling back restores the defective behavior, so retain a bounded repair interval if needed |
|
|
| Corrected premature CI cancellation | Ariadne's configured elapsed-time cutoff is now 120 minutes, verified in the running Deployment | Revert 0180ba66; this returns the 45-minute cutoff that aborted an active Pi build |
|
|
| Scoped Soteria permissions | Atlas 46596536 applied; intended state updates and credential reads allowed, unrelated secret reads/listing and Hermes job creation denied; policy/usage fingerprints unchanged | Revert only after reviewing why broader access is required; retain state Secrets |
|
|
| Recovered Monerod mount | Snapshot df99cdac, clean detach b119076c, restart 0984164f; existing ext4 filesystem mounted on Titan-19, LMDB opened and RPC initialized | Preserve the snapshot and PVC. Returning to Titan-08 requires storage-path validation; do not replay the temporary scale-zero commit as a normal rollback |
|
|
| Bounded Longhorn background I/O | b0be99b7 applied; live settings are least-effort balancing and one concurrent rebuild per node | Revert b0be99b7 and reconcile Longhorn; the versioned settings Job restores best-effort and two rebuilds per node |
|
|
|
|
Host backup implementation is in commit 80aff498 and
|
|
[its operational notes](../infrastructure/host-backup/NOTES.md).
|
|
The restore check validates database/schema restoration without global-role/ACL
|
|
replay or starting K3s. A complete control-plane disaster drill remains outstanding.
|
|
|
|
## Work being completed
|
|
|
|
* Soteria: source fix 7018d4c in the Soteria repository allows live RWO Longhorn
|
|
snapshots while preserving the restic mount guard. Manual Longhorn requests
|
|
also enforce exclusions. All Go package tests pass; normal image publication
|
|
and a representative backup verification remain required.
|
|
* Native application backups now run daily from Titan-0b. The timer is enabled,
|
|
the verified initial bundle timestamp is scraped, and the 36-hour freshness
|
|
alert is provisioned. The first scheduled run completed at 08:37 UTC with
|
|
service result success and exit status zero.
|
|
* Soteria CI: the coverage report was generated after Sonar analysis, causing a
|
|
false zero-coverage gate despite 96.4% measured test coverage. Commit f167d7d
|
|
orders tests before analysis and awaits the Sonar gate; Jenkins declarative
|
|
validation passed. Release build 1292 failed during Git checkout with a 404;
|
|
current credential/repository checks and build 1293's source checkout succeed.
|
|
Builds 1293/1294 subsequently passed. Release retry 1295 failed the supply-chain
|
|
gate on broad ClusterRole secret access, not a dependency vulnerability. Source
|
|
3c0c868 removes those permissions and the expired exception; Atlas 46596536
|
|
deploys namespaced, resource-named state access. A verified Trivy 0.70.0 local
|
|
configuration scan reports zero HIGH/CRITICAL findings using its embedded
|
|
checks. Jenkins remains the publication gate. Source 596b2df also replaces a
|
|
Docker client with no daemon using the cluster's existing digest-pinned Kaniko
|
|
builder, without privileged access. Declarative pipeline validation passed.
|
|
Redundant timer build 1296 was stopped before any stage/agent began; release
|
|
1297 runs the complete gate with PUBLISH_IMAGES=true. Do not claim the runtime
|
|
backup defect is repaired until the new image and a real eligible backup pass.
|
|
|
|
At the October 4 04:38 UTC follow-up, release 1297 had failed after 4369 seconds.
|
|
Its enforced quality stages passed, but `go build` in the Kaniko image build
|
|
failed with `error obtaining VCS status: exit status 128`. This is now the
|
|
publication blocker. Build 1298 subsequently succeeded with PUBLISH_IMAGES=false,
|
|
so it did not deploy the runtime fix; the live image remains 0.1.0-120. Sonar's
|
|
analysis also emitted a missing Node executable error and a Go parser error;
|
|
passing the current gate is not proof that every analysis completed. Repair
|
|
those toolchain/evidence checks without weakening enforcement.
|
|
|
|
### October 3 storage recovery and capacity evidence
|
|
|
|
Titan-08's kernel recorded I/O failure, an aborted ext4 journal and a read-only
|
|
remount of Monerod's device at 11:31 UTC. Longhorn still reported two healthy
|
|
replicas on Titan-15/17. Replica health therefore did not establish filesystem
|
|
health. A retained local snapshot, `monerod-before-recovery-20261003`, became
|
|
ready at 20:52 UTC. It is a recovery point, not an independent backup.
|
|
|
|
Monerod was stopped through its Flux-tracked Deployment; CSI completed unstage
|
|
at 21:04 UTC and the volume detached normally. Restoring one replica with
|
|
Titan-08 excluded moved it to Titan-19. An initial attachment lookup failure
|
|
resolved through the normal controller retry. The existing ext4 filesystem
|
|
mounted read/write, LMDB opened, RPC initialized at 21:16:58 UTC and the pod
|
|
became 2/2 Ready without a restart. The first ten-second RPC check timed out;
|
|
application responsiveness and steady synchronization remain explicit checks.
|
|
No PVC deletion, forced detach, database replacement or manual filesystem repair
|
|
was performed. Retain the snapshot until recovery verification is complete.
|
|
Subsequent `get_version` calls returned status OK with the existing height
|
|
3775922, then 3775942 against target 3776212. Synchronization is making progress;
|
|
the broader `get_info` call still timed out at 45 seconds during recovery.
|
|
|
|
The 21:12 UTC capacity snapshot showed Titan-07/11 CPU requests at 3.53/3.56 of
|
|
3.60 allocatable cores; Titan-11 memory requests were 6662 of 6785 MiB. Titan-17/19
|
|
working memory was 5994/6102 of approximately 6656 MiB. Longhorn instance-manager
|
|
pods have CPU requests but no memory requests in this installed configuration.
|
|
Their observed 24-hour memory peaks were 3525 MiB on Titan-17, 3176 on Titan-19,
|
|
2914 on Titan-13 and 2594 on Titan-15. Those bytes cannot be treated as spare
|
|
application capacity merely because Kubernetes has not reserved them.
|
|
|
|
At 21:16 UTC Titan-19 reported high I/O and memory pressure while the blockchain
|
|
opened and CI ran. The native Longhorn settings were changed from best-effort
|
|
to least-effort replica balancing and from two to one concurrent rebuild per
|
|
node, verified live at 21:18 UTC. This reduces elective replica movement and
|
|
concurrent recovery I/O; it does not repair slow media, disable required replica
|
|
recovery, or change desired replica counts. Multiple rebuilds may take longer.
|
|
The existing 600-second replica replenishment delay is unchanged. Per-volume
|
|
overrides remain unchanged (28 least-effort, three disabled at inspection).
|
|
Setting semantics were checked against the installed
|
|
[Longhorn 1.8.2 source](https://github.com/longhorn/longhorn-manager/blob/v1.8.2/types/setting.go).
|
|
|
|
All 57 active Flux Kustomizations were Ready again at 21:18 UTC, including
|
|
Hermes after a transient Longhorn webhook timeout cleared. That controller
|
|
snapshot is not a declaration of sufficient failure capacity or completed
|
|
hardware repair.
|
|
|
|
### Cloud backup blocker established
|
|
|
|
At approximately 21:29 UTC the Longhorn catalog contained 980 backup objects:
|
|
952 Error, 14 Completed and 14 without a reported state. The completed entries
|
|
were from June/July; 899 retained errors explicitly reported `storage cap
|
|
exceeded`. Recent failed attempts on October 3 reported the same condition.
|
|
This is not evidence that rotating an expired key will repair uploads.
|
|
|
|
An authenticated Backblaze authorization check succeeded. The synchronized
|
|
Longhorn/Soteria key is non-expiring and has writeFiles/readFiles/listFiles
|
|
permissions. Its scope also includes account/bucket/key management; replace
|
|
that broad shared scope with dedicated bucket-scoped credentials in a separate
|
|
controlled rotation, without revoking other consumers blindly. No credentials
|
|
were printed or committed.
|
|
|
|
Soteria's cached metadata scan at 21:27:15 UTC reported 64,692 visible objects
|
|
and 963,207,320,600 bytes in `atlas-soteria`, no new objects in 24 hours, and a
|
|
last object modification of July 15. This is a visible-object inventory, not
|
|
verified billable storage including historical versions or other buckets.
|
|
Backblaze's exact account cap and approved monthly budget are not known from
|
|
the supported API checks. The user has been asked to review Caps & Alerts and
|
|
provide the budget. No spending cap was raised, old recovery copy removed, or
|
|
local-only material uploaded. A readable bucket and an Available backup target
|
|
do not establish successful writes.
|
|
|
|
Hermes tenant-2 remains unresolved: its volume is detached with a share manager
|
|
waiting on Titan-14, while an older engine reports running on Titan-23 with an
|
|
ERR replica reference. Three replica objects remain on Titan-13/15/17. These are
|
|
controller observations, not evidence of empty or lost data. Both consumers
|
|
report Running despite the unavailable share. No workspace content was read,
|
|
consumer stopped, volume forcibly detached or replica removed during inspection.
|
|
Preserving an approved local recovery copy and repairing that specific engine/
|
|
share ownership remains an explicit recovery task.
|
|
|
|
|
|
## Confirmed problems needing further work
|
|
|
|
| Problem | Evidence / implication | Next action |
|
|
| --- | --- | --- |
|
|
| titan-05 / titan-06 offline | Nineteen of twenty-one nodes Ready at the start of implementation | Physical power/network check; recover or replace; no remote software claim of repair |
|
|
| Pi power instability | Fresh undervoltage observed on titan-04 and titan-11 during audit; titan-04 remains cordoned | Check supplies, cables and USB power demand; keep quarantine until a stability check passes |
|
|
| titan-14 runtime I/O | USB flash runtime disk at sustained saturation during audit | Replace/migrate to suitable SSD after data preservation and a controlled drain |
|
|
| titan-18 storage stalls | I/O some pressure near 99%, load near 39; Longhorn engine image fails to become ready | Reduce CI load, inspect storage/kernel health, move runtime off inadequate media if confirmed |
|
|
| titan-08 stale iSCSI session | Pushgateway engine could not log out an obsolete target; replica data remained usable on another host | Repair during a controlled storage maintenance window; do not mass-restart instance managers holding healthy volumes |
|
|
| Hermes tenant-2 shared workspace | Existing faulted/detached volume, repeated recovery attempts | Preserve replica/recovery evidence and perform component-supported repair; no source content in routine logs |
|
|
| Backup coverage and restore proof | Old Soteria coverage is insufficient; repairing its scheduler does not instantly create all backups | Confirm eligible data, completed backup objects, achievable schedule and representative restores service by service |
|
|
| Database collation drift | Native dumps reported stored collation 2.36 versus runtime 2.41 for several databases | Plan index rebuilds with compatible locale settings before refreshing version metadata; do not merely suppress warnings |
|
|
| Supported software baseline | Ubuntu 24.10 on titan-db and older Kubernetes cohorts are out of support | Backed-up, staged host/K3s/Longhorn upgrades, one compatible cohort at a time |
|
|
| Failure capacity | Pi pool is heavily reserved; unused x86 capacity has explicit simulation/GPU roles | Recalculate compatible N+1 capacity after repairs; do not silently take reserved GPUs or simulation capacity |
|
|
|
|
## Completion standard
|
|
|
|
Do not call the entire cluster fixed based on one green snapshot. Open storage
|
|
faults, physical power/media problems, unsupported host software and per-service
|
|
restore coverage remain visible until addressed. Validate application behavior,
|
|
backup freshness, actual restore results, and 7-14 days without the recurring
|
|
failures. No automatic real-suite inference jobs or outage drills are part of
|
|
this maintenance work.
|
|
|
|
## Operational lessons from this maintenance
|
|
|
|
The application database move took approximately nine minutes because both
|
|
pinned images had to be pulled on the destination. Its volume attached correctly;
|
|
PostgreSQL and its exporter became Ready afterward. Preload required images on
|
|
an eligible destination before another planned database move. The larger database
|
|
dumps also completed much more efficiently through native PostgreSQL than through
|
|
`kubectl exec` output streaming. The native recovery script records this path.
|
|
|
|
At 06:07 UTC, both Titan-04 and Titan-11 emitted fresh undervoltage messages.
|
|
Their instantaneous Pi 5 input samples were 4.9245 V and 4.89904 V; these are not
|
|
measurements of the transient minimum. Both had `get_throttled=0x50000` between
|
|
events. The fault remains active even when that individual sample reports only
|
|
historical bits. The user is checking power/boot-media issues physically.
|
|
|
|
Firefly recovered on Titan-08 at approximately 06:52 UTC after its old pod
|
|
terminated and its volume moved normally. OpenSearch recovered on Titan-22 around 07:03 UTC after the corrected placement
|
|
was applied. Commit 2f043d10 temporarily allowed the old offline-node rollback
|
|
to finish. Normal rollback readiness was restored after the corrected release
|
|
became Ready.
|
|
Titan-18 remains cordoned; rescheduling does not repair its runtime media.
|
|
|
|
Ananke was confirmed restarting jenkins-vault-sync approximately every minute
|
|
from retained credential-warning events, including after the original failure
|
|
cleared. The Ananke fixes require
|
|
a current matching pod UID, active image-pull failure and warning age below ten
|
|
minutes, with a persisted 30-minute cooldown per helper. Only the coordinator
|
|
runs periodic cluster repair; the peer retains UPS monitoring and forwarding.
|
|
Coordinator revision c73b8fe passed the native host's complete quality gate,
|
|
including the per-file 95% coverage check, and installed successfully. Its
|
|
temporary one-hour interval was restored to the original 60 seconds at 07:41 UTC.
|
|
The peer runs 53ad03a, which includes the credential guard and coordinator-only
|
|
repair. UPS services remain active; the coordinator reports line power. The
|
|
normal checks no longer roll the healthy credential helper.
|
|
|
|
The updater previously installed revisions even after failing verification.
|
|
Its default now keeps the running binary on failure. The host had a known Go
|
|
auto-toolchain coverage failure; the quality script now pins its resolved Go
|
|
version for child commands. A cancelled datastore preflight also now returns
|
|
cancellation before probing a live database. All runtime changes remain gated.
|
|
Updater commit 53ad03a also preserves the installer's nonzero exit code when
|
|
fallback is disabled. Three isolated shell regressions passed for strict failure,
|
|
strict success and explicitly enabled legacy fallback. The coordinator's installed
|
|
updater script matches that commit; the peer's normal update passed the full
|
|
quality gate and completed successfully at approximately 07:55 UTC.
|
|
|
|
At 07:21 UTC, Ariadne aborted Soteria build 1290 solely for exceeding its default
|
|
45-minute elapsed-time cap; Go compilation was still active. Atlas commit
|
|
0180ba66 raises that deterministic cutoff to 120 minutes. This is a time-budget
|
|
backstop, not proof that a build is hung. The next Soteria release was queued
|
|
with PUBLISH_IMAGES=true; its runtime fix remains unverified until publication.
|
|
|
|
At approximately 07:19 UTC, five-minute I/O pressure waiting fractions were
|
|
0.78 on Titan-19, 0.77 on Titan-08, 0.50 on Titan-17 and 0.48 on Titan-13.
|
|
Titan-08's SSH query timed out and Outline's cold image pull was still pending.
|
|
Outline subsequently became Ready on Titan-11. These are additional signs that
|
|
runtime/storage capacity must be addressed; reducing restart churn alone does
|
|
not establish healthy media.
|
|
|
|
At 07:40 UTC, all 57 active Flux Kustomizations were Ready. The intentionally
|
|
suspended bstein-dev-home-migrations Kustomization retains its historical
|
|
ArtifactFailed condition. This snapshot does not close the hardware, storage,
|
|
backup coverage or supported-version work above.
|