atlas-iac/docs/CLUSTER_STABILIZATION.md

39 KiB

Cluster stabilization - implementation record

This record begins on 2026-10-03 UTC, following authorization to implement the configuration-first plan. The October 2 audit remains a historical baseline, not a current health claim. Use Cluster operations for the ordinary operator path. The latest repair and settling cycles supersede the earlier point-in-time CI capacity and tenant-2 status below.

Six-priority progress tracker

Started against the user's approved list on 2026-10-03. Updated: 2026-10-04 13:50 UTC. "Verified" means the stated check passed; it does not imply a completed soak or failure drill. Physical repairs and disruptive recovery drills remain separate from routine software work.

The current first priority is a stable steady state. Stop recurring disruption, protect application and storage capacity, and verify recovery copies before expanding workload or scheduling platform upgrades. Keep necessary repairs small and inspectable; a new recovery controller is not a prerequisite.

Priority Status Completed evidence Next action / completion gate
1. Immediate software repairs Corrected Soteria release deployed; recovery checks remain Soteria 1305 succeeded and 0.1.0-128 is Ready with executable nonroot binary. Failed 127 rollout retained the serving old replica. Monerod previously recovered with existing data Validate an eligible backup and restore after resolving the cloud cap; continue Monerod latency/sync checks
2. Reliable nodes Software repairs applied; hardware and identity checks remain Full root administration verified on 21 reachable machines; Vault verification metadata saved. Repaired credentials on 12/13/19/20/21. Titan-04/08/11/12/14/18 excluded from new placement. Titan-19 log rotation repaired Verify changed SSH identities on 09/10; recover 05/06/16; address fresh power faults and runtime-media stalls. Metis 0.1.0-403 is live; complete physical replacement-image validation
3. Recovery coverage In progress; cloud uploads blocked External datastore restore and second LAN copy verified; 19 application DB dumps completed; scheduled application backup completed at 08:28 UTC; representative Gitea restore passed. Backblaze storage-cap failure established Resolve the account storage cap within an agreed budget, complete eligible PVC coverage and Soteria restore, verify application-level recovery after tenant-2 storage repair, and complete approved local-only recovery coverage
4. Capacity and placement Partial; build scratch moved off runtime media CI agent cap is one; all nine active project templates now exclude storage/quarantined workers. Disposable Longhorn scratch replaces inline emptyDir workspaces and build caches; Data Prepper image-layer storage corrected after a failed load check. Shared PostgreSQL/OpenSearch resource corrections deployed; Vault injectors spread. Longhorn uses least-effort balancing and one rebuild per node Soteria 1303/1305/1306, Ananke 384, Ariadne 563, Atlasbot 442 bstein-dev-home 581, Data Prepper 2826 and Metis 405 passed with scratch-backed agents; verify the remaining project builds; restore compatible healthy worker capacity. Account for the measured 2-3.5 GiB storage-process footprint and calculate spare capacity for one worker loss
5. Supported software Planned; backups established Installed K3s cohorts identified as 1.31.5 and 1.33.3; titan-db Ubuntu 24.10 is beyond support Select a compatible supported baseline; stage OS/K3s/storage upgrades with rollback. Plan collation/index repair
6. Maintainability and sustained health In progress Operator guide and rollback record published; Ananke stale-event loop fixed; updater failure reporting corrected; overlapping restart/eviction behavior removed Complete the service recovery checklist, validate remaining automation boundaries, arrange controlled recovery checks, then observe 7-14 days

Current work and acceptance checks

  • Correct Soteria release gate findings without disabling the gate; verify the local configuration scan.
  • Replace Soteria's daemonless Docker build with the existing pinned Kaniko builder; validate the pipeline.
  • Deploy and verify scoped Soteria state/credential permissions without altering stored policy or usage.
  • Publish the corrected Soteria image through the normal pipeline and Flux: build 1305, image 0.1.0-128, Ready at 11:49 UTC.
  • Verify an eligible backup and an isolated restore; retain Hermes exclusions.
  • Diagnose Monerod's host/attachment/filesystem error before selecting recovery.
  • Preserve a pre-recovery snapshot and remount the existing Monerod volume through Flux.
  • Recover Monerod with its existing data: get_version returned OK; height advanced from 3775922 to 3775942. Broader get_info requests remain slow.
  • Deploy and verify bounded Longhorn balancing/rebuild settings.
  • Verify the first scheduled native application PostgreSQL backup completed.
  • Recheck controller health: 57/57 active Flux Kustomizations and 19/21 nodes Ready at 2026-10-04 04:38 UTC; Titan-05/06 remain offline.
  • Complete application checks and the 7-14 day observation period.
  • Resolve Backblaze's storage cap before accepting cloud backup coverage; preserve existing copies.

Node access and software repair - October 4

The user prioritized node administration and software repairs before hardware replacement. No node was reflashed or rebooted during this work.

  • Verify full administrative execution on all 21 reachable machines, including the datastore and bastion; save the verification timestamp, admin username, access method and root-password state in Vault custom metadata.
  • Repair Titan-12/13/19: SSH keys worked, but atlas was locked and the root password differed from Vault. Both accounts now match their existing Vault passwords. Fresh independent password-backed sudo reached UID 0.
  • Repair the existing unlocked root passwords on Titan-20/21 to match Vault.
  • Remove the obsolete Metis passwordless command grants on 12/13/19 after verifying password-backed sudo; repeat the sudo verification afterward.
  • Fix the Vault bootstrap writer: preserve all fields, use compare-and-set, skip unchanged records, keep temporary files private, remove content-bearing errors, and retain the completed Job so Flux does not recreate it after TTL. Live Job -5 changed only Titan-23's incorrect .23 address to .24; 25 records were unchanged. Atlas commit 1bd8ae31.
  • Exclude Titan-11 from new placement through Flux (c2f41899); existing pods continue running. Its fresh voltage faults remain a hardware issue.
  • Repair Titan-19's logrotate post-hook failure: armbian-ramlog was inactive but ENABLED=true, copying between obsolete RAM-log locations although /var/log already points to external ext4 storage. The guarded versioned script disabled those hooks; the native logrotate service completed with result success/status 0.
  • Apply the same guarded repair on Titan-14 after confirming its RAM-log unit was already masked/inactive and logs use external ext4 storage. Logrotate then completed with result success/status 0. The later short I/O sample was quiet; this does not resolve its earlier intermittent storage stalls.
  • Migrate the active 50 MiB RAM-log filesystems on Titan-12/18 in a controlled node maintenance cycle. The obsolete-hook script intentionally refuses them. Their native stop action lazily unmounts /var/log while processes may still hold files open; do not treat changing ENABLED as a complete live migration. Preserve logs and arrange workload evacuation/restart before removing the mount.
  • Correct Ariadne's memory reservation/limit from 128/512 MiB to 512/1024 MiB after repeated cgroup OOM kills. The replacement is 2/2 Ready with zero restarts. This supplies headroom; it does not establish that its memory growth is bounded.
  • Publish Metis source fixes 1c110c9 and 6213159: native first-boot unit, per-injection pending marker, serialized setup, propagated password failures, removal of default passwordless grants, cleanup of applied bootstrap passwords, and correct file permissions/native systemd links in both image-writing paths.
  • Complete the normal Metis image publication and Flux rollout. Build 403 succeeded; controller and ARM/AMD sentinels are Ready on 0.1.0-403. Build 400's earlier cancellation was traced to Ariadne's elapsed-time rule and repaired. Queued/stalled older builds were stopped before the successful publication. A physical replacement-image boot still needs its separate acceptance check.
  • Enable Ananke's existing password-backed host-action path. Atlas 886dbeee adds 13 atlas-password mappings to the existing Vault CSI synchronizer; all matched Vault, with native two-minute rotation enabled. No root passwords are copied into that synchronizer. Source Ananke 41fd207 sets the namespace/name template and corrects Titan-24's administrator to tethys. Both live host configs were patched without changing unrelated settings and passed Ananke validation.
  • Verify read-only privileged actions using Ananke's own coordinator SSH identity on all 13 affected workers. Both Ananke services were restarted separately and returned active/running with zero automatic restarts. Jenkins credential-helper generation remained 796.
  • Restore the peer's missing SSH public key on 04/07/08/11/12/13/19/20/21. The key was derived on the authenticated Titan-24 host and matched the existing Vault recovery public key. New entries restrict their source to 192.168.22.26; existing administrator keys and SSH policies were preserved. The peer's own identity then passed password-backed sudo on all 13 workers, plus its existing privileged path on 0a/0b/0c/db/22. These are actual remote systemctl version checks, not just successful key installation.
  • Resolve the remaining network/physical findings below.

At the earlier 07:43 UTC check, 57/57 active Flux configurations and 19/21 Kubernetes nodes remained Ready. Ananke on both hosts had zero automatic restarts; Ariadne, Jenkins and the Grafana application container were Ready with zero restarts. The Grafana Vault sidecar had two earlier restarts. Storage acceptance remains open: Titan-15's engine-image installer was in CrashLoopBackOff with previous exit 137 (not established as an OOM), Titan-18's manager was unready, and the existing Titan-14 share manager was unready. Longhorn reported 74 attached healthy volumes and 61 detached volumes with unknown robustness; detached/unknown by itself is not a new volume failure.

Steady-state containment - October 4, 08:18 UTC

  • Quarantine/controller interaction: the Longhorn managers on 11/14/18 were repeatedly replaced at core Flux reconciliation times, followed by node-label restoration by the existing role reconciler. Atlas eb006075 declares their essential hardware/worker/storage labels and uses Flux SSA Merge on those three Node resources. Quarantine and prune protection remain. After the fix, all three manager pod UIDs were unchanged and Ready through two further explicit core reconciliations, with zero container restarts. This is a short stability check, not proof that the underlying power/media faults are repaired.
  • Ariadne's retained event for Metis 400 records hung_build and abort_requested=true. The deployed code canceled by elapsed age, without checking useful build progress. Atlas d47e7b8e sets ARIADNE_HERMES_HUNG_BUILD_MINUTES=0, the existing disable setting. Its new pod is Ready; a runtime check confirms long elapsed time alone no longer qualifies for cancellation. Other bounded diagnosis/repair remains enabled.
  • Atlas 8a6d3db1 sets Jenkins's cloud agent cap to one and makes storage/node exclusions mandatory in the shared template, IaC and Data Prepper pipelines. Metis 11ecd91 applies the same exclusions. These cover the inspected pipelines; unrelated projects with their own inline YAML still need a placement audit.
  • Jenkins was quieted for one controlled controller rollout. Metis 401 and Data Prepper 2814 were stopped through Jenkins to release Titan-15/19; their flows are complete and both agents are gone. Superseded queued Metis 402 and IaC 4230 were also stopped. Jenkins is Ready, quiet mode is off, and its runtime cloud cap is one. No application/storage pod was force-deleted for this change.
  • Metis 342a640 gives its lightweight Docker-client and publisher containers explicit 128 MiB reservations instead of namespace-default 256 MiB each. The resulting 1.25 GiB pod can fit the sampled Titan-08 reservation gap; actual build startup and peak consumption remain unverified. Limits and the build daemon's reservation were not reduced.
  • Capacity is still insufficient for some queued work: Data Prepper 2815 asks for 2 GiB and is Pending on the protected worker pool. A Pending agent also consumes the one-agent cap. Do not describe this as a fully functioning CI fleet, lower all requests to force placement, or return builds to overloaded storage nodes. Restore compatible worker capacity or measure and restructure the affected pipeline's reservations before increasing concurrency.

At the follow-up check, 57/57 active Flux configurations and 19/21 nodes were Ready; 05/06 remained unavailable. The intentionally suspended historical migration Kustomization is excluded from the active count. Titan-15's engine installer was Ready with its restart count unchanged at 109. Longhorn reported 73 attached healthy volumes and 61 detached/unknown volumes. The tenant-2 share manager remains unready; no workspace content was read. None of these snapshots establishes N+1 capacity or the required sustained-health observation period. Grafana's public HTTPS /api/health returned HTTP 200 with certificate checking enabled. This management workstation routes through 192.168.1.254; direct requests to the cluster LAN VIPs timed out, so this is not a successful laptop LAN-path test. Grafana's configured ingress VIP is 192.168.22.9; .50 is the separate LAN inference ingress.

Validation: Kustomize renders, client dry-runs and focused Flux diffs passed for core, maintenance, Jenkins and logging. The installed kubectl rejects combining --server-side with --dry-run=client; the valid client dry-run was used. Jenkins's native declarative validator accepted IaC, Data Prepper and both Metis pipeline revisions. Embedded JCasC parsed; live Flux and Jenkins checks above verify application of the configuration. No new real-data inference job ran.

Rollback: use focused Git reverts and Flux reconciliation. Restoring the older CI cap or placement permits renewed storage contention; restoring Ariadne's older threshold resumes cancellation of active builds by age. Preserve the explicit node role/storage labels when changing quarantine rather than restoring the minimal Node declarations that preceded eb006075. Jenkins JCasC changes trigger its tracked revision and controller rollout, so coordinate a quiet queue first. Metis resource requests can be reverted independently in its repo.

Current node distinctions

Nodes Established evidence Remaining action
12/13/19 atlas/root passwords restored; root SSH keys preserved; password-backed sudo works after legacy grant retirement Observe stability; use the corrected recovery image for future rebuilds
20/21 atlas sudo already worked; root passwords now match Vault Continue workload memory diagnosis; Jetson-21 recorded a Data Prepper cgroup OOM
04/11 Fresh undervoltage messages continued during this audit, despite sufficient disk space Onboard EXT5V samples were 4.82132 V (04) and 4.75968 V (11). Check delivered power/cables/peripherals; do not clear quarantine based on one sample
08 Image build drove runtime USB disk to 99.9% busy with 77% I/O pressure; pressure fell below 1% after the build stopped Quarantined in Git; repair or replace runtime media and verify loaded operation before return
12 SD-backed runtime correlated with probe stalls and an engine-image restart at 12:02; 2.7 GiB memory remained available New placement blocked without draining current services; inspect/migrate runtime media and perform loaded testing before return
14 Runtime USB flash; 39-45% iowait in the first sample, quiet later, no new kernel I/O error in the two-hour sample; stale RAM-log hooks corrected Preserve data; investigate intermittent storage stalls and validate sustained operation before return; hardware failure is not proven by I/O pressure alone
18 Quarantine has reduced current I/O pressure; its 50 MiB RAM-log filesystem is still active Keep quarantined until sustained storage checks pass; do not apply the inactive-RAM-log repair script to it
05 Answers LAN ping at .31; expected SSH, kubelet and common recovery ports refuse connections Console/user-space service check; this is not evidence that the machine is powered off
09/10 Answer LAN ping and SSH on port 22; recorded 2277 port is closed; host keys differ from saved keys Confirm reinstall/image identity before supplying passwords; no host-key checks bypassed
06/16 No LAN ping response; neighbor resolution incomplete; SSH unavailable Power, link and boot-media checks needed

The 12 unlocked root passwords match Vault. Nine machines retain locked direct root accounts (0a/0b/0c/04/07/08/11/db/jh); their administrator account plus sudo still provides full root control. Root-account lock state is recorded explicitly in Vault instead of implying that every root_password field is a usable login.

Validation: 19 Atlas regression tests passed (credential targeting, secret-safe output, legacy-grant preservation, Vault no-op/CAS behavior, preservation/idempotency of the Ananke config patch, and restricted peer-key enrollment); relevant Kustomize builds, client dry-runs and focused Flux diffs passed. Metis's complete Go tests and separate docs/LOC/vet/per-file coverage gate passed. Tests include a real throwaway ext4 image for systemd link injection and mocked password-setup failure, retry and stale image markers. No live node was burned to test recovery. All 57 active Flux Kustomizations remained Ready after these changes; 19/21 Kubernetes nodes are Ready.

Ananke still depends on an available Kubernetes API to retrieve its synchronized worker sudo credential. This is its existing implementation, not a new offline credential cache. Control-plane administration remains available through its existing path; use the dated manual SSH/Vault access record for intervention outside the automated recovery flow. Do not claim a full power-loss drill from read-only connection tests.

Rollback: Kubernetes changes use their focused Git commit and Flux. Titan-14/19's previous RAM-log config is root-only under /var/lib/atlas-maintenance/ramlog-before-20261004/; restoring it reintroduces the obsolete hooks. Retired sudo grants are root-only under /var/lib/atlas-maintenance/legacy-sudo-20261004/; do not restore passwordless access merely to avoid using the now-working password. Password repairs use the already stored Vault credentials; no secret values were changed or committed. Native Ananke config backups are root-only under /var/lib/atlas-maintenance/ananke-access-before-20261004/ on both hosts. Reverting its lookup settings would remove automated password-backed access; retain the working administrator passwords. Its source profiles and the versioned Atlas configuration helper document the change. Peer key enrollment uses scripts/install_ananke_peer_key.py. Rollback removes only the newly appended ananke-tethys-recovery line after checking another administrator session still works; do not replace the whole authorized_keys file.

Verified changes

Change Evidence Rollback
Native external PostgreSQL backup and second LAN copy A 54,083,817-byte custom dump restored into a disposable PostgreSQL 16 instance in 11 seconds with 22,350 kine rows; no TCP listener. Both timers enabled; checksum and matching-token checks pass Disable the two timers; retain all recovery bundles
Retired unconditional k3s-agent restart DaemonSet Flux preview removed only that DaemonSet; reconciliation completed and the DaemonSet is absent Revert f04c84ee, understanding that returning the DaemonSet restarts agents
Recovered metrics Pushgateway Existing volume attached healthy on titan-19; replacement pod Ready. No volume deletion or data replacement Revert the focused monitoring placement commits after the original host is repaired
Removed eviction for historical restart counts Descheduler manifest validates and Flux applied the change Revert fa9251ae
Restored GitOps UI Deployment Helm drift correction recreated weave-gitops; Deployment 1/1 Ready Revert 6c9398ea to disable ongoing drift correction; this does not remove recovered resources
Aligned Flux definition ownership Removed creation-only policy and adopted already-active service state; no service paths or source refs changed Revert the focused Flux commits; review suspension fields before doing so
Backup freshness alerts Both datastore-copy timestamps are scraped; native Grafana provisioning reload returned HTTP 200 Revert the focused monitoring change; backups continue independently
Quarantined stalled runtime media Titan-14 and Titan-18 are SchedulingDisabled via infrastructure/core/node-maintenance.yaml; no forced storage detach After repair, change unschedulable to false in Git and verify before removing the prune-disabled Node declaration
Bounded CI concurrency Jenkins controller Ready with runtime cap one (8a6d3db1); inspected pipeline placements exclude storage workers Revert the focused change and reconcile Jenkins; allow a controller restart and account for renewed storage load
Spread Vault injectors Two healthy replicas on Titan-08 and Titan-22; native anti-affinity and minAvailable=1 PDB Revert 5f3f3184
Protected application databases Nineteen logical dumps and globals completed with checksums on Titan-0b; isolated Gitea restore passed in six seconds with 111 tables Keep copies; stop the daily timer if needed
Application PostgreSQL resource/probe repair Pinned current PostgreSQL 15.18 image, 1 GiB request / 2 GiB limit, native startup/readiness; 2/2 Ready on Titan-17 and healthy attached volume Revert 103e2064 only after checking destination capacity; another database restart is required
Recovered Firefly and OpenSearch Firefly Ready on Titan-08; OpenSearch Ready with its existing healthy volume on Titan-22. Pi capacity could not fit OpenSearch, so the documented CPU fallback was used; no GPU reassignment Review capacity before placement rollback; preserve existing volumes
Protected local-only data policy Hermes PVCs excluded from Soteria's cloud policy through Flux Do not broaden this policy without reviewing data authorization
Stopped stale credential-event restart loop Ananke c73b8fe installed on Titan-db after its complete quality gate; helper generation 796 remained unchanged after normal one-minute checks resumed. Peer 53ad03a no longer runs periodic cluster repairs Native installer retains /usr/local/lib/ananke/rollback/ananke.previous; rolling back restores the defective behavior, so retain a bounded repair interval if needed
Disabled premature CI cancellation Ariadne's elapsed-only cutoff is disabled with value zero (d47e7b8e), verified in the running application Reverting resumes age-only cancellations; require progress-aware detection before re-enabling
Scoped Soteria permissions Atlas 46596536 applied; intended state updates and credential reads allowed, unrelated secret reads/listing and Hermes job creation denied; policy/usage fingerprints unchanged Revert only after reviewing why broader access is required; retain state Secrets
Recovered Monerod mount Snapshot df99cdac, clean detach b119076c, restart 0984164f; existing ext4 filesystem mounted on Titan-19, LMDB opened and RPC initialized Preserve the snapshot and PVC. Returning to Titan-08 requires storage-path validation; do not replay the temporary scale-zero commit as a normal rollback
Bounded Longhorn background I/O b0be99b7 applied; live settings are least-effort balancing and one concurrent rebuild per node Revert b0be99b7 and reconcile Longhorn; the versioned settings Job restores best-effort and two rebuilds per node

Host backup implementation is in commit 80aff498 and its operational notes. The restore check validates database/schema restoration without global-role/ACL replay or starting K3s. A complete control-plane disaster drill remains outstanding.

Work being completed

  • Soteria: source fix 7018d4c in the Soteria repository allows live RWO Longhorn snapshots while preserving the restic mount guard. Manual Longhorn requests also enforce exclusions. All Go package tests pass; normal image publication and a representative backup verification remain required.
  • Native application backups now run daily from Titan-0b. The timer is enabled, the verified initial bundle timestamp is scraped, and the 36-hour freshness alert is provisioned. The first scheduled run completed at 08:37 UTC with service result success and exit status zero.
  • Soteria CI: the coverage report was generated after Sonar analysis, causing a false zero-coverage gate despite 96.4% measured test coverage. Commit f167d7d orders tests before analysis and awaits the Sonar gate; Jenkins declarative validation passed. Release build 1292 failed during Git checkout with a 404; current credential/repository checks and build 1293's source checkout succeed. Builds 1293/1294 subsequently passed. Release retry 1295 failed the supply-chain gate on broad ClusterRole secret access, not a dependency vulnerability. Source 3c0c868 removes those permissions and the expired exception; Atlas 46596536 deploys namespaced, resource-named state access. A verified Trivy 0.70.0 local configuration scan reports zero HIGH/CRITICAL findings using its embedded checks. Jenkins remains the publication gate. Source 596b2df also replaces a Docker client with no daemon using the cluster's existing digest-pinned Kaniko builder, without privileged access. Declarative pipeline validation passed. Redundant timer build 1296 was stopped before any stage/agent began; release 1297 runs the complete gate with PUBLISH_IMAGES=true. Do not claim the runtime backup defect is repaired until the new image and a real eligible backup pass.

At the October 4 04:38 UTC follow-up, release 1297 had failed after 4369 seconds. Its enforced quality stages passed, but go build in the Kaniko image build failed with error obtaining VCS status: exit status 128. This is now the publication blocker. Build 1298 subsequently succeeded with PUBLISH_IMAGES=false, so it did not deploy the runtime fix; the live image remains 0.1.0-120. Sonar's analysis also emitted a missing Node executable error and a Go parser error; passing the current gate is not proof that every analysis completed. Repair those toolchain/evidence checks without weakening enforcement.

October 3 storage recovery and capacity evidence

Titan-08's kernel recorded I/O failure, an aborted ext4 journal and a read-only remount of Monerod's device at 11:31 UTC. Longhorn still reported two healthy replicas on Titan-15/17. Replica health therefore did not establish filesystem health. A retained local snapshot, monerod-before-recovery-20261003, became ready at 20:52 UTC. It is a recovery point, not an independent backup.

Monerod was stopped through its Flux-tracked Deployment; CSI completed unstage at 21:04 UTC and the volume detached normally. Restoring one replica with Titan-08 excluded moved it to Titan-19. An initial attachment lookup failure resolved through the normal controller retry. The existing ext4 filesystem mounted read/write, LMDB opened, RPC initialized at 21:16:58 UTC and the pod became 2/2 Ready without a restart. The first ten-second RPC check timed out; application responsiveness and steady synchronization remain explicit checks. No PVC deletion, forced detach, database replacement or manual filesystem repair was performed. Retain the snapshot until recovery verification is complete. Subsequent get_version calls returned status OK with the existing height 3775922, then 3775942 against target 3776212. Synchronization is making progress; the broader get_info call still timed out at 45 seconds during recovery.

The 21:12 UTC capacity snapshot showed Titan-07/11 CPU requests at 3.53/3.56 of 3.60 allocatable cores; Titan-11 memory requests were 6662 of 6785 MiB. Titan-17/19 working memory was 5994/6102 of approximately 6656 MiB. Longhorn instance-manager pods have CPU requests but no memory requests in this installed configuration. Their observed 24-hour memory peaks were 3525 MiB on Titan-17, 3176 on Titan-19, 2914 on Titan-13 and 2594 on Titan-15. Those bytes cannot be treated as spare application capacity merely because Kubernetes has not reserved them.

At 21:16 UTC Titan-19 reported high I/O and memory pressure while the blockchain opened and CI ran. The native Longhorn settings were changed from best-effort to least-effort replica balancing and from two to one concurrent rebuild per node, verified live at 21:18 UTC. This reduces elective replica movement and concurrent recovery I/O; it does not repair slow media, disable required replica recovery, or change desired replica counts. Multiple rebuilds may take longer. The existing 600-second replica replenishment delay is unchanged. Per-volume overrides remain unchanged (28 least-effort, three disabled at inspection). Setting semantics were checked against the installed Longhorn 1.8.2 source.

All 57 active Flux Kustomizations were Ready again at 21:18 UTC, including Hermes after a transient Longhorn webhook timeout cleared. That controller snapshot is not a declaration of sufficient failure capacity or completed hardware repair.

Cloud backup blocker established

At approximately 21:29 UTC the Longhorn catalog contained 980 backup objects: 952 Error, 14 Completed and 14 without a reported state. The completed entries were from June/July; 899 retained errors explicitly reported storage cap exceeded. Recent failed attempts on October 3 reported the same condition. This is not evidence that rotating an expired key will repair uploads.

An authenticated Backblaze authorization check succeeded. The synchronized Longhorn/Soteria key is non-expiring and has writeFiles/readFiles/listFiles permissions. Its scope also includes account/bucket/key management; replace that broad shared scope with dedicated bucket-scoped credentials in a separate controlled rotation, without revoking other consumers blindly. No credentials were printed or committed.

Soteria's cached metadata scan at 21:27:15 UTC reported 64,692 visible objects and 963,207,320,600 bytes in atlas-soteria, no new objects in 24 hours, and a last object modification of July 15. This is a visible-object inventory, not verified billable storage including historical versions or other buckets. Backblaze's exact account cap and approved monthly budget are not known from the supported API checks. The user has been asked to review Caps & Alerts and provide the budget. No spending cap was raised, old recovery copy removed, or local-only material uploaded. A readable bucket and an Available backup target do not establish successful writes.

Hermes tenant-2 remains unresolved: its volume is detached with a share manager waiting on Titan-14, while an older engine reports running on Titan-23 with an ERR replica reference. Three replica objects remain on Titan-13/15/17. These are controller observations, not evidence of empty or lost data. Both consumers report Running despite the unavailable share. No workspace content was read, consumer stopped, volume forcibly detached or replica removed during inspection. Preserving an approved local recovery copy and repairing that specific engine/ share ownership remains an explicit recovery task.

Confirmed problems needing further work

Problem Evidence / implication Next action
titan-05 / titan-06 offline Nineteen of twenty-one nodes Ready at the start of implementation Physical power/network check; recover or replace; no remote software claim of repair
Pi power instability Fresh undervoltage observed on titan-04 and titan-11 during audit; titan-04 remains cordoned Check supplies, cables and USB power demand; keep quarantine until a stability check passes
titan-14 runtime I/O USB flash runtime disk at sustained saturation during audit Replace/migrate to suitable SSD after data preservation and a controlled drain
titan-18 storage stalls I/O some pressure near 99%, load near 39; Longhorn engine image fails to become ready Reduce CI load, inspect storage/kernel health, move runtime off inadequate media if confirmed
titan-08 stale iSCSI session Pushgateway engine could not log out an obsolete target; replica data remained usable on another host Repair during a controlled storage maintenance window; do not mass-restart instance managers holding healthy volumes
Hermes tenant-2 shared workspace Existing faulted/detached volume, repeated recovery attempts Preserve replica/recovery evidence and perform component-supported repair; no source content in routine logs
Backup coverage and restore proof Old Soteria coverage is insufficient; repairing its scheduler does not instantly create all backups Confirm eligible data, completed backup objects, achievable schedule and representative restores service by service
Database collation drift Native dumps reported stored collation 2.36 versus runtime 2.41 for several databases Plan index rebuilds with compatible locale settings before refreshing version metadata; do not merely suppress warnings
Supported software baseline Ubuntu 24.10 on titan-db and older Kubernetes cohorts are out of support Backed-up, staged host/K3s/Longhorn upgrades, one compatible cohort at a time
Failure capacity Pi pool is heavily reserved; unused x86 capacity has explicit simulation/GPU roles Recalculate compatible N+1 capacity after repairs; do not silently take reserved GPUs or simulation capacity

Completion standard

Do not call the entire cluster fixed based on one green snapshot. Open storage faults, physical power/media problems, unsupported host software and per-service restore coverage remain visible until addressed. Validate application behavior, backup freshness, actual restore results, and 7-14 days without the recurring failures. No automatic real-suite inference jobs or outage drills are part of this maintenance work.

Operational lessons from this maintenance

The application database move took approximately nine minutes because both pinned images had to be pulled on the destination. Its volume attached correctly; PostgreSQL and its exporter became Ready afterward. Preload required images on an eligible destination before another planned database move. The larger database dumps also completed much more efficiently through native PostgreSQL than through kubectl exec output streaming. The native recovery script records this path.

At 06:07 UTC, both Titan-04 and Titan-11 emitted fresh undervoltage messages. Their instantaneous Pi 5 input samples were 4.9245 V and 4.89904 V; these are not measurements of the transient minimum. Both had get_throttled=0x50000 between events. The fault remains active even when that individual sample reports only historical bits. The user is checking power/boot-media issues physically.

Firefly recovered on Titan-08 at approximately 06:52 UTC after its old pod terminated and its volume moved normally. OpenSearch recovered on Titan-22 around 07:03 UTC after the corrected placement was applied. Commit 2f043d10 temporarily allowed the old offline-node rollback to finish. Normal rollback readiness was restored after the corrected release became Ready. Titan-18 remains cordoned; rescheduling does not repair its runtime media.

Ananke was confirmed restarting jenkins-vault-sync approximately every minute from retained credential-warning events, including after the original failure cleared. The Ananke fixes require a current matching pod UID, active image-pull failure and warning age below ten minutes, with a persisted 30-minute cooldown per helper. Only the coordinator runs periodic cluster repair; the peer retains UPS monitoring and forwarding. Coordinator revision c73b8fe passed the native host's complete quality gate, including the per-file 95% coverage check, and installed successfully. Its temporary one-hour interval was restored to the original 60 seconds at 07:41 UTC. The peer runs 53ad03a, which includes the credential guard and coordinator-only repair. UPS services remain active; the coordinator reports line power. The normal checks no longer roll the healthy credential helper.

The updater previously installed revisions even after failing verification. Its default now keeps the running binary on failure. The host had a known Go auto-toolchain coverage failure; the quality script now pins its resolved Go version for child commands. A cancelled datastore preflight also now returns cancellation before probing a live database. All runtime changes remain gated. Updater commit 53ad03a also preserves the installer's nonzero exit code when fallback is disabled. Three isolated shell regressions passed for strict failure, strict success and explicitly enabled legacy fallback. The coordinator's installed updater script matches that commit; the peer's normal update passed the full quality gate and completed successfully at approximately 07:55 UTC.

At 07:21 UTC, Ariadne aborted Soteria build 1290 solely for exceeding its default 45-minute elapsed-time cap; Go compilation was still active. Atlas commit 0180ba66 raises that deterministic cutoff to 120 minutes. This is a time-budget backstop, not proof that a build is hung. The next Soteria release was queued with PUBLISH_IMAGES=true; its runtime fix remains unverified until publication.

At approximately 07:19 UTC, five-minute I/O pressure waiting fractions were 0.78 on Titan-19, 0.77 on Titan-08, 0.50 on Titan-17 and 0.48 on Titan-13. Titan-08's SSH query timed out and Outline's cold image pull was still pending. Outline subsequently became Ready on Titan-11. These are additional signs that runtime/storage capacity must be addressed; reducing restart churn alone does not establish healthy media.

At 07:40 UTC, all 57 active Flux Kustomizations were Ready. The intentionally suspended bstein-dev-home-migrations Kustomization retains its historical ArtifactFailed condition. This snapshot does not close the hardware, storage, backup coverage or supported-version work above.