diff --git a/docs/CLUSTER_OPERATIONS.md b/docs/CLUSTER_OPERATIONS.md index 9ad6fc7c..4d4d30f9 100644 --- a/docs/CLUSTER_OPERATIONS.md +++ b/docs/CLUSTER_OPERATIONS.md @@ -14,6 +14,7 @@ separates completed repairs from open problems. The | Application deployment | `services//`, or `infrastructure//` for foundations | Images, resources, probes, storage and networking | | Durable application disks | Longhorn; `infrastructure/longhorn/` | Replication and volume attachment; not an application-consistent database backup by itself | | Kubernetes datastore recovery | Native PostgreSQL tools and systemd on titan-db and titan-0b | Hourly protected LAN recovery bundles, independent of Kubernetes | +| Application PostgreSQL recovery | `atlas-application-postgres-backup.timer` on titan-0b; `infrastructure/host-backup/` | Daily logical database dumps with checksums, kept on the approved LAN host | | Application backup orchestration | Soteria; `services/maintenance/apps/soteria-*` | Applies eligible-data policies through the chosen storage backend | | Power loss and exceptional node recovery | Ananke; host configuration and `scripts/ops/cluster_power_*` | Orders shutdown/startup and bounded node recovery; it does not replace storage or database recovery | | Node build/configuration | Metis and its sentinels; `services/maintenance/` | Approved node provisioning and host configuration | @@ -23,6 +24,40 @@ separates completed repairs from open problems. The There is no new coordinating framework to learn. Prefer the native owner above. Use a recovery tool only when its documented operation matches the failure. +## Automatic recovery that can change services + +Ananke's coordinator runs as `ananke.service` on Titan-db; its settings are in +`/etc/ananke/ananke.yaml`, with installed source in `/opt/ananke`. Titan-24 is +the UPS peer. Only the coordinator runs periodic cluster repair. UPS monitoring +is a separate responsibility within the same daemon: do not stop that daemon +casually to troubleshoot a Kubernetes helper. + +The credential-helper repair checks every 60 seconds but acts only on a current +pod with an active image-pull failure and a matching warning less than ten +minutes old. It records an attempt before restarting a helper and allows at +most one attempt per helper per 30 minutes, across daemon restarts. Old events +and warnings for replaced pods do not justify another restart. Check helper +generation and Ananke's journal when investigating unexpected rollouts: + +```bash +kubectl -n jenkins get deployment jenkins-vault-sync +ssh titan-db sudo journalctl -u ananke --since '10 minutes ago' --no-pager +ssh titan-db systemctl status ananke-update.timer +``` + +The native updater now preserves the running binary when its quality gate +fails. A binary-only maintenance install also preserves host/NUT configuration +and retains the prior binary for rollback. Read `scripts/install.sh` in the +Ananke repository before updating; `--binary-only --skip-deps` is the targeted +path, while the normal updater also applies host configuration templates. + +Ariadne's `ARIADNE_HERMES_HUNG_BUILD_MINUTES` setting is in +`services/maintenance/apps/ariadne-deployment.yaml`. It is currently 120 minutes. +Despite the name, this is an elapsed-time cancellation rule, not a detector of +CPU or log progress. It can abort an active build. Review the build stage and +node I/O before changing that cutoff; do not compensate for slow disks by +starting more concurrent agents. + ## Five-minute check Run from the management host with the existing administrator configuration: @@ -58,6 +93,17 @@ ssh titan-db sudo systemctl show atlas-k3s-backup.service -p Result -p ExecMainS ssh titan-0b sudo systemctl show atlas-k3s-replica.service -p Result -p ExecMainStatus ``` +Application database backups run separately each day at 08:15 UTC plus up to +five minutes of jitter. Check their timer and completion timestamp too: + +```bash +ssh titan-0b sudo systemctl status atlas-application-postgres-backup.timer +ssh titan-0b sudo cat /var/backups/atlas-postgres/latest/COMPLETE +``` + +Their freshness alert fires after 36 hours. See the host backup notes before +running the isolated restore verification or changing retention. + ## Find the cause before choosing a repair | Symptom | First check | Usual next action | diff --git a/docs/CLUSTER_STABILIZATION.md b/docs/CLUSTER_STABILIZATION.md index 39fee5fa..3cbe483d 100644 --- a/docs/CLUSTER_STABILIZATION.md +++ b/docs/CLUSTER_STABILIZATION.md @@ -23,6 +23,8 @@ Use [Cluster operations](CLUSTER_OPERATIONS.md) for the ordinary operator path. | Application PostgreSQL resource/probe repair | Pinned current PostgreSQL 15.18 image, 1 GiB request / 2 GiB limit, native startup/readiness; 2/2 Ready on Titan-17 and healthy attached volume | Revert 103e2064 only after checking destination capacity; another database restart is required | | Recovered Firefly and OpenSearch | Firefly Ready on Titan-08; OpenSearch Ready with its existing healthy volume on Titan-22. Pi capacity could not fit OpenSearch, so the documented CPU fallback was used; no GPU reassignment | Review capacity before placement rollback; preserve existing volumes | | Protected local-only data policy | Hermes PVCs excluded from Soteria's cloud policy through Flux | Do not broaden this policy without reviewing data authorization | +| Stopped stale credential-event restart loop | Ananke c73b8fe installed on Titan-db after its complete quality gate; helper generation 796 remained unchanged after normal one-minute checks resumed. Peer ec7c2ee no longer runs periodic cluster repairs | Native installer retains `/usr/local/lib/ananke/rollback/ananke.previous`; rolling back restores the defective behavior, so retain a bounded repair interval if needed | +| Corrected premature CI cancellation | Ariadne's configured elapsed-time cutoff is now 120 minutes, verified in the running Deployment | Revert 0180ba66; this returns the 45-minute cutoff that aborted an active Pi build | Host backup implementation is in commit 80aff498 and [its operational notes](../infrastructure/host-backup/NOTES.md). @@ -41,7 +43,9 @@ replay or starting K3s. A complete control-plane disaster drill remains outstand * Soteria CI: the coverage report was generated after Sonar analysis, causing a false zero-coverage gate despite 96.4% measured test coverage. Commit f167d7d orders tests before analysis and awaits the Sonar gate; Jenkins declarative - validation passed. Release build remains pending; do not claim the runtime + validation passed. Release build 1292 failed during Git checkout with a 404; + current credential/repository checks and build 1293's source checkout succeed. + The release retry remains queued; do not claim the runtime backup defect is repaired until the new image and a real eligible backup pass. @@ -93,7 +97,41 @@ Titan-18 remains cordoned; rescheduling does not repair its runtime media. Ananke was confirmed restarting jenkins-vault-sync approximately every minute from retained credential-warning events, including after the original failure -cleared. A current-pod/UID/freshness check and persistent 30-minute repair -cooldown are being validated in the Ananke repository. Its self-update fallback -also installed revisions after quality-gate failures; the new default will -leave the running binary in place when verification fails. +cleared. The Ananke fixes require +a current matching pod UID, active image-pull failure and warning age below ten +minutes, with a persisted 30-minute cooldown per helper. Only the coordinator +runs periodic cluster repair; the peer retains UPS monitoring and forwarding. +Coordinator revision c73b8fe passed the native host's complete quality gate, +including the per-file 95% coverage check, and installed successfully. Its +temporary one-hour interval was restored to the original 60 seconds at 07:41 UTC. +The peer runs ec7c2ee, which includes the credential guard and coordinator-only +repair. UPS services remain active; the coordinator reports line power. The +normal checks no longer roll the healthy credential helper. + +The updater previously installed revisions even after failing verification. +Its default now keeps the running binary on failure. The host had a known Go +auto-toolchain coverage failure; the quality script now pins its resolved Go +version for child commands. A cancelled datastore preflight also now returns +cancellation before probing a live database. All runtime changes remain gated. +Updater commit 53ad03a also preserves the installer's nonzero exit code when +fallback is disabled. Three isolated shell regressions passed for strict failure, +strict success and explicitly enabled legacy fallback. The coordinator's installed +updater script matches that commit; the peer's normal update is in progress. + +At 07:21 UTC, Ariadne aborted Soteria build 1290 solely for exceeding its default +45-minute elapsed-time cap; Go compilation was still active. Atlas commit +0180ba66 raises that deterministic cutoff to 120 minutes. This is a time-budget +backstop, not proof that a build is hung. The next Soteria release was queued +with PUBLISH_IMAGES=true; its runtime fix remains unverified until publication. + +At approximately 07:19 UTC, five-minute I/O pressure waiting fractions were +0.78 on Titan-19, 0.77 on Titan-08, 0.50 on Titan-17 and 0.48 on Titan-13. +Titan-08's SSH query timed out and Outline's cold image pull was still pending. +Outline subsequently became Ready on Titan-11. These are additional signs that +runtime/storage capacity must be addressed; reducing restart churn alone does +not establish healthy media. + +At 07:40 UTC, all 57 active Flux Kustomizations were Ready. The intentionally +suspended bstein-dev-home-migrations Kustomization retains its historical +ArtifactFailed condition. This snapshot does not close the hardware, storage, +backup coverage or supported-version work above.