docs: record recovery guards and current health

This commit is contained in:
jenkins 2026-10-03 02:50:38 -05:00
parent 0180ba6698
commit 6a0d0699d9
2 changed files with 89 additions and 5 deletions

View File

@ -14,6 +14,7 @@ separates completed repairs from open problems. The
| Application deployment | `services/<name>/`, or `infrastructure/<name>/` for foundations | Images, resources, probes, storage and networking |
| Durable application disks | Longhorn; `infrastructure/longhorn/` | Replication and volume attachment; not an application-consistent database backup by itself |
| Kubernetes datastore recovery | Native PostgreSQL tools and systemd on titan-db and titan-0b | Hourly protected LAN recovery bundles, independent of Kubernetes |
| Application PostgreSQL recovery | `atlas-application-postgres-backup.timer` on titan-0b; `infrastructure/host-backup/` | Daily logical database dumps with checksums, kept on the approved LAN host |
| Application backup orchestration | Soteria; `services/maintenance/apps/soteria-*` | Applies eligible-data policies through the chosen storage backend |
| Power loss and exceptional node recovery | Ananke; host configuration and `scripts/ops/cluster_power_*` | Orders shutdown/startup and bounded node recovery; it does not replace storage or database recovery |
| Node build/configuration | Metis and its sentinels; `services/maintenance/` | Approved node provisioning and host configuration |
@ -23,6 +24,40 @@ separates completed repairs from open problems. The
There is no new coordinating framework to learn. Prefer the native owner above.
Use a recovery tool only when its documented operation matches the failure.
## Automatic recovery that can change services
Ananke's coordinator runs as `ananke.service` on Titan-db; its settings are in
`/etc/ananke/ananke.yaml`, with installed source in `/opt/ananke`. Titan-24 is
the UPS peer. Only the coordinator runs periodic cluster repair. UPS monitoring
is a separate responsibility within the same daemon: do not stop that daemon
casually to troubleshoot a Kubernetes helper.
The credential-helper repair checks every 60 seconds but acts only on a current
pod with an active image-pull failure and a matching warning less than ten
minutes old. It records an attempt before restarting a helper and allows at
most one attempt per helper per 30 minutes, across daemon restarts. Old events
and warnings for replaced pods do not justify another restart. Check helper
generation and Ananke's journal when investigating unexpected rollouts:
```bash
kubectl -n jenkins get deployment jenkins-vault-sync
ssh titan-db sudo journalctl -u ananke --since '10 minutes ago' --no-pager
ssh titan-db systemctl status ananke-update.timer
```
The native updater now preserves the running binary when its quality gate
fails. A binary-only maintenance install also preserves host/NUT configuration
and retains the prior binary for rollback. Read `scripts/install.sh` in the
Ananke repository before updating; `--binary-only --skip-deps` is the targeted
path, while the normal updater also applies host configuration templates.
Ariadne's `ARIADNE_HERMES_HUNG_BUILD_MINUTES` setting is in
`services/maintenance/apps/ariadne-deployment.yaml`. It is currently 120 minutes.
Despite the name, this is an elapsed-time cancellation rule, not a detector of
CPU or log progress. It can abort an active build. Review the build stage and
node I/O before changing that cutoff; do not compensate for slow disks by
starting more concurrent agents.
## Five-minute check
Run from the management host with the existing administrator configuration:
@ -58,6 +93,17 @@ ssh titan-db sudo systemctl show atlas-k3s-backup.service -p Result -p ExecMainS
ssh titan-0b sudo systemctl show atlas-k3s-replica.service -p Result -p ExecMainStatus
```
Application database backups run separately each day at 08:15 UTC plus up to
five minutes of jitter. Check their timer and completion timestamp too:
```bash
ssh titan-0b sudo systemctl status atlas-application-postgres-backup.timer
ssh titan-0b sudo cat /var/backups/atlas-postgres/latest/COMPLETE
```
Their freshness alert fires after 36 hours. See the host backup notes before
running the isolated restore verification or changing retention.
## Find the cause before choosing a repair
| Symptom | First check | Usual next action |

View File

@ -23,6 +23,8 @@ Use [Cluster operations](CLUSTER_OPERATIONS.md) for the ordinary operator path.
| Application PostgreSQL resource/probe repair | Pinned current PostgreSQL 15.18 image, 1 GiB request / 2 GiB limit, native startup/readiness; 2/2 Ready on Titan-17 and healthy attached volume | Revert 103e2064 only after checking destination capacity; another database restart is required |
| Recovered Firefly and OpenSearch | Firefly Ready on Titan-08; OpenSearch Ready with its existing healthy volume on Titan-22. Pi capacity could not fit OpenSearch, so the documented CPU fallback was used; no GPU reassignment | Review capacity before placement rollback; preserve existing volumes |
| Protected local-only data policy | Hermes PVCs excluded from Soteria's cloud policy through Flux | Do not broaden this policy without reviewing data authorization |
| Stopped stale credential-event restart loop | Ananke c73b8fe installed on Titan-db after its complete quality gate; helper generation 796 remained unchanged after normal one-minute checks resumed. Peer ec7c2ee no longer runs periodic cluster repairs | Native installer retains `/usr/local/lib/ananke/rollback/ananke.previous`; rolling back restores the defective behavior, so retain a bounded repair interval if needed |
| Corrected premature CI cancellation | Ariadne's configured elapsed-time cutoff is now 120 minutes, verified in the running Deployment | Revert 0180ba66; this returns the 45-minute cutoff that aborted an active Pi build |
Host backup implementation is in commit 80aff498 and
[its operational notes](../infrastructure/host-backup/NOTES.md).
@ -41,7 +43,9 @@ replay or starting K3s. A complete control-plane disaster drill remains outstand
* Soteria CI: the coverage report was generated after Sonar analysis, causing a
false zero-coverage gate despite 96.4% measured test coverage. Commit f167d7d
orders tests before analysis and awaits the Sonar gate; Jenkins declarative
validation passed. Release build remains pending; do not claim the runtime
validation passed. Release build 1292 failed during Git checkout with a 404;
current credential/repository checks and build 1293's source checkout succeed.
The release retry remains queued; do not claim the runtime
backup defect is repaired until the new image and a real eligible backup pass.
@ -93,7 +97,41 @@ Titan-18 remains cordoned; rescheduling does not repair its runtime media.
Ananke was confirmed restarting jenkins-vault-sync approximately every minute
from retained credential-warning events, including after the original failure
cleared. A current-pod/UID/freshness check and persistent 30-minute repair
cooldown are being validated in the Ananke repository. Its self-update fallback
also installed revisions after quality-gate failures; the new default will
leave the running binary in place when verification fails.
cleared. The Ananke fixes require
a current matching pod UID, active image-pull failure and warning age below ten
minutes, with a persisted 30-minute cooldown per helper. Only the coordinator
runs periodic cluster repair; the peer retains UPS monitoring and forwarding.
Coordinator revision c73b8fe passed the native host's complete quality gate,
including the per-file 95% coverage check, and installed successfully. Its
temporary one-hour interval was restored to the original 60 seconds at 07:41 UTC.
The peer runs ec7c2ee, which includes the credential guard and coordinator-only
repair. UPS services remain active; the coordinator reports line power. The
normal checks no longer roll the healthy credential helper.
The updater previously installed revisions even after failing verification.
Its default now keeps the running binary on failure. The host had a known Go
auto-toolchain coverage failure; the quality script now pins its resolved Go
version for child commands. A cancelled datastore preflight also now returns
cancellation before probing a live database. All runtime changes remain gated.
Updater commit 53ad03a also preserves the installer's nonzero exit code when
fallback is disabled. Three isolated shell regressions passed for strict failure,
strict success and explicitly enabled legacy fallback. The coordinator's installed
updater script matches that commit; the peer's normal update is in progress.
At 07:21 UTC, Ariadne aborted Soteria build 1290 solely for exceeding its default
45-minute elapsed-time cap; Go compilation was still active. Atlas commit
0180ba66 raises that deterministic cutoff to 120 minutes. This is a time-budget
backstop, not proof that a build is hung. The next Soteria release was queued
with PUBLISH_IMAGES=true; its runtime fix remains unverified until publication.
At approximately 07:19 UTC, five-minute I/O pressure waiting fractions were
0.78 on Titan-19, 0.77 on Titan-08, 0.50 on Titan-17 and 0.48 on Titan-13.
Titan-08's SSH query timed out and Outline's cold image pull was still pending.
Outline subsequently became Ready on Titan-11. These are additional signs that
runtime/storage capacity must be addressed; reducing restart churn alone does
not establish healthy media.
At 07:40 UTC, all 57 active Flux Kustomizations were Ready. The intentionally
suspended bstein-dev-home-migrations Kustomization retains its historical
ArtifactFailed condition. This snapshot does not close the hardware, storage,
backup coverage or supported-version work above.