docs: record recovery guards and current health
This commit is contained in:
parent
0180ba6698
commit
6a0d0699d9
@ -14,6 +14,7 @@ separates completed repairs from open problems. The
|
||||
| Application deployment | `services/<name>/`, or `infrastructure/<name>/` for foundations | Images, resources, probes, storage and networking |
|
||||
| Durable application disks | Longhorn; `infrastructure/longhorn/` | Replication and volume attachment; not an application-consistent database backup by itself |
|
||||
| Kubernetes datastore recovery | Native PostgreSQL tools and systemd on titan-db and titan-0b | Hourly protected LAN recovery bundles, independent of Kubernetes |
|
||||
| Application PostgreSQL recovery | `atlas-application-postgres-backup.timer` on titan-0b; `infrastructure/host-backup/` | Daily logical database dumps with checksums, kept on the approved LAN host |
|
||||
| Application backup orchestration | Soteria; `services/maintenance/apps/soteria-*` | Applies eligible-data policies through the chosen storage backend |
|
||||
| Power loss and exceptional node recovery | Ananke; host configuration and `scripts/ops/cluster_power_*` | Orders shutdown/startup and bounded node recovery; it does not replace storage or database recovery |
|
||||
| Node build/configuration | Metis and its sentinels; `services/maintenance/` | Approved node provisioning and host configuration |
|
||||
@ -23,6 +24,40 @@ separates completed repairs from open problems. The
|
||||
There is no new coordinating framework to learn. Prefer the native owner above.
|
||||
Use a recovery tool only when its documented operation matches the failure.
|
||||
|
||||
## Automatic recovery that can change services
|
||||
|
||||
Ananke's coordinator runs as `ananke.service` on Titan-db; its settings are in
|
||||
`/etc/ananke/ananke.yaml`, with installed source in `/opt/ananke`. Titan-24 is
|
||||
the UPS peer. Only the coordinator runs periodic cluster repair. UPS monitoring
|
||||
is a separate responsibility within the same daemon: do not stop that daemon
|
||||
casually to troubleshoot a Kubernetes helper.
|
||||
|
||||
The credential-helper repair checks every 60 seconds but acts only on a current
|
||||
pod with an active image-pull failure and a matching warning less than ten
|
||||
minutes old. It records an attempt before restarting a helper and allows at
|
||||
most one attempt per helper per 30 minutes, across daemon restarts. Old events
|
||||
and warnings for replaced pods do not justify another restart. Check helper
|
||||
generation and Ananke's journal when investigating unexpected rollouts:
|
||||
|
||||
```bash
|
||||
kubectl -n jenkins get deployment jenkins-vault-sync
|
||||
ssh titan-db sudo journalctl -u ananke --since '10 minutes ago' --no-pager
|
||||
ssh titan-db systemctl status ananke-update.timer
|
||||
```
|
||||
|
||||
The native updater now preserves the running binary when its quality gate
|
||||
fails. A binary-only maintenance install also preserves host/NUT configuration
|
||||
and retains the prior binary for rollback. Read `scripts/install.sh` in the
|
||||
Ananke repository before updating; `--binary-only --skip-deps` is the targeted
|
||||
path, while the normal updater also applies host configuration templates.
|
||||
|
||||
Ariadne's `ARIADNE_HERMES_HUNG_BUILD_MINUTES` setting is in
|
||||
`services/maintenance/apps/ariadne-deployment.yaml`. It is currently 120 minutes.
|
||||
Despite the name, this is an elapsed-time cancellation rule, not a detector of
|
||||
CPU or log progress. It can abort an active build. Review the build stage and
|
||||
node I/O before changing that cutoff; do not compensate for slow disks by
|
||||
starting more concurrent agents.
|
||||
|
||||
## Five-minute check
|
||||
|
||||
Run from the management host with the existing administrator configuration:
|
||||
@ -58,6 +93,17 @@ ssh titan-db sudo systemctl show atlas-k3s-backup.service -p Result -p ExecMainS
|
||||
ssh titan-0b sudo systemctl show atlas-k3s-replica.service -p Result -p ExecMainStatus
|
||||
```
|
||||
|
||||
Application database backups run separately each day at 08:15 UTC plus up to
|
||||
five minutes of jitter. Check their timer and completion timestamp too:
|
||||
|
||||
```bash
|
||||
ssh titan-0b sudo systemctl status atlas-application-postgres-backup.timer
|
||||
ssh titan-0b sudo cat /var/backups/atlas-postgres/latest/COMPLETE
|
||||
```
|
||||
|
||||
Their freshness alert fires after 36 hours. See the host backup notes before
|
||||
running the isolated restore verification or changing retention.
|
||||
|
||||
## Find the cause before choosing a repair
|
||||
|
||||
| Symptom | First check | Usual next action |
|
||||
|
||||
@ -23,6 +23,8 @@ Use [Cluster operations](CLUSTER_OPERATIONS.md) for the ordinary operator path.
|
||||
| Application PostgreSQL resource/probe repair | Pinned current PostgreSQL 15.18 image, 1 GiB request / 2 GiB limit, native startup/readiness; 2/2 Ready on Titan-17 and healthy attached volume | Revert 103e2064 only after checking destination capacity; another database restart is required |
|
||||
| Recovered Firefly and OpenSearch | Firefly Ready on Titan-08; OpenSearch Ready with its existing healthy volume on Titan-22. Pi capacity could not fit OpenSearch, so the documented CPU fallback was used; no GPU reassignment | Review capacity before placement rollback; preserve existing volumes |
|
||||
| Protected local-only data policy | Hermes PVCs excluded from Soteria's cloud policy through Flux | Do not broaden this policy without reviewing data authorization |
|
||||
| Stopped stale credential-event restart loop | Ananke c73b8fe installed on Titan-db after its complete quality gate; helper generation 796 remained unchanged after normal one-minute checks resumed. Peer ec7c2ee no longer runs periodic cluster repairs | Native installer retains `/usr/local/lib/ananke/rollback/ananke.previous`; rolling back restores the defective behavior, so retain a bounded repair interval if needed |
|
||||
| Corrected premature CI cancellation | Ariadne's configured elapsed-time cutoff is now 120 minutes, verified in the running Deployment | Revert 0180ba66; this returns the 45-minute cutoff that aborted an active Pi build |
|
||||
|
||||
Host backup implementation is in commit 80aff498 and
|
||||
[its operational notes](../infrastructure/host-backup/NOTES.md).
|
||||
@ -41,7 +43,9 @@ replay or starting K3s. A complete control-plane disaster drill remains outstand
|
||||
* Soteria CI: the coverage report was generated after Sonar analysis, causing a
|
||||
false zero-coverage gate despite 96.4% measured test coverage. Commit f167d7d
|
||||
orders tests before analysis and awaits the Sonar gate; Jenkins declarative
|
||||
validation passed. Release build remains pending; do not claim the runtime
|
||||
validation passed. Release build 1292 failed during Git checkout with a 404;
|
||||
current credential/repository checks and build 1293's source checkout succeed.
|
||||
The release retry remains queued; do not claim the runtime
|
||||
backup defect is repaired until the new image and a real eligible backup pass.
|
||||
|
||||
|
||||
@ -93,7 +97,41 @@ Titan-18 remains cordoned; rescheduling does not repair its runtime media.
|
||||
|
||||
Ananke was confirmed restarting jenkins-vault-sync approximately every minute
|
||||
from retained credential-warning events, including after the original failure
|
||||
cleared. A current-pod/UID/freshness check and persistent 30-minute repair
|
||||
cooldown are being validated in the Ananke repository. Its self-update fallback
|
||||
also installed revisions after quality-gate failures; the new default will
|
||||
leave the running binary in place when verification fails.
|
||||
cleared. The Ananke fixes require
|
||||
a current matching pod UID, active image-pull failure and warning age below ten
|
||||
minutes, with a persisted 30-minute cooldown per helper. Only the coordinator
|
||||
runs periodic cluster repair; the peer retains UPS monitoring and forwarding.
|
||||
Coordinator revision c73b8fe passed the native host's complete quality gate,
|
||||
including the per-file 95% coverage check, and installed successfully. Its
|
||||
temporary one-hour interval was restored to the original 60 seconds at 07:41 UTC.
|
||||
The peer runs ec7c2ee, which includes the credential guard and coordinator-only
|
||||
repair. UPS services remain active; the coordinator reports line power. The
|
||||
normal checks no longer roll the healthy credential helper.
|
||||
|
||||
The updater previously installed revisions even after failing verification.
|
||||
Its default now keeps the running binary on failure. The host had a known Go
|
||||
auto-toolchain coverage failure; the quality script now pins its resolved Go
|
||||
version for child commands. A cancelled datastore preflight also now returns
|
||||
cancellation before probing a live database. All runtime changes remain gated.
|
||||
Updater commit 53ad03a also preserves the installer's nonzero exit code when
|
||||
fallback is disabled. Three isolated shell regressions passed for strict failure,
|
||||
strict success and explicitly enabled legacy fallback. The coordinator's installed
|
||||
updater script matches that commit; the peer's normal update is in progress.
|
||||
|
||||
At 07:21 UTC, Ariadne aborted Soteria build 1290 solely for exceeding its default
|
||||
45-minute elapsed-time cap; Go compilation was still active. Atlas commit
|
||||
0180ba66 raises that deterministic cutoff to 120 minutes. This is a time-budget
|
||||
backstop, not proof that a build is hung. The next Soteria release was queued
|
||||
with PUBLISH_IMAGES=true; its runtime fix remains unverified until publication.
|
||||
|
||||
At approximately 07:19 UTC, five-minute I/O pressure waiting fractions were
|
||||
0.78 on Titan-19, 0.77 on Titan-08, 0.50 on Titan-17 and 0.48 on Titan-13.
|
||||
Titan-08's SSH query timed out and Outline's cold image pull was still pending.
|
||||
Outline subsequently became Ready on Titan-11. These are additional signs that
|
||||
runtime/storage capacity must be addressed; reducing restart churn alone does
|
||||
not establish healthy media.
|
||||
|
||||
At 07:40 UTC, all 57 active Flux Kustomizations were Ready. The intentionally
|
||||
suspended bstein-dev-home-migrations Kustomization retains its historical
|
||||
ArtifactFailed condition. This snapshot does not close the hardware, storage,
|
||||
backup coverage or supported-version work above.
|
||||
|
||||
Loading…
x
Reference in New Issue
Block a user