atlas-iac/docs/CLUSTER_OPERATIONS.md
jenkins b41dab9886
Some checks failed
Tests / Declarative: Post Actions failed: 45, skipped: 89, passed: 4270
docs: record verified backup storage cap blocker
2026-10-03 16:33:02 -05:00

232 lines
13 KiB
Markdown

# Running Atlas without an AI assistant
Start here for ordinary operation. The repository describes desired state;
Kubernetes and Flux apply it. The [implementation record](CLUSTER_STABILIZATION.md)
separates completed repairs from open problems. The
[service inventory](cluster-audit-20261002/SERVICE_PLAN.md) covers the whole cluster.
## The small map
| Concern | Owner and configuration | What it does |
| --- | --- | --- |
| Desired Kubernetes configuration | Flux; `clusters/atlas/flux-system/` | Tracks `main`, applies service and infrastructure folders |
| Normal pod replacement and scheduling | Kubernetes; each workload manifest | Restarts failed containers and schedules replacements within placement and resource constraints |
| Application deployment | `services/<name>/`, or `infrastructure/<name>/` for foundations | Images, resources, probes, storage and networking |
| Durable application disks | Longhorn; `infrastructure/longhorn/` | Replication and volume attachment; not an application-consistent database backup by itself |
| Kubernetes datastore recovery | Native PostgreSQL tools and systemd on titan-db and titan-0b | Hourly protected LAN recovery bundles, independent of Kubernetes |
| Application PostgreSQL recovery | `atlas-application-postgres-backup.timer` on titan-0b; `infrastructure/host-backup/` | Daily logical database dumps with checksums, kept on the approved LAN host |
| Application backup orchestration | Soteria; `services/maintenance/apps/soteria-*` | Applies eligible-data policies through the chosen storage backend |
| Power loss and exceptional node recovery | Ananke; host configuration and `scripts/ops/cluster_power_*` | Orders shutdown/startup and bounded node recovery; it does not replace storage or database recovery |
| Node build/configuration | Metis and its sentinels; `services/maintenance/` | Approved node provisioning and host configuration |
| CI fault handling | Ariadne; `services/maintenance/apps/ariadne-*` | Bounded CI diagnosis/recovery; core cluster health must not depend on model calls |
| Health evidence | Grafana, VictoriaMetrics, Alertmanager; `services/monitoring/` | Shows application, node and storage evidence; a green controller alone is insufficient |
There is no new coordinating framework to learn. Prefer the native owner above.
Use a recovery tool only when its documented operation matches the failure.
## Automatic recovery that can change services
Ananke's coordinator runs as `ananke.service` on Titan-db; its settings are in
`/etc/ananke/ananke.yaml`, with installed source in `/opt/ananke`. Titan-24 is
the UPS peer. Only the coordinator runs periodic cluster repair. UPS monitoring
is a separate responsibility within the same daemon: do not stop that daemon
casually to troubleshoot a Kubernetes helper.
The credential-helper repair checks every 60 seconds but acts only on a current
pod with an active image-pull failure and a matching warning less than ten
minutes old. It records an attempt before restarting a helper and allows at
most one attempt per helper per 30 minutes, across daemon restarts. Old events
and warnings for replaced pods do not justify another restart. Check helper
generation and Ananke's journal when investigating unexpected rollouts:
```bash
kubectl -n jenkins get deployment jenkins-vault-sync
ssh titan-db sudo journalctl -u ananke --since '10 minutes ago' --no-pager
ssh titan-db systemctl status ananke-update.timer
```
The native updater now preserves the running binary when its quality gate
fails. A binary-only maintenance install also preserves host/NUT configuration
and retains the prior binary for rollback. Read `scripts/install.sh` in the
Ananke repository before updating; `--binary-only --skip-deps` is the targeted
path, while the normal updater also applies host configuration templates.
Ariadne's `ARIADNE_HERMES_HUNG_BUILD_MINUTES` setting is in
`services/maintenance/apps/ariadne-deployment.yaml`. It is currently 120 minutes.
Despite the name, this is an elapsed-time cancellation rule, not a detector of
CPU or log progress. It can abort an active build. Review the build stage and
node I/O before changing that cutoff; do not compensate for slow disks by
starting more concurrent agents.
## Five-minute check
Run from the management host with the existing administrator configuration:
```bash
kubectl get nodes -o wide
flux get kustomizations -A
flux get helmreleases -A
kubectl get deployments,statefulsets -A
kubectl get pods -A --field-selector=status.phase=Pending
kubectl -n longhorn-system get volumes.longhorn.io
```
Then check the affected service's health endpoint or UI. `Running` does not mean
ready; `Ready` does not prove useful application behavior. A detached Longhorn
volume may be intentional if its workload is parked. Distinguish that from an
attached volume that is faulted or an application waiting for its disk.
Check the control-plane backups separately:
```bash
ssh titan-db sudo systemctl status atlas-k3s-backup.timer
ssh titan-0b sudo systemctl status atlas-k3s-replica.timer
ssh titan-db sudo cat /var/backups/atlas-k3s/latest/COMPLETE
ssh titan-0b sudo cat /var/backups/atlas-k3s-replica/latest/COMPLETE
```
`COMPLETE` contains timestamps, not database contents. The backup and replica
should normally be less than two hours old. Check the last service result too:
```bash
ssh titan-db sudo systemctl show atlas-k3s-backup.service -p Result -p ExecMainStatus
ssh titan-0b sudo systemctl show atlas-k3s-replica.service -p Result -p ExecMainStatus
```
Application database backups run separately each day at 08:15 UTC plus up to
five minutes of jitter. Check their timer and completion timestamp too:
```bash
ssh titan-0b sudo systemctl status atlas-application-postgres-backup.timer
ssh titan-0b sudo cat /var/backups/atlas-postgres/latest/COMPLETE
```
Their freshness alert fires after 36 hours. See the host backup notes before
running the isolated restore verification or changing retention.
## Find the cause before choosing a repair
| Symptom | First check | Usual next action |
| --- | --- | --- |
| Node NotReady | Power/network, kubelet status, disk space, pressure | Repair/quarantine the node; do not repeatedly restart every application |
| Pod Pending / FailedScheduling | `kubectl describe pod` scheduling events | Correct resource requests or eligible capacity; free cluster-wide RAM does not guarantee a compatible destination |
| ContainerCreating / FailedMount | PVC, Longhorn volume, attachment and engine events | Preserve data; repair the attachment or use a verified healthy node. Do not delete the PVC |
| CrashLoopBackOff / OOMKilled | Previous container logs and last termination reason | Fix the application or its resource envelope; lifetime restart counts are not a reason for repeated eviction |
| Ingress 502/503 | Service endpoints and backend readiness | Restore the backend or its dependency before changing DNS/TLS |
| Flux Ready but missing Helm workload | Helm manifest and drift correction | Check that drift detection is enabled and suitable for that release |
| All cluster API operations fail | titan-db PostgreSQL and control-plane nodes | The datastore is external PostgreSQL; an etcd snapshot procedure will not restore it |
| CI waiting while applications work | Jenkins queue, agent capacity and node I/O | Let work queue; adding agents to saturated Pi disks makes both CI and services worse |
Useful scoped commands:
```bash
kubectl -n NAMESPACE describe pod POD
kubectl -n NAMESPACE get events --sort-by=.lastTimestamp
kubectl -n NAMESPACE get endpointslices
kubectl -n NAMESPACE logs POD -c CONTAINER --previous --tail=80
```
Application logs can contain private data. Inspect them locally and redact before
sharing. Never paste Secrets, tokens, source records or inference bodies into a
ticket or assistant conversation.
## Make a normal change
1. Edit the service's tracked manifest. Placement, requests, limits and probes
belong together in that workload's configuration.
2. Render and validate the affected folder; inspect the Flux preview.
3. Commit and push a small change to the tracked branch.
4. Reconcile that service and verify its application behavior and storage.
For example:
```bash
kubectl kustomize services/monitoring > /tmp/monitoring-render.yaml
kubectl apply --dry-run=client -k services/monitoring
flux diff kustomization monitoring --path services/monitoring
git diff --check
git diff
```
After committing and pushing the reviewed files:
```bash
flux reconcile kustomization monitoring -n flux-system --with-source
kubectl -n monitoring get deployments,statefulsets,pods
```
A Flux diff exits nonzero when it finds changes; read the output to distinguish
that from a validation failure. Kustomize renders a HelmRelease, not the chart's
workloads. For chart values, also inspect the chart's rendered Deployment or
StatefulSet. OpenSearch, for example, uses `values.nodeAffinity`.
Rollback a configuration change by reverting its focused commit and reconciling
the same service. Do not blindly revert storage migrations or database changes;
those require their recovery procedure. Do not delete data to clear a red status.
## Backups and recovery
The external datastore's [host backup notes](../infrastructure/host-backup/NOTES.md)
give the actual paths, schedule, credential scope, isolated restore check and
rollback. Backups contain sensitive data and remain root-only on approved LAN
hosts. They are not encrypted at rest or off-site disaster recovery copies.
Soteria is responsible for eligible application backups, not for making the
control plane boot. Live Longhorn snapshots and database-native dumps offer
different consistency guarantees. A recent backup indicator is not a restore
test. Keep database restore checks and service recovery drills explicit.
For cloud backups, inspect completed backup timestamps, not just the backup
target's Available status or a successful bucket listing. The October 3 audit
found a readable Backblaze bucket rejecting uploads with `storage cap exceeded`.
Check Backblaze Caps & Alerts and agree the storage budget before changing it.
Do not delete old recovery points merely to make uploads work. Distinguish
visible-object bytes from billed storage, including retained versions. The
current diagnosis and counts are in the implementation record.
The Hermes namespace is excluded from cloud backup policy. Its workspaces can
contain material restricted to local infrastructure. Do not remove that exclusion
to make a coverage dashboard greener.
## A one-day familiarization exercise
* Morning: follow one application from its Flux definition to its workload,
storage, service and ingress. Compare those manifests with live objects.
* Before lunch: inspect one scheduling incident and one storage incident using
the table above. Identify the responsible component without changing anything.
* Afternoon: run the isolated datastore restore verification, inspect backup
timestamps, then make and revert a harmless configuration change through Flux.
* Finish by finding each remaining issue in the implementation record and the
corresponding owner/runbook. Confirm access to Git, the management host, the
hosts' SSH path and protected recovery material without relying on SSO alone.
This provides a practical route to ownership; it cannot promise that every
hardware, storage or database disaster will be solvable in one day.
## Physical checks and return to service
For an offline Pi, record the board identity, supply model and its 5 V rating,
power/activity LEDs, HDMI boot result, Ethernet link and current router address.
Preserve the original SD/USB media. Test with separate known-good spare boot media
before concluding that a board has failed; do not clone a live node identity by
moving another cluster node's boot card into it.
Pi 5 power must be evaluated at its 5 V output mode, not a charger's headline
wattage. A controlled supply-and-cable substitution isolates the power path.
Record any inline adapters and USB loads. Do not unplug runtime USB storage on
a running node. Titan-11 carries important services; arrange workload movement
before disconnecting it. See the official Raspberry Pi power documentation for
5 V / 5 A capability and cable-loss requirements.
A repaired node returns only after stable power, clean storage/kernel logs,
reliable networking and representative-load testing. Observe it before restoring
normal placement. Titan-14/18 quarantine is explicitly tracked in
`infrastructure/core/node-maintenance.yaml`; `prune: disabled` protects the Node
objects from deletion when that maintenance declaration is eventually removed.
Clear `spec.unschedulable` through Git after validation, then reconcile core.
Container runtime media and Longhorn data disks are separate. A Pi with terabytes
of healthy application storage can still stall because containerd, image unpacking
and logs live on a small USB flash device. Include the runtime disk in hardware
repair and capacity checks.