Start here for ordinary operation. The repository describes desired state;
Kubernetes and Flux apply it. The [implementation record](CLUSTER_STABILIZATION.md)
separates completed repairs from open problems. The
[service inventory](cluster-audit-20261002/SERVICE_PLAN.md) covers the whole cluster.
## The small map
| Concern | Owner and configuration | What it does |
| --- | --- | --- |
| Desired Kubernetes configuration | Flux; `clusters/atlas/flux-system/` | Tracks `main`, applies service and infrastructure folders |
| Normal pod replacement and scheduling | Kubernetes; each workload manifest | Restarts failed containers and schedules replacements within placement and resource constraints |
| Application deployment | `services/<name>/`, or `infrastructure/<name>/` for foundations | Images, resources, probes, storage and networking |
| Durable application disks | Longhorn; `infrastructure/longhorn/` | Replication and volume attachment; not an application-consistent database backup by itself |
| Kubernetes datastore recovery | Native PostgreSQL tools and systemd on titan-db and titan-0b | Hourly protected LAN recovery bundles, independent of Kubernetes |
| Application backup orchestration | Soteria; `services/maintenance/apps/soteria-*` | Applies eligible-data policies through the chosen storage backend |
| Power loss and exceptional node recovery | Ananke; host configuration and `scripts/ops/cluster_power_*` | Orders shutdown/startup and bounded node recovery; it does not replace storage or database recovery |
| Node build/configuration | Metis and its sentinels; `services/maintenance/` | Approved node provisioning and host configuration |
| CI fault handling | Ariadne; `services/maintenance/apps/ariadne-*` | Bounded CI diagnosis/recovery; core cluster health must not depend on model calls |
| Health evidence | Grafana, VictoriaMetrics, Alertmanager; `services/monitoring/` | Shows application, node and storage evidence; a green controller alone is insufficient |
There is no new coordinating framework to learn. Prefer the native owner above.
Use a recovery tool only when its documented operation matches the failure.
## Five-minute check
Run from the management host with the existing administrator configuration:
```bash
kubectl get nodes -o wide
flux get kustomizations -A
flux get helmreleases -A
kubectl get deployments,statefulsets -A
kubectl get pods -A --field-selector=status.phase=Pending
kubectl -n longhorn-system get volumes.longhorn.io
```
Then check the affected service's health endpoint or UI. `Running` does not mean
ready; `Ready` does not prove useful application behavior. A detached Longhorn
volume may be intentional if its workload is parked. Distinguish that from an
attached volume that is faulted or an application waiting for its disk.
Check the control-plane backups separately:
```bash
ssh titan-db sudo systemctl status atlas-k3s-backup.timer
ssh titan-0b sudo systemctl status atlas-k3s-replica.timer
`COMPLETE` contains timestamps, not database contents. The backup and replica
should normally be less than two hours old. Check the last service result too:
```bash
ssh titan-db sudo systemctl show atlas-k3s-backup.service -p Result -p ExecMainStatus
ssh titan-0b sudo systemctl show atlas-k3s-replica.service -p Result -p ExecMainStatus
```
## Find the cause before choosing a repair
| Symptom | First check | Usual next action |
| --- | --- | --- |
| Node NotReady | Power/network, kubelet status, disk space, pressure | Repair/quarantine the node; do not repeatedly restart every application |
| Pod Pending / FailedScheduling | `kubectl describe pod` scheduling events | Correct resource requests or eligible capacity; free cluster-wide RAM does not guarantee a compatible destination |
| ContainerCreating / FailedMount | PVC, Longhorn volume, attachment and engine events | Preserve data; repair the attachment or use a verified healthy node. Do not delete the PVC |
| CrashLoopBackOff / OOMKilled | Previous container logs and last termination reason | Fix the application or its resource envelope; lifetime restart counts are not a reason for repeated eviction |
| Ingress 502/503 | Service endpoints and backend readiness | Restore the backend or its dependency before changing DNS/TLS |
| Flux Ready but missing Helm workload | Helm manifest and drift correction | Check that drift detection is enabled and suitable for that release |
| All cluster API operations fail | titan-db PostgreSQL and control-plane nodes | The datastore is external PostgreSQL; an etcd snapshot procedure will not restore it |
| CI waiting while applications work | Jenkins queue, agent capacity and node I/O | Let work queue; adding agents to saturated Pi disks makes both CI and services worse |
Useful scoped commands:
```bash
kubectl -n NAMESPACE describe pod POD
kubectl -n NAMESPACE get events --sort-by=.lastTimestamp
kubectl -n NAMESPACE get endpointslices
kubectl -n NAMESPACE logs POD -c CONTAINER --previous --tail=80
```
Application logs can contain private data. Inspect them locally and redact before
sharing. Never paste Secrets, tokens, source records or inference bodies into a
ticket or assistant conversation.
## Make a normal change
1. Edit the service's tracked manifest. Placement, requests, limits and probes
belong together in that workload's configuration.
2. Render and validate the affected folder; inspect the Flux preview.
3. Commit and push a small change to the tracked branch.
4. Reconcile that service and verify its application behavior and storage.