151 lines
8.2 KiB
Markdown
151 lines
8.2 KiB
Markdown
# Running Atlas without an AI assistant
|
|
|
|
Start here for ordinary operation. The repository describes desired state;
|
|
Kubernetes and Flux apply it. The [implementation record](CLUSTER_STABILIZATION.md)
|
|
separates completed repairs from open problems. The
|
|
[service inventory](cluster-audit-20261002/SERVICE_PLAN.md) covers the whole cluster.
|
|
|
|
## The small map
|
|
|
|
| Concern | Owner and configuration | What it does |
|
|
| --- | --- | --- |
|
|
| Desired Kubernetes configuration | Flux; `clusters/atlas/flux-system/` | Tracks `main`, applies service and infrastructure folders |
|
|
| Normal pod replacement and scheduling | Kubernetes; each workload manifest | Restarts failed containers and schedules replacements within placement and resource constraints |
|
|
| Application deployment | `services/<name>/`, or `infrastructure/<name>/` for foundations | Images, resources, probes, storage and networking |
|
|
| Durable application disks | Longhorn; `infrastructure/longhorn/` | Replication and volume attachment; not an application-consistent database backup by itself |
|
|
| Kubernetes datastore recovery | Native PostgreSQL tools and systemd on titan-db and titan-0b | Hourly protected LAN recovery bundles, independent of Kubernetes |
|
|
| Application backup orchestration | Soteria; `services/maintenance/apps/soteria-*` | Applies eligible-data policies through the chosen storage backend |
|
|
| Power loss and exceptional node recovery | Ananke; host configuration and `scripts/ops/cluster_power_*` | Orders shutdown/startup and bounded node recovery; it does not replace storage or database recovery |
|
|
| Node build/configuration | Metis and its sentinels; `services/maintenance/` | Approved node provisioning and host configuration |
|
|
| CI fault handling | Ariadne; `services/maintenance/apps/ariadne-*` | Bounded CI diagnosis/recovery; core cluster health must not depend on model calls |
|
|
| Health evidence | Grafana, VictoriaMetrics, Alertmanager; `services/monitoring/` | Shows application, node and storage evidence; a green controller alone is insufficient |
|
|
|
|
There is no new coordinating framework to learn. Prefer the native owner above.
|
|
Use a recovery tool only when its documented operation matches the failure.
|
|
|
|
## Five-minute check
|
|
|
|
Run from the management host with the existing administrator configuration:
|
|
|
|
```bash
|
|
kubectl get nodes -o wide
|
|
flux get kustomizations -A
|
|
flux get helmreleases -A
|
|
kubectl get deployments,statefulsets -A
|
|
kubectl get pods -A --field-selector=status.phase=Pending
|
|
kubectl -n longhorn-system get volumes.longhorn.io
|
|
```
|
|
|
|
Then check the affected service's health endpoint or UI. `Running` does not mean
|
|
ready; `Ready` does not prove useful application behavior. A detached Longhorn
|
|
volume may be intentional if its workload is parked. Distinguish that from an
|
|
attached volume that is faulted or an application waiting for its disk.
|
|
|
|
Check the control-plane backups separately:
|
|
|
|
```bash
|
|
ssh titan-db sudo systemctl status atlas-k3s-backup.timer
|
|
ssh titan-0b sudo systemctl status atlas-k3s-replica.timer
|
|
ssh titan-db sudo cat /var/backups/atlas-k3s/latest/COMPLETE
|
|
ssh titan-0b sudo cat /var/backups/atlas-k3s-replica/latest/COMPLETE
|
|
```
|
|
|
|
`COMPLETE` contains timestamps, not database contents. The backup and replica
|
|
should normally be less than two hours old. Check the last service result too:
|
|
|
|
```bash
|
|
ssh titan-db sudo systemctl show atlas-k3s-backup.service -p Result -p ExecMainStatus
|
|
ssh titan-0b sudo systemctl show atlas-k3s-replica.service -p Result -p ExecMainStatus
|
|
```
|
|
|
|
## Find the cause before choosing a repair
|
|
|
|
| Symptom | First check | Usual next action |
|
|
| --- | --- | --- |
|
|
| Node NotReady | Power/network, kubelet status, disk space, pressure | Repair/quarantine the node; do not repeatedly restart every application |
|
|
| Pod Pending / FailedScheduling | `kubectl describe pod` scheduling events | Correct resource requests or eligible capacity; free cluster-wide RAM does not guarantee a compatible destination |
|
|
| ContainerCreating / FailedMount | PVC, Longhorn volume, attachment and engine events | Preserve data; repair the attachment or use a verified healthy node. Do not delete the PVC |
|
|
| CrashLoopBackOff / OOMKilled | Previous container logs and last termination reason | Fix the application or its resource envelope; lifetime restart counts are not a reason for repeated eviction |
|
|
| Ingress 502/503 | Service endpoints and backend readiness | Restore the backend or its dependency before changing DNS/TLS |
|
|
| Flux Ready but missing Helm workload | Helm manifest and drift correction | Check that drift detection is enabled and suitable for that release |
|
|
| All cluster API operations fail | titan-db PostgreSQL and control-plane nodes | The datastore is external PostgreSQL; an etcd snapshot procedure will not restore it |
|
|
| CI waiting while applications work | Jenkins queue, agent capacity and node I/O | Let work queue; adding agents to saturated Pi disks makes both CI and services worse |
|
|
|
|
Useful scoped commands:
|
|
|
|
```bash
|
|
kubectl -n NAMESPACE describe pod POD
|
|
kubectl -n NAMESPACE get events --sort-by=.lastTimestamp
|
|
kubectl -n NAMESPACE get endpointslices
|
|
kubectl -n NAMESPACE logs POD -c CONTAINER --previous --tail=80
|
|
```
|
|
|
|
Application logs can contain private data. Inspect them locally and redact before
|
|
sharing. Never paste Secrets, tokens, source records or inference bodies into a
|
|
ticket or assistant conversation.
|
|
|
|
## Make a normal change
|
|
|
|
1. Edit the service's tracked manifest. Placement, requests, limits and probes
|
|
belong together in that workload's configuration.
|
|
2. Render and validate the affected folder; inspect the Flux preview.
|
|
3. Commit and push a small change to the tracked branch.
|
|
4. Reconcile that service and verify its application behavior and storage.
|
|
|
|
For example:
|
|
|
|
```bash
|
|
kubectl kustomize services/monitoring > /tmp/monitoring-render.yaml
|
|
kubectl apply --dry-run=client -k services/monitoring
|
|
flux diff kustomization monitoring --path services/monitoring
|
|
git diff --check
|
|
git diff
|
|
```
|
|
|
|
After committing and pushing the reviewed files:
|
|
|
|
```bash
|
|
flux reconcile kustomization monitoring -n flux-system --with-source
|
|
kubectl -n monitoring get deployments,statefulsets,pods
|
|
```
|
|
|
|
A Flux diff exits nonzero when it finds changes; read the output to distinguish
|
|
that from a validation failure. Kustomize renders a HelmRelease, not the chart's
|
|
workloads. For chart values, also inspect the chart's rendered Deployment or
|
|
StatefulSet. OpenSearch, for example, uses `values.nodeAffinity`.
|
|
|
|
Rollback a configuration change by reverting its focused commit and reconciling
|
|
the same service. Do not blindly revert storage migrations or database changes;
|
|
those require their recovery procedure. Do not delete data to clear a red status.
|
|
|
|
## Backups and recovery
|
|
|
|
The external datastore's [host backup notes](../infrastructure/host-backup/NOTES.md)
|
|
give the actual paths, schedule, credential scope, isolated restore check and
|
|
rollback. Backups contain sensitive data and remain root-only on approved LAN
|
|
hosts. They are not encrypted at rest or off-site disaster recovery copies.
|
|
|
|
Soteria is responsible for eligible application backups, not for making the
|
|
control plane boot. Live Longhorn snapshots and database-native dumps offer
|
|
different consistency guarantees. A recent backup indicator is not a restore
|
|
test. Keep database restore checks and service recovery drills explicit.
|
|
|
|
The Hermes namespace is excluded from cloud backup policy. Its workspaces can
|
|
contain material restricted to local infrastructure. Do not remove that exclusion
|
|
to make a coverage dashboard greener.
|
|
|
|
## A one-day familiarization exercise
|
|
|
|
* Morning: follow one application from its Flux definition to its workload,
|
|
storage, service and ingress. Compare those manifests with live objects.
|
|
* Before lunch: inspect one scheduling incident and one storage incident using
|
|
the table above. Identify the responsible component without changing anything.
|
|
* Afternoon: run the isolated datastore restore verification, inspect backup
|
|
timestamps, then make and revert a harmless configuration change through Flux.
|
|
* Finish by finding each remaining issue in the implementation record and the
|
|
corresponding owner/runbook. Confirm access to Git, the management host, the
|
|
hosts' SSH path and protected recovery material without relying on SSO alone.
|
|
|
|
This provides a practical route to ownership; it cannot promise that every
|
|
hardware, storage or database disaster will be solvable in one day.
|