338 lines
18 KiB
Markdown
338 lines
18 KiB
Markdown
# Running Atlas without an AI assistant
|
|
|
|
Start here for ordinary operation. The repository describes desired state;
|
|
Kubernetes and Flux apply it. The [implementation record](CLUSTER_STABILIZATION.md)
|
|
separates completed repairs from open problems. The
|
|
[service inventory](cluster-audit-20261002/SERVICE_PLAN.md) covers the whole cluster.
|
|
|
|
## The small map
|
|
|
|
| Concern | Owner and configuration | What it does |
|
|
| --- | --- | --- |
|
|
| Desired Kubernetes configuration | Flux; `clusters/atlas/flux-system/` | Tracks `main`, applies service and infrastructure folders |
|
|
| Normal pod replacement and scheduling | Kubernetes; each workload manifest | Restarts failed containers and schedules replacements within placement and resource constraints |
|
|
| Application deployment | `services/<name>/`, or `infrastructure/<name>/` for foundations | Images, resources, probes, storage and networking |
|
|
| Durable application disks | Longhorn; `infrastructure/longhorn/` | Replication and volume attachment; not an application-consistent database backup by itself |
|
|
| Kubernetes datastore recovery | Native PostgreSQL tools and systemd on titan-db and titan-0b | Hourly protected LAN recovery bundles, independent of Kubernetes |
|
|
| Application PostgreSQL recovery | `atlas-application-postgres-backup.timer` on titan-0b; `infrastructure/host-backup/` | Daily logical database dumps with checksums, kept on the approved LAN host |
|
|
| Application backup orchestration | Soteria; `services/maintenance/apps/soteria-*` | Applies eligible-data policies through the chosen storage backend |
|
|
| Power loss and exceptional node recovery | Ananke; host configuration and `scripts/ops/cluster_power_*` | Orders shutdown/startup and bounded node recovery; it does not replace storage or database recovery |
|
|
| Node build/configuration | Metis and its sentinels; `services/maintenance/` | Approved node provisioning and host configuration |
|
|
| CI fault handling | Ariadne; `services/maintenance/apps/ariadne-*` | Bounded CI diagnosis/recovery; core cluster health must not depend on model calls |
|
|
| Health evidence | Grafana, VictoriaMetrics, Alertmanager; `services/monitoring/` | Shows application, node and storage evidence; a green controller alone is insufficient |
|
|
|
|
There is no new coordinating framework to learn. Prefer the native owner above.
|
|
Use a recovery tool only when its documented operation matches the failure.
|
|
|
|
The first priority is steady service operation. Repair recurring restarts and
|
|
storage pressure before adding workload or upgrading the platform. Keep known
|
|
unstable nodes quarantined; retain recovery copies and verify application health.
|
|
CI may queue when compatible capacity is unavailable. That is an explicit
|
|
capacity problem to repair, not a reason to overcommit the storage workers.
|
|
|
|
## Node administrative access
|
|
|
|
From the management workstation, use the existing SSH aliases. Most nodes use
|
|
`atlas`; the exceptions are `titan-jh` (theia), `oceanus` (Titan-23, oceanus), and
|
|
`tethys` (Titan-24, tethys). The normal SSH port is 2277 through Titan-jh.
|
|
|
|
In Vault, open `kv/atlas/nodes/<hostname>`. For example,
|
|
`kv/atlas/nodes/titan-13` contains `atlas_password` for sudo and `root_password`
|
|
for the root account. Custom metadata now records `admin_username`,
|
|
`admin_access_method`, `admin_access_verified_utc` and `root_password_state`.
|
|
Those are dated audit results, not a continuously updated health guarantee.
|
|
|
|
A normal manual check is:
|
|
|
|
```bash
|
|
ssh titan-13
|
|
sudo -k
|
|
sudo -v
|
|
sudo id -u
|
|
sudo passwd -S root
|
|
```
|
|
|
|
Enter the node's atlas password at the sudo prompt; never put it into the shell
|
|
command or a shared log. UID 0 proves administrator access. Some original nodes
|
|
already have passwordless sudo; this work did not add such grants. Nine native
|
|
Ubuntu accounts retain a locked direct root login; use their administrator and
|
|
sudo. A nonempty Vault root_password field alone does not prove root login works.
|
|
|
|
For scoped repair, `scripts/node_admin_access.py` runs on the target as root and
|
|
accepts a hostname plus password mapping on stdin. It compares password hashes
|
|
locally, reports only booleans, and changes passwords only with `--apply`.
|
|
`--retire-legacy-sudo` removes only the two recognized historical Metis grants
|
|
after password verification, preserving a root-only rollback copy. Do not run
|
|
that migration before independently testing password-backed sudo.
|
|
|
|
Keep SSH host-key verification enabled. Titan-09/10 currently answer on port 22
|
|
with changed keys; their reinstall identity must be established before sending
|
|
passwords. Titan-05 answers ping but has no tested admin listener. Titan-06/16
|
|
did not answer the current LAN checks. See the dated implementation record.
|
|
|
|
Metis source now provisions a native one-shot identity service instead of
|
|
relying exclusively on cloud-init. After a future recovery, check the unit,
|
|
verify its pending marker cleared, then test a separate SSH/sudo session before
|
|
returning the node to Kubernetes service. Image publication and a real replacement
|
|
boot remain separate acceptance checks; do not assume a source commit is deployed.
|
|
|
|
Ananke's allowlisted host actions can now obtain the atlas password through the
|
|
existing CSI synchronization, configured in
|
|
`services/maintenance/bootstrap/secretproviderclass.yaml`. The synchronized
|
|
Secrets are named `maintenance/ananke-sudo-<hostname>`, with key `password`.
|
|
Only the 13 verified workers needing that path are included; root passwords
|
|
remain in Vault. The CSI driver refreshes them every two minutes.
|
|
|
|
The two native Ananke configurations use:
|
|
|
|
```yaml
|
|
startup:
|
|
host_sudo_secret_namespace: maintenance
|
|
host_sudo_secret_name_template: "ananke-sudo-{node}"
|
|
host_sudo_secret_password_key: password
|
|
```
|
|
|
|
These are fields within the existing startup section, not a replacement config.
|
|
`scripts/configure_ananke_node_access.py` applies that narrow change, validates it
|
|
with Ananke, and retains a root-only backup. The coordinator remains Titan-db;
|
|
Titan-24 remains its peer. Credentials go to sudo stdin for predefined actions,
|
|
not arbitrary command arguments or logs. No new controller was introduced.
|
|
|
|
The peer's public key must also be present in the worker's atlas authorized_keys.
|
|
It is recorded in `kv/atlas/maintenance/metis-ssh-keys`, field `ananke_tethys_pub`.
|
|
`scripts/install_ananke_peer_key.py` accepts a hostname and verified public key on
|
|
stdin, preserves existing keys, and restricts new entries to Titan-24's source
|
|
address 192.168.22.26. It grants no sudo permissions. Verify the public key through
|
|
an already authenticated connection; never replace trusted host keys just to
|
|
make SSH connect. On October 4, both the coordinator and peer passed actual
|
|
privileged read-only worker checks after this repair.
|
|
|
|
This automated worker-password path requires the Kubernetes API. It is not a
|
|
standalone cold-start credential store. A read-only access check is not a power
|
|
failure/recovery drill. When onboarding another recovered worker, first validate
|
|
its Vault password and SSH identity, then add its explicit CSI mapping.
|
|
|
|
## Automatic recovery that can change services
|
|
|
|
Ananke's coordinator runs as `ananke.service` on Titan-db; its settings are in
|
|
`/etc/ananke/ananke.yaml`, with installed source in `/opt/ananke`. Titan-24 is
|
|
the UPS peer. Only the coordinator runs periodic cluster repair. UPS monitoring
|
|
is a separate responsibility within the same daemon: do not stop that daemon
|
|
casually to troubleshoot a Kubernetes helper.
|
|
|
|
The credential-helper repair checks every 60 seconds but acts only on a current
|
|
pod with an active image-pull failure and a matching warning less than ten
|
|
minutes old. It records an attempt before restarting a helper and allows at
|
|
most one attempt per helper per 30 minutes, across daemon restarts. Old events
|
|
and warnings for replaced pods do not justify another restart. Check helper
|
|
generation and Ananke's journal when investigating unexpected rollouts:
|
|
|
|
```bash
|
|
kubectl -n jenkins get deployment jenkins-vault-sync
|
|
ssh titan-db sudo journalctl -u ananke --since '10 minutes ago' --no-pager
|
|
ssh titan-db systemctl status ananke-update.timer
|
|
```
|
|
|
|
The native updater now preserves the running binary when its quality gate
|
|
fails. A binary-only maintenance install also preserves host/NUT configuration
|
|
and retains the prior binary for rollback. Read `scripts/install.sh` in the
|
|
Ananke repository before updating; `--binary-only --skip-deps` is the targeted
|
|
path, while the normal updater also applies host configuration templates.
|
|
|
|
Ariadne's `ARIADNE_HERMES_HUNG_BUILD_MINUTES` setting is in
|
|
`services/maintenance/apps/ariadne-deployment.yaml`. It is currently zero, which
|
|
disables elapsed-time-only cancellation. Its retained event confirmed that the
|
|
previous 120-minute setting canceled an active Metis image build. Other bounded
|
|
diagnosis/repair remains enabled. Do not re-enable cancellation until it can
|
|
distinguish a stalled job from a long-running job making useful progress.
|
|
|
|
Jenkins's Kubernetes cloud cap is currently one in
|
|
`services/jenkins/configmap-jcasc.yaml`. The shared default, IaC, Data Prepper and
|
|
Metis templates exclude quarantined nodes and storage workers. Custom inline
|
|
templates in other repos may need the same correction. Inspect actual agent
|
|
placement rather than assuming the default template controls every pipeline.
|
|
A Pending agent consumes the cap; inspect FailedScheduling events and resource
|
|
requests before deciding that Jenkins is hung. Do not increase concurrency to
|
|
solve an unschedulable agent. The tracked JCasC revision triggers a controller
|
|
rollout, so quiet the queue and arrange active-build completion before editing.
|
|
|
|
Node quarantine is declared in `infrastructure/core/node-maintenance.yaml`.
|
|
Those Node resources explicitly preserve hardware/worker/storage labels and use
|
|
Flux SSA Merge to retain other runtime fields. A cordon blocks ordinary new
|
|
placement; it must not repeatedly remove existing storage DaemonSets. After a
|
|
quarantine change, verify manager pod UIDs stay unchanged across reconciliation.
|
|
Keep prune protection on Node declarations; repair, uncordon in Git and verify
|
|
before considering their removal.
|
|
|
|
## Five-minute check
|
|
|
|
Run from the management host with the existing administrator configuration:
|
|
|
|
```bash
|
|
kubectl get nodes -o wide
|
|
flux get kustomizations -A
|
|
flux get helmreleases -A
|
|
kubectl get deployments,statefulsets -A
|
|
kubectl get pods -A --field-selector=status.phase=Pending
|
|
kubectl -n longhorn-system get volumes.longhorn.io
|
|
```
|
|
|
|
Then check the affected service's health endpoint or UI. `Running` does not mean
|
|
ready; `Ready` does not prove useful application behavior. A detached Longhorn
|
|
volume may be intentional if its workload is parked. Distinguish that from an
|
|
attached volume that is faulted or an application waiting for its disk.
|
|
|
|
Check the control-plane backups separately:
|
|
|
|
```bash
|
|
ssh titan-db sudo systemctl status atlas-k3s-backup.timer
|
|
ssh titan-0b sudo systemctl status atlas-k3s-replica.timer
|
|
ssh titan-db sudo cat /var/backups/atlas-k3s/latest/COMPLETE
|
|
ssh titan-0b sudo cat /var/backups/atlas-k3s-replica/latest/COMPLETE
|
|
```
|
|
|
|
`COMPLETE` contains timestamps, not database contents. The backup and replica
|
|
should normally be less than two hours old. Check the last service result too:
|
|
|
|
```bash
|
|
ssh titan-db sudo systemctl show atlas-k3s-backup.service -p Result -p ExecMainStatus
|
|
ssh titan-0b sudo systemctl show atlas-k3s-replica.service -p Result -p ExecMainStatus
|
|
```
|
|
|
|
Application database backups run separately each day at 08:15 UTC plus up to
|
|
five minutes of jitter. Check their timer and completion timestamp too:
|
|
|
|
```bash
|
|
ssh titan-0b sudo systemctl status atlas-application-postgres-backup.timer
|
|
ssh titan-0b sudo cat /var/backups/atlas-postgres/latest/COMPLETE
|
|
```
|
|
|
|
Their freshness alert fires after 36 hours. See the host backup notes before
|
|
running the isolated restore verification or changing retention.
|
|
|
|
## Find the cause before choosing a repair
|
|
|
|
| Symptom | First check | Usual next action |
|
|
| --- | --- | --- |
|
|
| Node NotReady | Power/network, kubelet status, disk space, pressure | Repair/quarantine the node; do not repeatedly restart every application |
|
|
| Pod Pending / FailedScheduling | `kubectl describe pod` scheduling events | Correct resource requests or eligible capacity; free cluster-wide RAM does not guarantee a compatible destination |
|
|
| ContainerCreating / FailedMount | PVC, Longhorn volume, attachment and engine events | Preserve data; repair the attachment or use a verified healthy node. Do not delete the PVC |
|
|
| CrashLoopBackOff / OOMKilled | Previous container logs and last termination reason | Fix the application or its resource envelope; lifetime restart counts are not a reason for repeated eviction |
|
|
| Ingress 502/503 | Service endpoints and backend readiness | Restore the backend or its dependency before changing DNS/TLS |
|
|
| Flux Ready but missing Helm workload | Helm manifest and drift correction | Check that drift detection is enabled and suitable for that release |
|
|
| All cluster API operations fail | titan-db PostgreSQL and control-plane nodes | The datastore is external PostgreSQL; an etcd snapshot procedure will not restore it |
|
|
| CI waiting while applications work | Jenkins queue, agent capacity and node I/O | Let work queue; adding agents to saturated Pi disks makes both CI and services worse |
|
|
|
|
Useful scoped commands:
|
|
|
|
```bash
|
|
kubectl -n NAMESPACE describe pod POD
|
|
kubectl -n NAMESPACE get events --sort-by=.lastTimestamp
|
|
kubectl -n NAMESPACE get endpointslices
|
|
kubectl -n NAMESPACE logs POD -c CONTAINER --previous --tail=80
|
|
```
|
|
|
|
Application logs can contain private data. Inspect them locally and redact before
|
|
sharing. Never paste Secrets, tokens, source records or inference bodies into a
|
|
ticket or assistant conversation.
|
|
|
|
## Make a normal change
|
|
|
|
1. Edit the service's tracked manifest. Placement, requests, limits and probes
|
|
belong together in that workload's configuration.
|
|
2. Render and validate the affected folder; inspect the Flux preview.
|
|
3. Commit and push a small change to the tracked branch.
|
|
4. Reconcile that service and verify its application behavior and storage.
|
|
|
|
For example:
|
|
|
|
```bash
|
|
kubectl kustomize services/monitoring > /tmp/monitoring-render.yaml
|
|
kubectl apply --dry-run=client -k services/monitoring
|
|
flux diff kustomization monitoring --path services/monitoring
|
|
git diff --check
|
|
git diff
|
|
```
|
|
|
|
After committing and pushing the reviewed files:
|
|
|
|
```bash
|
|
flux reconcile kustomization monitoring -n flux-system --with-source
|
|
kubectl -n monitoring get deployments,statefulsets,pods
|
|
```
|
|
|
|
A Flux diff exits nonzero when it finds changes; read the output to distinguish
|
|
that from a validation failure. Kustomize renders a HelmRelease, not the chart's
|
|
workloads. For chart values, also inspect the chart's rendered Deployment or
|
|
StatefulSet. OpenSearch, for example, uses `values.nodeAffinity`.
|
|
|
|
Rollback a configuration change by reverting its focused commit and reconciling
|
|
the same service. Do not blindly revert storage migrations or database changes;
|
|
those require their recovery procedure. Do not delete data to clear a red status.
|
|
|
|
## Backups and recovery
|
|
|
|
The external datastore's [host backup notes](../infrastructure/host-backup/NOTES.md)
|
|
give the actual paths, schedule, credential scope, isolated restore check and
|
|
rollback. Backups contain sensitive data and remain root-only on approved LAN
|
|
hosts. They are not encrypted at rest or off-site disaster recovery copies.
|
|
|
|
Soteria is responsible for eligible application backups, not for making the
|
|
control plane boot. Live Longhorn snapshots and database-native dumps offer
|
|
different consistency guarantees. A recent backup indicator is not a restore
|
|
test. Keep database restore checks and service recovery drills explicit.
|
|
|
|
For cloud backups, inspect completed backup timestamps, not just the backup
|
|
target's Available status or a successful bucket listing. The October 3 audit
|
|
found a readable Backblaze bucket rejecting uploads with `storage cap exceeded`.
|
|
Check Backblaze Caps & Alerts and agree the storage budget before changing it.
|
|
Do not delete old recovery points merely to make uploads work. Distinguish
|
|
visible-object bytes from billed storage, including retained versions. The
|
|
current diagnosis and counts are in the implementation record.
|
|
|
|
The Hermes namespace is excluded from cloud backup policy. Its workspaces can
|
|
contain material restricted to local infrastructure. Do not remove that exclusion
|
|
to make a coverage dashboard greener.
|
|
|
|
## A one-day familiarization exercise
|
|
|
|
* Morning: follow one application from its Flux definition to its workload,
|
|
storage, service and ingress. Compare those manifests with live objects.
|
|
* Before lunch: inspect one scheduling incident and one storage incident using
|
|
the table above. Identify the responsible component without changing anything.
|
|
* Afternoon: run the isolated datastore restore verification, inspect backup
|
|
timestamps, then make and revert a harmless configuration change through Flux.
|
|
* Finish by finding each remaining issue in the implementation record and the
|
|
corresponding owner/runbook. Confirm access to Git, the management host, the
|
|
hosts' SSH path and protected recovery material without relying on SSO alone.
|
|
|
|
This provides a practical route to ownership; it cannot promise that every
|
|
hardware, storage or database disaster will be solvable in one day.
|
|
|
|
## Physical checks and return to service
|
|
|
|
For an offline Pi, record the board identity, supply model and its 5 V rating,
|
|
power/activity LEDs, HDMI boot result, Ethernet link and current router address.
|
|
Preserve the original SD/USB media. Test with separate known-good spare boot media
|
|
before concluding that a board has failed; do not clone a live node identity by
|
|
moving another cluster node's boot card into it.
|
|
|
|
Pi 5 power must be evaluated at its 5 V output mode, not a charger's headline
|
|
wattage. A controlled supply-and-cable substitution isolates the power path.
|
|
Record any inline adapters and USB loads. Do not unplug runtime USB storage on
|
|
a running node. Titan-11 carries important services; arrange workload movement
|
|
before disconnecting it. See the official Raspberry Pi power documentation for
|
|
5 V / 5 A capability and cable-loss requirements.
|
|
|
|
A repaired node returns only after stable power, clean storage/kernel logs,
|
|
reliable networking and representative-load testing. Observe it before restoring
|
|
normal placement. Titan-14/18 quarantine is explicitly tracked in
|
|
`infrastructure/core/node-maintenance.yaml`; `prune: disabled` protects the Node
|
|
objects from deletion when that maintenance declaration is eventually removed.
|
|
Clear `spec.unschedulable` through Git after validation, then reconcile core.
|
|
|
|
Container runtime media and Longhorn data disks are separate. A Pi with terabytes
|
|
of healthy application storage can still stall because containerd, image unpacking
|
|
and logs live on a small USB flash device. Include the runtime disk in hardware
|
|
repair and capacity checks.
|