# Running Atlas without an AI assistant Start here for ordinary operation. The repository describes desired state; Kubernetes and Flux apply it. The [implementation record](CLUSTER_STABILIZATION.md) separates completed repairs from open problems. The [service inventory](cluster-audit-20261002/SERVICE_PLAN.md) covers the whole cluster. ## The small map | Concern | Owner and configuration | What it does | | --- | --- | --- | | Desired Kubernetes configuration | Flux; `clusters/atlas/flux-system/` | Tracks `main`, applies service and infrastructure folders | | Normal pod replacement and scheduling | Kubernetes; each workload manifest | Restarts failed containers and schedules replacements within placement and resource constraints | | Application deployment | `services//`, or `infrastructure//` for foundations | Images, resources, probes, storage and networking | | Durable application disks | Longhorn; `infrastructure/longhorn/` | Replication and volume attachment; not an application-consistent database backup by itself | | Kubernetes datastore recovery | Native PostgreSQL tools and systemd on titan-db and titan-0b | Hourly protected LAN recovery bundles, independent of Kubernetes | | Application PostgreSQL recovery | `atlas-application-postgres-backup.timer` on titan-0b; `infrastructure/host-backup/` | Daily logical database dumps with checksums, kept on the approved LAN host | | Application backup orchestration | Soteria; `services/maintenance/apps/soteria-*` | Applies eligible-data policies through the chosen storage backend | | Power loss and exceptional node recovery | Ananke; host configuration and `scripts/ops/cluster_power_*` | Orders shutdown/startup and bounded node recovery; it does not replace storage or database recovery | | Node build/configuration | Metis and its sentinels; `services/maintenance/` | Approved node provisioning and host configuration | | CI fault handling | Ariadne; `services/maintenance/apps/ariadne-*` | Bounded CI diagnosis/recovery; core cluster health must not depend on model calls | | Health evidence | Grafana, VictoriaMetrics, Alertmanager; `services/monitoring/` | Shows application, node and storage evidence; a green controller alone is insufficient | There is no new coordinating framework to learn. Prefer the native owner above. Use a recovery tool only when its documented operation matches the failure. The first priority is steady service operation. Repair recurring restarts and storage pressure before adding workload or upgrading the platform. Keep known unstable nodes quarantined; retain recovery copies and verify application health. CI may queue when compatible capacity is unavailable. That is an explicit capacity problem to repair, not a reason to overcommit the storage workers. ## Node administrative access From the management workstation, use the existing SSH aliases. Most nodes use `atlas`; the exceptions are `titan-jh` (theia), `oceanus` (Titan-23, oceanus), and `tethys` (Titan-24, tethys). The normal SSH port is 2277 through Titan-jh. In Vault, open `kv/atlas/nodes/`. For example, `kv/atlas/nodes/titan-13` contains `atlas_password` for sudo and `root_password` for the root account. Custom metadata now records `admin_username`, `admin_access_method`, `admin_access_verified_utc` and `root_password_state`. Those are dated audit results, not a continuously updated health guarantee. A normal manual check is: ```bash ssh titan-13 sudo -k sudo -v sudo id -u sudo passwd -S root ``` Enter the node's atlas password at the sudo prompt; never put it into the shell command or a shared log. UID 0 proves administrator access. Some original nodes already have passwordless sudo; this work did not add such grants. Nine native Ubuntu accounts retain a locked direct root login; use their administrator and sudo. A nonempty Vault root_password field alone does not prove root login works. For scoped repair, `scripts/node_admin_access.py` runs on the target as root and accepts a hostname plus password mapping on stdin. It compares password hashes locally, reports only booleans, and changes passwords only with `--apply`. `--retire-legacy-sudo` removes only the two recognized historical Metis grants after password verification, preserving a root-only rollback copy. Do not run that migration before independently testing password-backed sudo. Keep SSH host-key verification enabled. Titan-09/10 currently answer on port 22 with changed keys; their reinstall identity must be established before sending passwords. Titan-05 answers ping but has no tested admin listener. Titan-06/16 did not answer the current LAN checks. See the dated implementation record. Metis source now provisions a native one-shot identity service instead of relying exclusively on cloud-init. After a future recovery, check the unit, verify its pending marker cleared, then test a separate SSH/sudo session before returning the node to Kubernetes service. Metis 0.1.0-403 is now published and Ready in the cluster. A real replacement boot remains a separate acceptance check; a successful image build does not prove hardware provisioning. Ananke's allowlisted host actions can now obtain the atlas password through the existing CSI synchronization, configured in `services/maintenance/bootstrap/secretproviderclass.yaml`. The synchronized Secrets are named `maintenance/ananke-sudo-`, with key `password`. Only the 13 verified workers needing that path are included; root passwords remain in Vault. The CSI driver refreshes them every two minutes. The two native Ananke configurations use: ```yaml startup: host_sudo_secret_namespace: maintenance host_sudo_secret_name_template: "ananke-sudo-{node}" host_sudo_secret_password_key: password ``` These are fields within the existing startup section, not a replacement config. `scripts/configure_ananke_node_access.py` applies that narrow change, validates it with Ananke, and retains a root-only backup. The coordinator remains Titan-db; Titan-24 remains its peer. Credentials go to sudo stdin for predefined actions, not arbitrary command arguments or logs. No new controller was introduced. The peer's public key must also be present in the worker's atlas authorized_keys. It is recorded in `kv/atlas/maintenance/metis-ssh-keys`, field `ananke_tethys_pub`. `scripts/install_ananke_peer_key.py` accepts a hostname and verified public key on stdin, preserves existing keys, and restricts new entries to Titan-24's source address 192.168.22.26. It grants no sudo permissions. Verify the public key through an already authenticated connection; never replace trusted host keys just to make SSH connect. On October 4, both the coordinator and peer passed actual privileged read-only worker checks after this repair. This automated worker-password path requires the Kubernetes API. It is not a standalone cold-start credential store. A read-only access check is not a power failure/recovery drill. When onboarding another recovered worker, first validate its Vault password and SSH identity, then add its explicit CSI mapping. ## Automatic recovery that can change services The `node-nofile` setup helper prepares systemd file limits and live inotify settings. It no longer restarts K3s when the file-limit definition changes. New unit limits take effect at the next planned K3s restart; do not assume a helper rollout activated them. Its small reservations reflect an idle setup process, while its burst limits are preserved. Ananke's coordinator runs as `ananke.service` on Titan-db; its settings are in `/etc/ananke/ananke.yaml`, with installed source in `/opt/ananke`. Titan-24 is the UPS peer. Only the coordinator runs periodic cluster repair. UPS monitoring is a separate responsibility within the same daemon: do not stop that daemon casually to troubleshoot a Kubernetes helper. The credential-helper repair checks every 60 seconds but acts only on a current pod with an active image-pull failure and a matching warning less than ten minutes old. It records an attempt before restarting a helper and allows at most one attempt per helper per 30 minutes, across daemon restarts. Old events and warnings for replaced pods do not justify another restart. Check helper generation and Ananke's journal when investigating unexpected rollouts: ```bash kubectl -n jenkins get deployment jenkins-vault-sync ssh titan-db sudo journalctl -u ananke --since '10 minutes ago' --no-pager ssh titan-db systemctl status ananke-update.timer ``` The native updater now preserves the running binary when its quality gate fails. A binary-only maintenance install also preserves host/NUT configuration and retains the prior binary for rollback. Read `scripts/install.sh` in the Ananke repository before updating; `--binary-only --skip-deps` is the targeted path, while the normal updater also applies host configuration templates. Ariadne's `ARIADNE_HERMES_HUNG_BUILD_MINUTES` setting is in `services/maintenance/apps/ariadne-deployment.yaml`. It is currently zero, which disables elapsed-time-only cancellation. Its retained event confirmed that the previous 120-minute setting canceled an active Metis image build. Other bounded diagnosis/repair remains enabled. Do not re-enable cancellation until it can distinguish a stalled job from a long-running job making useful progress. Jenkins's Kubernetes cloud cap is currently one in `services/jenkins/configmap-jcasc.yaml`. The active inline templates in Ariadne, Atlasbot, Ananke, Metis, Pegasus, Soteria, bstein-dev-home, IaC and Data Prepper now require Pi workers and exclude quarantined and storage workers. Inspect actual agent placement rather than assuming the shared default controls a custom template. A Pending agent consumes the cap; inspect FailedScheduling events and resource requests before deciding that Jenkins is hung. Do not increase concurrency to solve an unschedulable agent. The tracked JCasC revision triggers a controller rollout, so quiet the queue and arrange active-build completion before editing. Custom pipelines explicitly use `workspaceVolume dynamicPVC(...)` with the `ci-scratch` StorageClass in `infrastructure/core/storageclass-ci-scratch.yaml`. Workspaces and compiler/package caches use that volume. Docker-in-Docker stores its layers in the same volume under `.cache/docker`, outside the checkout. Ordinary workspaces request 20 GiB; Docker builds request 40 GiB. Builds remain serialized. This moves heavy scratch I/O off the Pi runtime USB drive. Soteria also compiles its UI and binary on scratch before Kaniko assembles the small `Dockerfile.runtime` image. Compiling inside Kaniko would write toolchains and compiler intermediates into the container runtime again, defeating that isolation. Data Prepper now uses the existing Docker-in-Docker path with a 40 GiB scratch PVC for image layers. A scratch workspace alone does not relocate Kaniko extraction out of a container root filesystem. The Docker API binds only to pod loopback; registry passwords go to stdin and the temporary client config is removed afterward. ARM64 output remains explicit. Keep registry-credential umask changes local to credential creation. Soteria explicitly sets the executable mode and nonroot ownership before image assembly; a successful compile alone does not prove the deployed user can execute it. `ci-scratch` is disposable, single-replica Longhorn storage on the existing `astreae` disks. Jenkins owns its PVC through the agent pod; deleting that pod removes its scratch PVC and volume. Never use this class for application data, credentials, recovery copies or retained artifacts. Test artifacts are still archived to the Jenkins controller with their existing 30-day policy. A lost scratch volume requires rerunning the build. Soteria excludes the ci-scratch class from application backups alongside local-path; preserve Hermes exclusions. Check orphaned PVCs before assuming cleanup occurred; do not delete unrelated retained volumes. Node quarantine is declared in `infrastructure/core/node-maintenance.yaml` for Titan-04/05/06/08/11/12/14/18. Offline nodes remain quarantined after reconnecting until their physical and loaded-operation checks pass. Those Node resources explicitly preserve hardware/worker/storage labels and use Flux SSA Merge to retain other runtime fields. A cordon blocks ordinary new placement; it must not repeatedly remove existing storage DaemonSets. After a quarantine change, verify manager pod UIDs stay unchanged across reconciliation. Keep prune protection on Node declarations; repair, uncordon in Git and verify before considering their removal. Keep one-off Flux Jobs completed without a TTL, or remove them from active resources after success. A TTL deletes the object and Flux then creates it again. The archived Titan-24 emergency root sweep is deliberately inactive; its zero-percent thresholds are not a normal maintenance schedule. The Metis SSH public-key bootstrap now retains its completion. Kubelet owns container-image garbage collection. All 19 reachable nodes were verified with 85% high / 80% low thresholds. The legacy-named node-image-sweeper no longer runs `crictl rmi --prune` or removes files from containerd/image-import directories. Its remaining role is bounded host-log/package-cache maintenance. The UID-aware Python helper preserves active pod logs in both `/var/log/pods` and Armbian `/var/log.hdd/pods`, skips unavailable inventories, and provides `--dry-run`. Native kubelet rotation controls active container logs. The node-image-sweeper DaemonSet uses a startup probe that passes after its first cleanup, then performs no continuous probe process. Its normal rollout permits one unavailable helper. Offline terminating pods can consume that budget; inspect those before raising it. This does not change application availability. ## Ingress and search capacity Traefik runs two replicas on separate NVMe control-plane nodes. Its internal `/ping` readiness check gates traffic. The public VIP remains `192.168.22.9`; Gitea SSH and Wolf share it using their existing ports. MetalLB's `ingress-pool` and `ingress-adv` restrict its announcer to the control planes. The private `192.168.22.50` pool does the same and retains `externalTrafficPolicy: Local`, source restrictions, TLS and service authentication. The communications pool retains `.4` through `.6`; its Local-policy services need a speaker on their actual endpoint node. Do not restrict all pools to ingress nodes indiscriminately. MetalLB uses native L2 mode, without unused FRR/BGP sidecars. There are no BGP peers or advertisements. Titan-05/06 are excluded from the speaker DaemonSet until their recovery checks pass. Remove that explicit exclusion in `infrastructure/metallb/helmrelease.yaml` when restoring them. Do not enable BGP features later without revisiting this choice. The speaker requests 128 MiB and uses five-second probes. Native L2 still has one announcing node per VIP. OpenSearch's existing PVC is now 1280 GiB. Its expansion reduced measured usage from 94% to 76%, allowing the blocked maintenance policy updates to resume. The existing scheduled tuner now reports failing operation/path/status without index contents. Its next scheduled run passed. No retention period or disk watermark was relaxed. Inspect free space and ingestion growth periodically; previous growth while writes were blocked is not a useful capacity forecast. The current volume still has one Longhorn replica; it is not storage HA. The scheduled tuner uses the native maintenance-batch priority class with preemptionPolicy Never. It prefers Pi 5 workers but permits Pi 4 workers when capacity is unavailable. Routine maintenance should wait rather than evict running work; check scheduling deadlines instead of raising its priority. **Do not shrink the PVC or blindly revert its expansion commit.** ## Five-minute check Run from the management host with the existing administrator configuration: ```bash kubectl get nodes -o wide flux get kustomizations -A flux get helmreleases -A kubectl get deployments,statefulsets -A kubectl get pods -A --field-selector=status.phase=Pending kubectl -n longhorn-system get volumes.longhorn.io ``` Then check the affected service's health endpoint or UI. `Running` does not mean ready; `Ready` does not prove useful application behavior. A detached Longhorn volume may be intentional if its workload is parked. Distinguish that from an attached volume that is faulted or an application waiting for its disk. Check the control-plane backups separately: ```bash ssh titan-db sudo systemctl status atlas-k3s-backup.timer ssh titan-0b sudo systemctl status atlas-k3s-replica.timer ssh titan-db sudo cat /var/backups/atlas-k3s/latest/COMPLETE ssh titan-0b sudo cat /var/backups/atlas-k3s-replica/latest/COMPLETE ``` `COMPLETE` contains timestamps, not database contents. The backup and replica should normally be less than two hours old. Check the last service result too: ```bash ssh titan-db sudo systemctl show atlas-k3s-backup.service -p Result -p ExecMainStatus ssh titan-0b sudo systemctl show atlas-k3s-replica.service -p Result -p ExecMainStatus ``` Application database backups run separately each day at 08:15 UTC plus up to five minutes of jitter. Check their timer and completion timestamp too: ```bash ssh titan-0b sudo systemctl status atlas-application-postgres-backup.timer ssh titan-0b sudo cat /var/backups/atlas-postgres/latest/COMPLETE ``` Their freshness alert fires after 36 hours. See the host backup notes before running the isolated restore verification or changing retention. ## Check a settling window Use the read-only helper after the rollout finishes, then compare 10-15 minutes later. It uses ordinary kubectl access; it does not repair, restart or delete anything and has no AI or provider dependency. ```bash python3 scripts/ops/cluster_settle_check.py --output /tmp/atlas-before.json # Wait 10-15 minutes with normal workload activity. python3 scripts/ops/cluster_settle_check.py --output /tmp/atlas-after.json --previous /tmp/atlas-before.json ``` Review restart increases, replaced service pods, unready workloads, requested but unavailable storage and new warning evidence. Expected short-lived CI agent turnover is different from an application controller repeatedly replacing pods. Known offline nodes are reported separately, not hidden. Pods already marked for deletion are not considered available merely because Ready remains true. Scheduler Preempted events are included even when their event type is Normal. A newly observed event has an unknown counter delta because its lifetime count may predate the window. The helper deliberately omits Secret values, environments, log bodies and event messages. Pair it with the service health endpoints and node I/O/memory metrics; Kubernetes can report a node Ready while its runtime disk is stalled. The helper checks Kustomizations and HelmReleases, not image-release policies or application incident history; inspect those separately before claiming all automation green. Current placement is deliberately conservative: Gitea and ClamAV use Titan-22's CPU/RAM as last-resort general capacity; they request no GPU. Nextcloud stays on a Pi worker. Titan-08 is also quarantined after a build saturated its runtime USB drive. Keep Jenkins at one agent until ordinary builds and application activity coexist without recurring pressure. See the dated cycle record for measurements. Nextcloud startup now preserves installed apps and their external-app keys. App upgrades are separate maintenance, rather than downloads/deletions on every pod replacement. Its single-writer volumes require maxSurge 0 / maxUnavailable 1; a replacement therefore has a short planned interruption. ClamAV readiness uses its native PING/PONG response; a running process alone does not open the endpoint. ## Find the cause before choosing a repair | Symptom | First check | Usual next action | | --- | --- | --- | | Node NotReady | Power/network, kubelet status, disk space, pressure | Repair/quarantine the node; do not repeatedly restart every application | | Pod Pending / FailedScheduling | `kubectl describe pod` scheduling events | Correct resource requests or eligible capacity; free cluster-wide RAM does not guarantee a compatible destination | | ContainerCreating / FailedMount | PVC, Longhorn volume, attachment and engine events | Preserve data; repair the attachment or use a verified healthy node. Do not delete the PVC | | CrashLoopBackOff / OOMKilled | Previous container logs and last termination reason | Fix the application or its resource envelope; lifetime restart counts are not a reason for repeated eviction | | Ingress 502/503 | Service endpoints and backend readiness | Restore the backend or its dependency before changing DNS/TLS | | Flux Ready but missing Helm workload | Helm manifest and drift correction | Check that drift detection is enabled and suitable for that release | | All cluster API operations fail | titan-db PostgreSQL and control-plane nodes | The datastore is external PostgreSQL; an etcd snapshot procedure will not restore it | | CI waiting while applications work | Jenkins queue, agent capacity and node I/O | Let work queue; adding agents to saturated Pi disks makes both CI and services worse | Useful scoped commands: ```bash kubectl -n NAMESPACE describe pod POD kubectl -n NAMESPACE get events --sort-by=.lastTimestamp kubectl -n NAMESPACE get endpointslices kubectl -n NAMESPACE logs POD -c CONTAINER --previous --tail=80 ``` Application logs can contain private data. Inspect them locally and redact before sharing. Never paste Secrets, tokens, source records or inference bodies into a ticket or assistant conversation. ## Make a normal change 1. Edit the service's tracked manifest. Placement, requests, limits and probes belong together in that workload's configuration. 2. Render and validate the affected folder; inspect the Flux preview. 3. Commit and push a small change to the tracked branch. 4. Reconcile that service and verify its application behavior and storage. For example: ```bash kubectl kustomize services/monitoring > /tmp/monitoring-render.yaml kubectl apply --dry-run=client -k services/monitoring flux diff kustomization monitoring --path services/monitoring git diff --check git diff ``` After committing and pushing the reviewed files: ```bash flux reconcile kustomization monitoring -n flux-system --with-source kubectl -n monitoring get deployments,statefulsets,pods ``` A Flux diff exits nonzero when it finds changes; read the output to distinguish that from a validation failure. Kustomize renders a HelmRelease, not the chart's workloads. For chart values, also inspect the chart's rendered Deployment or StatefulSet. OpenSearch, for example, uses `values.nodeAffinity`. Rollback a configuration change by reverting its focused commit and reconciling the same service. Do not blindly revert storage migrations or database changes; those require their recovery procedure. Do not delete data to clear a red status. ## Backups and recovery The external datastore's [host backup notes](../infrastructure/host-backup/NOTES.md) give the actual paths, schedule, credential scope, isolated restore check and rollback. Backups contain sensitive data and remain root-only on approved LAN hosts. They are not encrypted at rest or off-site disaster recovery copies. Soteria is responsible for eligible application backups, not for making the control plane boot. Live Longhorn snapshots and database-native dumps offer different consistency guarantees. A recent backup indicator is not a restore test. Keep database restore checks and service recovery drills explicit. For cloud backups, inspect completed backup timestamps, not just the backup target's Available status or a successful bucket listing. The October 3 audit found a readable Backblaze bucket rejecting uploads with `storage cap exceeded`. Check Backblaze Caps & Alerts and agree the storage budget before changing it. Do not delete old recovery points merely to make uploads work. Distinguish visible-object bytes from billed storage, including retained versions. The current diagnosis and counts are in the implementation record. The Hermes namespace is excluded from cloud backup policy. Its workspaces can contain material restricted to local infrastructure. Do not remove that exclusion to make a coverage dashboard greener. ## A one-day familiarization exercise * Morning: follow one application from its Flux definition to its workload, storage, service and ingress. Compare those manifests with live objects. * Before lunch: inspect one scheduling incident and one storage incident using the table above. Identify the responsible component without changing anything. * Afternoon: run the isolated datastore restore verification, inspect backup timestamps, then make and revert a harmless configuration change through Flux. * Finish by finding each remaining issue in the implementation record and the corresponding owner/runbook. Confirm access to Git, the management host, the hosts' SSH path and protected recovery material without relying on SSO alone. This provides a practical route to ownership; it cannot promise that every hardware, storage or database disaster will be solvable in one day. ## Physical checks and return to service For an offline Pi, record the board identity, supply model and its 5 V rating, power/activity LEDs, HDMI boot result, Ethernet link and current router address. Preserve the original SD/USB media. Test with separate known-good spare boot media before concluding that a board has failed; do not clone a live node identity by moving another cluster node's boot card into it. Pi 5 power must be evaluated at its 5 V output mode, not a charger's headline wattage. A controlled supply-and-cable substitution isolates the power path. Record any inline adapters and USB loads. Do not unplug runtime USB storage on a running node. Titan-11 carries important services; arrange workload movement before disconnecting it. See the official Raspberry Pi power documentation for 5 V / 5 A capability and cable-loss requirements. A repaired node returns only after stable power, clean storage/kernel logs, reliable networking and representative-load testing. Observe it before restoring normal placement. Titan-14/18 quarantine is explicitly tracked in `infrastructure/core/node-maintenance.yaml`; `prune: disabled` protects the Node objects from deletion when that maintenance declaration is eventually removed. Clear `spec.unschedulable` through Git after validation, then reconcile core. Container runtime media and Longhorn data disks are separate. A Pi with terabytes of healthy application storage can still stall because containerd, image unpacking and logs live on a small USB flash device. Include the runtime disk in hardware repair and capacity checks.