463 lines
26 KiB
Markdown
463 lines
26 KiB
Markdown
# Running Atlas without an AI assistant
|
|
|
|
Start here for ordinary operation. The repository describes desired state;
|
|
Kubernetes and Flux apply it. The [implementation record](CLUSTER_STABILIZATION.md)
|
|
separates completed repairs from open problems. The
|
|
[service inventory](cluster-audit-20261002/SERVICE_PLAN.md) covers the whole cluster.
|
|
|
|
## The small map
|
|
|
|
| Concern | Owner and configuration | What it does |
|
|
| --- | --- | --- |
|
|
| Desired Kubernetes configuration | Flux; `clusters/atlas/flux-system/` | Tracks `main`, applies service and infrastructure folders |
|
|
| Normal pod replacement and scheduling | Kubernetes; each workload manifest | Restarts failed containers and schedules replacements within placement and resource constraints |
|
|
| Application deployment | `services/<name>/`, or `infrastructure/<name>/` for foundations | Images, resources, probes, storage and networking |
|
|
| Durable application disks | Longhorn; `infrastructure/longhorn/` | Replication and volume attachment; not an application-consistent database backup by itself |
|
|
| Kubernetes datastore recovery | Native PostgreSQL tools and systemd on titan-db and titan-0b | Hourly protected LAN recovery bundles, independent of Kubernetes |
|
|
| Application PostgreSQL recovery | `atlas-application-postgres-backup.timer` on titan-0b; `infrastructure/host-backup/` | Daily logical database dumps with checksums, kept on the approved LAN host |
|
|
| Application backup orchestration | Soteria; `services/maintenance/apps/soteria-*` | Applies eligible-data policies through the chosen storage backend |
|
|
| Power loss and exceptional node recovery | Ananke; host configuration and `scripts/ops/cluster_power_*` | Orders shutdown/startup and bounded node recovery; it does not replace storage or database recovery |
|
|
| Node build/configuration | Metis and its sentinels; `services/maintenance/` | Approved node provisioning and host configuration |
|
|
| CI fault handling | Ariadne; `services/maintenance/apps/ariadne-*` | Bounded CI diagnosis/recovery; core cluster health must not depend on model calls |
|
|
| Health evidence | Grafana, VictoriaMetrics, Alertmanager; `services/monitoring/` | Shows application, node and storage evidence; a green controller alone is insufficient |
|
|
|
|
There is no new coordinating framework to learn. Prefer the native owner above.
|
|
Use a recovery tool only when its documented operation matches the failure.
|
|
|
|
The first priority is steady service operation. Repair recurring restarts and
|
|
storage pressure before adding workload or upgrading the platform. Keep known
|
|
unstable nodes quarantined; retain recovery copies and verify application health.
|
|
CI may queue when compatible capacity is unavailable. That is an explicit
|
|
capacity problem to repair, not a reason to overcommit the storage workers.
|
|
|
|
## Node administrative access
|
|
|
|
From the management workstation, use the existing SSH aliases. Most nodes use
|
|
`atlas`; the exceptions are `titan-jh` (theia), `oceanus` (Titan-23, oceanus), and
|
|
`tethys` (Titan-24, tethys). The normal SSH port is 2277 through Titan-jh.
|
|
|
|
In Vault, open `kv/atlas/nodes/<hostname>`. For example,
|
|
`kv/atlas/nodes/titan-13` contains `atlas_password` for sudo and `root_password`
|
|
for the root account. Custom metadata now records `admin_username`,
|
|
`admin_access_method`, `admin_access_verified_utc` and `root_password_state`.
|
|
Those are dated audit results, not a continuously updated health guarantee.
|
|
|
|
A normal manual check is:
|
|
|
|
```bash
|
|
ssh titan-13
|
|
sudo -k
|
|
sudo -v
|
|
sudo id -u
|
|
sudo passwd -S root
|
|
```
|
|
|
|
Enter the node's atlas password at the sudo prompt; never put it into the shell
|
|
command or a shared log. UID 0 proves administrator access. Some original nodes
|
|
already have passwordless sudo; this work did not add such grants. Nine native
|
|
Ubuntu accounts retain a locked direct root login; use their administrator and
|
|
sudo. A nonempty Vault root_password field alone does not prove root login works.
|
|
|
|
For scoped repair, `scripts/node_admin_access.py` runs on the target as root and
|
|
accepts a hostname plus password mapping on stdin. It compares password hashes
|
|
locally, reports only booleans, and changes passwords only with `--apply`.
|
|
`--retire-legacy-sudo` removes only the two recognized historical Metis grants
|
|
after password verification, preserving a root-only rollback copy. Do not run
|
|
that migration before independently testing password-backed sudo.
|
|
|
|
Keep SSH host-key verification enabled. Titan-09/10 currently answer on port 22
|
|
with changed keys; their reinstall identity must be established before sending
|
|
passwords. Titan-05 answers ping but has no tested admin listener. Titan-06/16
|
|
did not answer the current LAN checks. See the dated implementation record.
|
|
|
|
Metis source now provisions a native one-shot identity service instead of
|
|
relying exclusively on cloud-init. After a future recovery, check the unit,
|
|
verify its pending marker cleared, then test a separate SSH/sudo session before
|
|
returning the node to Kubernetes service. Metis 0.1.0-403 is now published and
|
|
Ready in the cluster. A real replacement boot remains a separate acceptance
|
|
check; a successful image build does not prove hardware provisioning.
|
|
|
|
Ananke's allowlisted host actions can now obtain the atlas password through the
|
|
existing CSI synchronization, configured in
|
|
`services/maintenance/bootstrap/secretproviderclass.yaml`. The synchronized
|
|
Secrets are named `maintenance/ananke-sudo-<hostname>`, with key `password`.
|
|
Only the 13 verified workers needing that path are included; root passwords
|
|
remain in Vault. The CSI driver refreshes them every two minutes.
|
|
|
|
The two native Ananke configurations use:
|
|
|
|
```yaml
|
|
startup:
|
|
host_sudo_secret_namespace: maintenance
|
|
host_sudo_secret_name_template: "ananke-sudo-{node}"
|
|
host_sudo_secret_password_key: password
|
|
```
|
|
|
|
These are fields within the existing startup section, not a replacement config.
|
|
`scripts/configure_ananke_node_access.py` applies that narrow change, validates it
|
|
with Ananke, and retains a root-only backup. The coordinator remains Titan-db;
|
|
Titan-24 remains its peer. Credentials go to sudo stdin for predefined actions,
|
|
not arbitrary command arguments or logs. No new controller was introduced.
|
|
|
|
The peer's public key must also be present in the worker's atlas authorized_keys.
|
|
It is recorded in `kv/atlas/maintenance/metis-ssh-keys`, field `ananke_tethys_pub`.
|
|
`scripts/install_ananke_peer_key.py` accepts a hostname and verified public key on
|
|
stdin, preserves existing keys, and restricts new entries to Titan-24's source
|
|
address 192.168.22.26. It grants no sudo permissions. Verify the public key through
|
|
an already authenticated connection; never replace trusted host keys just to
|
|
make SSH connect. On October 4, both the coordinator and peer passed actual
|
|
privileged read-only worker checks after this repair.
|
|
|
|
This automated worker-password path requires the Kubernetes API. It is not a
|
|
standalone cold-start credential store. A read-only access check is not a power
|
|
failure/recovery drill. When onboarding another recovered worker, first validate
|
|
its Vault password and SSH identity, then add its explicit CSI mapping.
|
|
|
|
## Automatic recovery that can change services
|
|
|
|
The `node-nofile` setup helper prepares systemd file limits and live inotify
|
|
settings. It no longer restarts K3s when the file-limit definition changes. New
|
|
unit limits take effect at the next planned K3s restart; do not assume a helper
|
|
rollout activated them. Its small reservations reflect an idle setup process,
|
|
while its burst limits are preserved.
|
|
|
|
Ananke's coordinator runs as `ananke.service` on Titan-db; its settings are in
|
|
`/etc/ananke/ananke.yaml`, with installed source in `/opt/ananke`. Titan-24 is
|
|
the UPS peer. Only the coordinator runs periodic cluster repair. UPS monitoring
|
|
is a separate responsibility within the same daemon: do not stop that daemon
|
|
casually to troubleshoot a Kubernetes helper.
|
|
|
|
The credential-helper repair checks every 60 seconds but acts only on a current
|
|
pod with an active image-pull failure and a matching warning less than ten
|
|
minutes old. It records an attempt before restarting a helper and allows at
|
|
most one attempt per helper per 30 minutes, across daemon restarts. Old events
|
|
and warnings for replaced pods do not justify another restart. Check helper
|
|
generation and Ananke's journal when investigating unexpected rollouts:
|
|
|
|
```bash
|
|
kubectl -n jenkins get deployment jenkins-vault-sync
|
|
ssh titan-db sudo journalctl -u ananke --since '10 minutes ago' --no-pager
|
|
ssh titan-db systemctl status ananke-update.timer
|
|
```
|
|
|
|
The native updater now preserves the running binary when its quality gate
|
|
fails. A binary-only maintenance install also preserves host/NUT configuration
|
|
and retains the prior binary for rollback. Read `scripts/install.sh` in the
|
|
Ananke repository before updating; `--binary-only --skip-deps` is the targeted
|
|
path, while the normal updater also applies host configuration templates.
|
|
|
|
Ariadne's `ARIADNE_HERMES_HUNG_BUILD_MINUTES` setting is in
|
|
`services/maintenance/apps/ariadne-deployment.yaml`. It is currently zero, which
|
|
disables elapsed-time-only cancellation. Its retained event confirmed that the
|
|
previous 120-minute setting canceled an active Metis image build. Other bounded
|
|
diagnosis/repair remains enabled. Do not re-enable cancellation until it can
|
|
distinguish a stalled job from a long-running job making useful progress.
|
|
|
|
Jenkins's Kubernetes cloud cap is currently one in
|
|
`services/jenkins/configmap-jcasc.yaml`. The active inline templates in Ariadne,
|
|
Atlasbot, Ananke, Metis, Pegasus,
|
|
Soteria, bstein-dev-home, IaC and Data Prepper now require Pi workers and exclude
|
|
quarantined and storage workers. Inspect actual agent placement rather than
|
|
assuming the shared default controls a custom template.
|
|
A Pending agent consumes the cap; inspect FailedScheduling events and resource
|
|
requests before deciding that Jenkins is hung. Do not increase concurrency to
|
|
solve an unschedulable agent. The tracked JCasC revision triggers a controller
|
|
rollout, so quiet the queue and arrange active-build completion before editing.
|
|
|
|
Custom pipelines explicitly use `workspaceVolume dynamicPVC(...)` with the
|
|
`ci-scratch` StorageClass in `infrastructure/core/storageclass-ci-scratch.yaml`.
|
|
Workspaces and compiler/package caches use that volume. Docker-in-Docker stores
|
|
its layers in the same volume under `.cache/docker`, outside the checkout.
|
|
Ordinary workspaces request 20 GiB; Docker builds request 40 GiB. Builds remain
|
|
serialized. This moves heavy scratch I/O off the Pi runtime USB drive. Soteria
|
|
also compiles its UI and binary on scratch before Kaniko assembles the small
|
|
`Dockerfile.runtime` image. Compiling inside Kaniko would write toolchains and
|
|
compiler intermediates into the container runtime again, defeating that isolation.
|
|
Data Prepper now uses the existing Docker-in-Docker path with a 40 GiB scratch
|
|
PVC for image layers. A scratch workspace alone does not relocate Kaniko
|
|
extraction out of a container root filesystem. The Docker API binds only to pod
|
|
loopback; registry passwords go to stdin and the temporary client config is
|
|
removed afterward. ARM64 output remains explicit.
|
|
Keep registry-credential umask changes local to credential creation. Soteria
|
|
explicitly sets the executable mode and nonroot ownership before image assembly;
|
|
a successful compile alone does not prove the deployed user can execute it.
|
|
|
|
`ci-scratch` is disposable, single-replica Longhorn storage on the existing
|
|
`astreae` disks. Jenkins owns its PVC through the agent pod; deleting that pod
|
|
removes its scratch PVC and volume. Never use this class for application data,
|
|
credentials, recovery copies or retained artifacts. Test artifacts are still
|
|
archived to the Jenkins controller with their existing 30-day policy. A lost
|
|
scratch volume requires rerunning the build. Soteria excludes the ci-scratch
|
|
class from application backups alongside local-path; preserve Hermes exclusions.
|
|
Check orphaned PVCs before assuming cleanup occurred; do not delete unrelated
|
|
retained volumes.
|
|
|
|
Node quarantine is declared in `infrastructure/core/node-maintenance.yaml` for
|
|
Titan-04/05/06/08/11/12/14/18. Offline nodes remain quarantined after reconnecting
|
|
until their physical and loaded-operation checks pass.
|
|
Those Node resources explicitly preserve hardware/worker/storage labels and use
|
|
Flux SSA Merge to retain other runtime fields. A cordon blocks ordinary new
|
|
placement; it must not repeatedly remove existing storage DaemonSets. After a
|
|
quarantine change, verify manager pod UIDs stay unchanged across reconciliation.
|
|
Keep prune protection on Node declarations; repair, uncordon in Git and verify
|
|
before considering their removal.
|
|
|
|
Keep one-off Flux Jobs completed without a TTL, or remove them from active
|
|
resources after success. A TTL deletes the object and Flux then creates it
|
|
again. The archived Titan-24 emergency root sweep is deliberately inactive;
|
|
its zero-percent thresholds are not a normal maintenance schedule. The Metis
|
|
SSH public-key bootstrap now retains its completion.
|
|
|
|
Kubelet owns container-image garbage collection. All 19 reachable nodes were
|
|
verified with 85% high / 80% low thresholds. The legacy-named node-image-sweeper
|
|
no longer runs `crictl rmi --prune` or removes files from containerd/image-import
|
|
directories. Its remaining role is bounded host-log/package-cache maintenance.
|
|
The UID-aware Python helper preserves active pod logs in both `/var/log/pods`
|
|
and Armbian `/var/log.hdd/pods`, skips unavailable inventories, and provides
|
|
`--dry-run`. Native kubelet rotation controls active container logs.
|
|
|
|
The node-image-sweeper DaemonSet uses a startup probe that passes after its first
|
|
cleanup, then performs no continuous probe process. Its normal rollout permits
|
|
one unavailable helper. Offline terminating pods can consume that budget;
|
|
inspect those before raising it. This does not change application availability.
|
|
|
|
## Ingress and search capacity
|
|
|
|
Traefik runs two replicas on separate NVMe control-plane nodes. Its internal
|
|
`/ping` readiness check gates traffic. The public VIP remains `192.168.22.9`;
|
|
Gitea SSH and Wolf share it using their existing ports. MetalLB's `ingress-pool`
|
|
and `ingress-adv` restrict its announcer to the control planes. The private
|
|
`192.168.22.50` pool does the same and retains `externalTrafficPolicy: Local`,
|
|
source restrictions, TLS and service authentication. The communications pool
|
|
retains `.4` through `.6`; its Local-policy services need a speaker on their
|
|
actual endpoint node. Do not restrict all pools to ingress nodes indiscriminately.
|
|
|
|
MetalLB uses native L2 mode, without unused FRR/BGP sidecars. There are no BGP
|
|
peers or advertisements. Titan-05/06 are excluded from the speaker DaemonSet
|
|
until their recovery checks pass. Remove that explicit exclusion in
|
|
`infrastructure/metallb/helmrelease.yaml` when restoring them. Do not enable BGP
|
|
features later without revisiting this choice. The speaker requests 128 MiB and
|
|
uses five-second probes. Native L2 still has one announcing node per VIP.
|
|
|
|
OpenSearch's existing PVC is now 1280 GiB. Its expansion reduced measured usage
|
|
from 94% to 76%, allowing the blocked maintenance policy updates to resume. The
|
|
existing scheduled tuner now reports failing operation/path/status without
|
|
index contents. Its next scheduled run passed. No retention period or disk
|
|
watermark was relaxed. Inspect free space and ingestion growth periodically;
|
|
previous growth while writes were blocked is not a useful capacity forecast.
|
|
The current volume still has one Longhorn replica; it is not storage HA.
|
|
The scheduled tuner uses the native maintenance-batch priority class with
|
|
preemptionPolicy Never. It prefers Pi 5 workers but permits Pi 4 workers when
|
|
capacity is unavailable. Routine maintenance should wait rather than evict
|
|
running work; check scheduling deadlines instead of raising its priority.
|
|
**Do not shrink the PVC or blindly revert its expansion commit.**
|
|
|
|
## Five-minute check
|
|
|
|
Run from the management host with the existing administrator configuration:
|
|
|
|
```bash
|
|
kubectl get nodes -o wide
|
|
flux get kustomizations -A
|
|
flux get helmreleases -A
|
|
kubectl get deployments,statefulsets -A
|
|
kubectl get pods -A --field-selector=status.phase=Pending
|
|
kubectl -n longhorn-system get volumes.longhorn.io
|
|
```
|
|
|
|
Then check the affected service's health endpoint or UI. `Running` does not mean
|
|
ready; `Ready` does not prove useful application behavior. A detached Longhorn
|
|
volume may be intentional if its workload is parked. Distinguish that from an
|
|
attached volume that is faulted or an application waiting for its disk.
|
|
|
|
Check the control-plane backups separately:
|
|
|
|
```bash
|
|
ssh titan-db sudo systemctl status atlas-k3s-backup.timer
|
|
ssh titan-0b sudo systemctl status atlas-k3s-replica.timer
|
|
ssh titan-db sudo cat /var/backups/atlas-k3s/latest/COMPLETE
|
|
ssh titan-0b sudo cat /var/backups/atlas-k3s-replica/latest/COMPLETE
|
|
```
|
|
|
|
`COMPLETE` contains timestamps, not database contents. The backup and replica
|
|
should normally be less than two hours old. Check the last service result too:
|
|
|
|
```bash
|
|
ssh titan-db sudo systemctl show atlas-k3s-backup.service -p Result -p ExecMainStatus
|
|
ssh titan-0b sudo systemctl show atlas-k3s-replica.service -p Result -p ExecMainStatus
|
|
```
|
|
|
|
Application database backups run separately each day at 08:15 UTC plus up to
|
|
five minutes of jitter. Check their timer and completion timestamp too:
|
|
|
|
```bash
|
|
ssh titan-0b sudo systemctl status atlas-application-postgres-backup.timer
|
|
ssh titan-0b sudo cat /var/backups/atlas-postgres/latest/COMPLETE
|
|
```
|
|
|
|
Their freshness alert fires after 36 hours. See the host backup notes before
|
|
running the isolated restore verification or changing retention.
|
|
|
|
## Check a settling window
|
|
|
|
Use the read-only helper after the rollout finishes, then compare 10-15 minutes
|
|
later. It uses ordinary kubectl access; it does not repair, restart or delete
|
|
anything and has no AI or provider dependency.
|
|
|
|
```bash
|
|
python3 scripts/ops/cluster_settle_check.py --output /tmp/atlas-before.json
|
|
# Wait 10-15 minutes with normal workload activity.
|
|
python3 scripts/ops/cluster_settle_check.py --output /tmp/atlas-after.json --previous /tmp/atlas-before.json
|
|
```
|
|
|
|
Review restart increases, replaced service pods, unready workloads, requested
|
|
but unavailable storage and new warning evidence. Expected short-lived CI agent
|
|
turnover is different from an application controller repeatedly replacing pods.
|
|
Known offline nodes are reported separately, not hidden. Pods already marked
|
|
for deletion are not considered available merely because Ready remains true.
|
|
Scheduler Preempted events are included even when their event type is Normal. A newly observed event
|
|
has an unknown counter delta because its lifetime count may predate the window.
|
|
The helper deliberately omits Secret values, environments, log bodies and event
|
|
messages. Pair it with the service health endpoints and node I/O/memory metrics;
|
|
Kubernetes can report a node Ready while its runtime disk is stalled. The helper
|
|
checks Kustomizations and HelmReleases, not image-release policies or application
|
|
incident history; inspect those separately before claiming all automation green.
|
|
|
|
Current placement is deliberately conservative: Gitea and ClamAV use Titan-22's
|
|
CPU/RAM as last-resort general capacity; they request no GPU. Nextcloud stays on
|
|
a Pi worker. Titan-08 is also quarantined after a build saturated its runtime USB
|
|
drive. Keep Jenkins at one agent until ordinary builds and application activity
|
|
coexist without recurring pressure. See the dated cycle record for measurements.
|
|
|
|
Nextcloud startup now preserves installed apps and their external-app keys.
|
|
App upgrades are separate maintenance, rather than downloads/deletions on every
|
|
pod replacement. Its single-writer volumes require maxSurge 0 / maxUnavailable 1;
|
|
a replacement therefore has a short planned interruption. ClamAV readiness uses
|
|
its native PING/PONG response; a running process alone does not open the endpoint.
|
|
|
|
## Find the cause before choosing a repair
|
|
|
|
| Symptom | First check | Usual next action |
|
|
| --- | --- | --- |
|
|
| Node NotReady | Power/network, kubelet status, disk space, pressure | Repair/quarantine the node; do not repeatedly restart every application |
|
|
| Pod Pending / FailedScheduling | `kubectl describe pod` scheduling events | Correct resource requests or eligible capacity; free cluster-wide RAM does not guarantee a compatible destination |
|
|
| ContainerCreating / FailedMount | PVC, Longhorn volume, attachment and engine events | Preserve data; repair the attachment or use a verified healthy node. Do not delete the PVC |
|
|
| CrashLoopBackOff / OOMKilled | Previous container logs and last termination reason | Fix the application or its resource envelope; lifetime restart counts are not a reason for repeated eviction |
|
|
| Ingress 502/503 | Service endpoints and backend readiness | Restore the backend or its dependency before changing DNS/TLS |
|
|
| Flux Ready but missing Helm workload | Helm manifest and drift correction | Check that drift detection is enabled and suitable for that release |
|
|
| All cluster API operations fail | titan-db PostgreSQL and control-plane nodes | The datastore is external PostgreSQL; an etcd snapshot procedure will not restore it |
|
|
| CI waiting while applications work | Jenkins queue, agent capacity and node I/O | Let work queue; adding agents to saturated Pi disks makes both CI and services worse |
|
|
|
|
Useful scoped commands:
|
|
|
|
```bash
|
|
kubectl -n NAMESPACE describe pod POD
|
|
kubectl -n NAMESPACE get events --sort-by=.lastTimestamp
|
|
kubectl -n NAMESPACE get endpointslices
|
|
kubectl -n NAMESPACE logs POD -c CONTAINER --previous --tail=80
|
|
```
|
|
|
|
Application logs can contain private data. Inspect them locally and redact before
|
|
sharing. Never paste Secrets, tokens, source records or inference bodies into a
|
|
ticket or assistant conversation.
|
|
|
|
## Make a normal change
|
|
|
|
1. Edit the service's tracked manifest. Placement, requests, limits and probes
|
|
belong together in that workload's configuration.
|
|
2. Render and validate the affected folder; inspect the Flux preview.
|
|
3. Commit and push a small change to the tracked branch.
|
|
4. Reconcile that service and verify its application behavior and storage.
|
|
|
|
For example:
|
|
|
|
```bash
|
|
kubectl kustomize services/monitoring > /tmp/monitoring-render.yaml
|
|
kubectl apply --dry-run=client -k services/monitoring
|
|
flux diff kustomization monitoring --path services/monitoring
|
|
git diff --check
|
|
git diff
|
|
```
|
|
|
|
After committing and pushing the reviewed files:
|
|
|
|
```bash
|
|
flux reconcile kustomization monitoring -n flux-system --with-source
|
|
kubectl -n monitoring get deployments,statefulsets,pods
|
|
```
|
|
|
|
A Flux diff exits nonzero when it finds changes; read the output to distinguish
|
|
that from a validation failure. Kustomize renders a HelmRelease, not the chart's
|
|
workloads. For chart values, also inspect the chart's rendered Deployment or
|
|
StatefulSet. OpenSearch, for example, uses `values.nodeAffinity`.
|
|
|
|
Rollback a configuration change by reverting its focused commit and reconciling
|
|
the same service. Do not blindly revert storage migrations or database changes;
|
|
those require their recovery procedure. Do not delete data to clear a red status.
|
|
|
|
## Backups and recovery
|
|
|
|
The external datastore's [host backup notes](../infrastructure/host-backup/NOTES.md)
|
|
give the actual paths, schedule, credential scope, isolated restore check and
|
|
rollback. Backups contain sensitive data and remain root-only on approved LAN
|
|
hosts. They are not encrypted at rest or off-site disaster recovery copies.
|
|
|
|
Soteria is responsible for eligible application backups, not for making the
|
|
control plane boot. Live Longhorn snapshots and database-native dumps offer
|
|
different consistency guarantees. A recent backup indicator is not a restore
|
|
test. Keep database restore checks and service recovery drills explicit.
|
|
|
|
For cloud backups, inspect completed backup timestamps, not just the backup
|
|
target's Available status or a successful bucket listing. The October 3 audit
|
|
found a readable Backblaze bucket rejecting uploads with `storage cap exceeded`.
|
|
Check Backblaze Caps & Alerts and agree the storage budget before changing it.
|
|
Do not delete old recovery points merely to make uploads work. Distinguish
|
|
visible-object bytes from billed storage, including retained versions. The
|
|
current diagnosis and counts are in the implementation record.
|
|
|
|
The Hermes namespace is excluded from cloud backup policy. Its workspaces can
|
|
contain material restricted to local infrastructure. Do not remove that exclusion
|
|
to make a coverage dashboard greener.
|
|
|
|
## A one-day familiarization exercise
|
|
|
|
* Morning: follow one application from its Flux definition to its workload,
|
|
storage, service and ingress. Compare those manifests with live objects.
|
|
* Before lunch: inspect one scheduling incident and one storage incident using
|
|
the table above. Identify the responsible component without changing anything.
|
|
* Afternoon: run the isolated datastore restore verification, inspect backup
|
|
timestamps, then make and revert a harmless configuration change through Flux.
|
|
* Finish by finding each remaining issue in the implementation record and the
|
|
corresponding owner/runbook. Confirm access to Git, the management host, the
|
|
hosts' SSH path and protected recovery material without relying on SSO alone.
|
|
|
|
This provides a practical route to ownership; it cannot promise that every
|
|
hardware, storage or database disaster will be solvable in one day.
|
|
|
|
## Physical checks and return to service
|
|
|
|
For an offline Pi, record the board identity, supply model and its 5 V rating,
|
|
power/activity LEDs, HDMI boot result, Ethernet link and current router address.
|
|
Preserve the original SD/USB media. Test with separate known-good spare boot media
|
|
before concluding that a board has failed; do not clone a live node identity by
|
|
moving another cluster node's boot card into it.
|
|
|
|
Pi 5 power must be evaluated at its 5 V output mode, not a charger's headline
|
|
wattage. A controlled supply-and-cable substitution isolates the power path.
|
|
Record any inline adapters and USB loads. Do not unplug runtime USB storage on
|
|
a running node. Titan-11 carries important services; arrange workload movement
|
|
before disconnecting it. See the official Raspberry Pi power documentation for
|
|
5 V / 5 A capability and cable-loss requirements.
|
|
|
|
A repaired node returns only after stable power, clean storage/kernel logs,
|
|
reliable networking and representative-load testing. Observe it before restoring
|
|
normal placement. Titan-14/18 quarantine is explicitly tracked in
|
|
`infrastructure/core/node-maintenance.yaml`; `prune: disabled` protects the Node
|
|
objects from deletion when that maintenance declaration is eventually removed.
|
|
Clear `spec.unschedulable` through Git after validation, then reconcile core.
|
|
|
|
Container runtime media and Longhorn data disks are separate. A Pi with terabytes
|
|
of healthy application storage can still stall because containerd, image unpacking
|
|
and logs live on a small USB flash device. Include the runtime disk in hardware
|
|
repair and capacity checks.
|