94 lines
5.0 KiB
Markdown
94 lines
5.0 KiB
Markdown
# ananke
|
|
|
|
Ananke gets Atlas back on its feet after power trouble.
|
|
|
|
It runs on both tethys (in cluster - titan-24) and titan-db (out of cluster), outside Kubernetes as the host level, because some failures start before the cluster is healthy enough to fix itself. Its job is to bring nodes, Flux, core workloads, ingresses, and service checks back into a known-good state.
|
|
|
|
It aspires to be boring software: do the checks, repair the known deadlocks, and stomp loudly when it exhausts is remedy library.
|
|
|
|
## How it works
|
|
|
|
Ananke walks the cluster through startup or shutdown gates:
|
|
|
|
- confirm the expected nodes and SSH access
|
|
- check that Flux is looking at the right repo and branch
|
|
- wait for required Flux kustomizations and namespaces
|
|
- repair known startup traps, including Harbor/Gitea/Flux coupling
|
|
- run ingress, service, endpoint, and soak checks before calling startup done
|
|
|
|
Recovery cordons are given short 1hr leases. If Ananke cordons a node to repair something, it must either clear the cordon within the configured window or mark the node for manual action.
|
|
|
|
The following are notes for future Brad.
|
|
|
|
## Environmental Dependencies
|
|
|
|
Ananke should be one of the first things working. It does not need Harbor, Gitea, Longhorn, Grafana, or the apps to be healthy before it starts; their mess is what Ananke was born to sort out.
|
|
|
|
Ananke needs:
|
|
|
|
- an Ananke host that came up on its own: usually `titan-db` (for the arm64 ups) and `tethys`/`titan-24` (for amd64 ups)
|
|
- `/etc/ananke/ananke.yaml`, the Ananke SSH key, and enough host ssh config to reach nodes on the Atlas SSH port
|
|
- Kubernetes API access once the control plane is answering; before that it can only do host-side checks
|
|
- Flux CRDs/controllers and the `titan-iac` source once the API is up, because most startup gates are Flux-shaped
|
|
- basic node hygiene that Ananke cannot fake forever: SSH, sudo for managed repairs, sane clocks, and Longhorn host packages like `cryptsetup`, `open-iscsi`, `dmsetup`, and `nfs-common`
|
|
- NUT/UPS access if this is making real shutdown decisions instead of just doing startup recovery
|
|
|
|
If this is a total bring-up, start Ananke after the host boots and before waiting on applications. If Ananke is not running, Atlas is missing the thing that knows the order of operations.
|
|
|
|
## Daily commands
|
|
|
|
```bash
|
|
sudo /usr/local/bin/ananke status --config /etc/ananke/ananke.yaml
|
|
sudo /usr/local/bin/ananke startup --config /etc/ananke/ananke.yaml --execute --force-flux-branch main
|
|
sudo /usr/local/bin/ananke shutdown --config /etc/ananke/ananke.yaml --execute --reason graceful-maintenance --mode cluster-only
|
|
```
|
|
|
|
Host files:
|
|
|
|
- `/var/lib/ananke/startup-progress.json`
|
|
- `/var/lib/ananke/last-startup-report.json`
|
|
- `/var/lib/ananke/last-shutdown-report.json`
|
|
- `/var/log/ananke/update.log`
|
|
|
|
## Terraform-Generated Inventory
|
|
|
|
Ananke can load a generated inventory fragment from `/etc/ananke/ananke.inventory.yaml` alongside `/etc/ananke/ananke.yaml`. The installer also accepts `ANANKE_INVENTORY_FRAGMENT=/path/to/ananke.inventory.yaml` and copies it into place.
|
|
|
|
That fragment is for non-secret declarative inputs only: node hosts, managed nodes, control planes, workers, required node labels, and ignored unavailable nodes. When Terraform owns node labels, the fragment should set `startup.required_node_labels_mode: validate`; Ananke will report label drift but will not apply `kubectl label` during normal startup or post-start auto-heal.
|
|
|
|
Keep recovery behavior in Ananke: UPS shutdown, SSH repair, k3s restart/reboot, Longhorn/runtime recovery, cordon/uncordon, pod recycling, and Flux suspend/resume remain imperative Ananke actions.
|
|
|
|
## Development
|
|
|
|
Local testing check before installing:
|
|
|
|
```bash
|
|
./scripts/quality_gate.sh
|
|
```
|
|
|
|
Emergency installs can bypass the gate with `ANANKE_ENFORCE_QUALITY_GATE=0` - try to avoid this. You should be treating failures as an instructive opportunity to improve Ananke.
|
|
|
|
## Credential recovery and targeted updates
|
|
|
|
Registry credential recovery requires a warning observed within ten minutes for
|
|
an existing pod UID whose container or init container is still in ErrImagePull
|
|
or ImagePullBackOff. Completed and deleting pods are ignored. A helper restart
|
|
is reserved in the existing run history before execution and may occur at most
|
|
once per helper in 30 minutes, including failed rollouts. History must remain
|
|
writable; otherwise this repair fails closed. It is a single-coordinator control,
|
|
not a cross-host distributed lock. Do not run two coordinators against one cluster.
|
|
|
|
A targeted daemon fix can use the existing installer without rewriting host
|
|
configuration, UPS settings, or systemd/bootstrap units:
|
|
|
|
```bash
|
|
sudo ./scripts/install.sh --binary-only --skip-deps
|
|
```
|
|
|
|
The quality gate still runs. The previous binary is retained at
|
|
`/usr/local/lib/ananke/rollback/ananke.previous`; a failed service start restores
|
|
it. Verify service health and journals after installation. Automated self-updates
|
|
now default to keeping the current binary when their quality gate fails. The
|
|
service unit and updater script both set that default; existing hosts need those
|
|
two files updated to adopt it.
|