7.6 KiB

External control-plane datastore recovery

K3s uses PostgreSQL 16 on titan-db (192.168.22.10), not embedded etcd. These host-side systemd jobs keep recovery independent of the cluster, Vault, SSO, Harbor and the homegrown controllers. Kubernetes manifests remain Flux-owned.

What runs where

Host Timer Protected directory Purpose
titan-db atlas-k3s-backup.timer, hourly at :00 plus 0-3 minutes /var/backups/atlas-k3s Consistent custom-format database dump, PostgreSQL globals, server token, checksums
titan-0b atlas-k3s-replica.timer, hourly at :20 plus 0-3 minutes /var/backups/atlas-k3s-replica Second LAN disk with the complete bundle

Both keep completed snapshots for seven days, publishing latest only after validation. Partial directories are removed on ordinary failure. Old snapshots are pruned only after a successful new snapshot. A backup older than 90 minutes cannot be exported as fresh. Both jobs refuse to start below 5 GiB free space.

The database snapshot is consistent at dump start; PostgreSQL globals are a separate dump. Avoid concurrent role or K3s token rotation during backup. The replica compares the saved token with the local live server token and fails if they differ. After an intentional token rotation, securely update the root-only /etc/atlas-k3s-backup/server-token on titan-db from a current control plane.

Files and directories are root-only. SSH encrypts transfer. These copies are not encrypted at rest and are on the same LAN/site; they are not an off-site or fire/theft recovery guarantee. Do not upload them to a general cloud backup: the datastore contains credentials and may contain restricted source-derived data.

Installation and credentials

Install scripts/ops/k3s_datastore_backup.sh as /usr/local/sbin/atlas-k3s-backup on titan-db, mode 0755. Install k3s_backup_replica.sh as /usr/local/sbin/atlas-k3s-replica on titan-0b. Install the corresponding service/timer files from this directory in /etc/systemd/system/, then run systemctl daemon-reload.

The dedicated private key exists only on titan-0b at /root/.ssh/atlas_k3s_backup (0600). Its pinned host key file is /root/.ssh/atlas_k3s_backup_known_hosts for [192.168.22.10]:2277. Verify replacement host keys out of band. The public key in titan-db's atlas authorized_keys has restrict, a source restriction to 192.168.22.12, and the forced command sudo -n /usr/local/sbin/atlas-k3s-backup export. It cannot run an arbitrary command or forward ports. Existing administrative keys remain separate. No keys, tokens, dumps or PostgreSQL globals belong in this repository.

Normal operation

On the appropriate host (sudo required):

systemctl status atlas-k3s-backup.timer atlas-k3s-backup.service
systemctl status atlas-k3s-replica.timer atlas-k3s-replica.service
systemctl list-timers 'atlas-k3s-*'

Only the matching pair exists on each host. A completed oneshot service is normally inactive with Result=success; the timer must remain active. Inspect latest/COMPLETE, verify SHA256SUMS, and check timestamps, not just whether the directory exists. Routine journal output contains status and byte counts. Restricted last-error.txt is for local diagnosis; do not paste it into chat or public logs without reviewing it.

The scripts atomically publish a non-sensitive success timestamp through the existing node exporter textfile collector. Titan-db uses the versioned node-exporter-textfile.conf systemd drop-in; titan-0b uses the cluster's node-exporter DaemonSet. The metric is atlas_k3s_backup_last_success_timestamp_seconds, labelled copy="datastore" or copy="lan_replica". These files contain no paths, keys or database records.

Run a new backup and then replicate it:

ssh titan-db sudo systemctl start atlas-k3s-backup.service
ssh titan-0b sudo systemctl start atlas-k3s-replica.service

Restore verification

Install scripts/ops/k3s_backup_verify.sh as /usr/local/sbin/atlas-k3s-verify on a host with PostgreSQL 16 binaries and a postgres system user. It restores into a disposable cluster, using a private Unix socket and no TCP listener. It never connects to or replaces the running datastore.

sudo /usr/local/sbin/atlas-k3s-verify /var/backups/atlas-k3s/latest

Success records RESTORE_CHECK with elapsed time and a nonzero kine row count. This proves archive readability and database/schema restoration. It deliberately does not replay global roles/ACLs or start K3s, so it is not a full disaster drill. Temporary database files are removed after the check.

For actual disaster recovery: fence the old datastore and stop K3s writers first; preserve any surviving data; provision compatible PostgreSQL; review and restore globals and the k3s dump; restore the matching server token; verify connectivity and start one server before the others. Do not overwrite a running database. See the K3s recovery requirements and PostgreSQL restore reference.

Rollback

Disable the matching timer with systemctl disable --now atlas-k3s-backup.timer or atlas-k3s-replica.timer. Do not delete existing recovery bundles. If retiring replication, remove only the atlas-k3s-backup-replica authorized-key line on titan-db. Disabling these jobs does not restart PostgreSQL or Kubernetes.

Application databases

Titan-0b also runs atlas-application-postgres-backup.timer daily at 08:15 UTC plus up to five minutes. Install PostgreSQL client 16 and the versioned postgres_application_backup.sh as /usr/local/sbin/atlas-application-postgres-backup. It uses the host's existing root K3s access to discover the database pod and read its mounted password, then connects with native PostgreSQL over the cluster LAN. The temporary password file is mode 0600 in /run and removed on exit. This is an application recovery copy; it cannot run while the Kubernetes API is down.

Recovery sets are root-only under /var/backups/atlas-postgres/run-*; latest changes only after all databases finish and checksum verification passes. databases.txt maps its ordered entries to db-1.dump, db-2.dump, and so on. Each database has its own consistent snapshot; the set is not one transaction across databases. Globals include sensitive role information. No contents are sent outside the approved LAN. Old run directories expire after seven days, but only after a successful new backup. A run is bounded to two hours.

ssh titan-0b sudo systemctl status atlas-application-postgres-backup.timer
ssh titan-0b sudo cat /var/backups/atlas-postgres/latest/COMPLETE
ssh titan-0b sudo systemctl start atlas-application-postgres-backup.service

errors.log remains inside the protected bundle. The only exported metric is atlas_application_postgres_backup_last_success_timestamp_seconds.

postgres_dump_verify.sh can verify one trusted dump on a host with PostgreSQL 16 server binaries and a postgres user. It starts a disposable instance without TCP, restores schema and data without role/ACL replay, counts application tables, and removes the instance. It does not overwrite a live application database. On titan-db it is installed as /usr/local/sbin/atlas-application-verify. The initial Gitea restore completed in six seconds with 111 application tables. This is representative restore evidence, not proof for every application.

To stop this schedule, disable atlas-application-postgres-backup.timer on Titan-0b. Retain recovery sets. The earlier incomplete pre-resources-* directory is not a valid complete backup; use only a set with COMPLETE and valid checksums.