101 lines
4.9 KiB
Markdown
101 lines
4.9 KiB
Markdown
# External control-plane datastore recovery
|
|
|
|
K3s uses PostgreSQL 16 on `titan-db` (192.168.22.10), not embedded etcd.
|
|
These host-side systemd jobs keep recovery independent of the cluster, Vault,
|
|
SSO, Harbor and the homegrown controllers. Kubernetes manifests remain Flux-owned.
|
|
|
|
## What runs where
|
|
|
|
| Host | Timer | Protected directory | Purpose |
|
|
| --- | --- | --- | --- |
|
|
| titan-db | atlas-k3s-backup.timer, hourly at :00 plus 0-3 minutes | /var/backups/atlas-k3s | Consistent custom-format database dump, PostgreSQL globals, server token, checksums |
|
|
| titan-0b | atlas-k3s-replica.timer, hourly at :20 plus 0-3 minutes | /var/backups/atlas-k3s-replica | Second LAN disk with the complete bundle |
|
|
|
|
Both keep completed snapshots for seven days, publishing `latest` only after
|
|
validation. Partial directories are removed on ordinary failure. Old snapshots
|
|
are pruned only after a successful new snapshot. A backup older than 90 minutes
|
|
cannot be exported as fresh. Both jobs refuse to start below 5 GiB free space.
|
|
|
|
The database snapshot is consistent at dump start; PostgreSQL globals are a
|
|
separate dump. Avoid concurrent role or K3s token rotation during backup.
|
|
The replica compares the saved token with the local live server token and fails
|
|
if they differ. After an intentional token rotation, securely update the root-only
|
|
`/etc/atlas-k3s-backup/server-token` on titan-db from a current control plane.
|
|
|
|
Files and directories are root-only. SSH encrypts transfer. These copies are
|
|
**not encrypted at rest** and are on the same LAN/site; they are not an off-site
|
|
or fire/theft recovery guarantee. Do not upload them to a general cloud backup:
|
|
the datastore contains credentials and may contain restricted source-derived data.
|
|
|
|
## Installation and credentials
|
|
|
|
Install `scripts/ops/k3s_datastore_backup.sh` as
|
|
`/usr/local/sbin/atlas-k3s-backup` on titan-db, mode 0755. Install
|
|
`k3s_backup_replica.sh` as `/usr/local/sbin/atlas-k3s-replica` on titan-0b.
|
|
Install the corresponding service/timer files from this directory in
|
|
`/etc/systemd/system/`, then run `systemctl daemon-reload`.
|
|
|
|
The dedicated private key exists only on titan-0b at
|
|
`/root/.ssh/atlas_k3s_backup` (0600). Its pinned host key file is
|
|
`/root/.ssh/atlas_k3s_backup_known_hosts` for `[192.168.22.10]:2277`.
|
|
Verify replacement host keys out of band.
|
|
The public key in titan-db's atlas authorized_keys has `restrict`, a source
|
|
restriction to 192.168.22.12, and the forced command
|
|
`sudo -n /usr/local/sbin/atlas-k3s-backup export`. It cannot run an arbitrary command
|
|
or forward ports. Existing administrative keys remain separate.
|
|
No keys, tokens, dumps or PostgreSQL globals belong in this repository.
|
|
|
|
## Normal operation
|
|
|
|
On the appropriate host (sudo required):
|
|
|
|
```bash
|
|
systemctl status atlas-k3s-backup.timer atlas-k3s-backup.service
|
|
systemctl status atlas-k3s-replica.timer atlas-k3s-replica.service
|
|
systemctl list-timers 'atlas-k3s-*'
|
|
```
|
|
|
|
Only the matching pair exists on each host. A completed oneshot service is
|
|
normally inactive with `Result=success`; the timer must remain active.
|
|
Inspect `latest/COMPLETE`, verify `SHA256SUMS`, and check timestamps, not just
|
|
whether the directory exists. Routine journal output contains status and byte
|
|
counts. Restricted `last-error.txt` is for local diagnosis; do not paste it into
|
|
chat or public logs without reviewing it.
|
|
|
|
Run a new backup and then replicate it:
|
|
|
|
```bash
|
|
ssh titan-db sudo systemctl start atlas-k3s-backup.service
|
|
ssh titan-0b sudo systemctl start atlas-k3s-replica.service
|
|
```
|
|
|
|
## Restore verification
|
|
|
|
Install `scripts/ops/k3s_backup_verify.sh` as `/usr/local/sbin/atlas-k3s-verify`
|
|
on a host with PostgreSQL 16 binaries and a postgres system user. It restores
|
|
into a disposable cluster, using a private Unix socket and no TCP listener.
|
|
It never connects to or replaces the running datastore.
|
|
|
|
```bash
|
|
sudo /usr/local/sbin/atlas-k3s-verify /var/backups/atlas-k3s/latest
|
|
```
|
|
|
|
Success records `RESTORE_CHECK` with elapsed time and a nonzero `kine` row count.
|
|
This proves archive readability and database/schema restoration. It deliberately
|
|
does not replay global roles/ACLs or start K3s, so it is not a full disaster drill.
|
|
Temporary database files are removed after the check.
|
|
|
|
For actual disaster recovery: fence the old datastore and stop K3s writers first;
|
|
preserve any surviving data; provision compatible PostgreSQL; review and restore
|
|
globals and the k3s dump; restore the matching server token; verify connectivity
|
|
and start one server before the others. Do not overwrite a running database.
|
|
See the [K3s recovery requirements](https://docs.k3s.io/datastore/backup-restore)
|
|
and [PostgreSQL restore reference](https://www.postgresql.org/docs/16/app-pgrestore.html).
|
|
|
|
## Rollback
|
|
|
|
Disable the matching timer with `systemctl disable --now atlas-k3s-backup.timer`
|
|
or `atlas-k3s-replica.timer`. Do not delete existing recovery bundles. If retiring
|
|
replication, remove only the `atlas-k3s-backup-replica` authorized-key line on
|
|
titan-db. Disabling these jobs does not restart PostgreSQL or Kubernetes.
|