147 lines
7.6 KiB
Markdown
147 lines
7.6 KiB
Markdown
# External control-plane datastore recovery
|
|
|
|
K3s uses PostgreSQL 16 on `titan-db` (192.168.22.10), not embedded etcd.
|
|
These host-side systemd jobs keep recovery independent of the cluster, Vault,
|
|
SSO, Harbor and the homegrown controllers. Kubernetes manifests remain Flux-owned.
|
|
|
|
## What runs where
|
|
|
|
| Host | Timer | Protected directory | Purpose |
|
|
| --- | --- | --- | --- |
|
|
| titan-db | atlas-k3s-backup.timer, hourly at :00 plus 0-3 minutes | /var/backups/atlas-k3s | Consistent custom-format database dump, PostgreSQL globals, server token, checksums |
|
|
| titan-0b | atlas-k3s-replica.timer, hourly at :20 plus 0-3 minutes | /var/backups/atlas-k3s-replica | Second LAN disk with the complete bundle |
|
|
|
|
Both keep completed snapshots for seven days, publishing `latest` only after
|
|
validation. Partial directories are removed on ordinary failure. Old snapshots
|
|
are pruned only after a successful new snapshot. A backup older than 90 minutes
|
|
cannot be exported as fresh. Both jobs refuse to start below 5 GiB free space.
|
|
|
|
The database snapshot is consistent at dump start; PostgreSQL globals are a
|
|
separate dump. Avoid concurrent role or K3s token rotation during backup.
|
|
The replica compares the saved token with the local live server token and fails
|
|
if they differ. After an intentional token rotation, securely update the root-only
|
|
`/etc/atlas-k3s-backup/server-token` on titan-db from a current control plane.
|
|
|
|
Files and directories are root-only. SSH encrypts transfer. These copies are
|
|
**not encrypted at rest** and are on the same LAN/site; they are not an off-site
|
|
or fire/theft recovery guarantee. Do not upload them to a general cloud backup:
|
|
the datastore contains credentials and may contain restricted source-derived data.
|
|
|
|
## Installation and credentials
|
|
|
|
Install `scripts/ops/k3s_datastore_backup.sh` as
|
|
`/usr/local/sbin/atlas-k3s-backup` on titan-db, mode 0755. Install
|
|
`k3s_backup_replica.sh` as `/usr/local/sbin/atlas-k3s-replica` on titan-0b.
|
|
Install the corresponding service/timer files from this directory in
|
|
`/etc/systemd/system/`, then run `systemctl daemon-reload`.
|
|
|
|
The dedicated private key exists only on titan-0b at
|
|
`/root/.ssh/atlas_k3s_backup` (0600). Its pinned host key file is
|
|
`/root/.ssh/atlas_k3s_backup_known_hosts` for `[192.168.22.10]:2277`.
|
|
Verify replacement host keys out of band.
|
|
The public key in titan-db's atlas authorized_keys has `restrict`, a source
|
|
restriction to 192.168.22.12, and the forced command
|
|
`sudo -n /usr/local/sbin/atlas-k3s-backup export`. It cannot run an arbitrary command
|
|
or forward ports. Existing administrative keys remain separate.
|
|
No keys, tokens, dumps or PostgreSQL globals belong in this repository.
|
|
|
|
## Normal operation
|
|
|
|
On the appropriate host (sudo required):
|
|
|
|
```bash
|
|
systemctl status atlas-k3s-backup.timer atlas-k3s-backup.service
|
|
systemctl status atlas-k3s-replica.timer atlas-k3s-replica.service
|
|
systemctl list-timers 'atlas-k3s-*'
|
|
```
|
|
|
|
Only the matching pair exists on each host. A completed oneshot service is
|
|
normally inactive with `Result=success`; the timer must remain active.
|
|
Inspect `latest/COMPLETE`, verify `SHA256SUMS`, and check timestamps, not just
|
|
whether the directory exists. Routine journal output contains status and byte
|
|
counts. Restricted `last-error.txt` is for local diagnosis; do not paste it into
|
|
chat or public logs without reviewing it.
|
|
|
|
The scripts atomically publish a non-sensitive success timestamp through the
|
|
existing node exporter textfile collector. Titan-db uses the versioned
|
|
`node-exporter-textfile.conf` systemd drop-in; titan-0b uses the cluster's
|
|
node-exporter DaemonSet. The metric is
|
|
`atlas_k3s_backup_last_success_timestamp_seconds`, labelled `copy="datastore"`
|
|
or `copy="lan_replica"`. These files contain no paths, keys or database records.
|
|
|
|
Run a new backup and then replicate it:
|
|
|
|
```bash
|
|
ssh titan-db sudo systemctl start atlas-k3s-backup.service
|
|
ssh titan-0b sudo systemctl start atlas-k3s-replica.service
|
|
```
|
|
|
|
## Restore verification
|
|
|
|
Install `scripts/ops/k3s_backup_verify.sh` as `/usr/local/sbin/atlas-k3s-verify`
|
|
on a host with PostgreSQL 16 binaries and a postgres system user. It restores
|
|
into a disposable cluster, using a private Unix socket and no TCP listener.
|
|
It never connects to or replaces the running datastore.
|
|
|
|
```bash
|
|
sudo /usr/local/sbin/atlas-k3s-verify /var/backups/atlas-k3s/latest
|
|
```
|
|
|
|
Success records `RESTORE_CHECK` with elapsed time and a nonzero `kine` row count.
|
|
This proves archive readability and database/schema restoration. It deliberately
|
|
does not replay global roles/ACLs or start K3s, so it is not a full disaster drill.
|
|
Temporary database files are removed after the check.
|
|
|
|
For actual disaster recovery: fence the old datastore and stop K3s writers first;
|
|
preserve any surviving data; provision compatible PostgreSQL; review and restore
|
|
globals and the k3s dump; restore the matching server token; verify connectivity
|
|
and start one server before the others. Do not overwrite a running database.
|
|
See the [K3s recovery requirements](https://docs.k3s.io/datastore/backup-restore)
|
|
and [PostgreSQL restore reference](https://www.postgresql.org/docs/16/app-pgrestore.html).
|
|
|
|
## Rollback
|
|
|
|
Disable the matching timer with `systemctl disable --now atlas-k3s-backup.timer`
|
|
or `atlas-k3s-replica.timer`. Do not delete existing recovery bundles. If retiring
|
|
replication, remove only the `atlas-k3s-backup-replica` authorized-key line on
|
|
titan-db. Disabling these jobs does not restart PostgreSQL or Kubernetes.
|
|
|
|
## Application databases
|
|
|
|
Titan-0b also runs `atlas-application-postgres-backup.timer` daily at 08:15 UTC
|
|
plus up to five minutes. Install PostgreSQL client 16 and the versioned
|
|
`postgres_application_backup.sh` as `/usr/local/sbin/atlas-application-postgres-backup`.
|
|
It uses the host's existing root K3s access to discover the database pod and read
|
|
its mounted password, then connects with native PostgreSQL over the cluster LAN.
|
|
The temporary password file is mode 0600 in `/run` and removed on exit. This is
|
|
an application recovery copy; it cannot run while the Kubernetes API is down.
|
|
|
|
Recovery sets are root-only under `/var/backups/atlas-postgres/run-*`; `latest`
|
|
changes only after all databases finish and checksum verification passes.
|
|
`databases.txt` maps its ordered entries to `db-1.dump`, `db-2.dump`, and so on.
|
|
Each database has its own consistent snapshot; the set is not one transaction
|
|
across databases. Globals include sensitive role information. No contents are
|
|
sent outside the approved LAN. Old run directories expire after seven days,
|
|
but only after a successful new backup. A run is bounded to two hours.
|
|
|
|
```bash
|
|
ssh titan-0b sudo systemctl status atlas-application-postgres-backup.timer
|
|
ssh titan-0b sudo cat /var/backups/atlas-postgres/latest/COMPLETE
|
|
ssh titan-0b sudo systemctl start atlas-application-postgres-backup.service
|
|
```
|
|
|
|
`errors.log` remains inside the protected bundle. The only exported metric is
|
|
`atlas_application_postgres_backup_last_success_timestamp_seconds`.
|
|
|
|
`postgres_dump_verify.sh` can verify one trusted dump on a host with PostgreSQL
|
|
16 server binaries and a postgres user. It starts a disposable instance without
|
|
TCP, restores schema and data without role/ACL replay, counts application tables,
|
|
and removes the instance. It does not overwrite a live application database.
|
|
On titan-db it is installed as `/usr/local/sbin/atlas-application-verify`.
|
|
The initial Gitea restore completed in six seconds with 111 application tables.
|
|
This is representative restore evidence, not proof for every application.
|
|
|
|
To stop this schedule, disable `atlas-application-postgres-backup.timer` on
|
|
Titan-0b. Retain recovery sets. The earlier incomplete `pre-resources-*` directory
|
|
is not a valid complete backup; use only a set with `COMPLETE` and valid checksums.
|