metis/docs/node-administration.md

55 lines
3.0 KiB
Markdown

# Node administration after recovery
Use the normal node SSH account and password-protected sudo. Passwords live in
Vault at `kv/atlas/nodes/<hostname>` under `atlas_password` and `root_password`.
The SSH key remains the normal login mechanism; a console root password does
not require allowing root password login over SSH.
Recovery injects a native `metis-node-identity.service`, enabled in the image.
It applies passwords even when cloud-init is absent. Each new injection creates
`/etc/metis/node-identity.pending`, so an old completion marker in the base image
cannot suppress the new identity. Cloud-init and the native unit serialize on
one lock. Failed credential setup remains a failed unit with the pending marker
intact; the Longhorn first-boot helper no longer ignores that failure.
Both administrator passwords must be supplied. The generated first-boot file is
root-readable only. After applying passwords, the script replaces it with the
non-secret boot metadata. This is removal from the active filesystem, not a
claim of forensic erasure from flash media or old image copies.
Default Metis passwordless command grants are removed. Existing sudo-group
administration remains password protected. Do not provision a substitute
NOPASSWD rule when password setup fails.
## Existing-node repair, October 4, 2026
Titan-12/13/19 retained root SSH keys but had locked atlas passwords; their root
passwords also differed from Vault. Both accounts were restored to their existing
Vault credentials without rebooting. A separate fresh SSH session authenticated
sudo with the Vault atlas password and reached UID 0 on each node. Titan-20/21's
existing unlocked root passwords were also restored to their Vault values.
The Atlas repository contains `scripts/node_admin_access.py` for an explicit
hostname-checked audit/repair. It accepts credentials only on stdin and emits
booleans, never passwords or hashes. Consult Atlas's cluster operator guide for
the current verification record and unavailable nodes. Existing native Ubuntu
root-account locks do not prevent full administration through atlas and sudo.
Do not apply a recovered node's old firstboot.env blindly to a live node: fetch
that node's current Vault record and verify the hostname first. Verify independent
password-backed sudo before retiring a legacy Metis grant.
## Checks after burning a replacement image
1. `systemctl status metis-node-identity.service`
2. Confirm `/etc/metis/node-identity.pending` is absent.
3. Log in over a separate SSH connection and run `sudo -k -v`, using the stored
atlas password; then `sudo id -u` must print `0`.
4. Confirm `firstboot.env` contains no password variables, without printing the
file into shared logs. Preserve the existing authorized keys.
5. Check host/storage health before uncordoning the Kubernetes node through Flux.
The image and script tests exercise password failure, successful retry, repeated
execution, ext4 symlink enablement, and credential-file permissions. These tests
are not a physical image boot validation; do that on the next replacement medium.