55 lines
3.0 KiB
Markdown
55 lines
3.0 KiB
Markdown
# Node administration after recovery
|
|
|
|
Use the normal node SSH account and password-protected sudo. Passwords live in
|
|
Vault at `kv/atlas/nodes/<hostname>` under `atlas_password` and `root_password`.
|
|
The SSH key remains the normal login mechanism; a console root password does
|
|
not require allowing root password login over SSH.
|
|
|
|
Recovery injects a native `metis-node-identity.service`, enabled in the image.
|
|
It applies passwords even when cloud-init is absent. Each new injection creates
|
|
`/etc/metis/node-identity.pending`, so an old completion marker in the base image
|
|
cannot suppress the new identity. Cloud-init and the native unit serialize on
|
|
one lock. Failed credential setup remains a failed unit with the pending marker
|
|
intact; the Longhorn first-boot helper no longer ignores that failure.
|
|
|
|
Both administrator passwords must be supplied. The generated first-boot file is
|
|
root-readable only. After applying passwords, the script replaces it with the
|
|
non-secret boot metadata. This is removal from the active filesystem, not a
|
|
claim of forensic erasure from flash media or old image copies.
|
|
|
|
Default Metis passwordless command grants are removed. Existing sudo-group
|
|
administration remains password protected. Do not provision a substitute
|
|
NOPASSWD rule when password setup fails.
|
|
|
|
## Existing-node repair, October 4, 2026
|
|
|
|
Titan-12/13/19 retained root SSH keys but had locked atlas passwords; their root
|
|
passwords also differed from Vault. Both accounts were restored to their existing
|
|
Vault credentials without rebooting. A separate fresh SSH session authenticated
|
|
sudo with the Vault atlas password and reached UID 0 on each node. Titan-20/21's
|
|
existing unlocked root passwords were also restored to their Vault values.
|
|
|
|
The Atlas repository contains `scripts/node_admin_access.py` for an explicit
|
|
hostname-checked audit/repair. It accepts credentials only on stdin and emits
|
|
booleans, never passwords or hashes. Consult Atlas's cluster operator guide for
|
|
the current verification record and unavailable nodes. Existing native Ubuntu
|
|
root-account locks do not prevent full administration through atlas and sudo.
|
|
|
|
Do not apply a recovered node's old firstboot.env blindly to a live node: fetch
|
|
that node's current Vault record and verify the hostname first. Verify independent
|
|
password-backed sudo before retiring a legacy Metis grant.
|
|
|
|
## Checks after burning a replacement image
|
|
|
|
1. `systemctl status metis-node-identity.service`
|
|
2. Confirm `/etc/metis/node-identity.pending` is absent.
|
|
3. Log in over a separate SSH connection and run `sudo -k -v`, using the stored
|
|
atlas password; then `sudo id -u` must print `0`.
|
|
4. Confirm `firstboot.env` contains no password variables, without printing the
|
|
file into shared logs. Preserve the existing authorized keys.
|
|
5. Check host/storage health before uncordoning the Kubernetes node through Flux.
|
|
|
|
The image and script tests exercise password failure, successful retry, repeated
|
|
execution, ext4 symlink enablement, and credential-file permissions. These tests
|
|
are not a physical image boot validation; do that on the next replacement medium.
|