3.0 KiB
Node administration after recovery
Use the normal node SSH account and password-protected sudo. Passwords live in
Vault at kv/atlas/nodes/<hostname> under atlas_password and root_password.
The SSH key remains the normal login mechanism; a console root password does
not require allowing root password login over SSH.
Recovery injects a native metis-node-identity.service, enabled in the image.
It applies passwords even when cloud-init is absent. Each new injection creates
/etc/metis/node-identity.pending, so an old completion marker in the base image
cannot suppress the new identity. Cloud-init and the native unit serialize on
one lock. Failed credential setup remains a failed unit with the pending marker
intact; the Longhorn first-boot helper no longer ignores that failure.
Both administrator passwords must be supplied. The generated first-boot file is root-readable only. After applying passwords, the script replaces it with the non-secret boot metadata. This is removal from the active filesystem, not a claim of forensic erasure from flash media or old image copies.
Default Metis passwordless command grants are removed. Existing sudo-group administration remains password protected. Do not provision a substitute NOPASSWD rule when password setup fails.
Existing-node repair, October 4, 2026
Titan-12/13/19 retained root SSH keys but had locked atlas passwords; their root passwords also differed from Vault. Both accounts were restored to their existing Vault credentials without rebooting. A separate fresh SSH session authenticated sudo with the Vault atlas password and reached UID 0 on each node. Titan-20/21's existing unlocked root passwords were also restored to their Vault values.
The Atlas repository contains scripts/node_admin_access.py for an explicit
hostname-checked audit/repair. It accepts credentials only on stdin and emits
booleans, never passwords or hashes. Consult Atlas's cluster operator guide for
the current verification record and unavailable nodes. Existing native Ubuntu
root-account locks do not prevent full administration through atlas and sudo.
Do not apply a recovered node's old firstboot.env blindly to a live node: fetch that node's current Vault record and verify the hostname first. Verify independent password-backed sudo before retiring a legacy Metis grant.
Checks after burning a replacement image
systemctl status metis-node-identity.service- Confirm
/etc/metis/node-identity.pendingis absent. - Log in over a separate SSH connection and run
sudo -k -v, using the stored atlas password; thensudo id -umust print0. - Confirm
firstboot.envcontains no password variables, without printing the file into shared logs. Preserve the existing authorized keys. - Check host/storage health before uncordoning the Kubernetes node through Flux.
The image and script tests exercise password failure, successful retry, repeated execution, ext4 symlink enablement, and credential-file permissions. These tests are not a physical image boot validation; do that on the next replacement medium.