metis/docs/node-administration.md

3.0 KiB

Node administration after recovery

Use the normal node SSH account and password-protected sudo. Passwords live in Vault at kv/atlas/nodes/<hostname> under atlas_password and root_password. The SSH key remains the normal login mechanism; a console root password does not require allowing root password login over SSH.

Recovery injects a native metis-node-identity.service, enabled in the image. It applies passwords even when cloud-init is absent. Each new injection creates /etc/metis/node-identity.pending, so an old completion marker in the base image cannot suppress the new identity. Cloud-init and the native unit serialize on one lock. Failed credential setup remains a failed unit with the pending marker intact; the Longhorn first-boot helper no longer ignores that failure.

Both administrator passwords must be supplied. The generated first-boot file is root-readable only. After applying passwords, the script replaces it with the non-secret boot metadata. This is removal from the active filesystem, not a claim of forensic erasure from flash media or old image copies.

Default Metis passwordless command grants are removed. Existing sudo-group administration remains password protected. Do not provision a substitute NOPASSWD rule when password setup fails.

Existing-node repair, October 4, 2026

Titan-12/13/19 retained root SSH keys but had locked atlas passwords; their root passwords also differed from Vault. Both accounts were restored to their existing Vault credentials without rebooting. A separate fresh SSH session authenticated sudo with the Vault atlas password and reached UID 0 on each node. Titan-20/21's existing unlocked root passwords were also restored to their Vault values.

The Atlas repository contains scripts/node_admin_access.py for an explicit hostname-checked audit/repair. It accepts credentials only on stdin and emits booleans, never passwords or hashes. Consult Atlas's cluster operator guide for the current verification record and unavailable nodes. Existing native Ubuntu root-account locks do not prevent full administration through atlas and sudo.

Do not apply a recovered node's old firstboot.env blindly to a live node: fetch that node's current Vault record and verify the hostname first. Verify independent password-backed sudo before retiring a legacy Metis grant.

Checks after burning a replacement image

  1. systemctl status metis-node-identity.service
  2. Confirm /etc/metis/node-identity.pending is absent.
  3. Log in over a separate SSH connection and run sudo -k -v, using the stored atlas password; then sudo id -u must print 0.
  4. Confirm firstboot.env contains no password variables, without printing the file into shared logs. Preserve the existing authorized keys.
  5. Check host/storage health before uncordoning the Kubernetes node through Flux.

The image and script tests exercise password failure, successful retry, repeated execution, ext4 symlink enablement, and credential-file permissions. These tests are not a physical image boot validation; do that on the next replacement medium.