nodes: document verified administration and remaining recovery work

This commit is contained in:
jenkins 2026-10-04 00:37:36 -05:00
parent c3f892fd0e
commit 36c7e76e56
4 changed files with 153 additions and 3 deletions

View File

@ -24,6 +24,52 @@ separates completed repairs from open problems. The
There is no new coordinating framework to learn. Prefer the native owner above.
Use a recovery tool only when its documented operation matches the failure.
## Node administrative access
From the management workstation, use the existing SSH aliases. Most nodes use
`atlas`; the exceptions are `titan-jh` (theia), `oceanus` (Titan-23, oceanus), and
`tethys` (Titan-24, tethys). The normal SSH port is 2277 through Titan-jh.
In Vault, open `kv/atlas/nodes/<hostname>`. For example,
`kv/atlas/nodes/titan-13` contains `atlas_password` for sudo and `root_password`
for the root account. Custom metadata now records `admin_username`,
`admin_access_method`, `admin_access_verified_utc` and `root_password_state`.
Those are dated audit results, not a continuously updated health guarantee.
A normal manual check is:
```bash
ssh titan-13
sudo -k
sudo -v
sudo id -u
sudo passwd -S root
```
Enter the node's atlas password at the sudo prompt; never put it into the shell
command or a shared log. UID 0 proves administrator access. Some original nodes
already have passwordless sudo; this work did not add such grants. Nine native
Ubuntu accounts retain a locked direct root login; use their administrator and
sudo. A nonempty Vault root_password field alone does not prove root login works.
For scoped repair, `scripts/node_admin_access.py` runs on the target as root and
accepts a hostname plus password mapping on stdin. It compares password hashes
locally, reports only booleans, and changes passwords only with `--apply`.
`--retire-legacy-sudo` removes only the two recognized historical Metis grants
after password verification, preserving a root-only rollback copy. Do not run
that migration before independently testing password-backed sudo.
Keep SSH host-key verification enabled. Titan-09/10 currently answer on port 22
with changed keys; their reinstall identity must be established before sending
passwords. Titan-05 answers ping but has no tested admin listener. Titan-06/16
did not answer the current LAN checks. See the dated implementation record.
Metis source now provisions a native one-shot identity service instead of
relying exclusively on cloud-init. After a future recovery, check the unit,
verify its pending marker cleared, then test a separate SSH/sudo session before
returning the node to Kubernetes service. Image publication and a real replacement
boot remain separate acceptance checks; do not assume a source commit is deployed.
## Automatic recovery that can change services
Ananke's coordinator runs as `ananke.service` on Titan-db; its settings are in

View File

@ -7,7 +7,7 @@ Use [Cluster operations](CLUSTER_OPERATIONS.md) for the ordinary operator path.
## Six-priority progress tracker
Started against the user's approved list on 2026-10-03. Updated: 2026-10-04 04:38 UTC.
Started against the user's approved list on 2026-10-03. Updated: 2026-10-04 05:35 UTC.
"Verified" means the stated check passed; it does not imply a completed soak or
failure drill. Physical repairs and disruptive recovery drills remain separate
from routine software work.
@ -15,7 +15,7 @@ from routine software work.
| Priority | Status | Completed evidence | Next action / completion gate |
| --- | --- | --- | --- |
| 1. Immediate software repairs | In progress; Soteria publication blocked | Monerod remains 2/2 Ready; its daemon has zero restarts, status sidecar one. Soteria release 1297 passed enforced gates but failed Go VCS stamping inside the image build; runtime remains 0.1.0-120 | Repair image build metadata, publish Soteria and validate an eligible backup and restore; continue Monerod latency/sync checks |
| 2. Reliable nodes | Partial; physical checks needed | Titan-04/14/18 excluded from new placement; Titan-05/06 remain offline. Fresh Titan-04/11 undervoltage was observed | Power/media checks, preservation or replacement of suspect media, then networking/storage/load validation before return. Establish the actual condition of 09/10/16 |
| 2. Reliable nodes | Software repairs applied; hardware and identity checks remain | Full root administration verified on 21 reachable machines; Vault verification metadata saved. Repaired credentials on 12/13/19/20/21. Titan-04/11/14/18 excluded from new placement. Titan-19 log rotation repaired | Verify changed SSH identities on 09/10; recover 05/06/16; address fresh power faults and runtime-media stalls. Complete Metis image publication and physical replacement-image validation |
| 3. Recovery coverage | In progress; cloud uploads blocked | External datastore restore and second LAN copy verified; 19 application DB dumps completed; scheduled application backup completed at 08:37 UTC; representative Gitea restore passed. Backblaze storage-cap failure established | Resolve the account storage cap within an agreed budget, complete eligible PVC coverage and Soteria restore, resolve Hermes tenant-2 safely, and provide approved local-only recovery copies |
| 4. Capacity and placement | Partial | CI agent cap reduced to two; shared PostgreSQL/OpenSearch resource corrections deployed; Vault injectors spread. Longhorn now uses least-effort balancing and one rebuild per node | Account for the measured 2-3.5 GiB storage-process footprint, correct remaining requests/placement, and calculate compatible spare capacity for one worker loss |
| 5. Supported software | Planned; backups established | Installed K3s cohorts identified as 1.31.5 and 1.33.3; titan-db Ubuntu 24.10 is beyond support | Select a compatible supported baseline; stage OS/K3s/storage upgrades with rollback. Plan collation/index repair |
@ -37,6 +37,77 @@ from routine software work.
- [ ] Complete application checks and the 7-14 day observation period.
- [ ] Resolve Backblaze's storage cap before accepting cloud backup coverage; preserve existing copies.
## Node access and software repair - October 4
The user prioritized node administration and software repairs before hardware
replacement. No node was reflashed or rebooted during this work.
- [x] Verify full administrative execution on all 21 reachable machines,
including the datastore and bastion; save the verification timestamp, admin
username, access method and root-password state in Vault custom metadata.
- [x] Repair Titan-12/13/19: SSH keys worked, but atlas was locked and the root
password differed from Vault. Both accounts now match their existing Vault
passwords. Fresh independent password-backed sudo reached UID 0.
- [x] Repair the existing unlocked root passwords on Titan-20/21 to match Vault.
- [x] Remove the obsolete Metis passwordless command grants on 12/13/19 after
verifying password-backed sudo; repeat the sudo verification afterward.
- [x] Fix the Vault bootstrap writer: preserve all fields, use compare-and-set,
skip unchanged records, keep temporary files private, remove content-bearing
errors, and retain the completed Job so Flux does not recreate it after TTL.
Live Job -5 changed only Titan-23's incorrect .23 address to .24; 25 records
were unchanged. Atlas commit 1bd8ae31.
- [x] Exclude Titan-11 from new placement through Flux (c2f41899); existing pods
continue running. Its fresh voltage faults remain a hardware issue.
- [x] Repair Titan-19's logrotate post-hook failure: armbian-ramlog was inactive
but ENABLED=true, copying between obsolete RAM-log locations although /var/log
already points to external ext4 storage. The guarded versioned script disabled
those hooks; the native logrotate service completed with result success/status 0.
- [x] Correct Ariadne's memory reservation/limit from 128/512 MiB to 512/1024 MiB
after repeated cgroup OOM kills. The replacement is 2/2 Ready with zero restarts.
This supplies headroom; it does not establish that its memory growth is bounded.
- [x] Publish Metis source fixes 1c110c9 and 6213159: native first-boot unit,
per-injection pending marker, serialized setup, propagated password failures,
removal of default passwordless grants, cleanup of applied bootstrap passwords,
and correct file permissions/native systemd links in both image-writing paths.
- [ ] Complete the normal Metis image publication and Flux rollout. Build 400 was
still active at the latest check; do not claim the running 0.1.0-399 contains
these source changes. The next replacement image still needs a physical boot
check before these changes are considered proven on hardware.
- [ ] Resolve the remaining network/physical findings below.
### Current node distinctions
| Nodes | Established evidence | Remaining action |
| --- | --- | --- |
| 12/13/19 | atlas/root passwords restored; root SSH keys preserved; password-backed sudo works after legacy grant retirement | Observe stability; use the corrected recovery image for future rebuilds |
| 20/21 | atlas sudo already worked; root passwords now match Vault | Continue workload memory diagnosis; Jetson-21 recorded a Data Prepper cgroup OOM |
| 04/11 | Fresh undervoltage messages continued during this audit, despite sufficient disk space | Check delivered power/cables/peripherals; do not clear quarantine based on a single idle voltage reading |
| 14 | Runtime USB flash, substantial I/O pressure; 39-45% iowait in the sample, no new kernel I/O error in the two-hour sample | Preserve data; repair or replace runtime media and validate before return |
| 18 | Quarantine has reduced current I/O pressure; its 50 MiB RAM-log filesystem is still active | Keep quarantined until sustained storage checks pass; do not apply the inactive-RAM-log repair script to it |
| 05 | Answers LAN ping at .31; expected SSH, kubelet and common recovery ports refuse connections | Console/user-space service check; this is not evidence that the machine is powered off |
| 09/10 | Answer LAN ping and SSH on port 22; recorded 2277 port is closed; host keys differ from saved keys | Confirm reinstall/image identity before supplying passwords; no host-key checks bypassed |
| 06/16 | No LAN ping response; neighbor resolution incomplete; SSH unavailable | Power, link and boot-media checks needed |
The 12 unlocked root passwords match Vault. Nine machines retain locked direct
root accounts (0a/0b/0c/04/07/08/11/db/jh); their administrator account plus sudo
still provides full root control. Root-account lock state is recorded explicitly
in Vault instead of implying that every root_password field is a usable login.
Validation: 13 Atlas regression tests passed (credential targeting, secret-safe
output, legacy-grant preservation, Vault no-op/CAS behavior); relevant Kustomize
builds, client dry-runs and focused Flux diffs passed. Metis's complete Go tests
and separate docs/LOC/vet/per-file coverage gate passed. Tests include a real
throwaway ext4 image for systemd link injection and mocked password-setup failure,
retry and stale image markers. No live node was burned to test recovery.
Rollback: Kubernetes changes use their focused Git commit and Flux. Titan-19's
previous RAM-log config is root-only under
`/var/lib/atlas-maintenance/ramlog-before-20261004/`; restoring it reintroduces the
obsolete hooks. Retired sudo grants are root-only under
`/var/lib/atlas-maintenance/legacy-sudo-20261004/`; do not restore passwordless
access merely to avoid using the now-working password. Password repairs use the
already stored Vault credentials; no secret values were changed or committed.
## Verified changes
| Change | Evidence | Rollback |

View File

@ -77,9 +77,10 @@ def retire_legacy_sudo(result):
raise ValueError("atlas_password_not_verified")
expected = ("atlas ALL=(ALL) NOPASSWD: /usr/bin/systemctl, /usr/sbin/poweroff, "
"/sbin/poweroff, /usr/local/bin/hecate, /usr/local/bin/k3s, /usr/bin/k3s")
known_grants = {expected, expected.split(", /usr/local/bin/k3s")[0]}
paths = [Path("/etc/sudoers.d/90-hecate-atlas"), Path("/etc/metis/sudoers-hecate")]
for path in paths:
if path.exists() and path.read_text().strip() != expected:
if path.exists() and path.read_text().strip() not in known_grants:
raise ValueError("legacy_grant_modified_requires_review")
if subprocess.run(["/usr/sbin/visudo", "-c"], capture_output=True).returncode:
raise ValueError("sudo_configuration_invalid")

View File

@ -2,6 +2,7 @@
import importlib.util
from pathlib import Path
import unittest
import tempfile
from unittest.mock import patch, Mock
spec = importlib.util.spec_from_file_location(
@ -70,6 +71,37 @@ class NodeAccessTests(unittest.TestCase):
with self.assertRaisesRegex(ValueError, "^password_update_failed$"):
module.run(self.payload, True)
def test_legacy_grant_retirement_keeps_a_rollback_copy(self):
with tempfile.TemporaryDirectory() as directory:
root = Path(directory)
grant = root / "etc/sudoers.d/90-hecate-atlas"
grant.parent.mkdir(parents=True)
grant.write_text("atlas ALL=(ALL) NOPASSWD: /usr/bin/systemctl, /usr/sbin/poweroff, "
"/sbin/poweroff, /usr/local/bin/hecate\n")
result = {"after": {"atlas": {"vault_password_matches": True}}}
with patch.object(module, "Path", side_effect=lambda p: root / p.lstrip("/")), \
patch.object(module.subprocess, "run", return_value=Mock(returncode=0)):
module.retire_legacy_sudo(result)
module.retire_legacy_sudo(result)
self.assertFalse(grant.exists())
self.assertTrue((root / "var/lib/atlas-maintenance/legacy-sudo-20261004/90-hecate-atlas").exists())
def test_custom_sudo_rule_is_never_removed(self):
with tempfile.TemporaryDirectory() as directory:
root = Path(directory)
grant = root / "etc/sudoers.d/90-hecate-atlas"
grant.parent.mkdir(parents=True)
grant.write_text("reviewed custom rule")
result = {"after": {"atlas": {"vault_password_matches": True}}}
with patch.object(module, "Path", side_effect=lambda p: root / p.lstrip("/")):
with self.assertRaisesRegex(ValueError, "legacy_grant_modified_requires_review"):
module.retire_legacy_sudo(result)
self.assertEqual(grant.read_text(), "reviewed custom rule")
def test_legacy_grant_requires_working_password(self):
with self.assertRaisesRegex(ValueError, "atlas_password_not_verified"):
module.retire_legacy_sudo({"after": {"atlas": {"vault_password_matches": False}}})
if __name__ == "__main__":
unittest.main()