docs: record Ananke access verification and node repair limits

This commit is contained in:
jenkins 2026-10-04 00:54:45 -05:00
parent 886dbeee10
commit d7dd48eee2
3 changed files with 101 additions and 5 deletions

View File

@ -70,6 +70,33 @@ verify its pending marker cleared, then test a separate SSH/sudo session before
returning the node to Kubernetes service. Image publication and a real replacement
boot remain separate acceptance checks; do not assume a source commit is deployed.
Ananke's allowlisted host actions can now obtain the atlas password through the
existing CSI synchronization, configured in
`services/maintenance/bootstrap/secretproviderclass.yaml`. The synchronized
Secrets are named `maintenance/ananke-sudo-<hostname>`, with key `password`.
Only the 13 verified workers needing that path are included; root passwords
remain in Vault. The CSI driver refreshes them every two minutes.
The two native Ananke configurations use:
```yaml
startup:
host_sudo_secret_namespace: maintenance
host_sudo_secret_name_template: "ananke-sudo-{node}"
host_sudo_secret_password_key: password
```
These are fields within the existing startup section, not a replacement config.
`scripts/configure_ananke_node_access.py` applies that narrow change, validates it
with Ananke, and retains a root-only backup. The coordinator remains Titan-db;
Titan-24 remains its peer. Credentials go to sudo stdin for predefined actions,
not arbitrary command arguments or logs. No new controller was introduced.
This automated worker-password path requires the Kubernetes API. It is not a
standalone cold-start credential store. A read-only access check is not a power
failure/recovery drill. When onboarding another recovered worker, first validate
its Vault password and SSH identity, then add its explicit CSI mapping.
## Automatic recovery that can change services
Ananke's coordinator runs as `ananke.service` on Titan-db; its settings are in

View File

@ -7,7 +7,7 @@ Use [Cluster operations](CLUSTER_OPERATIONS.md) for the ordinary operator path.
## Six-priority progress tracker
Started against the user's approved list on 2026-10-03. Updated: 2026-10-04 05:35 UTC.
Started against the user's approved list on 2026-10-03. Updated: 2026-10-04 05:55 UTC.
"Verified" means the stated check passed; it does not imply a completed soak or
failure drill. Physical repairs and disruptive recovery drills remain separate
from routine software work.
@ -73,6 +73,16 @@ replacement. No node was reflashed or rebooted during this work.
still active at the latest check; do not claim the running 0.1.0-399 contains
these source changes. The next replacement image still needs a physical boot
check before these changes are considered proven on hardware.
- [x] Enable Ananke's existing password-backed host-action path. Atlas 886dbeee
adds 13 atlas-password mappings to the existing Vault CSI synchronizer; all
matched Vault, with native two-minute rotation enabled. No root passwords are
copied into that synchronizer. Source Ananke 41fd207 sets the namespace/name
template and corrects Titan-24's administrator to tethys. Both live host configs
were patched without changing unrelated settings and passed Ananke validation.
- [x] Verify read-only privileged actions using Ananke's own coordinator SSH
identity on all 13 affected workers, and its peer identity on 12/13/19. Both
Ananke services were restarted separately and returned active/running with
zero automatic restarts. Jenkins credential-helper generation remained 796.
- [ ] Resolve the remaining network/physical findings below.
### Current node distinctions
@ -81,7 +91,7 @@ replacement. No node was reflashed or rebooted during this work.
| --- | --- | --- |
| 12/13/19 | atlas/root passwords restored; root SSH keys preserved; password-backed sudo works after legacy grant retirement | Observe stability; use the corrected recovery image for future rebuilds |
| 20/21 | atlas sudo already worked; root passwords now match Vault | Continue workload memory diagnosis; Jetson-21 recorded a Data Prepper cgroup OOM |
| 04/11 | Fresh undervoltage messages continued during this audit, despite sufficient disk space | Check delivered power/cables/peripherals; do not clear quarantine based on a single idle voltage reading |
| 04/11 | Fresh undervoltage messages continued during this audit, despite sufficient disk space | Onboard EXT5V samples were 4.82132 V (04) and 4.75968 V (11). Check delivered power/cables/peripherals; do not clear quarantine based on one sample |
| 14 | Runtime USB flash, substantial I/O pressure; 39-45% iowait in the sample, no new kernel I/O error in the two-hour sample | Preserve data; repair or replace runtime media and validate before return |
| 18 | Quarantine has reduced current I/O pressure; its 50 MiB RAM-log filesystem is still active | Keep quarantined until sustained storage checks pass; do not apply the inactive-RAM-log repair script to it |
| 05 | Answers LAN ping at .31; expected SSH, kubelet and common recovery ports refuse connections | Console/user-space service check; this is not evidence that the machine is powered off |
@ -93,12 +103,20 @@ root accounts (0a/0b/0c/04/07/08/11/db/jh); their administrator account plus sud
still provides full root control. Root-account lock state is recorded explicitly
in Vault instead of implying that every root_password field is a usable login.
Validation: 13 Atlas regression tests passed (credential targeting, secret-safe
output, legacy-grant preservation, Vault no-op/CAS behavior); relevant Kustomize
Validation: 16 Atlas regression tests passed (credential targeting, secret-safe
output, legacy-grant preservation, Vault no-op/CAS behavior, and preservation/idempotency of the Ananke config patch); relevant Kustomize
builds, client dry-runs and focused Flux diffs passed. Metis's complete Go tests
and separate docs/LOC/vet/per-file coverage gate passed. Tests include a real
throwaway ext4 image for systemd link injection and mocked password-setup failure,
retry and stale image markers. No live node was burned to test recovery.
retry and stale image markers. No live node was burned to test recovery. All 57 active Flux Kustomizations
remained Ready after these changes; 19/21 Kubernetes nodes are Ready.
Ananke still depends on an available Kubernetes API to retrieve its synchronized
worker sudo credential. This is its existing implementation, not a new offline
credential cache. Control-plane administration remains available through its
existing path; use the dated manual SSH/Vault access record for intervention
outside the automated recovery flow. Do not claim a full power-loss drill from
read-only connection tests.
Rollback: Kubernetes changes use their focused Git commit and Flux. Titan-19's
previous RAM-log config is root-only under
@ -107,6 +125,11 @@ obsolete hooks. Retired sudo grants are root-only under
`/var/lib/atlas-maintenance/legacy-sudo-20261004/`; do not restore passwordless
access merely to avoid using the now-working password. Password repairs use the
already stored Vault credentials; no secret values were changed or committed.
Native Ananke config backups are root-only under
`/var/lib/atlas-maintenance/ananke-access-before-20261004/` on both hosts. Reverting
its lookup settings would remove automated password-backed access; retain the
working administrator passwords. Its source profiles and the versioned Atlas
configuration helper document the change.
## Verified changes

View File

@ -0,0 +1,46 @@
"""Keep a narrow node-access update from replacing unrelated recovery settings."""
import importlib.util
from pathlib import Path
import unittest
import yaml
spec = importlib.util.spec_from_file_location('node_access_config',
Path(__file__).resolve().parents[2] / 'scripts/configure_ananke_node_access.py')
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
class AccessConfigTests(unittest.TestCase):
def test_unrelated_settings_and_comments_survive(self):
original = '''# operator comment
ssh_node_users:
titan-24: atlas
titan-12: atlas
workers:
- titan-12
startup:
host_sudo_secret_namespace: old
bounded_recovery: true
timeout: 45
power:
shutdown: false
'''
result = module.updated_config(original)
after = yaml.safe_load(result)
self.assertIn('# operator comment', result)
self.assertEqual(after['power'], {'shutdown': False})
self.assertEqual(after['startup']['timeout'], 45)
self.assertTrue(after['startup']['bounded_recovery'])
self.assertEqual(after['ssh_node_users']['titan-12'], 'atlas')
self.assertEqual(after['ssh_node_users']['titan-24'], 'tethys')
self.assertEqual(after['startup']['host_sudo_secret_namespace'], 'maintenance')
self.assertEqual(module.updated_config(result), result)
def test_unexpected_structure_fails_before_write(self):
with self.assertRaisesRegex(ValueError, 'unexpected_config_structure'):
module.updated_config('startup: {}\nssh_node_users: {}\n')
def test_missing_peer_override_is_added(self):
source = '\nssh_node_users:\n titan-12: atlas\nworkers: []\nstartup:\n timeout: 45\n'
result = yaml.safe_load(module.updated_config(source))
self.assertEqual(result['ssh_node_users'], {'titan-12': 'atlas', 'titan-24': 'tethys'})