docs: refresh cluster health and Metis rollout status
This commit is contained in:
parent
52dffae88c
commit
cc238c0b6c
@ -7,7 +7,7 @@ Use [Cluster operations](CLUSTER_OPERATIONS.md) for the ordinary operator path.
|
||||
|
||||
## Six-priority progress tracker
|
||||
|
||||
Started against the user's approved list on 2026-10-03. Updated: 2026-10-04 06:20 UTC.
|
||||
Started against the user's approved list on 2026-10-03. Updated: 2026-10-04 07:43 UTC.
|
||||
"Verified" means the stated check passed; it does not imply a completed soak or
|
||||
failure drill. Physical repairs and disruptive recovery drills remain separate
|
||||
from routine software work.
|
||||
@ -78,10 +78,12 @@ replacement. No node was reflashed or rebooted during this work.
|
||||
per-injection pending marker, serialized setup, propagated password failures,
|
||||
removal of default passwordless grants, cleanup of applied bootstrap passwords,
|
||||
and correct file permissions/native systemd links in both image-writing paths.
|
||||
- [ ] Complete the normal Metis image publication and Flux rollout. Build 400 was
|
||||
still active at the latest check; do not claim the running 0.1.0-399 contains
|
||||
these source changes. The next replacement image still needs a physical boot
|
||||
check before these changes are considered proven on hardware.
|
||||
- [ ] Complete the normal Metis image publication and Flux rollout. Build 400
|
||||
ended ABORTED after 7272.9 seconds; the exact abort cause is not established
|
||||
from the inspected metadata. Build 401 is running its quality gate. The live
|
||||
image is still 0.1.0-399 and does not contain these source changes. The next
|
||||
replacement image still needs a physical boot check before these changes are
|
||||
considered proven on hardware.
|
||||
- [x] Enable Ananke's existing password-backed host-action path. Atlas 886dbeee
|
||||
adds 13 atlas-password mappings to the existing Vault CSI synchronizer; all
|
||||
matched Vault, with native two-minute rotation enabled. No root passwords are
|
||||
@ -101,6 +103,16 @@ replacement. No node was reflashed or rebooted during this work.
|
||||
checks, not just successful key installation.
|
||||
- [ ] Resolve the remaining network/physical findings below.
|
||||
|
||||
At the 07:43 UTC check, 57/57 active Flux configurations and 19/21 Kubernetes
|
||||
nodes remained Ready. Ananke on both hosts had zero automatic restarts;
|
||||
Ariadne, Jenkins and the Grafana application container were Ready with zero
|
||||
restarts. The Grafana Vault sidecar had two earlier restarts. Storage acceptance
|
||||
remains open: Titan-15's engine-image installer was in CrashLoopBackOff with
|
||||
previous exit 137 (not established as an OOM), Titan-18's manager was unready,
|
||||
and the existing Titan-14 share manager was unready. Longhorn reported 74
|
||||
attached healthy volumes and 61 detached volumes with unknown robustness;
|
||||
detached/unknown by itself is not a new volume failure.
|
||||
|
||||
### Current node distinctions
|
||||
|
||||
| Nodes | Established evidence | Remaining action |
|
||||
|
||||
Loading…
x
Reference in New Issue
Block a user