docs: record stability repairs and settling evidence

This commit is contained in:
jenkins 2026-10-04 08:51:31 -05:00
parent 1aa556d2b1
commit cb2342bbe7
3 changed files with 562 additions and 21 deletions

View File

@ -73,8 +73,9 @@ did not answer the current LAN checks. See the dated implementation record.
Metis source now provisions a native one-shot identity service instead of
relying exclusively on cloud-init. After a future recovery, check the unit,
verify its pending marker cleared, then test a separate SSH/sudo session before
returning the node to Kubernetes service. Image publication and a real replacement
boot remain separate acceptance checks; do not assume a source commit is deployed.
returning the node to Kubernetes service. Metis 0.1.0-403 is now published and
Ready in the cluster. A real replacement boot remains a separate acceptance
check; a successful image build does not prove hardware provisioning.
Ananke's allowlisted host actions can now obtain the atlas password through the
existing CSI synchronization, configured in
@ -114,6 +115,12 @@ its Vault password and SSH identity, then add its explicit CSI mapping.
## Automatic recovery that can change services
The `node-nofile` setup helper prepares systemd file limits and live inotify
settings. It no longer restarts K3s when the file-limit definition changes. New
unit limits take effect at the next planned K3s restart; do not assume a helper
rollout activated them. Its small reservations reflect an idle setup process,
while its burst limits are preserved.
Ananke's coordinator runs as `ananke.service` on Titan-db; its settings are in
`/etc/ananke/ananke.yaml`, with installed source in `/opt/ananke`. Titan-24 is
the UPS peer. Only the coordinator runs periodic cluster repair. UPS monitoring
@ -147,16 +154,47 @@ diagnosis/repair remains enabled. Do not re-enable cancellation until it can
distinguish a stalled job from a long-running job making useful progress.
Jenkins's Kubernetes cloud cap is currently one in
`services/jenkins/configmap-jcasc.yaml`. The shared default, IaC, Data Prepper and
Metis templates exclude quarantined nodes and storage workers. Custom inline
templates in other repos may need the same correction. Inspect actual agent
placement rather than assuming the default template controls every pipeline.
`services/jenkins/configmap-jcasc.yaml`. The active inline templates in Ariadne,
Atlasbot, Ananke, Metis, Pegasus,
Soteria, bstein-dev-home, IaC and Data Prepper now require Pi workers and exclude
quarantined and storage workers. Inspect actual agent placement rather than
assuming the shared default controls a custom template.
A Pending agent consumes the cap; inspect FailedScheduling events and resource
requests before deciding that Jenkins is hung. Do not increase concurrency to
solve an unschedulable agent. The tracked JCasC revision triggers a controller
rollout, so quiet the queue and arrange active-build completion before editing.
Node quarantine is declared in `infrastructure/core/node-maintenance.yaml`.
Custom pipelines explicitly use `workspaceVolume dynamicPVC(...)` with the
`ci-scratch` StorageClass in `infrastructure/core/storageclass-ci-scratch.yaml`.
Workspaces and compiler/package caches use that volume. Docker-in-Docker stores
its layers in the same volume under `.cache/docker`, outside the checkout.
Ordinary workspaces request 20 GiB; Docker builds request 40 GiB. Builds remain
serialized. This moves heavy scratch I/O off the Pi runtime USB drive. Soteria
also compiles its UI and binary on scratch before Kaniko assembles the small
`Dockerfile.runtime` image. Compiling inside Kaniko would write toolchains and
compiler intermediates into the container runtime again, defeating that isolation.
Data Prepper now uses the existing Docker-in-Docker path with a 40 GiB scratch
PVC for image layers. A scratch workspace alone does not relocate Kaniko
extraction out of a container root filesystem. The Docker API binds only to pod
loopback; registry passwords go to stdin and the temporary client config is
removed afterward. ARM64 output remains explicit.
Keep registry-credential umask changes local to credential creation. Soteria
explicitly sets the executable mode and nonroot ownership before image assembly;
a successful compile alone does not prove the deployed user can execute it.
`ci-scratch` is disposable, single-replica Longhorn storage on the existing
`astreae` disks. Jenkins owns its PVC through the agent pod; deleting that pod
removes its scratch PVC and volume. Never use this class for application data,
credentials, recovery copies or retained artifacts. Test artifacts are still
archived to the Jenkins controller with their existing 30-day policy. A lost
scratch volume requires rerunning the build. Soteria excludes the ci-scratch
class from application backups alongside local-path; preserve Hermes exclusions.
Check orphaned PVCs before assuming cleanup occurred; do not delete unrelated
retained volumes.
Node quarantine is declared in `infrastructure/core/node-maintenance.yaml` for
Titan-04/05/06/08/11/12/14/18. Offline nodes remain quarantined after reconnecting
until their physical and loaded-operation checks pass.
Those Node resources explicitly preserve hardware/worker/storage labels and use
Flux SSA Merge to retain other runtime fields. A cordon blocks ordinary new
placement; it must not repeatedly remove existing storage DaemonSets. After a
@ -164,6 +202,56 @@ quarantine change, verify manager pod UIDs stay unchanged across reconciliation.
Keep prune protection on Node declarations; repair, uncordon in Git and verify
before considering their removal.
Keep one-off Flux Jobs completed without a TTL, or remove them from active
resources after success. A TTL deletes the object and Flux then creates it
again. The archived Titan-24 emergency root sweep is deliberately inactive;
its zero-percent thresholds are not a normal maintenance schedule. The Metis
SSH public-key bootstrap now retains its completion.
Kubelet owns container-image garbage collection. All 19 reachable nodes were
verified with 85% high / 80% low thresholds. The legacy-named node-image-sweeper
no longer runs `crictl rmi --prune` or removes files from containerd/image-import
directories. Its remaining role is bounded host-log/package-cache maintenance.
The UID-aware Python helper preserves active pod logs in both `/var/log/pods`
and Armbian `/var/log.hdd/pods`, skips unavailable inventories, and provides
`--dry-run`. Native kubelet rotation controls active container logs.
The node-image-sweeper DaemonSet uses a startup probe that passes after its first
cleanup, then performs no continuous probe process. Its normal rollout permits
one unavailable helper. Offline terminating pods can consume that budget;
inspect those before raising it. This does not change application availability.
## Ingress and search capacity
Traefik runs two replicas on separate NVMe control-plane nodes. Its internal
`/ping` readiness check gates traffic. The public VIP remains `192.168.22.9`;
Gitea SSH and Wolf share it using their existing ports. MetalLB's `ingress-pool`
and `ingress-adv` restrict its announcer to the control planes. The private
`192.168.22.50` pool does the same and retains `externalTrafficPolicy: Local`,
source restrictions, TLS and service authentication. The communications pool
retains `.4` through `.6`; its Local-policy services need a speaker on their
actual endpoint node. Do not restrict all pools to ingress nodes indiscriminately.
MetalLB uses native L2 mode, without unused FRR/BGP sidecars. There are no BGP
peers or advertisements. Titan-05/06 are excluded from the speaker DaemonSet
until their recovery checks pass. Remove that explicit exclusion in
`infrastructure/metallb/helmrelease.yaml` when restoring them. Do not enable BGP
features later without revisiting this choice. The speaker requests 128 MiB and
uses five-second probes. Native L2 still has one announcing node per VIP.
OpenSearch's existing PVC is now 1280 GiB. Its expansion reduced measured usage
from 94% to 76%, allowing the blocked maintenance policy updates to resume. The
existing scheduled tuner now reports failing operation/path/status without
index contents. Its next scheduled run passed. No retention period or disk
watermark was relaxed. Inspect free space and ingestion growth periodically;
previous growth while writes were blocked is not a useful capacity forecast.
The current volume still has one Longhorn replica; it is not storage HA.
The scheduled tuner uses the native maintenance-batch priority class with
preemptionPolicy Never. It prefers Pi 5 workers but permits Pi 4 workers when
capacity is unavailable. Routine maintenance should wait rather than evict
running work; check scheduling deadlines instead of raising its priority.
**Do not shrink the PVC or blindly revert its expansion commit.**
## Five-minute check
Run from the management host with the existing administrator configuration:
@ -210,6 +298,43 @@ ssh titan-0b sudo cat /var/backups/atlas-postgres/latest/COMPLETE
Their freshness alert fires after 36 hours. See the host backup notes before
running the isolated restore verification or changing retention.
## Check a settling window
Use the read-only helper after the rollout finishes, then compare 10-15 minutes
later. It uses ordinary kubectl access; it does not repair, restart or delete
anything and has no AI or provider dependency.
```bash
python3 scripts/ops/cluster_settle_check.py --output /tmp/atlas-before.json
# Wait 10-15 minutes with normal workload activity.
python3 scripts/ops/cluster_settle_check.py --output /tmp/atlas-after.json --previous /tmp/atlas-before.json
```
Review restart increases, replaced service pods, unready workloads, requested
but unavailable storage and new warning evidence. Expected short-lived CI agent
turnover is different from an application controller repeatedly replacing pods.
Known offline nodes are reported separately, not hidden. Pods already marked
for deletion are not considered available merely because Ready remains true.
Scheduler Preempted events are included even when their event type is Normal. A newly observed event
has an unknown counter delta because its lifetime count may predate the window.
The helper deliberately omits Secret values, environments, log bodies and event
messages. Pair it with the service health endpoints and node I/O/memory metrics;
Kubernetes can report a node Ready while its runtime disk is stalled. The helper
checks Kustomizations and HelmReleases, not image-release policies or application
incident history; inspect those separately before claiming all automation green.
Current placement is deliberately conservative: Gitea and ClamAV use Titan-22's
CPU/RAM as last-resort general capacity; they request no GPU. Nextcloud stays on
a Pi worker. Titan-08 is also quarantined after a build saturated its runtime USB
drive. Keep Jenkins at one agent until ordinary builds and application activity
coexist without recurring pressure. See the dated cycle record for measurements.
Nextcloud startup now preserves installed apps and their external-app keys.
App upgrades are separate maintenance, rather than downloads/deletions on every
pod replacement. Its single-writer volumes require maxSurge 0 / maxUnavailable 1;
a replacement therefore has a short planned interruption. ClamAV readiness uses
its native PING/PONG response; a running process alone does not open the endpoint.
## Find the cause before choosing a repair
| Symptom | First check | Usual next action |

View File

@ -4,10 +4,12 @@ This record begins on 2026-10-03 UTC, following authorization to implement the
[configuration-first plan](cluster-audit-20261002/INTEGRATED_PLAN.md).
The October 2 audit remains a historical baseline, not a current health claim.
Use [Cluster operations](CLUSTER_OPERATIONS.md) for the ordinary operator path.
The latest [repair and settling cycles](STABILITY_CYCLES_20261004.md) supersede
the earlier point-in-time CI capacity and tenant-2 status below.
## Six-priority progress tracker
Started against the user's approved list on 2026-10-03. Updated: 2026-10-04 08:18 UTC.
Started against the user's approved list on 2026-10-03. Updated: 2026-10-04 13:50 UTC.
"Verified" means the stated check passed; it does not imply a completed soak or
failure drill. Physical repairs and disruptive recovery drills remain separate
from routine software work.
@ -19,10 +21,10 @@ small and inspectable; a new recovery controller is not a prerequisite.
| Priority | Status | Completed evidence | Next action / completion gate |
| --- | --- | --- | --- |
| 1. Immediate software repairs | In progress; Soteria publication blocked | Monerod remains 2/2 Ready; its daemon has zero restarts, status sidecar one. Soteria release 1297 passed enforced gates but failed Go VCS stamping inside the image build; runtime remains 0.1.0-120 | Repair image build metadata, publish Soteria and validate an eligible backup and restore; continue Monerod latency/sync checks |
| 2. Reliable nodes | Software repairs applied; hardware and identity checks remain | Full root administration verified on 21 reachable machines; Vault verification metadata saved. Repaired credentials on 12/13/19/20/21. Titan-04/11/14/18 excluded from new placement. Titan-19 log rotation repaired | Verify changed SSH identities on 09/10; recover 05/06/16; address fresh power faults and runtime-media stalls. Complete Metis image publication and physical replacement-image validation |
| 3. Recovery coverage | In progress; cloud uploads blocked | External datastore restore and second LAN copy verified; 19 application DB dumps completed; scheduled application backup completed at 08:37 UTC; representative Gitea restore passed. Backblaze storage-cap failure established | Resolve the account storage cap within an agreed budget, complete eligible PVC coverage and Soteria restore, resolve Hermes tenant-2 safely, and provide approved local-only recovery copies |
| 4. Capacity and placement | Partial; CI waiting for compatible capacity | CI agent cap reduced to one; shared default, Metis, IaC and Data Prepper agent templates now exclude storage/quarantined workers. Shared PostgreSQL/OpenSearch resource corrections deployed; Vault injectors spread. Longhorn uses least-effort balancing and one rebuild per node | Restore compatible healthy worker capacity; inspect other projects' inline agent templates. Account for the measured 2-3.5 GiB storage-process footprint and calculate spare capacity for one worker loss |
| 1. Immediate software repairs | Corrected Soteria release deployed; recovery checks remain | Soteria 1305 succeeded and 0.1.0-128 is Ready with executable nonroot binary. Failed 127 rollout retained the serving old replica. Monerod previously recovered with existing data | Validate an eligible backup and restore after resolving the cloud cap; continue Monerod latency/sync checks |
| 2. Reliable nodes | Software repairs applied; hardware and identity checks remain | Full root administration verified on 21 reachable machines; Vault verification metadata saved. Repaired credentials on 12/13/19/20/21. Titan-04/08/11/12/14/18 excluded from new placement. Titan-19 log rotation repaired | Verify changed SSH identities on 09/10; recover 05/06/16; address fresh power faults and runtime-media stalls. Metis 0.1.0-403 is live; complete physical replacement-image validation |
| 3. Recovery coverage | In progress; cloud uploads blocked | External datastore restore and second LAN copy verified; 19 application DB dumps completed; scheduled application backup completed at 08:28 UTC; representative Gitea restore passed. Backblaze storage-cap failure established | Resolve the account storage cap within an agreed budget, complete eligible PVC coverage and Soteria restore, verify application-level recovery after tenant-2 storage repair, and complete approved local-only recovery coverage |
| 4. Capacity and placement | Partial; build scratch moved off runtime media | CI agent cap is one; all nine active project templates now exclude storage/quarantined workers. Disposable Longhorn scratch replaces inline emptyDir workspaces and build caches; Data Prepper image-layer storage corrected after a failed load check. Shared PostgreSQL/OpenSearch resource corrections deployed; Vault injectors spread. Longhorn uses least-effort balancing and one rebuild per node | Soteria 1303/1305/1306, Ananke 384, Ariadne 563, Atlasbot 442 bstein-dev-home 581, Data Prepper 2826 and Metis 405 passed with scratch-backed agents; verify the remaining project builds; restore compatible healthy worker capacity. Account for the measured 2-3.5 GiB storage-process footprint and calculate spare capacity for one worker loss |
| 5. Supported software | Planned; backups established | Installed K3s cohorts identified as 1.31.5 and 1.33.3; titan-db Ubuntu 24.10 is beyond support | Select a compatible supported baseline; stage OS/K3s/storage upgrades with rollback. Plan collation/index repair |
| 6. Maintainability and sustained health | In progress | Operator guide and rollback record published; Ananke stale-event loop fixed; updater failure reporting corrected; overlapping restart/eviction behavior removed | Complete the service recovery checklist, validate remaining automation boundaries, arrange controlled recovery checks, then observe 7-14 days |
@ -31,7 +33,7 @@ small and inspectable; a new recovery controller is not a prerequisite.
- [x] Correct Soteria release gate findings without disabling the gate; verify the local configuration scan.
- [x] Replace Soteria's daemonless Docker build with the existing pinned Kaniko builder; validate the pipeline.
- [x] Deploy and verify scoped Soteria state/credential permissions without altering stored policy or usage.
- [ ] Publish the corrected Soteria image through the normal pipeline and Flux.
- [x] Publish the corrected Soteria image through the normal pipeline and Flux: build 1305, image 0.1.0-128, Ready at 11:49 UTC.
- [ ] Verify an eligible backup and an isolated restore; retain Hermes exclusions.
- [x] Diagnose Monerod's host/attachment/filesystem error before selecting recovery.
- [x] Preserve a pre-recovery snapshot and remount the existing Monerod volume through Flux.
@ -83,14 +85,11 @@ replacement. No node was reflashed or rebooted during this work.
per-injection pending marker, serialized setup, propagated password failures,
removal of default passwordless grants, cleanup of applied bootstrap passwords,
and correct file permissions/native systemd links in both image-writing paths.
- [ ] Complete the normal Metis image publication and Flux rollout. Build 400
ended ABORTED after 7272.9 seconds; the later Ariadne event inspection established
that its elapsed-time rule requested this cancellation. Build 401 was stopped
deliberately to release its storage-node agent; queued build 402 was superseded
after placement/resource corrections. The live image is still 0.1.0-399 and
does not contain these source changes. The next
replacement image still needs a physical boot check before these changes are
considered proven on hardware.
- [x] Complete the normal Metis image publication and Flux rollout. Build 403
succeeded; controller and ARM/AMD sentinels are Ready on 0.1.0-403. Build 400's
earlier cancellation was traced to Ariadne's elapsed-time rule and repaired.
Queued/stalled older builds were stopped before the successful publication.
A physical replacement-image boot still needs its separate acceptance check.
- [x] Enable Ananke's existing password-backed host-action path. Atlas 886dbeee
adds 13 atlas-password mappings to the existing Vault CSI synchronizer; all
matched Vault, with native two-minute rotation enabled. No root passwords are
@ -192,6 +191,8 @@ queue first. Metis resource requests can be reverted independently in its repo.
| 12/13/19 | atlas/root passwords restored; root SSH keys preserved; password-backed sudo works after legacy grant retirement | Observe stability; use the corrected recovery image for future rebuilds |
| 20/21 | atlas sudo already worked; root passwords now match Vault | Continue workload memory diagnosis; Jetson-21 recorded a Data Prepper cgroup OOM |
| 04/11 | Fresh undervoltage messages continued during this audit, despite sufficient disk space | Onboard EXT5V samples were 4.82132 V (04) and 4.75968 V (11). Check delivered power/cables/peripherals; do not clear quarantine based on one sample |
| 08 | Image build drove runtime USB disk to 99.9% busy with 77% I/O pressure; pressure fell below 1% after the build stopped | Quarantined in Git; repair or replace runtime media and verify loaded operation before return |
| 12 | SD-backed runtime correlated with probe stalls and an engine-image restart at 12:02; 2.7 GiB memory remained available | New placement blocked without draining current services; inspect/migrate runtime media and perform loaded testing before return |
| 14 | Runtime USB flash; 39-45% iowait in the first sample, quiet later, no new kernel I/O error in the two-hour sample; stale RAM-log hooks corrected | Preserve data; investigate intermittent storage stalls and validate sustained operation before return; hardware failure is not proven by I/O pressure alone |
| 18 | Quarantine has reduced current I/O pressure; its 50 MiB RAM-log filesystem is still active | Keep quarantined until sustained storage checks pass; do not apply the inactive-RAM-log repair script to it |
| 05 | Answers LAN ping at .31; expected SSH, kubelet and common recovery ports refuse connections | Console/user-space service check; this is not evidence that the machine is powered off |

View File

@ -0,0 +1,415 @@
# Repair and settling cycles - 2026-10-04 UTC
This follows the user's request to repair, allow 10-15 minutes to settle, inspect
again, and repeat. The objective is stable operation with understandable native
configuration, not additional autonomous controllers. See
[operations](CLUSTER_OPERATIONS.md) for the everyday checks and
[the progress tracker](CLUSTER_STABILIZATION.md) for remaining work.
## Applied repairs
| Change | Evidence and effect | Revision |
| --- | --- | --- |
| Gitea placement and resources | Moved to Titan-22, reserved 2 GiB against a 1615 MiB peak, kept 3 GiB limit; pinned the exact running multiarch image and added HTTP readiness. Existing 25 GiB PVC remains healthy. Git and HTTPS health checks pass | 6ee8eb2e |
| Log maintenance and storage setup Jobs | Log guard requests 10m/16Mi instead of implicit limit-sized requests; all 11 eligible pods updated and Ready. Longhorn setup Jobs no longer expire and get recreated hourly; removed the Titan-11 pin. All three versioned Jobs completed | bb05ec9f |
| Hermes tenant-2 storage | Copied all three stopped opaque replicas locally before a guarded native reset of its one stranded engine process. Longhorn recovered all three replicas to RW; both consumer mounts respond to filesystem metadata checks. No case files were read and no inference job ran | 42969124, 56c959da; temporary permissions retired in 03089836 |
| Titan-08 containment | A build produced 99.9% runtime disk busy, 77% I/O pressure and delayed pod startup. Stopping the build lowered I/O pressure below 1%. Flux quarantine blocks new placement without draining existing services | 2857ca40 |
| ClamAV placement | Titan-22 CPU only; 2560 MiB reservation reflects a 2453 MiB reload peak, with the existing 3 GiB limit. Pinned the old running image, not the newer mutable tag. Existing PVC healthy; native PING/PONG checked | 64eb2601 |
| ClamAV readiness | Native bounded PING/PONG check gates service readiness; no liveness restart loop added | 570ef69b |
| Nextcloud resources/rollout | Kept Pi placement and all four existing PVCs. Memory request 512 MiB against a 394 MiB peak, limit still 3 GiB. Vault agent requests 25m against a 0.77m sampled peak. Single-writer rollout serializes replacement; status.php gates readiness | 5ba2211a |
| Nextcloud restart behavior | Existing apps survive restart; no routine removal of external-app keys or app reinstall. Missing-app downloads are bounded and extracted before replacement. Secret-setting output suppressed | 7a61fff8 |
| CI containment | Jenkins native stop/terminate/kill completed the stalled Data Prepper 2816 flow; native agent termination released its abandoned executor. Quiet mode was temporary; one-agent CI resumed and Metis 403 started on Titan-07 | Existing one-agent policy retained |
| Project CI gaps | All nine active projects now exclude storage/quarantined workers and Jetsons for ordinary CI. Ten pipeline definitions passed Jenkins validation. Soteria 1305 published successfully; Ananke 384 passed after its template correction | Atlas 21a4f677 and source commits below |
| Search disk capacity | OpenSearch data filesystem was 94% full and its policy endpoint returned 429 with create_index blocked. Expanded the existing PVC from 1024 to 1280 GiB; usage fell to 76%. A guarded check cleared the capacity block only below 80%. Normal 10:00, 10:30 and 11:00 UTC maintenance runs passed | ca8e3abc, cbf0c7f9, 53419f76; temporary check retired d6800b14 |
| Ingress isolation | Two Traefik replicas on separate NVMe control planes, native readiness and connection draining; same version pinned. Public .9 now announces from the ingress pool on control planes; private .50 keeps Local policy and authentication. Communications .4-.6 stay separate | f5ff949a, 689d6868, 4b087af1; temporary pool transition retired 785a01fc |
| Build scratch storage | Disposable single-replica ci-scratch class, 20/40 GiB workspaces, caches and Docker data off runtime USB. Soteria 1303 passed, mounted its Longhorn workspace and used the configured Go cache paths. Its agent, PVC and Longhorn scratch volume were removed automatically after completion. First observed runtime busy sample fell from about 97% to 19%; not a controlled benchmark | 2f158c6b, 70e1da36 and source commits below |
| L2 simplification | Removed unused FRR/BGP support containers: no BGP peers or advertisements exist. Eighteen eligible speakers are Ready in native L2 mode. Their sampled 94 MiB maximum has a 128 MiB request; five-second probes replace one-second probes | 63eafdd3 |
| Soteria image construction | Compile UI/ARM64 binary on the scratch PVC, then assemble only a 28 MB nonroot runtime image. Node/Go compiler layers no longer pass through Kaniko on runtime USB. Local UI/ARM build and image assembly passed; build 1305 succeeded in 839.351 seconds and published 0.1.0-128, Ready at 11:49 UTC | Soteria e9f24b1; explicit nonroot ownership/permission guard 4743b67 |
| Native observer readiness | Replaced repeated Python-process startup with an HTTP probe through the existing nginx proxy to a loopback-only health handler. The old listener was healthy while exec startup exceeded five seconds under build I/O. Public API paths and receipt checks remain unchanged | 7a81c3c4 |
| Backup scope for scratch | Added ci-scratch to Soteria's excluded storage classes and versioned its config rollout. Disposable build workspaces do not belong in application backup coverage; existing Hermes exclusions remain | 658ddadd |
| Persistent node quarantine | Existing cordons for Titan-04/05 are now declared in Git; Titan-06 will remain cordoned after reconnecting until recovery checks pass. SSA Merge and prune protection retain Node/storage identity. No drain or pod restart | d5fcbbfa |
| CI template repair | Removed Ananke's null `volumes:` mapping, which made the installed Jenkins Kubernetes plugin throw in `combineVolumes`. Pipeline validation and the actual plugin merge pass. Stopped only unallocated build 383; replacement 384 provisioned and completed successfully | Ananke 3d75387 |
| Retired repeating emergency cleanup | The completed Titan-24 root sweep expired hourly and Flux recreated it, running aggressive emergency settings again. Removed it from active resources; archived manifest remains | 59d206ba |
| Retained key bootstrap completion | Removed the one-hour TTL from the completed Metis SSH public-key bootstrap. It no longer authenticates to Vault and repeats bootstrap every hour; no key changed | b2a1761e |
| Titan-12 containment | Correlated probe stalls and an engine-image helper restart at 12:02 exposed SD-backed runtime latency. Cordoned through Git without draining existing services; every Longhorn manager UID stayed unchanged | 171979ea |
| Pod-log cleanup | Old code compared namespace_pod_UID names with bare UIDs and could delete active logs. New UID-aware helper fails closed on unreadable/empty inventories, checks file ages and rechecks activity. Ten synthetic filesystem regressions passed, 98.53% coverage; Titan-12 dry run preserved all 118 candidates | b4f0f6b4 |
| Native image lifecycle | Removed competing image pruning and direct containerd/image-import file deletion. Kubelet GC verified at 85%/80% on all 19 reachable nodes. The remaining helper handles host logs/package cache; its name remains for compatibility | a223b9fa |
| Node setup restart boundary | Removed automatic K3s restart after file-limit changes; definitions are reloaded and activate at planned maintenance. Initial and unchanged synthetic executions verified writes/reload without runtime restart. Setup-only helper rollout permits 25% replacement because settings persist outside the pod | baf6c3db |
| Node setup helper reservation | The idle node-nofile helper inherited 50m/96Mi namespace defaults, versus a 6.14 MiB measured memory peak. Explicit 5m/16Mi requests retain existing 500m/512Mi limits and free 45m/80Mi per updated node. Existing unit drop-ins were checked on every reachable node before rollout | 42d8f629 |
| Auth-helper placement | Two replicas each for Metis/Soteria, required hostname separation and quarantine exclusions. Soteria Vault-sidecar reservation reduced from 250m to 25m against a sampled peak below 1m. This also freed CPU needed by Titan-11's speaker | 82a0930d |
## Settling evidence
| Window | Outcome |
| --- | --- |
| 08:36:57-08:47:22, 10m25s | No application pod churn or restart increases; 57/57 active Flux Ready, 19/21 nodes Ready, all application controllers ready, attached volumes healthy. The known tenant-2 fault and pending log helper were still unresolved, so this was not whole-cluster acceptance |
| Follow-up after storage/log repairs | Caught Titan-08 runtime saturation under Data Prepper 2816. This was a failed load check, not a passed quiet window. It led to quarantine and capacity redistribution |
| 09:29:10-09:41:44, 12m35s | No restart increases; expected Metis publication and CI turnover. Probe delays and the still-failing OpenSearch maintenance job prevented full acceptance |
| From 09:54:27 | Further load revealed a Titan-19 speaker restart and lingering inline CI placement gaps. Corrected those gaps; did not count this as a clean window |
| 10:21:42-10:21:52 | Planned public VIP pool migration. One 15-second monitoring sample failed; the next succeeded. Flux fetched the complete two-wave change before the VIP moved, avoiding a Git/ingress dependency trap. All three shared services kept their fixed .9 address |
| 10:32-10:42 | Native L2 rollout. Titan-11's replacement could not fit with only 46m CPU unreserved, temporarily withdrawing its TURN announcement. Moving the Metis auth helper released 50m and restored it. HTTP ingress checks stayed successful. All 18 eligible speakers Ready afterward, each with one container and zero restarts |
| 10:42:15-10:58:16, 16m01s | No application replacements/restart increases, node readiness changes or storage failures. All active Kustomizations/HelmReleases and application controllers Ready. Soteria 1303 succeeded and cleaned up its scratch automatically. Ananke's next template exposed a null-volume merge error, repaired in source; CI then resumed with Soteria 1304. This window alone did not prove all queued templates could provision |
| From 10:57:30 | CI remained running after the template correction, without restart increases or readiness failures. The subsequent node-only quarantine change preserved every Longhorn manager UID |
| 11:06-11:19 | Image construction in Kaniko still saturated runtime USB: about 94% busy and 79% I/O wait, with Longhorn probe delays and two scheduled container-start timeouts. CI was paused and Soteria 1304 cancelled; this was a failed heavy-build acceptance window, not calm operation. Native observer readiness fixed separately. Both delayed scheduled jobs subsequently completed successfully |
| 11:35-11:49 | Late 0.1.0-127 publication failed nonroot execution; old Soteria replica stayed available. Corrected 0.1.0-128 replaced it successfully. This was not accepted as a quiet rollout window |
| 11:50-12:06 | Applications kept their desired replica targets, but Titan-12 had correlated probe stalls and one engine-image restart at 12:02. This was a failed quiet-window check and led to containment and the cleanup audit. CI agent/PVC turnover completed normally |
At 12:35:02 UTC the default scheduler preempted Hermes execution worker 0
on Titan-07 for the frontend rollout from bstein-dev-home build 581. The worker
had priority -10 (scavenger), while the incoming frontend had priority 0. Its
replacement was Ready again by 12:38. This was not a worker crash, failed
provider call or a recovery-controller restart. The scheduling event and pod UIDs
establish the cause; no workspaces or inference contents were inspected. The
node-helper request correction adds rollout headroom without changing worker
priority, restarting Keycloak or introducing another scheduling controller.
Raw evidence stays on the management host in
`/tmp/atlas-steady-state-20261004/`; summaries contain operational metadata only.
The committed `scripts/ops/cluster_settle_check.py` reproduces the readiness and
restart comparison. It does not inspect application contents or change anything.
## Acceptance boundary
A settling window passes the software containment check when the reachable
application controllers and requested data volumes stay available, no unexplained
application restart/replacement or node readiness transition occurs, and the
sampled service health paths continue responding. Scheduled CI pods and an
identified publication rollout are tracked separately. Probe warnings still
need investigation even when a sampled endpoint remains available.
This check does not declare quarantined hardware repaired. Stable service
placement with failed nodes excluded is a contained state, not restored spare
capacity. The user still needs a separate physical-repair and loaded-node
validation cycle before returning those nodes to placement. The platform and
backup follow-ups remain visible in the six-priority progress tracker.
## Validation and limits
Kustomize renders, client dry-runs and focused Flux diffs passed for every affected
stack. Twenty-four focused regressions passed across the two helpers, observer and
Nextcloud; the new read-only helper has 98.65%
statement coverage. Nextcloud preserves installed files and survives failed
downloads, and the observer distinguishes unavailable requested
storage, event lifetime counts, pod replacement and offline-node noise. Additional
checks prove that content-bearing API fields are omitted, failed API reads
cannot turn into partial green reports, and a serving old replica cannot hide a
failed rollout. Observer health requests cannot
access receipt state or open the existing API GET paths. Jenkins
accepted all ten updated pipeline definitions. The Soteria Docker build passed
on the management workstation; pipelines 1300/1301 had publication disabled.
Soteria 1303 passed using the new scratch volume. Publication-enabled 1304
was cancelled when its Kaniko toolchain build saturated runtime storage. Its
background builder nevertheless published 0.1.0-127 before agent cleanup. That
image had a root-owned mode-0700 binary because registry credential setup left
umask 077 active for compilation. Its rollout failed while the old replica kept
serving. Build 1305 succeeded in 839.351 seconds and published 0.1.0-128 with a
mode-0755 binary, which became Ready at 11:49 UTC. The prepared rollback commit was
superseded by this verified image before that rollback was pushed. Source 4743b67
additionally sets explicit executable permissions and nonroot ownership and confines umask
to registry-credential creation. Its local image validation and the normal
non-publishing pipeline 1306 passed (894.749 seconds including queue time). Metis 403
successfully published and its controller plus ARM/AMD sentinels are Ready on the new image.
Tenant-2 copies are root-only on Titan-13/15/17 under
`/mnt/astreae/atlas-recovery/tenant2-20261004/replica`. Copy completion and source
stability were verified; these are local preservation copies, not a completed
application restore drill. The one-time recovery Job, service account and scoped
exec permissions were pruned after success. Its archived manifest and script
remain outside the active Longhorn kustomization.
Titan-05/06 remain offline. Titan-04/11 power delivery, Titan-08/12/14/18 runtime
storage, changed identities on 09/10 and unreachable 16 still need their recorded
hardware/identity work. Backblaze still rejects uploads at its account storage
cap. A quiet 10-15 minute interval is not a 7-14 day soak, N+1 capacity proof, a
power-failure drill or proof that these physical faults are repaired.
The final runtime-only ARM64 image build passed locally: 27,996,620 bytes,
static aarch64 binary, UID/GID 65532 and the same /soteria entrypoint. This
measurement is image assembly on the management workstation, not the cluster
publication duration.
Public-network samples at 11:54:02 UTC (two requests) and 12:25:49 UTC
(three requests) timed out before connection, while other URLs in each sample
succeeded. Subsequent samples passed. There was no correlated node flap, ingress restart or unhealthy endpoint
in the retained cluster metadata. The management workstation is outside the
cluster LAN; direct private-VIP checks from that host cannot establish LAN
reachability. Separate TLS-verified checks through Titan-jh began at 11:56:46
UTC with explicit resolution to 192.168.22.9. The earlier failures did not capture DNS/TCP/TLS phase timing, so their exact
segment is unknown. The later 12:33:28 SSO failure had zero completed DNS, TCP
and TLS timing, with a five-second connect timeout. It failed before name
resolution completed. The timed public series at 12:28:58-12:41:43 had 207/208
successful requests. The concurrent LAN series at 12:29:00-12:40:36 passed
96/96, using the LAN VIP and certificate verification. The earlier LAN series
also passed 96/96. This does not prove uninterrupted public availability or the
Windows laptop route.
## Remaining operational warnings
- The platform-quality Pushgateway on Titan-19 still records occasional
five-second probe timeouts (including 12:27:34 UTC). It remained Ready and did
not restart during that observation. No probe threshold was relaxed.
- Titan-12's Longhorn manager and several Hermes readiness probes stalled
together around 12:02; one engine-image helper restarted. Application and
volume controllers recovered without data-volume failure, but SD runtime
latency remains unresolved. The node is cordoned; no application probe
threshold was loosened to hide the incident.
- Titan-07 had a Longhorn CSI probe warning at 12:56:19 during normal CI.
There was no CSI restart or requested-volume failure in the sampled window.
The runtime USB constraint remains; a transient probe failure is not proof of
an application outage, nor proof that the device has recovered.
- Titan-23's old `iptables-block` host-network pod repeatedly reports
`DNSConfigForming`. The DaemonSet is not currently declared in this repository.
This is inherited resolver-list truncation, not an observed service lookup
failure. Reconcile ownership and DNS configuration in a deliberate follow-up;
do not remove the firewall rules just to remove a warning.
- `hermes/hermes-chat-router-release` ImagePolicy has no matching validated
release tag. Its ImageRepository scan succeeds and the running digest remains
pinned. No tag was invented and the release filter was not broadened. Other
image policies, repositories and automation report Ready.
- Two Hermes triage critical alerts refer to Pegasus builds 313/326. Jenkins
confirms a later successful build 336; build 338 was subsequently aborted.
Ariadne's gauge refresh only suppresses older incidents when `lastBuild` is
SUCCESS, so it can revive historical failed-incident alerts after an aborted
build. This is a follow-up incident-state reporting defect, not evidence of a
new Pegasus service outage. Incident history and alert severity were preserved.
- Claude quota credential-expiry and missing Fable-weekly-quota warnings remain.
They concern the separate quota telemetry credential. No provider call, login,
inference request or credential change was made during these checks.
The scratch-volume `VolumeFailedDelete` events at 10:47 and 11:51 were transient detach
retries: the agent, PVC, PV and Longhorn volume were subsequently confirmed gone.
Initial `FailedScheduling` events for the next CI pod ended once its dynamic PVC
bound; it became Ready. Neither is an outstanding failed service. The 11:00
Cassandra snapshot warning concerned retirement: that object was subsequently
gone and all 18 current Cassandra snapshot objects were Ready.
## Source pipeline revisions
Scratch-volume changes are explicit in each project's Jenkinsfile; the IaC
pipeline has its existing mirrored file. No shared-library framework was added.
| Project | Scratch revision |
| --- | --- |
| Ariadne | cd69ae6; serial CPU shares and Docker-client reservation adjusted in fc39068, burst limits retained |
| bstein-dev-home | a082156 |
| Atlasbot | 8520f34 |
| Ananke | e2c04a7; null-volume merge correction 3d75387 |
| Metis | 3f30c30 |
| Pegasus | ef9e797 |
| Soteria | 8b6195c; scratch-backed release compilation e9f24b1; explicit permissions 4743b67 |
| IaC / Data Prepper | 70e1da36 |
A brief BFQ scheduler trial on Titan-07 was inconclusive and was reverted to its
original mq-deadline setting at 10:27 UTC. No scheduler tuning service was left
installed. Disk-pressure improvements should be checked under normal builds;
the replica/PVC mount and actual cache paths were directly verified. Kaniko
image construction still writes its container root filesystem on runtime media:
a later two-minute sample reached 82% disk busy and 51% I/O wait, while service
readiness and HTTPS checks stayed successful. The scratch move is not a complete
replacement for repairing runtime storage; build concurrency remains one.
The latest checked datastore backup completed at 12:01:08 UTC and its second
LAN copy completed at 12:21:34 UTC; the application database backup completed at 08:28:58 UTC. All
three native services reported success/status 0. These checks did not replace
the separately recorded restore drills or resolve the Backblaze account cap.
Certificate warnings on Titan-0a/0c refer to expiry on January 15, 2027. They
need a planned rolling K3s restart/renewal; they are not an immediate expired
certificate failure. No control plane was restarted during this settling cycle.
## Rollback
Revert the relevant focused commit, push the tracked branch and reconcile that
stack. Capacity reverts can make CI unschedulable again; do not restore the older
multi-agent cap or place builds on saturated storage merely to clear the queue.
Do not blindly revert Titan-08 quarantine until loaded I/O is verified. Preserve
Node role/storage labels and the SSA Merge/prune annotations on Node manifests.
Nextcloud resource and startup changes can be reverted independently; restoring
the old startup script reintroduces routine app removal/downloads. ClamAV/Gitea
placement reverts cause a single-writer volume handoff and need available Pi RAM.
Do not replay the retired tenant-2 engine reset. Its exact UID/state guards and
one-attempt markers intentionally prevent reuse against a different incident.
Use the preserved replicas and a fresh storage diagnosis for a future failure.
The OpenSearch capacity expansion cannot be undone by shrinking its PVC. Keep
1280 GiB in Git even when reverting other logging changes. The guarded temporary
block-release Job was retired; do not replay it without rechecking actual disk
space and the current block cause.
Keep ingress-pool and communication-pool ranges non-overlapping. Moving the VIP
back requires a staged Flux transition; blindly reverting the cleanup commit is
not a rollback procedure. Reverting only Traefik placement is independent, but
first keep at least one local endpoint on a node allowed by lan-ai-adv. Retain
TLS, source restrictions and Local traffic policy. FRR can be re-enabled through
the HelmRelease if BGP is deliberately introduced, with its resource impact
accounted for. Restoring 05/06 also requires removing their speaker exclusion.
For CI, revert the per-project workspace changes before removing ci-scratch,
and let existing agents/PVCs finish. Never reuse the Delete scratch class for
persistent data. The standard astreae class and all existing retained application
volumes were left unchanged. CPU-request changes preserve the existing limits;
verify scheduling and actual service latency before raising build concurrency.
Emergency cleanup rollback: leave `titan-24-rootfs-sweep-job.yaml` outside the
active kustomization. Re-adding it enables its aggressive emergency settings;
that is a fresh repair action, not routine cleanup. Keep completed bootstrap
Jobs without TTL while Flux manages them, or retire them explicitly. A cluster
inventory found all other remaining Flux Jobs with TTL suspended.
The first Playwright pull for bstein-dev-home 581 took about 12 minutes on
Titan-07; all six agent containers then became Ready. Its runtime lives on the
58 GiB USB device, separate from its 29 GiB SD root. Cold image extraction still
saturates that device despite scratch being offloaded. The custom image pruner
was triggered by root usage even when runtime images were on another filesystem.
Native kubelet image GC now owns that lifecycle; this is not a claim that the
USB device is adequate for sustained image construction. Keep CI serialized.
The sweeper rollout initially stalled behind terminating pods on offline 05/06.
Its budget was temporarily raised to three (two offline plus one replacement),
and startup readiness now waits for the initial cleanup to finish. Native
cleanup removed those old pods; the normal budget was restored to one in fd76b043.
Kubernetes recommends letting kubelet own image garbage collection rather than
running a competing external collector: [garbage collection documentation](https://kubernetes.io/docs/concepts/architecture/garbage-collection/).
At 12:26 UTC all 19 replacement cleanup helpers were Ready with zero
restarts. The later CI completion snapshot below supersedes the running/queued
state at that time. Cold-image and publication measurements above remain
separate from application availability.
The LAN bastion checks passed all 96 requests from 11:56:46 to 12:08:22 UTC
against the four explicit private-VIP health URLs with certificate verification.
The first four sample rows were inspected in tool output; the following twenty
rows are retained in `lan-health-observed.json`. This is not a laptop-side route
test. Private inference capabilities also returned the expected unauthenticated
401 through 192.168.22.50, with TLS verification successful; no inference ran.
At 12:46 UTC, Data Prepper 2820 had also completed SUCCESS. Ariadne 563,
Ananke 384, Atlasbot 442, bstein-dev-home 581 and Soteria 1306 passed on the
corrected CI setup. Metis 405 was provisioning; Pegasus 339 and IaC 4232 were
still queued. Earlier cancelled builds are retained as ABORTED, not counted as
new infrastructure failures.
The repository-suggested combination `--server-side --dry-run=client` is
rejected by the installed kubectl. Validation used
`kubectl apply --dry-run=client -k services/maintenance`, plus a successful
render and focused Flux diff. A diff exit status of one indicated the intended changes.
Do not restore the removed node-nofile automatic K3s restart as part of a
resource rollback. Changing its small reservation is independent of host-runtime
restart policy. The setup helper permits 25% unavailable because its host
settings persist when it exits; it is not an application-serving DaemonSet.
At 12:54 UTC Metis 405 completed SUCCESS. Its new publication entered the normal
Flux image-automation path. Data Prepper 2821 was running; Pegasus 339 and IaC
4232 remained queued. At 12:51:43 all 18 eligible node-nofile helpers were
updated/Ready; all old misplaced helper pods were removed by their native
controller. Titan-07 now reserves 5m/16Mi for this helper, freeing 45m/80Mi.
No application pod restart was caused by that rollout.
## Follow-up under scheduled and build load
The 12:51:59 window did not pass: at 13:00:00 the scheduled OpenSearch tuner
preempted worker 0 again. That job was restricted to Pi 5 workers, leaving only
Titan-07 eligible during hardware quarantine. Commit eb1eb369 gives it the
native maintenance-batch PriorityClass (value 0, preemptionPolicy Never), permits
Pi 4 placement while preferring Pi 5, removes its unused service-account token,
and pins the exact observed Python image. This is a targeted batch policy, not
a change to worker model, authentication, provider access or job state.
The read-only placement Job showed the admitted priority and Never policy,
but its first attempt exceeded its 300-second deadline during renewed runtime
I/O saturation. It is not recorded as a passed service check. The sampler now
includes Normal scheduler Preempted events and pods marked for deletion even
when Ready is still true; d85642ad adds that regression coverage (11 tests,
98.65% statement coverage). Only operational fields are retained.
Data Prepper 2822 passed, but Kaniko still unpacked its large base image into
the node's container filesystem. Titan-07 reached 99.7% disk busy and 59% I/O
pressure. Completing a build did not make this an acceptable service-load test.
CI was temporarily quieted and the next old-template Data Prepper build 2823
was stopped. Commit b51cbc2f switches this pipeline to the already deployed
Docker-in-Docker image, stores its layers on a 40 GiB disposable workspace PVC,
keeps ARM64 output explicit and binds the daemon to pod loopback. Registry
credentials enter via stdin and are removed with the temporary Docker config.
Jenkins accepted the full pipeline; a synthetic CLI test verified ARM64,
intended tags, credential stdin and cleanup. The actual replacement build and
its resource behavior must still be verified before closing this cycle.
The second placement check (4d9fe145) completed at 13:16:54 UTC. Its container
ran from 13:16:46 to 13:16:50 and exited 0 on Titan-07, with admitted
preemptionPolicy Never. It performed only a cluster-health GET and printed a
fixed success line, without reading index contents. The first failed attempt
remains part of the evidence; success after I/O subsided does not erase it.
Data Prepper 2824 exposed an implementation error in the new builder: the
Docker entrypoint added tcp://0.0.0.0:2375 while the supplied flags also requested
127.0.0.1:2375. Dockerd exited 1 with a listener conflict, and Jenkins's native
container-termination reaper removed the agent. This was not another automatic
Ariadne cancellation. b12991cb explicitly supplies `dockerd` as the first
entrypoint argument, retaining the image's setup while preventing default
listener insertion. Registry metadata verified the pinned image is ARM64 and
uses dockerd-entrypoint.sh. A successful corrected invocation remains required.
Pegasus 339 completed with a quality-gate failure; its reported aggregate
coverage was 98.855%, so that number alone does not establish which gate failed.
The regular 13:30 OpenSearch tuner completed at 13:30:10 on Titan-07 with
priority 0 and preemptionPolicy Never. The Hermes worker remained running.
This verifies the normal scheduled path, in addition to the read-only check.
The queued Data Prepper 2825 attempt loaded the old entrypoint configuration;
it was stopped while still waiting for an agent so the corrected attempt can
run. That deliberate cancellation is separate from 2824's startup defect.
Data Prepper 2826 passed after the entrypoint correction. Recorded duration was
406.619 seconds including its queue/allocation time, so this is not a pure build
benchmark. Both intended tags published digest
`sha256:62f00ee7d64c3f85446c282c11c387eb6f73d73ad9d6bc5a545ab9849e8f0bd0`.
Live Docker reported 27.5.1, aarch64, overlay2 and /var/lib/docker on the 39.1 GiB
Longhorn scratch filesystem. At 13:43:50 the two-minute runtime-disk busy sample
was 11.54%, versus 99.69% during the old builder; I/O pressure was 13.40%.
This is observational evidence under different build stages, not a controlled
throughput benchmark or a hardware repair.
Pegasus 339's sole gate-summary issue was Jenkinsfile length (516 > 500), caused
by the earlier pipeline additions. Sonar was OK. Pegasus 5fbc6d0 extracts the
existing report step into scripts/report_quality_gate.sh without weakening the
gate. Jenkinsfile is now 472 lines; its own production-file LOC check, Jenkins
pipeline validation, shell syntax and a failed-dependency/report regression
passed. A subsequent full project rerun remains separately observable.
## Final settling observation
From 13:16:24 to 13:45:59 UTC (29 minutes 35 seconds), all 145 application
controllers met their configured targets. All 57 active Flux Kustomizations
and 21 HelmReleases were Ready. There were no container restart increases,
same-name application pod replacements, new preemptions, node Ready transitions,
unhealthy attached volumes or unavailable requested volumes. Ordinary CI agents
continued to turn over. Titan-05/06 remained offline with 23 unavailable pods
on those nodes; those are not counted as recovered hardware.
The four LAN health checks passed 96/96 requests from 13:36:54 to 13:48:33 UTC,
using explicit 192.168.22.9 resolution and TLS verification from Titan-jh.
Public-path checks passed 350/352 from 13:16:56 to 13:38:41 UTC. The two timeouts
at 13:26:56 occurred before DNS resolution completed on the management host;
TCP and TLS had not started. Their cause is not established as a cluster outage.
These checks do not verify the Windows laptop's own route.
The completed read-only placement Job is now retired from the logging
kustomization. Its manifest remains available as an explicit validation example;
it will not recur automatically. The normal scheduled maintenance Job has already
passed with preemption disabled. Jenkins remains enabled with one agent. The
waiting IaC and Pegasus jobs retain eligible templates; their waiting state alone
does not establish a missing agent definition or failed service.
At 13:51 UTC the active agent was executing IaC quality-gate main build 84;
Jenkins was not quieted. Data Prepper 2827 and Metis 406 had also succeeded.
Pegasus 340 was queued with its corrected pipeline; its full rerun was not yet
verified. The final logging render and client dry-run passed; Flux diff showed
only deletion of the completed validation Job.
This closes the active application-churn repair cycle, not the physical repair
or backup programme. Native probe delays on weak media, offline/power faults,
limited spare worker capacity and Backblaze's storage cap remain tracked in
CLUSTER_STABILIZATION.md. No probe was weakened to hide those findings. A
7-14-day observation period and controlled restore/failure drills remain open.