atlas-iac/docs/STABILITY_CYCLES_20261004.md

34 KiB

Repair and settling cycles - 2026-10-04 UTC

This follows the user's request to repair, allow 10-15 minutes to settle, inspect again, and repeat. The objective is stable operation with understandable native configuration, not additional autonomous controllers. See operations for the everyday checks and the progress tracker for remaining work.

Applied repairs

Change Evidence and effect Revision
Gitea placement and resources Moved to Titan-22, reserved 2 GiB against a 1615 MiB peak, kept 3 GiB limit; pinned the exact running multiarch image and added HTTP readiness. Existing 25 GiB PVC remains healthy. Git and HTTPS health checks pass 6ee8eb2e
Log maintenance and storage setup Jobs Log guard requests 10m/16Mi instead of implicit limit-sized requests; all 11 eligible pods updated and Ready. Longhorn setup Jobs no longer expire and get recreated hourly; removed the Titan-11 pin. All three versioned Jobs completed bb05ec9f
Hermes tenant-2 storage Copied all three stopped opaque replicas locally before a guarded native reset of its one stranded engine process. Longhorn recovered all three replicas to RW; both consumer mounts respond to filesystem metadata checks. No case files were read and no inference job ran 42969124, 56c959da; temporary permissions retired in 03089836
Titan-08 containment A build produced 99.9% runtime disk busy, 77% I/O pressure and delayed pod startup. Stopping the build lowered I/O pressure below 1%. Flux quarantine blocks new placement without draining existing services 2857ca40
ClamAV placement Titan-22 CPU only; 2560 MiB reservation reflects a 2453 MiB reload peak, with the existing 3 GiB limit. Pinned the old running image, not the newer mutable tag. Existing PVC healthy; native PING/PONG checked 64eb2601
ClamAV readiness Native bounded PING/PONG check gates service readiness; no liveness restart loop added 570ef69b
Nextcloud resources/rollout Kept Pi placement and all four existing PVCs. Memory request 512 MiB against a 394 MiB peak, limit still 3 GiB. Vault agent requests 25m against a 0.77m sampled peak. Single-writer rollout serializes replacement; status.php gates readiness 5ba2211a
Nextcloud restart behavior Existing apps survive restart; no routine removal of external-app keys or app reinstall. Missing-app downloads are bounded and extracted before replacement. Secret-setting output suppressed 7a61fff8
CI containment Jenkins native stop/terminate/kill completed the stalled Data Prepper 2816 flow; native agent termination released its abandoned executor. Quiet mode was temporary; one-agent CI resumed and Metis 403 started on Titan-07 Existing one-agent policy retained
Project CI gaps All nine active projects now exclude storage/quarantined workers and Jetsons for ordinary CI. Ten pipeline definitions passed Jenkins validation. Soteria 1305 published successfully; Ananke 384 passed after its template correction Atlas 21a4f677 and source commits below
Search disk capacity OpenSearch data filesystem was 94% full and its policy endpoint returned 429 with create_index blocked. Expanded the existing PVC from 1024 to 1280 GiB; usage fell to 76%. A guarded check cleared the capacity block only below 80%. Normal 10:00, 10:30 and 11:00 UTC maintenance runs passed ca8e3abc, cbf0c7f9, 53419f76; temporary check retired d6800b14
Ingress isolation Two Traefik replicas on separate NVMe control planes, native readiness and connection draining; same version pinned. Public .9 now announces from the ingress pool on control planes; private .50 keeps Local policy and authentication. Communications .4-.6 stay separate f5ff949a, 689d6868, 4b087af1; temporary pool transition retired 785a01fc
Build scratch storage Disposable single-replica ci-scratch class, 20/40 GiB workspaces, caches and Docker data off runtime USB. Soteria 1303 passed, mounted its Longhorn workspace and used the configured Go cache paths. Its agent, PVC and Longhorn scratch volume were removed automatically after completion. First observed runtime busy sample fell from about 97% to 19%; not a controlled benchmark 2f158c6b, 70e1da36 and source commits below
L2 simplification Removed unused FRR/BGP support containers: no BGP peers or advertisements exist. Eighteen eligible speakers are Ready in native L2 mode. Their sampled 94 MiB maximum has a 128 MiB request; five-second probes replace one-second probes 63eafdd3
Soteria image construction Compile UI/ARM64 binary on the scratch PVC, then assemble only a 28 MB nonroot runtime image. Node/Go compiler layers no longer pass through Kaniko on runtime USB. Local UI/ARM build and image assembly passed; build 1305 succeeded in 839.351 seconds and published 0.1.0-128, Ready at 11:49 UTC Soteria e9f24b1; explicit nonroot ownership/permission guard 4743b67
Native observer readiness Replaced repeated Python-process startup with an HTTP probe through the existing nginx proxy to a loopback-only health handler. The old listener was healthy while exec startup exceeded five seconds under build I/O. Public API paths and receipt checks remain unchanged 7a81c3c4
Backup scope for scratch Added ci-scratch to Soteria's excluded storage classes and versioned its config rollout. Disposable build workspaces do not belong in application backup coverage; existing Hermes exclusions remain 658ddadd
Persistent node quarantine Existing cordons for Titan-04/05 are now declared in Git; Titan-06 will remain cordoned after reconnecting until recovery checks pass. SSA Merge and prune protection retain Node/storage identity. No drain or pod restart d5fcbbfa
CI template repair Removed Ananke's null volumes: mapping, which made the installed Jenkins Kubernetes plugin throw in combineVolumes. Pipeline validation and the actual plugin merge pass. Stopped only unallocated build 383; replacement 384 provisioned and completed successfully Ananke 3d75387
Retired repeating emergency cleanup The completed Titan-24 root sweep expired hourly and Flux recreated it, running aggressive emergency settings again. Removed it from active resources; archived manifest remains 59d206ba
Retained key bootstrap completion Removed the one-hour TTL from the completed Metis SSH public-key bootstrap. It no longer authenticates to Vault and repeats bootstrap every hour; no key changed b2a1761e
Titan-12 containment Correlated probe stalls and an engine-image helper restart at 12:02 exposed SD-backed runtime latency. Cordoned through Git without draining existing services; every Longhorn manager UID stayed unchanged 171979ea
Pod-log cleanup Old code compared namespace_pod_UID names with bare UIDs and could delete active logs. New UID-aware helper fails closed on unreadable/empty inventories, checks file ages and rechecks activity. Ten synthetic filesystem regressions passed, 98.53% coverage; Titan-12 dry run preserved all 118 candidates b4f0f6b4
Native image lifecycle Removed competing image pruning and direct containerd/image-import file deletion. Kubelet GC verified at 85%/80% on all 19 reachable nodes. The remaining helper handles host logs/package cache; its name remains for compatibility a223b9fa
Node setup restart boundary Removed automatic K3s restart after file-limit changes; definitions are reloaded and activate at planned maintenance. Initial and unchanged synthetic executions verified writes/reload without runtime restart. Setup-only helper rollout permits 25% replacement because settings persist outside the pod baf6c3db
Node setup helper reservation The idle node-nofile helper inherited 50m/96Mi namespace defaults, versus a 6.14 MiB measured memory peak. Explicit 5m/16Mi requests retain existing 500m/512Mi limits and free 45m/80Mi per updated node. Existing unit drop-ins were checked on every reachable node before rollout 42d8f629
Auth-helper placement Two replicas each for Metis/Soteria, required hostname separation and quarantine exclusions. Soteria Vault-sidecar reservation reduced from 250m to 25m against a sampled peak below 1m. This also freed CPU needed by Titan-11's speaker 82a0930d

Settling evidence

Window Outcome
08:36:57-08:47:22, 10m25s No application pod churn or restart increases; 57/57 active Flux Ready, 19/21 nodes Ready, all application controllers ready, attached volumes healthy. The known tenant-2 fault and pending log helper were still unresolved, so this was not whole-cluster acceptance
Follow-up after storage/log repairs Caught Titan-08 runtime saturation under Data Prepper 2816. This was a failed load check, not a passed quiet window. It led to quarantine and capacity redistribution
09:29:10-09:41:44, 12m35s No restart increases; expected Metis publication and CI turnover. Probe delays and the still-failing OpenSearch maintenance job prevented full acceptance
From 09:54:27 Further load revealed a Titan-19 speaker restart and lingering inline CI placement gaps. Corrected those gaps; did not count this as a clean window
10:21:42-10:21:52 Planned public VIP pool migration. One 15-second monitoring sample failed; the next succeeded. Flux fetched the complete two-wave change before the VIP moved, avoiding a Git/ingress dependency trap. All three shared services kept their fixed .9 address
10:32-10:42 Native L2 rollout. Titan-11's replacement could not fit with only 46m CPU unreserved, temporarily withdrawing its TURN announcement. Moving the Metis auth helper released 50m and restored it. HTTP ingress checks stayed successful. All 18 eligible speakers Ready afterward, each with one container and zero restarts
10:42:15-10:58:16, 16m01s No application replacements/restart increases, node readiness changes or storage failures. All active Kustomizations/HelmReleases and application controllers Ready. Soteria 1303 succeeded and cleaned up its scratch automatically. Ananke's next template exposed a null-volume merge error, repaired in source; CI then resumed with Soteria 1304. This window alone did not prove all queued templates could provision
From 10:57:30 CI remained running after the template correction, without restart increases or readiness failures. The subsequent node-only quarantine change preserved every Longhorn manager UID
11:06-11:19 Image construction in Kaniko still saturated runtime USB: about 94% busy and 79% I/O wait, with Longhorn probe delays and two scheduled container-start timeouts. CI was paused and Soteria 1304 cancelled; this was a failed heavy-build acceptance window, not calm operation. Native observer readiness fixed separately. Both delayed scheduled jobs subsequently completed successfully
11:35-11:49 Late 0.1.0-127 publication failed nonroot execution; old Soteria replica stayed available. Corrected 0.1.0-128 replaced it successfully. This was not accepted as a quiet rollout window
11:50-12:06 Applications kept their desired replica targets, but Titan-12 had correlated probe stalls and one engine-image restart at 12:02. This was a failed quiet-window check and led to containment and the cleanup audit. CI agent/PVC turnover completed normally

At 12:35:02 UTC the default scheduler preempted Hermes execution worker 0 on Titan-07 for the frontend rollout from bstein-dev-home build 581. The worker had priority -10 (scavenger), while the incoming frontend had priority 0. Its replacement was Ready again by 12:38. This was not a worker crash, failed provider call or a recovery-controller restart. The scheduling event and pod UIDs establish the cause; no workspaces or inference contents were inspected. The node-helper request correction adds rollout headroom without changing worker priority, restarting Keycloak or introducing another scheduling controller.

Raw evidence stays on the management host in /tmp/atlas-steady-state-20261004/; summaries contain operational metadata only. The committed scripts/ops/cluster_settle_check.py reproduces the readiness and restart comparison. It does not inspect application contents or change anything.

Acceptance boundary

A settling window passes the software containment check when the reachable application controllers and requested data volumes stay available, no unexplained application restart/replacement or node readiness transition occurs, and the sampled service health paths continue responding. Scheduled CI pods and an identified publication rollout are tracked separately. Probe warnings still need investigation even when a sampled endpoint remains available.

This check does not declare quarantined hardware repaired. Stable service placement with failed nodes excluded is a contained state, not restored spare capacity. The user still needs a separate physical-repair and loaded-node validation cycle before returning those nodes to placement. The platform and backup follow-ups remain visible in the six-priority progress tracker.

Validation and limits

Kustomize renders, client dry-runs and focused Flux diffs passed for every affected stack. Twenty-four focused regressions passed across the two helpers, observer and Nextcloud; the new read-only helper has 98.65% statement coverage. Nextcloud preserves installed files and survives failed downloads, and the observer distinguishes unavailable requested storage, event lifetime counts, pod replacement and offline-node noise. Additional checks prove that content-bearing API fields are omitted, failed API reads cannot turn into partial green reports, and a serving old replica cannot hide a failed rollout. Observer health requests cannot access receipt state or open the existing API GET paths. Jenkins accepted all ten updated pipeline definitions. The Soteria Docker build passed on the management workstation; pipelines 1300/1301 had publication disabled. Soteria 1303 passed using the new scratch volume. Publication-enabled 1304 was cancelled when its Kaniko toolchain build saturated runtime storage. Its background builder nevertheless published 0.1.0-127 before agent cleanup. That image had a root-owned mode-0700 binary because registry credential setup left umask 077 active for compilation. Its rollout failed while the old replica kept serving. Build 1305 succeeded in 839.351 seconds and published 0.1.0-128 with a mode-0755 binary, which became Ready at 11:49 UTC. The prepared rollback commit was superseded by this verified image before that rollback was pushed. Source 4743b67 additionally sets explicit executable permissions and nonroot ownership and confines umask to registry-credential creation. Its local image validation and the normal non-publishing pipeline 1306 passed (894.749 seconds including queue time). Metis 403 successfully published and its controller plus ARM/AMD sentinels are Ready on the new image.

Tenant-2 copies are root-only on Titan-13/15/17 under /mnt/astreae/atlas-recovery/tenant2-20261004/replica. Copy completion and source stability were verified; these are local preservation copies, not a completed application restore drill. The one-time recovery Job, service account and scoped exec permissions were pruned after success. Its archived manifest and script remain outside the active Longhorn kustomization.

Titan-05/06 remain offline. Titan-04/11 power delivery, Titan-08/12/14/18 runtime storage, changed identities on 09/10 and unreachable 16 still need their recorded hardware/identity work. Backblaze still rejects uploads at its account storage cap. A quiet 10-15 minute interval is not a 7-14 day soak, N+1 capacity proof, a power-failure drill or proof that these physical faults are repaired.

The final runtime-only ARM64 image build passed locally: 27,996,620 bytes, static aarch64 binary, UID/GID 65532 and the same /soteria entrypoint. This measurement is image assembly on the management workstation, not the cluster publication duration.

Public-network samples at 11:54:02 UTC (two requests) and 12:25:49 UTC (three requests) timed out before connection, while other URLs in each sample succeeded. Subsequent samples passed. There was no correlated node flap, ingress restart or unhealthy endpoint in the retained cluster metadata. The management workstation is outside the cluster LAN; direct private-VIP checks from that host cannot establish LAN reachability. Separate TLS-verified checks through Titan-jh began at 11:56:46 UTC with explicit resolution to 192.168.22.9. The earlier failures did not capture DNS/TCP/TLS phase timing, so their exact segment is unknown. The later 12:33:28 SSO failure had zero completed DNS, TCP and TLS timing, with a five-second connect timeout. It failed before name resolution completed. The timed public series at 12:28:58-12:41:43 had 207/208 successful requests. The concurrent LAN series at 12:29:00-12:40:36 passed 96/96, using the LAN VIP and certificate verification. The earlier LAN series also passed 96/96. This does not prove uninterrupted public availability or the Windows laptop route.

Remaining operational warnings

  • The platform-quality Pushgateway on Titan-19 still records occasional five-second probe timeouts (including 12:27:34 UTC). It remained Ready and did not restart during that observation. No probe threshold was relaxed.
  • Titan-12's Longhorn manager and several Hermes readiness probes stalled together around 12:02; one engine-image helper restarted. Application and volume controllers recovered without data-volume failure, but SD runtime latency remains unresolved. The node is cordoned; no application probe threshold was loosened to hide the incident.
  • Titan-07 had a Longhorn CSI probe warning at 12:56:19 during normal CI. There was no CSI restart or requested-volume failure in the sampled window. The runtime USB constraint remains; a transient probe failure is not proof of an application outage, nor proof that the device has recovered.
  • Titan-23's old iptables-block host-network pod repeatedly reports DNSConfigForming. The DaemonSet is not currently declared in this repository. This is inherited resolver-list truncation, not an observed service lookup failure. Reconcile ownership and DNS configuration in a deliberate follow-up; do not remove the firewall rules just to remove a warning.
  • hermes/hermes-chat-router-release ImagePolicy has no matching validated release tag. Its ImageRepository scan succeeds and the running digest remains pinned. No tag was invented and the release filter was not broadened. Other image policies, repositories and automation report Ready.
  • Two Hermes triage critical alerts refer to Pegasus builds 313/326. Jenkins confirms a later successful build 336; build 338 was subsequently aborted. Ariadne's gauge refresh only suppresses older incidents when lastBuild is SUCCESS, so it can revive historical failed-incident alerts after an aborted build. This is a follow-up incident-state reporting defect, not evidence of a new Pegasus service outage. Incident history and alert severity were preserved.
  • Claude quota credential-expiry and missing Fable-weekly-quota warnings remain. They concern the separate quota telemetry credential. No provider call, login, inference request or credential change was made during these checks.

The scratch-volume VolumeFailedDelete events at 10:47 and 11:51 were transient detach retries: the agent, PVC, PV and Longhorn volume were subsequently confirmed gone. Initial FailedScheduling events for the next CI pod ended once its dynamic PVC bound; it became Ready. Neither is an outstanding failed service. The 11:00 Cassandra snapshot warning concerned retirement: that object was subsequently gone and all 18 current Cassandra snapshot objects were Ready.

Source pipeline revisions

Scratch-volume changes are explicit in each project's Jenkinsfile; the IaC pipeline has its existing mirrored file. No shared-library framework was added.

Project Scratch revision
Ariadne cd69ae6; serial CPU shares and Docker-client reservation adjusted in fc39068, burst limits retained
bstein-dev-home a082156
Atlasbot 8520f34
Ananke e2c04a7; null-volume merge correction 3d75387
Metis 3f30c30
Pegasus ef9e797
Soteria 8b6195c; scratch-backed release compilation e9f24b1; explicit permissions 4743b67
IaC / Data Prepper 70e1da36

A brief BFQ scheduler trial on Titan-07 was inconclusive and was reverted to its original mq-deadline setting at 10:27 UTC. No scheduler tuning service was left installed. Disk-pressure improvements should be checked under normal builds; the replica/PVC mount and actual cache paths were directly verified. Kaniko image construction still writes its container root filesystem on runtime media: a later two-minute sample reached 82% disk busy and 51% I/O wait, while service readiness and HTTPS checks stayed successful. The scratch move is not a complete replacement for repairing runtime storage; build concurrency remains one.

The latest checked datastore backup completed at 12:01:08 UTC and its second LAN copy completed at 12:21:34 UTC; the application database backup completed at 08:28:58 UTC. All three native services reported success/status 0. These checks did not replace the separately recorded restore drills or resolve the Backblaze account cap.

Certificate warnings on Titan-0a/0c refer to expiry on January 15, 2027. They need a planned rolling K3s restart/renewal; they are not an immediate expired certificate failure. No control plane was restarted during this settling cycle.

Rollback

Revert the relevant focused commit, push the tracked branch and reconcile that stack. Capacity reverts can make CI unschedulable again; do not restore the older multi-agent cap or place builds on saturated storage merely to clear the queue. Do not blindly revert Titan-08 quarantine until loaded I/O is verified. Preserve Node role/storage labels and the SSA Merge/prune annotations on Node manifests. Nextcloud resource and startup changes can be reverted independently; restoring the old startup script reintroduces routine app removal/downloads. ClamAV/Gitea placement reverts cause a single-writer volume handoff and need available Pi RAM.

Do not replay the retired tenant-2 engine reset. Its exact UID/state guards and one-attempt markers intentionally prevent reuse against a different incident. Use the preserved replicas and a fresh storage diagnosis for a future failure.

The OpenSearch capacity expansion cannot be undone by shrinking its PVC. Keep 1280 GiB in Git even when reverting other logging changes. The guarded temporary block-release Job was retired; do not replay it without rechecking actual disk space and the current block cause.

Keep ingress-pool and communication-pool ranges non-overlapping. Moving the VIP back requires a staged Flux transition; blindly reverting the cleanup commit is not a rollback procedure. Reverting only Traefik placement is independent, but first keep at least one local endpoint on a node allowed by lan-ai-adv. Retain TLS, source restrictions and Local traffic policy. FRR can be re-enabled through the HelmRelease if BGP is deliberately introduced, with its resource impact accounted for. Restoring 05/06 also requires removing their speaker exclusion.

For CI, revert the per-project workspace changes before removing ci-scratch, and let existing agents/PVCs finish. Never reuse the Delete scratch class for persistent data. The standard astreae class and all existing retained application volumes were left unchanged. CPU-request changes preserve the existing limits; verify scheduling and actual service latency before raising build concurrency.

Emergency cleanup rollback: leave titan-24-rootfs-sweep-job.yaml outside the active kustomization. Re-adding it enables its aggressive emergency settings; that is a fresh repair action, not routine cleanup. Keep completed bootstrap Jobs without TTL while Flux manages them, or retire them explicitly. A cluster inventory found all other remaining Flux Jobs with TTL suspended.

The first Playwright pull for bstein-dev-home 581 took about 12 minutes on Titan-07; all six agent containers then became Ready. Its runtime lives on the 58 GiB USB device, separate from its 29 GiB SD root. Cold image extraction still saturates that device despite scratch being offloaded. The custom image pruner was triggered by root usage even when runtime images were on another filesystem. Native kubelet image GC now owns that lifecycle; this is not a claim that the USB device is adequate for sustained image construction. Keep CI serialized.

The sweeper rollout initially stalled behind terminating pods on offline 05/06. Its budget was temporarily raised to three (two offline plus one replacement), and startup readiness now waits for the initial cleanup to finish. Native cleanup removed those old pods; the normal budget was restored to one in fd76b043.

Kubernetes recommends letting kubelet own image garbage collection rather than running a competing external collector: garbage collection documentation.

At 12:26 UTC all 19 replacement cleanup helpers were Ready with zero restarts. The later CI completion snapshot below supersedes the running/queued state at that time. Cold-image and publication measurements above remain separate from application availability.

The LAN bastion checks passed all 96 requests from 11:56:46 to 12:08:22 UTC against the four explicit private-VIP health URLs with certificate verification. The first four sample rows were inspected in tool output; the following twenty rows are retained in lan-health-observed.json. This is not a laptop-side route test. Private inference capabilities also returned the expected unauthenticated 401 through 192.168.22.50, with TLS verification successful; no inference ran.

At 12:46 UTC, Data Prepper 2820 had also completed SUCCESS. Ariadne 563, Ananke 384, Atlasbot 442, bstein-dev-home 581 and Soteria 1306 passed on the corrected CI setup. Metis 405 was provisioning; Pegasus 339 and IaC 4232 were still queued. Earlier cancelled builds are retained as ABORTED, not counted as new infrastructure failures.

The repository-suggested combination --server-side --dry-run=client is rejected by the installed kubectl. Validation used kubectl apply --dry-run=client -k services/maintenance, plus a successful render and focused Flux diff. A diff exit status of one indicated the intended changes.

Do not restore the removed node-nofile automatic K3s restart as part of a resource rollback. Changing its small reservation is independent of host-runtime restart policy. The setup helper permits 25% unavailable because its host settings persist when it exits; it is not an application-serving DaemonSet.

At 12:54 UTC Metis 405 completed SUCCESS. Its new publication entered the normal Flux image-automation path. Data Prepper 2821 was running; Pegasus 339 and IaC 4232 remained queued. At 12:51:43 all 18 eligible node-nofile helpers were updated/Ready; all old misplaced helper pods were removed by their native controller. Titan-07 now reserves 5m/16Mi for this helper, freeing 45m/80Mi. No application pod restart was caused by that rollout.

Follow-up under scheduled and build load

The 12:51:59 window did not pass: at 13:00:00 the scheduled OpenSearch tuner preempted worker 0 again. That job was restricted to Pi 5 workers, leaving only Titan-07 eligible during hardware quarantine. Commit eb1eb369 gives it the native maintenance-batch PriorityClass (value 0, preemptionPolicy Never), permits Pi 4 placement while preferring Pi 5, removes its unused service-account token, and pins the exact observed Python image. This is a targeted batch policy, not a change to worker model, authentication, provider access or job state.

The read-only placement Job showed the admitted priority and Never policy, but its first attempt exceeded its 300-second deadline during renewed runtime I/O saturation. It is not recorded as a passed service check. The sampler now includes Normal scheduler Preempted events and pods marked for deletion even when Ready is still true; d85642ad adds that regression coverage (11 tests, 98.65% statement coverage). Only operational fields are retained.

Data Prepper 2822 passed, but Kaniko still unpacked its large base image into the node's container filesystem. Titan-07 reached 99.7% disk busy and 59% I/O pressure. Completing a build did not make this an acceptable service-load test. CI was temporarily quieted and the next old-template Data Prepper build 2823 was stopped. Commit b51cbc2f switches this pipeline to the already deployed Docker-in-Docker image, stores its layers on a 40 GiB disposable workspace PVC, keeps ARM64 output explicit and binds the daemon to pod loopback. Registry credentials enter via stdin and are removed with the temporary Docker config. Jenkins accepted the full pipeline; a synthetic CLI test verified ARM64, intended tags, credential stdin and cleanup. The actual replacement build and its resource behavior must still be verified before closing this cycle.

The second placement check (4d9fe145) completed at 13:16:54 UTC. Its container ran from 13:16:46 to 13:16:50 and exited 0 on Titan-07, with admitted preemptionPolicy Never. It performed only a cluster-health GET and printed a fixed success line, without reading index contents. The first failed attempt remains part of the evidence; success after I/O subsided does not erase it.

Data Prepper 2824 exposed an implementation error in the new builder: the Docker entrypoint added tcp://0.0.0.0:2375 while the supplied flags also requested 127.0.0.1:2375. Dockerd exited 1 with a listener conflict, and Jenkins's native container-termination reaper removed the agent. This was not another automatic Ariadne cancellation. b12991cb explicitly supplies dockerd as the first entrypoint argument, retaining the image's setup while preventing default listener insertion. Registry metadata verified the pinned image is ARM64 and uses dockerd-entrypoint.sh. A successful corrected invocation remains required. Pegasus 339 completed with a quality-gate failure; its reported aggregate coverage was 98.855%, so that number alone does not establish which gate failed.

The regular 13:30 OpenSearch tuner completed at 13:30:10 on Titan-07 with priority 0 and preemptionPolicy Never. The Hermes worker remained running. This verifies the normal scheduled path, in addition to the read-only check. The queued Data Prepper 2825 attempt loaded the old entrypoint configuration; it was stopped while still waiting for an agent so the corrected attempt can run. That deliberate cancellation is separate from 2824's startup defect.

Data Prepper 2826 passed after the entrypoint correction. Recorded duration was 406.619 seconds including its queue/allocation time, so this is not a pure build benchmark. Both intended tags published digest sha256:62f00ee7d64c3f85446c282c11c387eb6f73d73ad9d6bc5a545ab9849e8f0bd0. Live Docker reported 27.5.1, aarch64, overlay2 and /var/lib/docker on the 39.1 GiB Longhorn scratch filesystem. At 13:43:50 the two-minute runtime-disk busy sample was 11.54%, versus 99.69% during the old builder; I/O pressure was 13.40%. This is observational evidence under different build stages, not a controlled throughput benchmark or a hardware repair.

Pegasus 339's sole gate-summary issue was Jenkinsfile length (516 > 500), caused by the earlier pipeline additions. Sonar was OK. Pegasus 5fbc6d0 extracts the existing report step into scripts/report_quality_gate.sh without weakening the gate. Jenkinsfile is now 472 lines; its own production-file LOC check, Jenkins pipeline validation, shell syntax and a failed-dependency/report regression passed. A subsequent full project rerun remains separately observable.

Final settling observation

From 13:16:24 to 13:45:59 UTC (29 minutes 35 seconds), all 145 application controllers met their configured targets. All 57 active Flux Kustomizations and 21 HelmReleases were Ready. There were no container restart increases, same-name application pod replacements, new preemptions, node Ready transitions, unhealthy attached volumes or unavailable requested volumes. Ordinary CI agents continued to turn over. Titan-05/06 remained offline with 23 unavailable pods on those nodes; those are not counted as recovered hardware.

The four LAN health checks passed 96/96 requests from 13:36:54 to 13:48:33 UTC, using explicit 192.168.22.9 resolution and TLS verification from Titan-jh. Public-path checks passed 350/352 from 13:16:56 to 13:38:41 UTC. The two timeouts at 13:26:56 occurred before DNS resolution completed on the management host; TCP and TLS had not started. Their cause is not established as a cluster outage. These checks do not verify the Windows laptop's own route.

The completed read-only placement Job is now retired from the logging kustomization. Its manifest remains available as an explicit validation example; it will not recur automatically. The normal scheduled maintenance Job has already passed with preemption disabled. Jenkins remains enabled with one agent. The waiting IaC and Pegasus jobs retain eligible templates; their waiting state alone does not establish a missing agent definition or failed service. At 13:51 UTC the active agent was executing IaC quality-gate main build 84; Jenkins was not quieted. Data Prepper 2827 and Metis 406 had also succeeded. Pegasus 340 was queued with its corrected pipeline; its full rerun was not yet verified. The final logging render and client dry-run passed; Flux diff showed only deletion of the completed validation Job.

This closes the active application-churn repair cycle, not the physical repair or backup programme. Native probe delays on weak media, offline/power faults, limited spare worker capacity and Backblaze's storage cap remain tracked in CLUSTER_STABILIZATION.md. No probe was weakened to hide those findings. A 7-14-day observation period and controlled restore/failure drills remain open.