From f1fe01830a570f42c9e16e0a933cbff71f853be9 Mon Sep 17 00:00:00 2001 From: jenkins Date: Sat, 3 Oct 2026 02:03:50 -0500 Subject: [PATCH] logging: restore normal rollback checks after recovery --- docs/CLUSTER_STABILIZATION.md | 28 ++++++++++++-------- services/logging/opensearch-helmrelease.yaml | 4 --- 2 files changed, 17 insertions(+), 15 deletions(-) diff --git a/docs/CLUSTER_STABILIZATION.md b/docs/CLUSTER_STABILIZATION.md index 527a53dd..39fee5fa 100644 --- a/docs/CLUSTER_STABILIZATION.md +++ b/docs/CLUSTER_STABILIZATION.md @@ -21,6 +21,7 @@ Use [Cluster operations](CLUSTER_OPERATIONS.md) for the ordinary operator path. | Spread Vault injectors | Two healthy replicas on Titan-08 and Titan-22; native anti-affinity and minAvailable=1 PDB | Revert 5f3f3184 | | Protected application databases | Nineteen logical dumps and globals completed with checksums on Titan-0b; isolated Gitea restore passed in six seconds with 111 tables | Keep copies; stop the daily timer if needed | | Application PostgreSQL resource/probe repair | Pinned current PostgreSQL 15.18 image, 1 GiB request / 2 GiB limit, native startup/readiness; 2/2 Ready on Titan-17 and healthy attached volume | Revert 103e2064 only after checking destination capacity; another database restart is required | +| Recovered Firefly and OpenSearch | Firefly Ready on Titan-08; OpenSearch Ready with its existing healthy volume on Titan-22. Pi capacity could not fit OpenSearch, so the documented CPU fallback was used; no GPU reassignment | Review capacity before placement rollback; preserve existing volumes | | Protected local-only data policy | Hermes PVCs excluded from Soteria's cloud policy through Flux | Do not broaden this policy without reviewing data authorization | Host backup implementation is in commit 80aff498 and @@ -30,17 +31,13 @@ replay or starting K3s. A complete control-plane disaster drill remains outstand ## Work being completed -* OpenSearch: preserve the existing disk and 2 GiB heap; request 3 GiB memory - instead of 768 MiB, remove the offline titan-05 pin, and use the chart's actual - `nodeAffinity` setting. Pi5 then Pi4 are preferred; titan-22 is the explicit - last-resort CPU destination. The original Helm operation must converge before - the corrected generation can be verified. * Soteria: source fix 7018d4c in the Soteria repository allows live RWO Longhorn snapshots while preserving the restic mount guard. Manual Longhorn requests also enforce exclusions. All Go package tests pass; normal image publication and a representative backup verification remain required. -* Native application-backup schedule and its freshness alert are being installed - after the initial successful copy and restore check. +* Native application backups now run daily from Titan-0b. The timer is enabled, + the verified initial bundle timestamp is scraped, and the 36-hour freshness + alert is provisioned. The first scheduled run remains to be observed. * Soteria CI: the coverage report was generated after Sonar analysis, causing a false zero-coverage gate despite 96.4% measured test coverage. Commit f167d7d orders tests before analysis and awaits the Sonar gate; Jenkins declarative @@ -87,7 +84,16 @@ measurements of the transient minimum. Both had `get_throttled=0x50000` between events. The fault remains active even when that individual sample reports only historical bits. The user is checking power/boot-media issues physically. -Titan-18 remains a blocker: Firefly and OpenSearch pods are waiting for graceful -termination on its stalled runtime. Do not force detach their mounted volumes -while the old host may still have writers. Its replacement/repair is distinct -from Kubernetes scheduling; a cordon alone cannot repair hung I/O. +Firefly recovered on Titan-08 at approximately 06:52 UTC after its old pod +terminated and its volume moved normally. OpenSearch recovered on Titan-22 around 07:03 UTC after the corrected placement +was applied. Commit 2f043d10 temporarily allowed the old offline-node rollback +to finish. Normal rollback readiness was restored after the corrected release +became Ready. +Titan-18 remains cordoned; rescheduling does not repair its runtime media. + +Ananke was confirmed restarting jenkins-vault-sync approximately every minute +from retained credential-warning events, including after the original failure +cleared. A current-pod/UID/freshness check and persistent 30-minute repair +cooldown are being validated in the Ananke repository. Its self-update fallback +also installed revisions after quality-gate failures; the new default will +leave the running binary in place when verification fails. diff --git a/services/logging/opensearch-helmrelease.yaml b/services/logging/opensearch-helmrelease.yaml index 79a02989..0ea0e5ce 100644 --- a/services/logging/opensearch-helmrelease.yaml +++ b/services/logging/opensearch-helmrelease.yaml @@ -15,10 +15,6 @@ spec: remediation: retries: 3 strategy: rollback - # The previous release targets an offline node; do not wait for it to recover - # before Flux can retry the corrected placement. Upgrade readiness stays enabled. - rollback: - disableWait: true postRenderers: - kustomize: patches: