logging: restore normal rollback checks after recovery

This commit is contained in:
jenkins 2026-10-03 02:03:50 -05:00
parent 2f043d104c
commit f1fe01830a
2 changed files with 17 additions and 15 deletions

View File

@ -21,6 +21,7 @@ Use [Cluster operations](CLUSTER_OPERATIONS.md) for the ordinary operator path.
| Spread Vault injectors | Two healthy replicas on Titan-08 and Titan-22; native anti-affinity and minAvailable=1 PDB | Revert 5f3f3184 |
| Protected application databases | Nineteen logical dumps and globals completed with checksums on Titan-0b; isolated Gitea restore passed in six seconds with 111 tables | Keep copies; stop the daily timer if needed |
| Application PostgreSQL resource/probe repair | Pinned current PostgreSQL 15.18 image, 1 GiB request / 2 GiB limit, native startup/readiness; 2/2 Ready on Titan-17 and healthy attached volume | Revert 103e2064 only after checking destination capacity; another database restart is required |
| Recovered Firefly and OpenSearch | Firefly Ready on Titan-08; OpenSearch Ready with its existing healthy volume on Titan-22. Pi capacity could not fit OpenSearch, so the documented CPU fallback was used; no GPU reassignment | Review capacity before placement rollback; preserve existing volumes |
| Protected local-only data policy | Hermes PVCs excluded from Soteria's cloud policy through Flux | Do not broaden this policy without reviewing data authorization |
Host backup implementation is in commit 80aff498 and
@ -30,17 +31,13 @@ replay or starting K3s. A complete control-plane disaster drill remains outstand
## Work being completed
* OpenSearch: preserve the existing disk and 2 GiB heap; request 3 GiB memory
instead of 768 MiB, remove the offline titan-05 pin, and use the chart's actual
`nodeAffinity` setting. Pi5 then Pi4 are preferred; titan-22 is the explicit
last-resort CPU destination. The original Helm operation must converge before
the corrected generation can be verified.
* Soteria: source fix 7018d4c in the Soteria repository allows live RWO Longhorn
snapshots while preserving the restic mount guard. Manual Longhorn requests
also enforce exclusions. All Go package tests pass; normal image publication
and a representative backup verification remain required.
* Native application-backup schedule and its freshness alert are being installed
after the initial successful copy and restore check.
* Native application backups now run daily from Titan-0b. The timer is enabled,
the verified initial bundle timestamp is scraped, and the 36-hour freshness
alert is provisioned. The first scheduled run remains to be observed.
* Soteria CI: the coverage report was generated after Sonar analysis, causing a
false zero-coverage gate despite 96.4% measured test coverage. Commit f167d7d
orders tests before analysis and awaits the Sonar gate; Jenkins declarative
@ -87,7 +84,16 @@ measurements of the transient minimum. Both had `get_throttled=0x50000` between
events. The fault remains active even when that individual sample reports only
historical bits. The user is checking power/boot-media issues physically.
Titan-18 remains a blocker: Firefly and OpenSearch pods are waiting for graceful
termination on its stalled runtime. Do not force detach their mounted volumes
while the old host may still have writers. Its replacement/repair is distinct
from Kubernetes scheduling; a cordon alone cannot repair hung I/O.
Firefly recovered on Titan-08 at approximately 06:52 UTC after its old pod
terminated and its volume moved normally. OpenSearch recovered on Titan-22 around 07:03 UTC after the corrected placement
was applied. Commit 2f043d10 temporarily allowed the old offline-node rollback
to finish. Normal rollback readiness was restored after the corrected release
became Ready.
Titan-18 remains cordoned; rescheduling does not repair its runtime media.
Ananke was confirmed restarting jenkins-vault-sync approximately every minute
from retained credential-warning events, including after the original failure
cleared. A current-pod/UID/freshness check and persistent 30-minute repair
cooldown are being validated in the Ananke repository. Its self-update fallback
also installed revisions after quality-gate failures; the new default will
leave the running binary in place when verification fails.

View File

@ -15,10 +15,6 @@ spec:
remediation:
retries: 3
strategy: rollback
# The previous release targets an offline node; do not wait for it to recover
# before Flux can retry the corrected placement. Upgrade readiness stays enabled.
rollback:
disableWait: true
postRenderers:
- kustomize:
patches: