logging: restore normal rollback checks after recovery
This commit is contained in:
parent
2f043d104c
commit
f1fe01830a
@ -21,6 +21,7 @@ Use [Cluster operations](CLUSTER_OPERATIONS.md) for the ordinary operator path.
|
||||
| Spread Vault injectors | Two healthy replicas on Titan-08 and Titan-22; native anti-affinity and minAvailable=1 PDB | Revert 5f3f3184 |
|
||||
| Protected application databases | Nineteen logical dumps and globals completed with checksums on Titan-0b; isolated Gitea restore passed in six seconds with 111 tables | Keep copies; stop the daily timer if needed |
|
||||
| Application PostgreSQL resource/probe repair | Pinned current PostgreSQL 15.18 image, 1 GiB request / 2 GiB limit, native startup/readiness; 2/2 Ready on Titan-17 and healthy attached volume | Revert 103e2064 only after checking destination capacity; another database restart is required |
|
||||
| Recovered Firefly and OpenSearch | Firefly Ready on Titan-08; OpenSearch Ready with its existing healthy volume on Titan-22. Pi capacity could not fit OpenSearch, so the documented CPU fallback was used; no GPU reassignment | Review capacity before placement rollback; preserve existing volumes |
|
||||
| Protected local-only data policy | Hermes PVCs excluded from Soteria's cloud policy through Flux | Do not broaden this policy without reviewing data authorization |
|
||||
|
||||
Host backup implementation is in commit 80aff498 and
|
||||
@ -30,17 +31,13 @@ replay or starting K3s. A complete control-plane disaster drill remains outstand
|
||||
|
||||
## Work being completed
|
||||
|
||||
* OpenSearch: preserve the existing disk and 2 GiB heap; request 3 GiB memory
|
||||
instead of 768 MiB, remove the offline titan-05 pin, and use the chart's actual
|
||||
`nodeAffinity` setting. Pi5 then Pi4 are preferred; titan-22 is the explicit
|
||||
last-resort CPU destination. The original Helm operation must converge before
|
||||
the corrected generation can be verified.
|
||||
* Soteria: source fix 7018d4c in the Soteria repository allows live RWO Longhorn
|
||||
snapshots while preserving the restic mount guard. Manual Longhorn requests
|
||||
also enforce exclusions. All Go package tests pass; normal image publication
|
||||
and a representative backup verification remain required.
|
||||
* Native application-backup schedule and its freshness alert are being installed
|
||||
after the initial successful copy and restore check.
|
||||
* Native application backups now run daily from Titan-0b. The timer is enabled,
|
||||
the verified initial bundle timestamp is scraped, and the 36-hour freshness
|
||||
alert is provisioned. The first scheduled run remains to be observed.
|
||||
* Soteria CI: the coverage report was generated after Sonar analysis, causing a
|
||||
false zero-coverage gate despite 96.4% measured test coverage. Commit f167d7d
|
||||
orders tests before analysis and awaits the Sonar gate; Jenkins declarative
|
||||
@ -87,7 +84,16 @@ measurements of the transient minimum. Both had `get_throttled=0x50000` between
|
||||
events. The fault remains active even when that individual sample reports only
|
||||
historical bits. The user is checking power/boot-media issues physically.
|
||||
|
||||
Titan-18 remains a blocker: Firefly and OpenSearch pods are waiting for graceful
|
||||
termination on its stalled runtime. Do not force detach their mounted volumes
|
||||
while the old host may still have writers. Its replacement/repair is distinct
|
||||
from Kubernetes scheduling; a cordon alone cannot repair hung I/O.
|
||||
Firefly recovered on Titan-08 at approximately 06:52 UTC after its old pod
|
||||
terminated and its volume moved normally. OpenSearch recovered on Titan-22 around 07:03 UTC after the corrected placement
|
||||
was applied. Commit 2f043d10 temporarily allowed the old offline-node rollback
|
||||
to finish. Normal rollback readiness was restored after the corrected release
|
||||
became Ready.
|
||||
Titan-18 remains cordoned; rescheduling does not repair its runtime media.
|
||||
|
||||
Ananke was confirmed restarting jenkins-vault-sync approximately every minute
|
||||
from retained credential-warning events, including after the original failure
|
||||
cleared. A current-pod/UID/freshness check and persistent 30-minute repair
|
||||
cooldown are being validated in the Ananke repository. Its self-update fallback
|
||||
also installed revisions after quality-gate failures; the new default will
|
||||
leave the running binary in place when verification fails.
|
||||
|
||||
@ -15,10 +15,6 @@ spec:
|
||||
remediation:
|
||||
retries: 3
|
||||
strategy: rollback
|
||||
# The previous release targets an offline node; do not wait for it to recover
|
||||
# before Flux can retry the corrected placement. Upgrade readiness stays enabled.
|
||||
rollback:
|
||||
disableWait: true
|
||||
postRenderers:
|
||||
- kustomize:
|
||||
patches:
|
||||
|
||||
Loading…
x
Reference in New Issue
Block a user