docs: prioritize steady service operation and record containment
This commit is contained in:
parent
8a6d3db16a
commit
dfcf75866d
@ -24,6 +24,12 @@ separates completed repairs from open problems. The
|
||||
There is no new coordinating framework to learn. Prefer the native owner above.
|
||||
Use a recovery tool only when its documented operation matches the failure.
|
||||
|
||||
The first priority is steady service operation. Repair recurring restarts and
|
||||
storage pressure before adding workload or upgrading the platform. Keep known
|
||||
unstable nodes quarantined; retain recovery copies and verify application health.
|
||||
CI may queue when compatible capacity is unavailable. That is an explicit
|
||||
capacity problem to repair, not a reason to overcommit the storage workers.
|
||||
|
||||
## Node administrative access
|
||||
|
||||
From the management workstation, use the existing SSH aliases. Most nodes use
|
||||
@ -134,11 +140,29 @@ Ananke repository before updating; `--binary-only --skip-deps` is the targeted
|
||||
path, while the normal updater also applies host configuration templates.
|
||||
|
||||
Ariadne's `ARIADNE_HERMES_HUNG_BUILD_MINUTES` setting is in
|
||||
`services/maintenance/apps/ariadne-deployment.yaml`. It is currently 120 minutes.
|
||||
Despite the name, this is an elapsed-time cancellation rule, not a detector of
|
||||
CPU or log progress. It can abort an active build. Review the build stage and
|
||||
node I/O before changing that cutoff; do not compensate for slow disks by
|
||||
starting more concurrent agents.
|
||||
`services/maintenance/apps/ariadne-deployment.yaml`. It is currently zero, which
|
||||
disables elapsed-time-only cancellation. Its retained event confirmed that the
|
||||
previous 120-minute setting canceled an active Metis image build. Other bounded
|
||||
diagnosis/repair remains enabled. Do not re-enable cancellation until it can
|
||||
distinguish a stalled job from a long-running job making useful progress.
|
||||
|
||||
Jenkins's Kubernetes cloud cap is currently one in
|
||||
`services/jenkins/configmap-jcasc.yaml`. The shared default, IaC, Data Prepper and
|
||||
Metis templates exclude quarantined nodes and storage workers. Custom inline
|
||||
templates in other repos may need the same correction. Inspect actual agent
|
||||
placement rather than assuming the default template controls every pipeline.
|
||||
A Pending agent consumes the cap; inspect FailedScheduling events and resource
|
||||
requests before deciding that Jenkins is hung. Do not increase concurrency to
|
||||
solve an unschedulable agent. The tracked JCasC revision triggers a controller
|
||||
rollout, so quiet the queue and arrange active-build completion before editing.
|
||||
|
||||
Node quarantine is declared in `infrastructure/core/node-maintenance.yaml`.
|
||||
Those Node resources explicitly preserve hardware/worker/storage labels and use
|
||||
Flux SSA Merge to retain other runtime fields. A cordon blocks ordinary new
|
||||
placement; it must not repeatedly remove existing storage DaemonSets. After a
|
||||
quarantine change, verify manager pod UIDs stay unchanged across reconciliation.
|
||||
Keep prune protection on Node declarations; repair, uncordon in Git and verify
|
||||
before considering their removal.
|
||||
|
||||
## Five-minute check
|
||||
|
||||
|
||||
@ -7,17 +7,22 @@ Use [Cluster operations](CLUSTER_OPERATIONS.md) for the ordinary operator path.
|
||||
|
||||
## Six-priority progress tracker
|
||||
|
||||
Started against the user's approved list on 2026-10-03. Updated: 2026-10-04 07:43 UTC.
|
||||
Started against the user's approved list on 2026-10-03. Updated: 2026-10-04 08:18 UTC.
|
||||
"Verified" means the stated check passed; it does not imply a completed soak or
|
||||
failure drill. Physical repairs and disruptive recovery drills remain separate
|
||||
from routine software work.
|
||||
|
||||
The current first priority is a stable steady state. Stop recurring disruption,
|
||||
protect application and storage capacity, and verify recovery copies before
|
||||
expanding workload or scheduling platform upgrades. Keep necessary repairs
|
||||
small and inspectable; a new recovery controller is not a prerequisite.
|
||||
|
||||
| Priority | Status | Completed evidence | Next action / completion gate |
|
||||
| --- | --- | --- | --- |
|
||||
| 1. Immediate software repairs | In progress; Soteria publication blocked | Monerod remains 2/2 Ready; its daemon has zero restarts, status sidecar one. Soteria release 1297 passed enforced gates but failed Go VCS stamping inside the image build; runtime remains 0.1.0-120 | Repair image build metadata, publish Soteria and validate an eligible backup and restore; continue Monerod latency/sync checks |
|
||||
| 2. Reliable nodes | Software repairs applied; hardware and identity checks remain | Full root administration verified on 21 reachable machines; Vault verification metadata saved. Repaired credentials on 12/13/19/20/21. Titan-04/11/14/18 excluded from new placement. Titan-19 log rotation repaired | Verify changed SSH identities on 09/10; recover 05/06/16; address fresh power faults and runtime-media stalls. Complete Metis image publication and physical replacement-image validation |
|
||||
| 3. Recovery coverage | In progress; cloud uploads blocked | External datastore restore and second LAN copy verified; 19 application DB dumps completed; scheduled application backup completed at 08:37 UTC; representative Gitea restore passed. Backblaze storage-cap failure established | Resolve the account storage cap within an agreed budget, complete eligible PVC coverage and Soteria restore, resolve Hermes tenant-2 safely, and provide approved local-only recovery copies |
|
||||
| 4. Capacity and placement | Partial | CI agent cap reduced to two; shared PostgreSQL/OpenSearch resource corrections deployed; Vault injectors spread. Longhorn now uses least-effort balancing and one rebuild per node | Account for the measured 2-3.5 GiB storage-process footprint, correct remaining requests/placement, and calculate compatible spare capacity for one worker loss |
|
||||
| 4. Capacity and placement | Partial; CI waiting for compatible capacity | CI agent cap reduced to one; shared default, Metis, IaC and Data Prepper agent templates now exclude storage/quarantined workers. Shared PostgreSQL/OpenSearch resource corrections deployed; Vault injectors spread. Longhorn uses least-effort balancing and one rebuild per node | Restore compatible healthy worker capacity; inspect other projects' inline agent templates. Account for the measured 2-3.5 GiB storage-process footprint and calculate spare capacity for one worker loss |
|
||||
| 5. Supported software | Planned; backups established | Installed K3s cohorts identified as 1.31.5 and 1.33.3; titan-db Ubuntu 24.10 is beyond support | Select a compatible supported baseline; stage OS/K3s/storage upgrades with rollback. Plan collation/index repair |
|
||||
| 6. Maintainability and sustained health | In progress | Operator guide and rollback record published; Ananke stale-event loop fixed; updater failure reporting corrected; overlapping restart/eviction behavior removed | Complete the service recovery checklist, validate remaining automation boundaries, arrange controlled recovery checks, then observe 7-14 days |
|
||||
|
||||
@ -79,9 +84,11 @@ replacement. No node was reflashed or rebooted during this work.
|
||||
removal of default passwordless grants, cleanup of applied bootstrap passwords,
|
||||
and correct file permissions/native systemd links in both image-writing paths.
|
||||
- [ ] Complete the normal Metis image publication and Flux rollout. Build 400
|
||||
ended ABORTED after 7272.9 seconds; the exact abort cause is not established
|
||||
from the inspected metadata. Build 401 is running its quality gate. The live
|
||||
image is still 0.1.0-399 and does not contain these source changes. The next
|
||||
ended ABORTED after 7272.9 seconds; the later Ariadne event inspection established
|
||||
that its elapsed-time rule requested this cancellation. Build 401 was stopped
|
||||
deliberately to release its storage-node agent; queued build 402 was superseded
|
||||
after placement/resource corrections. The live image is still 0.1.0-399 and
|
||||
does not contain these source changes. The next
|
||||
replacement image still needs a physical boot check before these changes are
|
||||
considered proven on hardware.
|
||||
- [x] Enable Ananke's existing password-backed host-action path. Atlas 886dbeee
|
||||
@ -103,7 +110,7 @@ replacement. No node was reflashed or rebooted during this work.
|
||||
checks, not just successful key installation.
|
||||
- [ ] Resolve the remaining network/physical findings below.
|
||||
|
||||
At the 07:43 UTC check, 57/57 active Flux configurations and 19/21 Kubernetes
|
||||
At the earlier 07:43 UTC check, 57/57 active Flux configurations and 19/21 Kubernetes
|
||||
nodes remained Ready. Ananke on both hosts had zero automatic restarts;
|
||||
Ariadne, Jenkins and the Grafana application container were Ready with zero
|
||||
restarts. The Grafana Vault sidecar had two earlier restarts. Storage acceptance
|
||||
@ -113,6 +120,71 @@ and the existing Titan-14 share manager was unready. Longhorn reported 74
|
||||
attached healthy volumes and 61 detached volumes with unknown robustness;
|
||||
detached/unknown by itself is not a new volume failure.
|
||||
|
||||
### Steady-state containment - October 4, 08:18 UTC
|
||||
|
||||
- Quarantine/controller interaction: the Longhorn managers on 11/14/18 were
|
||||
repeatedly replaced at core Flux reconciliation times, followed by node-label
|
||||
restoration by the existing role reconciler. Atlas eb006075 declares their
|
||||
essential hardware/worker/storage labels and uses Flux SSA Merge on those
|
||||
three Node resources. Quarantine and prune protection remain. After the fix,
|
||||
all three manager pod UIDs were unchanged and Ready through two further
|
||||
explicit core reconciliations, with zero container restarts. This is a short
|
||||
stability check, not proof that the underlying power/media faults are repaired.
|
||||
- Ariadne's retained event for Metis 400 records `hung_build` and
|
||||
`abort_requested=true`. The deployed code canceled by elapsed age, without
|
||||
checking useful build progress. Atlas d47e7b8e sets
|
||||
`ARIADNE_HERMES_HUNG_BUILD_MINUTES=0`, the existing disable setting. Its new
|
||||
pod is Ready; a runtime check confirms long elapsed time alone no longer
|
||||
qualifies for cancellation. Other bounded diagnosis/repair remains enabled.
|
||||
- Atlas 8a6d3db1 sets Jenkins's cloud agent cap to one and makes storage/node
|
||||
exclusions mandatory in the shared template, IaC and Data Prepper pipelines.
|
||||
Metis 11ecd91 applies the same exclusions. These cover the inspected pipelines;
|
||||
unrelated projects with their own inline YAML still need a placement audit.
|
||||
- Jenkins was quieted for one controlled controller rollout. Metis 401 and Data
|
||||
Prepper 2814 were stopped through Jenkins to release Titan-15/19; their flows
|
||||
are complete and both agents are gone. Superseded queued Metis 402 and IaC 4230
|
||||
were also stopped. Jenkins is Ready, quiet mode is off, and its runtime cloud
|
||||
cap is one. No application/storage pod was force-deleted for this change.
|
||||
- Metis 342a640 gives its lightweight Docker-client and publisher containers
|
||||
explicit 128 MiB reservations instead of namespace-default 256 MiB each.
|
||||
The resulting 1.25 GiB pod can fit the sampled Titan-08 reservation gap;
|
||||
actual build startup and peak consumption remain unverified. Limits and the
|
||||
build daemon's reservation were not reduced.
|
||||
- Capacity is still insufficient for some queued work: Data Prepper 2815 asks
|
||||
for 2 GiB and is Pending on the protected worker pool. A Pending agent also
|
||||
consumes the one-agent cap. Do not describe this as a fully functioning CI
|
||||
fleet, lower all requests to force placement, or return builds to overloaded
|
||||
storage nodes. Restore compatible worker capacity or measure and restructure
|
||||
the affected pipeline's reservations before increasing concurrency.
|
||||
|
||||
At the follow-up check, 57/57 active Flux configurations and 19/21 nodes were
|
||||
Ready; 05/06 remained unavailable. The intentionally suspended historical
|
||||
migration Kustomization is excluded from the active count. Titan-15's engine
|
||||
installer was Ready with its restart count unchanged at 109. Longhorn reported
|
||||
73 attached healthy volumes and 61 detached/unknown volumes. The tenant-2 share
|
||||
manager remains unready; no workspace content was read. None of these snapshots
|
||||
establishes N+1 capacity or the required sustained-health observation period.
|
||||
Grafana's public HTTPS `/api/health` returned HTTP 200 with certificate checking
|
||||
enabled. This management workstation routes through 192.168.1.254; direct
|
||||
requests to the cluster LAN VIPs timed out, so this is not a successful laptop
|
||||
LAN-path test. Grafana's configured ingress VIP is 192.168.22.9; .50 is the
|
||||
separate LAN inference ingress.
|
||||
|
||||
Validation: Kustomize renders, client dry-runs and focused Flux diffs passed for
|
||||
core, maintenance, Jenkins and logging. The installed kubectl rejects combining
|
||||
`--server-side` with `--dry-run=client`; the valid client dry-run was used.
|
||||
Jenkins's native declarative validator accepted IaC, Data Prepper and both Metis
|
||||
pipeline revisions. Embedded JCasC parsed; live Flux and Jenkins checks above
|
||||
verify application of the configuration. No new real-data inference job ran.
|
||||
|
||||
Rollback: use focused Git reverts and Flux reconciliation. Restoring the older
|
||||
CI cap or placement permits renewed storage contention; restoring Ariadne's
|
||||
older threshold resumes cancellation of active builds by age. Preserve the
|
||||
explicit node role/storage labels when changing quarantine rather than restoring
|
||||
the minimal Node declarations that preceded eb006075. Jenkins JCasC changes
|
||||
trigger its tracked revision and controller rollout, so coordinate a quiet
|
||||
queue first. Metis resource requests can be reverted independently in its repo.
|
||||
|
||||
### Current node distinctions
|
||||
|
||||
| Nodes | Established evidence | Remaining action |
|
||||
@ -175,14 +247,14 @@ administrator session still works; do not replace the whole authorized_keys file
|
||||
| Aligned Flux definition ownership | Removed creation-only policy and adopted already-active service state; no service paths or source refs changed | Revert the focused Flux commits; review suspension fields before doing so |
|
||||
| Backup freshness alerts | Both datastore-copy timestamps are scraped; native Grafana provisioning reload returned HTTP 200 | Revert the focused monitoring change; backups continue independently |
|
||||
| Quarantined stalled runtime media | Titan-14 and Titan-18 are SchedulingDisabled via `infrastructure/core/node-maintenance.yaml`; no forced storage detach | After repair, change `unschedulable` to false in Git and verify before removing the prune-disabled Node declaration |
|
||||
| Bounded CI concurrency | Jenkins controller Ready after applying the two-agent cap; existing agents reconnected | Revert 440244f3 and reconcile Jenkins; allow a controller restart |
|
||||
| Bounded CI concurrency | Jenkins controller Ready with runtime cap one (8a6d3db1); inspected pipeline placements exclude storage workers | Revert the focused change and reconcile Jenkins; allow a controller restart and account for renewed storage load |
|
||||
| Spread Vault injectors | Two healthy replicas on Titan-08 and Titan-22; native anti-affinity and minAvailable=1 PDB | Revert 5f3f3184 |
|
||||
| Protected application databases | Nineteen logical dumps and globals completed with checksums on Titan-0b; isolated Gitea restore passed in six seconds with 111 tables | Keep copies; stop the daily timer if needed |
|
||||
| Application PostgreSQL resource/probe repair | Pinned current PostgreSQL 15.18 image, 1 GiB request / 2 GiB limit, native startup/readiness; 2/2 Ready on Titan-17 and healthy attached volume | Revert 103e2064 only after checking destination capacity; another database restart is required |
|
||||
| Recovered Firefly and OpenSearch | Firefly Ready on Titan-08; OpenSearch Ready with its existing healthy volume on Titan-22. Pi capacity could not fit OpenSearch, so the documented CPU fallback was used; no GPU reassignment | Review capacity before placement rollback; preserve existing volumes |
|
||||
| Protected local-only data policy | Hermes PVCs excluded from Soteria's cloud policy through Flux | Do not broaden this policy without reviewing data authorization |
|
||||
| Stopped stale credential-event restart loop | Ananke c73b8fe installed on Titan-db after its complete quality gate; helper generation 796 remained unchanged after normal one-minute checks resumed. Peer 53ad03a no longer runs periodic cluster repairs | Native installer retains `/usr/local/lib/ananke/rollback/ananke.previous`; rolling back restores the defective behavior, so retain a bounded repair interval if needed |
|
||||
| Corrected premature CI cancellation | Ariadne's configured elapsed-time cutoff is now 120 minutes, verified in the running Deployment | Revert 0180ba66; this returns the 45-minute cutoff that aborted an active Pi build |
|
||||
| Disabled premature CI cancellation | Ariadne's elapsed-only cutoff is disabled with value zero (d47e7b8e), verified in the running application | Reverting resumes age-only cancellations; require progress-aware detection before re-enabling |
|
||||
| Scoped Soteria permissions | Atlas 46596536 applied; intended state updates and credential reads allowed, unrelated secret reads/listing and Hermes job creation denied; policy/usage fingerprints unchanged | Revert only after reviewing why broader access is required; retain state Secrets |
|
||||
| Recovered Monerod mount | Snapshot df99cdac, clean detach b119076c, restart 0984164f; existing ext4 filesystem mounted on Titan-19, LMDB opened and RPC initialized | Preserve the snapshot and PVC. Returning to Titan-08 requires storage-path validation; do not replay the temporary scale-zero commit as a normal rollback |
|
||||
| Bounded Longhorn background I/O | b0be99b7 applied; live settings are least-effort balancing and one concurrent rebuild per node | Revert b0be99b7 and reconcile Longhorn; the versioned settings Job restores best-effort and two rebuilds per node |
|
||||
|
||||
Loading…
x
Reference in New Issue
Block a user