docs: prioritize steady service operation and record containment

This commit is contained in:
jenkins 2026-10-04 03:18:49 -05:00
parent 8a6d3db16a
commit dfcf75866d
2 changed files with 109 additions and 13 deletions

View File

@ -24,6 +24,12 @@ separates completed repairs from open problems. The
There is no new coordinating framework to learn. Prefer the native owner above.
Use a recovery tool only when its documented operation matches the failure.
The first priority is steady service operation. Repair recurring restarts and
storage pressure before adding workload or upgrading the platform. Keep known
unstable nodes quarantined; retain recovery copies and verify application health.
CI may queue when compatible capacity is unavailable. That is an explicit
capacity problem to repair, not a reason to overcommit the storage workers.
## Node administrative access
From the management workstation, use the existing SSH aliases. Most nodes use
@ -134,11 +140,29 @@ Ananke repository before updating; `--binary-only --skip-deps` is the targeted
path, while the normal updater also applies host configuration templates.
Ariadne's `ARIADNE_HERMES_HUNG_BUILD_MINUTES` setting is in
`services/maintenance/apps/ariadne-deployment.yaml`. It is currently 120 minutes.
Despite the name, this is an elapsed-time cancellation rule, not a detector of
CPU or log progress. It can abort an active build. Review the build stage and
node I/O before changing that cutoff; do not compensate for slow disks by
starting more concurrent agents.
`services/maintenance/apps/ariadne-deployment.yaml`. It is currently zero, which
disables elapsed-time-only cancellation. Its retained event confirmed that the
previous 120-minute setting canceled an active Metis image build. Other bounded
diagnosis/repair remains enabled. Do not re-enable cancellation until it can
distinguish a stalled job from a long-running job making useful progress.
Jenkins's Kubernetes cloud cap is currently one in
`services/jenkins/configmap-jcasc.yaml`. The shared default, IaC, Data Prepper and
Metis templates exclude quarantined nodes and storage workers. Custom inline
templates in other repos may need the same correction. Inspect actual agent
placement rather than assuming the default template controls every pipeline.
A Pending agent consumes the cap; inspect FailedScheduling events and resource
requests before deciding that Jenkins is hung. Do not increase concurrency to
solve an unschedulable agent. The tracked JCasC revision triggers a controller
rollout, so quiet the queue and arrange active-build completion before editing.
Node quarantine is declared in `infrastructure/core/node-maintenance.yaml`.
Those Node resources explicitly preserve hardware/worker/storage labels and use
Flux SSA Merge to retain other runtime fields. A cordon blocks ordinary new
placement; it must not repeatedly remove existing storage DaemonSets. After a
quarantine change, verify manager pod UIDs stay unchanged across reconciliation.
Keep prune protection on Node declarations; repair, uncordon in Git and verify
before considering their removal.
## Five-minute check

View File

@ -7,17 +7,22 @@ Use [Cluster operations](CLUSTER_OPERATIONS.md) for the ordinary operator path.
## Six-priority progress tracker
Started against the user's approved list on 2026-10-03. Updated: 2026-10-04 07:43 UTC.
Started against the user's approved list on 2026-10-03. Updated: 2026-10-04 08:18 UTC.
"Verified" means the stated check passed; it does not imply a completed soak or
failure drill. Physical repairs and disruptive recovery drills remain separate
from routine software work.
The current first priority is a stable steady state. Stop recurring disruption,
protect application and storage capacity, and verify recovery copies before
expanding workload or scheduling platform upgrades. Keep necessary repairs
small and inspectable; a new recovery controller is not a prerequisite.
| Priority | Status | Completed evidence | Next action / completion gate |
| --- | --- | --- | --- |
| 1. Immediate software repairs | In progress; Soteria publication blocked | Monerod remains 2/2 Ready; its daemon has zero restarts, status sidecar one. Soteria release 1297 passed enforced gates but failed Go VCS stamping inside the image build; runtime remains 0.1.0-120 | Repair image build metadata, publish Soteria and validate an eligible backup and restore; continue Monerod latency/sync checks |
| 2. Reliable nodes | Software repairs applied; hardware and identity checks remain | Full root administration verified on 21 reachable machines; Vault verification metadata saved. Repaired credentials on 12/13/19/20/21. Titan-04/11/14/18 excluded from new placement. Titan-19 log rotation repaired | Verify changed SSH identities on 09/10; recover 05/06/16; address fresh power faults and runtime-media stalls. Complete Metis image publication and physical replacement-image validation |
| 3. Recovery coverage | In progress; cloud uploads blocked | External datastore restore and second LAN copy verified; 19 application DB dumps completed; scheduled application backup completed at 08:37 UTC; representative Gitea restore passed. Backblaze storage-cap failure established | Resolve the account storage cap within an agreed budget, complete eligible PVC coverage and Soteria restore, resolve Hermes tenant-2 safely, and provide approved local-only recovery copies |
| 4. Capacity and placement | Partial | CI agent cap reduced to two; shared PostgreSQL/OpenSearch resource corrections deployed; Vault injectors spread. Longhorn now uses least-effort balancing and one rebuild per node | Account for the measured 2-3.5 GiB storage-process footprint, correct remaining requests/placement, and calculate compatible spare capacity for one worker loss |
| 4. Capacity and placement | Partial; CI waiting for compatible capacity | CI agent cap reduced to one; shared default, Metis, IaC and Data Prepper agent templates now exclude storage/quarantined workers. Shared PostgreSQL/OpenSearch resource corrections deployed; Vault injectors spread. Longhorn uses least-effort balancing and one rebuild per node | Restore compatible healthy worker capacity; inspect other projects' inline agent templates. Account for the measured 2-3.5 GiB storage-process footprint and calculate spare capacity for one worker loss |
| 5. Supported software | Planned; backups established | Installed K3s cohorts identified as 1.31.5 and 1.33.3; titan-db Ubuntu 24.10 is beyond support | Select a compatible supported baseline; stage OS/K3s/storage upgrades with rollback. Plan collation/index repair |
| 6. Maintainability and sustained health | In progress | Operator guide and rollback record published; Ananke stale-event loop fixed; updater failure reporting corrected; overlapping restart/eviction behavior removed | Complete the service recovery checklist, validate remaining automation boundaries, arrange controlled recovery checks, then observe 7-14 days |
@ -79,9 +84,11 @@ replacement. No node was reflashed or rebooted during this work.
removal of default passwordless grants, cleanup of applied bootstrap passwords,
and correct file permissions/native systemd links in both image-writing paths.
- [ ] Complete the normal Metis image publication and Flux rollout. Build 400
ended ABORTED after 7272.9 seconds; the exact abort cause is not established
from the inspected metadata. Build 401 is running its quality gate. The live
image is still 0.1.0-399 and does not contain these source changes. The next
ended ABORTED after 7272.9 seconds; the later Ariadne event inspection established
that its elapsed-time rule requested this cancellation. Build 401 was stopped
deliberately to release its storage-node agent; queued build 402 was superseded
after placement/resource corrections. The live image is still 0.1.0-399 and
does not contain these source changes. The next
replacement image still needs a physical boot check before these changes are
considered proven on hardware.
- [x] Enable Ananke's existing password-backed host-action path. Atlas 886dbeee
@ -103,7 +110,7 @@ replacement. No node was reflashed or rebooted during this work.
checks, not just successful key installation.
- [ ] Resolve the remaining network/physical findings below.
At the 07:43 UTC check, 57/57 active Flux configurations and 19/21 Kubernetes
At the earlier 07:43 UTC check, 57/57 active Flux configurations and 19/21 Kubernetes
nodes remained Ready. Ananke on both hosts had zero automatic restarts;
Ariadne, Jenkins and the Grafana application container were Ready with zero
restarts. The Grafana Vault sidecar had two earlier restarts. Storage acceptance
@ -113,6 +120,71 @@ and the existing Titan-14 share manager was unready. Longhorn reported 74
attached healthy volumes and 61 detached volumes with unknown robustness;
detached/unknown by itself is not a new volume failure.
### Steady-state containment - October 4, 08:18 UTC
- Quarantine/controller interaction: the Longhorn managers on 11/14/18 were
repeatedly replaced at core Flux reconciliation times, followed by node-label
restoration by the existing role reconciler. Atlas eb006075 declares their
essential hardware/worker/storage labels and uses Flux SSA Merge on those
three Node resources. Quarantine and prune protection remain. After the fix,
all three manager pod UIDs were unchanged and Ready through two further
explicit core reconciliations, with zero container restarts. This is a short
stability check, not proof that the underlying power/media faults are repaired.
- Ariadne's retained event for Metis 400 records `hung_build` and
`abort_requested=true`. The deployed code canceled by elapsed age, without
checking useful build progress. Atlas d47e7b8e sets
`ARIADNE_HERMES_HUNG_BUILD_MINUTES=0`, the existing disable setting. Its new
pod is Ready; a runtime check confirms long elapsed time alone no longer
qualifies for cancellation. Other bounded diagnosis/repair remains enabled.
- Atlas 8a6d3db1 sets Jenkins's cloud agent cap to one and makes storage/node
exclusions mandatory in the shared template, IaC and Data Prepper pipelines.
Metis 11ecd91 applies the same exclusions. These cover the inspected pipelines;
unrelated projects with their own inline YAML still need a placement audit.
- Jenkins was quieted for one controlled controller rollout. Metis 401 and Data
Prepper 2814 were stopped through Jenkins to release Titan-15/19; their flows
are complete and both agents are gone. Superseded queued Metis 402 and IaC 4230
were also stopped. Jenkins is Ready, quiet mode is off, and its runtime cloud
cap is one. No application/storage pod was force-deleted for this change.
- Metis 342a640 gives its lightweight Docker-client and publisher containers
explicit 128 MiB reservations instead of namespace-default 256 MiB each.
The resulting 1.25 GiB pod can fit the sampled Titan-08 reservation gap;
actual build startup and peak consumption remain unverified. Limits and the
build daemon's reservation were not reduced.
- Capacity is still insufficient for some queued work: Data Prepper 2815 asks
for 2 GiB and is Pending on the protected worker pool. A Pending agent also
consumes the one-agent cap. Do not describe this as a fully functioning CI
fleet, lower all requests to force placement, or return builds to overloaded
storage nodes. Restore compatible worker capacity or measure and restructure
the affected pipeline's reservations before increasing concurrency.
At the follow-up check, 57/57 active Flux configurations and 19/21 nodes were
Ready; 05/06 remained unavailable. The intentionally suspended historical
migration Kustomization is excluded from the active count. Titan-15's engine
installer was Ready with its restart count unchanged at 109. Longhorn reported
73 attached healthy volumes and 61 detached/unknown volumes. The tenant-2 share
manager remains unready; no workspace content was read. None of these snapshots
establishes N+1 capacity or the required sustained-health observation period.
Grafana's public HTTPS `/api/health` returned HTTP 200 with certificate checking
enabled. This management workstation routes through 192.168.1.254; direct
requests to the cluster LAN VIPs timed out, so this is not a successful laptop
LAN-path test. Grafana's configured ingress VIP is 192.168.22.9; .50 is the
separate LAN inference ingress.
Validation: Kustomize renders, client dry-runs and focused Flux diffs passed for
core, maintenance, Jenkins and logging. The installed kubectl rejects combining
`--server-side` with `--dry-run=client`; the valid client dry-run was used.
Jenkins's native declarative validator accepted IaC, Data Prepper and both Metis
pipeline revisions. Embedded JCasC parsed; live Flux and Jenkins checks above
verify application of the configuration. No new real-data inference job ran.
Rollback: use focused Git reverts and Flux reconciliation. Restoring the older
CI cap or placement permits renewed storage contention; restoring Ariadne's
older threshold resumes cancellation of active builds by age. Preserve the
explicit node role/storage labels when changing quarantine rather than restoring
the minimal Node declarations that preceded eb006075. Jenkins JCasC changes
trigger its tracked revision and controller rollout, so coordinate a quiet
queue first. Metis resource requests can be reverted independently in its repo.
### Current node distinctions
| Nodes | Established evidence | Remaining action |
@ -175,14 +247,14 @@ administrator session still works; do not replace the whole authorized_keys file
| Aligned Flux definition ownership | Removed creation-only policy and adopted already-active service state; no service paths or source refs changed | Revert the focused Flux commits; review suspension fields before doing so |
| Backup freshness alerts | Both datastore-copy timestamps are scraped; native Grafana provisioning reload returned HTTP 200 | Revert the focused monitoring change; backups continue independently |
| Quarantined stalled runtime media | Titan-14 and Titan-18 are SchedulingDisabled via `infrastructure/core/node-maintenance.yaml`; no forced storage detach | After repair, change `unschedulable` to false in Git and verify before removing the prune-disabled Node declaration |
| Bounded CI concurrency | Jenkins controller Ready after applying the two-agent cap; existing agents reconnected | Revert 440244f3 and reconcile Jenkins; allow a controller restart |
| Bounded CI concurrency | Jenkins controller Ready with runtime cap one (8a6d3db1); inspected pipeline placements exclude storage workers | Revert the focused change and reconcile Jenkins; allow a controller restart and account for renewed storage load |
| Spread Vault injectors | Two healthy replicas on Titan-08 and Titan-22; native anti-affinity and minAvailable=1 PDB | Revert 5f3f3184 |
| Protected application databases | Nineteen logical dumps and globals completed with checksums on Titan-0b; isolated Gitea restore passed in six seconds with 111 tables | Keep copies; stop the daily timer if needed |
| Application PostgreSQL resource/probe repair | Pinned current PostgreSQL 15.18 image, 1 GiB request / 2 GiB limit, native startup/readiness; 2/2 Ready on Titan-17 and healthy attached volume | Revert 103e2064 only after checking destination capacity; another database restart is required |
| Recovered Firefly and OpenSearch | Firefly Ready on Titan-08; OpenSearch Ready with its existing healthy volume on Titan-22. Pi capacity could not fit OpenSearch, so the documented CPU fallback was used; no GPU reassignment | Review capacity before placement rollback; preserve existing volumes |
| Protected local-only data policy | Hermes PVCs excluded from Soteria's cloud policy through Flux | Do not broaden this policy without reviewing data authorization |
| Stopped stale credential-event restart loop | Ananke c73b8fe installed on Titan-db after its complete quality gate; helper generation 796 remained unchanged after normal one-minute checks resumed. Peer 53ad03a no longer runs periodic cluster repairs | Native installer retains `/usr/local/lib/ananke/rollback/ananke.previous`; rolling back restores the defective behavior, so retain a bounded repair interval if needed |
| Corrected premature CI cancellation | Ariadne's configured elapsed-time cutoff is now 120 minutes, verified in the running Deployment | Revert 0180ba66; this returns the 45-minute cutoff that aborted an active Pi build |
| Disabled premature CI cancellation | Ariadne's elapsed-only cutoff is disabled with value zero (d47e7b8e), verified in the running application | Reverting resumes age-only cancellations; require progress-aware detection before re-enabling |
| Scoped Soteria permissions | Atlas 46596536 applied; intended state updates and credential reads allowed, unrelated secret reads/listing and Hermes job creation denied; policy/usage fingerprints unchanged | Revert only after reviewing why broader access is required; retain state Secrets |
| Recovered Monerod mount | Snapshot df99cdac, clean detach b119076c, restart 0984164f; existing ext4 filesystem mounted on Titan-19, LMDB opened and RPC initialized | Preserve the snapshot and PVC. Returning to Titan-08 requires storage-path validation; do not replay the temporary scale-zero commit as a normal rollback |
| Bounded Longhorn background I/O | b0be99b7 applied; live settings are least-effort balancing and one concurrent rebuild per node | Revert b0be99b7 and reconcile Longhorn; the versioned settings Job restores best-effort and two rebuilds per node |