diff --git a/docs/CLUSTER_OPERATIONS.md b/docs/CLUSTER_OPERATIONS.md index f223d8ae..eb299eef 100644 --- a/docs/CLUSTER_OPERATIONS.md +++ b/docs/CLUSTER_OPERATIONS.md @@ -24,6 +24,12 @@ separates completed repairs from open problems. The There is no new coordinating framework to learn. Prefer the native owner above. Use a recovery tool only when its documented operation matches the failure. +The first priority is steady service operation. Repair recurring restarts and +storage pressure before adding workload or upgrading the platform. Keep known +unstable nodes quarantined; retain recovery copies and verify application health. +CI may queue when compatible capacity is unavailable. That is an explicit +capacity problem to repair, not a reason to overcommit the storage workers. + ## Node administrative access From the management workstation, use the existing SSH aliases. Most nodes use @@ -134,11 +140,29 @@ Ananke repository before updating; `--binary-only --skip-deps` is the targeted path, while the normal updater also applies host configuration templates. Ariadne's `ARIADNE_HERMES_HUNG_BUILD_MINUTES` setting is in -`services/maintenance/apps/ariadne-deployment.yaml`. It is currently 120 minutes. -Despite the name, this is an elapsed-time cancellation rule, not a detector of -CPU or log progress. It can abort an active build. Review the build stage and -node I/O before changing that cutoff; do not compensate for slow disks by -starting more concurrent agents. +`services/maintenance/apps/ariadne-deployment.yaml`. It is currently zero, which +disables elapsed-time-only cancellation. Its retained event confirmed that the +previous 120-minute setting canceled an active Metis image build. Other bounded +diagnosis/repair remains enabled. Do not re-enable cancellation until it can +distinguish a stalled job from a long-running job making useful progress. + +Jenkins's Kubernetes cloud cap is currently one in +`services/jenkins/configmap-jcasc.yaml`. The shared default, IaC, Data Prepper and +Metis templates exclude quarantined nodes and storage workers. Custom inline +templates in other repos may need the same correction. Inspect actual agent +placement rather than assuming the default template controls every pipeline. +A Pending agent consumes the cap; inspect FailedScheduling events and resource +requests before deciding that Jenkins is hung. Do not increase concurrency to +solve an unschedulable agent. The tracked JCasC revision triggers a controller +rollout, so quiet the queue and arrange active-build completion before editing. + +Node quarantine is declared in `infrastructure/core/node-maintenance.yaml`. +Those Node resources explicitly preserve hardware/worker/storage labels and use +Flux SSA Merge to retain other runtime fields. A cordon blocks ordinary new +placement; it must not repeatedly remove existing storage DaemonSets. After a +quarantine change, verify manager pod UIDs stay unchanged across reconciliation. +Keep prune protection on Node declarations; repair, uncordon in Git and verify +before considering their removal. ## Five-minute check diff --git a/docs/CLUSTER_STABILIZATION.md b/docs/CLUSTER_STABILIZATION.md index c47f6b4e..dad0ac8b 100644 --- a/docs/CLUSTER_STABILIZATION.md +++ b/docs/CLUSTER_STABILIZATION.md @@ -7,17 +7,22 @@ Use [Cluster operations](CLUSTER_OPERATIONS.md) for the ordinary operator path. ## Six-priority progress tracker -Started against the user's approved list on 2026-10-03. Updated: 2026-10-04 07:43 UTC. +Started against the user's approved list on 2026-10-03. Updated: 2026-10-04 08:18 UTC. "Verified" means the stated check passed; it does not imply a completed soak or failure drill. Physical repairs and disruptive recovery drills remain separate from routine software work. +The current first priority is a stable steady state. Stop recurring disruption, +protect application and storage capacity, and verify recovery copies before +expanding workload or scheduling platform upgrades. Keep necessary repairs +small and inspectable; a new recovery controller is not a prerequisite. + | Priority | Status | Completed evidence | Next action / completion gate | | --- | --- | --- | --- | | 1. Immediate software repairs | In progress; Soteria publication blocked | Monerod remains 2/2 Ready; its daemon has zero restarts, status sidecar one. Soteria release 1297 passed enforced gates but failed Go VCS stamping inside the image build; runtime remains 0.1.0-120 | Repair image build metadata, publish Soteria and validate an eligible backup and restore; continue Monerod latency/sync checks | | 2. Reliable nodes | Software repairs applied; hardware and identity checks remain | Full root administration verified on 21 reachable machines; Vault verification metadata saved. Repaired credentials on 12/13/19/20/21. Titan-04/11/14/18 excluded from new placement. Titan-19 log rotation repaired | Verify changed SSH identities on 09/10; recover 05/06/16; address fresh power faults and runtime-media stalls. Complete Metis image publication and physical replacement-image validation | | 3. Recovery coverage | In progress; cloud uploads blocked | External datastore restore and second LAN copy verified; 19 application DB dumps completed; scheduled application backup completed at 08:37 UTC; representative Gitea restore passed. Backblaze storage-cap failure established | Resolve the account storage cap within an agreed budget, complete eligible PVC coverage and Soteria restore, resolve Hermes tenant-2 safely, and provide approved local-only recovery copies | -| 4. Capacity and placement | Partial | CI agent cap reduced to two; shared PostgreSQL/OpenSearch resource corrections deployed; Vault injectors spread. Longhorn now uses least-effort balancing and one rebuild per node | Account for the measured 2-3.5 GiB storage-process footprint, correct remaining requests/placement, and calculate compatible spare capacity for one worker loss | +| 4. Capacity and placement | Partial; CI waiting for compatible capacity | CI agent cap reduced to one; shared default, Metis, IaC and Data Prepper agent templates now exclude storage/quarantined workers. Shared PostgreSQL/OpenSearch resource corrections deployed; Vault injectors spread. Longhorn uses least-effort balancing and one rebuild per node | Restore compatible healthy worker capacity; inspect other projects' inline agent templates. Account for the measured 2-3.5 GiB storage-process footprint and calculate spare capacity for one worker loss | | 5. Supported software | Planned; backups established | Installed K3s cohorts identified as 1.31.5 and 1.33.3; titan-db Ubuntu 24.10 is beyond support | Select a compatible supported baseline; stage OS/K3s/storage upgrades with rollback. Plan collation/index repair | | 6. Maintainability and sustained health | In progress | Operator guide and rollback record published; Ananke stale-event loop fixed; updater failure reporting corrected; overlapping restart/eviction behavior removed | Complete the service recovery checklist, validate remaining automation boundaries, arrange controlled recovery checks, then observe 7-14 days | @@ -79,9 +84,11 @@ replacement. No node was reflashed or rebooted during this work. removal of default passwordless grants, cleanup of applied bootstrap passwords, and correct file permissions/native systemd links in both image-writing paths. - [ ] Complete the normal Metis image publication and Flux rollout. Build 400 - ended ABORTED after 7272.9 seconds; the exact abort cause is not established - from the inspected metadata. Build 401 is running its quality gate. The live - image is still 0.1.0-399 and does not contain these source changes. The next + ended ABORTED after 7272.9 seconds; the later Ariadne event inspection established + that its elapsed-time rule requested this cancellation. Build 401 was stopped + deliberately to release its storage-node agent; queued build 402 was superseded + after placement/resource corrections. The live image is still 0.1.0-399 and + does not contain these source changes. The next replacement image still needs a physical boot check before these changes are considered proven on hardware. - [x] Enable Ananke's existing password-backed host-action path. Atlas 886dbeee @@ -103,7 +110,7 @@ replacement. No node was reflashed or rebooted during this work. checks, not just successful key installation. - [ ] Resolve the remaining network/physical findings below. -At the 07:43 UTC check, 57/57 active Flux configurations and 19/21 Kubernetes +At the earlier 07:43 UTC check, 57/57 active Flux configurations and 19/21 Kubernetes nodes remained Ready. Ananke on both hosts had zero automatic restarts; Ariadne, Jenkins and the Grafana application container were Ready with zero restarts. The Grafana Vault sidecar had two earlier restarts. Storage acceptance @@ -113,6 +120,71 @@ and the existing Titan-14 share manager was unready. Longhorn reported 74 attached healthy volumes and 61 detached volumes with unknown robustness; detached/unknown by itself is not a new volume failure. +### Steady-state containment - October 4, 08:18 UTC + +- Quarantine/controller interaction: the Longhorn managers on 11/14/18 were + repeatedly replaced at core Flux reconciliation times, followed by node-label + restoration by the existing role reconciler. Atlas eb006075 declares their + essential hardware/worker/storage labels and uses Flux SSA Merge on those + three Node resources. Quarantine and prune protection remain. After the fix, + all three manager pod UIDs were unchanged and Ready through two further + explicit core reconciliations, with zero container restarts. This is a short + stability check, not proof that the underlying power/media faults are repaired. +- Ariadne's retained event for Metis 400 records `hung_build` and + `abort_requested=true`. The deployed code canceled by elapsed age, without + checking useful build progress. Atlas d47e7b8e sets + `ARIADNE_HERMES_HUNG_BUILD_MINUTES=0`, the existing disable setting. Its new + pod is Ready; a runtime check confirms long elapsed time alone no longer + qualifies for cancellation. Other bounded diagnosis/repair remains enabled. +- Atlas 8a6d3db1 sets Jenkins's cloud agent cap to one and makes storage/node + exclusions mandatory in the shared template, IaC and Data Prepper pipelines. + Metis 11ecd91 applies the same exclusions. These cover the inspected pipelines; + unrelated projects with their own inline YAML still need a placement audit. +- Jenkins was quieted for one controlled controller rollout. Metis 401 and Data + Prepper 2814 were stopped through Jenkins to release Titan-15/19; their flows + are complete and both agents are gone. Superseded queued Metis 402 and IaC 4230 + were also stopped. Jenkins is Ready, quiet mode is off, and its runtime cloud + cap is one. No application/storage pod was force-deleted for this change. +- Metis 342a640 gives its lightweight Docker-client and publisher containers + explicit 128 MiB reservations instead of namespace-default 256 MiB each. + The resulting 1.25 GiB pod can fit the sampled Titan-08 reservation gap; + actual build startup and peak consumption remain unverified. Limits and the + build daemon's reservation were not reduced. +- Capacity is still insufficient for some queued work: Data Prepper 2815 asks + for 2 GiB and is Pending on the protected worker pool. A Pending agent also + consumes the one-agent cap. Do not describe this as a fully functioning CI + fleet, lower all requests to force placement, or return builds to overloaded + storage nodes. Restore compatible worker capacity or measure and restructure + the affected pipeline's reservations before increasing concurrency. + +At the follow-up check, 57/57 active Flux configurations and 19/21 nodes were +Ready; 05/06 remained unavailable. The intentionally suspended historical +migration Kustomization is excluded from the active count. Titan-15's engine +installer was Ready with its restart count unchanged at 109. Longhorn reported +73 attached healthy volumes and 61 detached/unknown volumes. The tenant-2 share +manager remains unready; no workspace content was read. None of these snapshots +establishes N+1 capacity or the required sustained-health observation period. +Grafana's public HTTPS `/api/health` returned HTTP 200 with certificate checking +enabled. This management workstation routes through 192.168.1.254; direct +requests to the cluster LAN VIPs timed out, so this is not a successful laptop +LAN-path test. Grafana's configured ingress VIP is 192.168.22.9; .50 is the +separate LAN inference ingress. + +Validation: Kustomize renders, client dry-runs and focused Flux diffs passed for +core, maintenance, Jenkins and logging. The installed kubectl rejects combining +`--server-side` with `--dry-run=client`; the valid client dry-run was used. +Jenkins's native declarative validator accepted IaC, Data Prepper and both Metis +pipeline revisions. Embedded JCasC parsed; live Flux and Jenkins checks above +verify application of the configuration. No new real-data inference job ran. + +Rollback: use focused Git reverts and Flux reconciliation. Restoring the older +CI cap or placement permits renewed storage contention; restoring Ariadne's +older threshold resumes cancellation of active builds by age. Preserve the +explicit node role/storage labels when changing quarantine rather than restoring +the minimal Node declarations that preceded eb006075. Jenkins JCasC changes +trigger its tracked revision and controller rollout, so coordinate a quiet +queue first. Metis resource requests can be reverted independently in its repo. + ### Current node distinctions | Nodes | Established evidence | Remaining action | @@ -175,14 +247,14 @@ administrator session still works; do not replace the whole authorized_keys file | Aligned Flux definition ownership | Removed creation-only policy and adopted already-active service state; no service paths or source refs changed | Revert the focused Flux commits; review suspension fields before doing so | | Backup freshness alerts | Both datastore-copy timestamps are scraped; native Grafana provisioning reload returned HTTP 200 | Revert the focused monitoring change; backups continue independently | | Quarantined stalled runtime media | Titan-14 and Titan-18 are SchedulingDisabled via `infrastructure/core/node-maintenance.yaml`; no forced storage detach | After repair, change `unschedulable` to false in Git and verify before removing the prune-disabled Node declaration | -| Bounded CI concurrency | Jenkins controller Ready after applying the two-agent cap; existing agents reconnected | Revert 440244f3 and reconcile Jenkins; allow a controller restart | +| Bounded CI concurrency | Jenkins controller Ready with runtime cap one (8a6d3db1); inspected pipeline placements exclude storage workers | Revert the focused change and reconcile Jenkins; allow a controller restart and account for renewed storage load | | Spread Vault injectors | Two healthy replicas on Titan-08 and Titan-22; native anti-affinity and minAvailable=1 PDB | Revert 5f3f3184 | | Protected application databases | Nineteen logical dumps and globals completed with checksums on Titan-0b; isolated Gitea restore passed in six seconds with 111 tables | Keep copies; stop the daily timer if needed | | Application PostgreSQL resource/probe repair | Pinned current PostgreSQL 15.18 image, 1 GiB request / 2 GiB limit, native startup/readiness; 2/2 Ready on Titan-17 and healthy attached volume | Revert 103e2064 only after checking destination capacity; another database restart is required | | Recovered Firefly and OpenSearch | Firefly Ready on Titan-08; OpenSearch Ready with its existing healthy volume on Titan-22. Pi capacity could not fit OpenSearch, so the documented CPU fallback was used; no GPU reassignment | Review capacity before placement rollback; preserve existing volumes | | Protected local-only data policy | Hermes PVCs excluded from Soteria's cloud policy through Flux | Do not broaden this policy without reviewing data authorization | | Stopped stale credential-event restart loop | Ananke c73b8fe installed on Titan-db after its complete quality gate; helper generation 796 remained unchanged after normal one-minute checks resumed. Peer 53ad03a no longer runs periodic cluster repairs | Native installer retains `/usr/local/lib/ananke/rollback/ananke.previous`; rolling back restores the defective behavior, so retain a bounded repair interval if needed | -| Corrected premature CI cancellation | Ariadne's configured elapsed-time cutoff is now 120 minutes, verified in the running Deployment | Revert 0180ba66; this returns the 45-minute cutoff that aborted an active Pi build | +| Disabled premature CI cancellation | Ariadne's elapsed-only cutoff is disabled with value zero (d47e7b8e), verified in the running application | Reverting resumes age-only cancellations; require progress-aware detection before re-enabling | | Scoped Soteria permissions | Atlas 46596536 applied; intended state updates and credential reads allowed, unrelated secret reads/listing and Hermes job creation denied; policy/usage fingerprints unchanged | Revert only after reviewing why broader access is required; retain state Secrets | | Recovered Monerod mount | Snapshot df99cdac, clean detach b119076c, restart 0984164f; existing ext4 filesystem mounted on Titan-19, LMDB opened and RPC initialized | Preserve the snapshot and PVC. Returning to Titan-08 requires storage-path validation; do not replay the temporary scale-zero commit as a normal rollback | | Bounded Longhorn background I/O | b0be99b7 applied; live settings are least-effort balancing and one concurrent rebuild per node | Revert b0be99b7 and reconcile Longhorn; the versioned settings Job restores best-effort and two rebuilds per node |