56 lines
12 KiB
Markdown
56 lines
12 KiB
Markdown
# Service continuity plan for review
|
|
|
|
All services are included. This is proposed work, not a declaration that the current cluster provides high availability. Shared dependencies must be repaired first, and workload-specific state must be understood before adding replicas.
|
|
|
|
The complete controller list, images, resources, selectors, storage and current placements are in [workloads.csv](workloads.csv). Job retention and bootstrap jobs were also inspected; they are not represented as continuously available applications in that CSV. Inherited sidecars without probes are not automatically defects.
|
|
|
|
Proposed baseline targets for discussion: ordinary services recover automatically from a single worker failure within five minutes where their design supports it; long-running jobs retain progress and recover within a separately measured window; important transactional data has a much shorter recovery point than daily bulk-file backups. These targets are not verified current capabilities. Truly uninterrupted services need independent replicas plus state and dependency redundancy.
|
|
|
|
| Service family / namespace | Current issue or continuity risk | Proposed treatment and acceptance check |
|
|
|---|---|---|
|
|
| Control plane / external titan-db | Three API servers depend on one PostgreSQL host; stale discovered backup; unsupported host OS | Verified DB/token recovery first, supported host and independently recoverable standby design; prove API recovery in isolation before cutover |
|
|
| kube-system | CoreDNS has three instances; per-node CSI and ServiceLB depend on node health; local-path data is node-bound | Keep DNS spread, audit node resolver configuration and system reservations; document mail ServiceLB; verify DNS/CSI on a healthy replacement worker |
|
|
| flux-system | Kustomization-definition drift; controllers singleton; Weave UI release Ready but app absent | Normalize desired state and ownership; ensure controllers can reschedule; restore or explicitly park UI; successful render/reconcile must correspond to available objects |
|
|
| cert-manager | Two instances per controller/webhook class, with PDBs; certificate ownership conflict in Harbor | Preserve spread, resolve duplicate Certificate ownership, confirm renewal and admission health without disabling TLS |
|
|
| traefik | Two ingress instances on storage nodes 13/15; shared storage/network disturbance can affect ingress | Preserve independent placement and PDB, reserve small fixed resources, document 192.168.22.9 and .50 routes; later prove single ingress-node failure without client reconfiguration |
|
|
| metallb-system | L2 speakers across nodes, separate controller; noisy offline-node state and legacy LB mechanisms | Verify eligible speakers, L2 reachability and failover ownership for each VIP; keep tested VIP failover independent of broken nodes |
|
|
| longhorn-system | Two persistent attachment/engine faults; many retained volumes; bootstrap image dependency | Data-preserving repairs, replica/disk ownership, supported upgrade path, bounded rebuild traffic, verified backups/restores; test a replica-node loss after repair |
|
|
| postgres | Shared app DB singleton, 512 MiB limit, OOM evidence, no explicit probes | Measured resources, DB-aware readiness and backups; reduce common failure domain with Vault/SSO; introduce failover only with tested fencing and client behavior |
|
|
| vault | Singleton Vault; both injector replicas share Titan-12 | Spread injector replicas, reserve Vault capacity, verify recovery keys/backups and unseal procedure; consider supported replicated storage mode as a separate migration |
|
|
| sso | Keycloak and LDAP singletons; Keycloak shares node with DB/Vault | Spread front-end/auth dependencies, test supported Keycloak replication/session behavior, protect LDAP data and recovery; verify token issuance and existing sessions during failover |
|
|
| harbor | Core/registry/Redis/jobservice pinned to undervoltage Titan-11; duplicate certificate; boot dependency | Fix power and placement, protect registry data, keep independent bootstrap images; validate pull/push and restart order, without indiscriminately duplicating RWO writers |
|
|
| gitea | Singleton with PVC on Titan-08; health took ~5.8s once | Protect repositories and database together; reserve resources and test restart/recovery and Git fetch; add redundancy only with supported shared storage design |
|
|
| monitoring | Pushgateway unavailable; Grafana/Alertmanager co-located on Titan-11; VictoriaMetrics singleton | Repair volume, show stale/unknown metrics explicitly, spread components, back up config/data appropriately and provide independent heartbeat; prove query freshness and alert delivery |
|
|
| logging | OpenSearch pinned to offline Titan-05; heap exceeds memory request; Data Prepper restart evidence | Choose capacity, recover volume, align heap/request/limit, bound ingestion buffers and retention; verify log arrival and searchable recent entries rather than UI login only |
|
|
| maintenance | Several overlapping cleaners/recovery actors; Soteria backup failures | One ownership map, bounded action budgets, recoverability checks and incident history; prove recovery halts after repeated failures and never silently deletes required data |
|
|
| jenkins | Controller singleton; agents can exhaust compatible workers; controller/data on Titan-22 | Explicit agent concurrency and namespace budgets, protect controller resources and job state, bound caches; queued builds must not disrupt online applications |
|
|
| quality | SonarQube stateful singleton; relies on DB/SSO/metrics path | Reserve JVM memory, persist and back up DB/config, tolerate exporter failure; test analysis job, UI and recovery independently of incomplete telemetry |
|
|
| hermes | Many small-reservation workers/sidecars; stuck tenant-2 workspace; suite ledger local-path; GPU services compete | Fix workspace; measured tenant/agent limits and queues; isolate control API from model execution; preserve approved local job checkpoints and bounded retention; test recovery with synthetic metadata only |
|
|
| hermes-scm | Singleton broker with PVC on Titan-11 | Preserve broker authorization and local state; spread from unstable node, define safe interrupted-operation recovery; no added arbitrary execution or broader credentials |
|
|
| ai | Ollama tied to GPU nodes/local model storage; batch service parked | Explicit GPU lease ownership and admission; model-cache recoverability; approved local secondary backend only if capacity exists, otherwise honest queued/unavailable state; never introduce cloud fallback |
|
|
| bstein-dev-home | Singleton front/back end and chat gateway | Replicate stateless parts where compatible, keep shared DB/auth availability and route boundaries; test page load plus read-only API path, not only container status |
|
|
| cassandra | Application/simulation service, its own PostgreSQL and artifacts; Titan-23 intentionally reserved | Preserve simulations budget; back up application DB/artifacts, separate front-end availability from job capacity, verify job interruption semantics; pending cache PVC is WFFC, not automatically faulty |
|
|
| comms | Matrix, auth, Redis, LiveKit, TURN and web clients mostly singleton; many old Jobs | Protect Matrix DB/media/auth, use protocol-aware readiness and supported scaling; preserve TCP/UDP routes; test messaging/call establishment separately in an approved synthetic account test |
|
|
| mailu-mailserver | Most components singleton; mail data and queues stateful; alerts depend on mail | Reserve front/postfix/dovecot resources and storage; back up mailboxes/config, monitor queue age and dependency health; test SMTP/IMAP and controlled mail flow only after approval; do not blindly run duplicate mailbox writers |
|
|
| nextcloud | Stateful Nextcloud singleton on busy Titan-07; Collabora on Titan-24 | Reserve PHP/DB/Redis capacity, verify storage and background jobs; consistent DB/files backup, supported multi-instance design if required; prove restore and document office-session limits |
|
|
| outline | App and Redis singletons; DB/SSO/storage dependencies | Measured resources and session behavior, protect attachments/DB, replicate web layer if safe; verify document access and recovery with synthetic data |
|
|
| planka | Singleton app with stateful attachments/DB dependency | Separate stateless capacity from persisted data, backup/restore, app-level health; validate board read/write and restart in a synthetic test |
|
|
| finance | Firefly repeatedly unready; Actual Budget singleton with encrypted storage | Diagnose Firefly readiness safely, check DB/secret/probe path; preserve encrypted-storage keys and application backups; validate without reading financial records |
|
|
| health | Wger singleton on previously unreliable Titan-22 | Move or provide dependable compatible failover capacity, protect uploaded media/database; application-level health and tested restore |
|
|
| vaultwarden | Password-manager singleton and PVC on Titan-11 | Prioritize consistent encrypted vault/attachment backups and key recovery, healthy placement and safe restart; do not change auth/SSO integration as part of resource repair |
|
|
| jellyfin | Jellyfin/Pegasus singletons and media storage; GPU/network reliance | Reserve service resources, separate media transcode queue from web availability; preserve media mount availability and session state; verify supported alternate execution without stealing assigned GPU leases |
|
|
| game-stream | Wolf parked following GPU allocation; auth proxy availability does not mean streaming works | Record explicit parked state; allocate a stable GPU schedule/capacity before reactivation; test actual streaming path later; no silent reversal of the suite workload's allocation |
|
|
| crypto | Node/wallet state, a per-node miner; P2Pool and test wallet parked | Keep miners under explicit residual-capacity budgets so services remain available, back up private wallet state through approved handling, document parked capabilities; do not inspect wallets or transact during audit |
|
|
| climate | Typhon singleton and sync helper | Define hardware/external dependencies, modest reserved resources, safe recovery and stale-data indication; inspect device-control behavior before allowing duplicate active writers |
|
|
| sui-metrics | Singleton collector | Reserve small resources, ensure reschedulability, distinguish external-source outage from collector failure and preserve metric freshness |
|
|
| default / legacy objects | Old debug pods and oauth2-proxy-zot service with no endpoints | Inventory ownership; retire only confirmed obsolete objects through reviewed Git changes; preserve incident evidence and any required registry compatibility |
|
|
| CronJobs / bootstrap Jobs across namespaces | Old objects, suspended schedules and bootstrap migrations mingle with live-service health | Classify one-shot versus recurring work; set deadline/concurrency/retention and last-success checks; keep migrations explicit and prevent resubmission on routine reconciles |
|
|
|
|
## Capacity choices that need an explicit decision
|
|
|
|
1. **Keep current node reservations:** repair 04/05/06 and runtime media, right-size workloads and cap burst work. Demonstrate the compatible Pi pool can lose one node. If it cannot, this option cannot promise all-service continuity by configuration alone.
|
|
2. **Approve bounded x86 general-service capacity:** use a dedicated resource reservation on existing reliable hardware, preserving simulation/GPU budgets and validating ARM/x86 images first. This changes the existing placement policy and needs explicit approval.
|
|
3. **Add stable general-purpose capacity:** size SSD-backed workers from measured demand plus N+1 headroom. This avoids depending on repaired low-capacity nodes or taking reserved simulation/GPU resources, but has procurement and migration cost. No specific purchase or final node size is justified by this snapshot alone.
|
|
|
|
The plan does not remove a service to make the health dashboard green. Services that cannot run concurrently within the approved hardware envelope need a visible queue/schedule or additional capacity, with their availability impact agreed explicitly.
|