12 KiB

Service continuity plan for review

All services are included. This is proposed work, not a declaration that the current cluster provides high availability. Shared dependencies must be repaired first, and workload-specific state must be understood before adding replicas.

The complete controller list, images, resources, selectors, storage and current placements are in workloads.csv. Job retention and bootstrap jobs were also inspected; they are not represented as continuously available applications in that CSV. Inherited sidecars without probes are not automatically defects.

Proposed baseline targets for discussion: ordinary services recover automatically from a single worker failure within five minutes where their design supports it; long-running jobs retain progress and recover within a separately measured window; important transactional data has a much shorter recovery point than daily bulk-file backups. These targets are not verified current capabilities. Truly uninterrupted services need independent replicas plus state and dependency redundancy.

Service family / namespace Current issue or continuity risk Proposed treatment and acceptance check
Control plane / external titan-db Three API servers depend on one PostgreSQL host; stale discovered backup; unsupported host OS Verified DB/token recovery first, supported host and independently recoverable standby design; prove API recovery in isolation before cutover
kube-system CoreDNS has three instances; per-node CSI and ServiceLB depend on node health; local-path data is node-bound Keep DNS spread, audit node resolver configuration and system reservations; document mail ServiceLB; verify DNS/CSI on a healthy replacement worker
flux-system Kustomization-definition drift; controllers singleton; Weave UI release Ready but app absent Normalize desired state and ownership; ensure controllers can reschedule; restore or explicitly park UI; successful render/reconcile must correspond to available objects
cert-manager Two instances per controller/webhook class, with PDBs; certificate ownership conflict in Harbor Preserve spread, resolve duplicate Certificate ownership, confirm renewal and admission health without disabling TLS
traefik Two ingress instances on storage nodes 13/15; shared storage/network disturbance can affect ingress Preserve independent placement and PDB, reserve small fixed resources, document 192.168.22.9 and .50 routes; later prove single ingress-node failure without client reconfiguration
metallb-system L2 speakers across nodes, separate controller; noisy offline-node state and legacy LB mechanisms Verify eligible speakers, L2 reachability and failover ownership for each VIP; keep tested VIP failover independent of broken nodes
longhorn-system Two persistent attachment/engine faults; many retained volumes; bootstrap image dependency Data-preserving repairs, replica/disk ownership, supported upgrade path, bounded rebuild traffic, verified backups/restores; test a replica-node loss after repair
postgres Shared app DB singleton, 512 MiB limit, OOM evidence, no explicit probes Measured resources, DB-aware readiness and backups; reduce common failure domain with Vault/SSO; introduce failover only with tested fencing and client behavior
vault Singleton Vault; both injector replicas share Titan-12 Spread injector replicas, reserve Vault capacity, verify recovery keys/backups and unseal procedure; consider supported replicated storage mode as a separate migration
sso Keycloak and LDAP singletons; Keycloak shares node with DB/Vault Spread front-end/auth dependencies, test supported Keycloak replication/session behavior, protect LDAP data and recovery; verify token issuance and existing sessions during failover
harbor Core/registry/Redis/jobservice pinned to undervoltage Titan-11; duplicate certificate; boot dependency Fix power and placement, protect registry data, keep independent bootstrap images; validate pull/push and restart order, without indiscriminately duplicating RWO writers
gitea Singleton with PVC on Titan-08; health took ~5.8s once Protect repositories and database together; reserve resources and test restart/recovery and Git fetch; add redundancy only with supported shared storage design
monitoring Pushgateway unavailable; Grafana/Alertmanager co-located on Titan-11; VictoriaMetrics singleton Repair volume, show stale/unknown metrics explicitly, spread components, back up config/data appropriately and provide independent heartbeat; prove query freshness and alert delivery
logging OpenSearch pinned to offline Titan-05; heap exceeds memory request; Data Prepper restart evidence Choose capacity, recover volume, align heap/request/limit, bound ingestion buffers and retention; verify log arrival and searchable recent entries rather than UI login only
maintenance Several overlapping cleaners/recovery actors; Soteria backup failures One ownership map, bounded action budgets, recoverability checks and incident history; prove recovery halts after repeated failures and never silently deletes required data
jenkins Controller singleton; agents can exhaust compatible workers; controller/data on Titan-22 Explicit agent concurrency and namespace budgets, protect controller resources and job state, bound caches; queued builds must not disrupt online applications
quality SonarQube stateful singleton; relies on DB/SSO/metrics path Reserve JVM memory, persist and back up DB/config, tolerate exporter failure; test analysis job, UI and recovery independently of incomplete telemetry
hermes Many small-reservation workers/sidecars; stuck tenant-2 workspace; suite ledger local-path; GPU services compete Fix workspace; measured tenant/agent limits and queues; isolate control API from model execution; preserve approved local job checkpoints and bounded retention; test recovery with synthetic metadata only
hermes-scm Singleton broker with PVC on Titan-11 Preserve broker authorization and local state; spread from unstable node, define safe interrupted-operation recovery; no added arbitrary execution or broader credentials
ai Ollama tied to GPU nodes/local model storage; batch service parked Explicit GPU lease ownership and admission; model-cache recoverability; approved local secondary backend only if capacity exists, otherwise honest queued/unavailable state; never introduce cloud fallback
bstein-dev-home Singleton front/back end and chat gateway Replicate stateless parts where compatible, keep shared DB/auth availability and route boundaries; test page load plus read-only API path, not only container status
cassandra Application/simulation service, its own PostgreSQL and artifacts; Titan-23 intentionally reserved Preserve simulations budget; back up application DB/artifacts, separate front-end availability from job capacity, verify job interruption semantics; pending cache PVC is WFFC, not automatically faulty
comms Matrix, auth, Redis, LiveKit, TURN and web clients mostly singleton; many old Jobs Protect Matrix DB/media/auth, use protocol-aware readiness and supported scaling; preserve TCP/UDP routes; test messaging/call establishment separately in an approved synthetic account test
mailu-mailserver Most components singleton; mail data and queues stateful; alerts depend on mail Reserve front/postfix/dovecot resources and storage; back up mailboxes/config, monitor queue age and dependency health; test SMTP/IMAP and controlled mail flow only after approval; do not blindly run duplicate mailbox writers
nextcloud Stateful Nextcloud singleton on busy Titan-07; Collabora on Titan-24 Reserve PHP/DB/Redis capacity, verify storage and background jobs; consistent DB/files backup, supported multi-instance design if required; prove restore and document office-session limits
outline App and Redis singletons; DB/SSO/storage dependencies Measured resources and session behavior, protect attachments/DB, replicate web layer if safe; verify document access and recovery with synthetic data
planka Singleton app with stateful attachments/DB dependency Separate stateless capacity from persisted data, backup/restore, app-level health; validate board read/write and restart in a synthetic test
finance Firefly repeatedly unready; Actual Budget singleton with encrypted storage Diagnose Firefly readiness safely, check DB/secret/probe path; preserve encrypted-storage keys and application backups; validate without reading financial records
health Wger singleton on previously unreliable Titan-22 Move or provide dependable compatible failover capacity, protect uploaded media/database; application-level health and tested restore
vaultwarden Password-manager singleton and PVC on Titan-11 Prioritize consistent encrypted vault/attachment backups and key recovery, healthy placement and safe restart; do not change auth/SSO integration as part of resource repair
jellyfin Jellyfin/Pegasus singletons and media storage; GPU/network reliance Reserve service resources, separate media transcode queue from web availability; preserve media mount availability and session state; verify supported alternate execution without stealing assigned GPU leases
game-stream Wolf parked following GPU allocation; auth proxy availability does not mean streaming works Record explicit parked state; allocate a stable GPU schedule/capacity before reactivation; test actual streaming path later; no silent reversal of the suite workload's allocation
crypto Node/wallet state, a per-node miner; P2Pool and test wallet parked Keep miners under explicit residual-capacity budgets so services remain available, back up private wallet state through approved handling, document parked capabilities; do not inspect wallets or transact during audit
climate Typhon singleton and sync helper Define hardware/external dependencies, modest reserved resources, safe recovery and stale-data indication; inspect device-control behavior before allowing duplicate active writers
sui-metrics Singleton collector Reserve small resources, ensure reschedulability, distinguish external-source outage from collector failure and preserve metric freshness
default / legacy objects Old debug pods and oauth2-proxy-zot service with no endpoints Inventory ownership; retire only confirmed obsolete objects through reviewed Git changes; preserve incident evidence and any required registry compatibility
CronJobs / bootstrap Jobs across namespaces Old objects, suspended schedules and bootstrap migrations mingle with live-service health Classify one-shot versus recurring work; set deadline/concurrency/retention and last-success checks; keep migrations explicit and prevent resubmission on routine reconciles

Capacity choices that need an explicit decision

  1. Keep current node reservations: repair 04/05/06 and runtime media, right-size workloads and cap burst work. Demonstrate the compatible Pi pool can lose one node. If it cannot, this option cannot promise all-service continuity by configuration alone.
  2. Approve bounded x86 general-service capacity: use a dedicated resource reservation on existing reliable hardware, preserving simulation/GPU budgets and validating ARM/x86 images first. This changes the existing placement policy and needs explicit approval.
  3. Add stable general-purpose capacity: size SSD-backed workers from measured demand plus N+1 headroom. This avoids depending on repaired low-capacity nodes or taking reserved simulation/GPU resources, but has procurement and migration cost. No specific purchase or final node size is justified by this snapshot alone.

The plan does not remove a service to make the health dashboard green. Services that cannot run concurrently within the approved hardware envelope need a visible queue/schedule or additional capacity, with their availability impact agreed explicitly.