# Service continuity plan for review All services are included. This is proposed work, not a declaration that the current cluster provides high availability. Shared dependencies must be repaired first, and workload-specific state must be understood before adding replicas. The complete controller list, images, resources, selectors, storage and current placements are in [workloads.csv](workloads.csv). Job retention and bootstrap jobs were also inspected; they are not represented as continuously available applications in that CSV. Inherited sidecars without probes are not automatically defects. Proposed baseline targets for discussion: ordinary services recover automatically from a single worker failure within five minutes where their design supports it; long-running jobs retain progress and recover within a separately measured window; important transactional data has a much shorter recovery point than daily bulk-file backups. These targets are not verified current capabilities. Truly uninterrupted services need independent replicas plus state and dependency redundancy. | Service family / namespace | Current issue or continuity risk | Proposed treatment and acceptance check | |---|---|---| | Control plane / external titan-db | Three API servers depend on one PostgreSQL host; stale discovered backup; unsupported host OS | Verified DB/token recovery first, supported host and independently recoverable standby design; prove API recovery in isolation before cutover | | kube-system | CoreDNS has three instances; per-node CSI and ServiceLB depend on node health; local-path data is node-bound | Keep DNS spread, audit node resolver configuration and system reservations; document mail ServiceLB; verify DNS/CSI on a healthy replacement worker | | flux-system | Kustomization-definition drift; controllers singleton; Weave UI release Ready but app absent | Normalize desired state and ownership; ensure controllers can reschedule; restore or explicitly park UI; successful render/reconcile must correspond to available objects | | cert-manager | Two instances per controller/webhook class, with PDBs; certificate ownership conflict in Harbor | Preserve spread, resolve duplicate Certificate ownership, confirm renewal and admission health without disabling TLS | | traefik | Two ingress instances on storage nodes 13/15; shared storage/network disturbance can affect ingress | Preserve independent placement and PDB, reserve small fixed resources, document 192.168.22.9 and .50 routes; later prove single ingress-node failure without client reconfiguration | | metallb-system | L2 speakers across nodes, separate controller; noisy offline-node state and legacy LB mechanisms | Verify eligible speakers, L2 reachability and failover ownership for each VIP; keep tested VIP failover independent of broken nodes | | longhorn-system | Two persistent attachment/engine faults; many retained volumes; bootstrap image dependency | Data-preserving repairs, replica/disk ownership, supported upgrade path, bounded rebuild traffic, verified backups/restores; test a replica-node loss after repair | | postgres | Shared app DB singleton, 512 MiB limit, OOM evidence, no explicit probes | Measured resources, DB-aware readiness and backups; reduce common failure domain with Vault/SSO; introduce failover only with tested fencing and client behavior | | vault | Singleton Vault; both injector replicas share Titan-12 | Spread injector replicas, reserve Vault capacity, verify recovery keys/backups and unseal procedure; consider supported replicated storage mode as a separate migration | | sso | Keycloak and LDAP singletons; Keycloak shares node with DB/Vault | Spread front-end/auth dependencies, test supported Keycloak replication/session behavior, protect LDAP data and recovery; verify token issuance and existing sessions during failover | | harbor | Core/registry/Redis/jobservice pinned to undervoltage Titan-11; duplicate certificate; boot dependency | Fix power and placement, protect registry data, keep independent bootstrap images; validate pull/push and restart order, without indiscriminately duplicating RWO writers | | gitea | Singleton with PVC on Titan-08; health took ~5.8s once | Protect repositories and database together; reserve resources and test restart/recovery and Git fetch; add redundancy only with supported shared storage design | | monitoring | Pushgateway unavailable; Grafana/Alertmanager co-located on Titan-11; VictoriaMetrics singleton | Repair volume, show stale/unknown metrics explicitly, spread components, back up config/data appropriately and provide independent heartbeat; prove query freshness and alert delivery | | logging | OpenSearch pinned to offline Titan-05; heap exceeds memory request; Data Prepper restart evidence | Choose capacity, recover volume, align heap/request/limit, bound ingestion buffers and retention; verify log arrival and searchable recent entries rather than UI login only | | maintenance | Several overlapping cleaners/recovery actors; Soteria backup failures | One ownership map, bounded action budgets, recoverability checks and incident history; prove recovery halts after repeated failures and never silently deletes required data | | jenkins | Controller singleton; agents can exhaust compatible workers; controller/data on Titan-22 | Explicit agent concurrency and namespace budgets, protect controller resources and job state, bound caches; queued builds must not disrupt online applications | | quality | SonarQube stateful singleton; relies on DB/SSO/metrics path | Reserve JVM memory, persist and back up DB/config, tolerate exporter failure; test analysis job, UI and recovery independently of incomplete telemetry | | hermes | Many small-reservation workers/sidecars; stuck tenant-2 workspace; suite ledger local-path; GPU services compete | Fix workspace; measured tenant/agent limits and queues; isolate control API from model execution; preserve approved local job checkpoints and bounded retention; test recovery with synthetic metadata only | | hermes-scm | Singleton broker with PVC on Titan-11 | Preserve broker authorization and local state; spread from unstable node, define safe interrupted-operation recovery; no added arbitrary execution or broader credentials | | ai | Ollama tied to GPU nodes/local model storage; batch service parked | Explicit GPU lease ownership and admission; model-cache recoverability; approved local secondary backend only if capacity exists, otherwise honest queued/unavailable state; never introduce cloud fallback | | bstein-dev-home | Singleton front/back end and chat gateway | Replicate stateless parts where compatible, keep shared DB/auth availability and route boundaries; test page load plus read-only API path, not only container status | | cassandra | Application/simulation service, its own PostgreSQL and artifacts; Titan-23 intentionally reserved | Preserve simulations budget; back up application DB/artifacts, separate front-end availability from job capacity, verify job interruption semantics; pending cache PVC is WFFC, not automatically faulty | | comms | Matrix, auth, Redis, LiveKit, TURN and web clients mostly singleton; many old Jobs | Protect Matrix DB/media/auth, use protocol-aware readiness and supported scaling; preserve TCP/UDP routes; test messaging/call establishment separately in an approved synthetic account test | | mailu-mailserver | Most components singleton; mail data and queues stateful; alerts depend on mail | Reserve front/postfix/dovecot resources and storage; back up mailboxes/config, monitor queue age and dependency health; test SMTP/IMAP and controlled mail flow only after approval; do not blindly run duplicate mailbox writers | | nextcloud | Stateful Nextcloud singleton on busy Titan-07; Collabora on Titan-24 | Reserve PHP/DB/Redis capacity, verify storage and background jobs; consistent DB/files backup, supported multi-instance design if required; prove restore and document office-session limits | | outline | App and Redis singletons; DB/SSO/storage dependencies | Measured resources and session behavior, protect attachments/DB, replicate web layer if safe; verify document access and recovery with synthetic data | | planka | Singleton app with stateful attachments/DB dependency | Separate stateless capacity from persisted data, backup/restore, app-level health; validate board read/write and restart in a synthetic test | | finance | Firefly repeatedly unready; Actual Budget singleton with encrypted storage | Diagnose Firefly readiness safely, check DB/secret/probe path; preserve encrypted-storage keys and application backups; validate without reading financial records | | health | Wger singleton on previously unreliable Titan-22 | Move or provide dependable compatible failover capacity, protect uploaded media/database; application-level health and tested restore | | vaultwarden | Password-manager singleton and PVC on Titan-11 | Prioritize consistent encrypted vault/attachment backups and key recovery, healthy placement and safe restart; do not change auth/SSO integration as part of resource repair | | jellyfin | Jellyfin/Pegasus singletons and media storage; GPU/network reliance | Reserve service resources, separate media transcode queue from web availability; preserve media mount availability and session state; verify supported alternate execution without stealing assigned GPU leases | | game-stream | Wolf parked following GPU allocation; auth proxy availability does not mean streaming works | Record explicit parked state; allocate a stable GPU schedule/capacity before reactivation; test actual streaming path later; no silent reversal of the suite workload's allocation | | crypto | Node/wallet state, a per-node miner; P2Pool and test wallet parked | Keep miners under explicit residual-capacity budgets so services remain available, back up private wallet state through approved handling, document parked capabilities; do not inspect wallets or transact during audit | | climate | Typhon singleton and sync helper | Define hardware/external dependencies, modest reserved resources, safe recovery and stale-data indication; inspect device-control behavior before allowing duplicate active writers | | sui-metrics | Singleton collector | Reserve small resources, ensure reschedulability, distinguish external-source outage from collector failure and preserve metric freshness | | default / legacy objects | Old debug pods and oauth2-proxy-zot service with no endpoints | Inventory ownership; retire only confirmed obsolete objects through reviewed Git changes; preserve incident evidence and any required registry compatibility | | CronJobs / bootstrap Jobs across namespaces | Old objects, suspended schedules and bootstrap migrations mingle with live-service health | Classify one-shot versus recurring work; set deadline/concurrency/retention and last-success checks; keep migrations explicit and prevent resubmission on routine reconciles | ## Capacity choices that need an explicit decision 1. **Keep current node reservations:** repair 04/05/06 and runtime media, right-size workloads and cap burst work. Demonstrate the compatible Pi pool can lose one node. If it cannot, this option cannot promise all-service continuity by configuration alone. 2. **Approve bounded x86 general-service capacity:** use a dedicated resource reservation on existing reliable hardware, preserving simulation/GPU budgets and validating ARM/x86 images first. This changes the existing placement policy and needs explicit approval. 3. **Add stable general-purpose capacity:** size SSD-backed workers from measured demand plus N+1 headroom. This avoids depending on repaired low-capacity nodes or taking reserved simulation/GPU resources, but has procurement and migration cost. No specific purchase or final node size is justified by this snapshot alone. The plan does not remove a service to make the health dashboard green. Services that cannot run concurrently within the approved hardware envelope need a visible queue/schedule or additional capacity, with their availability impact agreed explicitly.