atlas-iac/docs/titan_20_21_22_capacity_optimization.md
Hermes Agent fbb401e4d5 monitoring(titan): add capacity guardrails for titan-20/21/22 audit
Evidence-led capacity/QoS audit of titan-20 (Hermes LLM fallback/
classifier), titan-21 (STT/TTS), and titan-22 (Jellyfin media-primary,
restored) for t_26da4c88. titan-20/21 are CPU-committed with no safe
headroom (titan-20 at 227 MiB free memory at its 24h worst point);
titan-22 has real idle CPU/RAM but its shared-GPU time-slicing has no
VRAM/engine isolation, so no workload is relocated. Adds alerting for
the sharpest gaps found (titan-20 memory exhaustion, titan-22 CPU/RAM/
GPU-VRAM pressure, Jellyfin CPU throttling, titan-21 CPU pressure) and
two Atlas GPU dashboard panels (VRAM, NVENC/NVDEC utilization) so a
future opportunistic-workload PR or a live transcode incident is
visible without a promql session.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-23 14:42:22 +00:00

21 KiB

Titan 20/21/22 capacity and placement optimization

Status: evidence-led audit + guardrails shipped. No workload was relocated onto titan-22 in this change. This document is the durable record for Kanban card t_26da4c88.

Evidence collected 2026-08-23 13:00-13:30 UTC via:

  • kubectl as system:serviceaccount:hermes:hermes-agent (per-namespace get/list/watch on pods/deployments/statefulsets/daemonsets; no mutate verbs at all — confirmed with kubectl auth can-i --list).
  • VictoriaMetrics, proxied read-only through Grafana's public datasource proxy (https://metrics.bstein.dev/api/datasources/proxy/uid/atlas-vm/...), covering the trailing 24h window ending ~2026-08-23T13:15Z. This is the longest valid window since titan-22's return to service (node manifests show the current jellyfin restore commit reconciling from 2026-08-23T01:11Z, i.e. the full 24h window is post-restore).
  • Git history (services/**, infrastructure/**) at origin/main@a019ecd55 (this branch was fast-forwarded from a stale base to pick up the in-flight titan-22 restoration commits before any analysis below).

No pods, deployments, taints, or labels were changed live. No kubectl port-forward/exec/mutate was available or attempted.

1. Node inventory (live)

Node Arch CPU alloc Mem alloc GPU Taint Labels of note
titan-20 arm64 Jetson Xavier 6 14.9 GiB nvidia.com/gpu.shared: 4 none jetson=true, accelerator=nvidia
titan-21 arm64 Jetson Xavier 6 14.9 GiB nvidia.com/gpu.shared: 4 none jetson=true, accelerator=nvidia, no longhorn-host
titan-22 amd64 (RTX 3050 Ti Laptop, 8 GiB VRAM) 20 31.1 GiB nvidia.com/gpu: 0, nvidia.com/gpu.shared: 4 atlas.bstein.dev/media-primary=true:PreferNoSchedule atlas.bstein.dev/media-primary=true, atlas.bstein.dev/general-compute=last-resort, jellyfin=true, longhorn-host=true

titan-22 really does report nvidia.com/gpu: 0 — confirmed live. It is a shared-GPU-only provider, as the card's snapshot assumed. DCGM identifies the card as an RTX 3050 Ti Laptop GPU, not the desktop 3050 8 GB the card text assumed; usable VRAM is ~7.95 GiB net of driver overhead (DCGM_FI_DEV_FB_FREE

  • DCGM_FI_DEV_FB_USED ≈ 7755 MiB at idle).

The three node labels/taints on titan-22 are produced every minute by infrastructure/core/node-prefer-noschedule-cronjob.yaml, a kubectl-based reconciler CronJob. Its own comment states the design intent directly: "Keep spare capacity available without letting routine pods displace media service." atlas.bstein.dev/general-compute=last-resort already exists for exactly the Phase B "Option B" question this card asks — but no workload in the repo currently targets or tolerates it. The taint is PreferNoSchedule (soft), so nothing is technically blocked from landing on titan-22 today; in practice nothing chooses to, because every other node in the same pools usually has room. This matters for the Jenkins finding in §5.

2. Workload placement and shared-GPU claims (live pods, Git-declared requests/limits)

Node Workload Priority class CPU req/lim Mem req/lim GPU
titan-20 ollama (ai-llm; Hermes local Qwen fallback + classifier) default (0) 4 / 8 10Gi / 14Gi gpu.shared: 1
titan-21 hermes-stt (Whisper small) default (0) 2 / 6 4Gi / 10Gi gpu.shared: 1
titan-21 hermes-tts (Piper, CPU-only) default (0) 1 / 4 512Mi / 2Gi none
titan-22 jellyfin media-core (400) 2 / 8 2Gi / 8Gi gpu.shared: 1
titan-22 DaemonSets (dcgm-exporter, node-exporter, nvidia-process-exporter, fluent-bit, metis-sentinel-amd64, node-nofile, node-image-sweeper, titan-22-link-keeper) default ~0.3 total ~0.9Gi total none

jellyfin's live pod spec already carries priorityClassName: media-core, runtimeClassName: nvidia, a hard nodeSelector: {kubernetes.io/hostname: titan-22}, and an explicit toleration for the media-primary taint — this is in the current services/jellyfin/deployment.yaml on origin/main (commit eec373dd5, "jellyfin: run media service on titan-22"). The branch this card started from was 89 commits behind origin/main and still had the pre-restore Jellyfin manifest (no GPU claim, no priority class, titan-24- preferring affinity); it was fast-forwarded before any of the analysis above. If you are reading a stale checkout of this repo, re-fetch before trusting any "Jellyfin isn't protected" read — it already is, as of PR #48's chain.

services/hermes/execution-worker-statefulset.yaml, execution-mediator.yaml, and switchyard-deployment.yaml all carry priorityClassName: scavenger (value -10) and a required node anti-affinity excluding titan-22/titan-24 (plus the storage-backbone nodes). This exclusion has no capacity justification found in this investigation — it looks like a defensive placement rule, not evidence that those workloads don't fit. It is left untouched per the card's explicit instruction not to relax it without workload-level evidence, but it is worth a follow-up card if Hermes execution capacity is ever tight (it currently is not — Hermes workers are healthy on their existing arm64 pool).

Shared-GPU device plugin (infrastructure/modules/profiles/components/device-plugin-config/configmap.yaml, mounted by all three per-node nvidia-device-plugin-* DaemonSets) uses:

sharing:
  timeSlicing:
    renameByDefault: true
    resources:
      - name: nvidia.com/gpu
        replicas: 4

This is plain NVIDIA time-slicing, not MPS and not MIG. It multiplies one physical device into 4 schedulable nvidia.com/gpu.shared slots so the Kubernetes scheduler can bin-pack GPU-requesting pods, but the driver time-slices SM access across whichever processes are co-resident with no priority, no VRAM isolation, and no per-tenant QoS. This is the load- bearing fact for the Option C verdict in §4.

3. 24h evidence (VictoriaMetrics, node_exporter/DCGM/kube-state-metrics)

CPU = 100% - idle%, averaged over 5m windows; percentiles are quantile_over_time(...)[24h:5m].

Node CPU p50 CPU p95 CPU p99 CPU max Mem now Mem p95/max Mem min-available
titan-20 68.2% 88.4% 96.9% 100%* 98.3% 98.2% / 98.5% 227 MiB of 14.9 GiB
titan-21 50.3% 59.9% 77.5% 61.1% 61.1% / 63.1% ~5.5 GiB
titan-22 2.8% 26.1% 47.0% 100%* 10.2% 10.9% / 11.0% ~27.8 GiB

* brief single-sample spikes; not sustained (p99 is the honest ceiling).

Declared kube_pod_container_resource_requests summed per node (live, includes DaemonSets):

Node CPU requested / allocatable Mem requested / allocatable
titan-20 4.88 / 6 (81%) 11.6 / 14.9 GiB (78%)
titan-21 5.58 / 6 (93%) 7.0 / 14.9 GiB (47%)
titan-22 5.20 / 20 (26%) 3.9 / 31.1 GiB (12.5%)

Two things follow directly from this table:

  1. titan-20 and titan-21 have essentially no scheduling headroom left, independent of the usage numbers above — the scheduler already considers 81% and 93% of their CPU committed. Adding anything there, even something that "should" be idle most of the time, risks failing to schedule at all or forcing evictions. This is true before accounting for the fact that titan-20's actual memory usage (98.3%) already runs well above its summed requests (78%), because ollama's container burst headroom (10Gi request → 14Gi limit) is being used in practice. titan-20 is the tightest resource in this entire audit and deserves its own alert (§6); it is not caused by this card and is not something this card's scope authorizes fixing (Hermes Chat first-pass model quality is explicitly protected — reducing ollama's memory/CPU footprint is a quality trade-off decision for Brad, not a capacity optimization).
  2. titan-22 has ~14.8 CPU and ~27 GiB of committed-but-unrequested headroom, and observed usage confirms it: p95 CPU is 26%, p50 is under 3%. Disk: 16.3% of ~915 GiB root filesystem used (~768 GiB free — this is local/ephemeral storage; Jellyfin's config/media/cache volumes are Longhorn PVCs and an emptyDir, not counted in that headroom). Network: 24h peak combined rx+tx ≈ 52 KB/s, i.e., not network-bound.

GPU (titan-22, DCGM 10.42.32.14:9400, 24h)

Metric p95 max
SM utilization 0% 3%
Encoder (NVENC) utilization 27%
Decoder (NVDEC) utilization 22%
Frame-buffer (VRAM) used 1276 MiB of ~7955 MiB

The encoder/decoder activity (peaking at 27%/22%, VRAM peaking at 1.27 GiB) is real evidence of at least one active hardware transcode/playback session inside the 24h window — the GPU is not idle-because-untested, it is idle-because-Jellyfin-usage-is-genuinely-light-relative-to-capacity right now. jellyfin shows 0 restarts and 0 OOMKills over the window, and CPU throttling (container_cpu_cfs_throttled_periods_total ratio) is 0 for jellyfin, ollama, stt, tts, and model-gate over the full 24h — none of the workloads audited are limit-throttled today.

4. Phase B — ranked distributions

A. Preserve the dedicated pattern (fallback/classifier on titan-20, voice on titan-21, titan-22 Jellyfin-only). This is close to the already-live state and remains correct for titan-20/21: both are already at 81-93% of CPU committed and titan-20 is within ~230 MiB of physical memory exhaustion at its worst point in the last 24h. There is no safe way to add load to either node without either starving Hermes Chat's protected first-pass model quality or risking an OOM on the node hosting it. Kept as-is for titan-20/ titan-21.

B. titan-22 CPU/RAM for isolated, preemptible, non-GPU opportunistic work, with a reserved Jellyfin budget. Evidence-backed and safe as a policy: titan-22 has ~14.8 idle CPU and ~27 GiB idle RAM, Kubernetes pod-priority preemption is a real, well-defined mechanism for CPU/RAM (unlike GPU time-slicing — see Option C), and the repo already ships the exact preemptible priority class needed (scavenger, value -10, already below Jellyfin's media-core at 400) plus the exact node label this design was built for (atlas.bstein.dev/general-compute=last-resort). This is the selected direction — see §5 for what shipped and what deliberately did not.

C. Conditional shared-GPU use of titan-22 when media is idle. Rejected. The device plugin config (§2) is pure NVIDIA time-slicing: 4 renamed replicas of one physical GPU, no MPS, no MIG, no per-client priority or VRAM quota. Kubernetes priority/preemption operates at pod scheduling granularity (evict-and-reschedule), not at GPU-engine granularity — it cannot arbitrate two already-running pods' NVENC/NVDEC/SM access mid-session, which is exactly the failure mode that would degrade an active Jellyfin transcode. There is no "idle detector + lease" primitive in this cluster (no DRA, no custom GPU operator, no MPS control daemon) to make this enforceable rather than a static promise. Until the cluster runs a GPU sharing mode with real isolation (MPS with per-client memory/thread limits, or a device-plugin generation with priority-aware time-slicing), Option C cannot meet the card's own bar ("reject it if it cannot protect active Jellyfin transcodes"). Not implemented.

D. Broaden Jenkins's hardware=rpi5 affinity to use titan-22. Not implemented; explicitly deferred to t_39cf1905. Two independent findings support this:

  • t_39cf1905's own read-only diagnosis (2026-08-23 12:53-13:19Z, its board comments) already root-caused the live Jenkins Pending state as a self-inflicted ConfigMap-hash rollout deadlock on the existing 3-node rpi5 pool (titan-07/08/11), unrelated to titan-22 capacity, and already shipped a fix as PR #49. Broadening Jenkins to titan-22 would not have fixed the actual problem.
  • Even setting that aside, moving Jenkins itself onto titan-22 is a poor fit for "opportunistic": Jenkins builds (especially Docker-in-Docker image builds) are bursty and CPU-hungry by design — the opposite of the tightly-bounded, preemptible workload profile Option B requires next to an active media host. It also has architecture (needs multi-arch image verification for jenkins/jenkins:2.528.3-jdk21 + plugin CLI on amd64), PVC-locality (three RWO PVCs currently only ever scheduled on the rpi5 pool), and recovery-isolation implications (the NotIn [titan-13,15,17,19] rule exists to keep Jenkins off storage-backbone nodes during Longhorn rebuilds — titan-22 is also a longhorn-host) that need their own evidence pass. This is a manual product decision for Brad via t_39cf1905, not something this card should change. See §7 for the board comment delivered.

5. Phase C — what shipped

No existing workload was relocated onto titan-22. Per the card's own acceptance criteria ("If no safe redistribution is proven, deliver the report/dashboard/guardrails only; do not move workloads for theoretical utilization"), and because this session's hermes-agent identity has zero mutate verbs on the live cluster (confirmed: kubectl auth can-i --list returns only get/list/watch across every resource type — there is no way to even trigger a live synthetic transcode to validate a placement change from this session), moving a real workload's placement without being able to observe the result live would violate the card's own validation bar ("New placement demonstrates Jellyfin remains healthy under a representative active transcode while the opportunistic workload runs, or does not ship"). Guardrails only:

  1. Alert rules (services/monitoring/vmalert-atlas-availability.yaml, new atlas.titan-media-capacity group): Jellyfin CPU/memory pressure at its own limits, titan-22 GPU VRAM saturation, titan-22 CPU/RAM near-exhaustion, CPU throttling on any of the audited media/AI containers, and — because it was the sharpest finding in this audit — titan-20 node memory exhaustion (it is currently 227 MiB from the edge with no alert covering it at all).
  2. Dashboard: two new panels on Atlas GPU (services/monitoring/grafana-dashboard-gpu.yaml, regenerated from scripts/render/dashboards_render_atlas.py) showing titan-22 VRAM used/free and encoder/decoder utilization by node, so a future opportunistic-workload PR (or a live transcode incident) is visible without a promql session.
  3. This document, as the durable capacity model, workload map, and validation plan referenced by the acceptance criteria.

What was deliberately not shipped, and why

A PriorityClass/quota primitive for Option B is not "new" — scavenger already exists and is already below media-core. What's missing is an actual opportunistic workload to place on titan-22 with it. Every existing CPU-only, architecture-portable, non-Jellyfin workload examined during this audit (Jenkins, Cassandra/Veles backends, collabora, CI/quality-gate jobs) either has its own node-pool pinning for reasons outside this card's scope, or — like Jenkins — has an incident already being independently worked. Manufacturing a placement change for a workload not already flagged as portable would be exactly the "blind migration" and "theoretical utilization" move the card prohibits. The safe next step is a follow-up card that names one specific, already-portable batch/CI workload and proves it end-to-end (including the live transcode-under-load validation this session cannot perform), using the reservation budget below.

Reservation budget for that follow-up (documented contract, not enforced by a new object)

Any future opportunistic workload placed on titan-22 must:

  • use priorityClassName: scavenger (preemptible below Jellyfin's media-core);
  • request zero nvidia.com/gpu* resources (Option C is rejected — see §4);
  • keep summed CPU/memory requests across all such workloads at or below 10 CPU / 18 GiB, which — added to Jellyfin's own limit ceiling (8 CPU / 8 GiB) and the ~0.3 CPU / 0.9 GiB of existing DaemonSets — stays within titan-22's 20 CPU / 31.1 GiB allocatable even if Jellyfin is transcoding at its full burst limit at the same time;
  • set a real CPU limit (not just a request) so a burst cannot starve Jellyfin's CFS shares even transiently.

Kubernetes has no native per-node ResourceQuota primitive, so this budget is a documented contract enforced by the new alerts in §6, not by an admission object. A namespace-scoped ResourceQuota keyed to scavenger pods was considered and rejected: it cannot be scoped to a specific node, so it would either do nothing (if the opportunistic workload's namespace also runs elsewhere) or wrongly cap unrelated scavenger workloads on other nodes (Hermes execution workers already use scavenger on the arm64 pool).

6. New alerts (services/monitoring/vmalert-atlas-availability.yaml)

Alert Fires when Why
Titan20NodeMemoryCritical titan-20 MemAvailable < 5% of MemTotal for 10m Sharpest finding in this audit: 227 MiB free at the worst point in 24h, no existing coverage.
TitanMediaHostCPUPressure titan-22 CPU busy > 85% for 10m Node-level early warning before Jellyfin's own limit-throttling would show up.
TitanMediaHostMemoryPressure titan-22 MemAvailable < 15% of MemTotal for 10m Same, for memory; also the number that must stay large if a future opportunistic workload lands here.
JellyfinCPUThrottled container_cpu_cfs_throttled_periods_total ratio > 5% for container="jellyfin" over 5m Direct transcode-quality signal — sustained throttling degrades encode.
TitanMediaGPUMemorySaturation titan-22 DCGM FB_USED / (FB_USED+FB_FREE) > 85% for 5m VRAM exhaustion is the concrete failure mode a second shared-GPU tenant would cause (Option C).
TitanVoiceHostCPUPressure titan-21 CPU busy > 90% for 15m titan-21 is already at 93% CPU requested; this catches sustained real contention that could add STT/TTS latency for Hermes Chat.

All use the existing atlas-vm VictoriaMetrics datasource/label set already scraped in this cluster (node_exporter, DCGM, cAdvisor container_cpu_cfs_*) — no new exporters required.

7. Board comments delivered

  • t_39cf1905 (Jenkins/WebUI release): told not to broaden Jenkins onto titan-22 as a workaround for the Pending state — their own root cause (ConfigMap-hash rollout deadlock on the rpi5 pool) is unrelated to titan-22 capacity, and titan-22 is a poor architectural fit for bursty CI/DinD builds next to a media-primary host with no enforceable GPU isolation.
  • t_60d2deb6 (Chat latency instrumentation): told that titan-20 (LLM fallback/classifier) is CPU/RAM-saturated (81% CPU requested, 98.3% memory used, 227 MiB free at worst) and titan-21 (STT/TTS) is CPU-committed at 93% though usage is lighter (50-60%) — so any voice-path latency variance they measure should be checked against titan-20/21 node pressure before being attributed to application code, and neither node has spare capacity for a node-level latency mitigation; also reminded that titan-22 is amd64/no-arm64-GPU and therefore not a relevant relocation target for any Jetson-pinned STT/TTS/classifier component.

8. Validation plan (live run pending — this session has no mutate access)

This session cannot apply the alert/dashboard PR, cannot trigger a synthetic transcode, and cannot restart or exec into jellyfin (RBAC is get/list/ watch only, confirmed). The plan below is what should be executed once the PR in §9 is merged and reconciled:

  1. Confirm the two new GPU dashboard panels and the six new alerts render (Atlas GPU dashboard; vmalert-atlas-availability targets in Grafana Alerting) with no parse error in vmalert's own logs.
  2. Start (or wait for) a real hardware-transcoding Jellyfin playback session (a client requesting a bitrate/codec the source doesn't natively support forces NVENC/NVDEC use). While it runs, confirm on the new panels: DCGM encoder/decoder utilization rises, jellyfin stays Ready, and none of the six new alerts fire.
  3. As a synthetic corroboration in the same window, run a short (~5 minute) CPU/memory-only load generator Job on titan-22 (priorityClassName: scavenger, request e.g. 4 CPU / 4Gi, no GPU resource — stress-ng or similar) and confirm: it schedules, Jellyfin's playback in step 2 shows no stutter/dropped-frame/error in its own logs, and TitanMediaHostCPUPressure/TitanMediaHostMemoryPressure do not fire (they are tuned to the node ceiling, not to this Job's own request, so a well-behaved opportunistic Job should not trip them).
  4. Record the result (pass/fail + screenshots or exported panel PNGs) as a comment on this card's board entry, or on the follow-up card once one exists, before any real opportunistic workload is proposed for titan-22.