atlas-iac/docs/titan_20_21_22_capacity_optimization.md
Hermes Agent 4f48a1d3de docs(titan): fold in live Jenkins/VM/Wger placement onto titan-22
Brad placed Jenkins, VictoriaMetrics, and Wger onto titan-22 directly
(main commits 271f3e8c3, 0ebab9d41) between this audit's first pass and
PR #50 going up. Revalidate with fresh 3h evidence instead of shipping a
report that describes a since-superseded snapshot: titan-22 headroom is
materially unchanged (22-26% CPU p95, mem 77%+ available) with the three
new residents, Jellyfin shows 0 restarts/throttling throughout, but no
GPU/transcode activity occurred in the window so the transcode-under-load
validation is still outstanding. Also flags VictoriaMetrics now being a
single point of failure for titan-22 observability (hard-pinned, no
fallback), and updates the board-comment record (§7) to match what was
actually posted to t_39cf1905/t_60d2deb6.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-23 14:56:43 +00:00

27 KiB

Titan 20/21/22 capacity and placement optimization

Status: evidence-led audit + guardrails shipped. No workload was relocated onto titan-22 by this card/PR. Between this audit's first pass and this revision, Brad independently placed three workloads (Jenkins, VictoriaMetrics, Wger) onto titan-22 via direct commits to main (outside this card's PR), implementing almost exactly the Option B pattern this document recommends. §4a documents that live change and the fresh evidence validating it. This document is the durable record for Kanban card t_26da4c88.

Evidence collected in two passes:

  • First pass 2026-08-23 13:00-13:30 UTC, at origin/main@a019ecd55, covering a trailing-24h VictoriaMetrics window ending ~13:15Z (the longest valid window since titan-22's return to service — node manifests show the current jellyfin restore commit reconciling from 2026-08-23T01:11Z).
  • Revalidation pass 2026-08-23 ~14:40-14:45 UTC, after git fetch --no-tags origin main showed this branch was 4 commits behind a fast-moving main (origin/main@42795f3d6); the branch was rebased onto it (clean, no file overlap with the audit's own commit). This pass queried VictoriaMetrics directly (victoria-metrics-single-server.monitoring.svc.cluster.local:8428, in-cluster) for current + trailing-3h data, since 3 of the 4 new commits changed placement for Jenkins/VictoriaMetrics/Wger only ~1-3h before this pass — see §4a.

Both passes used kubectl as system:serviceaccount:hermes:hermes-agent (per-namespace get/list/watch on pods/deployments/statefulsets/daemonsets; no mutate verbs at all — confirmed with kubectl auth can-i --list, both passes). No pods, deployments, taints, or labels were changed live by this card. No kubectl port-forward/exec/mutate was available or attempted.

1. Node inventory (live)

Node Arch CPU alloc Mem alloc GPU Taint Labels of note
titan-20 arm64 Jetson Xavier 6 14.9 GiB nvidia.com/gpu.shared: 4 none jetson=true, accelerator=nvidia
titan-21 arm64 Jetson Xavier 6 14.9 GiB nvidia.com/gpu.shared: 4 none jetson=true, accelerator=nvidia, no longhorn-host
titan-22 amd64 (RTX 3050 Ti Laptop, 8 GiB VRAM) 20 31.1 GiB nvidia.com/gpu: 0, nvidia.com/gpu.shared: 4 atlas.bstein.dev/media-primary=true:PreferNoSchedule atlas.bstein.dev/media-primary=true, atlas.bstein.dev/general-compute=last-resort, jellyfin=true, longhorn-host=true

titan-22 really does report nvidia.com/gpu: 0 — confirmed live. It is a shared-GPU-only provider, as the card's snapshot assumed. DCGM identifies the card as an RTX 3050 Ti Laptop GPU, not the desktop 3050 8 GB the card text assumed; usable VRAM is ~7.95 GiB net of driver overhead (DCGM_FI_DEV_FB_FREE

  • DCGM_FI_DEV_FB_USED ≈ 7755 MiB at idle).

The three node labels/taints on titan-22 are produced every minute by infrastructure/core/node-prefer-noschedule-cronjob.yaml, a kubectl-based reconciler CronJob. Its own comment states the design intent directly: "Keep spare capacity available without letting routine pods displace media service." atlas.bstein.dev/general-compute=last-resort already exists for exactly the Phase B "Option B" question this card asks — but no workload in the repo currently targets or tolerates it. The taint is PreferNoSchedule (soft), so nothing is technically blocked from landing on titan-22 today; in practice nothing chooses to, because every other node in the same pools usually has room. This matters for the Jenkins finding in §5.

2. Workload placement and shared-GPU claims (live pods, Git-declared requests/limits)

Node Workload Priority class CPU req/lim Mem req/lim GPU
titan-20 ollama (ai-llm; Hermes local Qwen fallback + classifier) default (0) 4 / 8 10Gi / 14Gi gpu.shared: 1
titan-21 hermes-stt (Whisper small) default (0) 2 / 6 4Gi / 10Gi gpu.shared: 1
titan-21 hermes-tts (Piper, CPU-only) default (0) 1 / 4 512Mi / 2Gi none
titan-22 jellyfin media-core (400) 2 / 8 2Gi / 8Gi gpu.shared: 1
titan-22 jenkins (added live, see §4a) scavenger (-10) 0.1 / 1.5 1Gi / 3Gi none
titan-22 victoria-metrics-single-server (added live, see §4a) scavenger (-10) 0.5 / 2 2Gi / 4Gi none
titan-22 wger (added live, see §4a) scavenger (-10) 0.25 / 1.2 0.6Gi / 2.25Gi none
titan-22 DaemonSets (dcgm-exporter, node-exporter, nvidia-process-exporter, fluent-bit, metis-sentinel-amd64, node-nofile, node-image-sweeper, titan-22-link-keeper) default ~0.3 total ~0.9Gi total none

jellyfin's live pod spec already carries priorityClassName: media-core, runtimeClassName: nvidia, a hard nodeSelector: {kubernetes.io/hostname: titan-22}, and an explicit toleration for the media-primary taint — this is in the current services/jellyfin/deployment.yaml on origin/main (commit eec373dd5, "jellyfin: run media service on titan-22"). The branch this card started from was 89 commits behind origin/main and still had the pre-restore Jellyfin manifest (no GPU claim, no priority class, titan-24- preferring affinity); it was fast-forwarded before any of the analysis above. If you are reading a stale checkout of this repo, re-fetch before trusting any "Jellyfin isn't protected" read — it already is, as of PR #48's chain.

services/hermes/execution-worker-statefulset.yaml, execution-mediator.yaml, and switchyard-deployment.yaml all carry priorityClassName: scavenger (value -10) and a required node anti-affinity excluding titan-22/titan-24 (plus the storage-backbone nodes). This exclusion has no capacity justification found in this investigation — it looks like a defensive placement rule, not evidence that those workloads don't fit. It is left untouched per the card's explicit instruction not to relax it without workload-level evidence, but it is worth a follow-up card if Hermes execution capacity is ever tight (it currently is not — Hermes workers are healthy on their existing arm64 pool).

Shared-GPU device plugin (infrastructure/modules/profiles/components/device-plugin-config/configmap.yaml, mounted by all three per-node nvidia-device-plugin-* DaemonSets) uses:

sharing:
  timeSlicing:
    renameByDefault: true
    resources:
      - name: nvidia.com/gpu
        replicas: 4

This is plain NVIDIA time-slicing, not MPS and not MIG. It multiplies one physical device into 4 schedulable nvidia.com/gpu.shared slots so the Kubernetes scheduler can bin-pack GPU-requesting pods, but the driver time-slices SM access across whichever processes are co-resident with no priority, no VRAM isolation, and no per-tenant QoS. This is the load- bearing fact for the Option C verdict in §4.

3. 24h evidence (VictoriaMetrics, node_exporter/DCGM/kube-state-metrics)

CPU = 100% - idle%, averaged over 5m windows; percentiles are quantile_over_time(...)[24h:5m].

Node CPU p50 CPU p95 CPU p99 CPU max Mem now Mem p95/max Mem min-available
titan-20 68.2% 88.4% 96.9% 100%* 98.3% 98.2% / 98.5% 227 MiB of 14.9 GiB
titan-21 50.3% 59.9% 77.5% 61.1% 61.1% / 63.1% ~5.5 GiB
titan-22 2.8% 26.1% 47.0% 100%* 10.2% 10.9% / 11.0% ~27.8 GiB

* brief single-sample spikes; not sustained (p99 is the honest ceiling).

Declared kube_pod_container_resource_requests summed per node (live, includes DaemonSets):

Node CPU requested / allocatable Mem requested / allocatable
titan-20 4.88 / 6 (81%) 11.6 / 14.9 GiB (78%)
titan-21 5.58 / 6 (93%) 7.0 / 14.9 GiB (47%)
titan-22 5.20 / 20 (26%) 3.9 / 31.1 GiB (12.5%)

Two things follow directly from this table:

  1. titan-20 and titan-21 have essentially no scheduling headroom left, independent of the usage numbers above — the scheduler already considers 81% and 93% of their CPU committed. Adding anything there, even something that "should" be idle most of the time, risks failing to schedule at all or forcing evictions. This is true before accounting for the fact that titan-20's actual memory usage (98.3%) already runs well above its summed requests (78%), because ollama's container burst headroom (10Gi request → 14Gi limit) is being used in practice. titan-20 is the tightest resource in this entire audit and deserves its own alert (§6); it is not caused by this card and is not something this card's scope authorizes fixing (Hermes Chat first-pass model quality is explicitly protected — reducing ollama's memory/CPU footprint is a quality trade-off decision for Brad, not a capacity optimization).
  2. titan-22 has ~14.8 CPU and ~27 GiB of committed-but-unrequested headroom, and observed usage confirms it: p95 CPU is 26%, p50 is under 3%. Disk: 16.3% of ~915 GiB root filesystem used (~768 GiB free — this is local/ephemeral storage; Jellyfin's config/media/cache volumes are Longhorn PVCs and an emptyDir, not counted in that headroom). Network: 24h peak combined rx+tx ≈ 52 KB/s, i.e., not network-bound.

GPU (titan-22, DCGM 10.42.32.14:9400, 24h)

Metric p95 max
SM utilization 0% 3%
Encoder (NVENC) utilization 27%
Decoder (NVDEC) utilization 22%
Frame-buffer (VRAM) used 1276 MiB of ~7955 MiB

The encoder/decoder activity (peaking at 27%/22%, VRAM peaking at 1.27 GiB) is real evidence of at least one active hardware transcode/playback session inside the 24h window — the GPU is not idle-because-untested, it is idle-because-Jellyfin-usage-is-genuinely-light-relative-to-capacity right now. jellyfin shows 0 restarts and 0 OOMKills over the window, and CPU throttling (container_cpu_cfs_throttled_periods_total ratio) is 0 for jellyfin, ollama, stt, tts, and model-gate over the full 24h — none of the workloads audited are limit-throttled today.

4. Phase B — ranked distributions

A. Preserve the dedicated pattern (fallback/classifier on titan-20, voice on titan-21, titan-22 Jellyfin-only). This is close to the already-live state and remains correct for titan-20/21: both are already at 81-93% of CPU committed and titan-20 is within ~230 MiB of physical memory exhaustion at its worst point in the last 24h. There is no safe way to add load to either node without either starving Hermes Chat's protected first-pass model quality or risking an OOM on the node hosting it. Kept as-is for titan-20/ titan-21.

B. titan-22 CPU/RAM for isolated, preemptible, non-GPU opportunistic work, with a reserved Jellyfin budget. Evidence-backed and safe as a policy: titan-22 has ~14.8 idle CPU and ~27 GiB idle RAM, Kubernetes pod-priority preemption is a real, well-defined mechanism for CPU/RAM (unlike GPU time-slicing — see Option C), and the repo already ships the exact preemptible priority class needed (scavenger, value -10, already below Jellyfin's media-core at 400) plus the exact node label this design was built for (atlas.bstein.dev/general-compute=last-resort). This is the selected direction — see §5 for what shipped and what deliberately did not.

C. Conditional shared-GPU use of titan-22 when media is idle. Rejected. The device plugin config (§2) is pure NVIDIA time-slicing: 4 renamed replicas of one physical GPU, no MPS, no MIG, no per-client priority or VRAM quota. Kubernetes priority/preemption operates at pod scheduling granularity (evict-and-reschedule), not at GPU-engine granularity — it cannot arbitrate two already-running pods' NVENC/NVDEC/SM access mid-session, which is exactly the failure mode that would degrade an active Jellyfin transcode. There is no "idle detector + lease" primitive in this cluster (no DRA, no custom GPU operator, no MPS control daemon) to make this enforceable rather than a static promise. Until the cluster runs a GPU sharing mode with real isolation (MPS with per-client memory/thread limits, or a device-plugin generation with priority-aware time-slicing), Option C cannot meet the card's own bar ("reject it if it cannot protect active Jellyfin transcodes"). Not implemented.

D. Broaden Jenkins's hardware=rpi5 affinity to use titan-22. Superseded by a live change Brad made directly — see §4a. This audit's first pass (13:00-13:30Z) found Jenkins still rpi5-pinned and recommended deferring this decision to t_39cf1905 as a manual product call, for the reasons below (kept for the record — they were the right concerns to raise, even though Brad has since made the call):

  • t_39cf1905's own read-only diagnosis (2026-08-23 12:53-13:19Z, its board comments) root-caused the then-live Jenkins Pending state as a self-inflicted ConfigMap-hash rollout deadlock on the existing 3-node rpi5 pool (titan-07/08/11), unrelated to titan-22 capacity, and shipped a fix as PR #49. Broadening Jenkins to titan-22 would not have fixed that specific problem on its own.
  • Moving Jenkins onto titan-22 is a nontrivial fit for "opportunistic": Jenkins builds (especially Docker-in-Docker image builds) are bursty and CPU-hungry by design. It also has architecture (multi-arch image verification for jenkins/jenkins:2.528.3-jdk21 + plugin CLI on amd64), PVC-locality (three RWO PVCs previously only ever scheduled on the rpi5 pool), and recovery-isolation implications (the previous NotIn [titan-13,15,17,19] rule kept Jenkins off storage-backbone nodes during Longhorn rebuilds — titan-22 is also a longhorn-host).

4a. Live update: Brad already implemented Option D (and Option B) directly

Between this audit's first pass and this revision, three direct commits to mainnot through this card's PR, authored jenkins <brad@bstein.dev> — re-pinned Jenkins, VictoriaMetrics, and Wger onto titan-22:

Commit Workload What changed
271f3e8c3 "ops: spread saturated node workloads" jenkins nodeSelector.arch rpi5→amd64; required affinity NotIn[titan-13,15,17,19]+hardware=rpi5 → required hostname In [titan-22]; added priorityClassName: scavenger and a PreferNoSchedule toleration for the media-primary taint
271f3e8c3 (same commit) victoria-metrics-single-server required affinity hostname NotIn [titan-12,20,21,22,24,...]In [titan-22]; added priorityClassName: scavenger + media-primary toleration; upgrade.force: true
0ebab9d41 "placement(wger): use preemptible titan-22 capacity" wger same pattern: NotIn[titan-13,15,17,19]hostname In [titan-22], arch arm64→amd64, priorityClassName: scavenger, media-primary toleration

This is exactly the Option B pattern this document independently arrived at (preemptible scavenger priority, CPU/RAM only, zero GPU requests) — applied to Jenkins as well, which the first pass had flagged as needing a separate evidence-based decision (Option D). Brad made that call directly. None of these three request nvidia.com/gpu* (confirmed live). This is a direct-to-main change outside the "all durable changes via reviewable titan-iac PRs" policy in this card's own body — noted here for the record, not something this card can undo or should second-guess; it's the owner's prerogative.

Fresh evidence (14:40-14:45Z revalidation pass, ~1-3h of run time for the new placements) shows it is safe so far:

Node 3h CPU p95 3h CPU max Mem available now Mem available min (3h)
titan-20 88.6% 90.5% 2.6% 1.7%
titan-21 65.1% 69.3% 38.9% 38.5%
titan-22 22.1% 26.2% 77.4% 77.5%
  • titan-22 CPU/RAM headroom is essentially unchanged from the first pass despite three new residents — confirms Option B's "~15 CPU / ~27 GiB idle" finding was real spare capacity, not measurement noise.
  • jellyfin (titan-22): 0 restarts, 0 CPU-throttled periods, Ready, no pod events, over the full window these three workloads have been co-resident.
  • jenkins (titan-22): CPU-throttled up to 30% of periods at its own 1.5-CPU limit — this is the CFS quota doing its job (Jenkins capped at its own ceiling), not evidence of node-level contention; wger and victoria-metrics-single-server show negligible throttling (<1%).
  • titan-22 GPU (DCGM): 0% SM/encoder/decoder utilization and ~1 MiB VRAM used for the entire 3h window — no active transcode occurred during this co-residency window, so this is not yet a live proof that Jellyfin survives a concurrent hardware transcode with these three neighbors present. That validation (§8) is still outstanding and still requires either a real playback session or synthetic load-gen, neither of which this read-only session can trigger.
  • titan-20/21 numbers reconfirm the first pass and got worse, not better: titan-20 is now down to 1.7% memory available at its 3h floor (vs. "227 MiB / ~1.5%" at its 24h floor in the first pass) — Option A (leave titan-20/21 alone) remains the only safe call for those two nodes.

New risk this live change introduces, not present in the first pass's Option B analysis: victoria-metrics-single-server is the cluster's only metrics backend, and it is now hard-pinned (required affinity, not preferred) to the same node as Jellyfin, Jenkins, and Wger. If titan-22 has a node-level incident, the dashboards and alerts needed to diagnose that exact incident go dark at the same time (vmalert-atlas-availability and grafana themselves stay up, on titan-0b/titan-11, but have nothing to query). This wasn't evaluated against the card's Jellyfin-protection bar because it's an observability-resilience question, not a Jellyfin-quality one — flagging it as a follow-up decision for Brad rather than changing it unilaterally.

5. Phase C — what shipped

No workload was relocated onto titan-22 by this card's PR. Per the card's own acceptance criteria ("If no safe redistribution is proven, deliver the report/dashboard/guardrails only; do not move workloads for theoretical utilization"), and because this session's hermes-agent identity has zero mutate verbs on the live cluster (confirmed: kubectl auth can-i --list returns only get/list/watch across every resource type — there is no way to even trigger a live synthetic transcode to validate a placement change from this session), this card does not itself move a workload's placement without being able to observe the result live — that would violate the card's own validation bar ("New placement demonstrates Jellyfin remains healthy under a representative active transcode while the opportunistic workload runs, or does not ship").

As documented in §4a, Brad independently placed Jenkins, VictoriaMetrics, and Wger onto titan-22 via direct main commits during this audit — using almost exactly the Option B pattern below, before this PR existed to review it. This PR does not touch, revert, or re-author that change; it documents it (§4a, workload table in §2) and extends the guardrails to cover it. Guardrails shipped by this PR:

  1. Alert rules (services/monitoring/vmalert-atlas-availability.yaml, new atlas.titan-media-capacity group): Jellyfin CPU/memory pressure at its own limits, titan-22 GPU VRAM saturation, titan-22 CPU/RAM near-exhaustion, CPU throttling on any of the audited media/AI containers, and — because it was the sharpest finding in this audit — titan-20 node memory exhaustion (it is currently 227 MiB from the edge with no alert covering it at all).
  2. Dashboard: two new panels on Atlas GPU (services/monitoring/grafana-dashboard-gpu.yaml, regenerated from scripts/render/dashboards_render_atlas.py) showing titan-22 VRAM used/free and encoder/decoder utilization by node, so a future opportunistic-workload PR (or a live transcode incident) is visible without a promql session.
  3. This document, as the durable capacity model, workload map, and validation plan referenced by the acceptance criteria.

What was deliberately not shipped, and why

A PriorityClass/quota primitive for Option B was not "new" to invent — scavenger already existed and was already below media-core. This audit's first pass found no already-portable opportunistic workload to place on titan-22 with it (Jenkins, Cassandra/Veles backends, collabora, CI/quality-gate jobs all had their own node-pool pinning or an in-flight incident) and recommended a follow-up card rather than manufacturing a placement. Brad then placed three workloads (Jenkins, VictoriaMetrics, Wger) directly — see §4a — which is the concrete instance of that follow-up, done outside this card. What is still genuinely outstanding, and still cannot be done from this read-only session, is the live transcode-under-load validation (§8): none of titan-22's new residents' ~1-3h of co-residency so far has overlapped with an active hardware transcode.

Reservation budget (documented contract, not enforced by a new object)

Any opportunistic workload placed on titan-22 — including the three already live (§4a), which all comply — must:

  • use priorityClassName: scavenger (preemptible below Jellyfin's media-core);
  • request zero nvidia.com/gpu* resources (Option C is rejected — see §4);
  • keep summed CPU/memory requests across all such workloads at or below 10 CPU / 18 GiB, which — added to Jellyfin's own limit ceiling (8 CPU / 8 GiB) and the ~0.3 CPU / 0.9 GiB of existing DaemonSets — stays within titan-22's 20 CPU / 31.1 GiB allocatable even if Jellyfin is transcoding at its full burst limit at the same time;
  • set a real CPU limit (not just a request) so a burst cannot starve Jellyfin's CFS shares even transiently.

Kubernetes has no native per-node ResourceQuota primitive, so this budget is a documented contract enforced by the new alerts in §6, not by an admission object. A namespace-scoped ResourceQuota keyed to scavenger pods was considered and rejected: it cannot be scoped to a specific node, so it would either do nothing (if the opportunistic workload's namespace also runs elsewhere) or wrongly cap unrelated scavenger workloads on other nodes (Hermes execution workers already use scavenger on the arm64 pool).

6. New alerts (services/monitoring/vmalert-atlas-availability.yaml)

Alert Fires when Why
Titan20NodeMemoryCritical titan-20 MemAvailable < 5% of MemTotal for 10m Sharpest finding in this audit: 227 MiB free at the worst point in 24h, no existing coverage.
TitanMediaHostCPUPressure titan-22 CPU busy > 85% for 10m Node-level early warning before Jellyfin's own limit-throttling would show up.
TitanMediaHostMemoryPressure titan-22 MemAvailable < 15% of MemTotal for 10m Same, for memory; also the number that must stay large if a future opportunistic workload lands here.
JellyfinCPUThrottled container_cpu_cfs_throttled_periods_total ratio > 5% for container="jellyfin" over 5m Direct transcode-quality signal — sustained throttling degrades encode.
TitanMediaGPUMemorySaturation titan-22 DCGM FB_USED / (FB_USED+FB_FREE) > 85% for 5m VRAM exhaustion is the concrete failure mode a second shared-GPU tenant would cause (Option C).
TitanVoiceHostCPUPressure titan-21 CPU busy > 90% for 15m titan-21 is already at 93% CPU requested; this catches sustained real contention that could add STT/TTS latency for Hermes Chat.

All use the existing atlas-vm VictoriaMetrics datasource/label set already scraped in this cluster (node_exporter, DCGM, cAdvisor container_cpu_cfs_*) — no new exporters required.

7. Board comments delivered

  • t_39cf1905 (Jenkins/WebUI release): told that Jenkins is already running on titan-22 as of a direct main commit (271f3e8c3, ~13:37Z, authored jenkins <brad@bstein.dev>, not through this card), that it is CPU-only/scavenger-priority (correctly preemptible below Jellyfin's media-core) and healthy so far (0 restarts, self-limited throttling only, titan-22 still at 22-26% p95 CPU / 77%+ mem available with it resident) — so their Part B (trigger the WebUI image builder now that Jenkins is healthy) should be unblocked from the node-placement side; and warned that no active hardware transcode has yet coincided with Jenkins's presence, so Jellyfin-under-load-plus-Jenkins-build is still unproven, not proven-safe.
  • t_60d2deb6 (Chat latency instrumentation): told that titan-20 (LLM fallback/classifier) is CPU/RAM-saturated (88.6% CPU p95 over the last 3h, memory available down to 1.7% at its floor) and titan-21 (STT/TTS) is at 65% CPU p95 with 38-39% memory available — so any voice-path latency variance they measure should be checked against titan-20/21 node pressure before being attributed to application code, and neither node has spare capacity for a node-level latency mitigation; also that titan-22 is amd64/no-arm64-GPU and therefore not a relevant relocation target for any Jetson-pinned STT/TTS/classifier component, and that titan-22 now also hosts the cluster's only VictoriaMetrics instance (hard-pinned, single point of failure for the dashboards this instrumentation work will build on) as a heads-up, not a blocker.

8. Validation plan (live run pending — this session has no mutate access)

This session cannot apply the alert/dashboard PR, cannot trigger a synthetic transcode, and cannot restart or exec into jellyfin (RBAC is get/list/ watch only, confirmed). As of this revision, Jenkins/VictoriaMetrics/Wger have been co-resident on titan-22 for ~1-3h (§4a) with zero Jellyfin impact, but zero GPU activity in that window either — so the one validation that actually matters (an active hardware transcode overlapping with real opportunistic load) still has not happened naturally and still cannot be forced from here. The plan below is what should be executed once the PR in §9 is merged and reconciled (or opportunistically, the next time Jellyfin is in active use):

  1. Confirm the two new GPU dashboard panels and the six new alerts render (Atlas GPU dashboard; vmalert-atlas-availability targets in Grafana Alerting) with no parse error in vmalert's own logs.
  2. Start (or wait for) a real hardware-transcoding Jellyfin playback session (a client requesting a bitrate/codec the source doesn't natively support forces NVENC/NVDEC use). While it runs, confirm on the new panels: DCGM encoder/decoder utilization rises, jellyfin stays Ready, and none of the six new alerts fire.
  3. As a synthetic corroboration in the same window, run a short (~5 minute) CPU/memory-only load generator Job on titan-22 (priorityClassName: scavenger, request e.g. 4 CPU / 4Gi, no GPU resource — stress-ng or similar) and confirm: it schedules, Jellyfin's playback in step 2 shows no stutter/dropped-frame/error in its own logs, and TitanMediaHostCPUPressure/TitanMediaHostMemoryPressure do not fire (they are tuned to the node ceiling, not to this Job's own request, so a well-behaved opportunistic Job should not trip them).
  4. Record the result (pass/fail + screenshots or exported panel PNGs) as a comment on this card's board entry, or on the follow-up card once one exists, before any real opportunistic workload is proposed for titan-22.