Hermes Agent
fbb401e4d5
monitoring(titan): add capacity guardrails for titan-20/21/22 audit
...
Evidence-led capacity/QoS audit of titan-20 (Hermes LLM fallback/
classifier), titan-21 (STT/TTS), and titan-22 (Jellyfin media-primary,
restored) for t_26da4c88. titan-20/21 are CPU-committed with no safe
headroom (titan-20 at 227 MiB free memory at its 24h worst point);
titan-22 has real idle CPU/RAM but its shared-GPU time-slicing has no
VRAM/engine isolation, so no workload is relocated. Adds alerting for
the sharpest gaps found (titan-20 memory exhaustion, titan-22 CPU/RAM/
GPU-VRAM pressure, Jellyfin CPU throttling, titan-21 CPU pressure) and
two Atlas GPU dashboard panels (VRAM, NVENC/NVDEC utilization) so a
future opportunistic-workload PR or a live transcode incident is
visible without a promql session.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-23 14:42:22 +00:00
jenkins
46c50814ad
fix(monitoring): retain short Jetson GPU activity
2026-08-09 14:38:51 -03:00
jenkins
f2328f12e4
monitoring(gpu): attribute Jetson activity by allocation
2026-08-02 04:26:33 -03:00
jenkins
b1ecfe96e4
ai(hermes): add operator guide and current GPU shares
2026-08-02 03:59:32 -03:00
jenkins
df953359b9
monitoring(gpu): use process-level pod attribution
2026-08-02 03:31:02 -03:00
jenkins
3255936aeb
monitoring(gpu): report time-weighted namespace usage
2026-08-02 03:08:20 -03:00
jenkins
04175d33ab
monitoring(gpu): show activity share by namespace
2026-05-22 04:22:51 -03:00
jenkins
250a850871
monitoring(gpu): count monitored GPU pool devices
2026-05-22 03:23:36 -03:00
jenkins
491da8a7f5
monitoring(gpu): add pool utilization counters
2026-05-22 03:09:10 -03:00
jenkins
8002181f36
monitoring(gpu): normalize utilization pie to pool capacity
2026-05-22 02:55:24 -03:00
jenkins
ddb49dc62d
monitoring(gpu): hide zero-utilization namespaces
2026-05-22 02:35:51 -03:00
jenkins
95dc2c5007
monitoring(gpu): add process-level utilization attribution
2026-05-22 02:28:08 -03:00
jenkins
b3c76d8707
monitoring(gpu): remove ambiguous shared wording
2026-05-22 01:55:25 -03:00
jenkins
2a1e4b3c49
monitoring(gpu): attribute utilization to namespaces
2026-05-22 01:46:32 -03:00
jenkins
5ddb428a96
monitoring(gpu): show utilization with idle fallback
2026-05-21 15:26:02 -03:00
jenkins
1cc2372791
monitoring(gpu): clarify reservation accounting
2026-05-21 13:04:58 -03:00
jenkins
8da1bb6d53
ops: harden ci placement and gpu idle reporting
2026-05-20 04:27:26 -03:00
jenkins
e5d00c4eab
monitoring: hide idle gpu share during activity
2026-05-19 23:30:22 -03:00
jenkins
c01c4fe9d9
monitoring: fix gpu share and overview legends
2026-05-16 05:58:59 -03:00
a255c60aed
monitoring: fix gpu idle label
2026-01-27 21:46:58 -03:00
b4f5fbeb2b
monitoring: unify gpu namespace usage
2026-01-27 21:43:37 -03:00
577e2a158d
monitoring: keep idle label in gpu share
2026-01-27 18:44:58 -03:00
86cd5194ea
monitoring: fix gpu idle share
2026-01-27 17:51:13 -03:00
0fbbbf39e9
monitoring: fix jetson gpu metrics
2026-01-27 16:19:54 -03:00
995050f544
monitoring: unify jetson gpu metrics
2026-01-26 22:26:24 -03:00
84710b99e8
monitoring: add glue dashboard and tag cronjobs
2026-01-18 02:50:07 -03:00
fddf58346d
monitoring: treat cert-manager as infrastructure
2026-01-12 00:26:46 -03:00
98d405bc42
monitoring: regenerate dashboards with expanded infra namespaces
2026-01-11 23:55:43 -03:00
879ff7c16b
monitoring: fix infra scopes and add jetson metrics
2026-01-11 23:46:24 -03:00
05a888aeb6
monitoring(dashboards): tune namespace share metrics
2026-01-05 13:30:51 -03:00
ceea2539bc
monitoring: per-panel namespace share filters
2026-01-01 14:44:33 -03:00
bcc1ceef6d
monitoring: ensure gpu idle share renders
2026-01-01 14:21:43 -03:00
91de1c1d8d
gpu: enable time-slicing and refresh dashboards
2026-01-01 14:16:08 -03:00
2baa537ec7
Use table format for namespace plurality panel
2025-12-13 18:23:19 -03:00
2e18a4e1c5
atlas dashboards: fix pod share display and zero/red stat thresholds
2025-12-12 20:40:32 -03:00
da8ed7a3b0
atlas dashboards: show pod counts (not %) and make zero-friendly stats
2025-12-12 20:30:00 -03:00
2906e3e5d9
monitoring: show GPU share over dashboard range
2025-12-02 20:28:35 -03:00
12fd5229dc
monitoring: fix gpu share query and root bar labels
2025-12-02 14:56:36 -03:00
1963fadec1
monitoring: polish dashboards and folders
2025-12-02 14:41:39 -03:00
d23e2fe78c
monitoring: regen dashboards with gpu details
2025-12-02 13:16:00 -03:00