jenkins
|
302fe78681
|
monitoring: preserve serving availability history
|
2026-08-04 18:39:36 -03:00 |
|
jenkins
|
368de58769
|
monitoring: fall back to live serving state
|
2026-08-04 18:16:13 -03:00 |
|
jenkins
|
bbab7330ce
|
monitoring: avoid unstable volume hosts
|
2026-08-04 18:06:43 -03:00 |
|
jenkins
|
5f9c313f0a
|
monitoring: render test category zero state
|
2026-08-04 12:34:08 -03:00 |
|
jenkins
|
17b25c4429
|
monitoring: reload vmalert rules automatically
|
2026-08-04 12:24:38 -03:00 |
|
jenkins
|
b0599f4318
|
monitoring: prevent query pool starvation
|
2026-08-04 12:22:23 -03:00 |
|
jenkins
|
e8f75f5cc9
|
feat(hermes): add actionable Atlas triage skills
|
2026-08-03 03:58:37 -03:00 |
|
jenkins
|
604f91cff3
|
monitoring(nodes): clamp CPU charts to physical range
|
2026-08-02 19:36:44 -03:00 |
|
jenkins
|
82b8a1c899
|
monitoring(gpu): attribute Jetson activity by allocation
|
2026-08-02 04:26:33 -03:00 |
|
jenkins
|
4991493d3b
|
ai(hermes): add operator guide and current GPU shares
|
2026-08-02 03:59:32 -03:00 |
|
jenkins
|
2fe327c7eb
|
monitoring(gpu): use process-level pod attribution
|
2026-08-02 03:31:02 -03:00 |
|
jenkins
|
e794ccf254
|
monitoring(gpu): report time-weighted namespace usage
|
2026-08-02 03:08:20 -03:00 |
|
jenkins
|
97ecb36b56
|
agent: replace OpenClaw with Hermes
|
2026-07-21 21:02:44 -03:00 |
|
jenkins
|
766c96f54d
|
monitoring: filter node dashboards to real nodes
|
2026-07-14 20:24:45 -03:00 |
|
jenkins
|
dad4448ebb
|
monitoring: collapse duplicate Typhon climate series
|
2026-07-14 19:01:53 -03:00 |
|
jenkins
|
c2a4b42b3d
|
monitoring: track titan-23 and pin tiny apps to workers
|
2026-07-14 18:14:38 -03:00 |
|
jenkins
|
a9ddc80e36
|
recovery(ananke): keep flux holds and place metrics on longhorn nodes
|
2026-06-18 22:08:14 -03:00 |
|
jenkins
|
6b777b8c74
|
recovery(ananke): reassert final flux hold after drain
|
2026-06-18 21:09:40 -03:00 |
|
jenkins
|
62d9791b76
|
recovery(ananke): leave flux root stopped on final hold
|
2026-06-18 20:54:16 -03:00 |
|
jenkins
|
ea333f6648
|
recovery(ananke): thaw critical flux without root apply
|
2026-06-18 20:37:00 -03:00 |
|
jenkins
|
6d3c59ad8d
|
recovery(ananke): fetch flux source before root thaw
|
2026-06-18 20:27:03 -03:00 |
|
jenkins
|
83b56488b4
|
recovery(ananke): quiet flux root apply window
|
2026-06-18 20:19:03 -03:00 |
|
jenkins
|
47b59a4f62
|
recovery(ananke): apply flux hold before suspension
|
2026-06-18 19:58:11 -03:00 |
|
jenkins
|
1648d392aa
|
recovery(ananke): apply root flux recovery hold
|
2026-06-18 19:40:19 -03:00 |
|
jenkins
|
32681728c0
|
recovery(ananke): finalize flux holds without races
|
2026-06-18 19:11:22 -03:00 |
|
jenkins
|
e893af2a55
|
recovery(ananke): verify flux suspension during thaw
|
2026-06-18 18:53:44 -03:00 |
|
jenkins
|
bb07f1598f
|
recovery(ananke): keep optional flux blocked during thaw
|
2026-06-18 18:42:37 -03:00 |
|
jenkins
|
761e4e4964
|
recovery(ananke): thaw flux critical path first
|
2026-06-18 18:35:13 -03:00 |
|
jenkins
|
0c2b59f7cc
|
recovery(ananke): avoid unnecessary longhorn sidecar churn
|
2026-06-18 18:20:22 -03:00 |
|
jenkins
|
8c45f9509e
|
recovery(ananke): use resident restart helper
|
2026-06-18 18:04:11 -03:00 |
|
jenkins
|
0f58aa16a9
|
recovery(ananke): handle longhorn harbor deadlock
|
2026-06-18 18:02:32 -03:00 |
|
jenkins
|
4fd8a00d4a
|
monitoring(testing): cap history panel ranges
|
2026-06-05 13:22:29 -03:00 |
|
jenkins
|
75d002dc88
|
monitoring(testing): cap expensive dashboard queries
|
2026-06-05 13:15:12 -03:00 |
|
jenkins
|
a2ecdef536
|
monitoring(testing): restore lesavka suite visibility
|
2026-06-05 01:04:56 -03:00 |
|
jenkins
|
fd7ec39a15
|
test(titan-iac): split dashboard trigger checks
|
2026-06-04 21:09:53 -03:00 |
|
jenkins
|
09e64c8ca4
|
ci(jenkins): refresh suite jobs twice daily
|
2026-06-04 20:38:02 -03:00 |
|
jenkins
|
f2ad8cca4c
|
monitoring(testing): clean up dashboard health signals
|
2026-06-04 16:09:08 -03:00 |
|
jenkins
|
5e27384ea2
|
monitoring(gpu): show activity share by namespace
|
2026-05-22 04:22:51 -03:00 |
|
jenkins
|
d21b61f6d9
|
monitoring(gpu): count monitored GPU pool devices
|
2026-05-22 03:23:36 -03:00 |
|
jenkins
|
6388ef5c6d
|
monitoring(gpu): add pool utilization counters
|
2026-05-22 03:09:10 -03:00 |
|
jenkins
|
570b1212d7
|
monitoring(gpu): normalize utilization pie to pool capacity
|
2026-05-22 02:55:24 -03:00 |
|
jenkins
|
b5dc723e02
|
monitoring(gpu): hide zero-utilization namespaces
|
2026-05-22 02:35:51 -03:00 |
|
jenkins
|
fd3da0e2ae
|
monitoring(gpu): add process-level utilization attribution
|
2026-05-22 02:28:08 -03:00 |
|
jenkins
|
5513608b1a
|
monitoring(gpu): remove ambiguous shared wording
|
2026-05-22 01:55:25 -03:00 |
|
jenkins
|
72e4dcd84b
|
monitoring(gpu): attribute utilization to namespaces
|
2026-05-22 01:46:32 -03:00 |
|
jenkins
|
4e82df6891
|
monitoring(gpu): show utilization with idle fallback
|
2026-05-21 15:26:02 -03:00 |
|
jenkins
|
d9955af899
|
monitoring(gpu): clarify reservation accounting
|
2026-05-21 13:04:58 -03:00 |
|
jenkins
|
f2ae3c1b0c
|
monitoring(testing): make branch filter static
|
2026-05-20 15:10:24 -03:00 |
|
jenkins
|
974955ac83
|
monitoring(testing): backfill category health rollups
|
2026-05-20 14:39:07 -03:00 |
|
jenkins
|
1c6c3992cf
|
monitoring(testing): reduce month-range query cost
|
2026-05-20 13:26:33 -03:00 |
|