261 Commits

Author SHA1 Message Date
01ecb75c5b scripts: default cleanup verifier to maintenance kustomization 2026-04-12 15:01:11 -03:00
fa30ea0ac2 scripts: add jenkins cleanup rollout verifier 2026-04-12 12:32:20 -03:00
3774b600ee scheduling: keep app workloads off control-plane 2026-04-12 04:27:43 -03:00
3ea296b552 maintenance: enforce Astraios + tmpfs /tmp on worker Pis 2026-04-11 11:55:49 -03:00
370ece5b60 testing(ci): centralize quality gate contract 2026-04-10 17:11:02 -03:00
b723382ff4 dashboards: unify suite pass-rate metrics on platform counters 2026-04-10 16:39:55 -03:00
32b6e55467 monitoring: use CI-only series for platform test success panels 2026-04-10 04:52:57 -03:00
5f4641553c monitoring: replace failure table with 24h suite pass snapshot 2026-04-09 20:16:44 -03:00
530f440679 monitoring: add suite probe metrics and align fan labels 2026-04-09 20:10:52 -03:00
5e3aadc640 monitoring: set overview platform test panel to 7d 2026-04-09 20:05:10 -03:00
ad1cbd6f85 monitoring: make test panel point-based and failure-by-suite 2026-04-09 19:27:48 -03:00
5cf9a16d97 monitoring: align overview panels with jobs and point-based suite rates 2026-04-09 16:35:14 -03:00
f8c1243dfd monitoring: add generic suite metric slots for platform tests 2026-04-09 16:16:35 -03:00
7b0e9acbb1 monitoring: make suite pass rate 30d rolling for sparse tests 2026-04-09 16:14:26 -03:00
0273727cb4 monitoring: make platform test success one line per suite 2026-04-09 15:21:59 -03:00
09fa3e716c monitoring/atlas: merge top rows and fix platform test pass-rate panel 2026-04-09 14:56:43 -03:00
293cd83999 monitoring/atlas: resize test/ops rows and source overview tests from atlas-jobs 2026-04-09 13:39:55 -03:00
764bfe189e monitoring/recovery: harden ananke checks and OIDC-gated service validation 2026-04-09 01:44:26 -03:00
e0b124ca4e monitoring: switch power telemetry to ananke metrics 2026-04-08 23:33:17 -03:00
fa160f5f9b ananke: harden recovery checks and finalize naming migration 2026-04-07 13:09:18 -03:00
525a0f9e71 harbor/bootstrap: pin via dynamic host label managed by recovery script 2026-04-06 21:32:43 -03:00
d168f02c7f harbor/recovery: remove fixed titan-05 pin and auto-select ready arm64 node 2026-04-06 21:27:23 -03:00
a5f405432b hecate: add bootstrap bundle manifests and helper build scripts 2026-04-06 05:01:17 -03:00
d880fac673 hecate: harden titan-24 cleanup and ups telemetry 2026-04-06 04:55:54 -03:00
1e891de7e8 recovery: load colocated kubeconfig on remote hosts 2026-04-06 01:05:18 -03:00
b7f6317fd2 recovery: make cluster power console self-contained 2026-04-06 01:02:30 -03:00
99bd68f61b recovery: unblock harbor cold start and add power console 2026-04-06 00:22:54 -03:00
96bc93670b monitoring(power): rename hecate UPS peers to Pyrphoros and Statera 2026-04-04 05:54:16 -03:00
82e1b87b8f monitoring(overview): refine ups-climate row and climate/fan stat display 2026-04-04 04:40:22 -03:00
5059d2918d monitoring(overview): swap jobs and power rows; tighten climate/fan display 2026-04-04 04:34:18 -03:00
55b96c0675 monitoring(overview): place six power/climate panels on one row and fix test/job data 2026-04-04 01:33:15 -03:00
cdc3c081f5 monitoring(overview): replace power/climate summary row with six-panel layout 2026-04-03 22:16:02 -03:00
69a02a3352 monitoring(power): implement six-panel UPS and climate layout 2026-04-03 20:45:40 -03:00
4167f0f988 monitoring(power): add UPS status snapshot table and climate placeholders 2026-04-03 17:53:42 -03:00
fd71c6644b monitoring(power): wire generated power dashboard and split per-UPS panels 2026-04-03 17:49:09 -03:00
7ae4746d10 monitoring: scope hecate power queries to hecate-power job 2026-04-03 15:23:27 -03:00
bc9bf0310a monitoring: add power dashboard and reorder atlas overview rows 2026-04-03 14:55:16 -03:00
5a577630df platform: expose metis on sentinel and move gitea to rpi5 2026-03-31 16:44:41 -03:00
9a8030bf68 maintenance: harden metis recovery and fix harbor rollout 2026-03-31 14:55:48 -03:00
10ae47110a monitoring: combine Ariadne and Metis tests 2026-03-31 14:54:54 -03:00
03ae79df3e maintenance: harden sd-write controls and recovery workflow 2026-03-31 00:06:44 -03:00
a255c60aed monitoring: fix gpu idle label 2026-01-27 21:46:58 -03:00
b4f5fbeb2b monitoring: unify gpu namespace usage 2026-01-27 21:43:37 -03:00
577e2a158d monitoring: keep idle label in gpu share 2026-01-27 18:44:58 -03:00
86cd5194ea monitoring: fix gpu idle share 2026-01-27 17:51:13 -03:00
0fbbbf39e9 monitoring: fix jetson gpu metrics 2026-01-27 16:19:54 -03:00
995050f544 monitoring: unify jetson gpu metrics 2026-01-26 22:26:24 -03:00
f7d4425740 ariadne: reduce comms noise, fix gpu labels 2026-01-26 20:54:33 -03:00
4f9479c7d5 atlasbot: add metrics kb and long timeout 2026-01-26 14:08:11 -03:00
b5e8192731 atlasbot: answer jetson nodes from knowledge 2026-01-26 12:06:48 -03:00