jenkins
1e6eda8633
fix: converge load-spread rollouts
2026-08-23 10:41:01 -03:00
jenkins
271f3e8c32
ops: spread saturated node workloads
2026-08-23 10:37:01 -03:00
Hermes Agent
c4eac8ceee
fix(monitoring): keep VictoriaMetrics writable
2026-08-19 10:51:44 +00:00
jenkins
612cefad23
monitoring(ai): add provider quota operations dashboard
2026-08-16 05:13:20 -03:00
jenkins
da128b0487
monitoring: retain Pi fallback for VictoriaMetrics
2026-08-11 01:40:28 -03:00
jenkins
e9a3b1fd0e
monitoring: keep VictoriaMetrics on Pi 5 workers
2026-08-11 01:36:50 -03:00
jenkins
84ff647898
fix(monitoring): Alertmanager must HELO with a FQDN for Mailu
...
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 21:33:11 -03:00
jenkins
664e436575
feat(hermes-triage): real-repo patch proposals + Alertmanager email escalation
...
- Ariadne: per-repo code config for metis, lesavka, soteria,
bstein-dev-home and ariadne, each with its own base branch, source path
prefixes and file suffixes so a proposal can only touch that repo's
source tree
- Alertmanager: the only receiver was an empty "default", so every alert
fired into a void. HermesTriageHumanRequired now routes to an email
receiver via Mailu's in-cluster local-domain relay, with resolved
notices; scoped to service=hermes-triage so nothing else mails yet.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 21:30:34 -03:00
jenkins
5fabd1a84b
monitoring: avoid unstable volume hosts
2026-08-04 18:06:43 -03:00
jenkins
e354f3f83a
monitoring: prevent query pool starvation
2026-08-04 12:22:23 -03:00
jenkins
1c070d96d4
monitoring(grafana): use Texas timezone
2026-08-03 15:58:16 -03:00
jenkins
8e7327a7ef
fix(monitoring): reload provisioned Grafana rules
2026-08-03 04:05:33 -03:00
jenkins
a155847915
feat(cassandra): add migration Grafana rollup
2026-07-25 00:29:03 -03:00
jenkins
618fd01cc5
monitoring: scope endpoint node relabel to node-exporter
2026-07-14 18:57:51 -03:00
jenkins
921edc257c
monitoring: relabel pod-network node-exporter names
2026-07-14 18:30:43 -03:00
jenkins
e03598caa2
monitoring: move node-exporter host port
2026-07-14 18:23:52 -03:00
jenkins
4f4a5f99b7
monitoring: avoid node-exporter host port collisions
2026-07-14 18:18:03 -03:00
jenkins
bea20b9c97
monitoring: track titan-23 and pin tiny apps to workers
2026-07-14 18:14:38 -03:00
jenkins
8d3402beb7
monitoring: give victoria metrics recovery headroom
2026-06-18 22:23:48 -03:00
jenkins
e7ec333fb0
recovery(ananke): keep flux holds and place metrics on longhorn nodes
2026-06-18 22:08:14 -03:00
jenkins
1d20fb35d2
veles: stage atlas infrastructure
2026-06-09 00:46:46 -03:00
jenkins
57e021425e
monitoring(grafana): lower recovery scheduling requests
2026-05-20 18:16:04 -03:00
jenkins
17094f1cd6
monitoring(grafana): keep off control plane spillover
2026-05-20 18:07:11 -03:00
jenkins
24b573f4f9
monitoring(grafana): avoid fragile placement and init pull
2026-05-20 17:25:39 -03:00
jenkins
13ad257875
monitoring(grafana): harden scheduling and readiness
2026-05-20 17:00:33 -03:00
jenkins
46890111b1
monitoring: retire duplicate jobs dashboard
2026-05-16 03:04:27 -03:00
jenkins
4f0748ac8e
monitoring: publish atlas testing dashboard
2026-05-16 02:56:52 -03:00
jenkins
5f00b8049b
monitoring(grafana): refresh provisioned dashboards
2026-04-22 15:13:26 -03:00
639f64b366
monitoring(vm): raise kube-state-metrics scrape size cap
2026-04-16 19:47:56 -03:00
cf6f03a5ee
monitoring(vm): avoid accelerator nodes for vmsingle
2026-04-16 19:39:35 -03:00
66fb9be159
monitoring(vm): auto-reload scrape config changes
2026-04-16 19:33:39 -03:00
2983edfccd
monitoring: fix grafana no-data scrape gaps
2026-04-16 19:30:31 -03:00
aef64c2473
monitoring(grafana): mount and provision atlas-testing dashboard
2026-04-12 22:13:58 -03:00
e0b124ca4e
monitoring: switch power telemetry to ananke metrics
2026-04-08 23:33:17 -03:00
96bc93670b
monitoring(power): rename hecate UPS peers to Pyrphoros and Statera
2026-04-04 05:54:16 -03:00
82e1b87b8f
monitoring(overview): refine ups-climate row and climate/fan stat display
2026-04-04 04:40:22 -03:00
1b682cc60f
monitoring(grafana): restart to pick up latest overview layout
2026-04-04 04:35:26 -03:00
d5fc6c89c4
monitoring(grafana): bump restart revision for overview dashboard reload
2026-04-04 01:34:36 -03:00
7ef4c895ba
monitoring(grafana): bump restart revision to reload provisioned dashboards
2026-04-03 20:54:12 -03:00
fd71c6644b
monitoring(power): wire generated power dashboard and split per-UPS panels
2026-04-03 17:49:09 -03:00
bc9bf0310a
monitoring: add power dashboard and reorder atlas overview rows
2026-04-03 14:55:16 -03:00
3cf426e23a
monitoring: roll grafana to apply latest alert rules
2026-03-30 18:41:26 -03:00
0aeb08d375
monitoring: fix noisy grafana email alerts and reload rules
2026-03-30 18:33:02 -03:00
a49fa6dd33
monitoring: restart grafana for alerting reload
2026-01-27 23:29:46 -03:00
32884e0b7e
monitoring: fix grafana smtp from address
2026-01-27 22:28:37 -03:00
7b43e8654f
monitoring: send grafana alerts via postmark
2026-01-27 22:00:19 -03:00
993702afee
monitoring: alert on VM outage
2026-01-23 11:51:28 -03:00
fc87432fdf
monitoring: refresh jobs dashboards
2026-01-21 13:37:36 -03:00
5fe70b1471
grafana: allow email-based oauth user lookup
2026-01-21 11:45:11 -03:00
14d75ccf7a
monitoring: label cronjob metrics and move grafana to arm64
2026-01-18 12:20:45 -03:00