101 Commits

Author SHA1 Message Date
Hermes Agent
c4eac8ceee fix(monitoring): keep VictoriaMetrics writable 2026-08-19 10:51:44 +00:00
jenkins
612cefad23 monitoring(ai): add provider quota operations dashboard 2026-08-16 05:13:20 -03:00
jenkins
da128b0487 monitoring: retain Pi fallback for VictoriaMetrics 2026-08-11 01:40:28 -03:00
jenkins
e9a3b1fd0e monitoring: keep VictoriaMetrics on Pi 5 workers 2026-08-11 01:36:50 -03:00
jenkins
84ff647898 fix(monitoring): Alertmanager must HELO with a FQDN for Mailu
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 21:33:11 -03:00
jenkins
664e436575 feat(hermes-triage): real-repo patch proposals + Alertmanager email escalation
- Ariadne: per-repo code config for metis, lesavka, soteria,
  bstein-dev-home and ariadne, each with its own base branch, source path
  prefixes and file suffixes so a proposal can only touch that repo's
  source tree
- Alertmanager: the only receiver was an empty "default", so every alert
  fired into a void. HermesTriageHumanRequired now routes to an email
  receiver via Mailu's in-cluster local-domain relay, with resolved
  notices; scoped to service=hermes-triage so nothing else mails yet.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 21:30:34 -03:00
jenkins
5fabd1a84b monitoring: avoid unstable volume hosts 2026-08-04 18:06:43 -03:00
jenkins
e354f3f83a monitoring: prevent query pool starvation 2026-08-04 12:22:23 -03:00
jenkins
1c070d96d4 monitoring(grafana): use Texas timezone 2026-08-03 15:58:16 -03:00
jenkins
8e7327a7ef fix(monitoring): reload provisioned Grafana rules 2026-08-03 04:05:33 -03:00
jenkins
a155847915 feat(cassandra): add migration Grafana rollup 2026-07-25 00:29:03 -03:00
jenkins
618fd01cc5 monitoring: scope endpoint node relabel to node-exporter 2026-07-14 18:57:51 -03:00
jenkins
921edc257c monitoring: relabel pod-network node-exporter names 2026-07-14 18:30:43 -03:00
jenkins
e03598caa2 monitoring: move node-exporter host port 2026-07-14 18:23:52 -03:00
jenkins
4f4a5f99b7 monitoring: avoid node-exporter host port collisions 2026-07-14 18:18:03 -03:00
jenkins
bea20b9c97 monitoring: track titan-23 and pin tiny apps to workers 2026-07-14 18:14:38 -03:00
jenkins
8d3402beb7 monitoring: give victoria metrics recovery headroom 2026-06-18 22:23:48 -03:00
jenkins
e7ec333fb0 recovery(ananke): keep flux holds and place metrics on longhorn nodes 2026-06-18 22:08:14 -03:00
jenkins
1d20fb35d2 veles: stage atlas infrastructure 2026-06-09 00:46:46 -03:00
jenkins
57e021425e monitoring(grafana): lower recovery scheduling requests 2026-05-20 18:16:04 -03:00
jenkins
17094f1cd6 monitoring(grafana): keep off control plane spillover 2026-05-20 18:07:11 -03:00
jenkins
24b573f4f9 monitoring(grafana): avoid fragile placement and init pull 2026-05-20 17:25:39 -03:00
jenkins
13ad257875 monitoring(grafana): harden scheduling and readiness 2026-05-20 17:00:33 -03:00
jenkins
46890111b1 monitoring: retire duplicate jobs dashboard 2026-05-16 03:04:27 -03:00
jenkins
4f0748ac8e monitoring: publish atlas testing dashboard 2026-05-16 02:56:52 -03:00
jenkins
5f00b8049b monitoring(grafana): refresh provisioned dashboards 2026-04-22 15:13:26 -03:00
639f64b366 monitoring(vm): raise kube-state-metrics scrape size cap 2026-04-16 19:47:56 -03:00
cf6f03a5ee monitoring(vm): avoid accelerator nodes for vmsingle 2026-04-16 19:39:35 -03:00
66fb9be159 monitoring(vm): auto-reload scrape config changes 2026-04-16 19:33:39 -03:00
2983edfccd monitoring: fix grafana no-data scrape gaps 2026-04-16 19:30:31 -03:00
aef64c2473 monitoring(grafana): mount and provision atlas-testing dashboard 2026-04-12 22:13:58 -03:00
e0b124ca4e monitoring: switch power telemetry to ananke metrics 2026-04-08 23:33:17 -03:00
96bc93670b monitoring(power): rename hecate UPS peers to Pyrphoros and Statera 2026-04-04 05:54:16 -03:00
82e1b87b8f monitoring(overview): refine ups-climate row and climate/fan stat display 2026-04-04 04:40:22 -03:00
1b682cc60f monitoring(grafana): restart to pick up latest overview layout 2026-04-04 04:35:26 -03:00
d5fc6c89c4 monitoring(grafana): bump restart revision for overview dashboard reload 2026-04-04 01:34:36 -03:00
7ef4c895ba monitoring(grafana): bump restart revision to reload provisioned dashboards 2026-04-03 20:54:12 -03:00
fd71c6644b monitoring(power): wire generated power dashboard and split per-UPS panels 2026-04-03 17:49:09 -03:00
bc9bf0310a monitoring: add power dashboard and reorder atlas overview rows 2026-04-03 14:55:16 -03:00
3cf426e23a monitoring: roll grafana to apply latest alert rules 2026-03-30 18:41:26 -03:00
0aeb08d375 monitoring: fix noisy grafana email alerts and reload rules 2026-03-30 18:33:02 -03:00
a49fa6dd33 monitoring: restart grafana for alerting reload 2026-01-27 23:29:46 -03:00
32884e0b7e monitoring: fix grafana smtp from address 2026-01-27 22:28:37 -03:00
7b43e8654f monitoring: send grafana alerts via postmark 2026-01-27 22:00:19 -03:00
993702afee monitoring: alert on VM outage 2026-01-23 11:51:28 -03:00
fc87432fdf monitoring: refresh jobs dashboards 2026-01-21 13:37:36 -03:00
5fe70b1471 grafana: allow email-based oauth user lookup 2026-01-21 11:45:11 -03:00
14d75ccf7a monitoring: label cronjob metrics and move grafana to arm64 2026-01-18 12:20:45 -03:00
60dee25f08 monitoring: add atlas testing dashboard folder 2026-01-18 12:07:45 -03:00
8b86c5dd67 monitoring: avoid titan-22 for core pods 2026-01-18 11:43:28 -03:00