jenkins
2c11b87fdb
fix(monitoring): Alertmanager must HELO with a FQDN for Mailu
...
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 21:33:11 -03:00
jenkins
84e9849e46
feat(hermes-triage): real-repo patch proposals + Alertmanager email escalation
...
- Ariadne: per-repo code config for metis, lesavka, soteria,
bstein-dev-home and ariadne, each with its own base branch, source path
prefixes and file suffixes so a proposal can only touch that repo's
source tree
- Alertmanager: the only receiver was an empty "default", so every alert
fired into a void. HermesTriageHumanRequired now routes to an email
receiver via Mailu's in-cluster local-domain relay, with resolved
notices; scoped to service=hermes-triage so nothing else mails yet.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 21:30:34 -03:00
jenkins
bbab7330ce
monitoring: avoid unstable volume hosts
2026-08-04 18:06:43 -03:00
jenkins
b0599f4318
monitoring: prevent query pool starvation
2026-08-04 12:22:23 -03:00
jenkins
95bc7b6f17
monitoring(grafana): use Texas timezone
2026-08-03 15:58:16 -03:00
jenkins
8db88d79f4
fix(monitoring): reload provisioned Grafana rules
2026-08-03 04:05:33 -03:00
jenkins
b26a4f8c66
feat(cassandra): add migration Grafana rollup
2026-07-25 00:29:03 -03:00
jenkins
8c0e848e30
monitoring: scope endpoint node relabel to node-exporter
2026-07-14 18:57:51 -03:00
jenkins
575cae700d
monitoring: relabel pod-network node-exporter names
2026-07-14 18:30:43 -03:00
jenkins
c02f36e24c
monitoring: move node-exporter host port
2026-07-14 18:23:52 -03:00
jenkins
d1c1172c1d
monitoring: avoid node-exporter host port collisions
2026-07-14 18:18:03 -03:00
jenkins
c2a4b42b3d
monitoring: track titan-23 and pin tiny apps to workers
2026-07-14 18:14:38 -03:00
jenkins
52920d0f8b
monitoring: give victoria metrics recovery headroom
2026-06-18 22:23:48 -03:00
jenkins
a9ddc80e36
recovery(ananke): keep flux holds and place metrics on longhorn nodes
2026-06-18 22:08:14 -03:00
jenkins
654900b8a2
veles: stage atlas infrastructure
2026-06-09 00:46:46 -03:00
jenkins
5ec561f620
monitoring(grafana): lower recovery scheduling requests
2026-05-20 18:16:04 -03:00
jenkins
f010d0547f
monitoring(grafana): keep off control plane spillover
2026-05-20 18:07:11 -03:00
jenkins
ed81d52dd9
monitoring(grafana): avoid fragile placement and init pull
2026-05-20 17:25:39 -03:00
jenkins
400077436b
monitoring(grafana): harden scheduling and readiness
2026-05-20 17:00:33 -03:00
jenkins
1cfc846ffc
monitoring: retire duplicate jobs dashboard
2026-05-16 03:04:27 -03:00
jenkins
8fb5831e00
monitoring: publish atlas testing dashboard
2026-05-16 02:56:52 -03:00
jenkins
a1f6758b95
monitoring(grafana): refresh provisioned dashboards
2026-04-22 15:13:26 -03:00
ff11f7ee65
monitoring(vm): raise kube-state-metrics scrape size cap
2026-04-16 19:47:56 -03:00
11d9c5eae3
monitoring(vm): avoid accelerator nodes for vmsingle
2026-04-16 19:39:35 -03:00
95dd0bbd56
monitoring(vm): auto-reload scrape config changes
2026-04-16 19:33:39 -03:00
72e7a39373
monitoring: fix grafana no-data scrape gaps
2026-04-16 19:30:31 -03:00
a1ab78b0c9
monitoring(grafana): mount and provision atlas-testing dashboard
2026-04-12 22:13:58 -03:00
e0b124ca4e
monitoring: switch power telemetry to ananke metrics
2026-04-08 23:33:17 -03:00
96bc93670b
monitoring(power): rename hecate UPS peers to Pyrphoros and Statera
2026-04-04 05:54:16 -03:00
82e1b87b8f
monitoring(overview): refine ups-climate row and climate/fan stat display
2026-04-04 04:40:22 -03:00
1b682cc60f
monitoring(grafana): restart to pick up latest overview layout
2026-04-04 04:35:26 -03:00
d5fc6c89c4
monitoring(grafana): bump restart revision for overview dashboard reload
2026-04-04 01:34:36 -03:00
7ef4c895ba
monitoring(grafana): bump restart revision to reload provisioned dashboards
2026-04-03 20:54:12 -03:00
fd71c6644b
monitoring(power): wire generated power dashboard and split per-UPS panels
2026-04-03 17:49:09 -03:00
bc9bf0310a
monitoring: add power dashboard and reorder atlas overview rows
2026-04-03 14:55:16 -03:00
3cf426e23a
monitoring: roll grafana to apply latest alert rules
2026-03-30 18:41:26 -03:00
0aeb08d375
monitoring: fix noisy grafana email alerts and reload rules
2026-03-30 18:33:02 -03:00
a49fa6dd33
monitoring: restart grafana for alerting reload
2026-01-27 23:29:46 -03:00
32884e0b7e
monitoring: fix grafana smtp from address
2026-01-27 22:28:37 -03:00
7b43e8654f
monitoring: send grafana alerts via postmark
2026-01-27 22:00:19 -03:00
993702afee
monitoring: alert on VM outage
2026-01-23 11:51:28 -03:00
fc87432fdf
monitoring: refresh jobs dashboards
2026-01-21 13:37:36 -03:00
5fe70b1471
grafana: allow email-based oauth user lookup
2026-01-21 11:45:11 -03:00
14d75ccf7a
monitoring: label cronjob metrics and move grafana to arm64
2026-01-18 12:20:45 -03:00
60dee25f08
monitoring: add atlas testing dashboard folder
2026-01-18 12:07:45 -03:00
8b86c5dd67
monitoring: avoid titan-22 for core pods
2026-01-18 11:43:28 -03:00
4bc57cf445
monitoring: restore grafana persistence
2026-01-18 11:37:01 -03:00
8fb73e023c
monitoring: disable grafana persistence to recover
2026-01-18 09:55:28 -03:00
b0698887a4
monitoring: add testing dashboard and switch postmark apikey
2026-01-18 09:21:33 -03:00
af86a610d9
fix ingress tls routing
2026-01-16 01:40:50 -03:00