jenkins
cd974baeeb
fix(demo): print the triage allowlist as a list, not raw config
...
Tests / Declarative: Post Actions testing.tests.test_repo_structure.test_knowledge_service_mirror_matches_source failed
The allowlist is the outermost safety boundary - a job absent from it is never
touched, whatever fails - so it is worth reading at a glance rather than as a
comma-separated setting value echoed verbatim.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 15:48:57 -03:00
jenkins
5364cc9666
fix(demo): source local credentials and stop waiting on a job that no longer exists
...
The script demanded JENKINS_USER and JENKINS_TOKEN in the environment and said
only 'set JENKINS_USER' when they were missing, which is not enough to act on.
It now sources scripts/ops/hermes_triage_demo.env, git-ignored so it can hold
real tokens, and names that file when credentials are absent. An example file
records what belongs in it.
The fixture command also still polled for a hermes-demo-repair-<build> Job.
That Job stopped existing when the repair became an in-process ConfigMap
patch, so the command would have waited its full 400 seconds and then reported
nothing. It now watches the fixture returning to healthy, which is what
actually happens.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 15:45:15 -03:00
jenkins
b233630007
feat(demo): preflight the agent-pool cap and open code-demo PRs
...
Two conditions silently break a rehearsal. A saturated Kubernetes agent pool
leaves the demo build queued reporting that all nodes are offline, and an open
hermes-repair PR makes the duplicate guard refuse a new proposal. Report both
so the operator sees them before starting rather than mid-demo.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 23:19:37 -03:00
jenkins
fc0d41056d
fix(scripts): dashboard renderer wrote outside the repo after the layout move
...
ROOT still used parents[1], which resolved to scripts/ once the renderer
moved into scripts/render/. Every --build run wrote a phantom
scripts/services/monitoring tree and silently left the real dashboards
untouched. Points at the repo root again and removes the stray tree.
Also adds Hermes triage panels to the Atlas Testing dashboard: open
escalations awaiting a human, automated actions succeeded, Hermes
diagnosis latency, actions by result, and incident state by job.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 21:41:02 -03:00
jenkins
a555fc0c96
feat(hermes-code): multibranch validation for hermes-repair/* proposal branches
...
Adds demo driver script and branch-level test gate so a Hermes-proposed
pull request carries a green build before human merge.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 19:46:30 -03:00
jenkins
f4097c6e59
Merge origin/main into layout migration
...
Reconciles the services layout migration with ~991 upstream commits:
- Remote content wins for cassandra/cassandra-auth, monitoring dashboards,
vmalert availability rules, veles, vault auth script, dashboard render
script and tests (request-v4 availability definition)
- Layout paths win for structure: keycloak/bstein-dev-home job dirs use
bootstrap-jobs/validation-jobs; cassandra realm jobs live in
cassandra-auth (removed keycloak duplicates)
- Union: applications CR list gains hermes-chat and cassandra
image-automation
- Fixed post-migration paths in dashboard test module loader and
hermes-access job header
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 16:26:36 -03:00
jenkins
0da9e4c82d
refactor: restructure services layout, retire oceanus, add aether scaffolding
...
- Move flat service manifests into structured subdirs (apps/, bootstrap-jobs/,
repair-jobs/, migration-jobs/, validation-jobs/, node-ops/, networking/)
- Retire oneoffs/ directories across services
- Remove oceanus cluster and its host roles; add aether cluster + terraform scaffolding
- Reorganize scripts/ into ops/, render/, sync/, manual-tests/
- Add Makefile with render/validate/test/flux targets and repo-structure tests
- Update flux-system application CRs to the new paths
- Add hermes-automated-triage-24h-plan knowledge doc (+ comms mirror)
- Refresh knowledge catalogs, dashboards, vmalert rules, quality contract
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 16:21:36 -03:00
jenkins
9b8a29022d
Revert "monitoring(network): organize Traefik traffic lanes"
...
This reverts commit eeec8b72dc885950dd116d6dc2ad6574c1cbe6b1.
2026-08-05 12:06:23 -03:00
jenkins
eeec8b72dc
monitoring(network): organize Traefik traffic lanes
2026-08-05 12:02:51 -03:00
jenkins
ea53bec75d
monitoring: remove legacy availability series
2026-08-04 22:49:48 -03:00
jenkins
663aa3e4f1
monitoring: retain daily availability through retries
2026-08-04 21:47:30 -03:00
jenkins
2691581cbc
monitoring: publish availability outside query pool
2026-08-04 21:40:24 -03:00
jenkins
a8dda9815d
monitoring: add daily availability rollups
2026-08-04 21:24:47 -03:00
jenkins
7c47a92610
monitoring: add compact availability fallback
2026-08-04 21:14:06 -03:00
jenkins
f724e72a96
monitoring: use request success for availability
2026-08-04 21:12:31 -03:00
jenkins
4906fb73d8
monitoring: measure gateway availability
2026-08-04 21:02:19 -03:00
jenkins
302fe78681
monitoring: preserve serving availability history
2026-08-04 18:39:36 -03:00
jenkins
368de58769
monitoring: fall back to live serving state
2026-08-04 18:16:13 -03:00
jenkins
bbab7330ce
monitoring: avoid unstable volume hosts
2026-08-04 18:06:43 -03:00
jenkins
5f9c313f0a
monitoring: render test category zero state
2026-08-04 12:34:08 -03:00
jenkins
17b25c4429
monitoring: reload vmalert rules automatically
2026-08-04 12:24:38 -03:00
jenkins
b0599f4318
monitoring: prevent query pool starvation
2026-08-04 12:22:23 -03:00
jenkins
e8f75f5cc9
feat(hermes): add actionable Atlas triage skills
2026-08-03 03:58:37 -03:00
jenkins
604f91cff3
monitoring(nodes): clamp CPU charts to physical range
2026-08-02 19:36:44 -03:00
jenkins
82b8a1c899
monitoring(gpu): attribute Jetson activity by allocation
2026-08-02 04:26:33 -03:00
jenkins
4991493d3b
ai(hermes): add operator guide and current GPU shares
2026-08-02 03:59:32 -03:00
jenkins
2fe327c7eb
monitoring(gpu): use process-level pod attribution
2026-08-02 03:31:02 -03:00
jenkins
e794ccf254
monitoring(gpu): report time-weighted namespace usage
2026-08-02 03:08:20 -03:00
jenkins
97ecb36b56
agent: replace OpenClaw with Hermes
2026-07-21 21:02:44 -03:00
jenkins
766c96f54d
monitoring: filter node dashboards to real nodes
2026-07-14 20:24:45 -03:00
jenkins
dad4448ebb
monitoring: collapse duplicate Typhon climate series
2026-07-14 19:01:53 -03:00
jenkins
c2a4b42b3d
monitoring: track titan-23 and pin tiny apps to workers
2026-07-14 18:14:38 -03:00
jenkins
a9ddc80e36
recovery(ananke): keep flux holds and place metrics on longhorn nodes
2026-06-18 22:08:14 -03:00
jenkins
6b777b8c74
recovery(ananke): reassert final flux hold after drain
2026-06-18 21:09:40 -03:00
jenkins
62d9791b76
recovery(ananke): leave flux root stopped on final hold
2026-06-18 20:54:16 -03:00
jenkins
ea333f6648
recovery(ananke): thaw critical flux without root apply
2026-06-18 20:37:00 -03:00
jenkins
6d3c59ad8d
recovery(ananke): fetch flux source before root thaw
2026-06-18 20:27:03 -03:00
jenkins
83b56488b4
recovery(ananke): quiet flux root apply window
2026-06-18 20:19:03 -03:00
jenkins
47b59a4f62
recovery(ananke): apply flux hold before suspension
2026-06-18 19:58:11 -03:00
jenkins
1648d392aa
recovery(ananke): apply root flux recovery hold
2026-06-18 19:40:19 -03:00
jenkins
32681728c0
recovery(ananke): finalize flux holds without races
2026-06-18 19:11:22 -03:00
jenkins
e893af2a55
recovery(ananke): verify flux suspension during thaw
2026-06-18 18:53:44 -03:00
jenkins
bb07f1598f
recovery(ananke): keep optional flux blocked during thaw
2026-06-18 18:42:37 -03:00
jenkins
761e4e4964
recovery(ananke): thaw flux critical path first
2026-06-18 18:35:13 -03:00
jenkins
0c2b59f7cc
recovery(ananke): avoid unnecessary longhorn sidecar churn
2026-06-18 18:20:22 -03:00
jenkins
8c45f9509e
recovery(ananke): use resident restart helper
2026-06-18 18:04:11 -03:00
jenkins
0f58aa16a9
recovery(ananke): handle longhorn harbor deadlock
2026-06-18 18:02:32 -03:00
jenkins
4fd8a00d4a
monitoring(testing): cap history panel ranges
2026-06-05 13:22:29 -03:00
jenkins
75d002dc88
monitoring(testing): cap expensive dashboard queries
2026-06-05 13:15:12 -03:00
jenkins
a2ecdef536
monitoring(testing): restore lesavka suite visibility
2026-06-05 01:04:56 -03:00