13 Commits

Author SHA1 Message Date
Hermes Agent
f997171b55 jenkins: rank the controller above its own build agents
The Jenkins controller and the ephemeral build agents it schedules both ran
at priority 0. Because the controller Deployment uses the Recreate strategy,
any JCasC change tears the controller down before its replacement is
scheduled, and while the eligible node pool is saturated that freed slot can
be taken by one of the controller's own pending agents.

That is what happened on 2026-08-23. Merging #48 changed
services/jenkins/configmap-jcasc.yaml, which rolled the hashed
jenkins-jcasc-revision ConfigMap and so rolled the Deployment. At 12:19:57Z
the old controller pod was deleted; at 12:19:59Z the build agent
typhon-133-1pmgb, pending since 11:42:15Z, took the freed capacity on
titan-11, and the new controller pod has been Pending ever since. The
scheduler cannot resolve it because every candidate victim is also priority
0: "No preemption victims found for incoming pod". The agent cannot exit on
its own either, since its jnlp container just loops on "Connection refused"
against the controller that cannot start. It is a self-deadlock, not a
transient capacity dip.

Give the controller a jenkins-core class (500) so it outranks its agents and
can preempt one when it has nowhere else to go, and mark the JCasC default
agent template jenkins-agent (50, preemptionPolicy: Never) so agents never
preempt anything themselves. Values follow the existing *-core/*-sim classes.

The controller-side class is what breaks the deadlock: agents declared inline
by a Jenkinsfile podTemplate, like typhon-133, do not inherit the JCasC
default template, so only the controller's own priority covers every agent.

PriorityClasses ship in infrastructure/modules/base via the core
Kustomization, so jenkins now dependsOn core, matching cassandra and veles.
Verified no dependency cycle: core itself has no dependsOn.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-23 13:12:40 +00:00
jenkins
5430c9ac01 jellyfin: prepare titan-22 media host 2026-08-22 22:06:39 -03:00
jenkins
8f45f47e60 refactor: restructure services layout, retire oceanus, add aether scaffolding
- Move flat service manifests into structured subdirs (apps/, bootstrap-jobs/,
  repair-jobs/, migration-jobs/, validation-jobs/, node-ops/, networking/)
- Retire oneoffs/ directories across services
- Remove oceanus cluster and its host roles; add aether cluster + terraform scaffolding
- Reorganize scripts/ into ops/, render/, sync/, manual-tests/
- Add Makefile with render/validate/test/flux targets and repo-structure tests
- Update flux-system application CRs to the new paths
- Add hermes-automated-triage-24h-plan knowledge doc (+ comms mirror)
- Refresh knowledge catalogs, dashboards, vmalert rules, quality contract

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 16:21:36 -03:00
jenkins
8c3e472e5f feat(cassandra): add parallel migration infrastructure 2026-07-25 00:18:48 -03:00
jenkins
8700593aa5 core: preserve Veles artifacts storageclass parameters 2026-06-27 07:48:46 -03:00
jenkins
6ef4ef161e Add Veles deployment IaC 2026-06-27 07:37:42 -03:00
jenkins
7d9d937b52 veles: harden app infrastructure contract 2026-06-09 11:59:27 -03:00
jenkins
1d20fb35d2 veles: stage atlas infrastructure 2026-06-09 00:46:46 -03:00
8d6d97e244 platform: restore cert-manager and encrypt budget storage 2026-01-17 07:38:38 -03:00
987dd126fa Fix Jetson device plugin args 2026-01-11 01:57:20 -03:00
91de1c1d8d gpu: enable time-slicing and refresh dashboards 2026-01-01 14:16:08 -03:00
dca749cc04 gpu: drop runtimeClass from minipc plugin 2025-11-09 13:28:40 -03:00
077654fa2d refactor: restructure atlas flux layout 2025-11-09 11:48:45 -03:00