13 Commits

Author SHA1 Message Date
jenkins
2bfdee6169 fix(hermes-agent): do NOT make titan-24 a general worker for the amd64 build
titan-24 is an accelerator node (co-hosts the out-of-cluster Sui validator), not
a general worker. The amd64 build leg was requiring node-role worker=true, which
forced labeling titan-24 as a worker and opened it to unrelated cluster
scheduling. It already pins by hostname+arch, so drop the worker requirement and
remove the titan-24 worker-join from the node-prefer CronJob entirely. The build
targets titan-24 specifically (hostname) and tolerates its taint; nothing else
in the cluster gets scheduled there.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvMSXH8VH2tMWXanb8SJdf
2026-08-25 11:50:52 -03:00
jenkins
8ddff9b626 infra(core): join titan-24 as an amd64 worker for the image build leg
The native amd64 hermes-agent image leg builds on titan-24. Worker membership
in this cluster is reconciled by the node-prefer-noschedule CronJob (kubectl
label), not Ansible, so add titan-24 there:

- clear_worker titan-24 amd64  -> node-role.kubernetes.io/worker=true + hardware=amd64
- a soft PreferNoSchedule guard taint (atlas.bstein.dev/sui-validator=true)
  mirroring titan-22's media guard, so routine pods do not crowd the
  out-of-cluster Sui validator that co-hosts titan-24. GPU workloads pinned to
  titan-24 by hostname are unaffected (PreferNoSchedule never blocks a pinned
  pod), and the amd64 build pod tolerates this taint explicitly.

Operator note: this reconciler does not manage cordons (owned by Ananke
recovery). titan-24 is on the recovery uncordon denylist, so the operator must
ensure titan-24 is uncordoned/schedulable before the first amd64 build.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvMSXH8VH2tMWXanb8SJdf
2026-08-25 11:05:42 -03:00
jenkins
5430c9ac01 jellyfin: prepare titan-22 media host 2026-08-22 22:06:39 -03:00
jenkins
46a44241c1 hermes: reclaim healthy worker capacity 2026-08-22 15:42:56 -03:00
jenkins
8a8df5ee4f hermes: add stateful accelerator fallback 2026-08-22 15:32:21 -03:00
jenkins
8c3e472e5f feat(cassandra): add parallel migration infrastructure 2026-07-25 00:18:48 -03:00
jenkins
75e2989ba7 core: repair node role reconciler 2026-06-19 15:45:45 -03:00
jenkins
1d20fb35d2 veles: stage atlas infrastructure 2026-06-09 00:46:46 -03:00
jenkins
11d4b2c013 maintenance: stabilize recovered worker nodes 2026-05-22 17:10:01 -03:00
jenkins
16a561e107 scheduling: target hdd storage node exclusions 2026-05-22 14:02:17 -03:00
jenkins
38a669ef8d core(nodes): mark rpi4 spillover workers 2026-05-20 18:14:49 -03:00
jenkins
712b97f64b agent(openclaw): expose oauth protected UI 2026-05-20 17:22:12 -03:00
jenkins
2c37ee4f84 recovery: keep storage nodes as spillover only 2026-05-15 11:52:26 -03:00