5375 Commits

Author SHA1 Message Date
jenkins
0ad2ffefce monitoring: prefer worker nodes for Alertmanager 2026-08-25 20:00:05 -03:00
jenkins
1f636489f7 monitoring: keep Alertmanager on available rpi5 workers 2026-08-25 19:52:55 -03:00
jenkins
8d3a1193b6 monitoring: move Alertmanager recovery to available worker 2026-08-25 19:38:48 -03:00
jenkins
35f650ef41 monitoring: recover Alertmanager placement drift 2026-08-25 19:34:42 -03:00
jenkins
db446c9244 monitoring(ai): alert before Claude quota auth expires 2026-08-25 19:16:01 -03:00
flux-bot
2a11c8d207 chore(bstein-dev-home): automated image update 2026-08-25 22:07:28 +00:00
flux-bot
c30a7c2d2b chore(bstein-dev-home): automated image update 2026-08-25 22:06:28 +00:00
jenkins
342677dde7 monitoring(ai): preserve quota data across rollouts 2026-08-25 18:33:18 -03:00
jenkins
6f783b7778 build(hermes-webui): multi-arch image (arm64 + amd64)
Repoints both Dockerfile.hermes-webui FROM bases to in-cluster Harbor mirrors
(webui base OCI index + the a68d1c4d multi-arch hermes-agent manifest list),
adds the suspended webui base-mirror Job, and gives the WebUI image build an
amd64 leg on titan-24 plus a manifest-list combine — so the hux sidecar can
schedule onto the amd64 accelerator node titan-22.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvMSXH8VH2tMWXanb8SJdf
2026-08-25 18:31:00 -03:00
jenkins
e7d6759140 fix(hermes-agent): reinstall codex when this arch's native dep is missing
The arch-specific CLI stamp assumed the shared node_modules keeps both arches'
codex native deps, but a sibling-arch npm install removes this arch's binary from
the shared volume. So after running on the other arch, the stamp exists yet the
native codex dep is gone -> configure-agent-clients fails -> churn. Gate the
install on the current arch's native codex package being present, so it self-heals.
2026-08-25 18:23:13 -03:00
jenkins
9c516b9808 build(hermes-webui): multi-arch image (arm64 + amd64)
Make registry.bstein.dev/bstein/hermes-webui a linux/amd64 + linux/arm64
manifest list so the agent pod's `hux` sidecar (which runs the webui image)
can schedule onto the amd64 node titan-22. Reuses the hermes-agent multi-arch
pattern already on main.

- Dockerfile.hermes-webui: repoint both FROMs to multi-arch, internal sources.
  The upstream WebUI base (ghcr sha256:a83a3893..., already a multi-arch OCI
  index) is now pulled from the in-cluster Harbor mirror; the agent base moves
  from the retired arm64-only leaf (81970563) to the multi-arch agent index
  (a68d1c4d). Kaniko selects the matching arch leaf per build node.
- services/harbor/hermes-webui-base-mirror-job.yaml: new suspended, operator-run
  skopeo `copy --all` Job mirroring the upstream WebUI base index into Harbor's
  `mirror` project (modeled on hermes-agent-base-mirror-job.yaml; reuses the
  generic ensure-project helper). Wired into the harbor kustomization.
- Jenkinsfile.hermes-webui-image: arm64 leg (titan-20) + amd64 leg (titan-24,
  hostname+arch pin, toleration Exists, resource-capped, own checkout scm) +
  Combine multi-arch index stage; per-arch evidence archived alongside the index.
- hermes_multiarch_combine.py: generalize the destination pattern/component to
  serve both hermes-agent and hermes-webui (fail-closed to just those two).
- Tests updated to the two-arch topology (two legs, combine, both FROM bases,
  the mirror Job, twelve archived evidence files).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvMSXH8VH2tMWXanb8SJdf
2026-08-25 18:17:35 -03:00
jenkins
70002aeff7 Revert "Reapply "hermes(agent): make titan-22 the strong primary home (no worker label)""
This reverts commit 9e4fbf4e0df364230588749f80710a8369a3108f.
2026-08-25 17:49:44 -03:00
jenkins
a01eac5059 Revert "Reapply "hermes(agent): also tolerate titan-22's media-primary taint (flip applied pre-classification)""
This reverts commit 2e38422478b0d569441176431f104faae0dd1703.
2026-08-25 17:49:44 -03:00
jenkins
2e38422478 Reapply "hermes(agent): also tolerate titan-22's media-primary taint (flip applied pre-classification)"
Some checks failed
Tests / Declarative: Post Actions failed: 49, skipped: 81, passed: 3753
This reverts commit 1c495a6e2c0a525b9e16608b16bc09c645e3c337.
2026-08-25 17:40:36 -03:00
jenkins
9e4fbf4e0d Reapply "hermes(agent): make titan-22 the strong primary home (no worker label)"
This reverts commit f8628e6ee0d2c33d287ec9086ff5328d982e88c3.
2026-08-25 17:40:36 -03:00
jenkins
12a6d2c4f5 hermes(agent): make runtime tooling install architecture-aware
The hermes-agent installs its CLI toolchain at runtime into the shared
/opt/data/tools Longhorn volume, but every download hardcoded arm64. On
the amd64 node titan-22 that left configure-agent-clients failing with
"Missing optional dependency @openai/codex-linux-x64" and the operator
toolchain fetching arm64 binaries, so the pod churned.

Detect the running node's arch (uname -m; fail closed on anything but
aarch64/x86_64) and resolve every asset per-arch:

- install-agent-tools init script (agent-deployment.yaml): ttyd and
  kubectl download the arch-correct asset with the arch-correct sha256
  (real ttyd 1.7.7 x86_64 and kubectl v1.33.3 amd64 checksums added; the
  arm64 ones kept). The npm CLI stamp is now arch-specific
  (.cli-versions-<vers>-${arch}) so a fresh arch re-runs npm install and
  pulls its own native optional deps; npm keeps both arches' packages.

- install_agent_tools.sh: flux/helm/kustomize/jq/yq/gh/vault/sops/age/
  k9s/terraform/go URLs, tarball subdirs (helm linux-${arch}, gh dir),
  and checksums are all arch-resolved with both arches pinned. Stamps
  and the Go tree are arch-specific, and an active-arch marker forces a
  republish of the single-arch ${bin} binaries when the pod moves
  between arches on the shared volume. Single fetch/verify helper kept.

Tests updated to assert the arch-aware form (both arches' Go checksums,
${dl_arch} templating) instead of the arm64-only literal.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvMSXH8VH2tMWXanb8SJdf
2026-08-25 17:37:57 -03:00
jenkins
f8628e6ee0 Revert "hermes(agent): make titan-22 the strong primary home (no worker label)"
This reverts commit 2de52ec3e2e42696ae411482cb16f27ff1f5d273.
2026-08-25 16:54:04 -03:00
jenkins
1c495a6e2c Revert "hermes(agent): also tolerate titan-22's media-primary taint (flip applied pre-classification)"
This reverts commit aa5898f9e33b371451c94af803cda1e0ed714591.
2026-08-25 16:54:04 -03:00
jenkins
aa5898f9e3 hermes(agent): also tolerate titan-22's media-primary taint (flip applied pre-classification)
Applying the titan-22 flip ahead of the full accelerator classification, so
titan-22 still carries its soft media-primary taint. Tolerate it too so the
strong titan-22 preference isn't penalised. Harmless once media-primary is gone.
2026-08-25 16:43:32 -03:00
jenkins
5631366c4e hermes(agent): make titan-22 the strong primary home (no worker label)
APPLY ONLY AFTER the multi-arch hermes-agent image is built + validated
(both arch leaves + promoted index). Supersedes the earlier titan-22
flip on feature/hermes-agent-multiarch (dfa50b75), which required
worker=true and tolerated the old media-primary taint.

Rewrites the runtime node affinity so hermes-agent runs on titan-22:
- Adds a second, OR'd nodeSelectorTerm matching amd64 + hostname
  titan-22 + node-role.kubernetes.io/accelerator=true. It does NOT
  require node-role.kubernetes.io/worker (titan-22 is no longer a
  generic worker).
- Keeps the arm64 pi-fleet term untouched as an OR'd fallback so the
  worker is never stranded if titan-22 is unavailable.
- Makes titan-22 the STRONG/primary preference: a hostname=titan-22
  preference at weight 100 (the scheduler maximum) outranks the pi-fleet
  rpi5 nudge, lowered to weight 50, so hermes actually lives on titan-22.
- Tolerates node-role.kubernetes.io/accelerator=true:NoSchedule so it
  can consider titan-22; this does not change jellyfin's media-core
  priority or preemption.

Updates test_hermes_agent_layout.py to the two-term topology, the
[100, 50] preference weights, the titan-22 primary preference, the
absence of a worker requirement on the titan-22 term, and the toleration.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvMSXH8VH2tMWXanb8SJdf
2026-08-25 16:43:32 -03:00
flux-bot
aae22ea798 chore(hermes): promote validated image release 2026-08-25 19:28:55 +00:00
jenkins
a7fc9b20a0 fix(harbor): mirror to external Harbor endpoint (valid TLS)
harbor-core internally advertises an HTTPS token realm, so skopeo could not push
over HTTP. Push to registry.bstein.dev (valid cert, the path kaniko already
uses); the image lands in the same 'mirror' project and stays internally pullable.
2026-08-25 14:23:51 -03:00
jenkins
e7ac6351a6 fix(harbor): mark harbor-core insecure so skopeo pushes over HTTP
skopeo derived the token realm as HTTPS and got 'HTTP response to HTTPS client'.
A registries.conf with insecure=true for harbor-core:80 makes the registry AND
its token request use HTTP.
2026-08-25 14:21:15 -03:00
jenkins
4f5fc44013 fix(harbor): push the mirror over Harbor's HTTP port 80
harbor-core serves http on :80 only; the skopeo dest omitted the port so it
dialed :443 and timed out. Pin the dest to :80 (with --dest-tls-verify=false).
2026-08-25 14:18:35 -03:00
jenkins
6a25a7681a fix(harbor): run Vault init first in the base-image mirror Job
The Job's ensure-project init container reads /vault/secrets/harbor-admin-password,
but Vault appended its init container AFTER ensure-project, so the secret file
was absent and the init failed. Force vault-agent-init to run first.
2026-08-25 14:10:57 -03:00
jenkins
8a71084585 build(hermes-agent): source base image + test deps from in-cluster mirrors
The hermes-agent-image pipeline failed intermittently on external network:
Kaniko's docker.io fallback for the base image is IPv6-broken from build
pods, and the "Validate reviewed release source" stage pip-installed pytest
from files.pythonhosted.org (DNS failures). Neither should touch the public
internet.

Base image: repoint the Dockerfile FROM from docker.io to the in-cluster
Harbor "mirror" project, keeping the exact content-addressed index digest
(9c841866...) and both arch leaves. A Flux-managed one-shot Job
(services/harbor/hermes-agent-base-mirror-job.yaml, suspend: true like the
cassandra bootstrap job) runs `skopeo copy --all` from docker.io into Harbor
using the same Vault-injected admin credential as the existing Harbor
immutability jobs; a tiny fail-closed helper ensures the public target
project first. Digest pinning and multi-arch are preserved; Kaniko pulls it
over the internal insecure registry with no docker.io fallback.

Test deps: install pytest/PyYAML fully offline (`pip --no-index
--find-links`) from a reviewed in-repo wheelhouse
(ci/vendor/hermes-agent-test-wheels) matching the arm64 python:3.12 build
container, so the validate stage never resolves a public index.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvMSXH8VH2tMWXanb8SJdf
2026-08-25 13:53:22 -03:00
jenkins
500741020a fix(hermes): revert hard rpi5 requirement (it stranded worker)
Requiring rpi5 while the historical hostname exclusion list removes the very rpi5
nodes that currently have headroom (titan-04/06) left only loaded rpi5s
(titan-05/07/11), so the agent could not schedule and worker went down. Revert to
rpi5-PREFERRED (soft) so it schedules again; proper rpi5 placement needs the
exclusion list refreshed against current node health/capacity, tracked separately.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvMSXH8VH2tMWXanb8SJdf
2026-08-25 10:25:54 -03:00
jenkins
c44f16e4ac fix(hermes): require rpi5 for worker (keep the heavy agent off rpi4)
Un-pinning let the scheduler land worker on titan-12 (rpi4). The agent is heavy
enough that an rpi4 risks the /api/status slowness that trips its liveness probe
- the exact flap we are avoiding. Make hardware=rpi5 a hard requirement so it
runs only on rpi5 storage workers (excluding the saturated/known-flaky ones);
the scheduler places it on a roomy rpi5 (titan-05).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvMSXH8VH2tMWXanb8SJdf
2026-08-25 10:10:24 -03:00
jenkins
e54d581ef5 fix(hermes): un-pin worker from titan-08; spread across arm64 storage workers
Worker (hermes-agent) was hard-pinned to titan-08 (a workaround after an earlier
attempt to place it on the amd64 titan-22 failed on architecture). That single-
node pin is exactly what makes it fragile: a titan-08 blip (as just happened when
the node's Longhorn CSI went down) strands worker, and the Recreate strategy then
deadlocks because the replacement can't schedule on the one tight node.

Restore the intended multi-node design: run on any arm64 storage worker except the
known-bad/weak ones (matching the repo's own affinity test, which was red). Its
Longhorn volumes have data-locality disabled with replicas on titan-15/17/19, so
there is no locality penalty to running on another node; the scheduler now places
it on a roomier Pi (e.g. titan-05) and a node blip simply reschedules it.

Also relax the gateway /api/status liveness probe (timeout 10s->15s,
failureThreshold 3->5) so a transient slowness (e.g. a brief storage hiccup) no
longer trips a kill-and-restart cascade.

Follow-up (separate): multi-arch agent image to enable the amd64 titan-22 target.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvMSXH8VH2tMWXanb8SJdf
2026-08-25 09:30:06 -03:00
flux-bot
8d781ea808 chore(bstein-dev-home): automated image update 2026-08-25 09:48:59 +00:00
flux-bot
ced73bf30d chore(bstein-dev-home): automated image update 2026-08-25 09:48:08 +00:00
flux-bot
5888850eba chore(hermes): promote validated image release
Some checks failed
Tests / Declarative: Post Actions failed: 49, skipped: 81, passed: 3727
2026-08-25 04:43:23 +00:00
flux-bot
9f49a1ba2c chore(hermes): promote validated image release 2026-08-25 03:37:13 +00:00
flux-bot
a59590be79 chore(hermes): promote validated image release 2026-08-25 02:56:09 +00:00
flux-bot
b59f754f5a chore(hermes): promote validated image release 2026-08-25 01:53:50 +00:00
flux-bot
29ba03105e chore(maintenance): automated image update 2026-08-25 01:50:53 +00:00
flux-bot
7826529273 chore(maintenance): automated image update 2026-08-25 01:50:07 +00:00
flux-bot
774d74c458 chore(maintenance): automated image update 2026-08-25 01:46:53 +00:00
flux-bot
770b812ae5 chore(maintenance): automated image update 2026-08-25 01:41:05 +00:00
flux-bot
63bd63c283 chore(hermes): promote validated image release 2026-08-25 01:35:01 +00:00
flux-bot
700301d554 chore(hermes): promote validated image release 2026-08-25 00:35:53 +00:00
flux-bot
77380bfc25 chore(hermes): promote validated image release 2026-08-25 00:11:52 +00:00
flux-bot
f96070f6d9 chore(hermes): promote validated image release 2026-08-24 23:48:49 +00:00
jenkins
b99952f16c hermes(stt): greedy single-temperature final decode for fast startup
Beam-2 with a temperature-fallback ladder made the final-model warmup
run all three temperature retries under beam search before the server
bound its port, so /health was refused for ~8 min and STT was down that
whole time on every roll (and hinted at slow per-utterance decodes).
Production now runs the accurate large-v3-turbo model greedily at a
single temperature, keeping the proper-noun priming prompt that fixes
names like Amy/Córdoba - fast startup, fast decodes, accuracy intact.
Beam stays env-tunable (HERMES_STT_FINAL_BEAM_SIZE) for a future pass;
serve-before-warmup is a recommended follow-up so cold start never
blocks readiness.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvMSXH8VH2tMWXanb8SJdf
2026-08-24 20:20:28 -03:00
jenkins
546e960185 hermes(stt-client): retry a brief STT outage, never dump a traceback
A transcription that landed while the private Whisper service was
restarting (an image roll) crashed hermes_stt_client.py with a raw
urllib ConnectionRefused traceback that got dumped into the
conversation. The client now retries the request with backoff (up to 5
attempts, ~10s - long enough to ride an STT pod restart) and, on a
persistent outage, exits with one concise line ('speech transcription
unavailable...') instead of a stack trace. Delivered via the coordinator
ConfigMap; the next reconcile picks it up.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvMSXH8VH2tMWXanb8SJdf
2026-08-24 20:16:57 -03:00
flux-bot
c54eeada1b chore(hermes): promote validated image release 2026-08-24 23:08:35 +00:00
flux-bot
88f8e9e629 chore(hermes): promote validated image release 2026-08-24 22:47:03 +00:00
flux-bot
4caf8942fd chore(bstein-dev-home): automated image update 2026-08-24 21:49:01 +00:00
flux-bot
a825d70d15 chore(bstein-dev-home): automated image update 2026-08-24 21:46:57 +00:00
jenkins
1fdfaff096 hermes(stt): accurate large-v3-turbo final decode, fast tiny partials
Proper nouns (Córdoba, Cancún) and dropped words came from decoding the
committed transcript with the small model. The image already ships
large-v3-turbo, so the final decode now uses it with beam_size=5, a
temperature fallback ladder, and a proper-noun/accents initial_prompt
that fixes first-pass capitalization and diacritics across EN/ES/RU;
the rolling previews stay on tiny at greedy so the on-the-fly feel is
unchanged. The accurate decode runs in the speculative predecode during
the end-of-speech silence and is cache-reused at commit, so perceived
latency stays low. All decode knobs are env-overridable for on-device
tuning (beam/temperature/prompt), with small as the guaranteed-present
rollback if turbo underperforms on the Jetson. 206 STT tests pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvMSXH8VH2tMWXanb8SJdf
2026-08-24 16:39:00 -03:00