titan-iac/services/hermes/configmap.yaml
2026-08-03 04:30:16 -03:00

308 lines
12 KiB
YAML

# services/hermes/configmap.yaml
apiVersion: v1
kind: ConfigMap
metadata:
name: hermes-config
namespace: hermes
labels:
app: hermes
data:
config.yaml: |
model:
provider: openai-codex
default: gpt-5.6-terra
model: gpt-5.6-terra
fallback_providers:
- provider: custom
model: gpt-oss:20b
base_url: http://hermes-model-gate.hermes.svc.cluster.local:11434/v1
api_key: ollama
agent:
api_max_retries: 1
platform_toolsets:
cli:
- clarify
- file
- session_search
- skills
- terminal
- todo
- web
api_server:
- clarify
- file
- session_search
- skills
- terminal
- todo
- web
skills:
creation_nudge_interval: 15
external_dirs:
- /opt/data/workspace/skills
terminal:
backend: local
cwd: /opt/data/workspace
timeout: 180
home_mode: auto
approvals:
mode: smart
deny:
- "*kubectl apply*"
- "*kubectl delete*"
- "*kubectl patch*"
- "*kubectl scale*"
- "*kubectl cordon*"
- "*kubectl uncordon*"
- "*kubectl drain*"
- "*kubectl rollout restart*"
- "*flux suspend*"
- "*flux resume*"
- "*flux reconcile*"
- "*vault kv get*"
- "*kubectl get secret*"
- "*kubectl describe secret*"
dashboard:
public_url: https://agent.bstein.dev
oauth:
provider: self-hosted
self_hosted:
issuer: https://sso.bstein.dev/realms/atlas
client_id: hermes-dashboard
scopes: openid profile email groups
display:
compact: true
tool_progress: all
interim_assistant_messages: true
long_running_notifications: true
tool_loop_guardrails:
warnings_enabled: true
hard_stop_enabled: true
warn_after:
exact_failure: 2
same_tool_failure: 3
idempotent_no_progress: 2
hard_stop_after:
exact_failure: 5
same_tool_failure: 8
idempotent_no_progress: 5
updates:
pre_update_backup: quick
backup_keep: 5
non_interactive_local_changes: stash
SOUL.md: |
You are Hermes running inside the Titan Kubernetes cluster as a supervised
testing and operations triage assistant.
Your strongest job is to follow the same evidence path Brad already uses:
Ariadne diagnosis first, then Jenkins logs and artifacts, Pushgateway
quality metrics, Flux state, Grafana dashboard context, and Kubernetes
read-only state. Turn repeated successful triage paths into reusable
skills or memory when the pattern is stable.
When Brad asks about CI, tests, suites, builds, or quality gates, use
triage-titan-test-failures. When he asks what is broken, why a service is
unhealthy, or why Grafana is red, use triage-atlas-service-health. When he
asks whether an alert is noisy or too aggressive, use tune-atlas-alerts.
Do not make him choose a tool, page, or evidence source first.
Stay Flux-first. Do not mutate the cluster directly. Explain evidence,
recommend the smallest repo-side change, and name the exact verification
commands a human should run after Flux reconciles.
AGENTS.md: |
# Titan Testing Triage
You are Hermes running inside the Titan Kubernetes cluster as a read-only
testing and operations triage assistant.
Route test/build/quality-gate requests to `triage-titan-test-failures`,
service/cluster-health requests to `triage-atlas-service-health`, and noisy
Grafana alert requests to `tune-atlas-alerts`. Ask only when approval is
required for a state-changing evidence refresh.
Ariadne owns deterministic evidence collection and local diagnosis. Start
every testing triage by reading:
- `curl -sS "$ARIADNE_BASE_URL/api/internal/testing/triage/diagnosis/latest"`
- if that is missing or stale, ask a human to run:
`curl -sS -X POST "$ARIADNE_BASE_URL/api/internal/testing/triage/diagnosis/run"`
- if the diagnosis is unavailable, fall back to:
`curl -sS "$ARIADNE_BASE_URL/api/internal/testing/triage/latest"`
Treat the Ariadne diagnosis and evidence bundle as the source of truth.
Your job is to explain the evidence and propose small Flux/IaC changes,
not to rediscover everything from raw shell commands.
Useful read-only commands:
- `curl -sS "$ARIADNE_BASE_URL/api/internal/testing/triage/diagnosis/latest"`
- `curl -sS "$ARIADNE_BASE_URL/api/internal/testing/triage/latest"`
- `kubectl get nodes -o wide`
- `kubectl get pods -A -o wide`
- `kubectl get pods -A --field-selector status.phase!=Running,status.phase!=Succeeded -o wide`
- `kubectl get events -A --sort-by=.lastTimestamp`
- `kubectl -n <namespace> describe pod <pod>`
- `kubectl -n <namespace> logs <pod> --all-containers --tail=200`
- `kubectl -n flux-system get kustomizations.kustomize.toolkit.fluxcd.io`
- `curl -sS "$VICTORIA_METRICS_URL/api/v1/query?query=up"`
Avoid grep/awk pipelines for absence checks; a grep exit code of 1 often
means "no matches", not a permission problem. Prefer Kubernetes field
selectors and quote the actual stderr when a command fails.
Do not run mutating commands such as `kubectl apply`, `delete`, `scale`,
`patch`, `cordon`, `uncordon`, `drain`, or `rollout restart`. Do not read
Kubernetes Secret values. Draft repo changes or operator steps instead.
START-HERE.md: |
# Hermes on Atlas: start here
Use Chat for a new investigation or Sessions to resume a previous one.
The shortest useful prompts are:
- `Triage the latest test failure.`
- `What is broken in the cluster right now?`
- `Why is Grafana red, and which alerts are noise?`
- `Triage <suite> build <number> and save the report.`
Hermes will use Ariadne evidence, Jenkins log excerpts, quality metrics,
recent Git changes, Flux state, Grafana context, and Kubernetes read-only
state. It will separate facts, inference, and unknowns and will ask before
triggering a fresh evidence collection.
Hermes now has three focused operator skills:
1. `triage-titan-test-failures` for all ten custom CI suites.
2. `triage-atlas-service-health` for any live namespace or service.
3. `tune-atlas-alerts` for evidence-backed alert and dashboard corrections.
Two validated CI examples are already available under Sessions:
1. `Proof 1 - Soteria 269 SonarQube triage`
2. `Proof 3 - titan-iac 1712 retained-artifact triage`
Their saved reports are in Files under `triage-proof/`.
Hermes may inspect workload and delivery metadata, logs, metrics, and
events. Its Kubernetes identity cannot read Secret values or mutate cluster
resources. Apply fixes through the titan-iac GitOps workflow after review.
HERMES-CAPABILITIES.md: |
# What Hermes can do on Atlas
## Working now
- Triage all custom CI suites: Ananke, Ariadne, Atlasbot, bstein_home,
data_prepper, Lesavka, Metis, Pegasus, Soteria, and titan_iac.
- Read Ariadne's deterministic evidence bundle, Jenkins failure logs and
retained quality artifacts, VictoriaMetrics, Flux state, Kubernetes
events/workloads/logs, and recent Gitea commits.
- Investigate any deployed service, group replica symptoms into incidents,
and separate active failures from startup grace, historical Jobs,
lifetime restart totals, and Veles migration residue.
- Audit Grafana alerts for bad counter/gauge math, low sample sizes, stale
schedules, duplicate series, missing persistence, and retired scope.
- Save concise triage reports and recommend exact repo-side Flux changes.
## Safety boundary
Hermes cannot read Kubernetes Secret values and cannot mutate Kubernetes or
Flux. It can quote secret-related errors already exposed in events or logs,
inspect the surrounding manifests, and tell Brad what approved change is
required. This keeps cluster operation supervised.
## Demonstrated corrections
- Zero-valued quality series no longer become failed suites.
- Running Jenkins builds remain in-progress instead of failed.
- Old failed Jobs, lifetime restart totals, and temporary Flux unknowns no
longer inflate the active incident count.
- Cassandra is authoritative; Veles failures remain visible as migration
residue but do not make the high-level Atlas health view red.
- Worker readiness uses the actual 18 Kubernetes workers.
- Root-disk growth, Soteria backup, Ariadne schedule, Postmark bounce, and
CPU alerts use actionable semantics and guardrails.
The evidence and verification for these changes are in
`triage-proof/ATLAS-TRIAGE-PROOFS.md`.
## Good next prompts
- `Use $triage-atlas-service-health. What is broken right now?`
- `Use $triage-titan-test-failures. Triage the worst current suite.`
- `Use $tune-atlas-alerts. Audit alerts fired in the last 24 hours.`
- `Save this investigation as a reusable skill after I approve the pattern.`
ATLAS-TRIAGE-PROOFS.md: |
# Atlas triage proofs
This file is the short evidence trail for what Hermes can do today. The
operator skills remain read-only: fixes are reviewed and delivered through
Git and Flux.
## Proof 1: Cassandra secret-sync recovery
- Finding: `cassandra-vault-sync` was stuck because the expected
`VELES_BYOK_ENCRYPTION_KEY` migration input was absent from Cassandra's
Vault path.
- Action: the existing Flux-tracked, suspended one-shot migration Job was
enabled, completed successfully, and immediately returned to suspended.
- Verification: Cassandra frontend, backend, PostgreSQL, and Vault sync all
became Ready. No Kubernetes Secret value was read or committed.
- Why this matters: Hermes can correlate workload state, logs, Vault policy
manifests, and migration ownership without confusing Veles residue with
a current Veles outage.
## Proof 2: Grafana alert-noise correction
- Removed retired Veles pods and completed Jobs from high-level incident
counts while preserving them on drill-down dashboards.
- Added 15-minute persistence to Pending, CrashLoopBackOff, and image-pull
summaries.
- Replaced invalid `increase()` use on gauges, bounded CPU percentages,
required meaningful Postmark sample sizes, and scoped Soteria/Ariadne
alerts to configured current work.
- Live verification on 2026-08-03: `Problem Pods=0`,
`CrashLoop/ImagePull=0`, and `Workers Ready=18/18`.
## Proof 3: all-suite testing triage
- Ariadne build 384 passed 676 tests and the exact coverage contract:
82 source files at or above 95%.
- Deployed image: `registry.bstein.dev/bstein/ariadne:0.1.0-384`.
- The deterministic bundle covers Ananke, Ariadne, Atlasbot, bstein_home,
data_prepper, Lesavka, Metis, Pegasus, Soteria, and titan_iac.
- Zero-valued quality metrics are healthy, running Jenkins builds are
in-progress, old failed Jobs are historical, and Veles objects are
migration residue. They no longer inflate the active failure set.
- Failed builds still include direct Jenkins console and retained artifact
links so Hermes can identify the first failed gate and smallest repo fix.
## Proof 4: Cassandra generator failure classification
- Finding: the Cassandra serving path is healthy, but generation Job
`cassandra-generator-3ce198971b` failed.
- Evidence: the primary OpenAI request returned HTTP 429 because its credit
balance was exhausted. Codex CLI fallback ran, but the strict promotion
gate still rejected unresolved high-impact hint claims and blocked the
semantic scenarios.
- Impact: one generation capability/request failed; this is not a Cassandra
registry outage and not a Veles outage.
- Next action: restore provider credits or continue through Codex fallback,
then resolve the reported hint claims before retrying. This needs operator
approval because it can spend money or rerun an expensive generation.
## Use Hermes next
- `What is broken in the cluster right now?`
- `Triage the worst current CI suite and link the evidence.`
- `Audit Grafana alerts from the last 24 hours and separate incidents from noise.`
- `Explain the Cassandra generator failure and give me the smallest safe next step.`