hermes: add operator mastery guide
This commit is contained in:
parent
92c2bbdb12
commit
7b99d12124
239
services/hermes/NOTES.md
Normal file
239
services/hermes/NOTES.md
Normal file
@ -0,0 +1,239 @@
|
||||
# Hermes on Atlas: operator guide
|
||||
|
||||
This is the mental model and demonstration script for the operator instance at
|
||||
`agent.bstein.dev`. Read it once, then prove each section in the live UI. The
|
||||
consumer instance at `chat.bstein.dev` is intentionally separate and is not the
|
||||
place to perform infrastructure triage.
|
||||
|
||||
## The one-sentence explanation
|
||||
|
||||
Hermes is the persistent agent runtime and control surface; Codex or the local
|
||||
GPU model supplies reasoning, while Hermes supplies sessions, tools, skills,
|
||||
approval rules, identity, and the repeatable procedure that connects the model
|
||||
to Atlas evidence.
|
||||
|
||||
Hermes is not the model. Replacing `gpt-5.6-terra` with `gpt-oss:20b` changes the
|
||||
reasoning engine, but it does not replace the Hermes sessions, tools, skills,
|
||||
files, permission boundary, or workflow.
|
||||
|
||||
## Request and evidence path
|
||||
|
||||
```text
|
||||
browser
|
||||
-> Keycloak login
|
||||
-> oauth2-proxy exact-user check
|
||||
-> operator Hermes pod and persistent workspace
|
||||
-> SOUL.md + AGENTS.md + selected skill
|
||||
-> Codex primary model
|
||||
-> read-only terminal/web/file tools
|
||||
-> Ariadne deterministic evidence
|
||||
-> retained Jenkins logs and artifacts
|
||||
-> Gitea commits and Flux state
|
||||
-> Kubernetes workloads, events, logs, and dependencies
|
||||
-> Pushgateway data queried through VictoriaMetrics
|
||||
-> Grafana query and alert context
|
||||
-> structured finding and proposed repo-side change
|
||||
-> human review and approval
|
||||
-> Git/Flux delivery outside Hermes
|
||||
|
||||
provider failure
|
||||
-> hermes-model-gate
|
||||
-> Hermes owns titan-24: Ollama serves gpt-oss:20b
|
||||
-> Wolf owns titan-24: gate returns 503; local inference yields
|
||||
```
|
||||
|
||||
The operator agent pod runs on an ARM worker and does not reserve the GPU.
|
||||
Ollama is the component on `titan-24` that owns GPU memory. The Lease
|
||||
`hermes/titan-24-gpu-owner` controls whether the model gate admits local
|
||||
inference. Normal operator conversations use independently authenticated Codex
|
||||
first, so Wolf ownership does not need to block the web agent.
|
||||
|
||||
## Who does what
|
||||
|
||||
| Component | Responsibility | What it does not prove |
|
||||
| --- | --- | --- |
|
||||
| Hermes | Maintains the conversation, chooses tools/skills, reasons over evidence, saves reports, and enforces the approval experience | It is not the inference model and does not automatically have cluster-admin access |
|
||||
| Ariadne | Collects and normalizes deterministic cluster, Jenkins, and quality evidence into a timestamped bundle; may optionally run a local diagnosis | Its diagnosis is not authoritative when stale, empty, or contradicted by the bundle |
|
||||
| Jenkins | Runs the suites and retains console/artifact evidence | A final `exit 1` alone does not identify the first failed gate |
|
||||
| Pushgateway | Receives suite/build quality telemetry | A cumulative counter is not the current build result |
|
||||
| VictoriaMetrics | Stores and answers PromQL for quality and environment telemetry | A Grafana color is not direct proof of root cause |
|
||||
| Grafana | Visualizes and alerts on VictoriaMetrics data | A red panel must be traced to its query, labels, time range, and raw series |
|
||||
| Flux | Applies reviewed Git state to the cluster | A recent reconciliation is correlation until the changed path matches the failure |
|
||||
| Kubernetes | Supplies workload, event, log, node, storage, and dependency evidence | An unrelated unhealthy pod is not proof that CI failed because of the cluster |
|
||||
|
||||
Direct Jenkins artifact requests currently require authorization and can return
|
||||
HTTP 403. That does not break the workflow: Ariadne retains selected Jenkins
|
||||
console tails and named artifact contents in its deterministic bundle. A report
|
||||
must say `retained Ariadne evidence` when that fallback is used; it must not
|
||||
pretend direct Jenkins access succeeded.
|
||||
|
||||
## The actual supervised triage algorithm
|
||||
|
||||
1. Classify the request as test/build triage, service health, or alert tuning.
|
||||
2. Read the latest Ariadne diagnosis and deterministic bundle with HTTP GET.
|
||||
3. Compare their timestamps. The deterministic bundle remains authoritative.
|
||||
4. Select a terminal failed build. Keep running or unknown builds out of the
|
||||
terminal-failure list.
|
||||
5. Identify the first enforced failed gate in this order:
|
||||
`style -> loc -> coverage -> tests -> gate_glue -> sonarqube -> supply_chain`.
|
||||
6. Cite the smallest decisive Jenkins log/artifact evidence. State whether it
|
||||
came directly from Jenkins or from Ariadne retention.
|
||||
7. Query the matching build/check/test series in VictoriaMetrics. Distinguish a
|
||||
current gauge from a cumulative counter and ignore zero-valued failure
|
||||
series.
|
||||
8. Correlate recent Gitea commits and the relevant Flux revision by timestamp
|
||||
and affected path. Recency alone is not causation.
|
||||
9. Check only the Kubernetes resources capable of explaining that failure.
|
||||
Classify observations as direct, contributing candidate, or background.
|
||||
10. Return `Finding`, `Confidence`, `Evidence`, `Likely cause`, `Blast radius`,
|
||||
`Next checks`, `Repo-side fix`, and `Approval required`.
|
||||
11. Stop at a proposal. A human reviews the patch or action; Git and Flux remain
|
||||
the delivery path.
|
||||
|
||||
## What Hermes can read and what it cannot do
|
||||
|
||||
The live `hermes-triage` ServiceAccount can get/list/watch ordinary workload,
|
||||
log, event, ingress, storage, Flux, and image-automation metadata. It cannot
|
||||
read Kubernetes Secret values, create exec sessions, patch Deployments, or
|
||||
patch Flux Kustomizations. The configured command deny list also blocks common
|
||||
Kubernetes and Flux mutations.
|
||||
|
||||
The security boundary is layered:
|
||||
|
||||
1. Keycloak authenticates the person.
|
||||
2. oauth2-proxy restricts the operator surface to Brad.
|
||||
3. The operator has a separate namespace, PVC, configuration, and ServiceAccount.
|
||||
4. Kubernetes RBAC is the hard API authorization boundary.
|
||||
5. NetworkPolicy limits reachable paths where configured.
|
||||
6. Hermes approvals and instructions provide a user-facing safety layer.
|
||||
|
||||
A skill is procedure and context, not a permission grant. Adding a sentence to
|
||||
a skill cannot bypass Kubernetes RBAC.
|
||||
|
||||
## What “Hermes learns the workflow” means
|
||||
|
||||
It does not silently retrain model weights on cluster data. Learning here means
|
||||
that a successful repeated procedure is written as a reusable skill on the
|
||||
operator PVC or versioned in Git. The skill describes when it should trigger,
|
||||
the evidence order, interpretation rules, output contract, and safety boundary.
|
||||
|
||||
The current workflow has a versioned top-level skill and persisted specialist
|
||||
skills for:
|
||||
|
||||
- orchestration and evidence reporting;
|
||||
- retained Jenkins evidence;
|
||||
- quality metrics;
|
||||
- Gitea/Flux correlation;
|
||||
- Kubernetes failure classification;
|
||||
- Grafana metric provenance;
|
||||
- Soteria backup health;
|
||||
- approval-required actions.
|
||||
|
||||
The model still evaluates fresh variable evidence on every incident. The skill
|
||||
makes the process repeatable; it does not freeze the answer.
|
||||
|
||||
## The four operator skills to remember
|
||||
|
||||
- `triage-titan-test-failures`: CI, tests, builds, quality gates, and suspected
|
||||
test-environment regressions.
|
||||
- `triage-atlas-service-health`: active service or cluster incidents and red
|
||||
high-level health panels.
|
||||
- `tune-atlas-alerts`: noisy alerts, impossible metrics, bad PromQL, and
|
||||
generator-owned Grafana corrections.
|
||||
- `master-hermes-on-atlas`: hands-on training, assessments, architecture, and
|
||||
claim audits.
|
||||
|
||||
Do not start by choosing raw tools. State the operational question and let the
|
||||
skill route to the narrow evidence source.
|
||||
|
||||
## Two live, repeatable proof cases
|
||||
|
||||
### Proof 6: Soteria build 272
|
||||
|
||||
- Terminal result: failure.
|
||||
- Tests: 318 passed, zero failed.
|
||||
- Local coverage: 96.195 percent.
|
||||
- First enforced failure: SonarQube `new_coverage=0.0` against threshold 80.
|
||||
- Runtime: Soteria remained Ready.
|
||||
- Correct conclusion: a Sonar new-code policy/reporting problem, not a test or
|
||||
Kubernetes capacity failure.
|
||||
- Correct restraint: do not weaken the policy until scanner import and baseline
|
||||
evidence identifies whether the issue is configuration or legitimate new
|
||||
uncovered code.
|
||||
|
||||
### Proof 7: Ananke build 242
|
||||
|
||||
- Terminal result: failure.
|
||||
- Local coverage gate: 61.8 percent and failed.
|
||||
- Other checks: tests, SonarQube, supply-chain, LOC, docs naming, and gate glue
|
||||
were healthy.
|
||||
- Sonar new-code coverage: 93.1 percent, which is a different scope.
|
||||
- Correct conclusion: local coverage input/scope caused the enforced failure;
|
||||
do not call the Sonar number contradictory without comparing inputs.
|
||||
- Correct restraint: do not lower the threshold until the retained local gate
|
||||
inputs are readable.
|
||||
|
||||
The full sessions and redacted Markdown exports are in Sessions and
|
||||
`Files/triage-proof`.
|
||||
|
||||
## Five-minute demonstration
|
||||
|
||||
1. Open Models and show `openai-codex/gpt-5.6-terra` as primary and
|
||||
`gpt-oss:20b` as local fallback.
|
||||
2. Open Skills and show the four operator skills plus the persisted component
|
||||
skills.
|
||||
3. Open `Proof 6 - Soteria 272 Sonar new coverage triage` in Sessions.
|
||||
4. Point out the timestamped bundle, exact build/artifact evidence, metric
|
||||
corroboration, healthy runtime, fact/inference separation, and approval
|
||||
boundary.
|
||||
5. Open Proof 7 and explain why local coverage and Sonar new-code coverage can
|
||||
differ without either number being fabricated.
|
||||
6. End by showing that Hermes proposes a repo-side correction but cannot patch
|
||||
the Deployment or Flux Kustomization with its ServiceAccount.
|
||||
|
||||
Use this short explanation:
|
||||
|
||||
> Hermes is the persistent, permissioned workflow layer around the model. In
|
||||
> this cluster Ariadne collects deterministic evidence, Hermes correlates it
|
||||
> with retained Jenkins artifacts, Git/Flux state, Kubernetes health, and
|
||||
> quality metrics, and then returns a supervised proposal. Repeated procedures
|
||||
> become explicit skills. Kubernetes RBAC prevents the agent from turning a
|
||||
> diagnosis into an unreviewed infrastructure change.
|
||||
|
||||
## Prompts that exercise the real system
|
||||
|
||||
- `Use $triage-titan-test-failures. Triage the worst current terminal CI failure and link every decisive piece of evidence.`
|
||||
- `Use $triage-titan-test-failures. Triage Ananke build 242. Separate local coverage from Sonar new-code coverage.`
|
||||
- `Use $triage-atlas-service-health. Explain the VictoriaMetrics outage from current state and retained events; separate recovered impact from current impact.`
|
||||
- `Use $tune-atlas-alerts. Trace one currently firing alert to its generated source and raw PromQL, but do not edit it.`
|
||||
- `Use $master-hermes-on-atlas. Assess me on the request path and permission boundary. One question at a time.`
|
||||
|
||||
## Honest limits
|
||||
|
||||
- Hermes does not currently apply production or cluster changes autonomously.
|
||||
- Direct Jenkins console/artifact access may be forbidden; Ariadne retention is
|
||||
the current fallback and must be named as provenance.
|
||||
- Local `gpt-oss:20b` is useful for bounded work but is slower and less reliable
|
||||
for multi-source triage than Codex.
|
||||
- The current Ariadne local diagnosis can return an empty model response. That
|
||||
does not invalidate the deterministic bundle or prevent Codex-backed Hermes
|
||||
from triaging it.
|
||||
- A single successful case demonstrates capability, not mastery. Mastery means
|
||||
leading different real incidents, improving the skill after failures, and
|
||||
teaching the architecture without prompts.
|
||||
|
||||
## Your shortest path to fluency
|
||||
|
||||
1. Explain the request diagram without looking.
|
||||
2. Explain why Ariadne and Hermes are separate components.
|
||||
3. Reproduce Proof 6 from the raw bundle and metrics.
|
||||
4. Reproduce Proof 7 and explain the two coverage scopes.
|
||||
5. Produce the live RBAC allow/deny matrix with `kubectl auth can-i`.
|
||||
6. Create one small writable skill from a repeated sub-workflow, test a trigger
|
||||
and non-trigger case, then improve it.
|
||||
7. Lead one new incident while Hermes coaches rather than answers for you.
|
||||
8. Teach the five-minute demonstration to another person.
|
||||
|
||||
At that point the interview claims are demonstrable. Continue repeating real
|
||||
incidents before describing yourself as fully autonomous or the system as
|
||||
self-correcting.
|
||||
@ -185,12 +185,14 @@ data:
|
||||
2. `triage-atlas-service-health` for any live namespace or service.
|
||||
3. `tune-atlas-alerts` for evidence-backed alert and dashboard corrections.
|
||||
|
||||
Two validated CI examples are already available under Sessions:
|
||||
Validated CI examples are available under Sessions, including:
|
||||
|
||||
1. `Proof 1 - Soteria 269 SonarQube triage`
|
||||
2. `Proof 3 - titan-iac 1712 retained-artifact triage`
|
||||
1. `Proof 6 - Soteria 272 Sonar new coverage triage`
|
||||
2. `Proof 7 - Ananke 242 local coverage triage`
|
||||
|
||||
Their saved reports are in Files under `triage-proof/`.
|
||||
Their saved reports are in Files under `triage-proof/`. Read
|
||||
`HERMES-OPERATOR-RUNBOOK.md` for the system mental model and five-minute
|
||||
demonstration.
|
||||
|
||||
Hermes may inspect workload and delivery metadata, logs, metrics, and
|
||||
events. Its Kubernetes identity cannot read Secret values or mutate cluster
|
||||
@ -309,6 +311,30 @@ data:
|
||||
- It kept running titan-iac, Data Prepper, and Lesavka builds out of the
|
||||
terminal-failure list and requested approval before any rerun or edit.
|
||||
|
||||
## Proof 6: Soteria build 272 Sonar new-coverage triage
|
||||
|
||||
- Hermes read the fresh 2026-08-04 Ariadne bundle and selected terminal
|
||||
Soteria build 272.
|
||||
- It proved that all 318 tests passed and local coverage was 96.195%, then
|
||||
isolated the enforced failure to SonarQube `new_coverage=0.0` against the
|
||||
threshold of 80.
|
||||
- It corroborated the retained Jenkins evidence with current quality
|
||||
metrics and healthy Soteria/Flux state, and did not blame Kubernetes.
|
||||
- It proposed no change until scanner import and baseline evidence can
|
||||
distinguish configuration error from genuinely uncovered new code.
|
||||
|
||||
## Proof 7: Ananke build 242 local-coverage triage
|
||||
|
||||
- Hermes selected terminal Ananke build 242 and isolated the first failure
|
||||
to the local coverage gate at 61.8%.
|
||||
- It showed that tests, SonarQube, supply chain, LOC, docs naming, gate glue,
|
||||
Flux, and cluster readiness were healthy.
|
||||
- It distinguished the local 61.8% scope from SonarQube's 93.1% new-code
|
||||
scope instead of treating the numbers as interchangeable.
|
||||
- Direct Jenkins artifact reads returned HTTP 403, so it explicitly named
|
||||
Ariadne's retained bundle as evidence provenance and stopped short of an
|
||||
unsupported fix.
|
||||
|
||||
## Use Hermes next
|
||||
|
||||
- `What is broken in the cluster right now?`
|
||||
|
||||
@ -24,7 +24,7 @@ spec:
|
||||
ai.bstein.dev/model: openai-codex/gpt-5.6-terra with local gpt-oss:20b fallback
|
||||
ai.bstein.dev/role: testing-triage
|
||||
ai.bstein.dev/placement: arm64 gateway lane (rpi5 preferred)
|
||||
ai.bstein.dev/config-rev: "20260803-all-service-triage-map"
|
||||
ai.bstein.dev/config-rev: "20260804-operator-guide"
|
||||
spec:
|
||||
serviceAccountName: hermes-triage
|
||||
automountServiceAccountToken: true
|
||||
@ -88,6 +88,7 @@ spec:
|
||||
cp /config/AGENTS.md /opt/data/workspace/AGENTS.md
|
||||
cp /config/START-HERE.md /opt/data/workspace/START-HERE.md
|
||||
cp /config/HERMES-CAPABILITIES.md /opt/data/workspace/HERMES-CAPABILITIES.md
|
||||
cp /guide/OPERATOR-RUNBOOK.md /opt/data/workspace/HERMES-OPERATOR-RUNBOOK.md
|
||||
cp /config/ATLAS-TRIAGE-PROOFS.md /opt/data/workspace/triage-proof/ATLAS-TRIAGE-PROOFS.md
|
||||
touch /opt/data/.env
|
||||
if ! grep -q '^API_SERVER_KEY=' /opt/data/.env; then
|
||||
@ -104,6 +105,8 @@ spec:
|
||||
mountPath: /opt/data
|
||||
- name: config
|
||||
mountPath: /config
|
||||
- name: operator-guide
|
||||
mountPath: /guide
|
||||
resources:
|
||||
requests:
|
||||
cpu: 25m
|
||||
@ -234,6 +237,9 @@ spec:
|
||||
- name: config
|
||||
configMap:
|
||||
name: hermes-config
|
||||
- name: operator-guide
|
||||
configMap:
|
||||
name: hermes-operator-guide
|
||||
- name: tools
|
||||
emptyDir: {}
|
||||
- name: triage-skill
|
||||
|
||||
@ -22,6 +22,12 @@ resources:
|
||||
- agent-ingress.yaml
|
||||
|
||||
configMapGenerator:
|
||||
- name: hermes-operator-guide
|
||||
namespace: hermes
|
||||
files:
|
||||
- OPERATOR-RUNBOOK.md=NOTES.md
|
||||
options:
|
||||
disableNameSuffixHash: true
|
||||
- name: hermes-triage-skill
|
||||
namespace: hermes
|
||||
files:
|
||||
|
||||
Loading…
x
Reference in New Issue
Block a user