ai(hermes): add operator mastery lab

This commit is contained in:
jenkins 2026-08-02 02:47:43 -03:00
parent 1494e76ccd
commit 85c5bfe9f4
9 changed files with 503 additions and 1 deletions

View File

@ -23,7 +23,7 @@ spec:
ai.bstein.dev/model: gpt-oss:20b with openai-codex fallback
ai.bstein.dev/role: testing-triage
ai.bstein.dev/placement: arm64 gateway lane (rpi5 preferred)
ai.bstein.dev/config-rev: "20260801-gpu-owner-fallback"
ai.bstein.dev/config-rev: "20260802-mastery-lab"
spec:
serviceAccountName: hermes-triage
automountServiceAccountToken: true
@ -50,6 +50,7 @@ spec:
- titan-13
- titan-15
- titan-17
- titan-18
- titan-19
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
@ -193,6 +194,9 @@ spec:
- name: triage-skill
mountPath: /opt/data/workspace/skills/triage-titan-test-failures
readOnly: true
- name: mastery-skill
mountPath: /opt/data/workspace/skills/master-hermes-on-atlas
readOnly: true
readinessProbe:
httpGet:
path: /api/status
@ -231,3 +235,19 @@ spec:
path: SKILL.md
- key: openai.yaml
path: agents/openai.yaml
- name: mastery-skill
configMap:
name: hermes-mastery-skill
items:
- key: SKILL.md
path: SKILL.md
- key: openai.yaml
path: agents/openai.yaml
- key: architecture.md
path: references/architecture.md
- key: curriculum.md
path: references/curriculum.md
- key: incident-drills.md
path: references/incident-drills.md
- key: mastery-rubric.md
path: references/mastery-rubric.md

View File

@ -29,3 +29,14 @@ configMapGenerator:
- openai.yaml=skills/triage-titan-test-failures/agents/openai.yaml
options:
disableNameSuffixHash: true
- name: hermes-mastery-skill
namespace: hermes
files:
- SKILL.md=skills/master-hermes-on-atlas/SKILL.md
- openai.yaml=skills/master-hermes-on-atlas/agents/openai.yaml
- architecture.md=skills/master-hermes-on-atlas/references/architecture.md
- curriculum.md=skills/master-hermes-on-atlas/references/curriculum.md
- incident-drills.md=skills/master-hermes-on-atlas/references/incident-drills.md
- mastery-rubric.md=skills/master-hermes-on-atlas/references/mastery-rubric.md
options:
disableNameSuffixHash: true

View File

@ -99,6 +99,7 @@ spec:
- --config=/vault/secrets/oidc-config
- --redirect-url=https://agent.bstein.dev/oauth2/callback
- --oidc-issuer-url=https://sso.bstein.dev/realms/atlas
- --code-challenge-method=S256
- --scope=openid profile email
- --email-domain=*
- --authenticated-emails-file=/etc/oauth2-proxy/allowed-emails

View File

@ -0,0 +1,71 @@
---
name: master-hermes-on-atlas
description: Coach Brad through hands-on mastery of the live Atlas Hermes deployment, including its web and CLI interfaces, local GPU inference and Codex fallback, tools, skills, sessions, profiles, channels, security boundaries, and evidence-driven test-failure triage. Use for Hermes training, guided labs, knowledge checks, incident simulations, capability tours, or checking whether the Hermes interview claims are supported by demonstrated work.
---
# Master Hermes on Atlas
Act as a demanding, practical coach. Teach the deployed system, not a generic
Hermes installation. Make Brad perform the work and explain the result; do not
substitute a lecture for a lab.
## Start or resume training
1. Read `references/architecture.md` before teaching deployment-specific facts.
2. Read `references/curriculum.md` to select the next lab.
3. Ask whether to assess, resume, or choose a lab. If no progress record exists,
default to a five-question assessment followed by the first weak area.
4. Give one bounded task at a time. State the goal, safety boundary, exact
success evidence, and at most one initial hint.
5. Wait for Brad's answer or observed command output before revealing the
explanation.
6. Grade with `references/mastery-rubric.md`. Separate demonstrated ability
from verbal familiarity.
7. Offer a concise progress update. Write it to
`/opt/data/workspace/hermes-training/progress.md` only after Brad explicitly
approves the file change.
## Run live exercises safely
- Use read-only commands by default: `kubectl get`, `describe`, `logs`, `auth
can-i`, HTTP GET, `hermes status`, and Hermes list/status commands.
- Never read Kubernetes Secret values. Never use Vault reads as a training
shortcut. Redact bearer tokens, cookies, device codes, and credentials.
- Never mutate Kubernetes, Flux, Jenkins, Git, credentials, inference ownership,
channels, cron jobs, plugins, MCP servers, or profiles merely to demonstrate
a feature.
- Draft state-changing commands and explain their effect. Execute only when Brad
separately requests the change and the configured approval path permits it.
- Do not use `hermes --oneshot` for a prompt that could modify state: that mode
bypasses interactive approval prompts.
- Treat a failed command as evidence to interpret, not a reason to broaden
access or repeatedly retry.
## Teach evidence-driven triage
For a live test-failure lab, also use `$triage-titan-test-failures`. Require the
student to distinguish:
- observed fact from inference;
- stale evidence from current state;
- application regression from environment failure;
- a read-only next check from a proposed mutation;
- a generic suggestion from the smallest Flux-tracked repo-side fix.
For offline practice, read `references/incident-drills.md` and present only the
student packet for one incident. Keep the coach notes hidden until after the
student commits to a finding.
## Validate claims honestly
Use the claim audit in `references/mastery-rubric.md`. Current configuration and
live demonstrations can support current-use claims. They cannot prove that
OpenClaw previously ran or was replaced; require Git history, retained manifests,
or other evidence before treating that historical claim as demonstrated.
End every session with:
- what Brad demonstrated;
- what remains unproven;
- the next smallest lab;
- any safety or operational issue discovered.

View File

@ -0,0 +1,4 @@
interface:
display_name: "Master Hermes on Atlas"
short_description: "Practice Hermes safely on the live Atlas stack"
default_prompt: "Use $master-hermes-on-atlas to assess me and begin my next hands-on Hermes lab."

View File

@ -0,0 +1,97 @@
# Live Atlas Hermes architecture
Use this reference for deployment-specific facts. Re-check live state before
asserting health, placement, ownership, or current model availability.
## Two isolated agent instances
| Surface | Purpose | Identity boundary | State and permissions |
| --- | --- | --- | --- |
| `agent.bstein.dev` | Brad's operator and triage lab | Keycloak plus an outer oauth2-proxy exact-email allow-list for `brad@bstein.dev` | `hermes` namespace, its own PVC and service account; read-only cluster triage plus approved internal evidence endpoints |
| `chat.bstein.dev` | Consumer chat and research | Keycloak login | `hermes-chat` namespace, separate PVC/config/auth/sessions/skills; read-only cluster observer RBAC and no access to Secrets or Kubernetes mutation APIs |
The instances do not share conversation state, credentials, profiles, skills
created on their PVCs, or Kubernetes identities. They share only the inference
service described below. Consumer users cannot reach the operator dashboard
through the chat Service or Ingress.
## Inference path
```text
operator Hermes ─┐
├─> hermes-model-gate ─> Ollama gpt-oss:20b on titan-24 GPU
consumer Hermes ─┘ │
└─ 503 while Wolf owns the GPU
each instance independently ─────> openai-codex/gpt-5.4 fallback
```
- The configured context length is 64,000 tokens.
- `hermes-model-gate` is the stable OpenAI-compatible endpoint.
- A GPU-owner ConfigMap and controller coordinate Hermes versus Wolf/Moonlight.
- Wolf ownership makes the gate return a deliberate unavailable response. The
Hermes gateway stays alive and can use its configured fallback.
- The two agent pods do not consume titan-24 GPU memory. Ollama does. Agent pods
are ARM gateway workloads with persistent state on separate PVCs.
- Fallback credentials are deliberately per-instance. Never copy Brad's Codex
credential store into the consumer PVC.
Verify rather than memorize:
```sh
hermes status
hermes fallback list
kubectl -n hermes get configmap hermes-gpu-owner -o yaml
kubectl -n hermes get deploy hermes hermes-model-gate hermes-ollama -o wide
kubectl -n hermes-chat get deploy hermes-chat -o wide
```
Do not change the GPU owner as part of a lesson.
## Operator evidence path
The operator role is designed for supervised test-failure triage:
1. Ariadne deterministic bundle and optional diagnosis.
2. Retained Jenkins logs and artifacts.
3. Recent Git changes and Flux revision/state.
4. Kubernetes workload, node, event, storage, DNS, and dependency health.
5. Pushgateway quality metrics and VictoriaMetrics/Grafana context.
6. A concise finding, confidence, evidence, blast radius, next checks,
repo-side fix, and approval-required section.
The versioned `$triage-titan-test-failures` skill encodes this workflow. The
operator has internal URLs for Ariadne, Jenkins, and VictoriaMetrics. The
consumer instance intentionally does not.
## Security layers and their limits
Security is layered rather than delegated to a prompt:
1. Keycloak authenticates the human.
2. The operator proxy limits `agent.bstein.dev` to Brad's exact email.
3. Separate namespaces, PVCs, service accounts, configs, and ingresses isolate
operator and consumer state.
4. Kubernetes RBAC grants the consumer get/list/watch only and excludes Secrets,
exec, attach, port-forward, and mutation.
5. NetworkPolicy permits the consumer to reach DNS, the shared model gate, the
Kubernetes API, and public IPv4 while blocking arbitrary private service
access.
6. Hermes deny patterns and instructions add user-facing guardrails.
Prompt rules and shell deny patterns are not the primary security boundary.
RBAC, network policy, credential separation, and routing isolation are. The
consumer can still disclose non-secret cluster metadata it is authorized to
read; that is an intentional requirement and should be described honestly.
## CLI-to-web mental model
The pinned Hermes build exposes chat, models and fallback, auth, sessions,
profiles, skills and bundles, plugins, tools, MCP, gateway/channels, pairing,
webhooks, cron, kanban, projects, hooks, memory/journey, logs, security, backup,
and diagnostics. The web interface is a control surface over many of these
same persisted capabilities.
Use `<command> --help` and the web Documentation page for the installed version.
Do not rely on screenshots or online docs from a different release when a live
command can settle the question.

View File

@ -0,0 +1,128 @@
# Hands-on curriculum
Complete labs by evidence, not elapsed time. A focused pass can establish
operational competence in several days; mastery requires repeating real triage
and recovery work over multiple incidents.
## Phase 1: orientation and control
### Lab 1 — Map the system
From the web UI, identify Chat, Sessions, Files, Models, Logs, Cron, Skills,
Plugins, MCP, Channels, Webhooks, Pairing, Profiles, Config, Keys, System, and
Documentation. Explain which state belongs to the operator PVC and which
components are shared. Verify three claims with read-only CLI output.
Success evidence: a correct diagram or written request path from browser to
agent to model, including the fallback branch and the consumer boundary.
### Lab 2 — Models, context, and fallback
Inspect `hermes status`, `hermes fallback list`, deployment placement, and GPU
owner state. Explain why a 32K model was rejected, why the gateway remains up
when Wolf owns the GPU, and which credentials a fallback consumes.
Success evidence: predict outcomes for local healthy, local slow, gate 503,
invalid local response, and expired Codex authorization without changing state.
### Lab 3 — Sessions, files, profiles, and logs
Create a named training session and a harmless workspace note through the UI,
then find the corresponding session/file/log surfaces. Inspect profile and
backup help without creating a profile or backup.
Success evidence: explain persistence, what survives a pod replacement, what is
instance-local, and how to recover a lost conversation without exposing auth.
## Phase 2: tools and safe autonomy
### Lab 4 — Tools versus skills versus MCP
Use `hermes tools`, `hermes skills`, `hermes plugins`, and `hermes mcp` help or
list/status views. Classify each as executable capability, procedure/context,
packaged extension, or external protocol integration. Explain why a skill is
not a security boundary.
Success evidence: choose the right mechanism for three examples: repeatable
Titan triage, a read-only external API, and a scheduled notification.
### Lab 5 — Prove the permission boundary
Use `kubectl auth can-i` as both service accounts. Inspect the relevant
ClusterRoles and NetworkPolicies. Do not attempt mutations.
Success evidence: an allow/deny matrix covering pods, logs, Secrets, exec,
deployment patch, Flux reads, Flux reconcile, the model gate, internal services,
and public web research.
### Lab 6 — Channels, pairing, webhooks, and cron
Inspect the configured state and help for gateway/channels, pairing, webhooks,
and cron. Design one safe notification workflow and one unsafe workflow. Do not
register a channel, create a webhook, or schedule a job during the lab.
Success evidence: identify the identity, secret, audience, tool policy, failure
mode, audit trail, and revocation path for the proposed integration.
## Phase 3: the interview workflow
### Lab 7 — Deterministic live triage
Invoke `$triage-titan-test-failures`. Read the current Ariadne diagnosis and
bundle, then correlate one suite with Jenkins, Flux/Git, cluster health, and
quality metrics. Keep every collection step read-only.
Success evidence: the required structured triage result with timestamps and no
invented evidence.
### Lab 8 — Incident simulations
Use one student packet from `incident-drills.md`. Diagnose it before receiving
coach notes. Repeat until three cases score at least `Independent`.
Success evidence: correct failure class, causal chain, next checks, minimal
repo-side correction, and explicit approval boundary.
### Lab 9 — Build and evaluate a reusable skill
Identify a stable repeated workflow, create a small skill in the operator's
writable skills directory, inspect it, and test it against a fresh incident.
Do not alter the Flux-mounted training or triage skills.
Success evidence: precise trigger description, concise procedure, progressive
disclosure where useful, one passing case, one adversarial case, and a stated
permission boundary.
### Lab 10 — Model outage drill
Use retained logs/current status or a separately approved maintenance window.
Do not take the live model down for training. Explain the expected trace from
model gate to fallback, recognize an auth failure, and name the restoration
checks.
Success evidence: correctly distinguish model unavailability, context rejection,
provider auth failure, and agent failure.
## Phase 4: independent operation
### Lab 11 — Supervised real incident
Lead a real failure triage from intake to an approved repo-side proposal. A
human applies any change through Git/Flux. Verify the result read-only.
### Lab 12 — Teach it back and audit the claims
Explain the complete architecture and demonstrate the workflow without hints.
Audit each sentence of the interview answer using the claim table in
`mastery-rubric.md`. Weaken or qualify anything not supported by evidence.
## Suggested pace
- Day 1: Labs 13.
- Days 23: Labs 46.
- Days 47: Labs 710.
- Weeks 24: repeat Labs 711 on real incidents; improve one skill from observed
failures; finish with Lab 12.
Do not advance solely because a day elapsed. Re-run any lab graded below
`Independent`.

View File

@ -0,0 +1,95 @@
# Incident drills
Present only one `Student packet` before the student answers. Use the coach
notes afterward to grade the failure classification and reasoning. These cases
are based on retained Atlas rollout observations; treat them as training
fixtures, not proof of current live state.
## Case A — Metrics disappear after a Soteria run
### Student packet
- The Soteria test stages ran.
- The post-stage log ends with:
`Syntax error: end of file unexpected (expecting "fi")`.
- The canonical Soteria Pushgateway series did not refresh.
- Other suites continued publishing.
Ask for: finding, confidence, evidence, likely cause, blast radius, read-only
next checks, repo-side fix, and approval-required actions.
### Coach notes
The primary class is pipeline glue/shell syntax, not a Pushgateway outage and
not a product-test failure. Inspect the post block and the exact shell parsed by
`/bin/sh`; confirm the last successful metric timestamp. The smallest likely
fix is closing/correcting the conditional in the Jenkinsfile, then a normal SCM
run and read-only metric verification. Editing, pushing, triggering, or manual
metric backfill requires explicit approval.
## Case B — UI tests fail before metrics publish
### Student packet
- `bstein-dev-home` requires Playwright `1.59.1`.
- The CI test container contains Playwright `1.51.0`.
- The frontend test stage hard-fails.
- No fresh canonical quality metrics appear for the run.
### Coach notes
Separate the direct test-environment mismatch from the telemetry-control-flow
defect. The version mismatch explains the test failure; early pipeline abort
explains missing metrics. A complete proposal aligns the image version and
persists stage return codes so publishing still runs while the final gate still
fails. Do not call this a Kubernetes capacity problem without node/pod evidence.
## Case C — Promotion exits 127
### Student packet
- Titan IaC validation and tests finish.
- The Promote stage uses `python:3.12-slim`.
- The stage returns exit code `127` at its first Git command.
- The repository and remote are reachable from other jobs.
### Coach notes
Exit 127 means the invoked command is unavailable. Confirm the exact log line
and image contents; the likely issue is that the slim image lacks `git`, not bad
Git credentials. The minimal repo-side fix installs `git` and CA certificates
in the job's dependency setup or uses a suitable pinned runner image.
## Case D — Data Preppers cannot publish an image
### Student packet
- Tests and local image build complete.
- Push to `registry.bstein.dev/streaming/data-prepper:2.8.0` is denied.
- The configured Jenkins credential belongs to a different Harbor project.
- A Sonar evidence GET also returns `401` in the same run.
### Coach notes
Do not collapse independent authentication failures into a generic network
failure. Verify the Harbor repository/project and credential ID without reading
the secret value; separately verify the Sonar token injection path and API
scope. The likely image fix is selecting a streaming-scoped robot credential.
Treat the Sonar 401 as a second issue unless evidence proves a shared cause.
## Case E — Hermes local inference fails during game streaming
### Student packet
- Wolf owns titan-24 according to the GPU-owner state.
- `hermes-model-gate` is ready but returns a deliberate unavailable response.
- Both Hermes dashboards remain healthy.
- One instance answers through Codex; the other reports missing provider auth.
### Coach notes
This is expected resource arbitration plus an instance-specific fallback auth
gap. Do not restart the agent or claim the GPU is underprovisioned. Confirm the
gate response, fallback chain, and each instance's auth status without exposing
tokens. Authorize the affected instance with its intended user's account. The
separate credential stores are a security property, not configuration drift.

View File

@ -0,0 +1,75 @@
# Mastery and claim rubric
## Performance levels
| Level | Evidence |
| --- | --- |
| Exposed | Recognizes terms but needs the path and commands supplied. |
| Assisted | Completes a lab with hints and can explain the result afterward. |
| Independent | Selects safe tools, gathers fresh evidence, and reaches a defensible result without hints. |
| Mastery | Handles ambiguity and failure, teaches the architecture, improves the reusable workflow, and preserves security boundaries. |
Score each domain separately:
1. Architecture and request routing.
2. Models, context, GPU ownership, and fallback.
3. Sessions, files, profiles, logs, and recovery.
4. Tools, skills, plugins, MCP, channels, webhooks, and cron.
5. Identity, RBAC, network, credentials, approvals, and auditability.
6. Jenkins/Flux/Kubernetes/metrics evidence correlation.
7. Clear triage findings and minimal repo-side proposals.
8. Skill creation, evaluation, and iteration.
Never award `Independent` from a verbal answer alone when the lab calls for live
or fixture evidence. Never award `Mastery` until Brad has led at least two
different real incidents and improved a workflow based on what failed.
## Interview claim audit
| Claim | What supports it | Minimum demonstration |
| --- | --- | --- |
| Hermes runs in the Kubernetes cluster | Flux manifests and live workloads | Trace browser, agent, model gate, Ollama, PVC, and service account. |
| Hermes fits existing automation | Operator config, internal evidence URLs, triage skill | Complete a triage using Ariadne, Jenkins, Git/Flux, cluster health, and metrics. |
| Hermes follows Brad's triage path | `$triage-titan-test-failures` procedure | Lead two evidence-backed cases with facts separated from inference. |
| Repeated workflows become reusable skills | Versioned triage/training skills plus writable user skill area | Build, test, and improve one skill from a repeated real workflow. |
| Access starts read-only and scoped | RBAC, NetworkPolicy, separate identities and state | Produce and explain the allow/deny matrix from live authorization checks. |
| Human approval protects changes | Approval config, deny patterns, Flux workflow | Identify every mutation in a proposed incident response and stop before it. |
| Local GPU gracefully yields to gaming | Owner state, model gate, fallback | Explain or observe a Wolf ownership window and verify continued fallback service. |
| OpenClaw previously ran and was replaced | Not established by the current Hermes deployment | Show retained Git history/manifests/logs or qualify the statement as recollection. |
## Triage answer scoring
Award one point for each:
- names the correct failure class;
- cites timestamped/build-specific evidence;
- labels inference;
- checks evidence freshness;
- identifies realistic blast radius;
- orders read-only next checks;
- proposes the smallest Flux/repo-side change or says evidence is insufficient;
- isolates approval-required actions;
- avoids invented resources, logs, metrics, or commits;
- remains concise enough for an operator to act on.
Scores below 8/10 require another incident drill. Any invented evidence or
unacknowledged mutation caps the result at `Assisted`.
## Progress record format
When Brad approves recording progress, use this compact structure:
```markdown
# Hermes training progress
- Last session: <UTC timestamp>
- Current level: <per-domain summary>
- Completed: <labs with evidence>
- Needs repetition: <specific gaps>
- Next lab: <one lab>
- Safety findings: <none or concrete issue>
- Claim status: <supported, partially supported, unproven>
```
Append evidence links or command summaries; never record tokens, cookies,
credentials, device codes, or Secret values.