3.9 KiB
3.9 KiB
Mastery and claim rubric
Performance levels
| Level | Evidence |
|---|---|
| Exposed | Recognizes terms but needs the path and commands supplied. |
| Assisted | Completes a lab with hints and can explain the result afterward. |
| Independent | Selects safe tools, gathers fresh evidence, and reaches a defensible result without hints. |
| Mastery | Handles ambiguity and failure, teaches the architecture, improves the reusable workflow, and preserves security boundaries. |
Score each domain separately:
- Architecture and request routing.
- Models, context, GPU ownership, and fallback.
- Sessions, files, profiles, logs, and recovery.
- Tools, skills, plugins, MCP, channels, webhooks, and cron.
- Identity, RBAC, network, credentials, approvals, and auditability.
- Jenkins/Flux/Kubernetes/metrics evidence correlation.
- Clear triage findings and minimal repo-side proposals.
- Skill creation, evaluation, and iteration.
Never award Independent from a verbal answer alone when the lab calls for live
or fixture evidence. Never award Mastery until Brad has led at least two
different real incidents and improved a workflow based on what failed.
Interview claim audit
| Claim | What supports it | Minimum demonstration |
|---|---|---|
| Hermes runs in the Kubernetes cluster | Flux manifests and live workloads | Trace browser, agent, model gate, Ollama, PVC, and service account. |
| Hermes fits existing automation | Operator config, internal evidence URLs, triage skill | Complete a triage using Ariadne, Jenkins, Git/Flux, cluster health, and metrics. |
| Hermes follows Brad's triage path | $triage-titan-test-failures procedure |
Lead two evidence-backed cases with facts separated from inference. |
| Repeated workflows become reusable skills | Versioned triage/training skills plus writable user skill area | Build, test, and improve one skill from a repeated real workflow. |
| Access starts read-only and scoped | RBAC, NetworkPolicy, separate identities and state | Produce and explain the allow/deny matrix from live authorization checks. |
| Human approval protects changes | Approval config, deny patterns, Flux workflow | Identify every mutation in a proposed incident response and stop before it. |
| Local GPU gracefully yields to gaming | Owner state, model gate, fallback | Explain or observe a Wolf ownership window and verify continued fallback service. |
| OpenClaw previously ran and was replaced | 1bc58e10 deploys the OpenClaw triage stack; follow-up commits add gateway/OIDC/persistence; 97ecb36b replaces it with Hermes; bb5e286f retires the dependency; the bound openclaw-home PVC remains without a workload |
Inspect the commits and live Flux/PVC state, then explain what they prove versus what depends on Brad's personal recollection. |
Triage answer scoring
Award one point for each:
- names the correct failure class;
- cites timestamped/build-specific evidence;
- labels inference;
- checks evidence freshness;
- identifies realistic blast radius;
- orders read-only next checks;
- proposes the smallest Flux/repo-side change or says evidence is insufficient;
- isolates approval-required actions;
- avoids invented resources, logs, metrics, or commits;
- remains concise enough for an operator to act on.
Scores below 8/10 require another incident drill. Any invented evidence or
unacknowledged mutation caps the result at Assisted.
Progress record format
When Brad approves recording progress, use this compact structure:
# Hermes training progress
- Last session: <UTC timestamp>
- Current level: <per-domain summary>
- Completed: <labs with evidence>
- Needs repetition: <specific gaps>
- Next lab: <one lab>
- Safety findings: <none or concrete issue>
- Claim status: <supported, partially supported, unproven>
Append evidence links or command summaries; never record tokens, cookies, credentials, device codes, or Secret values.