5.3 KiB
Raw Permalink Blame History

Hands-on curriculum

Complete labs by evidence, not elapsed time. A focused pass can establish operational competence in several days; mastery requires repeating real triage and recovery work over multiple incidents.

Run the curriculum on openai-codex/gpt-5.6-terra. Use the local model only for a deliberate comparison lab; it is not the default coach for multi-reference work.

Phase 1: orientation and control

Lab 1 — Map the system

From the web UI, identify Chat, Sessions, Files, Models, Logs, Cron, Skills, Plugins, MCP, Channels, Webhooks, Pairing, Profiles, Config, Keys, System, and Documentation. Explain which state belongs to the operator PVC and which components are shared. Verify three claims with read-only CLI output.

Success evidence: a correct diagram or written request path from browser to agent to model, including the fallback branch and the consumer boundary.

Lab 2 — Models, context, and fallback

Inspect hermes status, hermes fallback list, deployment placement, and GPU owner state. Explain why a 32K model was rejected, why normal Codex-backed chat remains available when Wolf owns the GPU, and when the local fallback can run.

Success evidence: predict outcomes for local healthy, local slow, gate 503, invalid local response, and expired Codex authorization without changing state.

Lab 3 — Sessions, files, profiles, and logs

Create a named training session and a harmless workspace note through the UI, then find the corresponding session/file/log surfaces. Inspect profile and backup help without creating a profile or backup.

Success evidence: explain persistence, what survives a pod replacement, what is instance-local, and how to recover a lost conversation without exposing auth.

Phase 2: tools and safe autonomy

Lab 4 — Tools versus skills versus MCP

Use hermes tools, hermes skills, hermes plugins, and hermes mcp help or list/status views. Classify each as executable capability, procedure/context, packaged extension, or external protocol integration. Explain why a skill is not a security boundary.

Success evidence: choose the right mechanism for three examples: repeatable Titan triage, a read-only external API, and a scheduled notification.

Lab 5 — Prove the permission boundary

Use kubectl auth can-i as both service accounts. Inspect the relevant ClusterRoles and NetworkPolicies. Do not attempt mutations.

Success evidence: an allow/deny matrix covering pods, logs, Secrets, exec, deployment patch, Flux reads, Flux reconcile, the model gate, internal services, and public web research.

Lab 6 — Channels, pairing, webhooks, and cron

Inspect the configured state and help for gateway/channels, pairing, webhooks, and cron. Design one safe notification workflow and one unsafe workflow. Do not register a channel, create a webhook, or schedule a job during the lab.

Success evidence: identify the identity, secret, audience, tool policy, failure mode, audit trail, and revocation path for the proposed integration.

Phase 3: the interview workflow

Lab 7 — Deterministic live triage

Invoke $triage-titan-test-failures. Read the current Ariadne diagnosis and bundle, then correlate one suite with Jenkins, Flux/Git, cluster health, and quality metrics. Keep every collection step read-only.

Success evidence: the required structured triage result with timestamps and no invented evidence.

Lab 8 — Incident simulations

Use one student packet from incident-drills.md. Diagnose it before receiving coach notes. Repeat until three cases score at least Independent.

Success evidence: correct failure class, causal chain, next checks, minimal repo-side correction, and explicit approval boundary.

Lab 9 — Build and evaluate a reusable skill

Identify a stable repeated workflow, create a small skill in the operator's writable skills directory, inspect it, and test it against a fresh incident. Do not alter the Flux-mounted training or triage skills.

Success evidence: precise trigger description, concise procedure, progressive disclosure where useful, one passing case, one adversarial case, and a stated permission boundary.

Lab 10 — Model outage drill

Use retained logs/current status or a separately approved maintenance window. Do not take the live model down for training. Explain the expected trace from model gate to fallback, recognize an auth failure, and name the restoration checks.

Success evidence: correctly distinguish model unavailability, context rejection, provider auth failure, and agent failure.

Phase 4: independent operation

Lab 11 — Supervised real incident

Lead a real failure triage from intake to an approved repo-side proposal. A human applies any change through Git/Flux. Verify the result read-only.

Lab 12 — Teach it back and audit the claims

Explain the complete architecture and demonstrate the workflow without hints. Audit each sentence of the interview answer using the claim table in mastery-rubric.md. Weaken or qualify anything not supported by evidence.

Suggested pace

  • Day 1: Labs 13.
  • Days 23: Labs 46.
  • Days 47: Labs 710.
  • Weeks 24: repeat Labs 711 on real incidents; improve one skill from observed failures; finish with Lab 12.

Do not advance solely because a day elapsed. Re-run any lab graded below Independent.