diff --git a/services/hermes/skills/master-hermes-on-atlas/SKILL.md b/services/hermes/skills/master-hermes-on-atlas/SKILL.md index 9258a05e3..e33e487ac 100644 --- a/services/hermes/skills/master-hermes-on-atlas/SKILL.md +++ b/services/hermes/skills/master-hermes-on-atlas/SKILL.md @@ -11,7 +11,7 @@ substitute a lecture for a lab. ## Select the training model -Use `openai-codex/gpt-5.4` for assessments, multi-reference labs, incident +Use `openai-codex/gpt-5.6-terra` for assessments, multi-reference labs, incident grading, and skill evaluation. If the active provider is the local `gpt-oss:20b`, stop before reading the reference files and ask Brad to select Codex in Models or start a Codex-backed session. Continue locally only when Brad diff --git a/services/hermes/skills/master-hermes-on-atlas/references/architecture.md b/services/hermes/skills/master-hermes-on-atlas/references/architecture.md index e1ff06308..c35188d6e 100644 --- a/services/hermes/skills/master-hermes-on-atlas/references/architecture.md +++ b/services/hermes/skills/master-hermes-on-atlas/references/architecture.md @@ -18,23 +18,25 @@ through the chat Service or Ingress. ## Inference path ```text -operator Hermes ─┐ - ├─> hermes-model-gate ─> Ollama gpt-oss:20b on titan-24 GPU -consumer Hermes ─┘ │ - └─ 503 while Wolf owns the GPU - -each instance independently ─────> openai-codex/gpt-5.4 fallback +operator Hermes ───> independently authenticated openai-codex/gpt-5.6-terra +consumer Hermes ───> independently authenticated openai-codex/gpt-5.6-terra + │ + └─ provider error or manual selection ─> hermes-model-gate + │ + ├─ Hermes owns GPU ─> Ollama gpt-oss:20b + └─ Wolf owns GPU ───> HTTP 503 ``` - The configured context length is 64,000 tokens. - `hermes-model-gate` is the stable OpenAI-compatible endpoint. - A coordination Lease stores whether Hermes or Wolf/Moonlight owns the GPU. Ariadne has narrowly scoped permission to update that Lease during handoff. -- Wolf ownership makes the gate return a deliberate unavailable response. The - Hermes gateway stays alive and can use its configured fallback. +- Wolf ownership makes the local gate return a deliberate unavailable response. + Normal interactive chat remains on its Codex primary and does not consume + titan-24 GPU resources. - The two agent pods do not consume titan-24 GPU memory. Ollama does. Agent pods are ARM gateway workloads with persistent state on separate PVCs. -- Fallback credentials are deliberately per-instance. Never copy Brad's Codex +- Codex credentials are deliberately per-instance. Never copy Brad's Codex credential store into the consumer PVC. Verify rather than memorize: diff --git a/services/hermes/skills/master-hermes-on-atlas/references/curriculum.md b/services/hermes/skills/master-hermes-on-atlas/references/curriculum.md index e39a456a2..3c3364bde 100644 --- a/services/hermes/skills/master-hermes-on-atlas/references/curriculum.md +++ b/services/hermes/skills/master-hermes-on-atlas/references/curriculum.md @@ -4,7 +4,7 @@ Complete labs by evidence, not elapsed time. A focused pass can establish operational competence in several days; mastery requires repeating real triage and recovery work over multiple incidents. -Run the curriculum on `openai-codex/gpt-5.4`. Use the local model only for a +Run the curriculum on `openai-codex/gpt-5.6-terra`. Use the local model only for a deliberate comparison lab; it is not the default coach for multi-reference work. ## Phase 1: orientation and control @@ -22,8 +22,8 @@ agent to model, including the fallback branch and the consumer boundary. ### Lab 2 — Models, context, and fallback Inspect `hermes status`, `hermes fallback list`, deployment placement, and GPU -owner state. Explain why a 32K model was rejected, why the gateway remains up -when Wolf owns the GPU, and which credentials a fallback consumes. +owner state. Explain why a 32K model was rejected, why normal Codex-backed chat +remains available when Wolf owns the GPU, and when the local fallback can run. Success evidence: predict outcomes for local healthy, local slow, gate 503, invalid local response, and expired Codex authorization without changing state. diff --git a/services/hermes/skills/master-hermes-on-atlas/references/two-hour-proof-sprint.md b/services/hermes/skills/master-hermes-on-atlas/references/two-hour-proof-sprint.md index 6db1f6581..ace25f926 100644 --- a/services/hermes/skills/master-hermes-on-atlas/references/two-hour-proof-sprint.md +++ b/services/hermes/skills/master-hermes-on-atlas/references/two-hour-proof-sprint.md @@ -9,12 +9,13 @@ Brad commits to a result. ## 0–15 minutes — control surface and request path -1. Have Brad select `openai-codex/gpt-5.4` for the training session. +1. Have Brad start a new session and confirm it uses + `openai-codex/gpt-5.6-terra`. 2. In the web UI, have him locate Chat, Sessions, Files, Models, Logs, Skills, Plugins, MCP, Channels, Webhooks, Cron, Profiles, Config, Keys, System, and Documentation. -3. Have him explain browser → operator agent → model gate → Ollama, including - Codex fallback and the separate consumer instance. +3. Have him explain browser → operator agent → Codex primary, including the + model-gate/Ollama fallback branch and the separate consumer instance. Evidence: Brad can identify persistent, shared, and instance-local state and can name the GPU owner check without changing it.