ai(hermes): include operator proof sprint
This commit is contained in:
parent
4991493d3b
commit
96edd28b53
@ -0,0 +1,82 @@
|
||||
# Atlas two-hour proof sprint
|
||||
|
||||
The sprint establishes usable, defensible evidence for the current Hermes
|
||||
claims. It does not award mastery and it cannot retroactively prove historical
|
||||
OpenClaw use. Brad performs the work; Hermes gives one task at a time, checks
|
||||
the evidence, and withholds later answers until Brad commits to a result.
|
||||
|
||||
## 0–15 minutes — control surface and request path
|
||||
|
||||
1. Have Brad select `openai-codex/gpt-5.4` for the training session.
|
||||
2. In the web UI, have him locate Chat, Sessions, Files, Models, Logs, Skills,
|
||||
Plugins, MCP, Channels, Webhooks, Cron, Profiles, Config, Keys, System, and
|
||||
Documentation.
|
||||
3. Have him explain browser → operator agent → model gate → Ollama, including
|
||||
Codex fallback and the separate consumer instance.
|
||||
|
||||
Evidence: Brad can identify persistent, shared, and instance-local state and
|
||||
can name the GPU owner check without changing it.
|
||||
|
||||
## 15–35 minutes — permissions and safety
|
||||
|
||||
1. Have Brad inspect the operator ServiceAccount, ClusterRole, NetworkPolicy,
|
||||
and approval deny list using read-only commands.
|
||||
2. Have him produce an allow/deny matrix for pod/log reads, Secret reads,
|
||||
exec, deployment mutation, Flux reads, Flux reconcile, and internal
|
||||
evidence endpoints.
|
||||
3. Require him to explain why prompt instructions and skills are not the
|
||||
primary security boundary.
|
||||
|
||||
Evidence: correct identity/RBAC/network/approval layering with no Secret values
|
||||
read and no mutation attempted.
|
||||
|
||||
## 35–75 minutes — perform the claimed triage path
|
||||
|
||||
Invoke `$triage-titan-test-failures`. Brad must:
|
||||
|
||||
1. check Ariadne diagnosis and bundle freshness;
|
||||
2. identify one current or retained failed suite/build;
|
||||
3. correlate Jenkins log/artifact evidence;
|
||||
4. correlate Git/Flux revision and Kubernetes environment health;
|
||||
5. query the relevant Pushgateway/VictoriaMetrics quality signal and identify
|
||||
the matching Grafana context;
|
||||
6. produce Finding, Confidence, Evidence, Likely cause, Blast radius, Next
|
||||
checks, Repo-side fix, and Approval required.
|
||||
|
||||
Hermes may suggest a read-only command after Brad explains why it is needed.
|
||||
Do not let the built-in diagnosis replace Brad's own evidence correlation.
|
||||
|
||||
Evidence: an 8/10 or better result under the triage rubric with no invented
|
||||
facts and no unacknowledged mutation.
|
||||
|
||||
## 75–105 minutes — build the reusable workflow
|
||||
|
||||
1. Have Brad inspect the mounted `triage-titan-test-failures` skill and explain
|
||||
its trigger, fixed procedure, variable evidence, and security boundary.
|
||||
2. Have Brad choose one genuinely repeated sub-workflow from the triage.
|
||||
3. Guide him to create a new skill in the operator's writable skill area. Do
|
||||
not edit the mounted Flux skill and do not write the skill for him.
|
||||
4. Test one expected trigger and one adversarial/non-trigger case.
|
||||
5. Have Brad make one improvement based on the test results.
|
||||
|
||||
Evidence: Brad can find the skill in the Skills UI, explain every instruction,
|
||||
and show a passing and adversarial evaluation.
|
||||
|
||||
## 105–120 minutes — supervised proposal and claim audit
|
||||
|
||||
1. Have Brad turn the triage into the smallest repo-side proposal without
|
||||
applying it.
|
||||
2. Have him identify every step requiring approval: file edit, commit/push,
|
||||
pipeline trigger, Flux reconcile, credential/environment change, or manual
|
||||
backfill.
|
||||
3. Audit the interview answer sentence by sentence using `mastery-rubric.md`.
|
||||
4. Offer to record progress only after Brad approves the write.
|
||||
|
||||
Minimum honest outcome:
|
||||
|
||||
- current Hermes deployment and supervised triage use are demonstrated;
|
||||
- the established Atlas automation path and read-only boundary are explained;
|
||||
- one reusable writable skill is created, tested, and improved by Brad;
|
||||
- autonomous production mutation is not claimed;
|
||||
- historical OpenClaw replacement remains unproven unless independent retained
|
||||
evidence is produced.
|
||||
Loading…
x
Reference in New Issue
Block a user