docs: state what each demo proves and record the two selection defects

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
jenkins 2026-08-06 01:39:51 -03:00
parent f44ed2f953
commit d0753c5eb1
2 changed files with 80 additions and 4 deletions

View File

@ -1004,8 +1004,31 @@ The handoff must state plainly whether the full Hermes Agent or the Ollama fallb
| Action registry: `repair_demo_fixture`, `retry_transient_infra` | live |
| Bounded source-patch proposal → branch → pull request | live (demo repo) |
| Human-required metric, vmalert rule, Alertmanager | live |
| Gitea issue filing for human-required incidents | built, enabling |
| Candidate-file selection for real service repos | built, enabling |
| Gitea issue filing for human-required incidents | live, proven on 3 real repos |
| Candidate-file selection for real service repos | live, PR not yet observed |
## What the fixture and code demos actually prove
Worth stating plainly before demonstrating either one, because the two demos
are not equally representative of the real-service path.
The **fixture demo** is end-to-end real. Ariadne detects a genuine red build,
Hermes diagnoses it from real evidence, the deterministic gates authorize, and
Ariadne executes and verifies via a rebuild. Nothing about it is staged beyond
the fixture itself being deliberately broken.
The **code demo** is real in its hard parts and staged in one: the repository
is real, the failing tests are real, Hermes genuinely locates the defect and
writes an anchored patch, the validator genuinely checks it, and the pull
request is genuinely opened and awaits review. What is staged is *which file*
Hermes is handed - `ARIADNE_HERMES_CODE_CANDIDATE_PATH=src/discount.py` pins
it. File selection, the step that has to work on a real service, is the one
step the demo does not exercise.
On real repositories there is no pin, and selection has to work from evidence
alone. Two defects blocked that and were fixed on 2026-08-06 (see below):
successful tool output crowding the real failure out of the evidence budget,
and the patcher being unable to reach a source file the console never names.
## Measured results
@ -1049,6 +1072,21 @@ Real homegrown service diagnoses (no action allowlisted, all escalated):
job went green, the alert self-cleared through supersession.
- `soteria` returned `unknown` at 0.58 — below the 0.85 threshold and with
no matching action, so two independent gates blocked it.
- Two defects were found by deliberately seeding a one-line break in a real
repository (`ariadne`) rather than by reasoning about the code. First, the
console evidence budget was filled entirely by *successful* tool output:
build 404 retained six regions spanning lines 259-476 - Trivy setup, Sonar
banners, and three lines whose only failure signal was the word `coverage` -
and dropped the pytest failure at line 1487. Markers are now ranked by
whether they assert a failure or merely name a tool that ran. Second, the
defective file `ariadne/utils/errors.py` appeared *zero times* in the entire
174 KB console, because a pytest assertion that fails inside a test names
only the test; candidate selection returned the wrong file. Ariadne now
reads the failing test and follows its imports back to the module under
test, widening what may be read without widening what may be written - the
patch validator still gates writes on `allowed_path_prefixes`, so a failing
test can be read for context and never patched, which also forecloses the
classic bad fix of silencing a test instead of repairing the code.
- `bstein-dev-home/446` was classified `known_demo_fixture_failure` at 0.99,
even though build 445 of the same job had been classified correctly as
`sonarqube_node_runtime_missing`. The evidence gate refused it again, so

View File

@ -1004,8 +1004,31 @@ The handoff must state plainly whether the full Hermes Agent or the Ollama fallb
| Action registry: `repair_demo_fixture`, `retry_transient_infra` | live |
| Bounded source-patch proposal → branch → pull request | live (demo repo) |
| Human-required metric, vmalert rule, Alertmanager | live |
| Gitea issue filing for human-required incidents | built, enabling |
| Candidate-file selection for real service repos | built, enabling |
| Gitea issue filing for human-required incidents | live, proven on 3 real repos |
| Candidate-file selection for real service repos | live, PR not yet observed |
## What the fixture and code demos actually prove
Worth stating plainly before demonstrating either one, because the two demos
are not equally representative of the real-service path.
The **fixture demo** is end-to-end real. Ariadne detects a genuine red build,
Hermes diagnoses it from real evidence, the deterministic gates authorize, and
Ariadne executes and verifies via a rebuild. Nothing about it is staged beyond
the fixture itself being deliberately broken.
The **code demo** is real in its hard parts and staged in one: the repository
is real, the failing tests are real, Hermes genuinely locates the defect and
writes an anchored patch, the validator genuinely checks it, and the pull
request is genuinely opened and awaits review. What is staged is *which file*
Hermes is handed - `ARIADNE_HERMES_CODE_CANDIDATE_PATH=src/discount.py` pins
it. File selection, the step that has to work on a real service, is the one
step the demo does not exercise.
On real repositories there is no pin, and selection has to work from evidence
alone. Two defects blocked that and were fixed on 2026-08-06 (see below):
successful tool output crowding the real failure out of the evidence budget,
and the patcher being unable to reach a source file the console never names.
## Measured results
@ -1049,6 +1072,21 @@ Real homegrown service diagnoses (no action allowlisted, all escalated):
job went green, the alert self-cleared through supersession.
- `soteria` returned `unknown` at 0.58 — below the 0.85 threshold and with
no matching action, so two independent gates blocked it.
- Two defects were found by deliberately seeding a one-line break in a real
repository (`ariadne`) rather than by reasoning about the code. First, the
console evidence budget was filled entirely by *successful* tool output:
build 404 retained six regions spanning lines 259-476 - Trivy setup, Sonar
banners, and three lines whose only failure signal was the word `coverage` -
and dropped the pytest failure at line 1487. Markers are now ranked by
whether they assert a failure or merely name a tool that ran. Second, the
defective file `ariadne/utils/errors.py` appeared *zero times* in the entire
174 KB console, because a pytest assertion that fails inside a test names
only the test; candidate selection returned the wrong file. Ariadne now
reads the failing test and follows its imports back to the module under
test, widening what may be read without widening what may be written - the
patch validator still gates writes on `allowed_path_prefixes`, so a failing
test can be read for context and never patched, which also forecloses the
classic bad fix of silencing a test instead of repairing the code.
- `bstein-dev-home/446` was classified `known_demo_fixture_failure` at 0.99,
even though build 445 of the same job had been classified correctly as
`sonarqube_node_runtime_missing`. The evidence gate refused it again, so