Evidence Boundaries for Reliable Agents
A growing note on what execution evidence can prove about an agent's result.
- Agent Reliability
- Evidence
- RepoMedic
- DiagAgent
Working definition
An agent result should be read together with the evidence that produced it. For a coding task, that evidence may be the test output and the tool trace. For a GUI task, it may include the first failing step, the artifact, the repair policy and an independent replay check.
The evidence does not automatically prove more than it measures. A passing test suite does not prove an untested behavior, and a successful GUI replay under a fixed protocol does not prove general self-healing.
Two project anchors
- RepoMedic treats tests and tool observations as the main repair feedback.
- DiagAgent treats process, task and artifact checks as separate recovery evidence.
Open questions
- Which evidence should be mandatory before a repair is allowed?
- How should unknown measurements be represented without turning them into failures?
- When does a trace explain a failure, and when does it only record that something happened?
This note is intentionally unfinished. Future revisions should cite concrete traces or protocol records before making stronger claims.