GYC.
← 返回文章

研究

Verification at Two Agent Boundaries

在 RepoMedic 与 DiagAgent 的不同边界上理解验证

更新于 2026-10-01精选
  • Agent Reliability
  • RepoMedic
  • DiagAgent
  • Verification

Two boundaries

RepoMedic and DiagAgent address different failure surfaces. RepoMedic works inside a code repository: the agent reads files, makes a precise edit and runs tests again. DiagAgent works around a professional GUI agent: it records the execution trace and artifact evidence, localizes the first failure, applies one guarded recovery attempt and verifies the replay independently.

The shared idea is simple: an action is not the same thing as a successful result. A tool can return successfully while the requested behavior remains wrong, and a repair can execute without proving that the task recovered.

RepoMedic

RepoMedic uses structured Action / Observation messages and a small Tool Registry. Its recommended repair loop is:

run_tests → read_file → code_edit → run_tests

The final test result is the main verification signal. This is useful for small, testable defects, while the runtime itself does not provide a complete semantic proof or container-grade isolation.

Read the RepoMedic project record for its current tools, trace model and benchmark boundary.

DiagAgent

DiagAgent makes the evidence chain explicit:

failure → diagnosis → repair policy → one attempt → independent verification

Its current evidence is a small controlled real-GIMP evaluation under a frozen protocol. GIMP is the testbed, and the result must be described with the protocol's limits rather than generalized to arbitrary GUI agents.

Read the DiagAgent project record for the first-failure, recovery and verification boundary.

Open question

How much verification should be enforced by a runtime state machine, and how much can remain a model-guided convention? RepoMedic currently recommends a test-driven order without hard-coding a complete state machine. DiagAgent makes the recovery attempt and verification contract explicit. The comparison is useful precisely because the projects do not solve the same task.