I am building **Flows**, an execution and verification layer for software-building agents.
The core rule: an agent should not convert “I think I finished” into “verified complete” without supporting proof.
A Flows project can contain implementation steps, checks, repair instructions, review, and release conditions.
An independent agent used one plan to build a real multi-module application with 59/59 automated checks passing.
The target metric is: **unsupported required claims shipped = 0 on real traffic.**
Should evidence enforcement live in the agent harness, repository CI, app platform, or a cross-agent workspace?
[link] [comments]