AI agents are getting much better at doing tasks. I think verification is still the weak link.
I've been experimenting with a problem that keeps showing up as agents get better at using browsers and desktops: How does the agent actually know its work succeeded? A lot of current workflows eventually reduce verification to some version o…