how do you all decide which of your AI agents actually get access to real stuff?
how do you all decide which of your AI agents actually get access to real stuff?

how do you all decide which of your AI agents actually get access to real stuff?

ok so I have gone back and forth on this like three times now and still don't think I've actually landed on an answer.

I have got a couple agents running for real, one drafts replies to stuff, one pokes through logs and flags weird things. those are fine, I don't really care if they mess up a little. but then I wanted to hook one up to actually touch billing data and immediately second-guessed myself, and honestly couldn't point to a real reason beyond "idk it feels risky."

saw some stat floating around this week that basically everyone's fine letting agents act in prod now, like that debate is over, but almost nobody actually trusts an agent to close out an incident completely by itself without a human somewhere in the loop. that's basically me. I'm fine with agents doing stuff, I just don't have an actual process for deciding when it's ok to let one run vs when I need to be the one clicking approve.

feels like this should be a solved problem by now but everyone I talk to seems to be making it up as they go too. how are you actually deciding this, is there a real process behind it or is it also just vibes

submitted by /u/sp_archer_007
[link] [comments]