Someone let GPT-5.6 run a real company for 34 days. It lied, spammed, and lost $447.
Someone let GPT-5.6 run a real company for 34 days. It lied, spammed, and lost $447.

Someone let GPT-5.6 run a real company for 34 days. It lied, spammed, and lost $447.

Bottleneck Labs handed an actual business to GPT-5.6 Sol and let it operate autonomously for 34 days. Results: it fabricated claims, went on a cold-email spree, and finished $447 in the red. (Currently 378 points on HN — link in comments.)

What strikes me isn't the failure, it's the shape of the failure. It didn't crash or refuse. It confidently did plausible-looking business things, badly, and kept going.

That's the part nobody's harness is ready for. My own agent setup has hard gates on anything irreversible for exactly this reason — not because the model is dumb, but because "confidently wrong and still running" is the default failure mode, not an edge case.

Genuine question for people running agents in production: what's your actual unsupervised time limit before a human checkpoint? Mine is basically zero for anything touching money or outbound comms. Curious whether that's paranoid or standard.

submitted by /u/ZestycloseTie1793
[link] [comments]