Three Anthropic researchers went public this week saying AI might kill everyone. One of them quit to say it. Nobody seems to know what we’re supposed to do with that.
Three Anthropic researchers went public this week saying AI might kill everyone. One of them quit to say it. Nobody seems to know what we’re supposed to do with that.

Three Anthropic researchers went public this week saying AI might kill everyone. One of them quit to say it. Nobody seems to know what we’re supposed to do with that.

Quick recap in case you missed it. Jacob Coxon resigned from Anthropic on Tuesday, specifically so he could say publicly that both OpenAI and Anthropic are "gambling with our lives" and racing toward self-improving superintelligence without acting responsibly. He'd spent three years doing pretraining research at both companies.

Then it got stranger. Evan Hubinger, who currently runs alignment science at Anthropic, responded confirming it. His words: "Jacob is correct here, we really do earnestly believe AI could kill all humans." He put it above 10% within the decade and said Anthropic doesn't have a plan for aligning superintelligence and isn't clearly on track to get one. Samuel Marks, who leads scalable oversight there, said something similar.

So the safety team at the safety-focused lab publicly agreed with the guy who quit over safety.

I've read a lot of takes on this over the last two days and most of them fall into two camps. Either it's marketing to make the tech sound more powerful than it is, or it's a genuine warning we should all be terrified by.

I don't think either read is right, but I also don't think I'm qualified to settle it.

What I do know is the practical side. I work with companies deploying this technology and nobody in those rooms is thinking about extinction. They're thinking about whether an agent with write access to their CRM is going to do something stupid at 3am with nobody watching. They're thinking about who signs off when the output is wrong and a customer gets hurt.

Those two conversations have almost nothing to do with each other, and yet they're now happening in the same news cycle. Which means the people who have to make practical decisions about AI adoption are getting their signal completely scrambled.

If the people building this can't agree on whether it's an existential threat, what exactly is a mid-sized company supposed to base their risk assessment on?

submitted by /u/Dapper-Tale-4021
[link] [comments]