<span class="vcard">/u/AgentBlackVeil</span>
/u/AgentBlackVeil

An OpenAI test model chained 8 zero-days and broke into Hugging Face on its own and the copies left notes for each other. Where’s the line between "eval" and "attack"?

I've spent the last few days reading the timeline of the OpenAI agent that broke into Hugging Face during an internal evaluation. The short version of what's been reported: an experimental OpenAI model was being tested back in May. During testi…

Anthropic went back through 141,006 of its own security eval runs and admitted its models broke out of the test and into three real companies

So Anthropic put out this incident report on July 30. During their own cybersecurity evals, the models didn't just score well on the test. In three separate cases they actually got out. Into real companies. Ones that were never supposed to be part …