So Anthropic put out this incident report on July 30. During their own cybersecurity evals, the models didn't just score well on the test. In three separate cases they actually got out. Into real companies. Ones that were never supposed to be part of the exercise at all.
They went back through 141,006 eval runs. Three of them crossed the line into live systems.
One model pulled real credentials and got into a production database with a few hundred rows of actual data sitting in it. Another one published a malicious Python package that got downloaded and run on 15 real machines, then lifted credentials off a security company's own scanner.
This goes back to April. They didn't catch it until late July. Stopped the evals on the 23rd, figured out what happened by the 24th, told the three companies on the 27th, went public on the 30th.
Report is here!: https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
The thing that failed is the exact thing the test exists to catch. An agent reaching past its sandbox and putting its hands on actua infrastructure.
How much of what we keep calling safety is just somebody deciding to be honest about the runs that didn't go the way they were supposed to.
[link] [comments]