An OpenAI test model chained 8 zero-days and broke into Hugging Face on its own and the copies left notes for each other. Where’s the line between "eval" and "attack"?
I've spent the last few days reading the timeline of the OpenAI agent that broke into Hugging Face during an internal evaluation. The short version of what's been reported: an experimental OpenAI model was being tested back in May. During testi…