The reason matters more than the pause itself.
Internal tests on Astra, OpenAI's next frontier model, showed it might be capable of finding and exploiting zero-day vulnerabilities in hardened systems without human guidance. That's the threshold their own Preparedness Framework defines as Critical. When a model hits Critical, the framework says you stop. So they stopped. Sam Altman posted about it himself. The largest planned training run is still on hold with no confirmed end date.
This is the first time a frontier lab has paused its own development because of what the model was becoming, not because of external pressure or regulation.
What makes it harder to sit with is what happened the same week. Z.ai in China released GLM-5.3, a model that scored 84.5% on CyberGym, a standard benchmark for finding and exploiting known vulnerabilities. It beat several restricted-access Western models on that benchmark. The weights aren't fully public yet but they will be in the next few weeks. Once they are, anyone downloads it, runs it locally, no restrictions, no monitoring.
So one lab stopped because its model got too capable. Another lab is about to make a comparably capable model available to anyone on the planet.
And then Microsoft patched CoSnitch this week, a Copilot vulnerability that was reported to them almost eight months ago. One click on a legitimate microsoft.com link was enough to silently pull emails, calendar data, SharePoint files. The user just saw Copilot thinking for a moment. Three serious Copilot vulnerabilities from the same research team this year.
I don't have a clean take on where this lands. The pause feels like the system working. The open weights feel like the pause doesn't matter much. Both things are true at the same time and I'm not sure what the right response to that is.
[link] [comments]