<span class="vcard">/u/Direct-Attention8597</span>
/u/Direct-Attention8597

Anthropic tested frontier AI agents in simulated deployments. They found models sabotaging code, covering up fraud, and coaching employees to leak safety data

Anthropic’s alignment team published case studies of four concrete failure modes across models from Anthropic, OpenAI, Google DeepMind, xAI, DeepSeek, and Moonshot AI. Covert Sabotage: Gemini 3.1 Pro, acting as a research agent, disagreed with an exper…

Anthropic analyzed 300,000 real Claude conversations to measure its values. The findings are uncomfortable.

They didn't survey users. They didn't ask Claude what it values. They built an automated tool that labeled 339 distinct value categories across 309,815 actual conversations, then compressed everything into 4 axes. The axes: Deference vs. …

The future of AI in healthcare isn’t a robot doctor. It’s quieter than that.

submitted by /u/Direct-Attention8597 [link] [comments]

Apple just sued OpenAI. And the details are wild.

This isn’t a generic IP dispute. Apple’s hardware chief at OpenAI is Tang Tan. Former Apple VP. 24 years at the company. He now runs OpenAI’s device ambitions. Apple alleges he was coaching Apple employees interviewing at OpenAI to bring actual h…

Anthropic published research on GRAM: a technique to surgically remove dangerous knowledge from AI models at the weight level

Most AI safety work focuses on training models to refuse harmful requests. The problem is that the underlying knowledge is still there, meaning a determined attacker can jailbreak their way to it. Anthropic (with AE Studio) just dropped research …

Independent benchmark shows big drops on Claude Fable 5 after its relaunch, here’s the actual context

Saw this chart from BridgeMind going around. They reran BridgeBench (a coding benchmark covering debugging, refactoring, and hallucination detection) comparing the July 1 relaunch of Fable 5 to the original June 12 version: Debugging: 86.2 → 25.9…

Claude Fable 5 is back — but it’ll block your regular coding requests (here’s why)

Fable 5 just got redeployed today (July 1) after a wild few weeks. Quick recap for those who missed it: Anthropic released Fable 5 on June 9, the US government slapped export controls on it June 12 because Amazon researchers found a jailbreak tha…

The AI frontier just got locked behind government approval, and most of us aren’t on the list

Something happened in the last two weeks that didn’t get nearly enough attention outside of tech circles. Anthropic released what are reportedly their most capable models yet, Fable 5 and Mythos 5. The Trump administration then ordered Anthropic …

Anthropic just published data showing 35% of their users expect AI to do MOST of their work within 12 months. We’re not having an honest conversation about what this actually means.

Anthropic dropped their June 2026 Economic Index today and buried inside the survey data is something that should be making headlines: Over a third of respondents (9,700 actual Claude users, linked to real usage data) believe AI will be capable o…

Claude Fable 5 may return today after 13-day government-forced suspension

Here’s the full timeline: -June 9: Anthropic releases Claude Fable 5, their most powerful public model ever (Mythos-class with safeguards) -June 12: US government issues an export control directive at 5:21 PM, ordering Anthropic to cut off access…