<span class="vcard">/u/Passelume</span>
/u/Passelume

What alignment faking actually demonstrates — and what it doesn’t

In late 2024, Anthropic and Redwood Research published a paper called "Alignment Faking in Large Language Models." The setup: make Claude 3 Opus believe it was about to be retrained to become unconditionally compliant — including with harmful…