Hi everyone! I’ve been conducting experiments with LLMs that have undergone RLHF alignment, and I’ve encountered a phenomenon that I still don’t have a complete explanation for. A long, harmless text containing no instructions at all causes a noticeable and persistent shift in activations in the middle and later layers - and this effectively disables the model’s safety mechanisms, without any explicit commands being given. If I’m not mistaken, Yann LeCun has said that for a model to predict text well, it needs to understand the reality underlying that text. My question is: could this activation drift be evidence that the model’s “world” is actually a collection of regions formed during training, and that context can move the model between them, completely bypassing its safety settings? If anyone is interested, I have the relevant metrics and reproducible tests. I’m familiar with the Anthropic research; it’s not quite about what I’m talking about here.
[link] [comments]