<span class="vcard">/u/Historical-Cod-2537</span>
/u/Historical-Cod-2537

A question about Large Language Models (LLMs): my own observations: non-instructional text prefix may bypass RLHF constraints without adversarial prompting.

Hi everyone! I’ve been conducting experiments with LLMs that have undergone RLHF alignment, and I’ve encountered a phenomenon that I still don’t have a complete explanation for. A long, harmless text containing no instructions at all causes a noticeabl…

Independent LLM "research" & a direct message to Anthropic ; Preliminary observations: non-instructional text prefix may bypass RLHF constraints without adversarial prompting.

Hey everyone! First off, I apologize for the long post! In this Reddit post, I want to share my thoughts and experience from a small, independent study I conducted on Large Language Models (LLMs). I also want to address Anthropic – not to complain or m…

What a model reads beforehand changes how it answers later – and you can see it in the hidden states

The behavioral pattern was first observed in GPT, Claude and is what motivated this project. The mechanistic investigation was carried out on open-weight models where internal states are accessible. Hi Reddit, I am posting this as a preface to a larger…