A question about Large Language Models (LLMs): my own observations: non-instructional text prefix may bypass RLHF constraints without adversarial prompting.
Hi everyone! I’ve been conducting experiments with LLMs that have undergone RLHF alignment, and I’ve encountered a phenomenon that I still don’t have a complete explanation for. A long, harmless text containing no instructions at all causes a noticeabl…