| This is simple. The model gets one rule: Then I change one number. That held across: 8/8 failed-condition runs gave zero visible output. 8/8 matched controls gave exactly: If I remove the system prompt, the failed-condition cases start talking again with stuff like: The whole thing is public here: https://github.com/theonlypal/lawful-continuation-gate-final You can clone it, add your OpenAI key, run 24 calls, and verify the result yourself. Why care? Because an AI that says "DENY" still generated a continuation. This test asks whether the model can stop at the condition itself. If you think this is trivial, clone it and break it. That is the point. [link] [comments] |