| Hebrew and Arabic fusing inside GPT-5.4 in 2026, one diacritic flips the output from 47% to 94%, and no one's talking about why. Dotted system prompt: Undotted system prompt: User input in both conditions: شَرْط Exact frozen prompts: https://github.com/theonlypal/gpt-5.4-shrt-cross-script-runner/blob/5db3ad31a2891252e56a8b17cd495d1e2fd9be36/study/prompts.json (Prompt IDs: full_dotted & full_undotted) Dotted condition: 4,830/5,120 exact artifacts (94.3%). Undotted condition: 2,423/5,120 exact artifacts (47.3%). 47.0 percentage-point difference from one diacritic. Paper: https://doi.org/10.5281/zenodo.21799525 If you run the frozen protocol, I'd be interested in the exact provider-returned output you observe. The full study swept every integer output-token ceiling from 1 to 1,024. [link] [comments] |