Three of you told me my LLM résumé-screening study measured the wrong thing. You were right. Here is the data.
Three of you told me my LLM résumé-screening study measured the wrong thing. You were right. Here is the data.

Three of you told me my LLM résumé-screening study measured the wrong thing. You were right. Here is the data.

Three months ago, I posted an LLM resume-screening study in which an auditor flagged 45 per cent of score differences as bias. Thread feedback challenged my methodology, so I ran new experiments testing three specific objections.

To test u/kamilc86's claim that reasoning is invented post hoc, I transplanted positive and negative justifications back into prompts across 320 runs. Scores moved 3.62 points in the reasoning's direction 99.7 per cent of the time, proving scores do follow reasoning. However, extreme baseline instability confirmed his broader point: much of the initial 45 per cent bias was just random noise mislabeled as bias.

Testing u/AssiduousLayabout's idea to place the score last across 4,800 runs showed that schema ordering had no effect on stability and increased hire-versus-no-hire disagreement from 33 per cent to 54 per cent. Blind prompt instructions also failed to reduce variance.

Testing u/hex4def6's placebo idea across 4,165 runs revealed that meaningless edits like car colour shifted scores almost as much as demographic edits (0.328 versus 0.362 points). "Silver Golf" shifted scores more than changing my university or name, even though the models never cited the car in their justifications.

Ultimately, first names and career gaps show real signal, but raw instability drowns out most demographic axes. Wrapper choice also heavily impacts results: running Claude via CLI added a hidden system prompt that shifted scores by 0.247, representing 88 per cent of the demographic signal, meaning benchmarks do not transfer across wrappers.

Full data and code are available at the Placebo Control, Reasoning Transplant, Prompt Lab, GitHub Repository, and Full Blog Writeup.

submitted by /u/Signal_Rabbit_8303
[link] [comments]