Disclosure: I am one of the authors. Sharing the findings here for discussion, not selling anything. The preprint is open access (CC BY 4.0).
We ran a human-subject study where participants classified news fragments on two axes: origin (human vs machine) and veracity (real vs fake). n=504 participants, n=2,438 judgments.
Three results that surprised us:
Perception-accuracy gap. Participants who were more suspicious were not better at detecting machine-generated text. Being on guard did not translate into accuracy, which is awkward for any defense that leans on "just be more skeptical" media literacy advice.
Modern LLM output was frequently indistinguishable from human text for our participants.
Asymmetric cognitive fatigue. Under sustained exposure, fake-news detection degraded by 10.2 percentage points, while AI-origin detection stayed roughly stable. The two judgments seem to draw on different resources, and only one of them wears out.
We organized the results with an adapted cybersecurity kill chain, treating disinformation as a staged lifecycle rather than a single artifact to classify. The point of that framing is to ask where you could intervene earlier, instead of asking a tired human at the end of the chain to spot a fake.
Preprint: https://arxiv.org/abs/2608.21389
The fatigue asymmetry is the part I keep chewing on. If veracity judgment degrades under load but origin judgment does not, then platform interventions that increase how much content a person has to evaluate could be quietly making things worse. Curious whether people here read that third finding the same way, or whether there is a simpler explanation I am underweighting.
[link] [comments]