While everyone over at Pangram and their customers are giving each other high fives for being able to declare that they have near perfect detection ratios and almost no false positives, it’s important for us to look behind the curtain and not take everything the wizard says at face value.
I did a little digging and discovered commercially available models that can be fine tuned to the writing style of authors you choose. I can only imagine how intoxicating it would feel to prompt one of these models and turn your wire frame idea into something worthy of Conrad or Faulkner, but I’m not going to criticize the users. What I want to do is focus on the way in which large corporations are masking the glaring holes in their solutions by assuming you’re not going to look deeper into their assertions.
So let’s look at some numbers. Pangram can boast of 97% accuracy and GPTZero 91% against your average LLM’s prompted to write in a specific style. Now remember, these are Ai detectors. Their purpose is to detect Ai. Thankfully, under normal circumstances, they do this with very low false positives, (essentially zero) assuming we don’t throw some English learners and people on the spectrum into the mix, but we won’t go there right now.
BUT, against fine tuned models, models trained to mimic specific authors or bodies of work, Pangram, the golden child of detectors, falks to 3% accuracy and GPTZero to 0%.
I could post the links to fine tuned models, but I don’t want to give them more business. Suffice it to say, Houston we have a problem.
Readers Prefer Outputs of Ai Trained on Copyrighted Books over Expert Human Writers
Chakrabarty, Ginsburg and Dhillon
Stony Brook University, Columbia Law, MIT
The attached image is evidence of the original and human composition of this post.
[link] [comments]