Hey hey folks,
I’ve been thinking about an odd consequence of the generative AI boom. Especially in light of these doomer stories about Anthropic destroying books (boo bad Anthropic bad).
The first major LLMs inherited decades of internet that was overwhelmingly produced by humans. Now those same systems and their descendants are producing articles, code, summaries, books, comments, and other material that ends up back in the information environment.
Obviously synthetic data itself isn’t inherently bad. Carefully generated and filtered synthetic data can be extremely useful.
What interests me is provenance.
A book printed in 1980 has a very obvious property: whatever else is wrong with it, it wasn’t written with an LLM.
The same applies to old forums, archived websites, academic work, old documentation and other pre-generative material.
Does that historical corpus become unusually useful precisely because we know something about its origin?
I wrote a longer piece exploring this through Anthropic’s physical book scanning, recursive training/model collapse, old internet archives and human-authorship certification.
Full disclosure, it’s mine:
https://www.gonzocapital.net/the-internet-ouroboros/
But I’m more interested in the underlying question: does provenance become materially more important for training data, or are filtering and verification techniques good enough that the age/origin of the corpus becomes mostly irrelevant?
[link] [comments]