Turns out the world’s top AI models can be nudged into spitting out near-verbatim chunks of bestselling books, and the copyright row is getting messier.
A run of recent studies suggests that large language models from OpenAI, Google, Meta, Anthropic, and xAI memorise far more of their training data than the industry has been letting on.
Imperial College London, professor of applied mathematics and computer science, Yves-Alexandre de Montjoye said, “There’s growing evidence that memorisation is a bigger thing than previously believed,”
That “memorisation” talent could land like a brick in the AI industry’s court fights, because it dents the line that models learn patterns rather than storing copyrighted text.
AI firms have insisted that memorisation does not occur, and in a 2023 letter to the US Copyright Office, Google said, “there is no copy of the training data, whether text, images, or other formats, present in the model itself.”
They have leaned hard on “fair use”, claiming training on copyrighted books transforms the source into something meaningfully new, even if the authors are not exactly cheering.
A study published last month said Stanford and Yale researchers used strategic prompting to pull thousands of words from 13 books using models from OpenAI, Google, Anthropic, and xAI.
By asking models to complete sentences, Gemini 2.5 regurgitated 76.8 per cent of Harry Potter and the Philosopher’s Stone with high accuracy, while Grok 3 generated 70.3 per cent.
The researchers also extracted almost the entirety of a novel “near-verbatim” from Anthropic’s Claude 3.7 Sonnet by jailbreaking it, pushing the model to ignore its safeguards.
It builds on last year’s work, claiming “open” models such as Meta’s Llama can memorise huge slabs of particular books from their training data.
Experts had been unsure whether “closed” models with heavier guardrails would be as prone to large-scale memorisation, which now looks a bit optimistic.
Yale University, researcher, A. Feder Cooper said, “It was a surprise that they could memorise entire texts” despite guardrails,
Researchers still have not pinned down why LLMs memorise some training data, or how much of it can be coaxed back out in normal use rather than lab-grade prompting.
That matters outside publishing, too, because if training data leaks in healthcare or education, privacy and confidentiality risks stop being theoretical.
The question of memorisation has already been a live wire in court, including a US ruling last year that Anthropic’s training on some copyrighted content could be fair use because it was deemed “transformative”.
In the same case, the court said storing pirated works was “inherently, irredeemably infringing”, pushing Anthropic into a $1.5bn settlement.
In a ruling from November last year, German courts found OpenAI infringed copyright because its model had memorised song lyrics, in a case brought by GEMA.
Husch Blackwell, partner, Rudy Telscher, said reproducing an entire book without jailbreaking is “clearly a copyright violation.” But “it’s a matter of whether this is happening enough that [AI models] could be vicariously liable for the infringement,” he added.
Anthropic said the jailbreaking method used by the Stanford and Yale team was impractical for normal users and would take more effort than simply buying the content.
It said its model does not store copies of specific datasets; instead, it learns patterns and relationships between words and strings in the training data.







