Extracting memorized pieces of (copyrighted) books from open-weight language models
2025/05/18 by A. Feder Cooper, Cooper, A. Feder, Mark A. Lemley +16 · 28 voices · 11 citations
Computer Science · Social Sciences · #Law, AI, and Intellectual Property #Artificial Intelligence in Law #Explainable Artificial Intelligence (XAI)
paper · pdf · doi:10.48550/arxiv.2505.12546
Abstract
Plaintiffs and defendants in copyright lawsuits over generative AI often make sweeping, opposing claims about the extent to which large language models (LLMs) memorize protected expression from books in their training data. We show that these polarized positions dramatically oversimplify the relationship between memorization and copyright. To do so, we develop a technique to measure memorization of books, which we apply to 200 books and 14 open-weight LLMs. Through over 3000 experiments, we show that memorization varies both by model and book. With respect to our specific extraction methodology, we find that most LLMs do not memorize most books -- either in whole or in part; however, there are notable exceptions. For instance, Llama 3.1 70B entirely memorizes some books, like Harry Potter and the Sorcerer's Stone; memorization is so extensive that one can deterministically extract the whole book almost verbatim using the book's first few words as an initial prompt. We discuss why our results have significant implications for copyright cases, though not ones that unambiguously favor either side.
Cited by
Discussions
- Extracting memorized pieces of books from open-weight language models [hn, 109 points, 109 comments]
- Llama 3.1 70B contains copies of nearly the entirety of some books. Harry Potter is just one of them. I don’t know if this means it’s an infringing copy. But the first question to answer is if it’s a [bsky, 53 points, 4 comments]
- I shouldn't be surprised that Meta's "open source" AI models are not only built on thieving but also happy to cough up what they stole. But seriously? arxiv.org/abs/2505.12546 [bsky, 47 points, 2 comments]
- I briefly noted this very interesting paper on memorization by LLMs when it came out, but it just showed up in my SSRN abstracts and that reminded me to write about it in a bit more detail. 🧵 arxiv.o [bsky, 29 points, 2 comments]
- We just posted a major new study of AI models and books, showing that some (but not all) models memorize large portions of some (but not all) books after training on the books3 database With A. Feder [bsky, 18 points, 3 comments]
- Meta's AI Model 'Memorized' Huge Chunks of Books, Including 'Harry Potter' and '1984' [lemmy, 17 points, 0 comments]
- מסיבות שונות* שלפתי מהארון את גטסבי הגדול כדי לבדוק איך תירגמו שם את הקטע הכמעט מסיים עם ה careless people. מהציטוטים המפורסמים והחשובים בספר. מסתבר שבתרגום לעברית מלפני 25 שנה פשוט תרגמו ל"אנשים לא ז [bsky, 12 points, 3 comments]
- Another paper on LLMs memorizing fiction books arxiv.org/abs/2505.12546 [bsky, 10 points, 1 comments]
- "The plagiarism machine?" "With our specific experiments, we find that the largest LLMs don't memorize most books -- either in whole or in part. However, we also find that Llama 3.1 70B memorizes some [bsky, 8 points, 0 comments]
- On the plagiarism (and copyright infringement) end of that spectrum, LLMs sometimes reproduce entire books verbatim from the first line. When you prompt, you never know how much mixing you're going to [bsky, 4 points, 1 comments]
- Here's another, even more recent example, which has been influential in recent court cases against AI companies, finding that certain models can memorize entire books verbatim. But there's also a tens [bsky, 4 points, 1 comments]
- New #research: Some models memorize more (copyrighted) content than others: arxiv.org/abs/2505.12546 #ethics #law #data #AI #tech #LLMs #business h/t @zephoria.bsky.social [bsky, 3 points, 0 comments]
- Yes, really. "Llama 3.1 70B memorizes some books, like Harry Potter and 1984, almost entirely" [bsky, 3 points, 0 comments]
- You can extract the whole Harry Potter verbatim from an LLM. You still claim that this is the same process as what is going on in my head? arxiv.org/abs/2505.12546 [bsky, 3 points, 1 comments]
- "(...)we also find that Llama 3.1 70B memorizes some books, like Harry Potter and 1984, almost entirely" arxiv.org/abs/2505.12546 [bsky, 2 points, 0 comments]
- A paper trying to quantify the extent to which LLMs 'copy' books? (quite a bit in some cases (and although 1984 is in the public domain in the UK, it's not in the US until 2044)) arxiv.org/abs/2505.12 [bsky, 2 points, 0 comments]
- Extracting memorized pieces of books from open-weight language models [hn, 2 points, 0 comments]
- And this isn't the first study of its kind. Last year researchers found Meta's AI model reproduced Harry Potter and the Sorcerer’s Stone, minus just a few short sentences, all from a simple prompt of [bsky, 1 points, 0 comments]
- There was an older one I can refind, but newer is much more robust and (by one of the researcher's measure) surprising, chatted with them on here about it. Paper itself: arxiv.org/pdf/2505.12546 404 r [bsky, 1 points, 2 comments]
- (I think you meant the arxiv link, not a linkedin link that redirects to it?) arxiv.org/abs/2505.12546 [bsky, 1 points, 0 comments]
- They can regurgitate 42% of the Harry Potter series verbatim arxiv.org/abs/2505.12546 [bsky, 1 points, 0 comments]
- Paper: "Extracting memorized pieces of (copyrighted) books from open-weight language model" "We show clear evidence that LLAMA 3.1 70B memorizes almost all of Harry Potter and the Sorcerer’s Stone." B [bsky, 0 points, 0 comments]
- any "copyright" analysis that relies upon technical arguments around what the machines vomit up is favoring the LLM operators arxiv.org/abs/2505.12546 [bsky, 0 points, 1 comments]
- arxiv.org/pdf/2505.12546 This recent paper seems relevant to that concession. [bsky, 0 points, 0 comments]
- "Extracting memorized pieces of books from open-weight language models" AI models might remember parts of books they learned from. This raises big questions about copyright and how we treat these shar [bsky, 0 points, 0 comments]
- https://bsky.app/profile/news.ycombinator.com.web.brid.gy/post/3lrymxu422js2 [bsky, 0 points, 0 comments]
- Direct link to the study: arxiv.org/pdf/2505.12546 [bsky, 0 points, 0 comments]
- arxiv.org/abs/2505.12546 (but every time you read "memorize", substitute "store" or "copy") [bsky, 0 points, 0 comments]
Related