About whether LLMs memorize copyrighted books, in a story on relicensing via AI rewrites.
Has this happened?
Yes 22 (16%)Not sure 64 (47%)No 50 (37%)
Votes cast 1–2 October 2026: 100,590 votes from 9,694 people.
Hacker News has set AI a lot of challenges over the years. Which ones has it met?
About whether LLMs memorize copyrighted books, in a story on relicensing via AI rewrites.
Votes cast 1–2 October 2026: 100,590 votes from 9,694 people.
Are you referring to this?
https://osyuksel.github.io/blog/reconstructing-moby-dick-llm...
I see a test where one model managed to 85% reproduce a paragraph given 3 input paragraphs under 50% of the time.
So it can't even produce 1 paragraph given 3 as input, and it can't even get close half the time.
"Contains Moby Dick" would be something like you give it the first paragraph and it produces the rest of the book. What we have here instead is a statistical model that when given passages can do an okay job at predicting a sentence or two, but otherwise quickly diverges.