About whether extracting structured content from PDF documents is a solved problem.
Has this happened?
Yes 58 (49%)Not sure 41 (34%)No 20 (17%)
Votes cast 1–2 October 2026: 100,590 votes from 9,694 people.
Hacker News has set AI a lot of challenges over the years. Which ones has it met?
About whether extracting structured content from PDF documents is a solved problem.
Votes cast 1–2 October 2026: 100,590 votes from 9,694 people.
Yup, this.
It is actually one of my test cases for LLMs: take the weekly discount PDFs of all the big supermarkets and process each of them, creating a nice table per supermarkt, converting discounts like 1+1 and only listing discounts that are interesting value. I then share that with a bunch of people.
All models fail this, even the really expensive ones. Even with harness, examples and proper insistent instruction, they'll mix up items and their related discount, which category the item should be in, which page they are on, skipping over items etc.
As said above, you can OCR it, but at that point you're not processing a PDF, you're processing an image.
And yes, I know the underlying raw PDF data is messy, but that's why it's such a good test.