About whether needle-in-haystack tests measure real understanding of long documents.
Has this happened?
Yes 67 (57%)Not sure 37 (31%)No 14 (12%)
Votes cast 1–2 October 2026: 100,590 votes from 9,694 people.
Hacker News has set AI a lot of challenges over the years. Which ones has it met?
About whether needle-in-haystack tests measure real understanding of long documents.
Votes cast 1–2 October 2026: 100,590 votes from 9,694 people.
There is no understanding, it can't do this.
GPT4o still can't do the intersection of two different ideas that are not in the training set. It can't even produce random variations on the intersection of two different ideas.
Further though, we shouldn't expect the model to do this. It is not fair to the model and its actual usefulness and how amazing what the models can do with zero understanding. To believe the model understands is to fool yourself.