Goalposts

Hacker News has set AI a lot of challenges over the years. Which ones has it met?

2024 Februaryxnorswap

About why asking LLMs to generate prime numbers is a useful test.

The value is showing how confidently is presents incorrect results.

Especially the lack of nuance or uncertainty in the language they use.

They extremely confidently present the incorrect information, and prime generation is interesting because it's information that isn't easy to spot as obviously incorrect to the user while being information that's possible to determine is wrong at small numbers and difficult to verify for large numbers.

It's my favourite test because it's a good demonstration of the lack of nuance or uncertainty in LLMs. They have no sense of how wrong the information they're giving out might be.

If they could give confidence intervals for any information then they could provide the context by how likely they think they might be correct, but they actually double-down on their incorrectness instead.

An LLM gives calibrated confidence for its answers, such as lists of primes, instead of confidently presenting wrong information.

Has this happened?

Yes 30 (29%)Not sure 29 (28%)No 44 (43%)

Votes cast 1–2 October 2026: 100,590 votes from 9,694 people.