About measuring AI intelligence by real-world tasks instead of benchmarks, after Gemini 2.5's launch.
Has this happened?
Yes 43 (38%)Not sure 53 (46%)No 18 (16%)
Votes cast 1–2 October 2026: 100,590 votes from 9,694 people.
Hacker News has set AI a lot of challenges over the years. Which ones has it met?
About measuring AI intelligence by real-world tasks instead of benchmarks, after Gemini 2.5's launch.
Votes cast 1–2 October 2026: 100,590 votes from 9,694 people.
I recently watched some Claude Plays Pokemon and believe it's better measure than all those AI benchmarks. The game could be beaten by a 8yo which obviously doesn't have all that knowledge that even small local LLMs posess, but has actual intelligence and could figure out the game within < 100h. So far Claude can't even get past the first half and I doubt any other AI could get much further.