Goalposts

Hacker News has set AI a lot of challenges over the years. Which ones has it met?

2025 Februaryrvz

About OpenAI research finding AI unable to solve most coding problems.

The benchmark for AI models to assess their 'coding' ability should be on actual real world production-grade repositories and fixing bugs in them such as the Linux kernel, Firefox, sqlite or other large scale well known repositories.

Not these Hackerrank, Leetcode or previous IOI and IMO problems which we already have the solutions to them and reproducing the most optimal solution copied from someone else.

If it can't manage most unseen coding problems with no previous solutions to them, what hope does it have against explaining and fixing bugs correctly on very complex repositories with over 1M-10M+ lines of code?

An AI correctly explains and fixes bugs in large real-world codebases like the Linux kernel, Firefox, or SQLite.

Has this happened?

Yes 96 (82%)Not sure 14 (12%)No 7 (6%)

Votes cast 1–2 October 2026: 100,590 votes from 9,694 people.