Goalposts

Hacker News has set AI a lot of challenges over the years. Which ones has it met?

2024 MarchJensson

About whether code execution tools count as an LLM's "pen and paper" for reasoning.

That is an LLM's "pen and paper".

No, that is an LLM's calculator or programming, it doesn't actually do the steps when it does that. When I use pen and paper to solve a problem I do all steps on my own, when I use a calculator or a programming language the tool does a lot of the work.

That difference is massive, since when I use a calculator that doesn't help me learn numbers and how they interact and how algorithms works, while if I do the steps myself I do. So getting an LLM that can reliably execute algorithms like us humans can is probably a critical step towards making them as reliable and smart as humans.

I do agree though that if LLMs could keep a hidden voice they used to reason before writing they could do better, but that voice being shown to the end user shouldn't make the model dumber, you would just see more spam.

An AI reliably executes algorithms step by step itself, as humans do with pen and paper, without calculators or code tools.

Has this happened?

Yes 56 (52%)Not sure 26 (24%)No 25 (23%)

Votes cast 1–2 October 2026: 100,590 votes from 9,694 people.