Goalposts

Hacker News has set AI a lot of challenges over the years. Which ones has it met?

2024 Novemberllm_trw

About using a simple graph problem to benchmark LLM math reasoning, re FrontierMath.

Not sure why you're being downvoted that is exactly why I'm using that simple problem to benchmark LLMs. If an LLM can't figure out how to traverse a graph in its working memory then it has no hope of figuring out how to structure a proof.

Under natural deduction all proofs are sub trees of the graph which is induced by the inference rules from the premise. Right now LLMs can't even do a linear proof if it gets too long when given all the induced vertices.

An LLM traverses a graph in its working memory well enough to complete long linear proofs.

Has this happened?

Yes 72 (50%)Not sure 63 (43%)No 10 (7%)

Votes cast 1–2 October 2026: 100,590 votes from 9,694 people.