Goalposts

Hacker News has set AI a lot of challenges over the years. Which ones has it met?

2023 Julybumby

About whether LLMs' weaker performance on counterfactual tasks like base-8 arithmetic shows a lack of reasoning.

Think about it. Do you genuinely belief you would score as accurately on a multiplication arithmetic test taken in base 8 ?

No, but I believe this is a different question. I think the more relevant question is whether a human can (even with the caveat of needing more time to reason about it). The larger question for a LLM is whether it can answer it at all and interpret why, without additional training data.

The paper seems to point that the ability of LLM to transfer is related to proximity to the default case. E.g., if default is base 10, is better at base 9 than base 2. I would interpret that as indicating more simple pattern recognition than deductive reasoning. The implication being that real transference is more dependent on the latter.

An LLM does arithmetic in unfamiliar bases like base 8 and explains why, without additional training data.

Has this happened?

Yes 53 (44%)Not sure 50 (41%)No 18 (15%)

Votes cast 1–2 October 2026: 100,590 votes from 9,694 people.