About whether LLMs' weaker performance on counterfactual tasks like base-8 arithmetic shows a lack of reasoning.
Has this happened?
Yes 53 (44%)Not sure 50 (41%)No 18 (15%)
Votes cast 1–2 October 2026: 100,590 votes from 9,694 people.
Hacker News has set AI a lot of challenges over the years. Which ones has it met?
About whether LLMs' weaker performance on counterfactual tasks like base-8 arithmetic shows a lack of reasoning.
Votes cast 1–2 October 2026: 100,590 votes from 9,694 people.
Think about it. Do you genuinely belief you would score as accurately on a multiplication arithmetic test taken in base 8 ?
No, but I believe this is a different question. I think the more relevant question is whether a human can (even with the caveat of needing more time to reason about it). The larger question for a LLM is whether it can answer it at all and interpret why, without additional training data.
The paper seems to point that the ability of LLM to transfer is related to proximity to the default case. E.g., if default is base 10, is better at base 9 than base 2. I would interpret that as indicating more simple pattern recognition than deductive reasoning. The implication being that real transference is more dependent on the latter.