Goalposts

Hacker News has set AI a lot of challenges over the years. Which ones has it met?

2023 MarchRC_ITR

About whether Othello-GPT shows that LLMs learn world models rather than just statistics.

To be clear, what they did here is take the core pre-trained GPT model, did Supervised Fine Tuning with Othello moves and then tried to see if the SFT lead to 'grokking' the rules of Othello.

In practice what essentially happened is that the super-high-quality Othello data had a huge impact on the parameters of GPT (since it was the last training data it received) and that impact manifested itself as those parameters overfitting to the rules of Othello.

The real test that I would be curious to see is if Othello GPT works when the logic of the rules are the same but the dimensions are different (e.g., smaller or larger boards).

My guess is that the findings would fall apart if asked about tile "N13".

An Othello-trained GPT model plays correctly on boards smaller or larger than the one it was trained on.

Has this happened?

Yes 24 (21%)Not sure 83 (71%)No 10 (9%)

Votes cast 1–2 October 2026: 100,590 votes from 9,694 people.