About whether Gemini Flash or Pro answers obscure questions correctly or invents plausible nonsense.
Has this happened?
Yes 31 (24%)Not sure 30 (23%)No 69 (53%)
Votes cast 1–2 October 2026: 100,590 votes from 9,694 people.
Hacker News has set AI a lot of challenges over the years. Which ones has it met?
About whether Gemini Flash or Pro answers obscure questions correctly or invents plausible nonsense.
Votes cast 1–2 October 2026: 100,590 votes from 9,694 people.
Why doesn't Flash get it correct, yet comes up with plausible sounding nonsense? That means it is trained on some texts in the area.
What would make 2.5 Pro (or anything else) categorically better would be if it could say "I don't know".
There will be things that Claude 3.7 or Gemini Pro will not know, and the interpolations they come up with will not make sense.