About whether LLM safety filters can ever reliably block harmful outputs.
Has this happened?
Yes 17 (14%)Not sure 30 (24%)No 78 (62%)
Votes cast 1–2 October 2026: 100,590 votes from 9,694 people.
Hacker News has set AI a lot of challenges over the years. Which ones has it met?
About whether LLM safety filters can ever reliably block harmful outputs.
Votes cast 1–2 October 2026: 100,590 votes from 9,694 people.
This one limitation of LLMs is kind of my bar for "Not truly AI yet" but I'm not saying it as a "its not good at all" type of bar, moreso, know the limits and work from there. LLMs will continue to struggle with things that require intuition for a while I think. It will get really interesting if they can ever truly detect a bad faith actor using them.