Goalposts

Hacker News has set AI a lot of challenges over the years. Which ones has it met?

2026 Septemberrkagerer

About whether AI reviewers can safely police always-on agents acting on user accounts.

Dots use auto-review[1] to check actions that could affect your accounts or share information against your instructions

So if I've got this right, their security relies on other agents that sit at the boundary and sentry whether a proposed action is allowed.

This means they have to interpret the purpose of the action, what effect it will have, whether those two things align, and what is the potential risk / splash zone for collateral damage.

Sorry, but all the evidence I've seen points to their models being nowhere near good enough to do this reliably, consistently and responsibly.

The architecture also feels ripe for becoming a cat and mouse game between the 'competing' agents. It's already pretty easy to see how humans are manipulating their AI to bypass the baked-in restrictions.

[1] https://learn.chatgpt.com/docs/sandboxing/auto-review

AI agents reliably and consistently judge whether another agent's proposed actions are safe and match the user's instructions.

Has this happened?

Yes 40 (31%)Not sure 34 (26%)No 57 (44%)

Votes cast 1–2 October 2026: 100,590 votes from 9,694 people.