Wake me up when AI can...

Oct 3rd, 2026

A few weeks back I was (once more) impressed with a new release of a frontier model, and realized it was still 2026, about three quarters through the year, the same year I discovered coding agents were becoming useful.

It's funny how quickly you get used to the status quo. Using a model for a few days is enough to develop an intuition for its capabilities, and that naturally leads to expectations that would have been unthinkable earlier.

Whenever I hear someone complain about an LLM, I try to picture having the same conversation a year ago. “I let Claude troubleshoot a massive performance regression overnight and told it not to use the profiler because it's too slow, guess what it did?” You let a machine troubleshoot an issue while you were asleep?!

If you use LLMs on a daily basis (and don't keep notes or a blog) it's hard to think back to what you would have been impressed with throughout the years.

Fortunately Hacker News provides a rich archive of statements of the form “wake me up when AI can…” and “I'll believe it when…”. I collected about 800 of these comments posted from 2016 to 2026, and thought it'd be interesting to let people judge whether the bar that others had set for AI had been cleared.

See all the results →

It made the Hacker News front page on October 1st. By the time voting closed on Friday night there were over 100,000 votes from almost 9,700 people.

Votes and voters per 30 minutes from Oct 1st 19:30 to Oct 2nd 22:00: a first peak of about 5,700 votes and over 700 voters per half hour right after it reached the front page, a quiet stretch before midnight, a second peak of about 3,600 votes after it returned to the front page around 00:45, then a steady 1,300 to 2,900 votes per half hour through the day, tapering off in the evening.

Below I'm providing some cherry-picked quotes and statements from the past, together with the results:  yes,  not sure,  no.

Turing test

About 8% of the statements you could vote on were some version of whether AI passes the Turing test. I think this is the clearest example of “moving the goalposts”, in the sense that the plain version gets a clear “yes”, but with further qualifiers the pass rate drops.

It should be noted that people used to talk more about the Turing test in the early days, so the topic may have lost relevance.

Clearly happened

I've somewhat arbitrarily defined “clearly happened” as at least two thirds of the voters saying “yes”. That was the case for 112 of the 826 statements. The topics are mostly code and text. Notice that quite a few of these statements were made before ChatGPT was released (November 2022).

Writing code

Understanding code

Reading and everyday tasks

Contested

A statement is contested if both “yes” and “no” got at least 30% of the votes, and are within 10 points of each other. This was the case for 78 of the 826 statements.

Knowing what's true

Thinking beyond the training data

Replacing programmers

Driving and robots

Natural conversation

Clearly not yet

On 89 of the 826 statements at least two thirds of the people voted “no”. Most of these have to do with the physical world, and with producing something truly novel on its own.

Plumbers, electricians and handymen

Around the house

Robots out in the world

Science

Making things on its own

Everything else

That leaves a relatively large group of 547 of the 826 statements. Skimming through them, most don't fall in the previous categories because of how they are phrased (for example: doing X reliably), because they're too niche for people to remember, nobody tried it, or it's impractical to verify.

A few more cherry-picked quotes in this larger group.

Famous last words

Oddly specific

Methods

Unsurprisingly, collecting comments and putting the voting page together was done largely using LLMs. About 22,000 candidate comments were selected using traditional keyword search. A cheap model (Haiku) flagged 5,200 of them as setting a bar for AI, and a more capable model (Opus) wrote the context and the statement to vote on and dropped the weaker ones, leaving 826.

The web server was a small Go application with an append-only text file as a database, hosted on a ten-year-old single-core VPS with 500 MB of memory.

At the peak about 1,400 votes came in every five minutes, so around five per second. The load average never went above 0.20. Before posting, I load-tested the server on my laptop pinned to one core, where it handled about 50,000 requests per second, so there was some room to spare.

Meanwhile, now that voting has closed, the Go server is gone and the results are plain static pages.

If one of these comments is yours and you'd rather not have it there, send me an email at firstname[at]lastname.ch and I'll remove it.

See all the results →