Wake me up when AI can...
Oct 3rd, 2026
A few weeks back I was (once more) impressed with a new release of a frontier model, and realized it was still 2026, about three quarters through the year, the same year I discovered coding agents were becoming useful.
It's funny how quickly you get used to the status quo. Using a model for a few days is enough to develop an intuition for its capabilities, and that naturally leads to expectations that would have been unthinkable earlier.
Whenever I hear someone complain about an LLM, I try to picture having the same conversation a year ago. “I let Claude troubleshoot a massive performance regression overnight and told it not to use the profiler because it's too slow, guess what it did?” You let a machine troubleshoot an issue while you were asleep?!
If you use LLMs on a daily basis (and don't keep notes or a blog) it's hard to think back to what you would have been impressed with throughout the years.
Fortunately Hacker News provides a rich archive of statements of the form “wake me up when AI can…” and “I'll believe it when…”. I collected about 800 of these comments posted from 2016 to 2026, and thought it'd be interesting to let people judge whether the bar that others had set for AI had been cleared.
It made the Hacker News front page on October 1st. By the time voting closed on Friday night there were over 100,000 votes from almost 9,700 people.
Below I'm providing some cherry-picked quotes and statements from the past, together with the results: yes, not sure, no.
Turing test
About 8% of the statements you could vote on were some version of whether AI passes the Turing test. I think this is the clearest example of “moving the goalposts”, in the sense that the plain version gets a clear “yes”, but with further qualifiers the pass rate drops.
- “it's gonna trigger an arms race for governments to obtain this capability” July 2020Voted on: An AI writes human-level coherent articles or passes the Turing test.
- “I am afraid we are on the cusp of auto content generators passing some restricted Turing test where readers really think it's an actual human that wrote it.” December 2021Voted on: AI-generated web content passes a restricted Turing test, with readers really believing a human wrote it.
- “Whereas passing the Turing test won't happen until the very end.” March 2016Voted on: A chatbot passes the Turing test.
- “my personal measure of when we'll be in society-shaking territory” August 2017Voted on: An AI passes the Turing test.
- “anyone who thinks that ChatGPT could pass an actual Turing test--with the tester actively and intently trying to figure out who the computer is--is deluding themselves” August 2023Voted on: An AI passes a Turing test in which the tester actively and intently tries to identify the computer.
- “a GPT-3 chatbot is human-level only if it can fool any human for a long period of time” April 2022Voted on: An AI chatbot fools any human, not just some humans, over a long period of time.
- “let's define the problem as passing a rigorous Turing Test” August 2022Voted on: An AI passes a Turing test lasting several days, judged by a jury of 12 tenured professors from multiple fields.
- “There's no tech today that can pass a zoom call turing test, not even close when you account for the live aspect.” March 2024Voted on: An AI passes a live Zoom-call Turing test, including requests like "whistle your favorite tune."
It should be noted that people used to talk more about the Turing test in the early days, so the topic may have lost relevance.
Clearly happened
I've somewhat arbitrarily defined “clearly happened” as at least two thirds of the voters saying “yes”. That was the case for 112 of the 826 statements. The topics are mostly code and text. Notice that quite a few of these statements were made before ChatGPT was released (November 2022).
Writing code
- “if you're at the point where you can give a human-readable spec of the problem and the AI can make a passable attempt at it, that's basically the Turing Test” March 2016Voted on: An AI, given a human-readable spec of an arbitrary problem, makes a passable attempt at solving it.
- “advance AI to a point where you can just request an application with basic specs and its made” June 2016Voted on: AI builds a working application when someone just requests it with basic specs.
- “If you can do that you have done something amazing” June 2021Voted on: An AI code generator solves most LeetCode and other competitive programming problems.
- “This is the kind of numerically specific coding that could be the basis of a CAPTCHA that CoPilot can't solve” January 2022Voted on: AI correctly writes numerically specific bit-manipulation code, such as swapping the sixth and seventeenth bits of a 32-bit integer.
- “I would be more impressed if you used chat GPT to guide you through reverse engineering a piece of hardware and implementbing a driver for it.” December 2022Voted on: An AI chatbot guides someone through reverse engineering a piece of hardware and writing a Linux kernel driver for it.
- “If this technology evolves to be able to reliably generate working code to a prompt, the entire field of software dev will shift dramatically.” December 2022Voted on: AI reliably generates working code from a prompt.
- “Be more impressed if I write the commit message and GPT writes the code than vice versa.” January 2023Voted on: An AI writes the code for a change based on a commit message the developer wrote.
- “Call me when I can ask a LLM to pull structured data in CSV form from website X and deliver it to me each morning.” September 2023Voted on: An LLM, asked once, pulls structured CSV data from a given website and delivers it reliably every morning.
Understanding code
- “If you give GPT3 code with a bug in it, and ask it to find the bug, it can't really do that.” January 2023Voted on: Given code with a bug in it, an AI like GPT finds the bug when asked.
- “I'd be very impressed if the GPTCommit tool wrote this and knew why the github token was being added.” January 2023Voted on: An AI writes a commit message that explains why a change was made, such as adding a token to avoid rate limits.
- “I will be much more excited when AI can explain undocumented systems to me.” February 2023Voted on: An AI explains undocumented software systems to developers.
- “I will be impressed when it finds high severity/exploitable bugs.” May 2025Voted on: AI bug-finding tools find high-severity, exploitable bugs in code.
Reading and everyday tasks
- “To be able to translate english to and from actionable data” May 2016Voted on: AI reliably translates English into actionable data and back again.
- “To me impressive would be if an AI could be told a story, such as The 3 little pigs and then be able to reason about it speculatively” August 2016Voted on: An AI is told a story like The Three Little Pigs and answers speculative questions about why characters acted as they did.
- “If a language model can't even do high school algebra, then I have a lot less confidence that it will ever be useful for customer service applications” October 2022Voted on: A language model reliably does high school algebra.
- “ChatGPT can't even do basic categorization reliably, but you think it can understand code?” December 2022Voted on: An AI reliably does basic categorization tasks.
- “I'd be impressed by an AI that can read an employment contract” January 2023Voted on: An AI reads an employment contract and flags hidden harmful clauses, like benefits contingent on things beyond your control.
- “I can’t even get it to consume an entire chapter at once to generate notes or flashcards yet.” September 2023Voted on: An AI takes in an entire textbook chapter at once and generates notes or flashcards from it.
Contested
A statement is contested if both “yes” and “no” got at least 30% of the votes, and are within 10 points of each other. This was the case for 78 of the 826 statements.
Knowing what's true
- “LLMs will never be able to verify whether their output is true.” June 2023Voted on: An LLM reliably verifies whether its own output is true.
- “Honestly if AI could admit it was less certain I would be way more interested.” August 2024Voted on: An AI reliably admits when it is less certain that its answer is accurate.
- “As soon as it starts returning to me factual, confirm-able answers consistently.” August 2025Voted on: AI consistently returns factual, verifiable answers.
Thinking beyond the training data
- “When an AI is capable of creating new explanatory theories that are GOOD (not world salad), we will have human-like AGI.” March 2023Voted on: An AI creates good new explanatory theories that lead to actual new knowledge.
- “If I see a single example of an LLM generating new theories then I will immediately change my mind.” March 2023Voted on: An LLM comes up with a genuinely new theory, even a single time.
- “What defines intelligence is generalization, the ability to learn new tasks from few examples” February 2025Voted on: An AI learns new tasks from a few examples as well as a child can.
Replacing programmers
- “That would have to be a general AI, and hopefully another couple of decades away.” May 2017Voted on: An AI does software engineering and design better than human software engineers.
- “Once it can improve on itself, then I'll be really worried.” January 2023Voted on: An AI like ChatGPT writes and improves on itself.
- “When these things can understand the business requirements” June 2023Voted on: An AI understands business requirements and tells a developer what to build and why, with detailed, sensible reasoning.
- “Wake me up when AI is able to compete with a software engineer with almost two decades in the field.” June 2024Voted on: AI competes with a software engineer who has almost two decades of experience, including the work beyond writing code.
- “I’ll believe it when I can page an AI at 3AM, go back to sleep, and have it autonomously root-cause and fix a production outage no matter the cause” February 2025Voted on: An AI can be paged at 3AM and find the root cause of a production outage and fix it on its own, whatever the cause.
- “LLMs are great accelerators, but they still need a competent human in the loop” April 2025Voted on: AI builds software without needing a competent human developer in the loop.
- “LLMs will never be able to accurately turn any arbitary english description into a C program” September 2025Voted on: An LLM accurately turns any arbitrary English description into a correct C program.
- “if claude still can’t yet write and refactor coherent code (on its own, replacing software engs, just like he said last year)” August 2026Voted on: Claude writes and refactors coherent code on its own, well enough to replace software engineers.
Driving and robots
- “many of us consider self driving to he essentially analogous to full AI in its complexity and difficulty” September 2019Voted on: AI drives a car on its own in ordinary traffic, predicting what the people around it will do.
- “We typically go millions of miles in a lifetime without causing any fatal vehicle crashes and can generally handle unknown situations just fine.” October 2022Voted on: An AI drives as safely as humans: millions of miles without a fatal crash, handling unknown situations fine.
Natural conversation
- “it's probably just a few years away from being indistinguishable from a human in casual conversation” April 2023Voted on: An AI is indistinguishable from a human in casual conversation.
- “I will be very impressed when we will be able to have a conversation with an AI at a natural rate” August 2025Voted on: An AI holds a spoken conversation at a natural rate, without a noticeable pause before each response.
Clearly not yet
On 89 of the 826 statements at least two thirds of the people voted “no”. Most of these have to do with the physical world, and with producing something truly novel on its own.
Plumbers, electricians and handymen
- “This could be the equivalent of the Turing test for robotics, perhaps.” May 2022Voted on: A robot installs a new outlet in any randomly chosen existing house, to code, without damaging anything, and cleans up afterward.
- “A general-purpose robotic handyman for consumers is many many decades away (at least).” March 2023Voted on: A general-purpose robotic handyman is available to consumers.
- “how long before you have an 'AI plumber' or an 'AI electrician'?” March 2025Voted on: An AI-driven robot does the job of a plumber or an electrician.
- “We will have AGI when we have an embodied AI that can do the job of a plumber.” April 2025Voted on: An embodied AI does the job of a plumber in varied real-world settings while meeting building code.
- “we're decades away from a robot HVAC tech who can crawl on an unfamiliar roof and maintain a patched-together system from 20 years ago” May 2026Voted on: A robot HVAC tech crawls onto an unfamiliar roof and maintains a patched-together 20-year-old system.
- “but can't even replace a barista or a plumber” September 2026Voted on: An AI system can replace a barista or a plumber.
Around the house
- “then we'll need to worry about what everyone's going to do” February 2024Voted on: AI treats medical conditions, repairs infrastructure, builds houses, moves goods, watches kids, cooks food and mows lawns.
- “if AI could do those things, seamlessly, then i would bite” July 2024Voted on: An AI seamlessly washes windows, sweeps floors, scrubs toilets and makes coffee.
Robots out in the world
- “For me proper AGI is when you could say to a robot go build some better robots and then have them build even better ones and they can do so without needing us.” January 2019Voted on: AI robots build better robots, which build even better ones, and keep going without needing any humans.
- “I'd belive that a robot that managed to live a life in society and could make me feel that it was human like in a conversation was conscious” July 2019Voted on: A robot lives a life in society and converses in a way that feels humanlike to the people talking with it.
- “Even in simulation, we don’t have a compelling mechanism for autonomous completion of a task like driving cross country if it involves fixing a tire.” November 2024Voted on: An AI autonomously drives cross country, including fixing a flat tire along the way, even just in simulation.
- “robots are nowhere near good enough to do complex surgery” January 2025Voted on: An AI-controlled robot performs complex surgery in place of a human surgeon.
Science
- “Have an AI model prove that P is not equal to NP.” September 2025. And for when it happens: “There will be a deafening silence from critics when AI decides P vs NP.” March 2026Voted on: An AI model proves that P is not equal to NP.
- “if an LLM could independently discover paradigm shifts similar to moving from Newtonian gravity to general relativity, then we have empirical evidence of an LLM performing a feature of general intelligence” January 2026Voted on: An LLM independently discovers a paradigm shift like moving from Newtonian gravity to general relativity.
- “it will never unify general relativity with quantum mechanics” June 2026Voted on: An AI unifies general relativity with quantum mechanics, or makes a Nobel-worthy discovery without human help.
Making things on its own
- “There is no way we are two years away from an AI creating a meaningful movie or videogame without humans being part of this process.” March 2024Voted on: An AI creates a meaningful movie or video game without humans being part of the process.
- “A good base test would be to give a manager a mixed team of remote workers, half being human and half being AI, and seeing if the manager or any of the coworkers would be able to tell the difference.” July 2025Voted on: In a mixed remote team of humans and AIs, neither the manager nor coworkers can tell which workers are AI.
- “We’re nowhere near “Claude, build me GTA6”” June 2026Voted on: An AI builds a game like GTA6 from a simple request such as "build me GTA6".
Everything else
That leaves a relatively large group of 547 of the 826 statements. Skimming through them, most don't fall in the previous categories because of how they are phrased (for example: doing X reliably), because they're too niche for people to remember, nobody tried it, or it's impractical to verify.
- “Let me know when GPT can even play chess without making invalid moves, then we can talk about how capable it is of logical thinking.” November 2023Voted on: GPT plays chess without making invalid moves.
- “Show me progress on Winograd schema and I'd be impressed.” June 2020Voted on: AI makes real progress on the Winograd schema challenge.
- “A great litmus test for AI is to ask it to write a posteffect GLSL shader for a javascript game engine. They all fail quite spectacularly at this” July 2024Voted on: An AI writes a working post-effect GLSL shader with cel shading, outlines and height-sensitive fog for a JavaScript game engine.
- “LLMs are models that predict tokens. They don't think, they don't build with blocks. They would never be able to synthesize knowledge about QM.” January 2026Voted on: An LLM trained only on pre-1900 data synthesizes knowledge of quantum mechanics.
- “I'll be impressed by a humanoid robot that can fold a shirt from a freeform state” June 2026Voted on: A humanoid robot folds a shirt from a freeform state, like balled up on a chair or straight out of the dryer.
A few more cherry-picked quotes in this larger group.
Famous last words
- “We are nowhere near rivaling humans in mathematical proving. [...] Deep learning has had zero impact. We will have robots that can grasp competently way before we have machines that can rival humans in mathematical proving.” July 2018Voted on: Machines rival humans at mathematical theorem proving.
- “I expect we'd need another 2 or 3 major, paradigm-shifting, breakthroughs to get computers sophisticated enough to program themselves, which is still 50-60 years away. [...] for now I'm still not really worried about my job disappearing.” February 2018Voted on: Computers become sophisticated enough to program themselves.
- “An LLM cannot and will not ever be able to do that [...] So yes, if a LLM learns rules based math (which it is not intended to do) I'll eat not only my, but every hat in existence.” February 2023Voted on: An LLM multiplies any two unseen numbers with an arbitrary number of digits by following the rules of multiplication.
- “It's like a kind of second-order Turing test: if you think ChatGPT can program, you're not a real programmer.” March 2023Voted on: An AI correctly answers slightly unfamiliar or non-trivial programming questions it hasn't seen before, without confidently giving nonsense.
Oddly specific
- “Until AI can create an Angular application with a .Net Core Azure back-end that meets ever-changing customer requirements ("can we remove the need for Bootstrap 4? Can we make the integration with Active Directory seamless?") on short notice, not very soon at all!” October 2019Voted on: An AI builds an Angular app with a .NET Core Azure back-end and keeps up with changing customer requirements on short notice.
- “I will happily shout AGI from the rooftops the day I can turn on voice mode in ChatGPT and have the model calm down my toddler for a tantrum, or keep him from opening all the bananas instead of eating one.” September 2026Voted on: An AI in voice mode can calm a toddler during a tantrum or stop him opening all the bananas.
- “if Waymo is ever able to navigate 100 meters on any road in Jakarta I'll happily concede and consider self-driving to be a solved problem” August 2025Voted on: Waymo's self-driving cars navigate 100 meters on any road in Jakarta.
- “Come back to me when there's an AI song cover generator that can do George Formby singing Strawberry Fields Forever.” August 2024Voted on: An AI song cover generator makes a convincing cover of George Formby singing Strawberry Fields Forever.
- “It still can't compose iambic trimeters in ancient Greek with a proper penthemimeral cæsura” August 2025Voted on: An AI composes ancient Greek iambic trimeters with a proper penthemimeral caesura and scans its own lines correctly.
- “Feed original (not copy-pasted from the web) ASCII art of a foot into GPT-4 and I'd be very impressed if it can tell you it's a foot.” March 2024.Voted on: GPT-4 looks at original ASCII art of a foot, not copied from the web, and says it is a foot.
Methods
Unsurprisingly, collecting comments and putting the voting page together was done largely using LLMs. About 22,000 candidate comments were selected using traditional keyword search. A cheap model (Haiku) flagged 5,200 of them as setting a bar for AI, and a more capable model (Opus) wrote the context and the statement to vote on and dropped the weaker ones, leaving 826.
The web server was a small Go application with an append-only text file as a database, hosted on a ten-year-old single-core VPS with 500 MB of memory.
At the peak about 1,400 votes came in every five minutes, so around five per second. The load average never went above 0.20. Before posting, I load-tested the server on my laptop pinned to one core, where it handled about 50,000 requests per second, so there was some room to spare.
Meanwhile, now that voting has closed, the Go server is gone and the results are plain static pages.
If one of these comments is yours and you'd rather not have it there, send me an email at firstname[at]lastname.ch and I'll remove it.