<- Back
Comments (103)
- hatthewMy comment about humanity's last exam being a misnomer is included, and I proposed better ideas about what a last exam could look like. One of the things I said was "solve an open math problem" which has conclusively been done with Navier-Stokes (regardless of the controversy surrounding that). However, in the spirit of clarifying the goalposts, AI has only passed 1/6 of the tests I proposed. 17% is not a passing grade, so I'd say no, my challenge has not been met.Another thing to note is that the (presumably AI-generated) summary of my challenge does not accurately represent what I wrote, listing only half the things I said and saying "or" rather than "and".
- joegibbsThere's one of mine in there where I predicted in 2023 that it would be 20 years until AI would be reliably able to entirely build and deploy arbitrary applications from a prompt. I was off by about 18 years on that one!
- bmenrighAt least 1/3rd of these predictions aren't clear enough to determine exactly what is being claimed/predicted. Even after reading the full comment multiple times, on a lot of them I couldn't tell where the author had set the goalposts well enough to say whether we've crossed it or not.
- Retr0idHeh, there's one of mine: https://stoppels.ch/goalposts/?c=39727943"GPT-4 looks at original ASCII art of a foot, not copied from the web, and says it is a foot."The vote is currently 64% yes, 18% no.Just now I asked Opus 5.5 to generate an ASCII art foot, and it did a passable job. It's not great, but it's a foot. Then I pasted it into ChatGPT (whatever they're serving to the free tier by default, which seems to be 5.6 Luna), and it said it was a "train/locomotive": https://chatgpt.com/share/6abeaa39-cc80-83ed-851f-29370db089...Maybe it's Opus's fault for drawing a bad foot but I think it's fair to say LLMs are still pretty bad at ASCII art (without additional tool calling etc).
- ben_wVery pleased one of my predictions was totally wrong: https://news.ycombinator.com/item?id=23252711Sure, sure, what LLMs make still isn't "efficient bug-free code": my prediction is falsified because while LLMs can write and train new models with machine learning, ML is fundamentally not advanced enough to throw arbitraty new tasks at like this.
- outloreThese questions could benefit from being rephrased to make it clear what is being voted for
- anonundefined
- ErrantXWhat is interesting to me is in 2016 people were like; pass Turing test, write code, order me a coffee.And even in 2024 the themes are similar, generally more complex or specific about the coding/turing/action test.But in 2026 a huge shift, we have things like; can open a physical door, emulates human pettiness convincingly, makes novel scientific breakthroughs.That alone tells you a lot IMO
- SoftTalkerToo many questions. I bailed after about 10, with no idea how many more there were.
- 6thbitNot sure why this thread got flagged ?Its fun. Can you add a sort by controversial? I'd like to know where people disagree the most between yes and no.
- mrweaselThe Turing test is interesting, because I believe that the current LLMs are perfectly capable of parsing the it in many situations. On the other hand we also have people are sound like they aren't real.Looking back, was the Turing test flawed perhaps? It failed to take into account that humans can be rather bad at telling actual people from a "parrot". Turing was perhaps a little to optimistic about people.
- Cider9986My test would be an AI agent has a constantly growing karma HN account that makes comments of various lengths without being detected or banned. Wait...
- AngryDataBased on the votes, I can only assume people are still deluding themselves on LLMs capabilities. Is it doing amazing stuff? Yes. But it seems like people still think coding is the ultimate and hardest possible job and so if it can do that it must surely be able to do everything else. My personal experience has show that it still regularly makes up garbage and throws in nonsense sources that do not back up its claims.Yeah maybe if your topic has 2 decades worth of text material to absorb it will get it mostly right like with coding, but anything that is less common? Complete crap shoot.Just today I wanted to know if platinum cure silicone will be inhibited by plaster. The first 20 results are all AI spam with 30 pages of fluff and thus unreliable at best, so I asked AI directly. At first it says sulfur and calcium will inhibit the reaction, which is bad because plaster contains those elements. Then it says it will be fine according to X sources. Check the sources, none of them have anything at all to do with curing silicone on plaster, the articles are about using silicone molds to cast plaster. Failure.Eventually I just had to search youtube videos until I found someone doing it in real life.I see the same bad, and sometimes catastrophic, takes on things I have a lot of experience in, like agriculture, construction, and mechanics. It is completely worthless for anything mechanical unless you are trying to start something extremely simple from the 40s or earlier, and even then it will still tell you stuff like "clean the carburetor" on an old hot bulb diesel.
- delichonIf for each mistaken prediction there was some mild accountability, like someone shows up and slaps you with a trout, it would improve the site. But it should be added to the terms of service first.
- travisgriggsHow was this assembled? From a meta point of view, how much AI was used to curate and highlite the goals; how much was used to assemble the site itself? Or deploy it?
- eternal_braidA chess scoresheet sometimes contains mistakes but chess players can figure out in many cases what was meant by thinking of what moves make sense and considering the level of play so far. Popular AIs tools fail at that.
- anonundefined
- User23Voting on this is ridiculous. Obviously we should have AI decide which AI challenges have been met.
- simianwordsI made a bet with a guy on HN that the market value of OpenAI + Anthropic would get to at least 2.5T by 2027. I think I'm on track to winning.https://news.ycombinator.com/item?id=48517353I also made a bet that API inference margins are greater than 10% for OpenAI and Anthropichttps://news.ycombinator.com/item?id=48500827I can make another prediction about Agentic Commerce and I think it will get big. Muse + Grok Bot + Dots.
- ofjcihenNot sure how questions are spread among people but so far all of mine have been “no” barring a few from before 2022.To be fair, none of them have actually been met. Mostly what’s stopping them is the “reliably” part.
- johnsmith1840All I learned from this is that 40% of hackernews are AI haters which maps pretty well from the overtly negative sentiment on it constantly.
- JBitsQuite a few of the challenges revolve around asking for LLMs to complete tasks reliably and aren't about whether an instance of an LLM completing the task exists. Quite a few of the goalposts are consequently completely changed without the surrounding context, are not the same as what the HN commenter requested and hence seem disingenuous to me.
- tamimioWell I said that before AI will soon make the pcb and electronics just like code, it seems some hw engineers didn’t like it, months later there are few products about the same idea :)
- simianwords[flagged]