<- Back
Comments (95)
- iamcoder18I've been waiting so long for something amazing to come out of the OpenAI and Cerebras collaboration.> In our evaluations, GPT-5.6 Sol on Ultrafast mode answered all 2,500 HLE questions in 11 hours and 11 minutes. Claude Fable 5 needed 78 hours and 27 minutes, more than three days of continuous compute, to arrive at the same conclusions. In other words, Ultrafast worked through the frontier of human knowledge in a single working day, achieving comparable accuracy nearly 7× faster.This is actually insane.Hopefully the release ultrafast of Terra and Luna too.
- TopfiUnless I have read over it, besides the animation in the intelligence vs speed graph which only mentions internal data and not whether they truly reran the AA suite, there is no actually solid statement on the important aspect of performance.Neither the Cerebras or OpenAI post [0] outright state that this performs exactly the same as regular 5.6 Sol. I feel if this was 1:1 just Sol but much faster, they'd (rightfully) scream that off the rooftops. A line such as "this is the same performance, just faster, with no downsides" would go a long way in clarity and communication. Along with no pricing information, I'll hold out on further information.[0] https://openai.com/index/previewing-ultrafast/
- GodelNumberingThe corresponding OpenAI post https://openai.com/index/previewing-ultrafast/There is no pricing info, which could mean it's "if you have to ask..." territory or they are simply gauging interest before deciding
- johnfnThis does look pretty incredible, but don't forget that incredible token thoroughput can only necessarily solve certain bottlenecks. If your e2e tests take an hour, they'll still take an hour after Ultracode. If the agent runs a 10 minute typecheck after a change, that will still take 10 minutes. grep over a massive codebase is still just as slow, etc. I say this not to take away from this accomplishment but just to ensure everyone here keeps a clear head about what it means - 14x faster tokens does not mean it completes every task 14x faster.I suspect Humanity's Last Exam is without tool-calls, making it kind of the perfect benchmark to highlight how fast Ultrafast is, but not really the same as the everyday work you or I do.
- wxw> Compared with output speeds reported by Artificial Analysis GPT-5.6 Sol on Ultrafast mode runs 11x faster than Fable 5, and 5x faster than Opus 4.8 on Fast mode.Awesome work. I'm personally very excited for faster models/inference.I think speed is underrated to some degree in the current conversation. For a while, I was using Cursor's Composer quite a lot, even over frontier models, just because of how darn fast it was.
- aenisGood news for Intel and AMD.Rught now on large scale codebases the bottleneck is both claude/codex inference, as well as time it takes to run tens of thousands of tests. We put those workloads on dedicated epyc 9005 build machines - but it still takes minutes per run. Those who can afford the fast tokens will be in the market for faster CPU that money can buy today.
- tristanMatthias> GPT-5.6 Sol on Ultrafast mode, delivering up to 750 output tokens per secondhttps://taalas.com/products/> delivering 17k tokens per second per user on Llama 3.1 8B model.Obviously this is a much smaller model, but I really can't wait for ASICs to take over the LLM space.Imagine running a model like Sol/Fable (even half the size with 60-70% of it's intelligence) on your own ASIC hardware.
- ricardobeatThe omission of Mimo v2.5-Pro Ultraspeed, released in June, which can achieve 1000tok/s is an interesting flaw in the comparison graphs.It is a bit outdated (scores ± 40% lower), but smart enough for a lot of coding tasks, and can cost under 1/10th of Sol.https://mimo.mi.com/models/en-US/mimo-v2.5-pro-ultraspeed
- sashank_1509I don’t know if this is that useful for coding. In some autonomous world, where no one check the code and the agent can just spend 10X more time checking its work and leading to better results, yes maybe it is useful.But if humans need to check its work, then 10X speed doesn’t really matter I guess.
- stillpointlabI haven't wrapped my head around what level of reasoning this involves. Is it equivalent to max?I didn't like Sol initially but it is growing on me the more I use it. Its personality is a bit flat and I caught it taking shortcuts a few times. But once I learned how to interact with it, I'm genuinely warming up to it. I find that it writes code that has fewer bugs even than Fable (although, to be fair I reach for Fable when the task is less well defined).If this has similar performance to Sol at max reasoning level, this would be a compelling reason to shift even more of my work (maybe the majority) to this model.
- owentbrownWhoa. This looks both powerful and expensive.My prediction is that, this time next year, top developers outside ai labs will be spending 50k USD+ on inference.Within labs, I've heard spend is already far beyond this per developer.
- anthonypasqI'd just like to point out that the largest model Cerebras has ever served is Kimi K2.6 which is 1T parameters, so that either means that theyve had a breakthrough on the hardware engineering side of things, or GPT-5.6 Sol is likely a lot smaller than people think.If it truly is only ~1-2T parameters, then this kinda kills 2 narratives for me.1. all the handwringing about open source catching up via Kimi K3 (3T params) is complete nonsense. All that matters imo for determining which labs are leading is intelligence per parameter. Anyone with a enough compute can train a giant model, but being able to squeeze capabilities into smaller models gives you a massive inference and training edge.2. Inference margins are clearly insane, and this explains why OpenAI was able to lower the price of Luna by 80%. Id guess that thing is probably 120b params based on the TPS they are serving it at.
- buybackoffThis is something I'm ready to pay for. Not more per token, but I will be happy to burn through 20x Pro subscription as fast as I consume my Plus weekly limit now, with 10x more tokens per unit of time. I've learned how to deal with and steer Sol medium quite efficiently, but at the same time I realize it's so slow for the small tasks it can do well, and still so unreliable for open-ended tasks.
- thraway3837This is really cool. Someone here commented about similarity between this and hardware advancements for AV encode/decode.I think it's only a matter of time before miniaturization can have a thumbnail sized user-replaceable accessory that contains the LLM built onto the hardware. I admit I don't know how any of that works, but would be amazing to experience. Fully local, fully offline, ultra fast local inference better than any personal computing product.
- crazysimGPT 5.6 Luna Ultrafast when?
- fg137> allowing Sol Ultrafast to accelerate your most time-sensitive, mission-critical workCurious, what are some of the use cases?
- ilakshDid Cerebras get rid of their like $1500 per month plans for open models?
- storusWow, that's even faster than diffusion LLMs but with the Fable-level quality! Congrats!
- HawtAdsTheir dinner plate chips are impressive.
- lostmsuStill no KV caching?
- scotty79I swear that now frontier AI stuff comes out few times a week.
- poly2itI guess Gemini 3.7 Flash is no longer at the pareto frontier of speed to intelligence.
- behnamohFast mode is already 1.5 times faster and 2x more expensive in the Codex subscription plan. If this thing is 14 times faster, then I can imagine running out of my quota in one session.
- pingouMeanwhile they are down 12,68% today because of disappointing earnings.
- Marciplan“our stock price went down today, here’s something to feed it”
- stephencoyner[dead]
- huflungdung[dead]
- applfanboysbgonThis kills the crab.Compilation time will be a genuine bottleneck for slop coding if this becomes the standard generation rate over the next few years. Go, Zig or even C99 with TCC for dev builds, any language that can get you systems-level performance (or close to it) in a dev environment where you can iterate in ms rather than minutes is going to be immensely more appealing than generating a potential prototype in 10 seconds and waiting 15 minutes for it to compile.