Need help?
<- Back

Comments (78)

  • GodelNumbering
    This is the golden age of model training. Some days ago, I decided I wanted a local CPU only model that can perform exceptionally well for English to Bash translation (to avoid the googling for command syntax). I got a bunch of subagents to generate large amount of training data (140k+ samples), got the Qwen 3 0.6B base model, pointed Astra at it, and off to the races. It trained for 2 days (on and off) and I got a surprisingly good model for my task! The total active time I spent was a few hours. And it is still improving, what a time to be alive!
  • netvarun
    Off topic:With sol pricing drop tbh kimi k3’s value prop has not been that great. For our internal use case/testing/benchmarks sol come out with way better quality and much cheaper costs. Kimi really needs to drop their pricing (I heard it’s set by them across all the neoclouds) Sol is at 2/10 vs kimi’s 3/15
  • jamienk
    Ignoring for the moment issues of what "counts" as open, won't open models rapidly advance due to stuff like this in ways that it's less possible for the proprietary ones to do? This is exactly how Linux & Wikipedia, for example, overtook their "frontiers", right?
  • srameshc
    > The problem: thinking models think too muchI see that with Opus 5, it started thinking like crazy in the last few days , I don't think my workflow is that complicated, still it gets into thinking mode and stays there
  • nico
    > The problem: thinking models think too muchThis is partly the appeal of Jev et al; having a quick model for simple tasks, that doesn’t require that much thinkingIt’s amazing all the workflows that models like that can unlock. And yes, classifiers and other ML models have been around for a while for these types of tasks, but Jev has made it easy and cheap to play and experiment. This in turn, is incentivizing people to try them for a bunch of stuff, unlocking creativity and producing a lot of new cool (and eventually potentially very useful) applications
  • Arcuru
    Over on /r/LocalLLaMA there's a group that's been getting popular doing the same thing for the Qwen 27B (and other) models. - https://huggingface.co/ukisai
  • andsoitis
    > The problem: thinking models think too muchAnalysis paralysis stifles not just human intelligence, but other intelligences too.
  • riquito
    Aside. I find the "cost per task" charts both useful and uncanny. Is It better a model that takes me to 90% in 1 dollar or one that takes me to 95% in 2 dollars? Or a different model that too scores 90% in 1 dollar? How much will it cost me the last 10% or 5%? At the end of the day, cost to 100% is what matters and the half (90%) backed solution may require more to reach 100% (or not, who knows?)
  • intothemild
    So they trained a model on open weights, and then aren't releasing the weights... am I reading this right?
  • tomrod
    Well done, and great iteration.The pareto frontier needs clearer distinction. Benchmarks miss half the story. What, if any, capability is lost by the token reduction (for example, was it like super awesome at Golang before and now kind of sucks? that kind of distinction).
  • erichocean
    Need this done for DeepSeek, ideally one of the Flash models.
  • tdhz77
    Does anybody know if this would be a good model for creative writing?
  • dbuxton
    Do they mean Opus 5.5 or Opus 5?
  • themgt
    The result? Ember-1 set a new Pareto frontier for Bedside Bench across both open and closed models including GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5 on cost/task."Pareto": 8 hits"Opus 5.5": zero hits
  • ls612
    On the smaller end, Quen 3.8, while being extraordinarily capable for a small local model, also suffers from extreme thinking. I wonder if the techniques described here generalize to other models too.
  • monkey_monkey
    I don't think the article mentions Pareto frontier enough.Also, did I miss a memo? Suddenly every article on AI seems to be talking about the Pareto frontier - or have I just not been paying attention?
  • logicallee
    This is really interesting. I think the Fireworks Serverless Training infrastructure they used to develop it is also unique and needed. Except if someone works at one of a handful of the largest labs, it is very difficult to set up or try any sort of training pipeline. The managed training infrastructure makes it available to more people.
  • esafak
    It looks like it would be similar to GLM 5.3 Flash, had they tested it...
  • justmeeew
    [dead]
  • huflungdung
    [dead]
  • fr2029
    [dead]