Need help?
<- Back

Comments (99)

  • postalcoder
    If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%).Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"
  • gruez
    Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked?https://www.youtube.com/watch?v=tNmgmwEtoWEAs others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in performance should be taken with a grain of salt.
  • nullbio
    Where are the model stats? Is this open-weights? If not, why would I use this over DeepSeek Flash 4.1?I think these competing labs need to realize that no one wants another closed-weight model provider... We aren't even happy with the two we have right now, and their days are entirely numbered. If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper).The big labs love to release their new model and quantize after the first week. You don't have that problem using dirt cheap API rates on OpenRouter. DS 4.1 flash is also faster than fast mode Astra. OAI's subscription rates are good value, but now these new open-weight models are nearly as cheap on API usage rates. I honestly can't wait for the day we're not beholden to the two big labs anymore. No wonder there's so much fear pumping happening at the moment from Anthropic and their funded NGOs.
  • TheJCDenton
    > SWE-2 is post-trained from Kimi K3On the one hand I would have expected a completely new model, on the other hand it's an RL-ed K3 go Fable 5 capabilities, which demonstrate that this is probably possible, which is nice.
  • pkilgore
    Not sure it matters when devin is the most consistently shit product I've used. And yes, I tried again, they wasted the money on the billboards.
  • andai
    Their benchmark used to show other metrics, like output tokens and time, but now only shows cost:https://cognition.com/frontiercodeWhich is too bad, since all of the gains here appear to be from massively reduced output tokens?The model SWE-2 is based on, Kimi K3, is cheaper per token than Sol, but costs more per task (ArtificialAnalysis) due to using way more tokens.Whereas, based on the graphs, SWE-2 appears even more token-efficient than Sol! That might have been worth showing off, if true.
  • sbseitz
    Why doesn't clickbait trash like this get moderated ?
  • CyLith
    I think I'm probably in the minority here, but for my line of work, the software engineering and coding is only a small part of the work. I write simulation software, so a deep understanding of physics, math, and how they can be applied to the software is absolutely crucial. I'm assuming this model is tuned to be more focused on SWE topics, and the very reason we seek "multidisciplinary" hires is the also why I actually need a jack-of-all-trades model to back my coding agents.
  • mydreamof
    Seems like benchmaxing? For example for Terminal-Bench 4 it doesn't have great results. And why not show other benchmarks?
  • bobtheborg
    SWE 1.6 was great for small tasks. Very fast and good enough. 1.7 was unusable for me. Took more time thinking than GLM 5.2 and seemed to be generally running in circles. I tried it but abandoned it.Looking forward to 2 -- maybe it'll be usable
  • Take8435
    Post made by account 2 days ago.
  • pelorat
    Unless it can do CAD via computer-use how can you say it rivals GPT-Astra?
  • alansaber
    Fair enough that they did a "propoganda and censorship" eval but not sure why i'd care about that in my highly juiced SWE kimi FT.
  • monkeydust
    As an Econ graduate, pretty cool seeing Pareto in the "AI-bro" zeitgeist. Slightly surreal watching a 1906 welfare economics idea get rediscovered as a plotting convention. The original, if anyone fancies 579 pages of Italian: https://archive.org/details/manualedieconomi00pareuoft. There is an English translation somewhere.
  • eyeris
    Wonder if this was the model that drove factoring the rsa-260The write-up from yesterday was by somebody from cognition using Devin to translate existing cpu sieving methods to gpu and to optimize the gpu sieve.
  • scronkfinkle
    Please correct me if I'm wrong, but this appears to require Devin to use? I'm disappointed to see I need to use a bespoke platform to interact with this agent, to the point that I probably won't be trying it.
  • llmslave
    At work I setup a cloud worker, where i can spin up as many concurrent agents I want, with unlimited fable 5.1 (thanks employer!!).I now just work from my phone, and speak into the agents as they run. I dont write code and I dont write documents. I work on very complicated distributed systems. I dont open my laptop most days. Its a legacy brick I carry around.Some of my coworkers are still doing things by hand, and are working long hours to produce 25% of the output (when considering hours worked). I stay quiet with my setup. We are in the end times for this job for the people that can see clearly how to automate their own job
  • Tsarp
    "SWE-2 is post-trained from Kimi K3"
  • ltsSmitty
    Well written and good diagrams. No idea the verity of the TMBB (trust me bro benchmarks) but it was pleasing to look at
  • m3kw9
    I'm using Codex, Gemini etc, they all have desktop apps and have a plan, how do i use SWE-2? Thats is a problem they have. I'm not about to switch out my workflow and plans with a shiny LLM that looks benchmaxxed and graph maxxed.
  • wqash71
    The horrible website is made by Claude or Cognition is distilled. I'm so tired of it all.
  • _doctor_love
    SWE-1.5 was surprisingly good when I used it last. I feel like Cognition is one of the solid players that’s flying a bit under the radar while Anthropic and OpenAI race to IPO.
  • bluelightning2k
    I like Cognition as a company and hope they succeed. Seemingly excellent engineering org.I used to really like Windsurf. (Now Devin. Kind of? But also now Antigravity.) I still use it as my editor but haven't touched the agent for a while simply due to the rise of Codex.