<- Back
Comments (227)
- vishvanandaThe reason people aren’t freaking out is because most people are using heavily subsidized subscriptions.I tried the cheapest provider on openrouter and burned through $50 in a few days. Quality was ok, seems slightly above Luna quality perhaps? But that $50 is 1/4 of my codex subscription where I could have burned that many tokens or more using Astra within my weekly reset.This won’t last forever but as long as the frontier labs are subsidizing this heavily the open models won’t matter.
- giancarlostoroCall me crazy but:VRAM & Memory Requirements by Precision• FP16 (Full Precision): Requires ~1,664 GB of VRAM (e.g., an 8x B300 288GB cluster).• INT8 Quantization: Requires ~832 GB of VRAM (e.g., 8x H200 141GB).• INT4 Quantization: Requires ~416 GB of VRAM (e.g., 8x A100 80GB)VRAM aint cheap, Sam Altman ruined the cost of memory, Nvidia doesnt make enough consumer GPUs letting the market go insane over them, I still have friends on 1070s or 1070 TIs because GPUs have been severely overpriced for too long. I remember when a gaming PC was only $1000.Even so why would anyone not sleep on a model they cannot run?
- p1neconeI have a pretty large, complex project I've been building with heavy AI use (new language + compiler). I was following a 'strong model as orchestrator launching cheap models as implementers' pattern, but I recently trialled just using Deepseek-V4.1-Flash as the model for both layers because of the cost savings (with mimo v2.6 flash on code review agents for some decorrelation).I was previously using GLM-5.3 as the orchestrator, after switching to DS anecdotally there was an unnacceptable quality loss, mostly around not taking all the relevant context into account when making decisions, pulling new design out of thin air without discussion too often, and being way too wordy and rambly in documentation despite prompting to avoid it. There's a lot of docs, rulings, core concepts, design philosophy to uphold and DS was just not cutting it.However, it's perfectly capable of being the sole agent for all of my well specced implementation tasks. I've gone back to GLM as the orchestrator.
- mlinseyI'm paying for the heavily-discounted subscriptions, not the API rates. There isn't really a cost gap for me. DeepSeek doesn't have a subscription to compare to, but when I compared the GLM 5.3 usage I got from a $100/mo Z.ai subscription compared to Opus 5.5 on a $100/mo Claude subscription, there wasn't a big gap. And GLM 5.3 is very clearly not a frontier model (deepseek v4 seemed a lot closer, but I didn't use it enough to really say for my workloads).I don't think those subscriptions nave negative contribution margins, either. I think we're seeing a lot of price discrimination by the big labs, and huge margins on their frontier models. The fact that they have been cutting prices to their second-biggest tier of models (Opus/Sol).Open models catching up and collapsing these margins would worry me if I were a shareholder in the big labs, but as a user, I really doubt that the western labs have bigger environmental impact just because they have higher API costs, I think they have a ton of efficiencies they aren't sharing with customers yet because demand is so high.
- lmf4lolOh man. v4.1-flash has been an sbolute game changer for us. We run all our Personal Assistants now on flash (thinking high) by default and it works incredibly well. There is really no need for basic agentic tasks that might require Kimi K.3 or GLM-5.3 levels.Once its gets juicier, we let flash launch specialized subagents with specific models. GLM-5.3 for coding or Kimi K.3 for research and critique.But as a main driver. I love flash. And it brought our bill down by A LOT :D
- gregwebsI have been using DeepSeek 4.1 flash intensively for over a month. If I run it all day long it costs $1-2. Its fast. Previously I was always quickly running up to my Claude/Codex 5 hour window (on the $20/month plan). The cost savings of DeepSeek is real as shown in this article and I am using subsidized plans.DeepSeek is horrible at grilling sessions (the /grill* skills to make technical decisions). It doesn't know how to explain things. Maybe the skill could be adjusted. It also doesn't come up with as good solutions as Opus/Sol.What I use it for is * the orchestator of my coding workflows * the tester/verifier of code changes * the sub agent that explores code or does web searches * putting together code base research reports Previously I planned with Opus/Sol/Astra and then I used DeepSeek for coding, and then reviewed with Opus/Sol/Astra. With the cost improvements to Opus/Sol I am trying to use them for coding instead now so there will be less back and forth review needed.They are all working together in Pi using the extension @tintinweb/pi-subagents where my workflow skill is calling different subagents that use different models.Luna is cost competitive, but doesn't score as well on intelligence. I do need the intelligence for most of what I use it for, so I am not motivated to use Luna. Haiku also doesn't seem like a competitive price/performance mix.
- user43928Because DeepSeek is not "a month or two" behind as claimed in the article.These open models still did not beat February's Mythos / Fable 5.DeepSeek 4.1 Flash is behind GPT 5.6 Sol, and that one is left in the dust by the excellent Opus 5.5.Rumors say Anthropic is holding in reserve the big improvement, Fable 5.5, for the IPO.It's plausible that open models are 6 - 12 months behind, and there is no "good enough". As long as progress doesn't slow down, leading labs have nothing to fear.
- arush15juneI am 4.1 maxxing on commandcode GOAT Plan + api rates with oh my pi for the last 4 weeks, it's absolutely amazing and crazy fast, it's alright if it makes a mistake, I have enough time to iterate again, I have also added an advisor layer of mimo 2.6 pro which does make it a notch smarter. Getting haiku 5.5/sonnet5.5 to work on plans and letting 4.1 flash work through it is helping a ton too.I am a big ChatGPT fan, all our team has ChatGPT Subs, but the TPS across all models including luna is just so damn slow.Commandcode giving 60$ worth of Deepseek for 10$ is just genuinely goat.And it never says no for cyber tasks so that's a big win
- hmontazeriI had the same experience using ds 4.1 last couple of weeks. It’s insanely good for the price. I’m doing mostly web dev it excels at everything I throw at it. The pricing is ridiculous. I canceled my gpt subscription and haven’t looked back hope the pricing stays like that. I almost never need a better model. I still keep my Claude 20$ sub for now but I feel like one more iteration and I won’t need even that anymore I hardly use it
- nerdypepperhttps://tangled.org/astrra.space/ds4-recipe is an incredibly cool writeup on making deepseek v4.1 flash run really fast.
- james2doyleBeen using Flash 4.1 via the ante harness to blast through a GBA recomp. The ante team has pushed hard to make Flash 4.1 perform well under it. So far, I've maybe spent $10 over the last 3 days. Its a real workhorse and works much better in this harness
- potsandpansI'm using it quite extensively in my PlayStation decompilation harness
- 0xbadcafebee[delayed]
- zug_zugI did a test on this a couple weeks ago. What I found was that the chinese models were far better than API rates, but about comparable on price vs the subsidized subcription model (chatgpt). Also it was my experience that codex completed tasks quicker.That said, it's my best understanding that these american companies aren't profitable and will eventually raise rates (the old uber trick) so I'm keeping myself ready to switch when that day comes.
- simpaticoderThe question seems rhetorical but I think there are two reasons in some combination. First is it there is some awareness lag here. That lag can be on the producer and consumer side. Software enterprises are pretty slow to adopt new things and slow to try new things so they might only be aware of openai and Claude as options. Plus there are some scariness because deep seek is a Chinese model and therefore export restricted - never mind that there are American in European providers.The other reason is more interesting. Maybe the frontier providers think that price performance is irrelevant in light of very powerful frontier models that can start the RSI loop and or a huge displacement of work and a winner take all economic situation. After all if frontier providers earn everyone's money then you won't have any money to spend on any model 100x cheaper or not.
- wg0While using DeepSeek v4.1 Flash I was architecting a system and I made a mistake of drawing the RPC boundaries at a wrong place that did cost me in so many ways.I realized that mistake and guided DeepSeek where it should be.Next I fired Fabble 5.5 set to high to check if the hype is real about Fabble. It exhausted 89% of quota and came up with NOTHING that DeepSeek hadn't flagged itself already in its notes.
- apitman> With my OpenCode Go sub of $10/month, DeepSeek is basically unlimitedMy OpenCode Go monthly window was scheduled to reset this morning. It was sitting at 22% used despite me using DeepSeek V4.1 Flash heavily as my implementation agent the past couple weeks (I use gpt-6.1-sol high for planning/orchestration).I had 1.5 hours left so I fired up first 10, then 20, and finally 50 concurrent subagents all working on reverse engineering C code from an old PC game. They found over 100 new functions.This is the first workload I've found that could make a dent in my sub. It got my 5 hour window to 85% used, but sadly my monthly was still only at about 35% when it reset. So that cost maybe $2.
- swiftcoderI think the interesting provider to cross-check this assumption with here is Meta, who is clearly freaking out, and is currently providing Muse 1.3 even cheaper so long as you are willing to share data with them
- RGS1811This model finally got me off my Claude Max subscription. I’ve found it superior to Opus 5.5 in certain use cases, and certainly faster.I’m convinced that I’ll have good enough inference on my laptop at reasonable speeds within the next year.
- alex-moonI think because we're all just using it thinking we have found the "model for me" and never mentioning it to anyone because what would we say? It's good. It's a bit like the Logitech MX Master, as more and more people assumed they had found the ideal mouse for their purposes, it quietly became the professional standard through sheer adoption.
- aguilaairWhat about MiMo v2.6 Pro? It’s throughput is slower by default (UltraSpeed is faster than DS4.1F) but is above the pareto line, and cheaper.see https://artificialanalysis.ai/models/releases/comparisons?co...
- anonundefined
- _jayhack_Enterprise is not freaking out because DeepSeek 4.1 Flash does not actually occupy a spot on the Pareto frontier for non-coding enterprise workflows. We see this at my employer, focused on non-technical knowledge work. Luna 6 and now Haiku 5.5 are both very competitive if not better on all axes that we care about
- ne01Deepseek V4.1 Flash is a hidden gem, really. Not to mention, you can easily get it through many providers that offer zero data retention and consistent speeds above 200 tokens per second!
- wren6991It's a solid little model, and I appreciate DeepSeek's commitment to the bit in releasing a brand new pretrain, double the size, numerous architectural innovations as a ".1" release over the excellent DeepSeek V4 Flash.
- jeffrallen[delayed]
- elmer2DeepSeek isn't even on my mind. I use the frontier models and can get the best in the industry for a relatively cheap price.
- bitfilpedBecause in two weeks someone will be asking why I'm not freaking out about AlphaDolphins 0.3 Zip and then in a month FrozenMonkey 2.5 Artic.
- LeBitI have subscriptions to OpenAI and Claude but use DeepSeek 4.1 Flash for my coding agents.It costs pennies and you got really great output.The author is spot on.
- pants2Probably because Luna is faster, cheaper, and approximately as smart
- try-workingI have used over 40B tokens and spent over $800 on DeepSeek API over the past 30 days, mostly on V4.1 Flash.It's good, and you can do most work with this. For complex software implementation you need to split your runs into various phases, build in verification, and use subagents so that work gets another audit and repair pass from the lead agent. You can do pretty much everything then. Frontier models can do without compelx workflows, that's the difference.
- browningstreetWhat would freaking out look like, or is this just a stupid bloggish title flourish?Is OpenAI coming in $20B under a sign of "freaking out"?
- smallmancontrovThey might be. They would delay public admission as long as possible, because public admission would make stocks go down.
- f6vMy anecdotal experience is that I can’t even trust DS4Pro let alone Flash. I always have to have Sol reviewing the code.
- jbellisI built mjolnir in large part so I could have Opus manage DeepSeek Flash subagents. It's phenomenal and extremely light on the Claude tokens. https://github.com/BrokkAi/mjolnir/And yes, Opus is enough smarter than DSF that it's worth the extra steps. This ranking is from live tickets, no contamination: https://slopcop.com/power-ranking
- booiBecause GLM 5.3 Flash is even cheaper?
- liuliuDeepSeek 4.1 Flash 0910 is perfect for M5 Ultra 256GiB. Running it fully resident in RAM, prefill at ~2500 tok/s and decode at ~40 tok/s. Probably tons of room to improve from there.
- wildsterI like GLM 5.3 Flash, it seems good enough for coding features if you have a good structure and a good AGENTS.md
- aussieguy1234What blows me away about this model is it's speed.It's way faster than Opus or any of the GPT models.I have a coding harness which is opencode plus a few skills relevant to my workflow. Deepseek 4.1 Flash does very well in this environment. I haven't noticed much difference quality wise compared to Opus 5, which I use in my day job as my employer pays for it (although I'm considering using DeepSeek here too given how cheap it is).
- xyzsparetimexyzThere was a moment 3 months back where the sentiment was that cheaper models were the way to go. Since then the pendulum has swung back.
- aszenBecause subscription plans are cheaper, only enterprises paying per tok pricing should be freaking out
- thefourthchimeFor non-coding tasks it may be fine. But for coding, Opus 5.5 is just a completely another level than something like Deepseek 4.1 Flash.Opus 5.5: TIME 9.3m COST / $1.99 / SCORE 99/100 https://jonclegg.github.io/pacman-bakeoff/#claude-opus-5-5Deepseek 4.1 Flash: TIME 2.8m / COST $1.89 / SCORE 72/100 https://jonclegg.github.io/pacman-bakeoff/dev/#deepseek-v4.1...
- pianopatrickI was just using a bunch of models in Cursor to review a project. I went looking for DeepSeek and it was not one of the options.Would be cool if they added it.
- tengbretsonI don't know about "freaking out", but I'd say I'm having a good time here with DS 4.1 flash.
- gskyAmerica bans Chinese models sooner or later just the China banned American big tech
- hypferIs it known why unsloth seems to not have touched DeepSeek 4.1 Flash?
- pizza234People have been raving since forever about Deepseek, but if one looks at the CoT, it's evident that it's way way stupider than frontier models (there's a reason why it's cheap). It's laughable to compare Deepseek 4.1 with Opus 5.5.I've benchmarked, rigorously, deepseek-v4-flash for programming and personal use, and it is definitely less smart than Qwen3.8-flash-next (which in turn, is not terribly smart).Local models are also really slow, unless one spends insane amounts of money.Having said that, Qwen3.8-flash-next is an impressive evolution; it reaches the small versions of the frontier models (like Sonnet) - but again, it's massively slower and not 100% reliable (including: stability).
- robertlane0Honestly for me the intelligence gap between DS 4.1 Flash and Muse Spark 1.3 makes Muse more worth it for me, especially on a $10 OpenCode Go sub, with the caveat that everything I use it on is open source which makes the fact that I'm sharing it with Meta a little moot because it's already published permissively on GitHub anyways.
- anonundefined
- kristianp> shrank the KV cache by roughly 437XCan't you just say "shrank to 1/437th the size"? It's not that hard.
- MisterMunchkinI had it make 25 different things today and it cost $0.70It’s disgustingly good value. I find it capable of doing anything I want.Obviously can’t use it at work, but for home projects it’s awesome.
- anguralbanish2I would love to get them more better, it's good not a bad thing.
- cactusplant7374Because engineers are lusting for 1000 tokens per second. You can only achieve something like that with OpenAI.
- pessimizerI'm no expert, but it think that it's the pricing on GPT-6 Luna. I'm also guessing that it's been underpriced just for this reason. I also don't think it's all that great, but it's definitely very cheap.If it's underpriced, it's a loss leader to sell the other models, so it actually can't be too good.I really put these things through their paces because I use them to review and work with new abstract game rules and models, so they're always flying blind. Luna misses the obvious (and more importantly, the clearly explained) consistently. My second prompt is listing all of the points in its first response, and saying "No, it doesn't work like that." The third prompt is picking out the two or three suggestions it made after correcting itself on all of the original points and saying "That's how it already works." The fourth prompt is "Now that we're done going over the rules, can we start?"I actually feel like 5.6 Luna seemed better.
- sergiotapiaIn my experience it just takes so much longer to arrive at "done" state for me. It thinks for soooooo long. I guess if you're running 12 sessions at once you don't really notice.
- AIblemblioNo they can't.And as long as I pay as little for claude opus 5.5 i do right now, i'm using it.But yes i'm glad that we have alternatives.
- m3kw9i thought 6.1sol copied the caching architecture so this isn't such a big deal no more
- doctorpanglossbecause it doesn't work very well?if you have a legitimate coding application, it isn't very good. if you have some kind of inauthentic activity, which could be what it is trained for for all sorts of reasons...
- verdvermWhy would we freak out? The systems we use have always gotten better, faster, cheaper with time
- kydanet[flagged]
- oh_noAA shows Luna at 1/4 the price, 1 point behind on intelligence matrix with a 38.Haiku 5.5 is 23% cheaper with a 4 point intelligence lead.I'm on subscription usage so I can't compare Flash 4.1 to them directly but the OP has his head up his ass if he thinks Opus 5.5 is the best point of comparison. Why is anyone using Opus if the new Haiku is indistinguishable /sJust absolutely terrible post, admits to using Opus for review but claims its intelligence isn't needed, why aren't you using Haiku or Sonnet then?
- CurbStomper4[dead]
- distantsoundsbecause we've all figured out that AI is just a huge grift?
- srousseyNot comparing to gpt-6-luna which seems comparable and priced well.
- wewewedxfgdfYou might also choose to pay money for a service that provides real value instead of actively choosing to support the Chinese deliberate effort to undermine this country.