Need help?
<- Back

Comments (157)

  • aftbit
    >In keeping with our commitment to user responsibility, following the official launch of V4.1 Flash and prior to the release of V4.1 Pro, all requests to the Pro model will be routed to V4.1 Flash and billed at Flash's pricePlease don't do this kind of thing. If a user has validated a workflow on V4 Pro, they might not want to suddenly start testing it in production on V4.1 Flash. Instead, keep V4 Pro around but deprecated for a defined period of time, then remove it.At least as open weights models, it's possible to use something like Together.ai or OpenRouter to run the V4 Pro model as long as other providers keep it up.
  • jiehong
    Sounds nice!But, the web ui chat version of flash has very poor language following abilities in my experience:You may ask it something in English, and get a thinking chain in Chinese with an answer in Chinese, or an English thinking chain and an English answer. Using the retry button on the same question has a 50/50 chance of any of those results.Sometimes, asking something in English, but where information are mostly in another language may make the answer in the language where data has been found. The other day, I asked something about a local German thing, in English, and I got an answer in German instead. It’s as if all the language data stirred it away from the language of the user’s question.
  • mmastrac
    I've been trying out the 4.1 flash preview for some bulk tasks: it did a pretty good job refactoring a bunch of .metal kernels to .cu. It needed less steering than Opus on refactoring, IMO, and writes better comments. It failed to port a root exploit from modern Android to an older Pixel 3, but I suspect part of that might have been harness configuration (it asked me to give it a longer timeout for tasks at some point, but I didn't have a chance to finish that).I was getting something like 300-400 tok/s which was just insanity. It was running so much faster than the toolcalls themselves. Honestly, even if it's not quite as strong in reasoning, it just throws so much so fast that it can do a lot more than you might expect.I'd say it was comparable with GLM5.3 Flash.
  • simonw
    > all requests to the Pro model will be routed to V4.1 Flash and billed at Flash's priceIf I'd carefully tested and optimized prompts against Pro I wouldn't be keen on this particular news. I feel like API model providers should lean towards not swapping out models on their paying customers, no matter how much "better" the new model is meant to be.
  • oefrha
    Source is apparently a banner announcement on https://platform.deepseek.com/usage. Had me searching for a couple minutes...
  • EbNar
    Since a few months, I almost exclusively use the Chinese "flash" models for my needs. They are a joy and they cost pennies per answer. Great job.
  • postalcoder
    I hope DeepSeek takes some time to improve their tuning for reasoning effort. Right now, there are only three reasoning efforts: low, high, and max.For all intents and purposes, "low" is pretty much the same as turning reasoning off, and "high" is similar to "max". "High/max" performs way too much reasoning, takes forever, and causes costs to balloon. They need a proper "medium" setting.I get it that they're probably focused on pushing performance right now, but the ergonomics of the model aren't great.
  • k__
    Beta testers report >400 TPS.https://www.geeky-gadgets.com/deepseek-v4-1-flash-review/I hope some of those speed increases will make it to production.
  • tarruda
    Hopefully it will be open weights and have the same architecture and size as the current v4 flash vision, which is probably the best LLM that can be run on 128G devices.
  • wg0
    DeepSeek v4 Flash with high is already a really great work horse. Reliable. But this time, not only that it is better but they are reducing the price by 50% so that's great.I also find the DeepSeek models to be more precise than Claude models (last I used 4.7) in that I yet had not the occasion where model did something unintentional that I did not direct it to.EDIT: Updated percentage reduction.
  • NitpickLawyer
    It's interesting that this is the third lab to find problems with larger models. Earlier last year oAI was rumoured to have failed their large pretrain. Now google has problems with their pro series, and ds just announced the same. There are some rumours on chinese forums talking about problems with the pretraining phase, so this is not mid/post training related.I wonder if this comes from using the bad architecture scaled up (and it hits some limits) or if this is a data problem (undertrained? bad data? bad pre-processing using smaller models?)...
  • edude03
    I've been watching a bunch of bycloud on YouTube recently, and although he's done a great job reviewing papers from the big AI labs, I feel like I'm missing something - how have all the labs seemingly made a model that's cheaper, faster AND has better performance? Historically `flash` variants (like codex spark as well) have been faster but perform worse
  • mrbonner
    Does anyone use the DS Flash models for general knowledge instead of coding focus tasks? If yes, how do you rate it?
  • eli
    You can test it now on the official deepseek api. Just set your model to deepseek-v4.1-flash-expires-on-0910It’s good and very fast.(Note that the deepseek API trains on your data)
  • nicman23
    qwen3.8-flash-next on a single rtx6000 (~1 euro per hour on a spot vm) with buun-llama is i think the cheapest reasoning / euro atm.hope deepseek makes me change my setup again
  • nickweb
    Via nitter: https://xcancel.com/JustinGorya/status/2097287080128708930Looks like the new model can be used if summoned via the API but the API won't list it.
  • a-ve
    I've been using deepseek-v4-flash as a "worker" model with Claude Code to implement a tool using Rust/Iroh for my personal use, and it works fairly nicely when I use Opus as the planner/reviewer model. It seems to follow the plan generated by Opus, albeit with a few misses here and there that it cleans up later after being reviewed by Opus.Fairly excited for the v4.1 launch. Input cache hit prices have been halved, which looks nice.
  • swiftcoder
    If they can keep up this cadence of Flash leap-frogging the previous Pro, we're in for a good time
  • igleria
    v4 pro was decent then a better cheaper faster model comes now?As a consumer I feel like hansel and gretel combined, deepseek could be the witch.
  • tensegrist
    In keeping with our commitment to user responsibility, following the official launch of V4.1 Flash and prior to the release of V4.1 Pro, all requests to the Pro model will be routed to V4.1 Flash and billed at Flash's price. just in terms of user perception when selling this sort of service, this is what they call a "good look"
  • damsta
    > all requests to the Pro model will be routed to V4.1 Flash and billed at Flash's priceWhile V4.1 Flash performance and cost looks promising this auto re-routing sounds concerning
  • stanac
    My problem with V4 flash is output limit. When I need to write or rewrite a larger file (~1000 lines of code) it will fail with message like output limit reached.
  • coopykins
    I really enjoy using V4 Flash for digging though data and such. Its a very good model for the price. Looking forward to this one.
  • aftbit
    Will V4.1 Flash and V4.1 pro be open-weights?
  • thrownaway561
    I will continue to be amazed by how much power you get from DeepSeek Flash for the cost. I have let that puppy lose on so many projects and it is has never let me down. It can build and entire Rails app in no time and even do the tests. For most things, I don't get why people pay the money for Claude. DeepSeek Flash is my default agent in Omarchy.
  • npn
    Crazy that they still keep the price -- or actually decrease it, even -- despite it is a big improvement. I hope it retains some of the tps speed of the preview release though, 300 tps means gemini flash is no longer "the fastest option" any more.
  • indigodaddy
    So, will it have vision? (based on deepseek-v4-flash-vision-exp ?)
  • nicce
    > In keeping with our commitment to user responsibility, following the official launch of V4.1 Flash and prior to the release of V4.1 Pro, all requests to the Pro model will be routed to V4.1 Flash and billed at Flash's price. If you encounter any issues during your comparative testing between V4 Pro and V4.1 Flash, please do not hesitate to reach out to us with your feedback. Thank you for your support!Wow. Imagine OpenAI/Google/Anthropic doing this! Nope.
  • c0rruptbytes
    are they releasing the weights too?
  • hinow
    But what no one mentions is that the price is going from a starting point of $0.16 to $0.60, so basically they're charging nearly four times as much.
  • ThouYS
    if this beats GLM 5.3 flash, I am sold
  • neugls
    Waiting to use it
  • asamadx
    [flagged]
  • donk8r
    [flagged]