Skip to content
B Bloger.fm

Product / DeepSeek

DeepSeek V4.1 Flash

This site has not written a definition for DeepSeek V4.1 Flash. The name appears in the feed but has no stable public identity the desk can verify, so what follows is only what was recorded — no description, and no claim about what it is.

DeepSeek V4.1 Flash was recorded in 31 items across 8 of the 8 briefings in the current window.

It lost ground: 17 items in the first half of the window and 14 in the second, while the feed as a whole grew about 2.2×.

It appeared most often alongside DeepSeek, DGX Spark and GitHub.

Losing ground
items
31
briefings
8
mentions
66
last seen
2026-09-19

Coverage timeline

Sat 12 Sept – Sat 19 Sept / 8 briefings

Everything recorded

19 September 2026 2 items

DeepSeek just made 1M-token context much less absurd to run. V4.1-Flash uses only 890 bytes of KV cache per token, 4x less than V4-Flash. This is the kind of boring-sounding breakthrough that could make long-context agents actually practical.

@HuggingPapers

DeepSeek just released DeepSeek-V4.1-Flash > A 552B MoE multimodal model with 1M-token context that compresses the KV cache to 890 bytes per token, slashing deployment costs for long-context agents.

I wonder how many evals are like this

@teortaxesTex

DeepSeek V4.1 Flash is the model with the biggest gap between its significance and the interest of the evaluator community. No ARC-AGI, no math-arena, no WeirdML… I guess a "0.1 flash" update doesn't sound like big news, plus AA score is middling. Disappointing.

18 September 2026 6 items

DeepSeek-V4.1-Flash at 1M context on one DGX Station GB300 — v15 "Pin Hot Experts": same 62 GB of experts in Grace, chosen by usage instead of layer order Release: v15 (2026-09-17) · Status: confirmed, two same-window pairs · 153 tok/s single-stream prose (v14: 89, +73%) · 165–250 tok/s on agent/code text (v14: 97–160) · ~955 agg tok/s at C16 (v14: ~400) Mechanism: profile routed experts on real traffic, pin the 295 hottest per layer in HBM, split the MoE into two unfinalized FlashInfer calls + one fp32-FMA finalize. Bind-mounted hook, no fork.

口喷剪辑时代来临,牛逼👍

@gengdaJ

把我珍藏已久的好东西分享给家人们,剪映11.4.2开源!!!🥳🥳🥳 > 配套神级Skill: > 只需要简单提示词,口播类视频,全程托管给Codex,不需要任何人类操作。 > 即使略微失误,由于调用豆包ASR,片段被切割到毫秒级别,导入剪映草稿也可以人工丝滑调整😋 > 放一个原本8分钟,剪辑后3分钟,完全由Codex+yichen-jianying-edit Skill自动剪辑完成的,前后视频对照,以及提示词截图👇 > 整个剪辑流程可能需要花费的地方: 1.Codex Token费用,可以用别的便宜模型平替,比如Workbuddy免费的hy3和DeepSeek-V4.1-Flash。 2.豆包ASR,一小时八毛钱,说实话,不能再便宜了。。。

Kimi K3 is now free in Cline Desktop to celebrate all the new users and give them more tokens to continue exploring the app ❤️

@cline

Introducing Cline Desktop - a native interface for working with open weights models. > Use with ClinePass and all our free models like DeepSeek-V4.1-Flash, Musespark-1.3, or BYOK with any provider!

truncated at source

You can run it too, easily. DeepSeek v4.1 Flash for 2x DGX Sparks Buttery smooth experience.

@HealthRanger

Special thanks to @MiaAI_lab for this recipe, which I have confirmed: "MiaAI-Lab/DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks" (on Github) indeed works quite well. > Installed it on 2x DGX Sparks, connected via the ConnectX-7 network ports, which took quite a bit of debugging to get to actually work at 200+ Gb/s. (The units ship with a bad settings that limits speed to around 12 Gb/s.) > Now getting nearly 48 tokens / second aggregate output (decode) across 4 lanes of concurrency, tested at theoretical allowance of 600K tokens context window but not actually using anywhere near that many tokens (as speed falls off when KV gets filled). > Have not benchmarked pre-fill, but the decode token speed is impressive for this hardware, and Mia AI Lab has once again achieved a huge milestone for run

Well, not quite frontier, but at least around Muse Spark 1.3 Pretty sensible results for a 40-layer 8/16B active

@j_dekoninck

We just added DeepSeek-v4.1-Flash on MathArena! Not quite on the level of Qwen-3.8, but boy is it cheap.

17 September 2026 3 items

You can do it too. Run DeepSeek v4.1 Flash locally: - 2x DGX Sparks ~50 tok/s on prose. - Smooth as a butter.

@jmurillocode

Well, well, well. > This DeepSeek v4.1 update but the one and only @MiaAI_lab is indeed insane. > DeepSeek-v4.1-Flash EXL3 (2.9bpw) serving TP=2 across two DGX Sparks. 600K context, 29.3K-token prompts, 2048 max tokens, thinking off. Median of 3 runs. > c=1: decode 49.9 tok/s · prefill 849 tok/s · TTFT 34.5 s > c=2: decode 28.7 tok/s per stream · 57.3 tok/s aggregate · wall 140.3 s · slowest TTFT 70.0 s > Opus4.8 intelligence level. > Let thank shink.

Zartbot is so wonderfully clear V4.1 is a bigger architecture advance than many imagine. It's an even wilder departure from the norm than V4 was, and more compelling. Is this weren't late 2026, I'd say this is the new default Transformer.

@zartbotF

A deepdive analysis on DeepSeek-V4.1-Flash,

While many are transitioning to DeepSeek v4.1 Flash and other models, @plotarmordev has continued working on important PRs and fixes. It's still the most widely used recipe for 2× DGX Sparks.

@plotarmordev

12 PRs merged on DeepSeek V4 Flash (2x DGX Spark), still the most used recipe, and we're keeping the improvements coming > The update fixes tool-call truncation crashes, tightens startup and benchmark scripts, and makes status checks report failures instead of silently passing 👇

16 September 2026 3 items

DeepSeek-V4.1-Flash is now available through DigitalOcean Inference Engine. 🆕 New Causal Encoder-Decoder architecture (8B params on input, 16B on output) beats the larger @deepseek_ai-V4-Pro on most agentic/coding benchmarks, cuts KV cache to 1/4 the memory and 1/8 the storage of the last Flash gen. Natively multimodal, 1M-token context.

DeepSeek v4.1 Flash improvements for 3-4x DGX Sparks ✨ - 14% speed in decode on chat-length replies. - Fixed a bug in thinking on/off. - Answers match what you asked for instead of silently reasoning in the background. - If the model gets stuck in a loop, it now stops. Expect further improvements. Get it here:

truncated at source

Deepseek 4.1 flash uncensored

@OrcaRouter

🐳 Run DeepSeek V4.1 Flash locally on your Mac — uncensored for security research. > We just released Orca’s official MLX weights for Apple Silicon. > This build is designed for AI security research, red teaming, alignment research, and agent-security testing — where refusal behavior itself can get in the way of measuring the model. > 4-bit — recommended → 458.7 GB → 0.9954 routed-expert fidelity → 512 GB Mac > 3-bit → 364.3 GB / 512 GB Mac 2-bit → 212.2 GB / 256 GB Mac > On our refusal evals, the uncensored build reduced harmful-prompt refusal by 87–96% across JBB, AdvBench, MaliciousInstruct, HarmBench, ForbiddenQuestions, StrongREJECT, and SimpleSafetyTests. > Built for security researchers who need to study what happens when the guardrails come off. > The downloadable weights are uncensored. Our hosted API remains guardrailed.🐳 > API:

15 September 2026 4 items

Something just launched that makes VS Code feel ancient. And im all in for it: Cline Desktop app, an open-source app for open-weight models (and you know how much i love open source) brings its coding agent into a standalone Cline Desktop app for Mac + Windows. Just open a project and tell it what you want to change. No need to open VS Code first. In Cline's demo, it adds priority filters to a task board, works through the code and runs the build. Glad to see the model choice stays open, too. You can bring your own API keys, use supported open-weight models and even switch models halfway through a project!

@cline

Introducing Cline Desktop - a native interface for working with open weights models. > Use with ClinePass and all our free models like DeepSeek-V4.1-Flash, Musespark-1.3, or BYOK with any provider!

Last week these models went live in Command Code. DeepSeek V4.1 Flash Muse Spark 1.3 (with max reasoning) Ling 3.0 Flash Sante (free model) 🐐

DeepSeek V4.1 Flash on max is ranked as the #3 open model on @arena. I'm using this model as my daily driver right now and am incredibly impressed. It's fast, it's smart and I have an enormous amount of context that remains coherent and accurate. The whale definitely cooked.

@YourLocalAILab

With today's release of Jovian Judgement R37, DeepSeek V4.1 Flash has full support with blazing speeds. > Decodes of 500 tok/s and prefills of 20k tok/s for single stream workloads on 4 RTX 6000 Pros. ⚡️ > To learn more point your agent at our GitHub here:

I know @kernelpool is cooking something cool on DeepSeek V4.1 Flash and DwarfStar/ds4 engine! 🚀

14 September 2026 5 items

If you've been looking for a best in class recipe for setting up DeepSeek V4.1 Flash on your RTX 6000 Pros, I have a treat for you from the team at @YourLocalAILab. 🐋

@YourLocalAILab

With today's release of Jovian Judgement R37, DeepSeek V4.1 Flash has full support with blazing speeds. > Decodes of 500 tok/s and prefills of 20k tok/s for single stream workloads on 4 RTX 6000 Pros. ⚡️ > To learn more point your agent at our GitHub here:

Extending $60 usage on DeepSeek V4.1 flash to 10 days. 🐐

@CommandCodeAI

DeepSeek V4.1 Flash is live in Command Code. > 6x usage on $10/mo GOAT for a week. $60 usage. 154K requests. 7.8B tokens. > GOAT plan is the best AI coding plan. > Available in all plans and API. > 🐐

With the last commit into DwarfStar now you can use DeepSeek v4.1 Flash in a single DGX Spark as well, with SSD streaming. Around 9 t/s generation. It works also dual-spark RDMA at ~22 t/s.

You can Run Locally DeepSeek V4.1 Flash It beats GPT-5.6 Sol, Claude Opus 5 & can run with 4-GPU at your home locally No NVLink. - Jovian Judgement R37/TP4/DCP1: - 21,400 tok/s uncached 32K prefill - 280 tok/s C1 decode (budget 50) - 500 tok/s single-stream decode - 417 tok/s sieve - +44% prefill vs R36 - DSpark K7 + Engram (RAM or SSD) - 552B MoE, vision, 1M context Engram in pinned RAM or disk, DSpark K7. -

@0x0SojalSec

You don’t need a GPU-cluster to run a Uncensored DeepSeek-V4.1-Flash 763B model locally at home. > You need a Single RTX PRO 6000 Blackwell. > - 316 tok/s output - 9,351 tok/s prefill - FP8 + vLLM + DSpark - 64-96GB class machine.

Time to test DeepSeek v4.1 Flash with DwarfStar TP-RDMA on 2xM3 Ultra 512GB!

13 September 2026 6 items

💀💀💀 - Abliterated with Cybersecurity Function unleashed Thank everybody for waiting i released a universal Abliteration with unleashed Cybersecurity for: DeepSeek v4.1 -Flash that other can apply with Stock, EXL3 3.5BPW, and TR3 hybrid. this is based off our long lines of DSV4F Dspark preview, 0731, Vision-EXP work saving all the anchors, preserve all the MTP, Dspark, and vision. now GO GET IT!!!!!. don't forget donate or contribute for support HF Repo: (please read instructions on how to apply to your model)

watch me buy more sparks the moment the price decreases one spark is cool but you get so much more intelligence by running two or more

@MiaAI_lab

Run DeepSeek v4.1 Flash on 2x DGX Sparks ✨ > One of the BEST models you can run today, now available for only two dgx sparks. > Ships with conservative settings for stabiliy: - 600k context, 775k KV cache pool - 2 concurrent connections - 550k prefill stress test passed - EXL3 quantization > Performance: ~32 tok/s on prose single stream ~42 tok/s on prose 2 concurrent streams ~1000 tok/s prefill 8k-128k ~872 tok/s prefill on 256k > Get it here:

What a crazy week in AI! 🚀 DeepSeek v4.1 Flash Marigold v2 Suno V6 Tencent AuK YuE2 MiniCPM-5 2B Unimate AlphaGenome Atlas Lingbot World 2 Isaac 0.5 World Sculpt Fire3D Navier-Stokes Show Harness UMR UnifoLM & more! Watch the full recap:

DeepSeek V4.1 Flash just landed on Token Harbor with a free API you can plug straight into your agentic OS, cutting the cost of running AI workflows and automation.

Here’s a list of all AI models launched this month already: * Claude Fable 5.1 * Claude Mythos 5.1 * Gemini 3.8 Flash * Gemini 3.8 Flash Cyber * Muse Spark 1.3 * Qwen3.8-Max-0902 * GPT-6 Astra * GPT-6 Astra Pro * Ling-3.0-flash-VL * Ling-3.0-flash-Sante * Lyria 3.5 * ChatGPT Images 2.5 * GPT-Image-2.5 Flare * GPT-Image-2.5 Sunburst * DeepSeek-V4.1-Flash * Fugu Max * Fugu Ultra v2 We’re not even half way through the month…

艹,居然忘记放白嫖WorkBuddy+ DeepSeek v4.1 Flash的链接👇

@servasyy_ai

我的天,WorkBuddy 海外版 居然可以直接用 GPT6 Astra !? > 除了HY3还在继续免费使用外, 最新上线的 DeepSeek V4.1-Flash 居然可以限时免费~ DeepSeek V4.1-Flash 在 WorkBuddy国内版 可是要钱的,太香了! > 兄弟们,赶紧去冲吧 👇 >

12 September 2026 2 items

truncated at source

2× DGX SPARK OWNERS REJOICE! DeepSeek-V4.1-Flash at 3.30 bpw EXL3, targeting just TWO DGX Sparks. 🔥 And yes, VISION is included. DeepSeek already ships its routed experts in FP4. I pushed that expert bank to a 3.30 bpw using my internal SAGE-EXL3 dynamic quant tool average while preserving the rest of the model, including Vision, Engram conditional memory, and the non-expert/source-format weights that are not part of the 3.30 bpw expert quant. TP4 was step one. Now TP2 is here. 𝗗𝗘𝗘𝗣𝗦𝗘𝗘𝗞-𝗩𝟰.𝟭 𝗙𝗟𝗔𝗦𝗛 𝗢𝗡 𝟮× 𝗦𝗣𝗔𝗥𝗞 SAGE-EXL3: 3.30 bpw routed-expert average Mixed precision: K2 → K8 Full pack: 415.8 GiB 31 shards 40 routed-expert layers Vision included. Engram preserved. Non-expert/source-format weights preserved outside the 3.30 bpw expert-bank average. SAGE is an internal quantization workflow I use so I’m not just applying one flat precision across every expert tensor. That’s about as much as I want to say about the method for now. 😁 𝗧𝗛𝗘 𝗧𝗣𝟮 𝗧𝗔𝗥𝗚𝗘𝗧 The full pack is ~415.8 GiB. Roughly ~189 GiB is the

DeepSeek v4.1 Flash on 2x DGX Sparks Coming soon