Skip to content
B Bloger.fm

Organisation

DeepSeek

A Chinese AI lab known for releasing open-weight reasoning and coding models. Its releases are watched for what they show about training costs as much as for what the models can do.

DeepSeek was recorded in 37 items across 8 of the 8 briefings in the current window.

It lost ground: 15 items in the first half of the window and 22 in the second, while the feed as a whole grew about 2.2×.

It appeared most often alongside DeepSeek V4.1 Flash, GLM and GPT-6 Astra.

Losing ground
items
37
briefings
8
mentions
92
last seen
2026-09-19
Official channel
deepseek.com

Coverage timeline

Sat 12 Sept – Sat 19 Sept / 8 briefings

Everything recorded

19 September 2026 5 items

Jina AI 发布文档解析模型 jina-ocr-v1,可以把 PDF、扫描件、表格和图表直接转成 Markdown。它基于 DeepSeek-OCR 做后训练,沿用其约 34 亿总参数、解码时每个 Token 激活约 5.7 亿参数的 MoE 架构,并加入 FastMTP 推测解码。 Jina 自测中,jina-ocr-v1 在 OmniDocBench v1.6 得 91.14,比 DeepSeek-OCR-2 高 0.89 分;在 olmOCR-Bench 得 83.4,比 DeepSeek-OCR 高 7.4 分。 吞吐量是 Jina 主打的卖点之一。单张 A100、并发 32 时,它达到 2.57 页/秒,在 Jina 测试的 14 个系统中最高,比 DeepSeek-OCR 的 2.10 页/秒高约 22%。 模型权重已经放上 Hugging Face,采用 CC BY-NC 4.0,商业使用需要联系 Jina。

@JinaAI_

Announcing jina-ocr-v1, our new visual document parser with 3.4B total parameters and 570M active parameters, with speculative decoding built in. Throw PDFs, scans, tables, charts, or invoices at it and get clean markdown back. Available on 🤗 & Jina Reader `x-respond-with` today

DeepSeek just made 1M-token context much less absurd to run. V4.1-Flash uses only 890 bytes of KV cache per token, 4x less than V4-Flash. This is the kind of boring-sounding breakthrough that could make long-context agents actually practical.

@HuggingPapers

DeepSeek just released DeepSeek-V4.1-Flash > A 552B MoE multimodal model with 1M-token context that compresses the KV cache to 890 bytes per token, slashing deployment costs for long-context agents.

truncated at source

YC S25 端侧 AI 公司 Cactus Compute 发布 Needle 3,专门做工具调用和结构化提取,不负责普通聊天。比如用户说「把客厅灯调到 30%」,它只需要找到对应功能,再把「客厅」和「30%」填进参数。 它和最近很火的 Jev 有点像,但路线不同。Jev 直接给候选决策打分;Needle 3 还是逐 Token 生成,只是输出被限制成程序能直接执行的工具调用。 Needle 3 一套权重可以裁成 2 到 20 层。2 层版本约 9MB,完整 20 层约 35MB,可以按设备性能选择不同大小。 针对安卓操作指令单独微调后,4 层、2900 万参数版本自测得分 62.5,超过 DeepSeek V4 Flash 的 60.5。

@cactuscompute

We release Needle 3: A Sliceable 8-29MB automation foundation model that can match DeepSeek V4 Flash. > One set of weights, every depth from 2 to 20 layers a model of its own, 25-121M parameters at CQ2-bit, built on our Simple Attention Networks and running locally at up to 4k tokens/sec decode speed on a Raspberry Pi 5. > Needle does not chat. Every turn is a function call: give it the tools your app exposes and it picks the right ones and fills every argument from what the user said, or hand it a schema and it returns a typed record. Ask for something no tool covers and you get an empty list, not a guess. > That trad

🚨 Deepseek v5 Leak: Beats Astra > DeepSeek is reportedly preparing an imminent V5 launch > it could match or beat Fable 5.1 and Astra > Expected to deliver much stronger performance at a lower cost > DeepSeek is reportedly keeping the open-weight strategy Could DeepSeek V5 become the new king of open-weight AI?

truncated at source

A significant benefit is the ability to compare different agents and models when performing the same task. While performance is important, cost is also a crucial factor. Identifying a model that achieves the desired outcome at a considerably lower price point provides the kind of adaptability required by users of artificial intelligence.

@quxiaoyin

We just launched Agentsky @agentsky_dev, world’s 1st Agent Market! OpenRouter is for models. AgentSky is for agents. > Use 40+ agents—Claude Code, Codex, OpenCode, Hermes, Pi in your browser(even your phone!) or via one API. All without installing or setting up anything. > Hit Codex Astra’s weekly limit? Hand off to another agent such as OpenCode + DeepSeek V4.1 in browser without losing any context. > You can compare any agent + model directly in browser and that's how I found Astra costs $5.3 while deepseek v4.1 cost $0.12 on the same dashboard task. (I actually preferred deepseek) > Try it at

18 September 2026 6 items

truncated at source

Cactus Compute just released Needle 3 today. Here's what you need to know. Needle 3 is a sliceable automation foundation model, sized just 8-29MB, that the company says matches DeepSeek V4 Flash on tool-calling and structured extraction tasks. It's one set of weights that works at every depth from 2 to 20 layers, so a single model acts as many models. Parameters range from 25M to 121M, quantized at CQ2-bit using Cactus's own Simple Attention Networks and Cactus Quants. Unlike a chatbot, Needle doesn't converse. Every interaction is a function call: give it your app's tools and it picks the right ones and fills in arguments from what the user said, or hand it a schema and it returns a typed record. If no tool fits, it returns an empty list instead of guessing. Cactus says the 121M version was trained on 360B tokens of structured data, letting it beat models 10x its size on mobile tool calls and match models 2-3x bigger on JSON extraction. Key numbers: - Model size: 8-29MB - Parameters: 25M-121M at CQ2-bit

DeepSeek Harness v0.1.6-alpha.2 Pre-release with another batch of updates. The app is coming together nicely…

I HAVE ALWAYS SAID THE NARRATIVE AND TIMELINE IS MOVING FROM USA LABS TO CHINESE LABS this is another proof of that look at the numbers Bolt shared from the first few days of open models on Forge - 1. GLM 5.3 Flash 54% 2. DeepSeek V4 Pro 17% 3. GLM 5.3 15% 4. Kimi K3 14% Chinese labs are moving fucking fast rn

@boltdotnew

11M+ people build on Bolt. On Monday, we gave them open models with up to 50x usage. > Which is #1 so far? The fastest, cheapest one, with prompts up to 2x the size. > 🥇 GLM 5.3 Flash 54% 🥈 DeepSeek V4 Pro 17% 🥉 GLM 5.3 15% 👏 Kimi K3 14% >

Open models are proving they can lead, not just fill the gaps. Real builders are putting them to work, and the numbers show it. GLM 5.3 Flash captured 54% of Forge prompts through Wednesday, with Flash getting up to 50x usage.

@boltdotnew

11M+ people build on Bolt. On Monday, we gave them open models with up to 50x usage. > Which is #1 so far? The fastest, cheapest one, with prompts up to 2x the size. > 🥇 GLM 5.3 Flash 54% 🥈 DeepSeek V4 Pro 17% 🥉 GLM 5.3 15% 👏 Kimi K3 14% >

AI for the people

@ShoBeiYong

deepseek迎来史诗级增强

让 AI Agent 操作浏览器,难点不在“能不能点”。 更关键的是,如何复用已经登录的真实会话,又不打断你正在使用的窗口。BrowserSkill 在 Agent 与浏览器之间加入本地桥接:Agent 调用 bsk CLI,本地 daemon 把任务交给扩展,再在独立 Agent Window 中执行。需要时也可借用现有标签页,是否允许借用和请求人工协助,都由扩展设置控制。 它可接入 Cursor、Claude Code、Codex、OpenClaw 等能执行 shell 的 Agent,还提供 DeepSeek Harness 插件、远程浏览器配对和可重复的能力评测。适合想把真实登录态、浏览器自动化与 Agent 工作流接起来,同时保留交互边界的开发者。

17 September 2026 4 items

You can do it too. Run DeepSeek v4.1 Flash locally: - 2x DGX Sparks ~50 tok/s on prose. - Smooth as a butter.

@jmurillocode

Well, well, well. > This DeepSeek v4.1 update but the one and only @MiaAI_lab is indeed insane. > DeepSeek-v4.1-Flash EXL3 (2.9bpw) serving TP=2 across two DGX Sparks. 600K context, 29.3K-token prompts, 2048 max tokens, thinking off. Median of 3 runs. > c=1: decode 49.9 tok/s · prefill 849 tok/s · TTFT 34.5 s > c=2: decode 28.7 tok/s per stream · 57.3 tok/s aggregate · wall 140.3 s · slowest TTFT 70.0 s > Opus4.8 intelligence level. > Let thank shink.

Union Alpha is GLM-5.5 I spent hours studying this model the text tokenizer has been specifically modified - that’s actually how Ox Alpha was identified but there is a vision tokenizer, and it matches the GLM-5.3 Flash and the GLM-5.5 was planned for release in September-October if other than the GLM-5.5, then it's DeepSeek or Qwen very strong model

@goodworse

> Opus 5.2 is COMING in the next TWO WEEKS > the model is already being tested as Opus 5 in Claude Code > a model that is not lazy at all and loves details > features a good conversational style and high speed > a cheaper Fable 5/5.1

While many are transitioning to DeepSeek v4.1 Flash and other models, @plotarmordev has continued working on important PRs and fixes. It's still the most widely used recipe for 2× DGX Sparks.

@plotarmordev

12 PRs merged on DeepSeek V4 Flash (2x DGX Spark), still the most used recipe, and we're keeping the improvements coming > The update fixes tool-call truncation crashes, tightens startup and benchmark scripts, and makes status checks report failures instead of silently passing 👇

DEEPSEEK-HARNESS IS A FREE OPEN-SOURCE FRAMEWORK FOR BUILDING CODING AGENTS WITH SWAPPABLE MODELS, TOOLS, SANDBOXES, UIS AND AGENT LOOPS. IT WORKS WITH DEEPSEEK, CLAUDE, GPT, GEMINI AND MORE.

16 September 2026 7 items

DeepSeek-V4.1-Flash is now available through DigitalOcean Inference Engine. 🆕 New Causal Encoder-Decoder architecture (8B params on input, 16B on output) beats the larger @deepseek_ai-V4-Pro on most agentic/coding benchmarks, cuts KV cache to 1/4 the memory and 1/8 the storage of the last Flash gen. Natively multimodal, 1M-token context.

The Forge gates are open. You all came running 🏃 We’re opening up access as fast as we can, with another wave coming soon. Join the Bolt Lite waitlist, our $9/month plan with Forge access → Or skip the wait entirely. Forge is live on all Pro plans.

@boltdotnew

Introducing Bolt Forge. Free until Oct 14th: > - Up to 50x more usage - The new frontier: GLM, DeepSeek, Kimi - Zero usage charges > Live now in your model picker on > And one more thing... 👇

DeepSeek Harness is a free open harness that handles research, writing, optimization, publishing and improvement in one place, routing strategy to stronger models and repetitive SEO work to faster ones with a rules gate before anything publishes.

Ox Alpha: The Most Brilliant Marketing Move In AI This Year 👀 >No logo, no branding, no PR just showed up on OpenRouter on Aug 20 as an anonymous "stealth model" >1M-token context, multimodal (text, images, video), zero data retention >Free for a week, with a claimed 100 trillion tokens/day of capacity behind it >Beat GPT-5.6-Sol and Claude Fable 5 on the DeepSWE coding benchmark >Usage blew past DeepSeek by more than 2x in days Stripe's CEO even called it "very impressive"

THIS FREE OPEN-SOURCE REPO LETS YOU RUN GLM-5.3 FLASH, DEEPSEEK V4 FLASH AND KIMI K3 LOCALLY WITHOUT A GPU OR HOSTED TOKEN QUOTAS. THE TRADEOFF: YOU’LL NEED A LOT OF STORAGE.

truncated at source

Deepseek 4.1 flash uncensored

@OrcaRouter

🐳 Run DeepSeek V4.1 Flash locally on your Mac — uncensored for security research. > We just released Orca’s official MLX weights for Apple Silicon. > This build is designed for AI security research, red teaming, alignment research, and agent-security testing — where refusal behavior itself can get in the way of measuring the model. > 4-bit — recommended → 458.7 GB → 0.9954 routed-expert fidelity → 512 GB Mac > 3-bit → 364.3 GB / 512 GB Mac 2-bit → 212.2 GB / 256 GB Mac > On our refusal evals, the uncensored build reduced harmful-prompt refusal by 87–96% across JBB, AdvBench, MaliciousInstruct, HarmBench, ForbiddenQuestions, StrongREJECT, and SimpleSafetyTests. > Built for security researchers who need to study what happens when the guardrails come off. > The downloadable weights are uncensored. Our hosted API remains guardrailed.🐳 > API:

Union Alpha points to ZAI's GLM family: 26 text+image probes match GLM-5.3's tokenizer; 13 image tests match GLM-5.3-Flash. Oddly, text-only counts swap between Llama/Qwen/DeepSeek-like patterns. Best guess: GLM-5.3-related variant/backend. Owner/model unconfirmed.

15 September 2026 5 items

DeepSeek Harness 最新的 v0.1.6-alpha.1 同时新增实验性 Browser Use 和 Computer Use。模型现在既能直接操作网页,也能查看并控制本地桌面。 Browser Use 支持 Playwright MCP、Chrome DevTools MCP 和 Stagehand,可以直接操作网页。 Computer Use 则接入 Cua Driver,可找到具体 App 和窗口、获取当前界面截图,再根据界面元素或坐标执行点击和输入。官方提供 MCP 和原生两种 Computer Use 接入方式,原生方式可以直接集成进 Harness,不需要另外安装独立的 Cua Driver App。 两项能力目前都属于实验性功能,需要手动启用。Computer Use 还需要给实际运行它的应用授予屏幕读取和电脑操作权限。

DeepSeek Harness 官方桌面端已经基本做完。官方仓库主分支新增完整的 apps/desktop,用 Electron 把现有 Harness Web UI 封装成桌面 App。应用自带 Node.js 和 pnpm,用户不需要再手动启动 Web 服务,桌面端也不会对外监听端口。 官方已经准备好 macOS Apple Silicon、Intel 和 Windows x64 三套安装包构建流程,还包括 macOS 签名与公证、Windows EV 签名、自动更新和安装失败恢复。生产更新服务器已经指向 DeepSeek 自己的 暂时不在官方发布目标里。 最近一轮桌面端提交主要在处理 macOS 公证提速、启动恢复、打包校验等发布前问题。 DeepSeek 目前还没有公布具体发布日期,GitHub Release 里也还没有桌面安装包。但从仓库状态来看,桌面端已经进入发布前收尾阶段。

Anyone who’s built with AI knows the annoying part: You’re finally in flow, then you start rationing prompts because every small change eats into your usage. @boltdotnew’s new Forge mode is built for exactly this. It’s an experimental mode in AI app builder with up to 50x more usage for building full-stack web apps. I took one from a prompt to a live app, then kept iterating without staring at the meter. You can also opt into Forge to help train open models. The research preview runs from Sep 14 to Oct 14, 2026. It’s free on Pro plans until Oct 14, or $9/month through the Bolt Lite early-access plan.

@boltdotnew

Introducing Bolt Forge. Free until Oct 14th: > - Up to 50x more usage - The new frontier: GLM, DeepSeek, Kimi - Zero usage charges > Live now in your model picker on > And one more thing... 👇

truncated at source

腾讯混元等团队发布 EvolveScaler,专门测试大模型能不能跟上不断变化的信息。它会在长文本里不断加入修改、撤回、补录和作废的信息,再让模型根据最新状态回答问题。 比如先给模型看 40 天游戏记录,再问:「如果第 7 天没打那个小 Boss,最后还能不能打赢?」这会连带改变后面的装备、血量和战斗结果。模型得从第 7 天开始,把后面几十天重新推一遍,原文里没有现成答案。 团队设计了 117 类任务、159 种问题,最长约 1200 个事件。14 个模型配置包括 GPT-5.5、Gemini 3.1 Pro、DeepSeek V4 Preview Pro、GLM-5.2 和 Qwen3.5 Plus。最难档里,成绩中位数只有 11.3%;表现最好的 GPT-5.5-xhigh 也只有 59.3%。 这套数据也能拿来训练模型。一个内部 A3B 模型加入 6000 条这类数据后,在 8 个外部测试上全部提升,平均提高 5.25 分。

@TencentHunyuan

🚀 EvolveScaler is here. > Read a 40-day RPG log. Now answer one question: if you skip the mini-boss on Day 7, do you still beat the final boss? > The answer isn't in the log. You have to replay the world. > That's Information Evolution — records get retracted, corrected, backfilled. The world keeps changing after you read it. > So we build it backwards: define the world as an executable state machine, then render it into natural language. Code guarantees the logic. Language delivers the mess. > ➡️ 117 prototypes. 159 q

DeepSeek V4.1 Flash on max is ranked as the #3 open model on @arena. I'm using this model as my daily driver right now and am incredibly impressed. It's fast, it's smart and I have an enormous amount of context that remains coherent and accurate. The whale definitely cooked.

@YourLocalAILab

With today's release of Jovian Judgement R37, DeepSeek V4.1 Flash has full support with blazing speeds. > Decodes of 500 tok/s and prefills of 20k tok/s for single stream workloads on 4 RTX 6000 Pros. ⚡️ > To learn more point your agent at our GitHub here:

14 September 2026 3 items

Did A BIG Speed Run Update Last Night ! 🔓⚡️ DeepSeek V4.1-Flash on 3-4 DGX Sparks just went uncensored AND got faster thanks to @u1tra_instinct 🚀 ~190 tok/s throughput (6 streams, +13%) 💻 ~85 tok/s on code, single stream 📥 ~2,000 tok/s cold prefill (+38%) ⏱️ 0.22s to first token (-20%) 📚 1M-token context, proven 🧊 3.76M-token KV pool Abliterated attention, experts untouched. Full recipe's updated👇 🐋🤖

DEEPSEEK-HARNESS IS A FREE OPEN-SOURCE FRAMEWORK FOR BUILDING CODING AGENTS WITH SWAPPABLE MODELS, TOOLS, SANDBOXES, UIS AND AGENT LOOPS. IT WORKS WITH DEEPSEEK, CLAUDE, GPT, GEMINI AND MORE.

truncated at source

I trained an AI model on my phone through Telegram > Using OpenClaw running on Hugging Face infra: ML Claw > It beat DeepSeek V4 Pro on the given task while having 80 thousand times less parameters > That's right. The model, GoePT-1-20m is only 20 million parameters. It runs in the browser, on the CPU. And it beats DeepSeek V4 Pro, a 1.6 trillion parameter model, in AlmanBench > This concludes my 4 year old side project (fun fact, I created a dataset for this pre AI agents, using SpaCy, and paid 50 euros from my own pocket to rent 4090s on runpod to train a T5 variant, *years* before I joined Hugging Face. That first attempt was not very successful. This one is. My first ML adventure 🤗) > The goal of this side project was to show that you can *vibe* machine learning now, on platforms with tightly integrated GPUs, storage and compute, like Hugging Face > Including dataset creation, autoresearch and the final training run (hat tip to ML-Intern which g

@onusoz

I trained an AI model on my phone through Telegram > Using OpenClaw running on Hugging Face infra: ML Claw > It beat DeepSeek V4 Pro on the given task while having 80 thousand times less parameters > That's right. The model, GoePT-1-20m is only 20 million parameters. It runs in the browser, on the CPU. And it beats DeepSeek V4 Pro, a 1.6 trillion parameter model, in AlmanBench > This concludes my 4 year old side project (fun fact, I created a dataset for this pre AI agents, using SpaCy, and paid 50 euros from my own pocket to rent 4090s on runpod to train a T5 variant, *years* before I joined Hugging Face. That first attempt was not very successful. This one is. My first ML adventure 🤗) > The goal of this side project was to show that you can *vibe* machine learning now, on platforms with tightly integrated GPUs, storage and compute, like Hugging Face > Including dataset creation, autoresearch and the final training run (hat tip to ML-Intern which g

13 September 2026 5 items

💀💀💀 - Abliterated with Cybersecurity Function unleashed Thank everybody for waiting i released a universal Abliteration with unleashed Cybersecurity for: DeepSeek v4.1 -Flash that other can apply with Stock, EXL3 3.5BPW, and TR3 hybrid. this is based off our long lines of DSV4F Dspark preview, 0731, Vision-EXP work saving all the anchors, preserve all the MTP, Dspark, and vision. now GO GET IT!!!!!. don't forget donate or contribute for support HF Repo: (please read instructions on how to apply to your model)

DeepSeek Code 2.0 Leak: Coming Soon 🔥 >expected to beat Mythos 5.1 and GPT-6 Astra >September release window reportedly targeted >Computer use expected to be a top priority >Expected to remain open-weight >1M-token context window reportedly carried over >Rumored 3T+ parameter model

DeepSeek 4 pro killer model You can run RTX 3090 Locally. - 33B hybrid attention (not MoE) - 262K context - image + video - Apache-2.0 - BF16 - 66GB - or Wait for Q4/Q5 or SGLang on 48–96GB. - Architecture is hybrid delta-rule + attention, not a sparse MoE cheat. DS V4 Pro on AA (36) is the closed model. 3090, 4090 only after quants. -

Here’s a list of all AI models launched this month already: * Claude Fable 5.1 * Claude Mythos 5.1 * Gemini 3.8 Flash * Gemini 3.8 Flash Cyber * Muse Spark 1.3 * Qwen3.8-Max-0902 * GPT-6 Astra * GPT-6 Astra Pro * Ling-3.0-flash-VL * Ling-3.0-flash-Sante * Lyria 3.5 * ChatGPT Images 2.5 * GPT-Image-2.5 Flare * GPT-Image-2.5 Sunburst * DeepSeek-V4.1-Flash * Fugu Max * Fugu Ultra v2 We’re not even half way through the month…

艹,居然忘记放白嫖WorkBuddy+ DeepSeek v4.1 Flash的链接👇

@servasyy_ai

我的天,WorkBuddy 海外版 居然可以直接用 GPT6 Astra !? > 除了HY3还在继续免费使用外, 最新上线的 DeepSeek V4.1-Flash 居然可以限时免费~ DeepSeek V4.1-Flash 在 WorkBuddy国内版 可是要钱的,太香了! > 兄弟们,赶紧去冲吧 👇 >

12 September 2026 2 items

truncated at source

This is happening “the boon in using open models for cost savings” is real. Just four months ago we launched @CommandCodeAI coding agent built specifically for open models. 60T tokens scale and 43K paying active customers later, more than 90% of our usage is open models. In the first 24hrs of DeepSeek V4.1 launch, Command Code processed 3.1T paid tokens that’s 3x more than entire market on OpenRouter. 1T tokens on SOTA Claude/GPT cost you ~$5M 1T tokens on SOTA open models cost you ~$50K I’m not joking. I’m literally looking at our data. These numbers are real. That’s 100 times cheaper. While nearly as good. Run twice as many. I believe top ten open models collectively beat single SOTA Claude/GPT model. And when you discover a workflow that works for you, there’s literally nothing stopping you. You are not bound by subscription subsidies. Regular API prices are ten times or more manageable. In near future, enterprise will wake up to this. Better, faster, cheaper, private, and available now

truncated at source

2× DGX SPARK OWNERS REJOICE! DeepSeek-V4.1-Flash at 3.30 bpw EXL3, targeting just TWO DGX Sparks. 🔥 And yes, VISION is included. DeepSeek already ships its routed experts in FP4. I pushed that expert bank to a 3.30 bpw using my internal SAGE-EXL3 dynamic quant tool average while preserving the rest of the model, including Vision, Engram conditional memory, and the non-expert/source-format weights that are not part of the 3.30 bpw expert quant. TP4 was step one. Now TP2 is here. 𝗗𝗘𝗘𝗣𝗦𝗘𝗘𝗞-𝗩𝟰.𝟭 𝗙𝗟𝗔𝗦𝗛 𝗢𝗡 𝟮× 𝗦𝗣𝗔𝗥𝗞 SAGE-EXL3: 3.30 bpw routed-expert average Mixed precision: K2 → K8 Full pack: 415.8 GiB 31 shards 40 routed-expert layers Vision included. Engram preserved. Non-expert/source-format weights preserved outside the 3.30 bpw expert-bank average. SAGE is an internal quantization workflow I use so I’m not just applying one flat precision across every expert tensor. That’s about as much as I want to say about the method for now. 😁 𝗧𝗛𝗘 𝗧𝗣𝟮 𝗧𝗔𝗥𝗚𝗘𝗧 The full pack is ~415.8 GiB. Roughly ~189 GiB is the