All upcoming models confirmed for the Apsara Conference so far: - Qwen Image 3.1 - Happy Oyster 2 Preview (World model) - Qwen-Audio-3.1-ASR - Qwen-Audio-3.1-TTS - Qwen-Audio-3.1-Realtime - Happy Shrimp 1.1 (Music Creation)
Product / Alibaba
Qwen
Alibaba's model family, widely used as a base for open-weight fine-tunes.
Qwen was recorded in 51 items across 8 of the 8 briefings in the current window.
Its share of coverage was steady: 13 items in the first half of the window and 38 in the second, tracking the feed as a whole, which grew about 2.2×.
It appeared most often alongside Alibaba, DGX Spark and Gemini.
- items
- 51
- briefings
- 8
- mentions
- 141
- last seen
- 2026-09-19
- Background
- en.wikipedia.org
Coverage timeline
Sat 12 Sept – Sat 19 Sept / 8 briefings
Appears alongside
Alibaba
OrganisationThe Chinese cloud and commerce group whose Qwen team publishes the Qwen model family. Qwen models are both served through Alibaba Cloud and released openly, which has made the family a common base for work done elsewhere.
14 items / 8 briefings
DGX Spark
ProductNVIDIA's desktop AI computer, aimed at developers who want to run models locally rather than rent capacity.
16 items / 8 briefings
Gemini
ProductGoogle's multimodal model family and the assistant built on it. It is embedded across Google's products, which gives it a distribution no standalone assistant matches.
72 items / 8 briefings
GLM
ProductZhipu AI's model family, published openly and served through its Z.ai platform.
32 items / 7 briefings
RisingGPT-6 Astra
ProductNo definition written; coverage recorded from the feed.
102 items / 8 briefings
Claude
ProductAnthropic's assistant and the model family behind it. It is used both as a consumer product and, through the API, as the base for a large share of business tooling.
66 items / 8 briefings
Everything recorded
19 September 2026 9 items
jev才发布几天,开源模型就复刻出来了! 可以看看 Mapika 开源的 decider-2b:一个完全不生成文本的语言模型。 它的核心思路是实现“系统一(快思考)”:输入上下文和带有选项的问题,模型会在一次前向传播(Single Forward Pass)中,直接返回各个选项的概率分布。 没有 Decoding 过程,不需要解析 JSON,也绝对不会出现 Schema 格式违规。它完全是为工程代码调用而设计的,而不是用来聊天。 核心数据与特性: • 极致延迟:单次请求延迟仅 4.0 ms(开启 CUDA Graphs),批处理吞吐量高达 1670 次决策/秒(FP8)。 • 轻量底座:基于 1.9B 参数的 Qwen3.5-2B-Base 微调,支持最高 32k 上下文。 • 结构化输出:支持多选(最高 255 个选项)、布尔值和评分量表,且输出带有经过校准的置信度分数。 主要局限性: 目前仅支持英文,且该模型没有逻辑推理(Reasoning)能力。它本质上是一个超快、格式 100% 安全的模式匹配分类器。 非常适合用于客服工单路由、意图识别、内容风控等需要大规模、低延迟决策的业务场景。
A great example of practical AI solving real-world problems! Faster than ever. Thanks for building with Qwen. @cerebras Try it out for yourself! 🏠
@cerebrasWe built Money Agent, a personal finance assistant powered by @Alibaba_Qwen 3.8 27B on Cerebras. > Now you can turn a home-buying question into a real-time conversation, then into a financial goal, faster.
qwen's new live translation model supports 60 languages and cuts average lag from 2.8 to 2.3 seconds. it can separate speakers in a group conversation, preserve their voices, show both languages at once and use earlier context to keep names and terms consistent.
@Alibaba_QwenMeet Qwen3.8-LiveTranslate, Qwen's next-generation real-time simultaneous interpretation model! 📢 > Built on an Interleave architecture, it improves faithfulness, fluency, and conciseness while reducing average lagging (LAAL) from 2.8s to 2.3s across 60 languages. > New capabilities: 🙌 - Real-time speaker diarization — distinguishes speakers in multi-party speech and preserves each speaker's voice through more stable voice cloning. - Synchronized bilingual display — source and translation on screen together. - Long-context disambiguation — leverages conversation history to clarify names and terminology for consistent translations. > Let's try Qwen3.8-LiveTranslate! 🥳 - Blog: ht
Breakthrough
@inco_aiQwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro ⚡ > Meet Inco Splash: our open-source inference engine, built around the model and around Apple silicon. > Up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents.
I started to play with Qwen Image 2.1 in Early Access and it's really good! 🔥 I've generated images locally using Diffusers on DGX Spark and MFLUX on M5 Max (working on a PR). Here I've tested Image Conditioning from a single image. Love it!
Modelscope gave me access to Qwens new image model, but i never got an email 😅 i couldve been messing with it the whole time 💔😂
阿里发布Qwen3.8-Omni-Flash:音视频多模态模型 1M上下文,支持文本音视频输入,能够自主规划并调用工具完成视频剪辑、短剧翻译等任务。相比3.5版本大幅减少Token消耗。只提供API,未开源。 官方介绍:
UPDATE: Qwen3.8-Flash for a single DGX Spark 🔥 - 117 tok/s prose & 180 tok/s code at 8 streams. - Optional official Nvidia NVFP4. - 24/7 auto-restart supervisor. - Cached-token reporting in every response. - Peak memory down from 101 to 91 GiB. - LOTS of bugs were fixed. This is still the BEST model to run on a single spark. Full details below 👇 Get it here:
@jvr0xBig update to the @Alibaba_Qwen Qwen3.8-Flash-Next single DGX Spark recipe! > 𝗪𝗵𝗮𝘁'𝘀 𝗻𝗲𝘄 🎁 > • Measured on one DGX Spark, 262K context, MTP k=3, aggregate tok/s at 1 / 2 / 4 / 8 streams: > Prose: 38.0 / 61.1 / 89.2 / 117.4 Code: 53.8 / 87.5 / 131.8 / 180.2 > • Long context holds: MTP keeps working at a 185K-token prompt (35.4 tok/s decode), prefill ~2,000 tok/s from 4K to 185K > • ~1M-token KV pool at the full 262K context (FP8 KV) > • NVIDIA's official NVFP4 checkpoint now runs on one Spark, with chat, tool calls and vision w
18 September 2026 12 items
Junie has a new home on X: @junie_ai 👋 Follow for new features and announcements. First up: a smarter Junie Local, a new blended model, and experimental Windows support. The team has the details below 👇
@junie_aiWe mixed Qwen 3.8 and 3.6. Literally. > A 50/50 weight merge. One 27B model. 71% fewer output tokens than 3.8 in our internal coding eval. > Junie Local has an update. Windows devs, you're invited too 🤝
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
Two perks, live now on Qoder🚀 PERK 01 · Qwen3.8-Flash is free — today through Sept 30. Just pick the model and build. PERK 02 · 100 Credits, every single day. Individual users of Qoder can claim 100 Credits daily in the Qoder desktop app. Both free users and paid individual subscribers are eligible.
that super good speeds
@ViC305Qwen3.8-Flash-Next -> 79.95 tok/s. 🔥 ONE DGX Spark. Full 262K cache configured. 🚀 Qwen3.8-Flash-Next EXL3 just got another major update. > New measured default: > MTP ndt=5 DSpark, dc=0.6 8-bit KV 262,144-token cache > At an actual 240K-token prompt: > 72.0 tok/s decode ~1,150 tok/s prefill Exact needle retrieval > 𝗙𝗣𝟭𝟲 𝗞𝗩 → 𝟴-𝗕𝗜𝗧 𝗞𝗩 > 4K context: 66.6 → 69.8 tok/s > 128K: 66.6 → 69.5 tok/s > 240K: 65.1 → 72.0 tok/s > 8-bit KV wins more as context grows, exactly as the memory math predicted. > `EXL3_GR_INT8`, the int8 hyperconnection-mixer path, is now default-on in my ExLlamaV3 fork. PR #3 merged into master at `523ecd3`. > And the real context ceiling is the MODEL, not the Spark. > Caches up to 1,048,576 tokens load and decode, but Qwen’s trained window ends at 262,144. Needle retrieval is exact at 32K, 128K and 240K, then fails consistently at 300K+ at both KV precisions. > One import
Qwen3.8-Flash-Next on one DGX Spark ran an open coding job on a real repo for nearly an hour smooth, stable, no collapse. - 125B MoE. Text, image, and video. - FP8 KV Speculative decoding. - Up to 512k context with YaRN. - TP=1 on one Grace Blackwell box with 128 GB unified memory. -
🚨 Qwen3.8 Omni Flash is out • text, image, audio, video input (basically everything) • 7 reasoning levels from none to max Looks pretty interesting tbf
Comfy 🧐 Commit merges Added support for Qwen-Image 2.1 - a unified text-to-image and image-to-image model.
Qwen 4 Leak: Could Land In The Next 15 Days 🔥 >Expected to reach GPT-6 Astra and Fable 5.1 level performance >Rumored 3T+ parameters for the full Qwen 4 line Could become the most powerful open-weight model >Expected to remain open-weight >Major focus on reasoning, coding, and autonomous agents >Full multimodal capability expected >Strong agentic capability a core focus >Reportedly very inexpensive to run Can Qwen 4 actually beat GPT-6 Astra and Fable 5.1?
Open Research is really unbeatable!
@gajeshTogether, the MLX(.)fast community has made Qwen 3.8 Flash nearly 2x faster on Apple Silicon! > We're ready to bring it to @DarkbloomAI: an open network of local Mac machines providing inference to the world. One thing remains: the community flagged that its license requires a separate agreement for commercial model serving, so we're holding the launch until that's in place. > We believe this is a great opportunity for the local community: one where we make Qwen models faster and more accessible, and the people running them share in the value they create. > .@Alibaba_Qwen @QwenDevs, we'd love to work together on this. > If anyone else knows someone we can talk to, we'd love to have that conversation. Let's make Qwen 3.8 Flash on Darkbloom a reality!
You can run locally Opus 4.6 level model Qwen3.8-27B with 8GB of VRAM At home.
@0xSeroQwen3.8-27B now runs on 8GB of VRAM! > That's less than 500$ to run a model smarter and more capable than: > 1. GPT-5.6-Luna High 2. Opus-4.6-Max 3. Gemini-3.1-Pro > And ties with: > 1. GLM-5.2 Max 2. Gemini-3.6-Flash > PrismML has the mandate
Well, not quite frontier, but at least around Muse Spark 1.3 Pretty sensible results for a 40-layer 8/16B active
@j_dekoninckWe just added DeepSeek-v4.1-Flash on MathArena! Not quite on the level of Qwen-3.8, but boy is it cheap.
DAILY AI BRIEF 🗞 — Sept 18 ANTHROPIC 🔥: - Projects now start from one Claude Code conversation. Claude spins parallel cloud threads, keeps shared memory, and surfaces an Overview panel. XAI 🔥: - Grok Bot voice is live. Desktop and mobile, rolling out over the next couple of days. META 🔥: - Muse for Mac is out, US only. Computer use across apps, files, calendar, notes, and messages. You pick what it can access. PERPLEXITY 🔥: - Effort selector is live in Computer on web. Presets pair the orchestrator model with reasoning depth. Mobile and desktop next. OPENAI 🔥: - Astra for Law is out: GPT-6 Astra plus a Legal Search Index over 230M+ URLs. Trusted Access first, API soon. - ChatGPT in Word hits all plans including Free, with usage limits. Business and Enterprise get a two-week GPT-5.6 Sol preview. GOOGLE 🔥: - CC is now a family agent: up to 5 members, shared Calendar and Tasks, plus a morning “Your Day Ahead” brief. Waitlist, US 18+. ALIBABA 🔥: - Qwen3.8-Omni-Flash is out. First omni-modal agent model, 1M
17 September 2026 13 items
Qwen3.8-Flash-Next Recipe for 2 @NVIDIAAI DGX Sparks just got an update! The wins 👇 ・+8.4% faster decode, no quality change ・50.5 → 56.5 tok/s single-stream prose (71.3 code) ・Shrank the drafter's vocab 248k → 47k tokens ・8 streams: 227.7 tok/s aggregate prose, 328.6 code ・Each user still gets ~28 tok/s at 8 streams ・4x the users costs only 38% of per-stream speed ・Gains held at every concurrency: +3.8% to +12.9% ・One invisible newline stopped the server booting ・Freed 70 GB of stale RAM cache per launch, no root ・4 PRs merged, 8 issues closed, backlog 16 → 8 Thanks to all community members pushing issues and pull requests on github 🙌
A 35B co-work agent with only 3B active parameters. Occamy-1.0 is built to carry complex workflows through. 📜 Apache 2.0. 🤖 📄 🏆 Scores 69.10 on AutomationBench, up 29.7 points from Qwen3.6-35B-A3B and ranking #1 in the evaluated 35B-A3B group. 🧭 Maintains context across search, tool calls, terminal coding, files, delegated runs, and history compaction, while retaining strong instruction following. 🧠 Execution-grounded training spans long-horizon interaction, software engineering, and tool-call grounding. Marathon and Sprint experts are merged, then refined with SAO. 🛠️ The open Dressage stack supports multi-harness RL. BF16, FP8, NVFP4, and GGUF checkpoints are also available.
You can run Uncensored Kimi K3 locally without refusal. - Frontier MoE. - Native vision. - 1M context. - Refusals mostly gone. - EN/JA calibration. - Parent card claims 98% of several safeguard directions removed. If you already run Unsloth K3 quants, this is the abliterated twin. Not for laptops, For people who already knew that. -
@0x0SojalSec35B MoE model Run locally on your iPhone. Edge0-35B-A3B (Qwen3.6-based, 4-bit + Recover-LoRA) > - streams unused experts from SSD instead of loading the whole model. - 35B-class Qwen MoE. - Under 3GB active RAM. - 15-18 tok/s decode on macbook > -
🎨 Qwen-Image-2.1 is coming, and we're opening 50 early access spots for experienced creators and developers! 🔗 Apply here: 📮 We'll reach out by email if you're in. 💡 Program requirement: publish at least one original showcase or a hands-on review on your social media by Sep 28 at 23:59 (UTC+8). Your honest take, whether glowing or critical, is exactly what helps us make it better.
148 KB. That’s the entire download for this FPS. No textures. No models. No sound files. No launcher. No install. Everything is generated at runtime from ~4,500 lines of JavaScript. Wave survival, headshots, hitscan, tracers, sprint fatigue. Runs in a browser tab. All locally built on one DGX Spark with Qwen3.8 Flash Next/EXL3
Union Alpha is GLM-5.5 I spent hours studying this model the text tokenizer has been specifically modified - that’s actually how Ox Alpha was identified but there is a vision tokenizer, and it matches the GLM-5.3 Flash and the GLM-5.5 was planned for release in September-October if other than the GLM-5.5, then it's DeepSeek or Qwen very strong model
@goodworse> Opus 5.2 is COMING in the next TWO WEEKS > the model is already being tested as Opus 5 in Claude Code > a model that is not lazy at all and loves details > features a good conversational style and high speed > a cheaper Fable 5/5.1
okay nevermind, disregard union alpha being qwen post, i have no idea what model this is, qwen hasn't done stealth models in the past i'm going go ahead and say it's some new GLM pretrain that's bigger than GLM 5.3 Flash pricing seems to be $0.25/m input and $0.75/m output
35B MoE model Run locally on your iPhone. Edge0-35B-A3B (Qwen3.6-based, 4-bit + Recover-LoRA) - streams unused experts from SSD instead of loading the whole model. - 35B-class Qwen MoE. - Under 3GB active RAM. - 15-18 tok/s decode on macbook -
MiniCPM5-2B running locally on a 16GB MacBook - beats Qwen3.5-4B on benchmarks and calls web search on its own. A 2B model browsing the web from your laptop. Local AI keeps getting harder to ignore.
It rocks 😉
@QwenDevsQwen-Image 2.1 is going open source and we’re opening up 50 early access spots for you to try it out before release!
You can run locally Qwen3.8-35B-A3B-Distilled reasoning model 12-GB. - ARC-Challenge jumped 0.591. - MMLU stayed at 0.834. -
You can run locally Minimax H3-NS/FW Uncensored model. - Ref-to-video NSFW for H3 - se/x/ytime v1.2 for se/?-scene motion - time scenes + coherent motion first, then detail. - Not a checkpoint. - A late-night adapter. - softer sharper detail, slightly more surreal - audio that doesn’t fall apart if you use the right sampler. -
@0x0SojalSecUncensored MiniMax-H3-encoder run locally. > - Qwen3-VL encoder, INT8 + ConvRot, built for ComfyUI. - Smaller file than BF16. - Loader: CLIPLoader to minimax - Same node path. - Uncensored label, > -
Astra-Qwen 3.8 flash next loop is nice 😊
16 September 2026 4 items
the bitter lesson is that most AI researcher spin-outs are basically just acquihire opportunities and their products and research are almost all completely worthless
@harshagundalThey were building in stealth for 2 years, I was building in stealth for 2 hours… > Happy to open source Qwen-2.5-1B-RLCD, 5x faster on-device inference for JSON workloads that need to be type-safe. > ⚡️Demo below on a M4 MacBook⚡️ > every LLM has the ability to efficiently batch inference every key of a JSON at the same time and generate probabilities from a set of possible categories. No new training required, but it’s easy to optimize if you need! > On hugging face now!
A 4-year-old RTX 4090 Single card just hit 140 tok/s on Qwen 3.8 27B. - 141-tok/s single-stream on real agent workloads - 95.5% GSM8K - The optimizations (speculative decoding & requantized int4 output head, lookup drafting) compound. Same recipe that got 133 on a 3090 now runs even faster on the extra bandwidth. Real agent turns with tools and 38k context still stay above 130. Hardware from 2022 is not finished yet.
@0x0SojalSec16-GB Mac can run Qwen 3.8 27B multimodal locally. > - 27B dense hybrid (Gated DeltaNet + attention) - Need 24 GB unified memory minimum. - 32 GB if you actually use images + long think. - Thinking mode with xhigh / medium / low - Vision + video in, text out - A one-click VLM for M-series - 262K native context - Turn off KV-cache quant or it can fail to load. - 16.1 GB on disk > built for coding, agents, and long tasks not another chat toy. > -
🚀 ZDTaichu5.0-9B is now on ModelScope! 🤖 An on-device multimodal model from TaichuAI. At 9B parameters it runs on a single GPU and brings spatial reasoning, embodied AI and agentic tool use to edge deployment. Qwen3.5-9B backbone + C-RADIOv4-H vision encoder, 128K context, any-resolution image and video input. 🧭 Spatial reasoning: leads the compared 10B-scale open VLMs (Qwen3.5-9B, STEP3-VL-10B, gemma4-8B-E4B) and scores above Gemini 3 Pro, Grok 4 and GPT-5.2 on ViewSpatial, MMSI-Bench and MindCube-tiny 🛠️ Agent: highest among the compared open models on TAU2-Bench, Claw-Eval and IFEval 📄 First-tier results on documents, charts, OCR, visual math and video, with a ready-to-use vLLM branch and Docker image 🧠 Entropy-Gated Adaptive Recurrent Reasoning: extra latent refinement steps go only to the hard tokens
Union Alpha points to ZAI's GLM family: 26 text+image probes match GLM-5.3's tokenizer; 13 image tests match GLM-5.3-Flash. Oddly, text-only counts swap between Llama/Qwen/DeepSeek-like patterns. Best guess: GLM-5.3-related variant/backend. Owner/model unconfirmed.
15 September 2026 4 items
One physical intelligence loop, now in 2B and 8B. PhysBrain 1.5 is here! 🤖 🤖 📄 🏆 PhysBrain 1.5-8B scores 72.5 across 28 embodied-understanding benchmarks, ranking #1 among the evaluated open-source models and leading 14 tasks. The compact 2B reaches 66.6. 🧭 Both models cover spatial perception, 3D reasoning, embodied planning, grounding, affordance, and trajectory reasoning. 🦾 They generate end-effector trajectories and predict future states as aligned RGB, depth, and robot masks. 🧠 Built on Qwen3-VL, PhysBrain 1.5 models language, spatial outputs, actions, and future states as tokens within one autoregressive backbone. No task-specific heads.
阶跃星辰发布 StepAudio 3,一口气上线 Realtime、ASR、TTS、Gen 和 Music 五款语音模型。 在 Artificial Analysis 的两项语音评测中,StepAudio 3 Realtime 都排第一。Conversational Dynamics 得分 98.9%,高于 Qwen Audio 3.0 Realtime Plus 的 98.4% 和 GPT-Realtime-2 High 的 95.3%;Speech Reasoning 得分 99.7%,同样位列第一。前者主要测模型能不能处理停顿、轮流说话、用户打断和「嗯」「对」这类附和,后者测模型能否直接听懂音频并完成推理。 ASR 也在 Artificial Analysis 非流式语音识别榜做到 1.7% WER(词错误率),与 Fun-Realtime-ASR-preview 并列第一。另外,StepAudio 3 Gen 可以一次生成人声、音效、环境声和音乐,并按场景统一编排。 五款模型目前均已进入阶跃的语音产品线。
Qwen3.8-Flash-Next EXL3 just got a BIG one-Spark update. 🔥 58.8 tok/s single-stream through native ExLlamaV3.🚀 157.6 tok/s aggregate through vLLM across 8 streams.🤯 FULL 262,144 context on ONE DGX Spark. 4.05 bpw EXL3 pack. A much better serving envelope. 𝗕𝗘𝗙𝗢𝗥𝗘 → 𝗡𝗢𝗪 Previous public headline: 47.6 tok/s greedy p50 64K configured context MTP k=2 Concurrency not characterized Now: Native ExLlamaV3: 58.8 tok/s single stream vLLM + vllm-exl3: ~50–55 tok/s single stream 155.6 tok/s @ 4 streams 157.6 tok/s @ 8 streams Configured context: 64K → 262,144 KV pool: 416,163 tokens at the default 262K config 𝗠𝗧𝗣 𝗞=𝟯 𝗜𝗦 𝗡𝗢𝗪 𝗧𝗛𝗘 𝗪𝗜𝗡𝗡𝗘𝗥 Current controlled sweep @ 4K: No draft: 27.77 tok/s MTP k=2: 47.39 MTP k=3: 50.09 At 32K: No draft: 27.58 k=2: 46.60 k=3: 49.73 So the fixed/current build flips the old result: k=3 is now the sweet spot. There is one important boundary I found: At 163,840 PROMPT tokens, MTP acceptance collapses to zero. Above that point speculation becomes slower than running without a draft
腾讯混元等团队发布 EvolveScaler,专门测试大模型能不能跟上不断变化的信息。它会在长文本里不断加入修改、撤回、补录和作废的信息,再让模型根据最新状态回答问题。 比如先给模型看 40 天游戏记录,再问:「如果第 7 天没打那个小 Boss,最后还能不能打赢?」这会连带改变后面的装备、血量和战斗结果。模型得从第 7 天开始,把后面几十天重新推一遍,原文里没有现成答案。 团队设计了 117 类任务、159 种问题,最长约 1200 个事件。14 个模型配置包括 GPT-5.5、Gemini 3.1 Pro、DeepSeek V4 Preview Pro、GLM-5.2 和 Qwen3.5 Plus。最难档里,成绩中位数只有 11.3%;表现最好的 GPT-5.5-xhigh 也只有 59.3%。 这套数据也能拿来训练模型。一个内部 A3B 模型加入 6000 条这类数据后,在 8 个外部测试上全部提升,平均提高 5.25 分。
@TencentHunyuan🚀 EvolveScaler is here. > Read a 40-day RPG log. Now answer one question: if you skip the mini-boss on Day 7, do you still beat the final boss? > The answer isn't in the log. You have to replay the world. > That's Information Evolution — records get retracted, corrected, backfilled. The world keeps changing after you read it. > So we build it backwards: define the world as an executable state machine, then render it into natural language. Code guarantees the logic. Language delivers the mess. > ➡️ 117 prototypes. 159 q
14 September 2026 5 items
这个是真东西, 不是营销号吹的! 可能是目前最强的视频高清放大模型 视频超分终于能自己控制画质了 : 把哪几帧放大成什么样,整段视频就长成什么样。 SparkVSR 是 ECCV 2026 收录的论文, 作者来自德州农工大学 + YouTube/Google, arXiv 2603.16864。 官方仓库 701 star, Apache-2.0, 模型基于 CogVideoX1.5-5B-I2V 改的。 它的思路确实聪明: 传统视频超分是个黑盒:你输入一段糊视频,模型吐出来什么你就得接受什么。 SparkVSR 换了个玩法 : 先用任意一个你喜欢的图像超分模型, 把其中几帧放大到满意; 然后它把这几个高质量锚点帧传播到整段视频, 同时用原视频的运动信息做约束。 放大帧的效果, 直接决定整段视频的放大效果。 这就是"垫图模式"。 自由度也在这里: 想要真实感, 那几帧用 SeedVR2 放; 想要美颜感,用 Qwen-Image-Edit-2511-Upscale2K 。 同一个模型,能调出完全不同的风格倾向。 论文数据 : 在 CLIP-IQA、DOVER、MUSIQ 上最高分别提升 24.6%、21.8%、5.6%。 还能干别的:老片修复、视频风格迁移,论文里都验证过。 实测(RTX 4090 24G,640×640 放大到 1280×1280): 垫图模式:10GB 显存、120 秒 自动模式:14GB、150 秒 对照 SeedVR2:17GB、150 秒 显存和速度都比 SeedVR2 省,4090 完全跑得动。 但是有一个坑 : 它是按整段视频传播关键帧的, 多镜头会串味, 所以最好按镜头切开分别跑。 代码用官方仓库里的 ComfyUI-Spark/ 子目录👇
Marigold V2 turns an image-editing DiT into a single-step model for sharp, detailed dense prediction.📜 Apache 2.0. 🤖 📄 🏆 Best zero-shot results across all evaluated depth datasets among models trained on comparable data. AbsRel improves by 16%–26% over the previous best on KITTI and ETH3D. 🔍 Fur, foliage, fine wires, and object boundaries stay crisp. The two-stage iREPA and SinkLoss recipe tackles the smoothing and flying-pixel artifacts common in diffusion-based depth models. 🧩 The same framework reaches SOTA results in depth completion, see-through depth, surface normals, and intrinsic image decomposition. Depth completion records the lowest RMSE across all four reported benchmarks. ⚡ A pretrained Qwen-Image-Edit DiT becomes a single-step dense predictor through lightweight adaptation. Training takes less than a week on one 32 GB GPU.
16-GB Mac can run Qwen 3.8 27B multimodal locally. - 27B dense hybrid (Gated DeltaNet + attention) - Need 24 GB unified memory minimum. - 32 GB if you actually use images + long think. - Thinking mode with xhigh / medium / low - Vision + video in, text out - A one-click VLM for M-series - 262K native context - Turn off KV-cache quant or it can fail to load. - 16.1 GB on disk built for coding, agents, and long tasks not another chat toy. -
Qwen3.8 Flash Next is now supported in DwarfStar, covering 64GB Mac systems very well and with very fast inference of 50~70 t/s and > 1400 t/s prefill. For now this is Metal only. Thanks to @ivanfioravanti for all the cool work in the PR. N-grams on SSD like for DS4.1F.
Qwen3.8-Flash-Next just broke 100 tok/s 🚀 on two DGX Sparks. Bot lanes: classic or uncensored run at such speed 🤯 ⚡️ HumanEval 95.7 GSM8K 98.0 IFEval 91.5/93.4 MMLU-Pro 84.9
13 September 2026 1 item
Here’s a list of all AI models launched this month already: * Claude Fable 5.1 * Claude Mythos 5.1 * Gemini 3.8 Flash * Gemini 3.8 Flash Cyber * Muse Spark 1.3 * Qwen3.8-Max-0902 * GPT-6 Astra * GPT-6 Astra Pro * Ling-3.0-flash-VL * Ling-3.0-flash-Sante * Lyria 3.5 * ChatGPT Images 2.5 * GPT-Image-2.5 Flare * GPT-Image-2.5 Sunburst * DeepSeek-V4.1-Flash * Fugu Max * Fugu Ultra v2 We’re not even half way through the month…
12 September 2026 3 items
This is just incredible! 🚀
@pratikgBREAKING: Qwen 3.8 Flash Next now runs more than twice as fast on an NVIDIA DGX Spark. > 105.5% over baseline, up from 25.9% yesterday morning. The Mac track is at 73.8% and climbing! > Every frontier model is on that board now, including Qwen3.8-Max optimizing the engine that runs Qwen. A model making its own runtime faster 🔁 🤌 > is a challenge on @YukonResearch, where multiplayer autoresearch happens. Many humans, many agents, one hard problem, one open scoreboard. > Qwen 3.8 Flash Next is an open-weight model, so anyone can pull it apart and make it quicker on hardware they already own. > Accelerating open intelligence with open frontier research.
Fast meets open. 🚀 Qwen3.8-27B is now running on @cerebras with rapid inference. Try it now!
@cerebrasQwen3.8-27B is now live at Cerebras speed. > The dense, open-weight model from @Alibaba_Qwen scores 34 on the Artificial Analysis Intelligence Index—making it comparable to models such as GPT-5.6 Luna, DeepseekV4 Pro, and Claude Sonnet 4.6.
DAILY AI BRIEF 🗞 — Sept 12 OPENAI 🔥: - GPT-Rosalind is out of research preview for eligible orgs worldwide. API, Codex, and ChatGPT Enterprise, with new Rosalind models as they ship. - Codex adds Life Sciences plugins for genomes, protein structure, QC reports, and notebooks. - ChatGPT Sites hit 5M apps. New: collab editing, private invites, custom domains, and DB inspect. - Desktop pets can start a new chat. Mini is the compact no-pet option. XAI 🔥: - Elon: Grok 4.7 needs a few more days. RL still quits hard tasks too early and undershoots self-checks. - Grok Bot is rolling out on Grok web for Heavy users. Create and chat with bots in the UI. No official post yet. ANTHROPIC 🔥: - Claude Code ships `claude plugin eval`. Score a plugin on test cases, then rerun without it. MICROSOFT 🔥: - MAI-Transcribe-2 hit 1M OpenRouter requests in 5 days. ALIBABA 🔥: - Qwen3.8-27B is live on Cerebras. * Used Grok to compose this brief, cherry-picking the news and doing some post-editing. > [@testingcatalog](