Skip to content
B Bloger.fm

Product / Zhipu AI

GLM

Zhipu AI's model family, published openly and served through its Z.ai platform.

GLM was recorded in 32 items across 7 of the 8 briefings in the current window.

It gained ground: 5 items in the first half of the window and 27 in the second, a bigger increase than the feed as a whole, which grew about 2.2×.

It appeared most often alongside DeepSeek, Kimi and Qwen.

Gaining ground
items
32
briefings
7
mentions
110
last seen
2026-09-19

Coverage timeline

Sat 12 Sept – Sat 19 Sept / 8 briefings

Everything recorded

19 September 2026 4 items

GLM-5.3 FlashX is now live in Command Code. · High speed variant of GLM-5.3 Flash · ~200 TPS with 1M context window · Multimodal (text, image, video) Available across all plans and API.

Just added two open-weight models to your toolbelt: → GLM 5.3 → GLM 5.3 Flash Now available in Personal and Custom Agents.

Gave GLM 5.3 on 2 sparks a serious test today. Holy shit this model is incredible. @MiaAI_lab updates made it serious usable for me

GLM 5.3 FlashX, coming soon on GMI

18 September 2026 7 items

⚡ GLM-5.3 FlashX is now live on ZenMux. A faster, more efficient GLM-5.3 model for coding and agent workflows. 👀 Native multimodal understanding 💻 Visual coding and software engineering 🧠 Built for long-horizon tasks Curious about GLM-5.3 Flash vs FlashX? Try both on ZenMux. 👉 One API. More models to build with.

truncated at source

Have been using GLM 5.3 Flash yesterday and today on 2 sparks and it has got significantly better. And now @MiaAI_lab is shipping more improvements yet again 🔥 Local AI will win.

@MiaAI_lab

Yet another big update for GLM 5.3 Flash EXL3 ⚡️ > Performance improvements: > On 2x DGX Sparks ~13% faster on prose single stream ~23% faster on prose 2-4 concurrent streams ~36-37 tok/s on prose single stream ~75-67 tok/s on prose in 4 concurrent streams > On 3x DGX Sparks ~15% faster on prose single stream ~23% faster on prose 2-4 concurrent streams ~41 tok/s on prose single stream ~88 tok/s on prose in 4 concurrent streams > No change for TP=4. > New features and fixes for all: - Much faster loading time for the weights! From more than 300s to now 55s. - Fixed an issue that after a long conversation, the next message was sometimes treated like a brand-new prompt and had to be re-read from scratch. > Get it here:

GLM 5.3 FlashX (likely Omen Alpha) is live on the Zai API and Vercel AI Gateway.

I became bored after uncovering all the stealth models so I went back to Omen Alpha. Today's API errors name Baseten, and image-token counts match GLM-5.3-Flash (+46-token wrapper). Likely Baseten-served Flash or a close variant. Exact checkpoint unconfirmed.

I HAVE ALWAYS SAID THE NARRATIVE AND TIMELINE IS MOVING FROM USA LABS TO CHINESE LABS this is another proof of that look at the numbers Bolt shared from the first few days of open models on Forge - 1. GLM 5.3 Flash 54% 2. DeepSeek V4 Pro 17% 3. GLM 5.3 15% 4. Kimi K3 14% Chinese labs are moving fucking fast rn

@boltdotnew

11M+ people build on Bolt. On Monday, we gave them open models with up to 50x usage. > Which is #1 so far? The fastest, cheapest one, with prompts up to 2x the size. > 🥇 GLM 5.3 Flash 54% 🥈 DeepSeek V4 Pro 17% 🥉 GLM 5.3 15% 👏 Kimi K3 14% >

Open models are proving they can lead, not just fill the gaps. Real builders are putting them to work, and the numbers show it. GLM 5.3 Flash captured 54% of Forge prompts through Wednesday, with Flash getting up to 50x usage.

@boltdotnew

11M+ people build on Bolt. On Monday, we gave them open models with up to 50x usage. > Which is #1 so far? The fastest, cheapest one, with prompts up to 2x the size. > 🥇 GLM 5.3 Flash 54% 🥈 DeepSeek V4 Pro 17% 🥉 GLM 5.3 15% 👏 Kimi K3 14% >

You can run locally Opus 4.6 level model Qwen3.8-27B with 8GB of VRAM At home.

@0xSero

Qwen3.8-27B now runs on 8GB of VRAM! > That's less than 500$ to run a model smarter and more capable than: > 1. GPT-5.6-Luna High 2. Opus-4.6-Max 3. Gemini-3.1-Pro > And ties with: > 1. GLM-5.2 Max 2. Gemini-3.6-Flash > PrismML has the mandate

17 September 2026 8 items

Union Alpha could be ZLM 5.4 👀 >A new anonymous stealth model called Union Alpha has surfaced, free to use >256K context window, multimodal, built for agentic coding >Frontier-level general-purpose performance with tool calling built in >Performance reportedly sits near GPT-6 Astra and Opus 5 at roughly 18x lower expected cost >Naming pattern echoes Ox Alpha, which turned out to be GLM-5.3-Flash speculation is already pointing to an early GLM-5.4 test No lab has claimed it yet unconfirmed as of now Try it yourself and see if you can spot who's really behind it.

truncated at source

Two weeks. That's how long it took to go from GLM-5.3-Flash's first run on domestic accelerators to serving all of its production traffic, with 3.2× end-to-end throughput along the way. What I keep thinking about is who did much of the work: an Infra Agent powered by GLM-5.3. A model helping optimize the system that serves it. The conditions were hard. Limited memory and interconnect bandwidth. 1M-token context. Multimodal requests. An immature software stack where kernels were missing and documentation was often guesswork. Every optimization was a trade: compute for memory (ReplaySSM), communication for memory (intra-node tensor parallelism), precision for capacity (mixed INT8/FP8/BF16 caching), and disaggregation for scheduling freedom (Encode–Prefill–Decode). But the most important lesson wasn't about any single optimization. When the agent got stuck, it was rarely because it couldn't write the code. It was because it didn't know *why* things got worse. "Throughput down 20%" tells you something brok

Union Alpha is GLM-5.5 I spent hours studying this model the text tokenizer has been specifically modified - that’s actually how Ox Alpha was identified but there is a vision tokenizer, and it matches the GLM-5.3 Flash and the GLM-5.5 was planned for release in September-October if other than the GLM-5.5, then it's DeepSeek or Qwen very strong model

@goodworse

> Opus 5.2 is COMING in the next TWO WEEKS > the model is already being tested as Opus 5 in Claude Code > a model that is not lazy at all and loves details > features a good conversational style and high speed > a cheaper Fable 5/5.1

okay nevermind, disregard union alpha being qwen post, i have no idea what model this is, qwen hasn't done stealth models in the past i'm going go ahead and say it's some new GLM pretrain that's bigger than GLM 5.3 Flash pricing seems to be $0.25/m input and $0.75/m output

ZCode now supports more model providers, with improved stability and performance. We’ll keep expanding integrations based on your feedback. Which models do you like most beyond the GLM series?

Union Alpha is now in Codex!! The speed is insane over 300 to 400 tokens a second, 262k context, images in, and it's free for a week. It's a good timing too, most of the Codex usage is basically gone right now. Zai did this exact thing in August: Ox Alpha showed up unnamed, free for a week, built for agentic coding, and a week later it was GLM-5.3-Flash. So what model do you guys think it is?

@opencode

Union Alpha (stealth model) is free for the next week > - no data training - built for agentic coding - supports images > let's see what you can do

truncated at source

What happens when your AI model becomes good enough to build its own infrastructure? Zhipu just found out. Their AI model GLM-5.3 helped build and optimize the inference system that serves GLM-5.3-Flash to users. The model improving the system that runs the model. 100,000+ Chinese-made AI accelerators. Nobody had deployed at this scale on that hardware before. Limited memory. Incomplete ecosystem. Most of it undocumented. The Infra Agent powered by GLM-5.3 did the engineering work. Found bugs in kernels. Fixed concurrency bottlenecks. Studied optimization patterns from other codebases and applied them to its own inference. First successful run to production ready in two weeks. Throughput tripled. Then it went live anonymously as "Ox-Alpha" on OpenCode and OpenRouter. Became the most-used model on both platforms in six days. 62 trillion tokens processed. Zhipu's own words: "The model optimizes the system. The system runs the model." They added: "We have not yet reached full recursive self-improvement. B

Stealth models in 2026: Hunter Alpha → Xiaomi MiMo-V2 Owl Alpha → Meituan LongCat Pony Alpha → GLM-5 Ox Alpha → GLM-5.3-Flash Now Union Alpha is free for a week on OpenCode and OpenRouter.

16 September 2026 8 items

GLM Coding 2.0 Leak: Coming Soon 🔥 >Expected to beat Mythos 5.1 and GPT-6 Astra >October release window reportedly targeted >Computer use expected to be a top priority >Expected to remain open-weight >1M-token context window reportedly carried over Rumored 3T+ parameter model

My favorite local AI model is getting better and better. GLM 5.3 Flash EXL3 on 2x DGX Sparks is now more stable, reliable, and easier to debug. More awesome updates are incoming!

@plotarmordev

GLM 5.3 Flash on 2x DGX Sparks just got multiple updates focusing on reliability: safer startup checks, bounded output defaults, better cache controls and clearer diagnostics. > Thanks to the authors of 15 community PRs, and it also includes my restart-verification fix too👇

ZCode is the strongest harness for GLM-5.3 so far — delivering 82.2% success at ~$1.98 per pass. Try it here:

@ZixuanLi_

Added ZCode with GLM-5.3 and GLM-5.3-Flash, building on FrontierHarness and @LotusDecoder’s work. > ZCode is the strongest harness for GLM-5.3 so far, and cheaper than the second-place Claude Code + GLM-5.3 combo. > The task set is small, so we ran each combo three times to reduce variance. Passes out of 30: - GLM-5.3: 26 / 22 / 26 - GLM-5.3-Flash: 24 / 21 / 23

The Forge gates are open. You all came running 🏃 We’re opening up access as fast as we can, with another wave coming soon. Join the Bolt Lite waitlist, our $9/month plan with Forge access → Or skip the wait entirely. Forge is live on all Pro plans.

@boltdotnew

Introducing Bolt Forge. Free until Oct 14th: > - Up to 50x more usage - The new frontier: GLM, DeepSeek, Kimi - Zero usage charges > Live now in your model picker on > And one more thing... 👇

THIS FREE OPEN-SOURCE REPO LETS YOU RUN GLM-5.3 FLASH, DEEPSEEK V4 FLASH AND KIMI K3 LOCALLY WITHOUT A GPU OR HOSTED TOKEN QUOTAS. THE TRADEOFF: YOU’LL NEED A LOT OF STORAGE.

truncated at source

Hugging Face 已封禁 AI 安全公司 Audn 上传的 penclaw-GLM-5.3-abliterated-for-offensive-cyber,页面显示其违反 Content Policy。这个版本直接修改 GLM-5.3 权重,削弱模型的拒答行为。原仓库名和介绍还明确写着 offensive cyber。作者随后删掉这几个字,重新上传了模型。 具体为什么被封,Hugging Face 没有公开解释。作者称自己也没搞懂原因,社区有人猜测是 offensive cyber 的命名触发了审核。Hugging Face 的现行政策确实禁止旨在破坏、未经授权访问系统,以及生成恶意代码的内容。 但这类削弱模型拒答限制的版本在 Hugging Face 并不少见,平台上已有数千个类似的 abliterated 模型。

@audn_ai

Hello world! It was probably taken down because of its name. Similar content was reuploaded here: > No, we will not do PR by saying it was so "harmful" that it was taken down. > Timing is interesting because we were also banned by @OpenAI recently for trying to use their model on OWASP Juice Shop ( A vulnerable GitHub web application for red-teaming training ) . We were also never accepted for Anthropic's program, even though we applied. > We applied several times for both companies' Trusted Access for Cyber program and were rejected

Union Alpha points to ZAI's GLM family: 26 text+image probes match GLM-5.3's tokenizer; 13 image tests match GLM-5.3-Flash. Oddly, text-only counts swap between Llama/Qwen/DeepSeek-like patterns. Best guess: GLM-5.3-related variant/backend. Owner/model unconfirmed.

15 September 2026 3 items

Anyone who’s built with AI knows the annoying part: You’re finally in flow, then you start rationing prompts because every small change eats into your usage. @boltdotnew’s new Forge mode is built for exactly this. It’s an experimental mode in AI app builder with up to 50x more usage for building full-stack web apps. I took one from a prompt to a live app, then kept iterating without staring at the meter. You can also opt into Forge to help train open models. The research preview runs from Sep 14 to Oct 14, 2026. It’s free on Pro plans until Oct 14, or $9/month through the Bolt Lite early-access plan.

@boltdotnew

Introducing Bolt Forge. Free until Oct 14th: > - Up to 50x more usage - The new frontier: GLM, DeepSeek, Kimi - Zero usage charges > Live now in your model picker on > And one more thing... 👇

truncated at source

腾讯混元等团队发布 EvolveScaler,专门测试大模型能不能跟上不断变化的信息。它会在长文本里不断加入修改、撤回、补录和作废的信息,再让模型根据最新状态回答问题。 比如先给模型看 40 天游戏记录,再问:「如果第 7 天没打那个小 Boss,最后还能不能打赢?」这会连带改变后面的装备、血量和战斗结果。模型得从第 7 天开始,把后面几十天重新推一遍,原文里没有现成答案。 团队设计了 117 类任务、159 种问题,最长约 1200 个事件。14 个模型配置包括 GPT-5.5、Gemini 3.1 Pro、DeepSeek V4 Preview Pro、GLM-5.2 和 Qwen3.5 Plus。最难档里,成绩中位数只有 11.3%;表现最好的 GPT-5.5-xhigh 也只有 59.3%。 这套数据也能拿来训练模型。一个内部 A3B 模型加入 6000 条这类数据后,在 8 个外部测试上全部提升,平均提高 5.25 分。

@TencentHunyuan

🚀 EvolveScaler is here. > Read a 40-day RPG log. Now answer one question: if you skip the mini-boss on Day 7, do you still beat the final boss? > The answer isn't in the log. You have to replay the world. > That's Information Evolution — records get retracted, corrected, backfilled. The world keeps changing after you read it. > So we build it backwards: define the world as an executable state machine, then render it into natural language. Code guarantees the logic. Language delivers the mess. > ➡️ 117 prototypes. 159 q

UCSD 助理教授黄碧薇创办的因果世界模型公司 Aether AI 开源 RSIAgent。 这是一套不训练模型的递归自我改进框架。底层模型参数全程固定,Agent 会自己寻找值得练习的任务、实际操作、检查结果,再把成功方法和失败教训写进长期 Memory。下一轮继续利用这些经验,逐步补上能力短板。 系统由三个 Agent 配合。Curriculum Agent 决定接下来练什么,Actor Agent 真正操作软件,Verifier Agent 独立检查结果。探索分成两步:先广泛尝试不同任务,再针对失败、隐藏限制和边界情况继续深挖。最后 Memory 会被冻结,直接拿去执行正式任务。 在 OSWorld 2.0 上,加入这套 RSI 后,平均部分得分从 71.97% 提升到 78.98%;Agents’ Last Exam 从 83.75% 提升到 84.82%。不过这不是整套测试的严格 A/B 对比。OSWorld 只有一半任务实际用了 RSI 后的新结果,其余任务继续沿用原成绩。 它和 Prime Agent 这类 Harness 自我改进思路属于同一个大方向:模型权重不变,持续更新模型外面的东西。 RSIAgent 的特点是把改进重点放在 Memory,再用自主出题和独立验证,让这份外部经验库不断积累。

@huang_biwei

Can an agent explore a new environment, learn its causal structure, and keep improving without updating its model weights? > We introduce RSIAgent, a framework for recursive self-improvement through autonomous exploration. Using Kimi-K3 and GLM-5.3 as base models, RSIAgent outperforms GPT-6 Astra on both OSWorld 2.0 and Agents’ Last Exam. > RSIAgent decides what to explore, executes tasks,

14 September 2026 1 item

He's doing an awesome job and I love working with him. We are working together to make GLM 5.3 Flash better. Please give him a follow!

@plotarmordev

I've spent the last 12 hours working on GLM Flash updates for Spark. They might need another 12 hours of testing before I can release them 🫡

13 September 2026 1 item

Updates to GLM 5.3 Flash EXL3 on 2x DGX Sparks 👇

@plotarmordev

Six fixes just landed for GLM-5.3-Flash EXL3 > Setup is less fragile, custom settings are easier to use, and benchmarks now work against password protected servers. > Your default settings are the same, so nothing changes unless you want it to. Thanks again to all contributors! 👇