AI briefing
14 September 2026
84 items were recorded on 14 September 2026, filed under 43 organisations and products.
That is up from 69 the day before.
Most covered: GPT-6 Astra (8), Grok (7) and GitHub (6).
22 of the day's items named no organisation or product this site tracks; they are listed under “Also recorded”.
12 items reached the source feed's 1024-character limit and are cut off mid-text; each is marked “truncated at source”.
Compiled by Bloger.fm Editorial Desk
Compiled from a monitored feed of public AI announcements. Items are quoted or summarised as recorded and are not independently verified — see the editorial policy.
GPT-6 Astra
OpenAI / 8 items
V4.1 paints its take on "Sunrise by the ocean" (Vladimir Kush) with SDFs in Rust, no reference image, just a description of the piece from the previous context. It has some sense of beauty I'd say. Astra is mildly impressed.
@teortaxesTexPainting in Rust is hard. It got the broad strokes correct within ≈5 minutes but then just kept rabbitholing. Work to be done on agent cooperation, too. Still, I'm impressed. V4.1 can reproduce gists of paintings in arbitrary language.
Elon just gave a huge update on the Grok model roadmap and Grok 4.8 is the major step forward • Grok 4.7 — roughly on par with Opus 5.0, better in some areas and worse in others....multimodal performance still needs work • Grok 4.8 — a massive 2.5T model trained on SpaceXAI’s new C++ software stack....training finishes this week and then moves straight into reinforcement learning • Grok 4.9 — Astra/Fable class • Grok 5 — potentially better than anything The crazy part is 4.7 hasn’t even launched yet and 4.8 is already finishing training Then 4.9 and Grok 5 are lined up right behind it Elon is moving the entire Grok roadmap insanely fast
@elonmusk@itslueul Grok 4.7 should be roughly on par with Opus 5.0, not 5.1. Better in some ways, worse in others. We need to fix multimodal performance. > Grok 4.8 will be a noticeable improvement. > Grok 4.9 is probably Astra/Fable class. > Grok 5 maybe better than anything. We shall see.
Amazing! A full 3D anatomy atlas has now been developed with GPT-6 Astra and WebMCP support! Where was this when we were in medical school? 😅
@ZentrixHQGPT-6 Astra has solved the problem of 3D anatomy tools and created an Anatomical Atlas with WebMCP support > 3D anatomy tools have had a control problem for years, and almost nobody built a real fix. > Every viewer forces the same workflow: click, drag, zoom, repeat, just to reach the angle a lecture or a case actually needs. A new WebMCP-enabled Anatomy Atlas skips that entirely. > The interface still supports manual exploration - full 3D anatomy, cross-section scrolling, tissue retraction by hand. But WebMCP exposes those same actions to an AI agent, so setting up a specific view, isolating a structure, or producing a short walkthrough video becomes a single instruction instead of a sequence of clicks. > It is not a UI overhaul. It is the same tool, plus a layer that lets software operate it the way a per
🚨 We're about to get some crazy Grok releases soon Grok 4.7 hopefully this week as Cursor employees are dropping hints and vagueposting Grok 4.8 will be a noticeable improvement according to Elon, my guess is early October unless they finish it sooner Grok 4.9 will apparently be Astra/Fable level 💀 I genuinely can't wait for Grok models to become more competitive, we're getting close
@LuminaBenchCrazy, so we're probably going to get Grok 4.7 this week now > Also Grok 4.8 (2.5T) finishes training this week and will start RL, this is going to be on their new stack too so will be incredibly efficient as well as fast > I wouldn't be surprised if Grok 4.8 drops at the end of September or early October as a guess
I asked GPT-6 to help me stay on top of my emails, bills, subscriptions, schedule, things like that. It sent me a checklist of about 15 different things I had to approve to grant it access. I instantly clicked approve all and hit yes. Astra, take the wheel!
🚨 BREAKING: Clone any successful company GPT-6 Astra in @shipper_now can take any app and make it yours: design, code, business plan... you can now one-shot the next duolingo / twitter / airbnb / etc this is THE END of vibe coding.
+ 2 more items − collapse
GPT-6 Sol just drop in the API : and it might be the model people actually use. - Up to 6x faster than Astra - Similar raw intelligence - Cheaper + better usage limits - Built for daily work, not just max benchmarks - Astra may stay the flagship, Sol could become the default. - Still unofficial pricing, or model ID yet.
@0x0SojalSecGPT-6 Sol might be the real everyday model : > - OpenAI reportedly treats Astra as the foundation, then post-trains Sol on top - Early talk: better daily quality, lower cost, Plus-friendly limits - Low-effort Astra already dunks on max-effort GPT-5.6 Sol - Astra on “low” already beats GPT-5.6 Sol on “max” - That gap is the whole story: this generation is not incremental If it ships, Sol is the one most people will actually live in
AIエージェント「Devin」を開発するCognitionが、AIモデル「GPT-6 Astra」や「Claude Fable」の性能を最大限に引き出すハーネス(AIモデルの実行環境・仕組み)「Fusion」を開発した。
Grok
xAI / 7 items
official siteGrok Bot新增了一个"本地出口路由"(local egress routing)功能 可以让流量从你本机转一圈再出去 Grok Bot 平时执行任务云端虚拟机里跑的,所以对外请求的 IP 地址是数据中心的 IP,不是你自己的 IP。 现在可以把 Grok Bot 虚拟机的网络流量绕一圈,从你自己的电脑(本地机器)出去 这样对方网站看到的访问来源就是你自己的 IP,而不是数据中心的 IP... 有什么用: 有些网站/服务会屏蔽数据中心IP(比如很多云服务商、VPS的IP段),只允许"看起来像真人"的家庭/办公网络IP访问,这种"挑剔"的网站以前 Grok Bot 直连会被拦 还有些服务本来就要求必须从你本人的IP发起请求(比如一些有IP白名单限制的内部系统、银行、区域限制服务等),只有走本地IP才能用 怎么开启:去 Settings(设置)> Computer(电脑)> Route egress(路由出口) 打开这个开关 注意:需要先把 Grok Bot 的"computer"组件更新到最新版本,这个功能才能用。
Elon Musk just gave a huge update on Grok 4.8 Grok 4.8 is a massive 2.5T model trained using SpaceXAI’s new C++ software stack Pretraining finishes this week.....then it moves straight into reinforcement learning While Grok 4.7 is getting its final refinements.....and Grok 4.8 is already moving into its next phase
Elon just gave a huge update on the Grok model roadmap and Grok 4.8 is the major step forward • Grok 4.7 — roughly on par with Opus 5.0, better in some areas and worse in others....multimodal performance still needs work • Grok 4.8 — a massive 2.5T model trained on SpaceXAI’s new C++ software stack....training finishes this week and then moves straight into reinforcement learning • Grok 4.9 — Astra/Fable class • Grok 5 — potentially better than anything The crazy part is 4.7 hasn’t even launched yet and 4.8 is already finishing training Then 4.9 and Grok 5 are lined up right behind it Elon is moving the entire Grok roadmap insanely fast
@elonmusk@itslueul Grok 4.7 should be roughly on par with Opus 5.0, not 5.1. Better in some ways, worse in others. We need to fix multimodal performance. > Grok 4.8 will be a noticeable improvement. > Grok 4.9 is probably Astra/Fable class. > Grok 5 maybe better than anything. We shall see.
🚨 We're about to get some crazy Grok releases soon Grok 4.7 hopefully this week as Cursor employees are dropping hints and vagueposting Grok 4.8 will be a noticeable improvement according to Elon, my guess is early October unless they finish it sooner Grok 4.9 will apparently be Astra/Fable level 💀 I genuinely can't wait for Grok models to become more competitive, we're getting close
@LuminaBenchCrazy, so we're probably going to get Grok 4.7 this week now > Also Grok 4.8 (2.5T) finishes training this week and will start RL, this is going to be on their new stack too so will be incredibly efficient as well as fast > I wouldn't be surprised if Grok 4.8 drops at the end of September or early October as a guess
Grok 4.8 to complete training and start reinforcement learning this week. Grok 4.7 is due any day now, according to Elon.
@elonmusk@techdevnotes Grok 4.8, which is a 2.5T model trained with our new C++ software stack, will finish training this week and start RL
Only 0.2 versions left ...
@elonmusk@AdamLowisz That will be Grok 5
+ 1 more items − collapse
Grok 4.7 is right around the corner
GitHub
6 items
official siteNGL ...this is fucking insane a solo developer just dropped a fully free, open-source ElevenLabs replacement no subscription, no catch and it's already at 28K stars on GitHub. what it can do: → clone a voice off a single clean audio clip → dub full videos into 646 languages → generate audiobooks, handle dictation and transcription → let you swap between 14 different TTS engines ElevenLabs? 32 languages. this one? 646. No per-character pricing. No caps on usage. Everything processes locally, so nothing leaves your machine. Bookmark it now repo's linked below.
If you've been looking for a best in class recipe for setting up DeepSeek V4.1 Flash on your RTX 6000 Pros, I have a treat for you from the team at @YourLocalAILab. 🐋
@YourLocalAILabWith today's release of Jovian Judgement R37, DeepSeek V4.1 Flash has full support with blazing speeds. > Decodes of 500 tok/s and prefills of 20k tok/s for single stream workloads on 4 RTX 6000 Pros. ⚡️ > To learn more point your agent at our GitHub here:
Crazy how T3 Code is still the only tool that cares about the relationship between a thread and a PR
@jullerinoThe 1:1 Thread to PR relationship is now gone. It was always wrong. A thread can now be linked to multiple PRs in @t3dotcodes, oh and we added support for GitHub Stacks as well!
下了一本 300 页的 PDF 想系统读一遍,让 AI 总结出来的东西又不敢全信,哪句是书里写的、哪句是它自己补的分不清。 learn-from-materials 是一个学习用的 Agent Skill,把 PDF、EPUB、Word、PPT、网页这些材料变成一个能交互的学习网页,Claude Code、Codex、Copilot 命令行都能装。 它先完整读完材料建一个知识库,页面里每条内容都标出处,PDF 精确到页码,PPT 精确到第几张,EPUB 精确到章节。 GitHub: 生成的页面有核心框架、内容导学、术语大全、行动规则几个板块,还带自检题和笔记,学到哪不明白直接在页面上记。 我最看重的是它把「材料里有的」「材料没写的」「模型补充的」分开标,读的时候心里有底,这个设计比多数 AI 读书工具老实。 分「快速了解」和「系统学习」两档深度,生成完还会审计一遍有没有漏读的章节。 默认不联网,笔记和错题只存在自己浏览器里,README 和页面都是中文。 手里堆着几本电子书和论文一直没读的,可以让它先啃一遍。
让 Agent 用 Three.js 搭个 3D 场景,出来的东西常常是几个方块加一盏灯,或者跑起来一片黑,它自己还说搭好了。 3dviz-pro-max 是一个专门做 3D 可视化的 Agent Skill,Claude Code 和 Codex 都能装,一句话描述想要的场景,它带着 Agent 从美术方向、物体结构一路做到能交互的成品。 它最不一样的地方,是逼 Agent 看自己渲染出来的画面。自带一个截图脚本,把搭好的场景按预设机位跑一遍截图,Agent 对着真实的帧改。 GitHub: 背后是一个能检索的资料库,223 个配方、440 条知识记录、22 套验证过的组件套件,覆盖 24 个方向。 从奇幻村庄、产品拆解到心脏剖面、线性代数、天体轨道都有,配方里调色板、天空、雾、灯光、相机参数都给了起始值,Agent 拿来改就行,不用从零猜。 仓库里还有 37 个能直接跑起来的示例,解剖、数学、物理、灯光练习分了 5 章,每个都附当初的提示词。装了 Blender 能烘出更细的模型,没装也能跑。 想拿 Agent 做 3D 演示、教学动画或者小场景的,装上试试,比裸奔靠谱。
平时在 Mac 上干活,偶尔要回家里那台 Windows 上取个文件、跑个只有 Windows 才有的软件,商业远程软件不是限速就是要买会员。 YourDesk 是一款 macOS 和 Windows 互连的远程桌面工具,输入对方 ID 和密码就连上,源码在 GitHub 公开,个人和公司内部用都免费。 编解码走硬件加速,Mac 上 H.264 和 HEVC,Windows 上 H.264 和 AV1,两端按各自能力自动选,工具栏能看到当前用的是哪种。 跨电脑复制粘贴是我觉得最实用的一点,文字、图片、文件、整个文件夹都能复制过去直接粘贴,一批最大 2GB。 GitHub: 多显示器在工具栏中间直接切,最多 4 个屏。远端分辨率高的话可以开画面增强,先压低传输再本地放大,4K 桌面效果更明显。 新版加了 MCP,AI Agent 能自己连上远程电脑,操作桌面或命令行,干完活断开。这个功能默认关着,开了也只放行本机来源。 除了桌面模式还有远程命令行模式,Linux 机器也能当命令行被控端。Mac 端是签过名过了苹果公证的 DMG,Apple Silicon 专用。
Qwen
Alibaba / 5 items
这个是真东西, 不是营销号吹的! 可能是目前最强的视频高清放大模型 视频超分终于能自己控制画质了 : 把哪几帧放大成什么样,整段视频就长成什么样。 SparkVSR 是 ECCV 2026 收录的论文, 作者来自德州农工大学 + YouTube/Google, arXiv 2603.16864。 官方仓库 701 star, Apache-2.0, 模型基于 CogVideoX1.5-5B-I2V 改的。 它的思路确实聪明: 传统视频超分是个黑盒:你输入一段糊视频,模型吐出来什么你就得接受什么。 SparkVSR 换了个玩法 : 先用任意一个你喜欢的图像超分模型, 把其中几帧放大到满意; 然后它把这几个高质量锚点帧传播到整段视频, 同时用原视频的运动信息做约束。 放大帧的效果, 直接决定整段视频的放大效果。 这就是"垫图模式"。 自由度也在这里: 想要真实感, 那几帧用 SeedVR2 放; 想要美颜感,用 Qwen-Image-Edit-2511-Upscale2K 。 同一个模型,能调出完全不同的风格倾向。 论文数据 : 在 CLIP-IQA、DOVER、MUSIQ 上最高分别提升 24.6%、21.8%、5.6%。 还能干别的:老片修复、视频风格迁移,论文里都验证过。 实测(RTX 4090 24G,640×640 放大到 1280×1280): 垫图模式:10GB 显存、120 秒 自动模式:14GB、150 秒 对照 SeedVR2:17GB、150 秒 显存和速度都比 SeedVR2 省,4090 完全跑得动。 但是有一个坑 : 它是按整段视频传播关键帧的, 多镜头会串味, 所以最好按镜头切开分别跑。 代码用官方仓库里的 ComfyUI-Spark/ 子目录👇
Marigold V2 turns an image-editing DiT into a single-step model for sharp, detailed dense prediction.📜 Apache 2.0. 🤖 📄 🏆 Best zero-shot results across all evaluated depth datasets among models trained on comparable data. AbsRel improves by 16%–26% over the previous best on KITTI and ETH3D. 🔍 Fur, foliage, fine wires, and object boundaries stay crisp. The two-stage iREPA and SinkLoss recipe tackles the smoothing and flying-pixel artifacts common in diffusion-based depth models. 🧩 The same framework reaches SOTA results in depth completion, see-through depth, surface normals, and intrinsic image decomposition. Depth completion records the lowest RMSE across all four reported benchmarks. ⚡ A pretrained Qwen-Image-Edit DiT becomes a single-step dense predictor through lightweight adaptation. Training takes less than a week on one 32 GB GPU.
16-GB Mac can run Qwen 3.8 27B multimodal locally. - 27B dense hybrid (Gated DeltaNet + attention) - Need 24 GB unified memory minimum. - 32 GB if you actually use images + long think. - Thinking mode with xhigh / medium / low - Vision + video in, text out - A one-click VLM for M-series - 262K native context - Turn off KV-cache quant or it can fail to load. - 16.1 GB on disk built for coding, agents, and long tasks not another chat toy. -
Qwen3.8 Flash Next is now supported in DwarfStar, covering 64GB Mac systems very well and with very fast inference of 50~70 t/s and > 1400 t/s prefill. For now this is Metal only. Thanks to @ivanfioravanti for all the cool work in the PR. N-grams on SSD like for DS4.1F.
Qwen3.8-Flash-Next just broke 100 tok/s 🚀 on two DGX Sparks. Bot lanes: classic or uncensored run at such speed 🤯 ⚡️ HumanEval 95.7 GSM8K 98.0 IFEval 91.5/93.4 MMLU-Pro 84.9
DeepSeek V4.1 Flash
DeepSeek / 5 items
If you've been looking for a best in class recipe for setting up DeepSeek V4.1 Flash on your RTX 6000 Pros, I have a treat for you from the team at @YourLocalAILab. 🐋
@YourLocalAILabWith today's release of Jovian Judgement R37, DeepSeek V4.1 Flash has full support with blazing speeds. > Decodes of 500 tok/s and prefills of 20k tok/s for single stream workloads on 4 RTX 6000 Pros. ⚡️ > To learn more point your agent at our GitHub here:
Extending $60 usage on DeepSeek V4.1 flash to 10 days. 🐐
@CommandCodeAIDeepSeek V4.1 Flash is live in Command Code. > 6x usage on $10/mo GOAT for a week. $60 usage. 154K requests. 7.8B tokens. > GOAT plan is the best AI coding plan. > Available in all plans and API. > 🐐
With the last commit into DwarfStar now you can use DeepSeek v4.1 Flash in a single DGX Spark as well, with SSD streaming. Around 9 t/s generation. It works also dual-spark RDMA at ~22 t/s.
You can Run Locally DeepSeek V4.1 Flash It beats GPT-5.6 Sol, Claude Opus 5 & can run with 4-GPU at your home locally No NVLink. - Jovian Judgement R37/TP4/DCP1: - 21,400 tok/s uncached 32K prefill - 280 tok/s C1 decode (budget 50) - 500 tok/s single-stream decode - 417 tok/s sieve - +44% prefill vs R36 - DSpark K7 + Engram (RAM or SSD) - 552B MoE, vision, 1M context Engram in pinned RAM or disk, DSpark K7. -
@0x0SojalSecYou don’t need a GPU-cluster to run a Uncensored DeepSeek-V4.1-Flash 763B model locally at home. > You need a Single RTX PRO 6000 Blackwell. > - 316 tok/s output - 9,351 tok/s prefill - FP8 + vLLM + DSpark - 64-96GB class machine.
Time to test DeepSeek v4.1 Flash with DwarfStar TP-RDMA on 2xM3 Ultra 512GB!
xAI
4 items
Busiest day official siteSpaceXAI just announced a 3-day livestream to build a company from scratch using Grok Bot. Here's what you need to know. Elon Musk shared the announcement today, promoting "Grok Bot Galaxy," a live event where three people start with just an idea and try to build an entire product or company from nothing, using Grok Bot as their AI teammate. The livestream runs September 15 to 17, roughly 8:30 AM to 6:00 PM PT each day, streamed from The Howard at 661 Howard St in San Francisco. Day one opens with a Grok Bot 101 session, followed by a Grok Bot for Engineering session. The company name has not been chosen yet. This builds on Grok Bot, xAI's AI agent that launched in beta on August 11. Each Bot gets its own persistent cloud computer, with a desktop, files, terminal, browser, and apps, and can keep working around the clock without supervision. Key numbers: - Livestream dates: September 15 to 17, 2026 - Daily hours: about 8:30 AM to 6:00 PM PT - 3 builders, 1 company, built from scratch - Grok Bot beta launch
Try Grok Build
@XFreezeI literally use Grok Build to edit images and videos....not just generate them, which it can also do > I mean actual image and video editing....the kind of stuff you would normally sit in Photoshop or editing software and do manually > I can tell Grok Build what I want changed, have it find the right footage, cut clips, edit everything together and handle a lot of the repetitive work for me > You can even save your preferred designs and workflows as skills....so it knows exactly how you like things done every time > It saves me an insane amount of time
Elon Musk just gave a huge update on Grok 4.8 Grok 4.8 is a massive 2.5T model trained using SpaceXAI’s new C++ software stack Pretraining finishes this week.....then it moves straight into reinforcement learning While Grok 4.7 is getting its final refinements.....and Grok 4.8 is already moving into its next phase
Elon just gave a huge update on the Grok model roadmap and Grok 4.8 is the major step forward • Grok 4.7 — roughly on par with Opus 5.0, better in some areas and worse in others....multimodal performance still needs work • Grok 4.8 — a massive 2.5T model trained on SpaceXAI’s new C++ software stack....training finishes this week and then moves straight into reinforcement learning • Grok 4.9 — Astra/Fable class • Grok 5 — potentially better than anything The crazy part is 4.7 hasn’t even launched yet and 4.8 is already finishing training Then 4.9 and Grok 5 are lined up right behind it Elon is moving the entire Grok roadmap insanely fast
@elonmusk@itslueul Grok 4.7 should be roughly on par with Opus 5.0, not 5.1. Better in some ways, worse in others. We need to fix multimodal performance. > Grok 4.8 will be a noticeable improvement. > Grok 4.9 is probably Astra/Fable class. > Grok 5 maybe better than anything. We shall see.
Gemini
Google / 4 items
official siteCongratulations to the @OpenAI team on the GPT-Live-1 launch. Here is the early read from our RWVoiceEQ STS evals. Voice naturalness is where it moves: up about 70% on their previous flagship, from near the bottom to level with Gemini 3.1 Flash Live at the top. What was OpenAI's weakest factor is now their strongest. Emotion understanding also improved, about 10% up and the highest score in the set. Emotion alignment is where it doesn't. Whether the model notices when tone contradicts the words, a hesitant or frustrated yes to a critical confirmation: unchanged from the previous flagship, under 1% apart. Sounding human and understanding humans are separate problems. The industry is solving the first one quickly. The second one is where the work is. Measuring the gap is how we close it. Early numbers. Full leaderboard update to follow.
new Google "model" appeared in the API antigravity-preview-09-2026 it isn't a model it is the antigravity agent endpoint hosted on linux sandbox, code execution, files, browser, etc... default model is Gemini 3.8 flash
Gemini just landed on Windows - Alt + Space and it's right there alongside everything you're already doing. The desktop AI war is no longer just a Mac fight.
@GeminiAppThe Gemini app is now available for Windows. > Stay in your flow with help that’s only one shortcut away. Press Alt + Space to polish drafts, summarize long documents, brainstorm new ideas, and create custom images and videos right alongside your favorite tools and daily apps.
DEEPSEEK-HARNESS IS A FREE OPEN-SOURCE FRAMEWORK FOR BUILDING CODING AGENTS WITH SWAPPABLE MODELS, TOOLS, SANDBOXES, UIS AND AGENT LOOPS. IT WORKS WITH DEEPSEEK, CLAUDE, GPT, GEMINI AND MORE.
4 items
official sitenew Google "model" appeared in the API antigravity-preview-09-2026 it isn't a model it is the antigravity agent endpoint hosted on linux sandbox, code execution, files, browser, etc... default model is Gemini 3.8 flash
这个是真东西, 不是营销号吹的! 可能是目前最强的视频高清放大模型 视频超分终于能自己控制画质了 : 把哪几帧放大成什么样,整段视频就长成什么样。 SparkVSR 是 ECCV 2026 收录的论文, 作者来自德州农工大学 + YouTube/Google, arXiv 2603.16864。 官方仓库 701 star, Apache-2.0, 模型基于 CogVideoX1.5-5B-I2V 改的。 它的思路确实聪明: 传统视频超分是个黑盒:你输入一段糊视频,模型吐出来什么你就得接受什么。 SparkVSR 换了个玩法 : 先用任意一个你喜欢的图像超分模型, 把其中几帧放大到满意; 然后它把这几个高质量锚点帧传播到整段视频, 同时用原视频的运动信息做约束。 放大帧的效果, 直接决定整段视频的放大效果。 这就是"垫图模式"。 自由度也在这里: 想要真实感, 那几帧用 SeedVR2 放; 想要美颜感,用 Qwen-Image-Edit-2511-Upscale2K 。 同一个模型,能调出完全不同的风格倾向。 论文数据 : 在 CLIP-IQA、DOVER、MUSIQ 上最高分别提升 24.6%、21.8%、5.6%。 还能干别的:老片修复、视频风格迁移,论文里都验证过。 实测(RTX 4090 24G,640×640 放大到 1280×1280): 垫图模式:10GB 显存、120 秒 自动模式:14GB、150 秒 对照 SeedVR2:17GB、150 秒 显存和速度都比 SeedVR2 省,4090 完全跑得动。 但是有一个坑 : 它是按整段视频传播关键帧的, 多镜头会串味, 所以最好按镜头切开分别跑。 代码用官方仓库里的 ComfyUI-Spark/ 子目录👇
🚨 New Google Model antigravity-preview-09-2026 i think this is coding focused model and they are making updates in antigravity too and jules agents
Meet MobileNet-v3-small: a tiny but mighty image classifier. Built for LiteRT/TFLite, it brings Google's efficient vision research to edge devices. Perfect for mobile apps that need fast, on-device AI without the cloud.
OpenAI
4 items
official siteCongratulations to the @OpenAI team on the GPT-Live-1 launch. Here is the early read from our RWVoiceEQ STS evals. Voice naturalness is where it moves: up about 70% on their previous flagship, from near the bottom to level with Gemini 3.1 Flash Live at the top. What was OpenAI's weakest factor is now their strongest. Emotion understanding also improved, about 10% up and the highest score in the set. Emotion alignment is where it doesn't. Whether the model notices when tone contradicts the words, a hesitant or frustrated yes to a critical confirmation: unchanged from the previous flagship, under 1% apart. Sounding human and understanding humans are separate problems. The industry is solving the first one quickly. The second one is where the work is. Measuring the gap is how we close it. Early numbers. Full leaderboard update to follow.
I built an AI startup simulator from scratch using Hy4 preview and wanted to see how far I could push it you start with $1M + 12 months of runway and have to build the next big AI company before you die hire researchers + engineers, rent GPUs, ship features, acquire users, raise funding and deal with random events like: > OpenAI ships your core feature for free > GPU prices spike 40% > your launch goes viral > your best researcher gets a $900k offer from Meta all while burn, ARR, users, compute, valuation + founder ownership update in real time Hy4 preview is a pretty interesting model for this kind of work too: - 770B parameters / 49B active - 1M-token context window - Apache 2.0 weights they also upgraded the preview yesterday to use significantly fewer tokens, which makes the whole build loop feel a lot faster. prompt you can try yourself: "Build me a playable AI startup simulator. Start me with $1M and 12 months of runway. Let me hire a team, allocate compute, build products, acquire users and ra
i just hope that we get gpt-6 spark
@imjustnewataiI’m hearing rumors of OpenAI gpt 6 spark releasing on OpenAI dev day 👀👀
GPT-6 Sol just drop in the API : and it might be the model people actually use. - Up to 6x faster than Astra - Similar raw intelligence - Cheaper + better usage limits - Built for daily work, not just max benchmarks - Astra may stay the flagship, Sol could become the default. - Still unofficial pricing, or model ID yet.
@0x0SojalSecGPT-6 Sol might be the real everyday model : > - OpenAI reportedly treats Astra as the foundation, then post-trains Sol on top - Early talk: better daily quality, lower cost, Plus-friendly limits - Low-effort Astra already dunks on max-effort GPT-5.6 Sol - Astra on “low” already beats GPT-5.6 Sol on “max” - That gap is the whole story: this generation is not incremental If it ships, Sol is the one most people will actually live in
Claude Code
Anthropic / 4 items
official siteIn what way is this a confirmation?
@notjazii🚨opus 5 is confirmed routing opus 5.2 > ran some more tests and it looks like routed mode is > way faster > gives really clean output > not lazy and loves to do longer tasks > run this prompt in claude code: > "do you know who is "tibo" the reset guy, don't search" > if it knows who tibo is, you most likely have opus 5.2 > run it and lemme know what you get
BREAKING: Today, we killed Claude Computer. I just watched my Mac alone one-shot a $31B company in 186 seconds. This is, without exaggeration, completely scary.
@claudeaiClaude can now use your computer in the background in Claude Cowork and Claude Code. > Give it something to do on your desktop and Claude clicks, types, and opens apps just like you would, while you work on something else.
下了一本 300 页的 PDF 想系统读一遍,让 AI 总结出来的东西又不敢全信,哪句是书里写的、哪句是它自己补的分不清。 learn-from-materials 是一个学习用的 Agent Skill,把 PDF、EPUB、Word、PPT、网页这些材料变成一个能交互的学习网页,Claude Code、Codex、Copilot 命令行都能装。 它先完整读完材料建一个知识库,页面里每条内容都标出处,PDF 精确到页码,PPT 精确到第几张,EPUB 精确到章节。 GitHub: 生成的页面有核心框架、内容导学、术语大全、行动规则几个板块,还带自检题和笔记,学到哪不明白直接在页面上记。 我最看重的是它把「材料里有的」「材料没写的」「模型补充的」分开标,读的时候心里有底,这个设计比多数 AI 读书工具老实。 分「快速了解」和「系统学习」两档深度,生成完还会审计一遍有没有漏读的章节。 默认不联网,笔记和错题只存在自己浏览器里,README 和页面都是中文。 手里堆着几本电子书和论文一直没读的,可以让它先啃一遍。
让 Agent 用 Three.js 搭个 3D 场景,出来的东西常常是几个方块加一盏灯,或者跑起来一片黑,它自己还说搭好了。 3dviz-pro-max 是一个专门做 3D 可视化的 Agent Skill,Claude Code 和 Codex 都能装,一句话描述想要的场景,它带着 Agent 从美术方向、物体结构一路做到能交互的成品。 它最不一样的地方,是逼 Agent 看自己渲染出来的画面。自带一个截图脚本,把搭好的场景按预设机位跑一遍截图,Agent 对着真实的帧改。 GitHub: 背后是一个能检索的资料库,223 个配方、440 条知识记录、22 套验证过的组件套件,覆盖 24 个方向。 从奇幻村庄、产品拆解到心脏剖面、线性代数、天体轨道都有,配方里调色板、天空、雾、灯光、相机参数都给了起始值,Agent 拿来改就行,不用从零猜。 仓库里还有 37 个能直接跑起来的示例,解剖、数学、物理、灯光练习分了 5 章,每个都附当初的提示词。装了 Blender 能烘出更细的模型,没装也能跑。 想拿 Agent 做 3D 演示、教学动画或者小场景的,装上试试,比裸奔靠谱。
MCP
3 items
official siteGoodbye, Postman Developers use your APIs. Now agents do too. Today, we’re launching @elvabytheneo , our biggest launch ever. Connect your codebase and Elva: → Discovers your APIs, no spec required → Builds and categorizes your API catalog → Maps usage across services → Checks security and AI readiness → Creates controlled API contracts for humans and agents → Publishes hosted MCP servers with auth and analytics → Tracks changes as your code evolves Some of our customers already receive more API calls from agents than humans. API management needs to catch up. Elva is now open to everyone. Try it with your own repo and tell us what we’re missing.
Grok Build just got another update focused on making multi-agent sessions easier to read and navigate Subagent groups now clearly show how many agents are still running versus finished, dashboard search and resume selection behave more predictably, cancellation banners show the actual cause, and long paths are cleaned up across headers and dashboards MCP handling also gets a useful fix....tools with long server prefixes are no longer silently dropped Another small release that makes busy agent sessions much easier to understand at a glance Release Notes: v1.0.31 Bug Fixes: • Folded subagent groups in scrollback now correctly label how many are still running versus completed. • Dashboard search and the resume picker now clearly indicate the active text field and restore list selection after search. • Dock sections no longer keep a selection highlight after you collapse them with a click. • Worktree paths no longer show an extra suffix in the header or dashboard. • MCP tools with long server prefixes are n
平时在 Mac 上干活,偶尔要回家里那台 Windows 上取个文件、跑个只有 Windows 才有的软件,商业远程软件不是限速就是要买会员。 YourDesk 是一款 macOS 和 Windows 互连的远程桌面工具,输入对方 ID 和密码就连上,源码在 GitHub 公开,个人和公司内部用都免费。 编解码走硬件加速,Mac 上 H.264 和 HEVC,Windows 上 H.264 和 AV1,两端按各自能力自动选,工具栏能看到当前用的是哪种。 跨电脑复制粘贴是我觉得最实用的一点,文字、图片、文件、整个文件夹都能复制过去直接粘贴,一批最大 2GB。 GitHub: 多显示器在工具栏中间直接切,最多 4 个屏。远端分辨率高的话可以开画面增强,先压低传输再本地放大,4K 桌面效果更明显。 新版加了 MCP,AI Agent 能自己连上远程电脑,操作桌面或命令行,干完活断开。这个功能默认关着,开了也只放行本机来源。 除了桌面模式还有远程命令行模式,Linux 机器也能当命令行被控端。Mac 端是签过名过了苹果公证的 DMG,Apple Silicon 专用。
Microsoft
3 items
official siteMicrosoft CEO Satya Nadella said superintelligence is only worth building if it stays under human control and helps people... he also pushed for companies to keep their own knowledge and learning loops instead of depending on one model provider also announced Microsoft’s own first-party MAI models will get a public Code of Conduct tomorrow..
@satyanadellaAny pursuit of superintelligence has to be grounded in the core principle that if the AI we build is not helping humanity and under human control, it's not worth pursuing. > We also need to accelerate and spread the benefits of AI, such that they are diffused broadly across countries, communities, and companies. This requires a frontier ecosystem in which both closed and open-source models can thrive. > And for firms, it’s imperative that they retain full control over their unique and tacit knowledge. Every organization should be able to build its own continuous learning loop/hill climbing machine, without be
RIP =COPILOT() in Excel. It put AI inside a formula. Had its moment. Paved the way. Now Copilot can edit the workbook and handle multi-step tasks. The function served its purpose. Now it has to go. Retired today, September 14. On to the next.
We’re expanding our work with @nvidia to bring fully local AI to Microsoft Windows PCs with RTX GPUs. Unmetered local intelligence on every Windows PC running on NVIDIA hardware and Perplexity harness. Enjoy!
@perplexity_aiPortable Computer is now available on Windows PCs with @NVIDIA RTX GPUs. > Run the harness, agents, and models locally on your PC. > Work with local files and connected apps without sending tasks to the cloud. Use frontier cloud models when needed.
Hermes Agent
3 items
ChatGPT-Image-2.5 is now available on Nous Portal to use in Hermes Agent, as well as it's previously supported Fal and ChatGPT direct routes!
Crazyy that this even works. Agent Reach gives your Hermes and OpenClaw agents access to X, LinkedIn, Facebook, Instagram and more. And it's 100% free and Opensource.
Reasoning effort is now a seperate selector in the Hermes Agent desktop app's composer now!
DeepSeek
3 items
official siteDid A BIG Speed Run Update Last Night ! 🔓⚡️ DeepSeek V4.1-Flash on 3-4 DGX Sparks just went uncensored AND got faster thanks to @u1tra_instinct 🚀 ~190 tok/s throughput (6 streams, +13%) 💻 ~85 tok/s on code, single stream 📥 ~2,000 tok/s cold prefill (+38%) ⏱️ 0.22s to first token (-20%) 📚 1M-token context, proven 🧊 3.76M-token KV pool Abliterated attention, experts untouched. Full recipe's updated👇 🐋🤖
DEEPSEEK-HARNESS IS A FREE OPEN-SOURCE FRAMEWORK FOR BUILDING CODING AGENTS WITH SWAPPABLE MODELS, TOOLS, SANDBOXES, UIS AND AGENT LOOPS. IT WORKS WITH DEEPSEEK, CLAUDE, GPT, GEMINI AND MORE.
I trained an AI model on my phone through Telegram > Using OpenClaw running on Hugging Face infra: ML Claw > It beat DeepSeek V4 Pro on the given task while having 80 thousand times less parameters > That's right. The model, GoePT-1-20m is only 20 million parameters. It runs in the browser, on the CPU. And it beats DeepSeek V4 Pro, a 1.6 trillion parameter model, in AlmanBench > This concludes my 4 year old side project (fun fact, I created a dataset for this pre AI agents, using SpaCy, and paid 50 euros from my own pocket to rent 4090s on runpod to train a T5 variant, *years* before I joined Hugging Face. That first attempt was not very successful. This one is. My first ML adventure 🤗) > The goal of this side project was to show that you can *vibe* machine learning now, on platforms with tightly integrated GPUs, storage and compute, like Hugging Face > Including dataset creation, autoresearch and the final training run (hat tip to ML-Intern which g
@onusozI trained an AI model on my phone through Telegram > Using OpenClaw running on Hugging Face infra: ML Claw > It beat DeepSeek V4 Pro on the given task while having 80 thousand times less parameters > That's right. The model, GoePT-1-20m is only 20 million parameters. It runs in the browser, on the CPU. And it beats DeepSeek V4 Pro, a 1.6 trillion parameter model, in AlmanBench > This concludes my 4 year old side project (fun fact, I created a dataset for this pre AI agents, using SpaCy, and paid 50 euros from my own pocket to rent 4090s on runpod to train a T5 variant, *years* before I joined Hugging Face. That first attempt was not very successful. This one is. My first ML adventure 🤗) > The goal of this side project was to show that you can *vibe* machine learning now, on platforms with tightly integrated GPUs, storage and compute, like Hugging Face > Including dataset creation, autoresearch and the final training run (hat tip to ML-Intern which g
Codex
OpenAI / 3 items
Codex tasks aren't tied to one Mac anymore 👀 Tested Handoff both ways: ☁️ Push it to the cloud 💻 Move it, files included, to another Mac Start anywhere. Finish anywhere. The work follows you 🔥 (Feature flagged for now, so you may not see it yet)
下了一本 300 页的 PDF 想系统读一遍,让 AI 总结出来的东西又不敢全信,哪句是书里写的、哪句是它自己补的分不清。 learn-from-materials 是一个学习用的 Agent Skill,把 PDF、EPUB、Word、PPT、网页这些材料变成一个能交互的学习网页,Claude Code、Codex、Copilot 命令行都能装。 它先完整读完材料建一个知识库,页面里每条内容都标出处,PDF 精确到页码,PPT 精确到第几张,EPUB 精确到章节。 GitHub: 生成的页面有核心框架、内容导学、术语大全、行动规则几个板块,还带自检题和笔记,学到哪不明白直接在页面上记。 我最看重的是它把「材料里有的」「材料没写的」「模型补充的」分开标,读的时候心里有底,这个设计比多数 AI 读书工具老实。 分「快速了解」和「系统学习」两档深度,生成完还会审计一遍有没有漏读的章节。 默认不联网,笔记和错题只存在自己浏览器里,README 和页面都是中文。 手里堆着几本电子书和论文一直没读的,可以让它先啃一遍。
让 Agent 用 Three.js 搭个 3D 场景,出来的东西常常是几个方块加一盏灯,或者跑起来一片黑,它自己还说搭好了。 3dviz-pro-max 是一个专门做 3D 可视化的 Agent Skill,Claude Code 和 Codex 都能装,一句话描述想要的场景,它带着 Agent 从美术方向、物体结构一路做到能交互的成品。 它最不一样的地方,是逼 Agent 看自己渲染出来的画面。自带一个截图脚本,把搭好的场景按预设机位跑一遍截图,Agent 对着真实的帧改。 GitHub: 背后是一个能检索的资料库,223 个配方、440 条知识记录、22 套验证过的组件套件,覆盖 24 个方向。 从奇幻村庄、产品拆解到心脏剖面、线性代数、天体轨道都有,配方里调色板、天空、雾、灯光、相机参数都给了起始值,Agent 拿来改就行,不用从零猜。 仓库里还有 37 个能直接跑起来的示例,解剖、数学、物理、灯光练习分了 5 章,每个都附当初的提示词。装了 Blender 能烘出更细的模型,没装也能跑。 想拿 Agent 做 3D 演示、教学动画或者小场景的,装上试试,比裸奔靠谱。
GPT-5.6 Sol
OpenAI / 3 items
You can Run Locally DeepSeek V4.1 Flash It beats GPT-5.6 Sol, Claude Opus 5 & can run with 4-GPU at your home locally No NVLink. - Jovian Judgement R37/TP4/DCP1: - 21,400 tok/s uncached 32K prefill - 280 tok/s C1 decode (budget 50) - 500 tok/s single-stream decode - 417 tok/s sieve - +44% prefill vs R36 - DSpark K7 + Engram (RAM or SSD) - 552B MoE, vision, 1M context Engram in pinned RAM or disk, DSpark K7. -
@0x0SojalSecYou don’t need a GPU-cluster to run a Uncensored DeepSeek-V4.1-Flash 763B model locally at home. > You need a Single RTX PRO 6000 Blackwell. > - 316 tok/s output - 9,351 tok/s prefill - FP8 + vLLM + DSpark - 64-96GB class machine.
GPT-6 Sol just drop in the API : and it might be the model people actually use. - Up to 6x faster than Astra - Similar raw intelligence - Cheaper + better usage limits - Built for daily work, not just max benchmarks - Astra may stay the flagship, Sol could become the default. - Still unofficial pricing, or model ID yet.
@0x0SojalSecGPT-6 Sol might be the real everyday model : > - OpenAI reportedly treats Astra as the foundation, then post-trains Sol on top - Early talk: better daily quality, lower cost, Plus-friendly limits - Low-effort Astra already dunks on max-effort GPT-5.6 Sol - Astra on “low” already beats GPT-5.6 Sol on “max” - That gap is the whole story: this generation is not incremental If it ships, Sol is the one most people will actually live in
Open weights. Shared progress. MiniMax H3 is moving fast. We built MiniMax H3 for video generation with native stereo audio and multimodal reference control. The open-source community is making that capability faster, more accessible, and easier to build on. Recent highlights: • FastH3 — FastVideo, Nuva Lab and NVIDIA: 4-step distillation, now running on DGX Spark and Apple Silicon. • Sol-H3 — NVIDIA’s SANA team: now on DGX Spark with a two-stage H3 + LTX-2.5 pipeline. On 8×B300, the team reports 15 seconds of 768p video + audio in 6.6 seconds of warm inference.* • VDN — Haocheng Xi and the OpenVDN team: rethinking attention for faster H3 inference, with weights, training and inference code released. • PDD — NVIDIA’s distillation method, brought to H3 by Alibaba PAI as 8-step Acc-LoRAs, now supported in ComfyUI. • LightX2V — 4- and 8-step Turbo LoRAs, with workflows for text, image and reference-conditioned video + audio. Behind every release are people training, optimizing, quantizing, testing and sh
Grok Bot
xAI / 2 items
SpaceXAI just announced a 3-day livestream to build a company from scratch using Grok Bot. Here's what you need to know. Elon Musk shared the announcement today, promoting "Grok Bot Galaxy," a live event where three people start with just an idea and try to build an entire product or company from nothing, using Grok Bot as their AI teammate. The livestream runs September 15 to 17, roughly 8:30 AM to 6:00 PM PT each day, streamed from The Howard at 661 Howard St in San Francisco. Day one opens with a Grok Bot 101 session, followed by a Grok Bot for Engineering session. The company name has not been chosen yet. This builds on Grok Bot, xAI's AI agent that launched in beta on August 11. Each Bot gets its own persistent cloud computer, with a desktop, files, terminal, browser, and apps, and can keep working around the clock without supervision. Key numbers: - Livestream dates: September 15 to 17, 2026 - Daily hours: about 8:30 AM to 6:00 PM PT - 3 builders, 1 company, built from scratch - Grok Bot beta launch
Grok Bot新增了一个"本地出口路由"(local egress routing)功能 可以让流量从你本机转一圈再出去 Grok Bot 平时执行任务云端虚拟机里跑的,所以对外请求的 IP 地址是数据中心的 IP,不是你自己的 IP。 现在可以把 Grok Bot 虚拟机的网络流量绕一圈,从你自己的电脑(本地机器)出去 这样对方网站看到的访问来源就是你自己的 IP,而不是数据中心的 IP... 有什么用: 有些网站/服务会屏蔽数据中心IP(比如很多云服务商、VPS的IP段),只允许"看起来像真人"的家庭/办公网络IP访问,这种"挑剔"的网站以前 Grok Bot 直连会被拦 还有些服务本来就要求必须从你本人的IP发起请求(比如一些有IP白名单限制的内部系统、银行、区域限制服务等),只有走本地IP才能用 怎么开启:去 Settings(设置)> Computer(电脑)> Route egress(路由出口) 打开这个开关 注意:需要先把 Grok Bot 的"computer"组件更新到最新版本,这个功能才能用。
Grok Build
xAI / 2 items
Grok Build just got another update focused on making multi-agent sessions easier to read and navigate Subagent groups now clearly show how many agents are still running versus finished, dashboard search and resume selection behave more predictably, cancellation banners show the actual cause, and long paths are cleaned up across headers and dashboards MCP handling also gets a useful fix....tools with long server prefixes are no longer silently dropped Another small release that makes busy agent sessions much easier to understand at a glance Release Notes: v1.0.31 Bug Fixes: • Folded subagent groups in scrollback now correctly label how many are still running versus completed. • Dashboard search and the resume picker now clearly indicate the active text field and restore list selection after search. • Dock sections no longer keep a selection highlight after you collapse them with a click. • Worktree paths no longer show an extra suffix in the header or dashboard. • MCP tools with long server prefixes are n
Try Grok Build
@XFreezeI literally use Grok Build to edit images and videos....not just generate them, which it can also do > I mean actual image and video editing....the kind of stuff you would normally sit in Photoshop or editing software and do manually > I can tell Grok Build what I want changed, have it find the right footage, cut clips, edit everything together and handle a lot of the repetitive work for me > You can even save your preferred designs and workflows as skills....so it knows exactly how you like things done every time > It saves me an insane amount of time
Cursor
2 items
official site🚨 We're about to get some crazy Grok releases soon Grok 4.7 hopefully this week as Cursor employees are dropping hints and vagueposting Grok 4.8 will be a noticeable improvement according to Elon, my guess is early October unless they finish it sooner Grok 4.9 will apparently be Astra/Fable level 💀 I genuinely can't wait for Grok models to become more competitive, we're getting close
@LuminaBenchCrazy, so we're probably going to get Grok 4.7 this week now > Also Grok 4.8 (2.5T) finishes training this week and will start RL, this is going to be on their new stack too so will be incredibly efficient as well as fast > I wouldn't be surprised if Grok 4.8 drops at the end of September or early October as a guess
composer 2.5 is kinda flying under the radar Cursor reports 79.8% on SWE-bench Multilingual and 69.3% on Terminal-Bench 2.0, at $0.50/M input and $2.50/M output how are we not talking about this one more?
Claude
Anthropic / 2 items
official siteBREAKING: Today, we killed Claude Computer. I just watched my Mac alone one-shot a $31B company in 186 seconds. This is, without exaggeration, completely scary.
@claudeaiClaude can now use your computer in the background in Claude Cowork and Claude Code. > Give it something to do on your desktop and Claude clicks, types, and opens apps just like you would, while you work on something else.
DEEPSEEK-HARNESS IS A FREE OPEN-SOURCE FRAMEWORK FOR BUILDING CODING AGENTS WITH SWAPPABLE MODELS, TOOLS, SANDBOXES, UIS AND AGENT LOOPS. IT WORKS WITH DEEPSEEK, CLAUDE, GPT, GEMINI AND MORE.
MiniMax
2 items
MiniMax Code 2.0 just launched - rebuilt from the core on the Pi Agent framework with remote control baked in. Connect your phone to a live desktop session, follow agent output, send instructions, and handle approvals from anywhere. Your coding agent now fits in your pocket.
Open weights. Shared progress. MiniMax H3 is moving fast. We built MiniMax H3 for video generation with native stereo audio and multimodal reference control. The open-source community is making that capability faster, more accessible, and easier to build on. Recent highlights: • FastH3 — FastVideo, Nuva Lab and NVIDIA: 4-step distillation, now running on DGX Spark and Apple Silicon. • Sol-H3 — NVIDIA’s SANA team: now on DGX Spark with a two-stage H3 + LTX-2.5 pipeline. On 8×B300, the team reports 15 seconds of 768p video + audio in 6.6 seconds of warm inference.* • VDN — Haocheng Xi and the OpenVDN team: rethinking attention for faster H3 inference, with weights, training and inference code released. • PDD — NVIDIA’s distillation method, brought to H3 by Alibaba PAI as 8-step Acc-LoRAs, now supported in ComfyUI. • LightX2V — 4- and 8-step Turbo LoRAs, with workflows for text, image and reference-conditioned video + audio. Behind every release are people training, optimizing, quantizing, testing and sh
Copilot
Microsoft / 2 items
official siteRIP =COPILOT() in Excel. It put AI inside a formula. Had its moment. Paved the way. Now Copilot can edit the workbook and handle multi-step tasks. The function served its purpose. Now it has to go. Retired today, September 14. On to the next.
下了一本 300 页的 PDF 想系统读一遍,让 AI 总结出来的东西又不敢全信,哪句是书里写的、哪句是它自己补的分不清。 learn-from-materials 是一个学习用的 Agent Skill,把 PDF、EPUB、Word、PPT、网页这些材料变成一个能交互的学习网页,Claude Code、Codex、Copilot 命令行都能装。 它先完整读完材料建一个知识库,页面里每条内容都标出处,PDF 精确到页码,PPT 精确到第几张,EPUB 精确到章节。 GitHub: 生成的页面有核心框架、内容导学、术语大全、行动规则几个板块,还带自检题和笔记,学到哪不明白直接在页面上记。 我最看重的是它把「材料里有的」「材料没写的」「模型补充的」分开标,读的时候心里有底,这个设计比多数 AI 读书工具老实。 分「快速了解」和「系统学习」两档深度,生成完还会审计一遍有没有漏读的章节。 默认不联网,笔记和错题只存在自己浏览器里,README 和页面都是中文。 手里堆着几本电子书和论文一直没读的,可以让它先啃一遍。
Perplexity
2 items
Busiest day official siteComputer is turning into a workspace of humans and AIs.
@AskPerplexityYou can now tag coworkers in Computer sessions. > Use @ to mention and share a session with someone. > Available now on the web for all Computer users.
We’re expanding our work with @nvidia to bring fully local AI to Microsoft Windows PCs with RTX GPUs. Unmetered local intelligence on every Windows PC running on NVIDIA hardware and Perplexity harness. Enjoy!
@perplexity_aiPortable Computer is now available on Windows PCs with @NVIDIA RTX GPUs. > Run the harness, agents, and models locally on your PC. > Work with local files and connected apps without sending tasks to the cloud. Use frontier cloud models when needed.
OpenClaw
2 items
Crazyy that this even works. Agent Reach gives your Hermes and OpenClaw agents access to X, LinkedIn, Facebook, Instagram and more. And it's 100% free and Opensource.
I trained an AI model on my phone through Telegram > Using OpenClaw running on Hugging Face infra: ML Claw > It beat DeepSeek V4 Pro on the given task while having 80 thousand times less parameters > That's right. The model, GoePT-1-20m is only 20 million parameters. It runs in the browser, on the CPU. And it beats DeepSeek V4 Pro, a 1.6 trillion parameter model, in AlmanBench > This concludes my 4 year old side project (fun fact, I created a dataset for this pre AI agents, using SpaCy, and paid 50 euros from my own pocket to rent 4090s on runpod to train a T5 variant, *years* before I joined Hugging Face. That first attempt was not very successful. This one is. My first ML adventure 🤗) > The goal of this side project was to show that you can *vibe* machine learning now, on platforms with tightly integrated GPUs, storage and compute, like Hugging Face > Including dataset creation, autoresearch and the final training run (hat tip to ML-Intern which g
@onusozI trained an AI model on my phone through Telegram > Using OpenClaw running on Hugging Face infra: ML Claw > It beat DeepSeek V4 Pro on the given task while having 80 thousand times less parameters > That's right. The model, GoePT-1-20m is only 20 million parameters. It runs in the browser, on the CPU. And it beats DeepSeek V4 Pro, a 1.6 trillion parameter model, in AlmanBench > This concludes my 4 year old side project (fun fact, I created a dataset for this pre AI agents, using SpaCy, and paid 50 euros from my own pocket to rent 4090s on runpod to train a T5 variant, *years* before I joined Hugging Face. That first attempt was not very successful. This one is. My first ML adventure 🤗) > The goal of this side project was to show that you can *vibe* machine learning now, on platforms with tightly integrated GPUs, storage and compute, like Hugging Face > Including dataset creation, autoresearch and the final training run (hat tip to ML-Intern which g
Hugging Face
2 items
official sitezoom 256× into a photo, one 4× step at a time 🔍 OracleZoom approaches super-resolution by writing its own prompt at every level and re-rendering ▶️
I trained an AI model on my phone through Telegram > Using OpenClaw running on Hugging Face infra: ML Claw > It beat DeepSeek V4 Pro on the given task while having 80 thousand times less parameters > That's right. The model, GoePT-1-20m is only 20 million parameters. It runs in the browser, on the CPU. And it beats DeepSeek V4 Pro, a 1.6 trillion parameter model, in AlmanBench > This concludes my 4 year old side project (fun fact, I created a dataset for this pre AI agents, using SpaCy, and paid 50 euros from my own pocket to rent 4090s on runpod to train a T5 variant, *years* before I joined Hugging Face. That first attempt was not very successful. This one is. My first ML adventure 🤗) > The goal of this side project was to show that you can *vibe* machine learning now, on platforms with tightly integrated GPUs, storage and compute, like Hugging Face > Including dataset creation, autoresearch and the final training run (hat tip to ML-Intern which g
@onusozI trained an AI model on my phone through Telegram > Using OpenClaw running on Hugging Face infra: ML Claw > It beat DeepSeek V4 Pro on the given task while having 80 thousand times less parameters > That's right. The model, GoePT-1-20m is only 20 million parameters. It runs in the browser, on the CPU. And it beats DeepSeek V4 Pro, a 1.6 trillion parameter model, in AlmanBench > This concludes my 4 year old side project (fun fact, I created a dataset for this pre AI agents, using SpaCy, and paid 50 euros from my own pocket to rent 4090s on runpod to train a T5 variant, *years* before I joined Hugging Face. That first attempt was not very successful. This one is. My first ML adventure 🤗) > The goal of this side project was to show that you can *vibe* machine learning now, on platforms with tightly integrated GPUs, storage and compute, like Hugging Face > Including dataset creation, autoresearch and the final training run (hat tip to ML-Intern which g
DGX Spark
NVIDIA / 2 items
With the last commit into DwarfStar now you can use DeepSeek v4.1 Flash in a single DGX Spark as well, with SSD streaming. Around 9 t/s generation. It works also dual-spark RDMA at ~22 t/s.
Open weights. Shared progress. MiniMax H3 is moving fast. We built MiniMax H3 for video generation with native stereo audio and multimodal reference control. The open-source community is making that capability faster, more accessible, and easier to build on. Recent highlights: • FastH3 — FastVideo, Nuva Lab and NVIDIA: 4-step distillation, now running on DGX Spark and Apple Silicon. • Sol-H3 — NVIDIA’s SANA team: now on DGX Spark with a two-stage H3 + LTX-2.5 pipeline. On 8×B300, the team reports 15 seconds of 768p video + audio in 6.6 seconds of warm inference.* • VDN — Haocheng Xi and the OpenVDN team: rethinking attention for faster H3 inference, with weights, training and inference code released. • PDD — NVIDIA’s distillation method, brought to H3 by Alibaba PAI as 8-step Acc-LoRAs, now supported in ComfyUI. • LightX2V — 4- and 8-step Turbo LoRAs, with workflows for text, image and reference-conditioned video + audio. Behind every release are people training, optimizing, quantizing, testing and sh
Apple
2 items
official site平时在 Mac 上干活,偶尔要回家里那台 Windows 上取个文件、跑个只有 Windows 才有的软件,商业远程软件不是限速就是要买会员。 YourDesk 是一款 macOS 和 Windows 互连的远程桌面工具,输入对方 ID 和密码就连上,源码在 GitHub 公开,个人和公司内部用都免费。 编解码走硬件加速,Mac 上 H.264 和 HEVC,Windows 上 H.264 和 AV1,两端按各自能力自动选,工具栏能看到当前用的是哪种。 跨电脑复制粘贴是我觉得最实用的一点,文字、图片、文件、整个文件夹都能复制过去直接粘贴,一批最大 2GB。 GitHub: 多显示器在工具栏中间直接切,最多 4 个屏。远端分辨率高的话可以开画面增强,先压低传输再本地放大,4K 桌面效果更明显。 新版加了 MCP,AI Agent 能自己连上远程电脑,操作桌面或命令行,干完活断开。这个功能默认关着,开了也只放行本机来源。 除了桌面模式还有远程命令行模式,Linux 机器也能当命令行被控端。Mac 端是签过名过了苹果公证的 DMG,Apple Silicon 专用。
Open weights. Shared progress. MiniMax H3 is moving fast. We built MiniMax H3 for video generation with native stereo audio and multimodal reference control. The open-source community is making that capability faster, more accessible, and easier to build on. Recent highlights: • FastH3 — FastVideo, Nuva Lab and NVIDIA: 4-step distillation, now running on DGX Spark and Apple Silicon. • Sol-H3 — NVIDIA’s SANA team: now on DGX Spark with a two-stage H3 + LTX-2.5 pipeline. On 8×B300, the team reports 15 seconds of 768p video + audio in 6.6 seconds of warm inference.* • VDN — Haocheng Xi and the OpenVDN team: rethinking attention for faster H3 inference, with weights, training and inference code released. • PDD — NVIDIA’s distillation method, brought to H3 by Alibaba PAI as 8-step Acc-LoRAs, now supported in ComfyUI. • LightX2V — 4- and 8-step Turbo LoRAs, with workflows for text, image and reference-conditioned video + audio. Behind every release are people training, optimizing, quantizing, testing and sh
NVIDIA
2 items
official siteWe’re expanding our work with @nvidia to bring fully local AI to Microsoft Windows PCs with RTX GPUs. Unmetered local intelligence on every Windows PC running on NVIDIA hardware and Perplexity harness. Enjoy!
@perplexity_aiPortable Computer is now available on Windows PCs with @NVIDIA RTX GPUs. > Run the harness, agents, and models locally on your PC. > Work with local files and connected apps without sending tasks to the cloud. Use frontier cloud models when needed.
Open weights. Shared progress. MiniMax H3 is moving fast. We built MiniMax H3 for video generation with native stereo audio and multimodal reference control. The open-source community is making that capability faster, more accessible, and easier to build on. Recent highlights: • FastH3 — FastVideo, Nuva Lab and NVIDIA: 4-step distillation, now running on DGX Spark and Apple Silicon. • Sol-H3 — NVIDIA’s SANA team: now on DGX Spark with a two-stage H3 + LTX-2.5 pipeline. On 8×B300, the team reports 15 seconds of 768p video + audio in 6.6 seconds of warm inference.* • VDN — Haocheng Xi and the OpenVDN team: rethinking attention for faster H3 inference, with weights, training and inference code released. • PDD — NVIDIA’s distillation method, brought to H3 by Alibaba PAI as 8-step Acc-LoRAs, now supported in ComfyUI. • LightX2V — 4- and 8-step Turbo LoRAs, with workflows for text, image and reference-conditioned video + audio. Behind every release are people training, optimizing, quantizing, testing and sh
ElevenLabs
1 item
official siteNGL ...this is fucking insane a solo developer just dropped a fully free, open-source ElevenLabs replacement no subscription, no catch and it's already at 28K stars on GitHub. what it can do: → clone a voice off a single clean audio clip → dub full videos into 646 languages → generate audiobooks, handle dictation and transcription → let you swap between 14 different TTS engines ElevenLabs? 32 languages. this one? 646. No per-character pricing. No caps on usage. Everything processes locally, so nothing leaves your machine. Bookmark it now repo's linked below.
Kimi
Moonshot AI / 1 item
Kimi-K3 vs new Kimi-K2.8 Preview build the exact same game using both models with the same prompts i believe the whole point of releasing kimi-k2.8 after kimi-k3 was to fix the issues kimi-k3 has which is speed and cost efficiency keeping kimi-k3 as a premium option and kimi-k2.8 as mid tier option would you switch kimi-k3 for kimi-k2.8?
@TimJayaswe asked Moonshot to drop Kimi-k3.1 but they ended up dropping Kimi-k2.8 preview instead? if the whole point was to fix speed then Kimi-k3-Flash would've been a better name
Meta
1 item
official siteI built an AI startup simulator from scratch using Hy4 preview and wanted to see how far I could push it you start with $1M + 12 months of runway and have to build the next big AI company before you die hire researchers + engineers, rent GPUs, ship features, acquire users, raise funding and deal with random events like: > OpenAI ships your core feature for free > GPU prices spike 40% > your launch goes viral > your best researcher gets a $900k offer from Meta all while burn, ARR, users, compute, valuation + founder ownership update in real time Hy4 preview is a pretty interesting model for this kind of work too: - 770B parameters / 49B active - 1M-token context window - Apache 2.0 weights they also upgraded the preview yesterday to use significantly fewer tokens, which makes the whole build loop feel a lot faster. prompt you can try yourself: "Build me a playable AI startup simulator. Start me with $1M and 12 months of runway. Let me hire a team, allocate compute, build products, acquire users and ra
Muse
Meta / 1 item
something people may have missed: we have been building our models (muse spark 1 through 1.3) specifically to be exceptional for muse over many months. in each of our releases, we made big gains on agentic and multimodal capability. each of these steps were secretly to set the stage for our personal agent, @Muse. (they’re named the same thing for a reason… we planned ahead!) when we started MSL last year, the goal was always to develop personal superintelligence. that always meant giving AI superpowers to everyone in the world, and the personal agent is the perfect manifestation of that. building great models takes time, and it takes an enormous amount of planning and foresight to develop unique capabilities into these models. it often requires months of careful research and work across infra and data to build excellent models. it is very sweet to see the culmination of over a year of hard work come together into a product that people are loving. we are so glad that people are enjoying muse, and we can’t
Muse Spark
Meta / 1 item
something people may have missed: we have been building our models (muse spark 1 through 1.3) specifically to be exceptional for muse over many months. in each of our releases, we made big gains on agentic and multimodal capability. each of these steps were secretly to set the stage for our personal agent, @Muse. (they’re named the same thing for a reason… we planned ahead!) when we started MSL last year, the goal was always to develop personal superintelligence. that always meant giving AI superpowers to everyone in the world, and the personal agent is the perfect manifestation of that. building great models takes time, and it takes an enormous amount of planning and foresight to develop unique capabilities into these models. it often requires months of careful research and work across infra and data to build excellent models. it is very sweet to see the culmination of over a year of hard work come together into a product that people are loving. we are so glad that people are enjoying muse, and we can’t
ChatGPT
OpenAI / 1 item
official siteChatGPT-Image-2.5 is now available on Nous Portal to use in Hermes Agent, as well as it's previously supported Fal and ChatGPT direct routes!
GLM
Zhipu AI / 1 item
He's doing an awesome job and I love working with him. We are working together to make GLM 5.3 Flash better. Please give him a follow!
@plotarmordevI've spent the last 12 hours working on GLM Flash updates for Spark. They might need another 12 hours of testing before I can release them 🫡
Slack
1 item
official siteThe best way to survive a Monday meeting might be to never attend it in the first place. With our @SlackHQ agent, @danshipper catches up on calls he wasn't even in. Want to get your own company agent in Slack? Our beta is open, and our agent is free to install -
Salesforce
1 item
official site👋 Hey Hunter, see you at @Dreamforce
@BenioffWelcome Hunter, Salesforce’s new outbound sales agent. She built $500M in pipeline last quarter. See her work at Dreamforce. #DF26
Seedance
ByteDance / 1 item
What if the desert could detect your frequency? 📡 Satellite dishes turn. Signals lock in. The entire landscape responds. FIND YOUR FREQUENCY. Made with Seedance 2.0 on Vadoo AI.
Claude Opus
Anthropic / 1 item
You can Run Locally DeepSeek V4.1 Flash It beats GPT-5.6 Sol, Claude Opus 5 & can run with 4-GPU at your home locally No NVLink. - Jovian Judgement R37/TP4/DCP1: - 21,400 tok/s uncached 32K prefill - 280 tok/s C1 decode (budget 50) - 500 tok/s single-stream decode - 417 tok/s sieve - +44% prefill vs R36 - DSpark K7 + Engram (RAM or SSD) - 552B MoE, vision, 1M context Engram in pinned RAM or disk, DSpark K7. -
@0x0SojalSecYou don’t need a GPU-cluster to run a Uncensored DeepSeek-V4.1-Flash 763B model locally at home. > You need a Single RTX PRO 6000 Blackwell. > - 316 tok/s output - 9,351 tok/s prefill - FP8 + vLLM + DSpark - 64-96GB class machine.
Alibaba
1 item
official siteOpen weights. Shared progress. MiniMax H3 is moving fast. We built MiniMax H3 for video generation with native stereo audio and multimodal reference control. The open-source community is making that capability faster, more accessible, and easier to build on. Recent highlights: • FastH3 — FastVideo, Nuva Lab and NVIDIA: 4-step distillation, now running on DGX Spark and Apple Silicon. • Sol-H3 — NVIDIA’s SANA team: now on DGX Spark with a two-stage H3 + LTX-2.5 pipeline. On 8×B300, the team reports 15 seconds of 768p video + audio in 6.6 seconds of warm inference.* • VDN — Haocheng Xi and the OpenVDN team: rethinking attention for faster H3 inference, with weights, training and inference code released. • PDD — NVIDIA’s distillation method, brought to H3 by Alibaba PAI as 8-step Acc-LoRAs, now supported in ComfyUI. • LightX2V — 4- and 8-step Turbo LoRAs, with workflows for text, image and reference-conditioned video + audio. Behind every release are people training, optimizing, quantizing, testing and sh
MiniMax H3
MiniMax / 1 item
Open weights. Shared progress. MiniMax H3 is moving fast. We built MiniMax H3 for video generation with native stereo audio and multimodal reference control. The open-source community is making that capability faster, more accessible, and easier to build on. Recent highlights: • FastH3 — FastVideo, Nuva Lab and NVIDIA: 4-step distillation, now running on DGX Spark and Apple Silicon. • Sol-H3 — NVIDIA’s SANA team: now on DGX Spark with a two-stage H3 + LTX-2.5 pipeline. On 8×B300, the team reports 15 seconds of 768p video + audio in 6.6 seconds of warm inference.* • VDN — Haocheng Xi and the OpenVDN team: rethinking attention for faster H3 inference, with weights, training and inference code released. • PDD — NVIDIA’s distillation method, brought to H3 by Alibaba PAI as 8-step Acc-LoRAs, now supported in ComfyUI. • LightX2V — 4- and 8-step Turbo LoRAs, with workflows for text, image and reference-conditioned video + audio. Behind every release are people training, optimizing, quantizing, testing and sh
Claude Fable
Anthropic / 1 item
AIエージェント「Devin」を開発するCognitionが、AIモデル「GPT-6 Astra」や「Claude Fable」の性能を最大限に引き出すハーネス(AIモデルの実行環境・仕組み)「Fusion」を開発した。
Devin
1 item
AIエージェント「Devin」を開発するCognitionが、AIモデル「GPT-6 Astra」や「Claude Fable」の性能を最大限に引き出すハーネス(AIモデルの実行環境・仕組み)「Fusion」を開発した。
Also recorded
22 items that named no organisation or product this site tracks.
Put your team’s knowledge to work. Developed by Tencent’s Weixin team, WeKnora is an open-source, LLM-powered knowledge framework built for enterprise-grade document understanding, semantic retrieval, and autonomous reasoning. Build assistants that find answers in your docs, show their sources and use skills to create files you can preview and download. Keep useful context across chats with opt-in long-term memory. You approve suggested memories before they’re used. Self-host it and run skills in a sandbox with network limits you set. Explore WeKnora: Have a project in mind? Join our community to share ideas and build with us:
Benchmark-leading scientific reasoning meets long-horizon agent capabilities. Intern-S2-397B is now available in BF16 and FP8.📜 Apache 2.0. 🤖 🏆 Scores 87.0 on FrontierScience-Olympiad and 84.0 on SWE-bench Multilingual, leading the reported comparison on both. Several specialized scientific benchmarks also see top results. 🔬 Raw scientific pages become training material directly, preserving text, visuals, symbols, and their relationships without intermediate parsing. 🧪 Joint training across 20+ scientific domains covers scientific reasoning, biomolecular interaction design, material generation, and time-series forecasting. 🛠 Long-horizon agent RL strengthens tool use and sustained execution, with thinking and non-thinking modes available.
We open-sourced CubeSandbox for a reason: so builders tell us what's broken. They did: → Cross-node pause/resume. Pause on one node, resume on another. Snapshot restore and stability both improved. → Component failure used to take the cluster down. CubeMaster multi-replica, template service split out. No single point of failure left on the core path. → LLM timeout: 60s proxy cutoff was killing slow-first-token models. Now 2h. Next thing builders told us @CubeSandbox : pause/resume memory cost is too high, on it.
@TencentAI_News🥳We just open-sourced Cube Sandbox! An instant, concurrent, secure and lightweight sandbox runtime for AI Agents. > Built with RustVMM and KVM, it achieves the perfect balance of security and performance: > → Sub-60ms cold start (2.5-50x faster) → Under 5MB memory overhead per instance (6x less memory) → Dedicated kernel per sandbox (hardware-level isolation) → Thousands of concurre
Sunday ship: built-in herdr support in Command Code Zero setup. Shipped v1.54.0. · working / idle / blocked in the sidebar · works in cmd -p, not just the TUI · releases the pane on exit, even on kill · /herdr to check status 🐐
This is pretty cool for artists, have an idea? Hum it and get a full song to work with and iterate on!
@ModelScope2022🎶Drop in a beat, a hum, or a vocal. DiffSynth-Music turns it into a complete song. 📜 Apache 2.0. 🤖 📄 > 🎛 Five audio controls cover beats, vocals, accompaniment, prosody, and reference audio, alongside native text-and-lyrics music generation. 🎤 Keep an input vocal and generate the backing track, or preserve the accompaniment and create new vocals around it. 🎼 Prosody control transfers pitch, rhythm, and syllable timing to new lyrics. Reference control captures broader style, melody, vocal character, and timbre. 🧠 Composable KV-cache adapters compute each audio condition once and reuse it throughout sampling. Evaluations show stronger adherence across all five control types while retaining broadly comparable music quality.
With 40M+ users served and 5M business apps created, Miaoda's latest upgrade expands its no-code platform for both enterprises and individual creators. 🚀 The upgrade brings enhanced AI agents for design, app generation, and testing, along with enterprise tools for private deployment and collaboration, and a marketplace connecting businesses with creators for templates and custom development. A single prompt can now be your first step toward building a business.
Crash recovery for agent payments is here. Raxol helps restarted agents track pending transfers instead of sending funds twice, a foundation for agentic trading.
@axol_ioTry out Raxol now at > Most agents are a Python process wrapped around a model. Raxol is an OTP runtime. That difference decides what an agent can survive, how many can run at once, where they can render, and what they can safely be allowed to do.
Yifan has big ideas but I wouldn't overindex on this until a sizable experiment
@yifanzhang_We are at the dawn of Superintelligence. > Introducing the Recurrent Looped Transformer (RLT), > We now have Transformers with Infinite Reasoning depth. > From now on, we should pace progress at the Open Frontier of Superintelligence, > Until Safe Superintelligence is achieved. >
Habibi, welcome back to the desert. Uncle said the money would follow. Then the whole desert started dancing 😂 Built this one with Ray 3.2 inside @LumaLabsAI Used Ray 3.2 to generate this video, and I was honestly amazed by how much control I had over the camera, pacing and character beats in one flow. Orbit, tracking, push-ins, everything felt more directed and less random🔥
This is actually pretty massive and I suspect not many understand what it means. TL:DR - you can train LORAs now on your own music (or any music) and then generate your own novel songs or do audio covers etc. It works so incredibly well too.... its 100% worth testing
@sin_ceriously> I trained the missing encoder for YuE2, so we can all bring our own music into it.
卧槽!我的MacDuo有应用价值了😄 今天看到这位老哥的这个idea,我没有做AirPods的追踪来判断的模糊效果。 我做了用户不看屏幕或者离开电脑就会进行模糊的效果,支持自定义时间以及对应的锁屏密码开启登。 只要眼睛不看屏幕就可以模糊,其实可以做个番茄钟模式😁 感兴趣的去试试,GitHub已经刷版。 地址:
@bryllim_Inspired by the iPhone Duo, I built a Mac app that blurs the side of your screen when you’re not facing your laptop screen to hide what you’re working on and uses Airpods for head tracking.
crazy.. the first AI Design Agent Team just dropped. you give OJO one rough idea → its AI Design Agent Team turns it into product logic, interaction design, visual systems, motion, a working prototype, and dev handoff.. all on one canvas even a Steve Jobs perspective is just another stackable skill Tutorial below:
How do you build personalized AI tech that truly understands what matters to you? With Dreambeans from @GoogleLabs, we're safely connecting the dots across your digital life. Instead of siloing pieces of data one app at a time, Dreambeans gathers details from connected sources like @Gmail, Calendar, Search, @GeminiApp, and through face grouping in your @googlephotos. It might notice your friend Beth's birthday is coming up, generate a one-of-a-kind illustrated story of the two of you, and surface the perfect gift ideas. Rather than doomscrolling an endless feed, you’ll get a daily in-app notification from Dreambeans when your customized “stories” are ready. And when a story clicks with you, you can tap the illustrated tile to find additional info and direct links to the next step, whether that’s watching a movie trailer or buying a suggested gift. This experience is opt-in, transparent, and protected by strict privacy filters. And because you are always in control, you can simply tap the "thumbs down" butt
上海 AI Lab 和上海交大发布 89 亿参数模型 NCP-ArchPreview。普通大模型主要训练「猜下一个 token」,NCP 还会预测后面一小段对应的内部概念信号,再用它辅助 token 生成。最终仍然逐 token 输出,但模型内部已经会提前猜后文。 这套信号还能拿来给投机解码提速。DFlash、DSpark 这类方案会让一个小模型先猜后文,再交给大模型统一检查。NCP 团队把内部概念信号也交给这个小模型,相当于先给它一点「后面大概怎么写」的提示,再让它一次并行猜 16 个 token。 这样一来,小模型猜中的内容更多。四项测试里,大模型每验证一次,平均能通过的 token 从 5.933 个升到 6.180 个,提升 4.17%;HumanEval 上提升 7.59%。接入这套信号只给草稿模型增加约 4 万参数。 同样是 89 亿参数,NCP 用普通 Transformer 约 85% 的训练计算量,就能把训练误差降到差不多的水平。
@mark_kA new AI architecture may have just cut pretraining almost in half. > NCP-ArchPreview is an 8.9B-parameter latent-space language model that doesn't just predict the next token. It also learns to predict discrete concepts spanning multiple tokens, then feeds those concepts back into token generation. > The researchers say it matched OLMo-3-7B's final pretraining loss using only 51.3% of the training tokens, while finishing 2.45 points ahead across downstream benchmarks, including +5.99 on GSM8K. The model weights, training recipes, and
I asked one question about a 7,000-file repo from the terminal. Then I ran the same agent with no interface at all and got a JSON file back. @minimaxagent CLI is the desktop app with the window removed. Same agent, same account, same sessions -- pointed at a repo instead of a chat box. Inside the TUI: > mcode init reads the repo once and writes AGENTS.md, so every later session starts informed. > Ctrl+O opens its reasoning and the files it touched. > Shift+Tab cycles Ask, Auto and Full access. > /usage shows the bill, /sessions brings back any old thread. I asked where Django decides which middleware runs on a request, and in what order. It came back with django/core/handlers/base.py, BaseHandler.load_middleware, line 41 -- and it was right. Then the part the desktop app cannot do: > mcode exec runs it headless. --permission sets what it may touch unattended, --output-format json writes straight to a file. > The same call drops into CI, a Makefile, or cron. Three ranked risks with file and line numbers
LTX-2.5 is out. Open weights, your hardware, your IP. It runs on any GPU and comes with low latency generation for live pipelines, native 4K and HDR, stable motion and auto duration set. It’s truly crazy what this model can do and that we can just run it on our own hardware.
Install this. Your UI will thank you.
@JakubantalikOne skill for all UI motion. > The Transitions-dev skill uses content and knowledge from the entire Transitions library and applies it to your project wherever it makes sense, for a more polished and snappy UI.
You're on a job. A new leads comes through. This skill lets your AI agent reply immediately, ask follow up questions, and book the visit. You come back to more qualified leads on your calendar. Get the skill on Hyperagent Marketplace:
Great ideas don’t have to stay in your head🔥 GPT-Image 2.5 + Lovart Pencil Turn any-size sketches into finished images, multiple styles.
I give up. We now stream responses in T3 Code by default. (Not token by token because that makes zero sense)
We gave GPT-Image 2.5 A MESS. 🧱 Random bricks. Not-so-random results. Now on WaveSpeedAI.
Runs a 1-billion parameter LLM on a $10 board with zero dependencies