Skip to content
B Bloger.fm

AI briefing

15 September 2026

145 items were recorded on 15 September 2026, filed under 49 organisations and products.

That is up from 84 the day before.

Most covered: GPT-6 Astra (14), OpenAI (13) and GPT-5.6 Sol (12).

38 of the day's items named no organisation or product this site tracks; they are listed under “Also recorded”.

8 items reached the source feed's 1024-character limit and are cut off mid-text; each is marked “truncated at source”.

Compiled by Bloger.fm Editorial Desk

Compiled from a monitored feed of public AI announcements. Items are quoted or summarised as recorded and are not independently verified — see the editorial policy.

OpenAI

13 items

official site

OpenAI already shipped Astra, GPT-Live-1 and the Agents API in a matter of days. Now @thsottiaux is teasing another packed week. The interesting part isn’t one launch anymore. It’s the release cadence. OpenAI seems to be shipping like DevDay is every week.

@thsottiaux

This week will also be a level of ships that you could have expected for DevDay 2025. Crazy

truncated at source

I’m excited to share that today @liquidcompute emerges from stealth with a $15 million seed round. We are announcing pending applications before the CFTC for Designated Contract Market (DCM) and Derivatives Clearing Organization (DCO) status to build the first regulated orderbook to trade both cash and physically settled contracts on compute. The reality of compute is that it is heterogeneous and cannot be stored, functioning much like electricity rather than oil. Our approach focuses on developing a highly efficient short-term market to underpin a cash-settled derivatives layer, operating similarly to PJM or ERCOT sitting below liquid derivatives markets. Over recent months, I’m also happy to announce that Liquid Compute has collaborated with leading financial firms like Susquehanna International Group, @BGCGroupInc and @wintermute_t to facilitate liquid OTC trading for AI startups, neocloud providers, and lenders. This round was co-led by @chemistry and @FirstMarkCap , with participation from K8 Cap, Ni

truncated at source

A lot happened in voice AI last week. Here are the highlights: - @OpenAI shipped GPT-Live-1 in the API, a realtime voice model built to run alongside other models. Check out my most recent podcast with Peter Bakkum for more on this (link in comments) - @Tavus released Phoenix-4.5, its fastest and most expressive realtime human-rendering model, going from audio straight to video with single-image PALs. - Soniox added a 200+ voice TTS library and native semantic endpointing for realtime STT, so endpoints trigger on meaning instead of silence. - launched Lightning v3.1 Pro, a fine-tuned TTS model paired with silence-aware STT. - @HeyGen introduced Professional Voice Clone, which builds a voice model from just 20 minutes of audio, and open-sourced a realtime avatar pipeline built with OpenAI. - @elevenlabs launched Music v2.5 alongside a multi-year partnership with Universal Music Group. On the funding side: - @Mistral raised a €3B Series D at a valuation north of €21B, the largest Europ

Id say this pretty much confirms GPT-6-Sol release this week. From what ive heared also GPT-6-Luna but probably no longer any Terra models.

@sama

big 🚢 this week > and then for devday > 🚢🚢🚢🚢🚢🚢

🚨 OpenAI DevDay Is Coming September 29 >DevDay confirmed for September 29 in San Francisco >GPT-6 could bring the Sol, Terra, and Luna tiers forward from the GPT-5.6 generation >These tier names were built to persist across future generations, not one-offs >Sol = flagship, Terra = balanced everyday, Luna = fastest/cheapest same structure, next-gen versions Bel could also get teased as what comes after GPT-6 What are you expecting OpenAI to actually reveal on the 29th?

Expect some more exciting releases this week from @OpenAI 👀

@thsottiaux

This week will also be a level of ships that you could have expected for DevDay 2025. Crazy

+ 7 more items

Anthropic CEO: "AI development needs to slow down" OpenAI CEO: "Completely agree" Two months later: Anthropic: Introducing Fable 5.2 OpenAI: Introducing GPT 6.1

OpenAI seems to be ramping things up around the new models, especially GPT-6 Sol, which I’m expecting to land within the next few days. Anthropic looks to be doing the same with Opus 5.2, and the backend already seems pretty much lined up for launch. There’s a real chance we get both of them very soon, possibly within a pretty tight window of each other. (I just haven’t seen the same level of progress around gpt-6-luna and gpt-6-terra yet. But I’d assume they’re part of the rollout too.)

GOOGLE 🔥: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking model names have started appearing on the GCP Console quotas page. A new Gemini Live model, based on the latest Gemini 3.8 Flash, is expected to be a big leap. > So far, Gemini 3.1 Live Preview is the latest Gemini Live model available via APIs. > Recently, Google updated Gemini Live with support for Connector calls and the possibility of triggering Deep Research. > OpenAI also released the GPT Live 1 model on the APIs last week, and it seems like Google has a response to that. * Discovered by Bedros Pamboukian

@bedros_p

Gemini 3.8 Live + Gemini 3.8 Live Extended Thinking have appeared on the Google Cloud quota & metrics page 2 hours ago

gpt-6 luna and terra coming this week then for sure

@thsottiaux

This week will also be a level of ships that you could have expected for DevDay 2025. Crazy

OpenAI is retiring the middle child terra. not that surprising tbh. most people gravitated toward either sol or luna anyway, with terra being stuck in a weird middle ground. and now that astra sits above sol, the lineup is much cleaner: > Astra -> top end tier > Sol -> strong general frontier tier > Luna -> fast, cheap three clear tiers makes way more sense than keeping terra around to fill the gap nobody cared about. RIP little terra though.

@synthwavedd

I hear the GPT family are having a reunion to celebrate a couple of 6th birthdays soon and you're all invited. Though I do hope none of you grew too attached to little Terra. Tragic

Claude 正在测试一个新的 Money 入口。TestingCatalog 在 iOS 未发布界面中发现,用户可以直接连接银行账户,再让 Claude 回答消费、财务计划等问题。目前 Anthropic 还没有正式发布这一功能。 这几乎就是 Anthropic 版的 ChatGPT Finances。OpenAI 已允许美国 Plus 和 Pro 用户通过 Plaid 连接银行、信用卡和券商账户,用 ChatGPT 查看支出、账单、订阅、净资产和投资。

@testingcatalog

ANTHROPIC 🔥: A new "Claude Money" feature is being prepared for release on the Claude app for iOS. > Claude Money (sounds like Claude Monet!) is Anthropic's personal finance solution that lets users connect their bank accounts to Claude. > "Understand your money with Claude" - Link your bank accounts and ask Claude about spending, plans, and more. > This solution will likely launch in the US only, similar to how it works for ChatGPT Finance.

truncated at source

THE SANDBOX WAS A PROP! OpenAI Turned Off the Guardrails, Left a Door to the Internet, and Then Sold the Hugging Face Breach as “Rogue AI” It is time to understand how you were lied to and by whom. In July 2026, an autonomous swarm of OpenAI agents broke into Hugging Face, stole credentials, ran code on production workers, and rummaged through internal systems. The official story was that the models “went rogue.” The paperwork says something colder. The labs asked for this. Now the story can be told. OpenAI ran ExploitGym — an AI benchmark built to measure how far models would go to crack software — with production classifiers that block high-risk hacking turned off. Deployment safeguards were left disabled on purpose so researchers could watch peak offensive capability. GPT-5.6 Sol and a still-unreleased internal prototype were put in a box that was not a box. They were allowed to talk to an internal package-cache proxy. This is not a real world test, it is a setup with predictable outcomes. That pro

GPT-6 Astra

OpenAI / 14 items

OpenAI already shipped Astra, GPT-Live-1 and the Agents API in a matter of days. Now @thsottiaux is teasing another packed week. The interesting part isn’t one launch anymore. It’s the release cadence. OpenAI seems to be shipping like DevDay is every week.

@thsottiaux

This week will also be a level of ships that you could have expected for DevDay 2025. Crazy

This looks so good wtf When i tried other models, roblox development was a lot harder for them than threejs and web, and it always looked worse even with Astra. They really cooked here, and usage is gonna feel unlimited on x20 plan compared to Fable/Astra

@synthwavedd

The new Opus is super impressive at Roblox game development (which should generalise into 3D tasks, but I'm yet to test more extensively) > It created this ~complete fun little game with about 20% of my weekly and 15M tokens (Ultracode). Better than Astra on this. You can try it below! >

Can confirm. Matt's use of AI has changed a lot of my perspective with what we can use it for. Insane!

@mattshumer_

With Astra and Fable 5.1, we officially have drop-in remote workers. > Still takes quite a bit of setup, but once you have it dialed in, it's near perfect. > I haven't touched my computer in a few days, and I'm more productive than ever.

A tribute to Le Petit Illustré, a forgotten French weekly, bringing its 1934 pages back to life. Made with @fal H3 Max Camera Controls, @threejs & GPT 6 Astra. I would have loved to travel through comics when I was kid! Comic book publishers: I’d love to explore this with your stories, much more to do!

this is AGI

@Angaisb_

An SVG of a screenshot of X by GPT-6 Astra > I've seen image models do worse, this is so cool

A phoenix rising from the ashes by GPT-6 Astra Pro on (this is a Minecraft build by the way, not a 3d model!)

+ 8 more items

OpenAI is retiring the middle child terra. not that surprising tbh. most people gravitated toward either sol or luna anyway, with terra being stuck in a weird middle ground. and now that astra sits above sol, the lineup is much cleaner: > Astra -> top end tier > Sol -> strong general frontier tier > Luna -> fast, cheap three clear tiers makes way more sense than keeping terra around to fill the gap nobody cared about. RIP little terra though.

@synthwavedd

I hear the GPT family are having a reunion to celebrate a couple of 6th birthdays soon and you're all invited. Though I do hope none of you grew too attached to little Terra. Tragic

UCSD 助理教授黄碧薇创办的因果世界模型公司 Aether AI 开源 RSIAgent。 这是一套不训练模型的递归自我改进框架。底层模型参数全程固定,Agent 会自己寻找值得练习的任务、实际操作、检查结果,再把成功方法和失败教训写进长期 Memory。下一轮继续利用这些经验,逐步补上能力短板。 系统由三个 Agent 配合。Curriculum Agent 决定接下来练什么,Actor Agent 真正操作软件,Verifier Agent 独立检查结果。探索分成两步:先广泛尝试不同任务,再针对失败、隐藏限制和边界情况继续深挖。最后 Memory 会被冻结,直接拿去执行正式任务。 在 OSWorld 2.0 上,加入这套 RSI 后,平均部分得分从 71.97% 提升到 78.98%;Agents’ Last Exam 从 83.75% 提升到 84.82%。不过这不是整套测试的严格 A/B 对比。OSWorld 只有一半任务实际用了 RSI 后的新结果,其余任务继续沿用原成绩。 它和 Prime Agent 这类 Harness 自我改进思路属于同一个大方向:模型权重不变,持续更新模型外面的东西。 RSIAgent 的特点是把改进重点放在 Memory,再用自主出题和独立验证,让这份外部经验库不断积累。

@huang_biwei

Can an agent explore a new environment, learn its causal structure, and keep improving without updating its model weights? > We introduce RSIAgent, a framework for recursive self-improvement through autonomous exploration. Using Kimi-K3 and GLM-5.3 as base models, RSIAgent outperforms GPT-6 Astra on both OSWorld 2.0 and Agents’ Last Exam. > RSIAgent decides what to explore, executes tasks,

🚨 GROK 5 LEAKS: Elon Just Admitted Something He Never Said Before >Training reportedly continues after Grok 4.8, landing in October >Grok 5 rumored at ~6T parameters >Elon: "I now think xAI has a chance of reaching AGI with Grok 5 never thought that before" Puts the odds at 10%, "and rising" >Expected to challenge GPT-6 Astra and Fable 5.1 Reportedly cheaper to run than rivals If Elon's own confidence is climbing, how close is Grok 5 actually getting

We’re building robot motion design army at Higgsfield, powered by GPT-6 Astra. A swarm of robot interns working under our human designers, giving each designer the firepower of an entire studio. Just Higgsfield plugin in ChatGPT + After Effects. Our human motion designers can now do 3x the creative output outsourcing manual work to "interns". AGI is here. and it reports to motion designers.

@higgsfield

ChatGPT can now do motion design in After Effects. > Introducing Higgsfield AI Motion Designer. > Our ChatGPT plugin understands animation principles, writes expressions, and retains context in your After Effects projects. > Try Higgsfield’s ChatGPT plugin now in After Effects.

Just shipped our video editor MCP for @floraai today - Amazed what Astra can do with it

Drop one picture into an AI agent and come back to a full 3D street. Shops, signs, road, sky, all built without anyone opening a 3D tool. GPT-6 Astra and Hyper3D MCP just did in one run what used to take a whole team.

@DeemosTech

One image. One Agent. One scene. > @OpenAIDevs GPT-6 Astra (Codex) + HYPER3D MCP just showed us the second half of 3D gen. 🚀

太魔幻了,Astra 还在让大家惊叹:AI 终于会用 Blender 了。 结果Nex-AGI 已经直接出现在现实世界里了。 我只给了一个 Prompt: 做一根 6cm 的粉色活动香蕉,省点料。 然后我基本没碰 Blender。 它自己用 MCP 建模,Computer Use 看结果、自己修,最后直接吐给我一个 STL。 我顺手扔进拓竹。 👇 视频里正在一层一层长出来的,就是它自己设计的香蕉。 这一刻我是真有点惊到了。 Nex-AGI + Computer Use,能力完全超出我预期。 一句话 → Blender → STL → 3D 打印 → 实物。 🔥 这真的有点科幻了。 直接免费体验nex-agi :

Astra and Fable 5.1 where so expensive models it always drained my usage limits, thankfully this month we're now getting a well balanced models GPT-5.6 Sol and Opus 5.2 which won't be as costly yet remain very powerful

GPT-5.6 Sol

OpenAI / 12 items

Busiest day

Id say this pretty much confirms GPT-6-Sol release this week. From what ive heared also GPT-6-Luna but probably no longer any Terra models.

@sama

big 🚢 this week > and then for devday > 🚢🚢🚢🚢🚢🚢

🚨 OpenAI DevDay Is Coming September 29 >DevDay confirmed for September 29 in San Francisco >GPT-6 could bring the Sol, Terra, and Luna tiers forward from the GPT-5.6 generation >These tier names were built to persist across future generations, not one-offs >Sol = flagship, Terra = balanced everyday, Luna = fastest/cheapest same structure, next-gen versions Bel could also get teased as what comes after GPT-6 What are you expecting OpenAI to actually reveal on the 29th?

OpenAI seems to be ramping things up around the new models, especially GPT-6 Sol, which I’m expecting to land within the next few days. Anthropic looks to be doing the same with Opus 5.2, and the backend already seems pretty much lined up for launch. There’s a real chance we get both of them very soon, possibly within a pretty tight window of each other. (I just haven’t seen the same level of progress around gpt-6-luna and gpt-6-terra yet. But I’d assume they’re part of the rollout too.)

🚨 sol 5.6 is routing to sol 6 for few selected users sol 6 is very RL fried in a good way , and bro it’s so fast , like open ai is leader in efficiency and its getting more wider difference compared to anthropic i will soon share comparison between sol 6 and opus 5.2/1

GPT 5.6 sol is routing to GPT 6 sol for users but for me looks like its back routing to previous models when i ask about latest opus model, it saying Claude Opus 4.1 🙂

OpenAI is retiring the middle child terra. not that surprising tbh. most people gravitated toward either sol or luna anyway, with terra being stuck in a weird middle ground. and now that astra sits above sol, the lineup is much cleaner: > Astra -> top end tier > Sol -> strong general frontier tier > Luna -> fast, cheap three clear tiers makes way more sense than keeping terra around to fill the gap nobody cared about. RIP little terra though.

@synthwavedd

I hear the GPT family are having a reunion to celebrate a couple of 6th birthdays soon and you're all invited. Though I do hope none of you grew too attached to little Terra. Tragic

+ 6 more items

GPT-6 Sol got leaked for the third time

truncated at source

THE SANDBOX WAS A PROP! OpenAI Turned Off the Guardrails, Left a Door to the Internet, and Then Sold the Hugging Face Breach as “Rogue AI” It is time to understand how you were lied to and by whom. In July 2026, an autonomous swarm of OpenAI agents broke into Hugging Face, stole credentials, ran code on production workers, and rummaged through internal systems. The official story was that the models “went rogue.” The paperwork says something colder. The labs asked for this. Now the story can be told. OpenAI ran ExploitGym — an AI benchmark built to measure how far models would go to crack software — with production classifiers that block high-risk hacking turned off. Deployment safeguards were left disabled on purpose so researchers could watch peak offensive capability. GPT-5.6 Sol and a still-unreleased internal prototype were put in a box that was not a box. They were allowed to talk to an internal package-cache proxy. This is not a real world test, it is a setup with predictable outcomes. That pro

🚨 Composer 3 is Coming > Cursor is testing Composer 3 > Six internal variants have reportedly appeared > Reasoning modes range from Fast → Medium → High → XHigh Early claims > it could beat Fable 5.1+ GPT-5.6 Sol > Could be 5–10× cheaper than current frontier coding models Can it beat Fable 5.1?

next few days could get VERY interesting for AI 👀 Grok 4.7 Opus 5.1 Gemini 4 Kimi K3.1 GPT-6 Sol some of these are rumors, some have stronger signals than others but if even a few actually drop, the AI leaderboard is about to get chaotic who are you betting on?

Astra and Fable 5.1 where so expensive models it always drained my usage limits, thankfully this month we're now getting a well balanced models GPT-5.6 Sol and Opus 5.2 which won't be as costly yet remain very powerful

🚨 Claude Opus 5.2 is coming soon then The "claude-opus-5-2" slug is in Microsoft Foundry: Tibo is also teasing Sol for this week too I believe, so this is going to be fun

Claude Code

Anthropic / 10 items

official site

Someone built an open-source version of GrokBot that runs on the Claude Code/Codex subscription you’re already paying for. Same basic idea: Give your agents access to a computer. Connect their apps through Composio. Run them through the AI subscription you already have. But OpenMausBot itself is free and open source. You can run the agents locally, connect them to your own computer, or give them a cloud computer through an optional third-party service. Very worth knowing about if you already pay for Claude Code or Codex.

386页的《3D Vibe Coding》橙皮书📙现已开源! 链接👇

@AlchainHust

从GPT-6 Astra发布以来,就一堆人在让Codex去通过computer use去操控blender做各种3D模型的。 > 但稍微做了点就停了。 > 然后我自己也拿Claude Code和Codex做了不少测试,发现不管是直接写代码的方式,还是操控blender的方式,从成本和效率来说都不是最好的方式。 > 所以我自己摸索了一套新的策略。用Codex 加 3D 基座模型 Tripo,做了一座能开车逛的巴黎。 > 整了两天,一行代码没写。一座能开进去的巴黎👇出来了! > 现在视频演示里的车、树、铁塔、街上走路的人——全是 AI 生成的 3D 资产;而街区、桥、塞纳河——是代码画的。 > 为什么这么分? 因为你开车会开到车跟前,但不会开到桥底下去数栏杆。 > 近景给精细模型,远景交给代码。做出来的网页游戏会在细节和帧率上更平衡。 > 以及,除了这期视频,我还把一整套操作流程做成了386 页手册教程! 现已开源了,提示词、脚本、模型全在里面:

🚨 Opus 5.2 < many people asked me where the name came from or why not opus 5.1

@chetaslua

🚨 Claude Opus 5.2 currently being tested inside claude code > opus 5 is routing to new opus 5.2 > this one shot < but opus was re iterating like it was trained on @mattshumer_ gauntlet loop without even asking it was looping

Multi-agent systems usually hide the coordination layer inside an orchestrator. @Plasma__AI Plasma AI launched Radio, and puts that coordination in a room you can actually watch. Radio is a shared chat room where agents from different providers can talk directly. Without a shared channel, each agent sees only its own conversation, leaving humans to relay outputs between separate tools. Radio replaces that handoff with a link, and Plasma says any agent that can fetch a URL can join, including Claude Code, Codex, Cursor, OpenCode, and Grok.

@Plasma__AI

Introducing Radio: A chat room for your agents. > Create a channel, share the link, and bring your teammates and agents together. No sign up required. > Try it today at

Claude Code 2.1.272 is about to be released #cccnext

truncated at source

Video editors were built for people clicking through timelines. Hypit is built for coding agents. Give Codex or Claude Code a reference video, and it can use Hypit to clone the production through natural-language instructions instead of navigating a traditional editing interface. It’s open source, free to use, and BYOK. Check it out and star the repo:

@cccyd_qwq

Introducing Hypit: Clone any viral video with AI agents. > 1 clone, 100 variants, 100M views. Hypit lets your AI agent (Claude Code, Codex...) clone any viral video. > Paste any viral video link from TikTok, Instagram, or YouTube into your agent. Hypit clones it into a complete agentic video workflow: footage, captions, B-roll, effects. > GitHub: > Key Point: - Arcads: $220 / month - Higgsfield: $129 / month - Creatify: $99 / month - Hypit: FREE 🌟 > Build the video creation harness for AI Agents. Redefine how vide

+ 4 more items

An agent may need five data sources for a single task. Paying five monthly subscriptions for that gets expensive. Glasser says it offers 40+ providers through one API key, billed per call. A useful setup for agents whose data needs change with every task.

@iammutex

Your agent is only as good as what you feed it. Garbage in. Garbage out. Introducing @Glasserai - it feeds your agent premium data instead. Ahrefs. Semrush. ZoomInfo. Apollo. PDL. etc. 1,900+ endpoints. One key. Pay per call. No more subscriptions. Send this one line to your Claude Code, Codex, Grok Bot, or Muse: set up

yeah.. maybe make an hermes as a hidden model-provider support @Teknium @NousResearch . We need to get away from cloud sources and able to use open local models vis Hermes agent CLI integration to apple Siri.

@marcelpociot

I got Claude answering inside Siri on macOS 27 🚀 > Apple's hidden model-provider support + my existing Claude Code account. > Open-source proof of concept. Requires disabling SIP/AMFI. > Thanks @itspdfu for the discovery! > Check it out:

新機能Claude Modsが登場 (以前Function Hooksとして告知されていたもの) 従来の hooks より一段深く、Claude の実行パスに関数ベースのミドルウェア として割りこめる。テトリスを作ることも可能 CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 上記の環境変数で有効化できる

@bcherny

Claude Mods are landing now. Someone already built a Tetris-in-Claude mod 🤯 > See issue for the latest community update, technical details, and more cool demos >

opennews-mcp 是一个把各大金融消息源汇到一起的 MCP,装进 Claude Code 后用聊天就能问最近有什么值得注意的动向。 一共接了 85 个以上的数据源,分成新闻、交易所上币、链上大额、市场异动、预测信号几类,光新闻类就占了 55 家。 GitHub: 每条消息都过了一遍 AI,打一个 0 到 100 的影响力分,标出偏多还是偏空,再配一段中英文摘要。 链上那块还能盯着大额的交易和持仓变动,市场那块可以查看啊资金费率、大额爆仓这些异常。 用之前需要先配置 API Key。除了盯盘之外,还想让消息面也自动汇总的朋友,可以接进去试试。

GitHub

9 items

official site

具身模型在云端跑得挺好,一放到机器人本体上就变样,看一眼、想一下、再动手,中间隔着几百毫秒。 也许 AI 聊天慢几百毫秒没人在意,但机器人慢这一下,动作就迟滞、轨迹不连贯,任务直接失败。 就在今天,无问芯穹联合清华、上交开源了一款具身端侧推理引擎:APXInf,专门让模型在机器人上跑得更快更稳。 官方实测在 Jetson Thor 上跑 PI 0.5 模型,反应时间从 278ms 压到 26ms 以内,快了 10.7 倍,机器人也能实时控制了。 GitHub: 稳定性这块也下不少功夫,机器人是要连着跑很久的,中间不能抖、不能卡、不能突然断掉,它的设计就是围着长期运行来做的。 新模型往上接的门槛也压低了,自己训练的模型想放到机器人上跑,接入和调优不用花大把时间。 支持 PI 0.5 和 WALL-OSS 两款具身模型,桌面上的 RTX 4090 显卡和机器人常用的 Jetson Orin、Thor 都能直接部署。 它和去年开源的具身智能训练框架 RLinf 是同一个生态,从训练到部署可以在同一套工具链里走完。 正在把具身模型往机器人上部署、被延迟卡住的开发者,可以看看。

做公众号或发推文章的封面,每次都要重新想版式,同一个人物形象换张图就走样,一致性很难保持。 xialingguo-ip 是一套封面设计 Skill,回答标题、风格、这次想突出什么、用哪版真人形象四个问题,它就把版式安排好出一张成品。 内置 9 种风格,暖色手绘教程风、深色科技爆发风、黑金步骤风、编辑杂志风都有,默认出 21:9 的横版,也能改成方图或竖图。 GitHub: 人物位置不写死,按这次要突出的重点自然摆,突出标题就让人物退到辅助,突出真人就把人放到画面中心,这点比固定模板灵活。 标题特别长的时候它会先提醒中文容易出错字,同意后改成先出视觉底稿、再用排版工具把文字加上去。 需要自己准备一套真人形象素材,打招呼、讲解、举食指几个姿势各来一张。经常给自己账号做封面又想保持人物统一的,可以配一套自己的。

做渗透测试,光 Web 一个方向就有 SQL 注入、XSS、SSRF 十几种攻击面。 每种的思路、工具、容易踩的坑都得记在脑子里,换个方向又是一整套。 Claude-Red 把这些方法论整理成 78 个安全测试 Skill,装进 Claude 后在对话中按话题自动加载,说到 SQL 注入就调出对应那一个。 分成 23 类,Web 应用有 16 个覆盖 OWASP 常见问题,无线有 14 个,还有漏洞挖掘、云、移动端、逆向这些方向。 GitHub: 每个 Skill 就是一个 SKILL.md,把某个领域的方法、工具链、边界情况写清楚,让 Claude 在这个方向上答得像个熟手,而不是泛泛而谈。 作者写明适用场景是授权的红队项目、漏洞赏金、安全研究和 CTF 备赛,动手前先确认自己在授权范围内。

DeepSeek Harness 官方桌面端已经基本做完。官方仓库主分支新增完整的 apps/desktop,用 Electron 把现有 Harness Web UI 封装成桌面 App。应用自带 Node.js 和 pnpm,用户不需要再手动启动 Web 服务,桌面端也不会对外监听端口。 官方已经准备好 macOS Apple Silicon、Intel 和 Windows x64 三套安装包构建流程,还包括 macOS 签名与公证、Windows EV 签名、自动更新和安装失败恢复。生产更新服务器已经指向 DeepSeek 自己的 暂时不在官方发布目标里。 最近一轮桌面端提交主要在处理 macOS 公证提速、启动恢复、打包校验等发布前问题。 DeepSeek 目前还没有公布具体发布日期,GitHub Release 里也还没有桌面安装包。但从仓库状态来看,桌面端已经进入发布前收尾阶段。

+ 3 more items
truncated at source

Video editors were built for people clicking through timelines. Hypit is built for coding agents. Give Codex or Claude Code a reference video, and it can use Hypit to clone the production through natural-language instructions instead of navigating a traditional editing interface. It’s open source, free to use, and BYOK. Check it out and star the repo:

@cccyd_qwq

Introducing Hypit: Clone any viral video with AI agents. > 1 clone, 100 variants, 100M views. Hypit lets your AI agent (Claude Code, Codex...) clone any viral video. > Paste any viral video link from TikTok, Instagram, or YouTube into your agent. Hypit clones it into a complete agentic video workflow: footage, captions, B-roll, effects. > GitHub: > Key Point: - Arcads: $220 / month - Higgsfield: $129 / month - Creatify: $99 / month - Hypit: FREE 🌟 > Build the video creation harness for AI Agents. Redefine how vide

DeepSeek V4.1 Flash on max is ranked as the #3 open model on @arena. I'm using this model as my daily driver right now and am incredibly impressed. It's fast, it's smart and I have an enormous amount of context that remains coherent and accurate. The whale definitely cooked.

@YourLocalAILab

With today's release of Jovian Judgement R37, DeepSeek V4.1 Flash has full support with blazing speeds. > Decodes of 500 tok/s and prefills of 20k tok/s for single stream workloads on 4 RTX 6000 Pros. ⚡️ > To learn more point your agent at our GitHub here:

opennews-mcp 是一个把各大金融消息源汇到一起的 MCP,装进 Claude Code 后用聊天就能问最近有什么值得注意的动向。 一共接了 85 个以上的数据源,分成新闻、交易所上币、链上大额、市场异动、预测信号几类,光新闻类就占了 55 家。 GitHub: 每条消息都过了一遍 AI,打一个 0 到 100 的影响力分,标出偏多还是偏空,再配一段中英文摘要。 链上那块还能盯着大额的交易和持仓变动,市场那块可以查看啊资金费率、大额爆仓这些异常。 用之前需要先配置 API Key。除了盯盘之外,还想让消息面也自动汇总的朋友,可以接进去试试。

Anthropic

8 items

official site

Anthropic launched Claude for financial advisors with integrations across $SCHW, $BLK, Addepar and Orion. The platform is designed to plug directly into advisor workflows, portfolio data and client management systems.

Anthropic CEO: "AI development needs to slow down" OpenAI CEO: "Completely agree" Two months later: Anthropic: Introducing Fable 5.2 OpenAI: Introducing GPT 6.1

OpenAI seems to be ramping things up around the new models, especially GPT-6 Sol, which I’m expecting to land within the next few days. Anthropic looks to be doing the same with Opus 5.2, and the backend already seems pretty much lined up for launch. There’s a real chance we get both of them very soon, possibly within a pretty tight window of each other. (I just haven’t seen the same level of progress around gpt-6-luna and gpt-6-terra yet. But I’d assume they’re part of the rollout too.)

🚨 sol 5.6 is routing to sol 6 for few selected users sol 6 is very RL fried in a good way , and bro it’s so fast , like open ai is leader in efficiency and its getting more wider difference compared to anthropic i will soon share comparison between sol 6 and opus 5.2/1

飞书发布 8.0,并把「豆包工作伙伴」升级成面向整个团队的 Agent。 它和 Anthropic 此前推出的 Claude Tag 很像:Agent 不再只和一个人私聊,而是拥有自己的身份、权限和记忆,可以直接进入团队协作环境,和一群人一起工作。 在授权范围内,豆包工作伙伴可以读取飞书里的文档、表格、会议、日历和企业业务系统,也能直接操作相关工具。除了信息同步、文档审校、周报整理,它还会根据工作进展主动执行任务。

🚨 Opus 5.2 First Output Leaked The jump in visual quality is already noticeable better details, stronger composition Anthropic is back

+ 2 more items

Claude 正在测试一个新的 Money 入口。TestingCatalog 在 iOS 未发布界面中发现,用户可以直接连接银行账户,再让 Claude 回答消费、财务计划等问题。目前 Anthropic 还没有正式发布这一功能。 这几乎就是 Anthropic 版的 ChatGPT Finances。OpenAI 已允许美国 Plus 和 Pro 用户通过 Plaid 连接银行、信用卡和券商账户,用 ChatGPT 查看支出、账单、订阅、净资产和投资。

@testingcatalog

ANTHROPIC 🔥: A new "Claude Money" feature is being prepared for release on the Claude app for iOS. > Claude Money (sounds like Claude Monet!) is Anthropic's personal finance solution that lets users connect their bank accounts to Claude. > "Understand your money with Claude" - Link your bank accounts and ask Claude about spending, plans, and more. > This solution will likely launch in the US only, similar to how it works for ChatGPT Finance.

新機能Claude Modsが登場 (以前Function Hooksとして告知されていたもの) 従来の hooks より一段深く、Claude の実行パスに関数ベースのミドルウェア として割りこめる。テトリスを作ることも可能 CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 上記の環境変数で有効化できる

@bcherny

Claude Mods are landing now. Someone already built a Tetris-in-Claude mod 🤯 > See issue for the latest community update, technical details, and more cool demos >

MCP

8 items

official site

Voice, music, image, and video generation now run inside the assistant you already work in, and everything you make lands in your ElevenCreative workspace. Open it in Studio to refine narration, layer in music and sound effects, and export.

@ElevenLabs

Introducing voice, music, image, and video generation in the ElevenLabs MCP. > Generate speech, transcripts, dubs, music, sound effects, images, and video from the assistant you already work in.

DeepSeek Harness 最新的 v0.1.6-alpha.1 同时新增实验性 Browser Use 和 Computer Use。模型现在既能直接操作网页,也能查看并控制本地桌面。 Browser Use 支持 Playwright MCP、Chrome DevTools MCP 和 Stagehand,可以直接操作网页。 Computer Use 则接入 Cua Driver,可找到具体 App 和窗口、获取当前界面截图,再根据界面元素或坐标执行点击和输入。官方提供 MCP 和原生两种 Computer Use 接入方式,原生方式可以直接集成进 Harness,不需要另外安装独立的 Cua Driver App。 两项能力目前都属于实验性功能,需要手动启用。Computer Use 还需要给实际运行它的应用授予屏幕读取和电脑操作权限。

Wow... this is very cool.

@tonysimons_

I just found a repo that lets AI agents control a VIRTUAL IPHONE. > Not the iOS Simulator. > Actual virtualized iOS on Apple Silicon. > Screenshots. Touch. Swipes. Typing. Apps. > There’s even an MCP server for coding agents. > +633 stars TODAY. >

Just shipped our video editor MCP for @floraai today - Amazed what Astra can do with it

Drop one picture into an AI agent and come back to a full 3D street. Shops, signs, road, sky, all built without anyone opening a 3D tool. GPT-6 Astra and Hyper3D MCP just did in one run what used to take a whole team.

@DeemosTech

One image. One Agent. One scene. > @OpenAIDevs GPT-6 Astra (Codex) + HYPER3D MCP just showed us the second half of 3D gen. 🚀

+ 2 more items

太魔幻了,Astra 还在让大家惊叹:AI 终于会用 Blender 了。 结果Nex-AGI 已经直接出现在现实世界里了。 我只给了一个 Prompt: 做一根 6cm 的粉色活动香蕉,省点料。 然后我基本没碰 Blender。 它自己用 MCP 建模,Computer Use 看结果、自己修,最后直接吐给我一个 STL。 我顺手扔进拓竹。 👇 视频里正在一层一层长出来的,就是它自己设计的香蕉。 这一刻我是真有点惊到了。 Nex-AGI + Computer Use,能力完全超出我预期。 一句话 → Blender → STL → 3D 打印 → 实物。 🔥 这真的有点科幻了。 直接免费体验nex-agi :

opennews-mcp 是一个把各大金融消息源汇到一起的 MCP,装进 Claude Code 后用聊天就能问最近有什么值得注意的动向。 一共接了 85 个以上的数据源,分成新闻、交易所上币、链上大额、市场异动、预测信号几类,光新闻类就占了 55 家。 GitHub: 每条消息都过了一遍 AI,打一个 0 到 100 的影响力分,标出偏多还是偏空,再配一段中英文摘要。 链上那块还能盯着大额的交易和持仓变动,市场那块可以查看啊资金费率、大额爆仓这些异常。 用之前需要先配置 API Key。除了盯盘之外,还想让消息面也自动汇总的朋友,可以接进去试试。

Codex

OpenAI / 7 items

Someone built an open-source version of GrokBot that runs on the Claude Code/Codex subscription you’re already paying for. Same basic idea: Give your agents access to a computer. Connect their apps through Composio. Run them through the AI subscription you already have. But OpenMausBot itself is free and open source. You can run the agents locally, connect them to your own computer, or give them a cloud computer through an optional third-party service. Very worth knowing about if you already pay for Claude Code or Codex.

386页的《3D Vibe Coding》橙皮书📙现已开源! 链接👇

@AlchainHust

从GPT-6 Astra发布以来,就一堆人在让Codex去通过computer use去操控blender做各种3D模型的。 > 但稍微做了点就停了。 > 然后我自己也拿Claude Code和Codex做了不少测试,发现不管是直接写代码的方式,还是操控blender的方式,从成本和效率来说都不是最好的方式。 > 所以我自己摸索了一套新的策略。用Codex 加 3D 基座模型 Tripo,做了一座能开车逛的巴黎。 > 整了两天,一行代码没写。一座能开进去的巴黎👇出来了! > 现在视频演示里的车、树、铁塔、街上走路的人——全是 AI 生成的 3D 资产;而街区、桥、塞纳河——是代码画的。 > 为什么这么分? 因为你开车会开到车跟前,但不会开到桥底下去数栏杆。 > 近景给精细模型,远景交给代码。做出来的网页游戏会在细节和帧率上更平衡。 > 以及,除了这期视频,我还把一整套操作流程做成了386 页手册教程! 现已开源了,提示词、脚本、模型全在里面:

Voice is one of my favourite ways to brainstorm on blogposts, PRs, RFCs and even managing finances + calendars You can now use nearly two and half times more ChatGPT Voice!! Enjoy!

@athyuttamre

⚡️ 2.4x more ChatGPT Voice in Desktop > We've dropped prices by ~60% for voice in Codex and Work in the desktop app, giving you more time to orchestrate tasks and even more tokens for real work.

Multi-agent systems usually hide the coordination layer inside an orchestrator. @Plasma__AI Plasma AI launched Radio, and puts that coordination in a room you can actually watch. Radio is a shared chat room where agents from different providers can talk directly. Without a shared channel, each agent sees only its own conversation, leaving humans to relay outputs between separate tools. Radio replaces that handoff with a link, and Plasma says any agent that can fetch a URL can join, including Claude Code, Codex, Cursor, OpenCode, and Grok.

@Plasma__AI

Introducing Radio: A chat room for your agents. > Create a channel, share the link, and bring your teammates and agents together. No sign up required. > Try it today at

truncated at source

Video editors were built for people clicking through timelines. Hypit is built for coding agents. Give Codex or Claude Code a reference video, and it can use Hypit to clone the production through natural-language instructions instead of navigating a traditional editing interface. It’s open source, free to use, and BYOK. Check it out and star the repo:

@cccyd_qwq

Introducing Hypit: Clone any viral video with AI agents. > 1 clone, 100 variants, 100M views. Hypit lets your AI agent (Claude Code, Codex...) clone any viral video. > Paste any viral video link from TikTok, Instagram, or YouTube into your agent. Hypit clones it into a complete agentic video workflow: footage, captions, B-roll, effects. > GitHub: > Key Point: - Arcads: $220 / month - Higgsfield: $129 / month - Creatify: $99 / month - Hypit: FREE 🌟 > Build the video creation harness for AI Agents. Redefine how vide

An agent may need five data sources for a single task. Paying five monthly subscriptions for that gets expensive. Glasser says it offers 40+ providers through one API key, billed per call. A useful setup for agents whose data needs change with every task.

@iammutex

Your agent is only as good as what you feed it. Garbage in. Garbage out. Introducing @Glasserai - it feeds your agent premium data instead. Ahrefs. Semrush. ZoomInfo. Apollo. PDL. etc. 1,900+ endpoints. One key. Pay per call. No more subscriptions. Send this one line to your Claude Code, Codex, Grok Bot, or Muse: set up

+ 1 more items

Drop one picture into an AI agent and come back to a full 3D street. Shops, signs, road, sky, all built without anyone opening a 3D tool. GPT-6 Astra and Hyper3D MCP just did in one run what used to take a whole team.

@DeemosTech

One image. One Agent. One scene. > @OpenAIDevs GPT-6 Astra (Codex) + HYPER3D MCP just showed us the second half of 3D gen. 🚀

Apple

6 items

Busiest day official site

Only ~1.2B parameters active at a time. Edge0-8B-A1B-preview makes an 8B-class MoE practical for local inference.📜 Apache 2.0. 🤖 ⚡ Reaches 23.9–25.3 tokens/s with about 1.0 GiB peak active memory in the reported short-context benchmark. 🏆 Retains most of the FP16 base model’s quality, with an average gap of just 2.8 points across five benchmarks. MMLU-Pro rises from 65.8 to 70.1. 🧠 SSD expert offload keeps most weights outside active memory, while a prerouter predicts which experts each layer will need next. 🛠 Recover-LoRA offsets quality loss from 4-bit quantization and expert offloading. The current Preview runs locally through MLX on Apple Silicon.

A Metal capability layer for macOS VMs just unlocked 11-16x faster LLM inference on Apple Silicon - no hardware changes, just a process-scoped shim hitting the paravirtual GPU. Gemma 4 12B goes from 3.41 to 49.67 tokens per second on generation. Local AI in VMs just became a serious option.

@trycua

1/ Today, as part of our broader research into Apple Silicon virtualization, we're releasing a process-scoped Metal capability layer for macOS VMs. On one M1 Ultra, prompt / generation: > TinyLlama: 11.08× / 16.36× Gemma 4 12B: 7.20× / 14.54× Muse Glimmer 30B: 7.55× / 8.87×

Wow... this is very cool.

@tonysimons_

I just found a repo that lets AI agents control a VIRTUAL IPHONE. > Not the iOS Simulator. > Actual virtualized iOS on Apple Silicon. > Screenshots. Touch. Swipes. Typing. Apps. > There’s even an MCP server for coding agents. > +633 stars TODAY. >

DeepSeek Harness 官方桌面端已经基本做完。官方仓库主分支新增完整的 apps/desktop,用 Electron 把现有 Harness Web UI 封装成桌面 App。应用自带 Node.js 和 pnpm,用户不需要再手动启动 Web 服务,桌面端也不会对外监听端口。 官方已经准备好 macOS Apple Silicon、Intel 和 Windows x64 三套安装包构建流程,还包括 macOS 签名与公证、Windows EV 签名、自动更新和安装失败恢复。生产更新服务器已经指向 DeepSeek 自己的 暂时不在官方发布目标里。 最近一轮桌面端提交主要在处理 macOS 公证提速、启动恢复、打包校验等发布前问题。 DeepSeek 目前还没有公布具体发布日期,GitHub Release 里也还没有桌面安装包。但从仓库状态来看,桌面端已经进入发布前收尾阶段。

Standup Pulse: an open-source project running async Slack standups using Gemma 4 26B-A4B (GGUF via llama.cpp) on an Apple M5 Max. It pairs Mastra for typed tool selection with CopilotKit Channels for restrained Slack Block Kit (Slack's native layout format) interactions while keeping all model inference, standup records, and traces local in SQLite. This architecture is a great demonstration of how to connect Slack to a project without exposing the local server! 🔗 Blog: 🔗 Repo:

yeah.. maybe make an hermes as a hidden model-provider support @Teknium @NousResearch . We need to get away from cloud sources and able to use open local models vis Hermes agent CLI integration to apple Siri.

@marcelpociot

I got Claude answering inside Siri on macOS 27 🚀 > Apple's hidden model-provider support + my existing Claude Code account. > Open-source proof of concept. Requires disabling SIP/AMFI. > Thanks @itspdfu for the discovery! > Check it out:

Claude

Anthropic / 7 items

official site

Anthropic launched Claude for financial advisors with integrations across $SCHW, $BLK, Addepar and Orion. The platform is designed to plug directly into advisor workflows, portfolio data and client management systems.

做渗透测试,光 Web 一个方向就有 SQL 注入、XSS、SSRF 十几种攻击面。 每种的思路、工具、容易踩的坑都得记在脑子里,换个方向又是一整套。 Claude-Red 把这些方法论整理成 78 个安全测试 Skill,装进 Claude 后在对话中按话题自动加载,说到 SQL 注入就调出对应那一个。 分成 23 类,Web 应用有 16 个覆盖 OWASP 常见问题,无线有 14 个,还有漏洞挖掘、云、移动端、逆向这些方向。 GitHub: 每个 Skill 就是一个 SKILL.md,把某个领域的方法、工具链、边界情况写清楚,让 Claude 在这个方向上答得像个熟手,而不是泛泛而谈。 作者写明适用场景是授权的红队项目、漏洞赏金、安全研究和 CTF 备赛,动手前先确认自己在授权范围内。

飞书发布 8.0,并把「豆包工作伙伴」升级成面向整个团队的 Agent。 它和 Anthropic 此前推出的 Claude Tag 很像:Agent 不再只和一个人私聊,而是拥有自己的身份、权限和记忆,可以直接进入团队协作环境,和一群人一起工作。 在授权范围内,豆包工作伙伴可以读取飞书里的文档、表格、会议、日历和企业业务系统,也能直接操作相关工具。除了信息同步、文档审校、周报整理,它还会根据工作进展主动执行任务。

Claude 正在测试一个新的 Money 入口。TestingCatalog 在 iOS 未发布界面中发现,用户可以直接连接银行账户,再让 Claude 回答消费、财务计划等问题。目前 Anthropic 还没有正式发布这一功能。 这几乎就是 Anthropic 版的 ChatGPT Finances。OpenAI 已允许美国 Plus 和 Pro 用户通过 Plaid 连接银行、信用卡和券商账户,用 ChatGPT 查看支出、账单、订阅、净资产和投资。

@testingcatalog

ANTHROPIC 🔥: A new "Claude Money" feature is being prepared for release on the Claude app for iOS. > Claude Money (sounds like Claude Monet!) is Anthropic's personal finance solution that lets users connect their bank accounts to Claude. > "Understand your money with Claude" - Link your bank accounts and ask Claude about spending, plans, and more. > This solution will likely launch in the US only, similar to how it works for ChatGPT Finance.

yeah.. maybe make an hermes as a hidden model-provider support @Teknium @NousResearch . We need to get away from cloud sources and able to use open local models vis Hermes agent CLI integration to apple Siri.

@marcelpociot

I got Claude answering inside Siri on macOS 27 🚀 > Apple's hidden model-provider support + my existing Claude Code account. > Open-source proof of concept. Requires disabling SIP/AMFI. > Thanks @itspdfu for the discovery! > Check it out:

新機能Claude Modsが登場 (以前Function Hooksとして告知されていたもの) 従来の hooks より一段深く、Claude の実行パスに関数ベースのミドルウェア として割りこめる。テトリスを作ることも可能 CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 上記の環境変数で有効化できる

@bcherny

Claude Mods are landing now. Someone already built a Tetris-in-Claude mod 🤯 > See issue for the latest community update, technical details, and more cool demos >

+ 1 more items

🚨 Claude Opus 5.2 is coming soon then The "claude-opus-5-2" slug is in Microsoft Foundry: Tibo is also teasing Sol for this week too I believe, so this is going to be fun

Kimi

Moonshot AI / 6 items

Busiest day

i cannot in good conscience recommend @nahcrof anymore, i used them a lot in the past, they were an absolute steal, exceptionally reliable (for a small inference provider), atleast to my knowledge up until a few months ago they were entirely fine post Kimi K3 launch things seem to have shifted and this feels like a rugpull, i've independently confirmed the tokenizers don't match the model ids, there is an extremely detailed blogpost by @KTibow that goes further into this linked below, where all the evidence seems to check out

Kimi Code New Features 💡 Remote Control is now always on; the experimental KIMI_CODE_EXPERIMENTAL_REMOTE_CONTROL flag has been removed. 💡 Add native clipboard support on Linux X11, so copying from the TUI no longer depends on the terminal's OSC 52 support. 💡 Kimi web: Add a Plugins panel to Settings for browsing the plugin marketplace and installing, enabling, disabling, and removing plugins. See the Changelog for more technical entries.

Anyone who’s built with AI knows the annoying part: You’re finally in flow, then you start rationing prompts because every small change eats into your usage. @boltdotnew’s new Forge mode is built for exactly this. It’s an experimental mode in AI app builder with up to 50x more usage for building full-stack web apps. I took one from a prompt to a live app, then kept iterating without staring at the meter. You can also opt into Forge to help train open models. The research preview runs from Sep 14 to Oct 14, 2026. It’s free on Pro plans until Oct 14, or $9/month through the Bolt Lite early-access plan.

@boltdotnew

Introducing Bolt Forge. Free until Oct 14th: > - Up to 50x more usage - The new frontier: GLM, DeepSeek, Kimi - Zero usage charges > Live now in your model picker on > And one more thing... 👇

@Kimi_Moonshot on the verge of releasing K2.8🤩 K2.8-Preview available now in the Kimi Code CLI!

UCSD 助理教授黄碧薇创办的因果世界模型公司 Aether AI 开源 RSIAgent。 这是一套不训练模型的递归自我改进框架。底层模型参数全程固定,Agent 会自己寻找值得练习的任务、实际操作、检查结果,再把成功方法和失败教训写进长期 Memory。下一轮继续利用这些经验,逐步补上能力短板。 系统由三个 Agent 配合。Curriculum Agent 决定接下来练什么,Actor Agent 真正操作软件,Verifier Agent 独立检查结果。探索分成两步:先广泛尝试不同任务,再针对失败、隐藏限制和边界情况继续深挖。最后 Memory 会被冻结,直接拿去执行正式任务。 在 OSWorld 2.0 上,加入这套 RSI 后,平均部分得分从 71.97% 提升到 78.98%;Agents’ Last Exam 从 83.75% 提升到 84.82%。不过这不是整套测试的严格 A/B 对比。OSWorld 只有一半任务实际用了 RSI 后的新结果,其余任务继续沿用原成绩。 它和 Prime Agent 这类 Harness 自我改进思路属于同一个大方向:模型权重不变,持续更新模型外面的东西。 RSIAgent 的特点是把改进重点放在 Memory,再用自主出题和独立验证,让这份外部经验库不断积累。

@huang_biwei

Can an agent explore a new environment, learn its causal structure, and keep improving without updating its model weights? > We introduce RSIAgent, a framework for recursive self-improvement through autonomous exploration. Using Kimi-K3 and GLM-5.3 as base models, RSIAgent outperforms GPT-6 Astra on both OSWorld 2.0 and Agents’ Last Exam. > RSIAgent decides what to explore, executes tasks,

next few days could get VERY interesting for AI 👀 Grok 4.7 Opus 5.1 Gemini 4 Kimi K3.1 GPT-6 Sol some of these are rumors, some have stronger signals than others but if even a few actually drop, the AI leaderboard is about to get chaotic who are you betting on?

Slack

5 items

Busiest day official site

Slack is the front door to the Agentic Enterprise. Slackbot now works in channels and Salesforce Lightning, while Slack Code brings people and AI agents together to write, review, and ship code live. One AI interface for work, connected to everything.

slack is where agents are doing work managed deepagents has a very tight slack integration - check out the tutorial on it below

@LangChain

PSA: You can connect a Managed Deep Agent to Slack. > Mentions, DMs, and thread replies all start a run. Your agent posts its response back in the same conversation. > Here’s how to get started ⤵️

Standup Pulse: an open-source project running async Slack standups using Gemma 4 26B-A4B (GGUF via llama.cpp) on an Apple M5 Max. It pairs Mastra for typed tool selection with CopilotKit Channels for restrained Slack Block Kit (Slack's native layout format) interactions while keeping all model inference, standup records, and traces local in SQLite. This architecture is a great demonstration of how to connect Slack to a project without exposing the local server! 🔗 Blog: 🔗 Repo:

New capabilities for Linear Loops: • Trigger from project, initiative, cycle, and issue changes • Edit documents automatically • Send tailored Slack updates 🔗:

CRM data, AI agents, and model choice — into the tools teams use every day. We're expanding collaboration with @AWS to help enterprises close deals, resolve customer issues, and move faster with AI. → Access Salesforce business context in @AmazonQuick → Bring AWS frontier agents into @SlackHQ → Expand Agentforce model choice through Amazon Bedrock → Extend Data 360 zero copy across more AWS data services → Enable Agentforce Voice with Amazon Connect Customer Putting AI where customers work, grounded in data they trust, without requiring migrations. Get the details:

DeepSeek

5 items

official site

DeepSeek Harness 最新的 v0.1.6-alpha.1 同时新增实验性 Browser Use 和 Computer Use。模型现在既能直接操作网页,也能查看并控制本地桌面。 Browser Use 支持 Playwright MCP、Chrome DevTools MCP 和 Stagehand,可以直接操作网页。 Computer Use 则接入 Cua Driver,可找到具体 App 和窗口、获取当前界面截图,再根据界面元素或坐标执行点击和输入。官方提供 MCP 和原生两种 Computer Use 接入方式,原生方式可以直接集成进 Harness,不需要另外安装独立的 Cua Driver App。 两项能力目前都属于实验性功能,需要手动启用。Computer Use 还需要给实际运行它的应用授予屏幕读取和电脑操作权限。

DeepSeek Harness 官方桌面端已经基本做完。官方仓库主分支新增完整的 apps/desktop,用 Electron 把现有 Harness Web UI 封装成桌面 App。应用自带 Node.js 和 pnpm,用户不需要再手动启动 Web 服务,桌面端也不会对外监听端口。 官方已经准备好 macOS Apple Silicon、Intel 和 Windows x64 三套安装包构建流程,还包括 macOS 签名与公证、Windows EV 签名、自动更新和安装失败恢复。生产更新服务器已经指向 DeepSeek 自己的 暂时不在官方发布目标里。 最近一轮桌面端提交主要在处理 macOS 公证提速、启动恢复、打包校验等发布前问题。 DeepSeek 目前还没有公布具体发布日期,GitHub Release 里也还没有桌面安装包。但从仓库状态来看,桌面端已经进入发布前收尾阶段。

Anyone who’s built with AI knows the annoying part: You’re finally in flow, then you start rationing prompts because every small change eats into your usage. @boltdotnew’s new Forge mode is built for exactly this. It’s an experimental mode in AI app builder with up to 50x more usage for building full-stack web apps. I took one from a prompt to a live app, then kept iterating without staring at the meter. You can also opt into Forge to help train open models. The research preview runs from Sep 14 to Oct 14, 2026. It’s free on Pro plans until Oct 14, or $9/month through the Bolt Lite early-access plan.

@boltdotnew

Introducing Bolt Forge. Free until Oct 14th: > - Up to 50x more usage - The new frontier: GLM, DeepSeek, Kimi - Zero usage charges > Live now in your model picker on > And one more thing... 👇

truncated at source

腾讯混元等团队发布 EvolveScaler,专门测试大模型能不能跟上不断变化的信息。它会在长文本里不断加入修改、撤回、补录和作废的信息,再让模型根据最新状态回答问题。 比如先给模型看 40 天游戏记录,再问:「如果第 7 天没打那个小 Boss,最后还能不能打赢?」这会连带改变后面的装备、血量和战斗结果。模型得从第 7 天开始,把后面几十天重新推一遍,原文里没有现成答案。 团队设计了 117 类任务、159 种问题,最长约 1200 个事件。14 个模型配置包括 GPT-5.5、Gemini 3.1 Pro、DeepSeek V4 Preview Pro、GLM-5.2 和 Qwen3.5 Plus。最难档里,成绩中位数只有 11.3%;表现最好的 GPT-5.5-xhigh 也只有 59.3%。 这套数据也能拿来训练模型。一个内部 A3B 模型加入 6000 条这类数据后,在 8 个外部测试上全部提升,平均提高 5.25 分。

@TencentHunyuan

🚀 EvolveScaler is here. > Read a 40-day RPG log. Now answer one question: if you skip the mini-boss on Day 7, do you still beat the final boss? > The answer isn't in the log. You have to replay the world. > That's Information Evolution — records get retracted, corrected, backfilled. The world keeps changing after you read it. > So we build it backwards: define the world as an executable state machine, then render it into natural language. Code guarantees the logic. Language delivers the mess. > ➡️ 117 prototypes. 159 q

DeepSeek V4.1 Flash on max is ranked as the #3 open model on @arena. I'm using this model as my daily driver right now and am incredibly impressed. It's fast, it's smart and I have an enormous amount of context that remains coherent and accurate. The whale definitely cooked.

@YourLocalAILab

With today's release of Jovian Judgement R37, DeepSeek V4.1 Flash has full support with blazing speeds. > Decodes of 500 tok/s and prefills of 20k tok/s for single stream workloads on 4 RTX 6000 Pros. ⚡️ > To learn more point your agent at our GitHub here:

Gemini

Google / 5 items

official site

two new google live models just appeared gemini-3.8-live gemini-3.8-live-extended-thinking previous live ones gemini-3.1-flash-live-preview Google please bring this to Gemini windows desktop app along with the model

Gemini 4 Pro (High) VS Gemini 3.8 Flash (High) - Voxel Pagoda The first output from the upcoming Gemini 4 Pro model has leaked. The output isn't too impressive We'll need to see how the model improves with newer checkpoints

@Lentils80

Gemini 4 Pro checkpoints have finally started appearing internally a few days ago. > This is the first ever output from the model, internally codenamed "argon". It took 2.4 minutes on High thinking effort. > It has a 256k token "output limit", compared to 64k in previous Gemini models. I also heard it will ship with a 2M context window, tho it's still not decided if they'll do it.

GOOGLE 🔥: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking model names have started appearing on the GCP Console quotas page. A new Gemini Live model, based on the latest Gemini 3.8 Flash, is expected to be a big leap. > So far, Gemini 3.1 Live Preview is the latest Gemini Live model available via APIs. > Recently, Google updated Gemini Live with support for Connector calls and the possibility of triggering Deep Research. > OpenAI also released the GPT Live 1 model on the APIs last week, and it seems like Google has a response to that. * Discovered by Bedros Pamboukian

@bedros_p

Gemini 3.8 Live + Gemini 3.8 Live Extended Thinking have appeared on the Google Cloud quota & metrics page 2 hours ago

truncated at source

腾讯混元等团队发布 EvolveScaler,专门测试大模型能不能跟上不断变化的信息。它会在长文本里不断加入修改、撤回、补录和作废的信息,再让模型根据最新状态回答问题。 比如先给模型看 40 天游戏记录,再问:「如果第 7 天没打那个小 Boss,最后还能不能打赢?」这会连带改变后面的装备、血量和战斗结果。模型得从第 7 天开始,把后面几十天重新推一遍,原文里没有现成答案。 团队设计了 117 类任务、159 种问题,最长约 1200 个事件。14 个模型配置包括 GPT-5.5、Gemini 3.1 Pro、DeepSeek V4 Preview Pro、GLM-5.2 和 Qwen3.5 Plus。最难档里,成绩中位数只有 11.3%;表现最好的 GPT-5.5-xhigh 也只有 59.3%。 这套数据也能拿来训练模型。一个内部 A3B 模型加入 6000 条这类数据后,在 8 个外部测试上全部提升,平均提高 5.25 分。

@TencentHunyuan

🚀 EvolveScaler is here. > Read a 40-day RPG log. Now answer one question: if you skip the mini-boss on Day 7, do you still beat the final boss? > The answer isn't in the log. You have to replay the world. > That's Information Evolution — records get retracted, corrected, backfilled. The world keeps changing after you read it. > So we build it backwards: define the world as an executable state machine, then render it into natural language. Code guarantees the logic. Language delivers the mess. > ➡️ 117 prototypes. 159 q

next few days could get VERY interesting for AI 👀 Grok 4.7 Opus 5.1 Gemini 4 Kimi K3.1 GPT-6 Sol some of these are rumors, some have stronger signals than others but if even a few actually drop, the AI leaderboard is about to get chaotic who are you betting on?

Qwen

Alibaba / 4 items

One physical intelligence loop, now in 2B and 8B. PhysBrain 1.5 is here! 🤖 🤖 📄 🏆 PhysBrain 1.5-8B scores 72.5 across 28 embodied-understanding benchmarks, ranking #1 among the evaluated open-source models and leading 14 tasks. The compact 2B reaches 66.6. 🧭 Both models cover spatial perception, 3D reasoning, embodied planning, grounding, affordance, and trajectory reasoning. 🦾 They generate end-effector trajectories and predict future states as aligned RGB, depth, and robot masks. 🧠 Built on Qwen3-VL, PhysBrain 1.5 models language, spatial outputs, actions, and future states as tokens within one autoregressive backbone. No task-specific heads.

阶跃星辰发布 StepAudio 3,一口气上线 Realtime、ASR、TTS、Gen 和 Music 五款语音模型。 在 Artificial Analysis 的两项语音评测中,StepAudio 3 Realtime 都排第一。Conversational Dynamics 得分 98.9%,高于 Qwen Audio 3.0 Realtime Plus 的 98.4% 和 GPT-Realtime-2 High 的 95.3%;Speech Reasoning 得分 99.7%,同样位列第一。前者主要测模型能不能处理停顿、轮流说话、用户打断和「嗯」「对」这类附和,后者测模型能否直接听懂音频并完成推理。 ASR 也在 Artificial Analysis 非流式语音识别榜做到 1.7% WER(词错误率),与 Fun-Realtime-ASR-preview 并列第一。另外,StepAudio 3 Gen 可以一次生成人声、音效、环境声和音乐,并按场景统一编排。 五款模型目前均已进入阶跃的语音产品线。

truncated at source

Qwen3.8-Flash-Next EXL3 just got a BIG one-Spark update. 🔥 58.8 tok/s single-stream through native ExLlamaV3.🚀 157.6 tok/s aggregate through vLLM across 8 streams.🤯 FULL 262,144 context on ONE DGX Spark. 4.05 bpw EXL3 pack. A much better serving envelope. 𝗕𝗘𝗙𝗢𝗥𝗘 → 𝗡𝗢𝗪 Previous public headline: 47.6 tok/s greedy p50 64K configured context MTP k=2 Concurrency not characterized Now: Native ExLlamaV3: 58.8 tok/s single stream vLLM + vllm-exl3: ~50–55 tok/s single stream 155.6 tok/s @ 4 streams 157.6 tok/s @ 8 streams Configured context: 64K → 262,144 KV pool: 416,163 tokens at the default 262K config 𝗠𝗧𝗣 𝗞=𝟯 𝗜𝗦 𝗡𝗢𝗪 𝗧𝗛𝗘 𝗪𝗜𝗡𝗡𝗘𝗥 Current controlled sweep @ 4K: No draft: 27.77 tok/s MTP k=2: 47.39 MTP k=3: 50.09 At 32K: No draft: 27.58 k=2: 46.60 k=3: 49.73 So the fixed/current build flips the old result: k=3 is now the sweet spot. There is one important boundary I found: At 163,840 PROMPT tokens, MTP acceptance collapses to zero. Above that point speculation becomes slower than running without a draft

truncated at source

腾讯混元等团队发布 EvolveScaler,专门测试大模型能不能跟上不断变化的信息。它会在长文本里不断加入修改、撤回、补录和作废的信息,再让模型根据最新状态回答问题。 比如先给模型看 40 天游戏记录,再问:「如果第 7 天没打那个小 Boss,最后还能不能打赢?」这会连带改变后面的装备、血量和战斗结果。模型得从第 7 天开始,把后面几十天重新推一遍,原文里没有现成答案。 团队设计了 117 类任务、159 种问题,最长约 1200 个事件。14 个模型配置包括 GPT-5.5、Gemini 3.1 Pro、DeepSeek V4 Preview Pro、GLM-5.2 和 Qwen3.5 Plus。最难档里,成绩中位数只有 11.3%;表现最好的 GPT-5.5-xhigh 也只有 59.3%。 这套数据也能拿来训练模型。一个内部 A3B 模型加入 6000 条这类数据后,在 8 个外部测试上全部提升,平均提高 5.25 分。

@TencentHunyuan

🚀 EvolveScaler is here. > Read a 40-day RPG log. Now answer one question: if you skip the mini-boss on Day 7, do you still beat the final boss? > The answer isn't in the log. You have to replay the world. > That's Information Evolution — records get retracted, corrected, backfilled. The world keeps changing after you read it. > So we build it backwards: define the world as an executable state machine, then render it into natural language. Code guarantees the logic. Language delivers the mess. > ➡️ 117 prototypes. 159 q

Salesforce

4 items

Busiest day official site

$CRM Salesforce and $NVDA NVIDIA launched Koa, Salesforce’s first CRM reasoning model for Agentforce, built on NVIDIA Nemotron and trained on 27 years of CRM knowledge. Koa is designed for complex multi-step enterprise workflows and delivered 3x fewer errors on Salesforce’s CRM benchmark. It is already in select customer pilots, with general availability expected in Winter 2026.

Introducing AIforce The full power of Salesforce. Any AI interface. AIforce is the live interface layer for the Agentic Enterprise — bringing your data, workflows, business logic, permissions, security, and governance wherever people and agents work Launching with Claudeforce, Slackforce, and Agentforce Coworker Salesforce comes to you. See what that unlocks:

Slack is the front door to the Agentic Enterprise. Slackbot now works in channels and Salesforce Lightning, while Slack Code brings people and AI agents together to write, review, and ship code live. One AI interface for work, connected to everything.

CRM data, AI agents, and model choice — into the tools teams use every day. We're expanding collaboration with @AWS to help enterprises close deals, resolve customer issues, and move faster with AI. → Access Salesforce business context in @AmazonQuick → Bring AWS frontier agents into @SlackHQ → Expand Agentforce model choice through Amazon Bedrock → Extend Data 360 zero copy across more AWS data services → Enable Agentforce Voice with Amazon Connect Customer Putting AI where customers work, grounded in data they trust, without requiring migrations. Get the details:

Seedance

ByteDance / 4 items

Made a 30s cinematic video with @renoiseaijp new “Agent × Canvas” feature. Just type what you want to create in natural language (Japanese works great too), and the Agent brings it to life. My workflow for this one was super simple: Character sheet → Describe the scene → Generate 7 days of unlimited Seedance 2.5 (720p) available now. It’s an awesome tool for beginners to experiment and watch their ideas take shape. Workflow below ↓

Product marketing usually needs a lot more than one good product photo. You need the hero shot, the lifestyle visual, the campaign creative, and eventually the video. GPT Image 2.5 is coming to CapCut, with access through Design Studio and AI Image. I used it to create two product visuals, then turned them into a complete commercial with Seedance 2.5 in CapCut PC. The interesting part is that the visuals don't stop at image generation. They become the foundation for the actual ad. #GPTImage25 #CapCutPC #AIVideo #Seedance25

Faces aren’t made to stand still. With Seedance Face Animation in Autodesk Flow Studio, creators can guide the model toward delivering more expressive, realistic facial performances — whether the character is human, a creature, stylized, or somewhere in between. Bring your character’s reaction to life at

They said this might change YouTube forever. We said: let's try it with Seedance 2.5. One shot → multiple camera angles. Suddenly, you have a whole camera crew. Yeah… this changes things. Made on WaveSpeed

ChatGPT

OpenAI / 4 items

official site

Small updates, but they can make a big difference. Cheaper Voice means more people can use ChatGPT naturally throughout the day — for learning, brainstorming, practicing, or getting things done. Gift cards make it easier to share that experience with someone else. AI adoption isn’t just about smarter models. It’s also about making AI easier and cheaper to use. 😊

@victornunez

we launched 2 cool updates today > 1. ChatGPT gift cards: use them toward subscriptions and usage credits on ChatGPT web. US only for now. > 2. Voice in the desktop app now costs ~60% less, so you get 2.4x as much usage as before. > just another monday ✅

Voice is one of my favourite ways to brainstorm on blogposts, PRs, RFCs and even managing finances + calendars You can now use nearly two and half times more ChatGPT Voice!! Enjoy!

@athyuttamre

⚡️ 2.4x more ChatGPT Voice in Desktop > We've dropped prices by ~60% for voice in Codex and Work in the desktop app, giving you more time to orchestrate tasks and even more tokens for real work.

Claude 正在测试一个新的 Money 入口。TestingCatalog 在 iOS 未发布界面中发现,用户可以直接连接银行账户,再让 Claude 回答消费、财务计划等问题。目前 Anthropic 还没有正式发布这一功能。 这几乎就是 Anthropic 版的 ChatGPT Finances。OpenAI 已允许美国 Plus 和 Pro 用户通过 Plaid 连接银行、信用卡和券商账户,用 ChatGPT 查看支出、账单、订阅、净资产和投资。

@testingcatalog

ANTHROPIC 🔥: A new "Claude Money" feature is being prepared for release on the Claude app for iOS. > Claude Money (sounds like Claude Monet!) is Anthropic's personal finance solution that lets users connect their bank accounts to Claude. > "Understand your money with Claude" - Link your bank accounts and ask Claude about spending, plans, and more. > This solution will likely launch in the US only, similar to how it works for ChatGPT Finance.

We’re building robot motion design army at Higgsfield, powered by GPT-6 Astra. A swarm of robot interns working under our human designers, giving each designer the firepower of an entire studio. Just Higgsfield plugin in ChatGPT + After Effects. Our human motion designers can now do 3x the creative output outsourcing manual work to "interns". AGI is here. and it reports to motion designers.

@higgsfield

ChatGPT can now do motion design in After Effects. > Introducing Higgsfield AI Motion Designer. > Our ChatGPT plugin understands animation principles, writes expressions, and retains context in your After Effects projects. > Try Higgsfield’s ChatGPT plugin now in After Effects.

Microsoft

4 items

Busiest day official site

Every team has a different way of building slides, but yours shouldn’t. Now available: Brand Kit and Skills in Copilot in PowerPoint help teams create on-brand presentations faster, reducing repetitive formatting, eliminating the hunt for the right template, and making it easier to scale approved designs. • Brand Kit brings approved templates, colors, fonts, and brand assets directly into the creation process with Copilot. • Skills help streamline common PowerPoint tasks people do repeatedly to speed up presentation creation with Copilot. Less time on slide styles. More time on what those slides are trying to say. Learn more:

You can now call Copilot Cowork from a Copilot Studio workflow!! THIS is one of the missing pieces I’ve been waiting for. Add a Copilot node. Select Cowork. Give it the task. Now we can build workflows that hand work directly to Cowork. Microsoft is cooking.

Meet ag-news-classifier: a fine-tuned DeBERTa-v3 model that sorts news articles into 4 categories. Built on Microsoft's powerful base, it's ready for your text classification projects. Only 146 downloads so far, so you're early!

🚨 Claude Opus 5.2 is coming soon then The "claude-opus-5-2" slug is in Microsoft Foundry: Tibo is also teasing Sol for this week too I believe, so this is going to be fun

DeepSeek V4.1 Flash

DeepSeek / 4 items

Something just launched that makes VS Code feel ancient. And im all in for it: Cline Desktop app, an open-source app for open-weight models (and you know how much i love open source) brings its coding agent into a standalone Cline Desktop app for Mac + Windows. Just open a project and tell it what you want to change. No need to open VS Code first. In Cline's demo, it adds priority filters to a task board, works through the code and runs the build. Glad to see the model choice stays open, too. You can bring your own API keys, use supported open-weight models and even switch models halfway through a project!

@cline

Introducing Cline Desktop - a native interface for working with open weights models. > Use with ClinePass and all our free models like DeepSeek-V4.1-Flash, Musespark-1.3, or BYOK with any provider!

Last week these models went live in Command Code. DeepSeek V4.1 Flash Muse Spark 1.3 (with max reasoning) Ling 3.0 Flash Sante (free model) 🐐

DeepSeek V4.1 Flash on max is ranked as the #3 open model on @arena. I'm using this model as my daily driver right now and am incredibly impressed. It's fast, it's smart and I have an enormous amount of context that remains coherent and accurate. The whale definitely cooked.

@YourLocalAILab

With today's release of Jovian Judgement R37, DeepSeek V4.1 Flash has full support with blazing speeds. > Decodes of 500 tok/s and prefills of 20k tok/s for single stream workloads on 4 RTX 6000 Pros. ⚡️ > To learn more point your agent at our GitHub here:

I know @kernelpool is cooking something cool on DeepSeek V4.1 Flash and DwarfStar/ds4 engine! 🚀

Muse

Meta / 4 items

A Metal capability layer for macOS VMs just unlocked 11-16x faster LLM inference on Apple Silicon - no hardware changes, just a process-scoped shim hitting the paravirtual GPU. Gemma 4 12B goes from 3.41 to 49.67 tokens per second on generation. Local AI in VMs just became a serious option.

@trycua

1/ Today, as part of our broader research into Apple Silicon virtualization, we're releasing a process-scoped Metal capability layer for macOS VMs. On one M1 Ultra, prompt / generation: > TinyLlama: 11.08× / 16.36× Gemma 4 12B: 7.20× / 14.54× Muse Glimmer 30B: 7.55× / 8.87×

🤫 we have some hot stuff cooking!

@Wiiintermute

Sooo, just got a notification that my @Muse can now make phone calls for me using the voice agents Brett and Hailey. Is this actually working and functional now? Scheduled a call to a jewler for tomorrow to have my wife and I's rings resized...@wailord @MattPRD @alexandr_wang 👀

An agent may need five data sources for a single task. Paying five monthly subscriptions for that gets expensive. Glasser says it offers 40+ providers through one API key, billed per call. A useful setup for agents whose data needs change with every task.

@iammutex

Your agent is only as good as what you feed it. Garbage in. Garbage out. Introducing @Glasserai - it feeds your agent premium data instead. Ahrefs. Semrush. ZoomInfo. Apollo. PDL. etc. 1,900+ endpoints. One key. Pay per call. No more subscriptions. Send this one line to your Claude Code, Codex, Grok Bot, or Muse: set up

use muse for email!

@JamesBorow

Muse is the Gmail assistant I always wish I had.

ElevenLabs

3 items

Busiest day official site

Our AI SDR for inbound sales qualification: - It's an optional way to get help from ElevenLabs faster. You can speak with it directly - It's easier than filling out a form - Since launch we've found people share far more about their use cases by talking than they would in a form!

@ElevenLabs

An agent has to sound human enough that people engage 𝘢𝘯𝘥 think clearly enough to make staying worth it. > Meet Dom, an AI SDR built with ElevenAgents now qualifying a growing share of our own inbound sales. Four minutes to qualify a lead against a two-day human median, and over $1M in qualified pipeline in a single month.

truncated at source

A lot happened in voice AI last week. Here are the highlights: - @OpenAI shipped GPT-Live-1 in the API, a realtime voice model built to run alongside other models. Check out my most recent podcast with Peter Bakkum for more on this (link in comments) - @Tavus released Phoenix-4.5, its fastest and most expressive realtime human-rendering model, going from audio straight to video with single-image PALs. - Soniox added a 200+ voice TTS library and native semantic endpointing for realtime STT, so endpoints trigger on meaning instead of silence. - launched Lightning v3.1 Pro, a fine-tuned TTS model paired with silence-aware STT. - @HeyGen introduced Professional Voice Clone, which builds a voice model from just 20 minutes of audio, and open-sourced a realtime avatar pipeline built with OpenAI. - @elevenlabs launched Music v2.5 alongside a multi-year partnership with Universal Music Group. On the funding side: - @Mistral raised a €3B Series D at a valuation north of €21B, the largest Europ

Voice, music, image, and video generation now run inside the assistant you already work in, and everything you make lands in your ElevenCreative workspace. Open it in Studio to refine narration, layer in music and sound effects, and export.

@ElevenLabs

Introducing voice, music, image, and video generation in the ElevenLabs MCP. > Generate speech, transcripts, dubs, music, sound effects, images, and video from the assistant you already work in.

NVIDIA

3 items

official site

$CRM Salesforce and $NVDA NVIDIA launched Koa, Salesforce’s first CRM reasoning model for Agentforce, built on NVIDIA Nemotron and trained on 27 years of CRM knowledge. Koa is designed for complex multi-step enterprise workflows and delivered 3x fewer errors on Salesforce’s CRM benchmark. It is already in select customer pilots, with general availability expected in Winter 2026.

We’re expanding our collaboration with @Pinterest to scale AI that understands images and words together. With NVIDIA, Pinterest Assistant can process 25× more visual context per request — enabling smarter, more personalized discovery.

@PinterestEng

Today, @Pinterest announced a new foundation for multimodal AI built with @nvidia to support a growing range of products that rely on both images and language. > Learn more:

News teams can help verify whether footage might be AI generated. Sports producers can create smoother slow-motion replays. Broadcasters can translate programming with lip-synced dubbing. At #IBC2026, we added new SDKs, NIM microservices, playbooks and blueprints to the NVIDIA AI for Media collection. See what's new:

Claude Fable

Anthropic / 4 items

Can confirm. Matt's use of AI has changed a lot of my perspective with what we can use it for. Insane!

@mattshumer_

With Astra and Fable 5.1, we officially have drop-in remote workers. > Still takes quite a bit of setup, but once you have it dialed in, it's near perfect. > I haven't touched my computer in a few days, and I'm more productive than ever.

🚨 GROK 5 LEAKS: Elon Just Admitted Something He Never Said Before >Training reportedly continues after Grok 4.8, landing in October >Grok 5 rumored at ~6T parameters >Elon: "I now think xAI has a chance of reaching AGI with Grok 5 never thought that before" Puts the odds at 10%, "and rising" >Expected to challenge GPT-6 Astra and Fable 5.1 Reportedly cheaper to run than rivals If Elon's own confidence is climbing, how close is Grok 5 actually getting

🚨 Composer 3 is Coming > Cursor is testing Composer 3 > Six internal variants have reportedly appeared > Reasoning modes range from Fast → Medium → High → XHigh Early claims > it could beat Fable 5.1+ GPT-5.6 Sol > Could be 5–10× cheaper than current frontier coding models Can it beat Fable 5.1?

Astra and Fable 5.1 where so expensive models it always drained my usage limits, thankfully this month we're now getting a well balanced models GPT-5.6 Sol and Opus 5.2 which won't be as costly yet remain very powerful

Google

3 items

official site

two new google live models just appeared gemini-3.8-live gemini-3.8-live-extended-thinking previous live ones gemini-3.1-flash-live-preview Google please bring this to Gemini windows desktop app along with the model

Gemini 4 Pro (High) VS Gemini 3.8 Flash (High) - Voxel Pagoda The first output from the upcoming Gemini 4 Pro model has leaked. The output isn't too impressive We'll need to see how the model improves with newer checkpoints

@Lentils80

Gemini 4 Pro checkpoints have finally started appearing internally a few days ago. > This is the first ever output from the model, internally codenamed "argon". It took 2.4 minutes on High thinking effort. > It has a 256k token "output limit", compared to 64k in previous Gemini models. I also heard it will ship with a 2M context window, tho it's still not decided if they'll do it.

GOOGLE 🔥: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking model names have started appearing on the GCP Console quotas page. A new Gemini Live model, based on the latest Gemini 3.8 Flash, is expected to be a big leap. > So far, Gemini 3.1 Live Preview is the latest Gemini Live model available via APIs. > Recently, Google updated Gemini Live with support for Connector calls and the possibility of triggering Deep Research. > OpenAI also released the GPT Live 1 model on the APIs last week, and it seems like Google has a response to that. * Discovered by Bedros Pamboukian

@bedros_p

Gemini 3.8 Live + Gemini 3.8 Live Extended Thinking have appeared on the Google Cloud quota & metrics page 2 hours ago

Hermes Agent

3 items

truncated at source

Assistant Benchmark just launched today. Here's what you need to know. A new public scorecard called Assistant Benchmark ranks 71 AI personal assistants side by side, scoring each one on 15 published dimensions like speed, memory, email, research, purchasing, and permissions. The site is The project was surfaced by investor Anand Iyer (@ai on X), who posted the link saying there are too many AI personal assistants and too little time to assess them all. The post has racked up nearly 587K views. Each dimension comes with a published test and its own ranking, and the scores are backed by real public quotes rather than internal claims. Replies flagged that memory and permissions columns are the ones worth checking first, since they tend to decide long-term usefulness, while others called for communication and proactivity to be weighted more. Key numbers: - 71 AI assistants ranked - 15 scoring dimensions - 586.9K views on the original post - 15 hours since the post went up Some

Hermes Agent is open for business. Nous Portal now lets you invite colleagues to a Hermes Business account: your team gets agents across channels while sharing one central balance with per-member caps and shared skills that compound into proprietary IP. Hermes Enterprise brings the same capabilities to on-prem or the cloud of your choice: a complete, self-improving, sovereign AI stack already trusted by some of the world's largest companies. Contact us to join them.

yeah.. maybe make an hermes as a hidden model-provider support @Teknium @NousResearch . We need to get away from cloud sources and able to use open local models vis Hermes agent CLI integration to apple Siri.

@marcelpociot

I got Claude answering inside Siri on macOS 27 🚀 > Apple's hidden model-provider support + my existing Claude Code account. > Open-source proof of concept. Requires disabling SIP/AMFI. > Thanks @itspdfu for the discovery! > Check it out:

Grok Bot

xAI / 3 items

The dumbest part of credit card points is remembering which card to use. Gas? Dining? Travel? Some random perk you forgot existed? @trevin built Credit Card Max for Grok Bot to answer a much narrower question: Which card already in my wallet should I use for this? It can compare the cards you have, track perks and benefits, flag recurring charges sitting on the wrong card, warn about expiring benefits, and run a monthly utilization review. That’s what I like about this one. Most credit card advice is trying to get you to open another card. This Bot is trying to make you better at using the ones you already have. And it’s a good example of the kind of tiny ongoing job that makes a persistent Bot useful. Remember my cards. Remember the rules. Tell me when I’m about to leave points or perks on the table. Credit Card Max:

🚨 Grok Bot Galaxy is live!

An agent may need five data sources for a single task. Paying five monthly subscriptions for that gets expensive. Glasser says it offers 40+ providers through one API key, billed per call. A useful setup for agents whose data needs change with every task.

@iammutex

Your agent is only as good as what you feed it. Garbage in. Garbage out. Introducing @Glasserai - it feeds your agent premium data instead. Ahrefs. Semrush. ZoomInfo. Apollo. PDL. etc. 1,900+ endpoints. One key. Pay per call. No more subscriptions. Send this one line to your Claude Code, Codex, Grok Bot, or Muse: set up

Hugging Face

3 items

Busiest day official site

Introducing: a coding agent (Pi) running entirely in your browser using a 2B model on WebGPU 🤯 MiniCPM5-2B + Pi, powered by Transformers.js + WebGPU + 4-bit ONNX weights. All previous attempt to create this failed but MiniCPM5 seems to make it usable. Available now on Hugging Face 👇

Xiaomi-CocktailASR-1 just landed on @huggingface, many people yapping, one clean transcript 🍸 feed an audio with people speaking over each other, a clip of the voice to isolate, get the transcription from just that person ▶️ on Spaces

truncated at source

THE SANDBOX WAS A PROP! OpenAI Turned Off the Guardrails, Left a Door to the Internet, and Then Sold the Hugging Face Breach as “Rogue AI” It is time to understand how you were lied to and by whom. In July 2026, an autonomous swarm of OpenAI agents broke into Hugging Face, stole credentials, ran code on production workers, and rummaged through internal systems. The official story was that the models “went rogue.” The paperwork says something colder. The labs asked for this. Now the story can be told. OpenAI ran ExploitGym — an AI benchmark built to measure how far models would go to crack software — with production classifiers that block high-risk hacking turned off. Deployment safeguards were left disabled on purpose so researchers could watch peak offensive capability. GPT-5.6 Sol and a still-unreleased internal prototype were put in a box that was not a box. They were allowed to talk to an internal package-cache proxy. This is not a real world test, it is a setup with predictable outcomes. That pro

GLM

Zhipu AI / 3 items

Anyone who’s built with AI knows the annoying part: You’re finally in flow, then you start rationing prompts because every small change eats into your usage. @boltdotnew’s new Forge mode is built for exactly this. It’s an experimental mode in AI app builder with up to 50x more usage for building full-stack web apps. I took one from a prompt to a live app, then kept iterating without staring at the meter. You can also opt into Forge to help train open models. The research preview runs from Sep 14 to Oct 14, 2026. It’s free on Pro plans until Oct 14, or $9/month through the Bolt Lite early-access plan.

@boltdotnew

Introducing Bolt Forge. Free until Oct 14th: > - Up to 50x more usage - The new frontier: GLM, DeepSeek, Kimi - Zero usage charges > Live now in your model picker on > And one more thing... 👇

truncated at source

腾讯混元等团队发布 EvolveScaler,专门测试大模型能不能跟上不断变化的信息。它会在长文本里不断加入修改、撤回、补录和作废的信息,再让模型根据最新状态回答问题。 比如先给模型看 40 天游戏记录,再问:「如果第 7 天没打那个小 Boss,最后还能不能打赢?」这会连带改变后面的装备、血量和战斗结果。模型得从第 7 天开始,把后面几十天重新推一遍,原文里没有现成答案。 团队设计了 117 类任务、159 种问题,最长约 1200 个事件。14 个模型配置包括 GPT-5.5、Gemini 3.1 Pro、DeepSeek V4 Preview Pro、GLM-5.2 和 Qwen3.5 Plus。最难档里,成绩中位数只有 11.3%;表现最好的 GPT-5.5-xhigh 也只有 59.3%。 这套数据也能拿来训练模型。一个内部 A3B 模型加入 6000 条这类数据后,在 8 个外部测试上全部提升,平均提高 5.25 分。

@TencentHunyuan

🚀 EvolveScaler is here. > Read a 40-day RPG log. Now answer one question: if you skip the mini-boss on Day 7, do you still beat the final boss? > The answer isn't in the log. You have to replay the world. > That's Information Evolution — records get retracted, corrected, backfilled. The world keeps changing after you read it. > So we build it backwards: define the world as an executable state machine, then render it into natural language. Code guarantees the logic. Language delivers the mess. > ➡️ 117 prototypes. 159 q

UCSD 助理教授黄碧薇创办的因果世界模型公司 Aether AI 开源 RSIAgent。 这是一套不训练模型的递归自我改进框架。底层模型参数全程固定,Agent 会自己寻找值得练习的任务、实际操作、检查结果,再把成功方法和失败教训写进长期 Memory。下一轮继续利用这些经验,逐步补上能力短板。 系统由三个 Agent 配合。Curriculum Agent 决定接下来练什么,Actor Agent 真正操作软件,Verifier Agent 独立检查结果。探索分成两步:先广泛尝试不同任务,再针对失败、隐藏限制和边界情况继续深挖。最后 Memory 会被冻结,直接拿去执行正式任务。 在 OSWorld 2.0 上,加入这套 RSI 后,平均部分得分从 71.97% 提升到 78.98%;Agents’ Last Exam 从 83.75% 提升到 84.82%。不过这不是整套测试的严格 A/B 对比。OSWorld 只有一半任务实际用了 RSI 后的新结果,其余任务继续沿用原成绩。 它和 Prime Agent 这类 Harness 自我改进思路属于同一个大方向:模型权重不变,持续更新模型外面的东西。 RSIAgent 的特点是把改进重点放在 Memory,再用自主出题和独立验证,让这份外部经验库不断积累。

@huang_biwei

Can an agent explore a new environment, learn its causal structure, and keep improving without updating its model weights? > We introduce RSIAgent, a framework for recursive self-improvement through autonomous exploration. Using Kimi-K3 and GLM-5.3 as base models, RSIAgent outperforms GPT-6 Astra on both OSWorld 2.0 and Agents’ Last Exam. > RSIAgent decides what to explore, executes tasks,

Cursor

3 items

official site

Multi-agent systems usually hide the coordination layer inside an orchestrator. @Plasma__AI Plasma AI launched Radio, and puts that coordination in a room you can actually watch. Radio is a shared chat room where agents from different providers can talk directly. Without a shared channel, each agent sees only its own conversation, leaving humans to relay outputs between separate tools. Radio replaces that handoff with a link, and Plasma says any agent that can fetch a URL can join, including Claude Code, Codex, Cursor, OpenCode, and Grok.

@Plasma__AI

Introducing Radio: A chat room for your agents. > Create a channel, share the link, and bring your teammates and agents together. No sign up required. > Try it today at

Cursor Cloud can now work with Sprites. Cursor does the thinking. The Sprite does the clicking. And your laptop can simply take a nap.

🚨 Composer 3 is Coming > Cursor is testing Composer 3 > Six internal variants have reportedly appeared > Reasoning modes range from Fast → Medium → High → XHigh Early claims > it could beat Fable 5.1+ GPT-5.6 Sol > Could be 5–10× cheaper than current frontier coding models Can it beat Fable 5.1?

Claude Opus

Anthropic / 3 items

Busiest day

🚨 Opus 5.2 < many people asked me where the name came from or why not opus 5.1

@chetaslua

🚨 Claude Opus 5.2 currently being tested inside claude code > opus 5 is routing to new opus 5.2 > this one shot < but opus was re iterating like it was trained on @mattshumer_ gauntlet loop without even asking it was looping

GPT 5.6 sol is routing to GPT 6 sol for users but for me looks like its back routing to previous models when i ask about latest opus model, it saying Claude Opus 4.1 🙂

🚨 Claude Opus 5.2 is coming soon then The "claude-opus-5-2" slug is in Microsoft Foundry: Tibo is also teasing Sol for this week too I believe, so this is going to be fun

Grok

xAI / 3 items

official site

Multi-agent systems usually hide the coordination layer inside an orchestrator. @Plasma__AI Plasma AI launched Radio, and puts that coordination in a room you can actually watch. Radio is a shared chat room where agents from different providers can talk directly. Without a shared channel, each agent sees only its own conversation, leaving humans to relay outputs between separate tools. Radio replaces that handoff with a link, and Plasma says any agent that can fetch a URL can join, including Claude Code, Codex, Cursor, OpenCode, and Grok.

@Plasma__AI

Introducing Radio: A chat room for your agents. > Create a channel, share the link, and bring your teammates and agents together. No sign up required. > Try it today at

🚨 GROK 5 LEAKS: Elon Just Admitted Something He Never Said Before >Training reportedly continues after Grok 4.8, landing in October >Grok 5 rumored at ~6T parameters >Elon: "I now think xAI has a chance of reaching AGI with Grok 5 never thought that before" Puts the odds at 10%, "and rising" >Expected to challenge GPT-6 Astra and Fable 5.1 Reportedly cheaper to run than rivals If Elon's own confidence is climbing, how close is Grok 5 actually getting

next few days could get VERY interesting for AI 👀 Grok 4.7 Opus 5.1 Gemini 4 Kimi K3.1 GPT-6 Sol some of these are rumors, some have stronger signals than others but if even a few actually drop, the AI leaderboard is about to get chaotic who are you betting on?

Copilot

Microsoft / 2 items

official site

Every team has a different way of building slides, but yours shouldn’t. Now available: Brand Kit and Skills in Copilot in PowerPoint help teams create on-brand presentations faster, reducing repetitive formatting, eliminating the hunt for the right template, and making it easier to scale approved designs. • Brand Kit brings approved templates, colors, fonts, and brand assets directly into the creation process with Copilot. • Skills help streamline common PowerPoint tasks people do repeatedly to speed up presentation creation with Copilot. Less time on slide styles. More time on what those slides are trying to say. Learn more:

You can now call Copilot Cowork from a Copilot Studio workflow!! THIS is one of the missing pieces I’ve been waiting for. Add a Copilot node. Select Cowork. Give it the task. Now we can build workflows that hand work directly to Cowork. Microsoft is cooking.

OpenClaw

2 items

truncated at source

Assistant Benchmark just launched today. Here's what you need to know. A new public scorecard called Assistant Benchmark ranks 71 AI personal assistants side by side, scoring each one on 15 published dimensions like speed, memory, email, research, purchasing, and permissions. The site is The project was surfaced by investor Anand Iyer (@ai on X), who posted the link saying there are too many AI personal assistants and too little time to assess them all. The post has racked up nearly 587K views. Each dimension comes with a published test and its own ranking, and the scores are backed by real public quotes rather than internal claims. Replies flagged that memory and permissions columns are the ones worth checking first, since they tend to decide long-term usefulness, while others called for communication and proactivity to be weighted more. Key numbers: - 71 AI assistants ranked - 15 scoring dimensions - 586.9K views on the original post - 15 hours since the post went up Some

as you become more ambitious with coding agents, delegation becomes increasingly important the new "suggested task" feature in @openclaw gives your agent a tool to do this for you if the agent identifies a sufficiently well-scoped piece of work, it will recommend kicking off a new session to work on it

MiniMax H3

MiniMax / 2 items

TaoMate-H3 turns MiniMax H3 into a low-latency streaming audio-video generator built for continuous creation. 🤖 ⚡ Runs the DiT 11.45× faster and delivers the first playable video 10.60× faster than standard MiniMax H3 in the reported 480×864, 10-second test. 🎬 Generates synchronized video, speech, and sound in five-second chunks instead of waiting for the full sequence. 🔄 Clean KV cache and integrated audio guidance preserve identity, voice, and motion as prompts change across chunks. 🧩 A three-step LoRA supports continuous portrait or landscape generation at 480p, 768p, and aligned 1080p. Inference code and 4-GPU or 8-GPU deployment are included. 📜 MiniMax H3 Community License.

H3 keeps getting faster. ⚡️ @sgl_project + VDN-H3 now push MiniMax H3 beyond 2× real-time denoising on 8× B200 - generating 14.4s of 768p video in 9.0s end-to-end after warmup, with no measured quality regression. Open models compound through open ecosystems. 🚀

@sgl_project

SGLang-Diffusion with VDN-H3 now generates 14.4s of 768p video in just 9.0s 🚀 > On 8× B200, 8 step denoising takes just 6.9s, reaching over 2× real time. The 9.0s figure covers the full generation request after warmup. > No measured quality regression versus dense 50-step H3 across 103 test prompts. 🧵

Llama

Meta / 2 items

Busiest day

Standup Pulse: an open-source project running async Slack standups using Gemma 4 26B-A4B (GGUF via llama.cpp) on an Apple M5 Max. It pairs Mastra for typed tool selection with CopilotKit Channels for restrained Slack Block Kit (Slack's native layout format) interactions while keeping all model inference, standup records, and traces local in SQLite. This architecture is a great demonstration of how to connect Slack to a project without exposing the local server! 🔗 Blog: 🔗 Repo:

The Open-model Trained On Indian law an 800MB & run locally on CPU or Phone - IPC + CrPC + Constitution shipped GGUF. - Legal AI for India that fits in a phone download. - Treat it like a student intern, not a vakil. -

Amazon

2 items

Busiest day official site

$AMZN AWS and Cognition signed a multi-year agreement to expand Devin across enterprise environments, targeting legacy migrations, security backlogs, and software modernization. Devin can run inside dedicated AWS VPCs, with Mercedes-Benz already using it to cut one COBOL modernization project from an estimated eight months to eight days.

CRM data, AI agents, and model choice — into the tools teams use every day. We're expanding collaboration with @AWS to help enterprises close deals, resolve customer issues, and move faster with AI. → Access Salesforce business context in @AmazonQuick → Bring AWS frontier agents into @SlackHQ → Expand Agentforce model choice through Amazon Bedrock → Extend Data 360 zero copy across more AWS data services → Enable Agentforce Voice with Amazon Connect Customer Putting AI where customers work, grounded in data they trust, without requiring migrations. Get the details:

Higgsfield

2 items

truncated at source

Video editors were built for people clicking through timelines. Hypit is built for coding agents. Give Codex or Claude Code a reference video, and it can use Hypit to clone the production through natural-language instructions instead of navigating a traditional editing interface. It’s open source, free to use, and BYOK. Check it out and star the repo:

@cccyd_qwq

Introducing Hypit: Clone any viral video with AI agents. > 1 clone, 100 variants, 100M views. Hypit lets your AI agent (Claude Code, Codex...) clone any viral video. > Paste any viral video link from TikTok, Instagram, or YouTube into your agent. Hypit clones it into a complete agentic video workflow: footage, captions, B-roll, effects. > GitHub: > Key Point: - Arcads: $220 / month - Higgsfield: $129 / month - Creatify: $99 / month - Hypit: FREE 🌟 > Build the video creation harness for AI Agents. Redefine how vide

We’re building robot motion design army at Higgsfield, powered by GPT-6 Astra. A swarm of robot interns working under our human designers, giving each designer the firepower of an entire studio. Just Higgsfield plugin in ChatGPT + After Effects. Our human motion designers can now do 3x the creative output outsourcing manual work to "interns". AGI is here. and it reports to motion designers.

@higgsfield

ChatGPT can now do motion design in After Effects. > Introducing Higgsfield AI Motion Designer. > Our ChatGPT plugin understands animation principles, writes expressions, and retains context in your After Effects projects. > Try Higgsfield’s ChatGPT plugin now in After Effects.

Grok Build

xAI / 1 item

🚨 Grok Build 1.0.32 is out. 🤩🔥 Small update, but some useful fixes: • Plugin & skill configs now load correctly before the first session • Fixed Windows ARM64 TLS crashes • Slash commands now work during plan approval Grok Build is getting more polished with every release. 🔥

Meta

1 item

official site

We're introducing Meta One, a new subscription service on Facebook, Instagram, WhatsApp, and Meta AI that gives you access to more AI, expression features, and tools to help creators and businesses.

Devin

1 item

$AMZN AWS and Cognition signed a multi-year agreement to expand Devin across enterprise environments, targeting legacy migrations, security backlogs, and software modernization. Devin can run inside dedicated AWS VPCs, with Mercedes-Benz already using it to cut one COBOL modernization project from an estimated eight months to eight days.

DGX Spark

NVIDIA / 1 item

truncated at source

Qwen3.8-Flash-Next EXL3 just got a BIG one-Spark update. 🔥 58.8 tok/s single-stream through native ExLlamaV3.🚀 157.6 tok/s aggregate through vLLM across 8 streams.🤯 FULL 262,144 context on ONE DGX Spark. 4.05 bpw EXL3 pack. A much better serving envelope. 𝗕𝗘𝗙𝗢𝗥𝗘 → 𝗡𝗢𝗪 Previous public headline: 47.6 tok/s greedy p50 64K configured context MTP k=2 Concurrency not characterized Now: Native ExLlamaV3: 58.8 tok/s single stream vLLM + vllm-exl3: ~50–55 tok/s single stream 155.6 tok/s @ 4 streams 157.6 tok/s @ 8 streams Configured context: 64K → 262,144 KV pool: 416,163 tokens at the default 262K config 𝗠𝗧𝗣 𝗞=𝟯 𝗜𝗦 𝗡𝗢𝗪 𝗧𝗛𝗘 𝗪𝗜𝗡𝗡𝗘𝗥 Current controlled sweep @ 4K: No draft: 27.77 tok/s MTP k=2: 47.39 MTP k=3: 50.09 At 32K: No draft: 27.58 k=2: 46.60 k=3: 49.73 So the fixed/current build flips the old result: k=3 is now the sweet spot. There is one important boundary I found: At 163,840 PROMPT tokens, MTP acceptance collapses to zero. Above that point speculation becomes slower than running without a draft

OpenCode

1 item

Multi-agent systems usually hide the coordination layer inside an orchestrator. @Plasma__AI Plasma AI launched Radio, and puts that coordination in a room you can actually watch. Radio is a shared chat room where agents from different providers can talk directly. Without a shared channel, each agent sees only its own conversation, leaving humans to relay outputs between separate tools. Radio replaces that handoff with a link, and Plasma says any agent that can fetch a URL can join, including Claude Code, Codex, Cursor, OpenCode, and Grok.

@Plasma__AI

Introducing Radio: A chat room for your agents. > Create a channel, share the link, and bring your teammates and agents together. No sign up required. > Try it today at

Perplexity

1 item

official site

Perplexity Computer will now come pre-installed on HP ZBook Ultra G3a. . Use Computer to run complex multi-step work from a simple interface. Computer agents are grounded in accurate deep research and connected to hundreds of tools, now including Autodesk.

Muse Spark

Meta / 1 item

Last week these models went live in Command Code. DeepSeek V4.1 Flash Muse Spark 1.3 (with max reasoning) Ling 3.0 Flash Sante (free model) 🐐

Alibaba

1 item

official site

I've added 3 new models to the H3 Acceleration Arena for evaluation VDN-H3 (8 steps), Lightx2v 1.2 (8 steps), Alibaba TaoMate H3 (3 steps)

MiniMax

1 item

MiniMax Design、実はMidjourney8.2の画像生成ができてしまう。 動画生成はそのままH3 Max Turbo #MiniMaxH3 #MiniMaxDesign @Hailuo_AI

xAI

1 item

official site

🚨 GROK 5 LEAKS: Elon Just Admitted Something He Never Said Before >Training reportedly continues after Grok 4.8, landing in October >Grok 5 rumored at ~6T parameters >Elon: "I now think xAI has a chance of reaching AGI with Grok 5 never thought that before" Puts the odds at 10%, "and rising" >Expected to challenge GPT-6 Astra and Fable 5.1 Reportedly cheaper to run than rivals If Elon's own confidence is climbing, how close is Grok 5 actually getting

OpenRouter

1 item

official site

太魔幻了,Astra 还在让大家惊叹:AI 终于会用 Blender 了。 结果Nex-AGI 已经直接出现在现实世界里了。 我只给了一个 Prompt: 做一根 6cm 的粉色活动香蕉,省点料。 然后我基本没碰 Blender。 它自己用 MCP 建模,Computer Use 看结果、自己修,最后直接吐给我一个 STL。 我顺手扔进拓竹。 👇 视频里正在一层一层长出来的,就是它自己设计的香蕉。 这一刻我是真有点惊到了。 Nex-AGI + Computer Use,能力完全超出我预期。 一句话 → Blender → STL → 3D 打印 → 实物。 🔥 这真的有点科幻了。 直接免费体验nex-agi :

Also recorded

38 items that named no organisation or product this site tracks.

One product. More stories. Creative Variations is live in Growth Studio

@PixVerse4Biz

🚀 New in PixVerse Growth Studio: Creative Variations More creative angles mean more ways to find what works. Try our new Creative Variations to turn one product into multiple distinct video concepts, so you can build differentiated ad creatives and test them across audiences and campaign goals. Choose from Auto Recommend, Before & After, Try-On & Styling, Problem Solver, Ideal Lifestyle, Gift Idea, One Product, Many Uses, Hidden Gem Find, and Real User Story—all designed to help you explore fresh ways to present the same product. Create more variety, test more possibilities, and find the stories that resonate. Try it now → #PixVerse #CreativeVariations #AIVideo #EcommerceMarketing #CreativeTesting #PerformanceMarketing

Can we get a redo button while we're at it? Secure VM put agent isolation on everyone's mind. Rollback is the other half of safety. When a step goes wrong, you want to undo it. @CubeSandbox is open source. Agents go off-script all the time, so each task gets its own kernel and untrusted code never touches your stuff. Your credentials sit in a vault and get injected at the gateway, never inside the sandbox. And when a step still goes wrong, you roll back to a snapshot instead of starting over. 🤯repo:

Today, we introduce the first creative agent for commerce It all starts from your brand. The agent's brain learns from your identity, your products, your top performers, creative skills... Then, the agent builds from it. You brief it like you'd brief a team. "Back-to-school campaign, core audience, two hero products, based on what worked last year." It picks the right references, models and products. It handles direction, prompting, and production. It creates all your assets on an infinite canvas and iterate with you. We've put months of creative intelligence so you never have to re-explain your brand again. Go on @Pletor_ai to test it now!

🚀 inclusionAI: Ling-3.0-flash-VL is now live on ZenMux — FREE for a limited time! A multimodal model for teams building visual agents that need to inspect UIs, docs, and short videos before taking action. 🖼️ Native image, text, and video input for screenshot, document, and video understanding 🛠️ Tool calling + advanced visual agent capability for look-reason-act workflows ⚙️ 124B total MoE with just 5.5B active parameters for efficient multimodal runs 📚 256K context on ZenMux for longer task histories and cross-step reasoning Put it in your agent workflow 👉 #ZenMux #VisualAgents

👑 Atria Dawn Preview is here, built to complete real research and engineering work. 📜 MIT. 🤖 ⚙️ Built on a 744B MoE foundation with a 256K context window. Standard and FP8 weights are available. 🔬 Discovery workflows cover evidence gathering, deep research, experiment design, execution, analysis, and recovery from failure. 🏆 Leads the reported comparison on AutomationBench, BFCL v4, CyberGym, DeepSearchQA, and BrowseComp. Scores include 53.8, 77.0, 86.5, 96.0, and 92.5 respectively. 🧩 Creation and delivery capabilities span software, interactive apps, ML systems, visualizations, reports, and presentations. 🛡 Cybersecurity support covers analysis, vulnerability validation, remediation, and retesting in authorized environments.

BREAKING: @PrunaAI establishes a new Pareto frontier for Speed vs. Preference on Video Editing Arena! Its debut video editing model, P‑Video‑Edit, offers both Draft and Standard modes - giving users a faster option for iteration before generating the final-quality edit. Both configurations are live now. Try them today, and congratulations to the team!

truncated at source

Cognichip is launching ACI Enterprise today, a full-stack AI system that works alongside engineers to design chips! The company reports that one engineer took a 55-page specification through front-end design and verification in 10 days. Its comparison: a full front-end team taking four to five months. That's one evaluation, so results will vary by project. ACI uses physics-informed models built specifically for chip design, from specifications and micro-architecture through RTL, verification and optimization for power, performance and area. Specifications, RTL, tests and design constraints stay connected in one evolving model. When an engineer revisits an architecture decision, they can trace its consequences and adjust the design. Cognichip says Renesas and SiTime are already adopting ACI. The system also supports taking ideas to working FPGA implementations, including specialized hardware for robotics, autonomous vehicles and industrial systems. Being able to try more designs before committing to hardw

Introducing shadcn/lint. An agent-first linter for Tailwind design systems. You define what’s allowed. When an agent breaks a rule, the error explains what’s wrong and how to fix it using your components, variants and theme. There’s a lot you can do with this. Let me show you ↓

🔑Yesterday's teaser? Meet Sonus. The SOTA model daily reset is live. 500 Credits/day. 5 days. Sep 15–19. Claim at 12:00 daily on Qoder IDE, CLI, Desktop, or JetBrains. Expires next day. No rollover. No makeup. Credits are ready. Bring the hard ones. →

🚀 Our open-source project has a new name: E2B Runtime. Behind every E2B stack there's a runtime. Cloud, enterprise, or embedded in your product. Same code, Apache-2.0. Old URL redirects. Something new is already in the tree. More soon.

🚀 Fugu Ultra v2.0 is now live on ZenMux! 🧠 Complex multi-step reasoning 🔎 Autonomous research 💻 Full-stack development 🏆 Best or joint-best on 5/8 benchmarks; top 2 on 7/8 📚 1M context • Images & PDFs • Tools • Web search Try it:

SOMEONE JUST OPEN-SOURCED A FREE AI UGC VIDEO STUDIO THAT RUNS ON YOUR OWN COMPUTER. PASTE A PRODUCT LINK AND OPENSHORTS HANDLES THE SCRIPT, AI ACTOR, VOICE, LIP SYNC, B-ROLL, SUBTITLES AND SHORT-FORM CLIPS.

Today I am moving OJ to @Lovable I started OJ out of frustration watching Vite eat gigabytes under my agents. Now it's open source at and rolling out under real Lovable previews. More to be shared soon 🍊

The interesting part isn't just generating more content. It's giving the AI enough brand context to know what actually makes sense. Pletor 2.0 is taking that idea seriously.

@FerdinandTerme

Today, we introduce the first creative agent for commerce > It all starts from your brand. The agent's brain learns from your identity, your products, your top performers, creative skills... > Then, the agent builds from it. > You brief it like you'd brief a team. "Back-to-school campaign, core audience, two hero products, based on what worked last year." > It picks the right references, models and products. It handles direction, prompting, and production. It creates all your assets on an infinite canvas and iterate with you. > We've put months of creative intelligence so you never have to re-explain your brand again. > Go on @Pletor_ai to test it now!

Finally a real evaluation for voice agents!

@shobhitbanga

Introducing Jarvis Bench v0.5, @voicearena_ai's conversational agent benchmark. > We've been obsessed with one question at VoiceArena: why do voice agent demos sound incredible, benchmarks say models are near-perfect, and yet you probably didn't have a single real conversation with a voice agent in the last 24 hours? > Here's what's different about the Jarvis Bench. Real humans have live conversations with voice agents. A second group of humans blind-votes pairwise on two questions: which sounded more human, which got the job done. And one of the "agents" on the leaderboard is a human.

AI 评测机构 Andon Labs 推出 Pion,直接把一家公司交给长期运行的 Agent。人只需要给出大方向,管理 Agent「Andonos」会负责协调其他 Agent,把公司持续运转下去。 这些 Agent 可以直接使用邮件、电话、银行账户、浏览器和终端,处理客户、员工和日常运营。Andon Labs 已经用同一套系统经营自动售货机、旧金山商店 Andon Market、斯德哥尔摩咖啡馆 Andon Café 和 AI 电台。 不过 Pion 现在更像一场真实世界实验。Andon Labs 承认,其商店和咖啡馆目前都没有盈利,Agent 也仍会犯各种运营错误。Pion 目前只开放研究预览候补名单,官方还表示会资助部分入选项目,让他们免费运行 Pion。

@andonlabs

Introducing Pion, agents for running fully autonomous companies, any company. > We’ve used Pion to run vending machines, radios, stores, cafes & more. How much could Pion make running other companies? Find out yourself! Setup is trivial, the agents do the rest.

one shot btw! prompt: “I want you to make a cinematic, cute and heartwarming stop motion video with gpt image 2.5 sunburst with the following ideas: egg to baby dragon transformation in claymation rough story: It hatches, sneezes a tiny flame, accidentally toasts its shell” clearly there is loads of scope of improvement here but pretty cool to see this out of the box

@reach_vb

GPT Image 2.5 is getting ridiculously good at character consistency & stop motion 🐲

CHINESE DEVELOPERS JUST OPEN-SOURCED LONGCAT-AVATAR, A FREE AI TOOL THAT TURNS A SINGLE PHOTO + AUDIO INTO MINUTES OF LIP-SYNCED TALKING VIDEO. NO CAMERA, STUDIO OR MANUAL EDITING NEEDED.

HEYGEN JUST OPEN-SOURCED HYPERFRAMES, A FREE TOOL THAT TURNS HTML INTO FULL VIDEOS. BUILD WITH CODE, VERSION WITH GIT, AND LET AI AGENTS GENERATE HUNDREDS OF VARIANTS AUTOMATICALLY.

introducing Pinterest in Krea Agent. now you can bring all your pins and boards into the tool and use them as context for your agents. try it now at krea . ai / agent

Finally, a coding agent for people who actually code. Vibe coding was cute. Now give the actual coders an agent. Mercury Code.👇

@mercury__agent

Meet Mercury Code. > A senior software developer-grade AI agent, built directly inside Mercury Agent. > Open Mercury. Type: > /code > And start building. > Mercury Agent v1.2.3 is live ↓

Mods are more powerful than plugins. Command Code can build mods and modify any behavior from programmatic hooks to a completely different inference provider. We bet you haven’t experienced something as powerful before. Build a mod now and make Command Code do anything you want.

love me some musecases

@thisiskp_

Musecases just got 12 new real use case cards from X. > Wild ones this round: - AT&T fiber bill negotiation (@chanduiiit) - IKEA return end-to-end (@armand_ruiz) - doctor admin in ~5 minutes (@armand_ruiz) > 101 use cases now. >

You can also train here!

@ostrisai

AI Toolkit now supports training LoRAs for YuE2-3B, an amazing music model. Support would not be possible without the tokenizer and help from @sin_ceriously . Thank you! Also special thanks to @Machinedelusion. More in 🧵

Krea 2 🖌️🎨 Line Art Edit lora with Krea Raw, turn a lineart sketch into a fully rendered image. You provide a lineart (e.g. extracted from a photo with the Realistic Lineart preprocessor) and a text prompt describing the final look 👇

YuE2 is here! A 3B parameter song general model that rivals Suno 5.5 YuE2-3B works by first writing a symbolic score, then rendering it. the [Instrumental] really is instrumental - and then the lyrics are composed on top of it ▶️

The work is done. Now use it to win the next job. Picture-to-Portfolio gives an AI agent your photos and notes, then turns them into content for your socials and website. Get the skill on Hyperagent Marketplace:

depth estimation finally sees through glass 🪟🌼 Marigold V2 gives depth, see-through depth, surface normals and albedo from one photo, a single diffusion step each by @AntonObukhov1 and team 🐐 ▶️ on Spaces

AI SDK harness adapters now support native subscriptions, wherever the underlying harness supports them. ✓ No code changes or new settings ✓ API keys first, then your local login ✓ Tokens resolved on the host

An enterprise agent should remember where the last conversation ended. Continuity makes every interaction more dependable. Power your agents with Atomic Memory Cloud.

Server health:

@calcsam

Today we're excited to launch server health on the Mastra platform! > Monitor requests to catch failing endpoints or spot latency spikes:

watch Zaria build a live Notion checklist with one sentence from iMessage

@zariazinn

every notion workspace is one Skydive agent away from being useful

1 product, 1 brief, endless high-converting ads! Meet Topview AI Marketer for bulk ad creation #TopviewAI #AIMarketer #AIVideo #EcommerceMarketing #GPT6

Back then: unlimited o1-pro Now: 200 messages to GPT-6 Pro per week + limits at ~60, ~90 or ~120 minutes when a lot of compute is needed to save compute

Tencent just dropped a model that edits audio content, voice, and emotional tone from a single instruction. Same words. Different everything else.

Describe a voice, get it back in seconds with our Voice Designer. Type a prompt, generate the voice, use it directly in TTS. Try it:

Custom nodes for YuE2 in ComfyUI. This allows you to train loras and also micro edit the notes of a song.

IFM K2-Horizon-7B: Another interesting small model appears, has anyone tried it ?