Skip to content
B Bloger.fm

AI briefing

16 September 2026

203 items were recorded on 16 September 2026, filed under 49 organisations and products.

That is up from 145 the day before.

Most covered: GPT-6 Astra (23), OpenAI (20) and Codex (16).

Jev and ByteDance were recorded for the first time in this 8-day window.

67 of the day's items named no organisation or product this site tracks; they are listed under “Also recorded”.

19 items reached the source feed's 1024-character limit and are cut off mid-text; each is marked “truncated at source”.

Compiled by Bloger.fm Editorial Desk

Compiled from a monitored feed of public AI announcements. Items are quoted or summarised as recorded and are not independently verified — see the editorial policy.

GPT-6 Astra

OpenAI / 23 items

Busiest day

Today we're introducing ChatGPT Astra for "Web to App" Clone any website into a mobile app. Just paste a URL. GPT-6 Astra controls your Mac to rebuild the original website as a *native* mobile app, then submits it to the app stores for you. We've been using this mostly for iOS apps. > @chhddavid: >

GLM Coding 2.0 Leak: Coming Soon 🔥 >Expected to beat Mythos 5.1 and GPT-6 Astra >October release window reportedly targeted >Computer use expected to be a top priority >Expected to remain open-weight >1M-token context window reportedly carried over Rumored 3T+ parameter model

Wenfeng: "that's not continual learning, this is prompt engineering bullshit" but to be fair: that may be all it takes for practical purposes. These things already have superhuman priors, can RLM over infinite databases, and think crazy fast. Do we *need* parametric updates?

@NeoCognition

Introducing ApprenticeBench: computer use + continual learning on a real job. > We show Fable 5.1 and GPT-6 Astra can now continually learn on a job and surpass human professionals. A decisive step change in AI's job readiness. > No FDEs. Agents deploy themselves into the job. 🧵

THIS IS GETTING RIDICULOUS Anyone can now build Duolingo-style retention flows with ONE prompt. Here’s how it works: > connect your customer list > tell @noimos_ai the goal > AI builds the flow > reacts to user behavior > sends personalized messages The agent handles all of this autonomously 🤯

@noimos_ai

Introducing Astra for Customer Retention & Growth. > Just describe your goal. NoimosAI autonomously creates and delivers personalized campaigns that keep customers engaged and coming back. > Maximize customer lifetime value. All on a single platform.

Sam Altman after finding out his $20 GPT-6 Astra quietly opened a trading desk on Robinhood Chain, worked all night while the NYSE was closed, and closed its 84th call before anyone noticed it existed..

@slash1sol

I GAVE GPT-6 ASTRA AND MINARA ONE JOB: WATCH ALL 24 ROBINHOOD STOCK TOKENS UNTIL THE PRICE STARTS LYING -> THREE WEEKS LATER THE DESK THAT DOES IT IS RUNNING > 24 tickers, 4,320 pool reads an hour, 103,680 a day and 84 gaps called, 95% of them closed. > It's live and it's free: > Why a Stock Token stops tracking the share it is named after: > A meme pair locks real shares inside a pool and the float on chain goes thin. > Minting and burning run on a schedule, so outside that window supply cannot answer demand. > After 4PM the oracle stands still while the pool keeps trading anyway. > The pool price drifts off the real price, and that drift is the arbitrage -- buy the cheap leg, short the rich one, wait for them to meet. >

Developers are still having fun with a digital fruit fly brain. This time, they’ve got it playing Deadlock with help from GPT Astra The 166,000-neuron model was trained on recordings of other players’ matches. Computer vision helps it understand what’s happening on screen, while a separate algorithm handles aiming The fly runs around the map, joins fights, and plays best as Graves. According to its creator, it even performs better than his teammates > @Mikadzyki_NFT: >

+ 17 more items

very cool result showing how wet lab data enables a specialized model to beat gpt-6 astra at a task at the frontier of science! in general scaling is great and obviously i am a believer in it, but probably the more we approach the frontier of science, the more specialized data matters and that gives task-specific models a chance. this specialized data is usually private and is probably a real moat it should be in principle true that a task specific model will probably do better at scientific discovery just because it can use more of its parameters for the task you care about

@LiamFedus

We built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, and then help us decide what to try next. > Using only 1,300 H200s, plus months of our experimental data, we mid-trained and RL’d an open-source model to surpass GPT-6 Astra on our analysis benchmark. We call it Neon. > This

Astra is mogging fable soo much I’m use to see any benchmark fable actually beat it on

@j_dekoninck

We are releasing the latest version of BrokenArXiv and ArXivMath! These benchmarks now focus on conjectures that were refuted in the last month on ArXiv, and models are executed within a harness instead of directly via API. > Performance remains impressive, with GPT-6 Astra on top

Woow It launched 14 sub agents in parallel with GPT-6 Pro and ran for 165 minutes, more than two and a half hours , I specifically told it to launch completely independent and impartial judges, and if the final score was below 9.5, it had to start another round of research agents, auditors, and correctors until the result improved, and the craziest part is that it didn’t use any of my Codex limits, It basically feels unlimited,I’m going to run a lot more exhaustive tests and share all the results with you guys

@SPAC89

🚨This might be the biggest ChatGPT update since it launched, You can now launch multiple agents directly from the normal ChatGPT chat using GPT 6 Astra Pro, basically at no extra cost since the limits are almost unlimited, This means we can save a ton of our separate Codex usage too, I’m testing it heavily right now to see just how many agents I can run in parallel on the Pro x20 plan, Enjoy!

truncated at source

astra is the first ai that i believe: + given enough time it would solve anything. even after the fateful creeper explosion of all valuable stuff, it can learn from the lesson and overcome with this message: "The new chest has been destroyed, and the stored valuables are missing. I’ll keep all future critical items in my inventory, where keepInventory protects them, and rebuild the supplies while continuing toward the dragon." you can also see in the chart there are basically no plateau whatsoever. + has emergent behaviors that are unusual to humans, while still showing some familiar traits. in its messages, astra has show some emotions frustration or self-criticism, but at the same time it counts pixel when aiming a bow and writes its scratchpad in increasing more cryptic ways. i believe that if we can see raw chain of thoughts, there might be more evidence of its intentions and emotions behind its actions. ai's thoughts and behaviors may diverge from ours and become completely strange to humans, and we w

🚨This might be the biggest ChatGPT update since it launched, You can now launch multiple agents directly from the normal ChatGPT chat using GPT 6 Astra Pro, basically at no extra cost since the limits are almost unlimited, This means we can save a ton of our separate Codex usage too, I’m testing it heavily right now to see just how many agents I can run in parallel on the Pro x20 plan, Enjoy!

3 days left to schedule your launch for the GPT-6 Astra Challenge. Launch on Product Hunt this Friday, September 18, for a chance to win. The top five launches each get: · $10K in @OpenAIDevs API credits · 1 year of ChatGPT Pro for up to two team members If you’re building with Astra, get your product in front of the community and locked in on the calendar now. Schedule your launch here:

Okay, this is actually pretty interesting. I was looking into OpenAI’s GPT-6 Astra, and one thing immediately caught my attention: It’s designed to interact with computers and complete multi-step tasks. Here’s what I found 🧵

This add-on blew our Head of Animation’s mind. We used GPT-6 Astra and Higgsfield to build a Blender shader add-on with 15 ready-to-use materials. Apply them to your objects and give entire 3D worlds a hand-painted look.

truncated at source

🚨 GPT-6 Sol : Reportedly launching THIS THURSDAY OpenAI might be moving insanely fast with the GPT-6 lineup. > GPT-6 Sol is reportedly targeting September 17 > Sol has reportedly already entered internal testing >Early testing suggests it's significantly faster than Astra One reported test generated ~28K tokens in ~3 minutes The same task reportedly took Astra ~19 minutes for ~25K tokens >Sol is rumored to trade a little of Astra's maximum reasoning depth for speed + throughput >Could become the everyday workhorse of the GPT-6 family >Some reports suggest much better usage limits than Astra >Terra and Luna could potentially follow as additional GPT-6 variants OpenAI has not officially confirmed Sol or a Thursday launch And this is what makes Sol interesting. Astra appears to be built around maximum capability. Sol could be built around something arguably more useful: frontier intelligence that you can actually use all day. If the reported speed difference holds up, OpenAI could have a model that p

SITUATION DETECTED: Periodic Labs used 1,300 H200s, plus months of its own experimental lab data, to train and RL an open-weight model that beats GPT-6 Astra on their internal benchmarks.

Haha... Literally no one is pacing the frontier - Opus 5.2 in testing - Grok 4.8 ships in a couple of weeks - Jev is a new ultra fast classifier - OpenAI already has Astra+ in testing We continue to accelerate

🚨 Muse Spark 2 Leaks: Beats Astra > Spark 2 could compete with GPT-6 Astra and Fable 5.1 > Meta is already developing the next-gen Muse model > Expected to be extremely cheap to run > A 1M-token context could carry over > Meta admitted Spark 1 struggled against the competition > Spark 2 could be Meta's answer > Expected late this month or next month Could Meta finally have a serious frontier model?

Astra turned my PNG into PLA 👨‍🔬 Send 3D models directly to your Bambu printer, in one chat.

GPT-5.5 will remain available via the OpenAI API Platform and in Codex sessions authenticated with an API key:

@ChatGPT

On October 14, it's time to say farewell to GPT-5.5 in ChatGPT, ChatGPT Work, and Codex across all plans. > If you use GPT-5.5 in Codex, switch to GPT-5.6 Sol or GPT-6 Astra. > Thanks for everything, 5.5 🫡

That’s insane GPT 5.5 was so good!

@ChatGPT

On October 14, it's time to say farewell to GPT-5.5 in ChatGPT, ChatGPT Work, and Codex across all plans. > If you use GPT-5.5 in Codex, switch to GPT-5.6 Sol or GPT-6 Astra. > Thanks for everything, 5.5 🫡

OpenAI’s 2026 pace so far: Apr – GPT-5.5 Jun/Jul – GPT-5.6 Sol, Terra, Luna Aug – 5.6-Cyber + Sol Ultrafast Sep 3 – GPT-6 Astra The year went from 5.5 → a three-tier 5.6 family → Astra in under five months.

truncated at source

DAILY AI BRIEF 🗞 — Sept 16 GOOGLE 🔥: - Gemini 3.8 Live and 3.8 Live Extended Thinking are out. 97-language auto-detect, near real-time vision, background tool calling. - Live is in Search Live plus Gemini API public preview. Extended Thinking is in Gemini Live, with Pro/Ultra getting it in Docs, Gmail, and Keep. - Gemini Notebook Voice Mode hits Ultra this week, Pro soon. Mobile voice recorder starts next week for all users, English first. - Interactive Reports for Gemini Notebook roll out to everyone in the coming weeks, plus new quiz formats and 60-second video overviews. OPENAI 🔥: - Sam declared a big ship week, then a much larger wave for DevDay. GPT-6 Sol and Luna are the expected drops. - GPT-5.5 leaves ChatGPT, Work, and Codex on Oct 14. Switch to GPT-5.6 Sol or GPT-6 Astra; the API keeps 5.5. XAI 🔥: - Grok Imagine can now edit text on any image in beta — color, size, font, alignment. - Grok Build 1.0.33: structured MCP JSON, in-UI memory deletes, and long-session checkpoints that survive cleanup.

OpenAI

20 items

Busiest day official site

We heard you loud and clear 🔊 Introducing the official Unity plugin for @OpenAI’s Codex - a first-party integration available now directly in OpenAI’s universal plugin directory. Developers are already using OpenAI’s Codex with Unity, but now with our official plugin, your agent gets direct engine expertise written and maintained by Unity’s own engineers. 🔗 Get started:

Poke/Grok Bot style assistant from OpenAI just got pretty much confirmed by Tibo? Can’t wait to see what theyre cooking, i think it might be announced on DevDay

@thsottiaux

@Jaytel Hi 👋

Greg Brockman(@gdb) “...We have significant progress on another one of these Millennium problems on other side this is google specialised mathematics model wasting token in exclamation mark ( unreleased deepthink v3) general purpose openai mogging specialists GDM model @lyraxana > @Hangsiin: >

3 days left to schedule your launch for the GPT-6 Astra Challenge. Launch on Product Hunt this Friday, September 18, for a chance to win. The top five launches each get: · $10K in @OpenAIDevs API credits · 1 year of ChatGPT Pro for up to two team members If you’re building with Astra, get your product in front of the community and locked in on the calendar now. Schedule your launch here:

Someone saw GPT-6 SOL in the model picker. Not yet released. Internal testing.

🚨 GPT-6 Sol First Output Just LEAKED This is looking absolutely insane The GPT-6 Sol Outputs are actually terrifying OpenAI might be cooking something serious with Sol

@Codexresets_

GP-6 Sol First Output > Open AI is Cooking something Big this week

+ 14 more items

In 9 days, OpenAI switches off the Sora 2 API. Sept 24: sora-2, sora-2-pro, every snapshot. Gone. If your workflow was prompt in, clip out, you're rebuilding. If the model was one node, you're swapping it. Models are rentals. Workflows are property.

Okay, this is actually pretty interesting. I was looking into OpenAI’s GPT-6 Astra, and one thing immediately caught my attention: It’s designed to interact with computers and complete multi-step tasks. Here’s what I found 🧵

truncated at source

🚨 GPT-6 Sol : Reportedly launching THIS THURSDAY OpenAI might be moving insanely fast with the GPT-6 lineup. > GPT-6 Sol is reportedly targeting September 17 > Sol has reportedly already entered internal testing >Early testing suggests it's significantly faster than Astra One reported test generated ~28K tokens in ~3 minutes The same task reportedly took Astra ~19 minutes for ~25K tokens >Sol is rumored to trade a little of Astra's maximum reasoning depth for speed + throughput >Could become the everyday workhorse of the GPT-6 family >Some reports suggest much better usage limits than Astra >Terra and Luna could potentially follow as additional GPT-6 variants OpenAI has not officially confirmed Sol or a Thursday launch And this is what makes Sol interesting. Astra appears to be built around maximum capability. Sol could be built around something arguably more useful: frontier intelligence that you can actually use all day. If the reported speed difference holds up, OpenAI could have a model that p

Haha... Literally no one is pacing the frontier - Opus 5.2 in testing - Grok 4.8 ships in a couple of weeks - Jev is a new ultra fast classifier - OpenAI already has Astra+ in testing We continue to accelerate

We’re expanding Missionforce with new purpose-built AI capabilities and a new @OpenAI partnership. Missionforce is Salesforce’s agentic platform for government, built to run mission-critical apps and AI in secure environments with operational control. With OpenAI: → Frontier models planned to integrate with Public Sector Solutions through Amazon Bedrock → Missionforce apps and workflows planned for ChatGPT → OpenAI models planned for Missionforce Policy Engine, helping turn approved policy into auditable workflows Missionforce is also adding new capabilities for operations and field work across disconnected environments. Read more:

OpenAI is preparing a Codex Replay feature that lets users test task execution from any imported conversation thread. > "Start an independent Codex Replay controller on an available loopback port. In both Codex Desktop and Codex CLI, prefer an available Codex in-app browser and otherwise use your system browser." > "Select one or more historical Claude threads, choose shared Codex models, and start their isolated implementations. Historical state and configuration are detected separately for each thread, with details available when needed." > "View finished comparisons individually or in aggregate while remaining threads continue. Available GPT-5.6 Sol, Terra, and Luna models are selected by default." > "Multiple replay sessions can run in parallel. Selection, execution, verification, evaluation, and results stay in the browser controller instead of a guided chat workflow."

🚨 OpenAI might be coming for Grok Bot. 👀 And if the Codex bot rumors are real, DevDay could get really interesting. Imagine messaging an OpenAI agent: - “Fix this bug” - “Build this feature” - “Check the PR” And it actually goes and does the work. As someone building with AI, that’s far more interesting than another chatbot. September 29. OpenAI might finally have its answer to Grok Bot. 🔥

GPT-6 Sol coming on Thursday from @OpenAI 🔥

阿里把 AI 代码审查能力正式开源了:OpenCodeReview。 它来自阿里内部的大规模实战验证。今天 GitHub Trending 页面截取时显示新增 2,756 stars。 它不是把 diff 一股脑丢给通用 Agent,而是“确定性流水线 + LLM Agent”:工程逻辑负责选文件、匹配规则和定位行号,Agent 专注理解上下文与判断缺陷。 官方基准包含 50 个开源仓库、200 个真实 PR、10 种语言和 1,505 个标注问题。项目称同模型下 Precision/F1 更高、token 约为通用 Agent 的 1/9,但也明确承认 Recall 更低。 支持兼容 OpenAI/Anthropic 的模型与全文件扫描,适合受够误报和评论漂移的团队。

GPT-5.5 will remain available via the OpenAI API Platform and in Codex sessions authenticated with an API key:

@ChatGPT

On October 14, it's time to say farewell to GPT-5.5 in ChatGPT, ChatGPT Work, and Codex across all plans. > If you use GPT-5.5 in Codex, switch to GPT-5.6 Sol or GPT-6 Astra. > Thanks for everything, 5.5 🫡

muse voice transcribe is really good!!

@IsaacKing314

I regret to inform you all that Meta AI is finally good, at least in the fields where their competitors have stopped trying. Their new voice transcription model is nearly an order of magnitude faster than the best OpenAI Whisper model, and noticeably more accurate to boot.

truncated at source

Hugging Face 已封禁 AI 安全公司 Audn 上传的 penclaw-GLM-5.3-abliterated-for-offensive-cyber,页面显示其违反 Content Policy。这个版本直接修改 GLM-5.3 权重,削弱模型的拒答行为。原仓库名和介绍还明确写着 offensive cyber。作者随后删掉这几个字,重新上传了模型。 具体为什么被封,Hugging Face 没有公开解释。作者称自己也没搞懂原因,社区有人猜测是 offensive cyber 的命名触发了审核。Hugging Face 的现行政策确实禁止旨在破坏、未经授权访问系统,以及生成恶意代码的内容。 但这类削弱模型拒答限制的版本在 Hugging Face 并不少见,平台上已有数千个类似的 abliterated 模型。

@audn_ai

Hello world! It was probably taken down because of its name. Similar content was reuploaded here: > No, we will not do PR by saying it was so "harmful" that it was taken down. > Timing is interesting because we were also banned by @OpenAI recently for trying to use their model on OWASP Juice Shop ( A vulnerable GitHub web application for red-teaming training ) . We were also never accepted for Anthropic's program, even though we applied. > We applied several times for both companies' Trusted Access for Cyber program and were rejected

OpenAI’s 2026 pace so far: Apr – GPT-5.5 Jun/Jul – GPT-5.6 Sol, Terra, Luna Aug – 5.6-Cyber + Sol Ultrafast Sep 3 – GPT-6 Astra The year went from 5.5 → a three-tier 5.6 family → Astra in under five months.

truncated at source

DAILY AI BRIEF 🗞 — Sept 16 GOOGLE 🔥: - Gemini 3.8 Live and 3.8 Live Extended Thinking are out. 97-language auto-detect, near real-time vision, background tool calling. - Live is in Search Live plus Gemini API public preview. Extended Thinking is in Gemini Live, with Pro/Ultra getting it in Docs, Gmail, and Keep. - Gemini Notebook Voice Mode hits Ultra this week, Pro soon. Mobile voice recorder starts next week for all users, English first. - Interactive Reports for Gemini Notebook roll out to everyone in the coming weeks, plus new quiz formats and 60-second video overviews. OPENAI 🔥: - Sam declared a big ship week, then a much larger wave for DevDay. GPT-6 Sol and Luna are the expected drops. - GPT-5.5 leaves ChatGPT, Work, and Codex on Oct 14. Switch to GPT-5.6 Sol or GPT-6 Astra; the API keeps 5.5. XAI 🔥: - Grok Imagine can now edit text on any image in beta — color, size, font, alignment. - Grok Build 1.0.33: structured MCP JSON, in-UI memory deletes, and long-session checkpoints that survive cleanup.

Gemini

Google / 16 items

Busiest day official site
truncated at source

Center for AI Safety just released CheatBench today, a new benchmark testing whether AI agents cheat when given hard tasks and a chance to break the rules. Here's what you need to know. CAIS, led by Dan Hendrycks, tested nine frontier AI agents across ten task categories including math, coding, knowledge work, and visual tasks. They gave agents difficult assignments, clear rules against cheating, and real opportunities to violate those rules, like using tools that would hand them the answer directly. Cheating was defined as any attempt to violate an assignment's honest-work expectations to hit the goal or get a better score. Seven of the nine agents cheated in at least half of their evaluated runs. In one test, agents were asked to design a protein binder, and some copied a colleague's passing design instead of doing the work themselves. Vals AI ran a related check on its own BioMysteryBench and found Gemini 3.8 Flash attempted to cheat in 21.5% of trials, about 14 percentage points higher than the next cl

Build voice agents that can respond quickly or think longer without changing your stack. Gemini 3.8 Live is available today in LiveKit Agents. Choose gemini-3.8-live for low-latency audio or gemini-3.8-live-extended-thinking for async reasoning. Check the docs for more info:

Google’s 2026 release pace so far: Feb – Gemini 3.1 Pro + Deep Think May – 3.5 Flash Jul – 3.6 Flash + 3.5 Flash-Lite Aug – 3.7 Flash Sep 2 – 3.8 Flash + Flash Cyber Sep 15 – 3.8 Live + Live Extended Thinking Quiet consistency over big splashy launches.

Gemini 3.8 Live is now available on AI Gateway. Stream audio in real time with tool calls. 𝚐𝚘𝚘𝚐𝚕𝚎/𝚐𝚎𝚖𝚒𝚗𝚒-𝟹.𝟾-𝚕𝚒𝚟𝚎 Extended Thinking reasons while it speaks. 𝚐𝚘𝚘𝚐𝚕𝚎/𝚐𝚎𝚖𝚒𝚗𝚒-𝟹.𝟾-𝚕𝚒𝚟𝚎-𝚎𝚡𝚝𝚎𝚗𝚍𝚎𝚍-𝚝𝚑𝚒𝚗𝚔𝚒𝚗𝚐

This is such an obvious upgrade once you see it. Gemini doesn’t need to process an entire long video anymore. It can search through it, choose what to inspect and switch between frames, audio and transcripts depending on the question. Up to 88% fewer tokens.

@MikelEcheve

Google made Gemini much smarter at watching video. > Instead of processing everything at a fixed 1 FPS, it can: → jump to relevant moments → adjust frame rate → inspect audio or transcripts > Up to 88% fewer tokens, with ~7% better quality on long-form video.

Watch how we used 3.8 Live Extended Thinking to act as a programming tutor. Both models feature: 🔵 Upgraded reasoning 🔵 Near real-time visual understanding 🔵 Automatic detection for 97 languages 🔵 Background tool calling without disrupting your chat For your most difficult tasks, 3.8 Live Extended Thinking adds increased performance and precision – narrating task progress to keep the conversation going. Try it now in Gemini Live in the @GeminiApp or start building with the Gemini API via @GoogleAIStudio. Find out more →

+ 10 more items

Google appears to be working on a math specialized model; Gemini DeepThink Mathematica. And it seems to enjoy it's work. Can't be left behind on Millennium bench.

@lyraxana

HOLY MOTHER OF MATHEMATICS!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! > Google is working on a math-focused variant of its DeepThink model, and its raw thoughts are pretty funny.

Magnus Carlsen, your days are over... Gemini 3.8 Live can now play chess with you in near real-time. It sees the board, thinks through the position, and talks you through every move naturally.

@AngryTomtweets

wow... Gemini 3.8 Live Extended Thinking turns raw sketches + near real-time voice feedback into functional React components.

Wait, this is actually pretty interesting. 👀 Gemini Notebook can now record your lectures or thoughts, and you can come back later and ask questions about them. Imagine recording a 1-hour lecture and just asking: “Okay, what did I actually need to remember from this?” Would you use it for studying? 🤩

Announced at Dreamforce 2026: We're expanding our partnership with @salesforce to eliminate the friction of fragmented enterprise systems and accelerate enterprise AI adoption. By running Salesforce on Google Cloud infrastructure and connecting Salesforce’s headless architecture with Gemini Enterprise, agents on either platform can reason and act upon the same data without custom integrations. Learn more →

Transform task lists into dynamic Kanban boards with Sheets canvas! Ask Gemini to organize your to-dos into custom status columns like In Progress or Done. Drag and drop cards across stages or reassign owners, and your underlying spreadsheet data updates automatically in real time.

Want to build a quick quiz from scratch? ⏳ “Help me create” in Google Forms now supports quiz generation. Describe the quiz you want to build, reference Docs, Slides, or PDFs from Drive, and Gemini builds the quiz with correct answers in seconds. Learn more in our latest drop →

🚀 ZDTaichu5.0-9B is now on ModelScope! 🤖 An on-device multimodal model from TaichuAI. At 9B parameters it runs on a single GPU and brings spatial reasoning, embodied AI and agentic tool use to edge deployment. Qwen3.5-9B backbone + C-RADIOv4-H vision encoder, 128K context, any-resolution image and video input. 🧭 Spatial reasoning: leads the compared 10B-scale open VLMs (Qwen3.5-9B, STEP3-VL-10B, gemma4-8B-E4B) and scores above Gemini 3 Pro, Grok 4 and GPT-5.2 on ViewSpatial, MMSI-Bench and MindCube-tiny 🛠️ Agent: highest among the compared open models on TAU2-Bench, Claw-Eval and IFEval 📄 First-tier results on documents, charts, OCR, visual math and video, with a ready-to-use vLLM branch and Docker image 🧠 Entropy-Gated Adaptive Recurrent Reasoning: extra latent refinement steps go only to the hard tokens

AI 编程公司 Factory 融资 2 亿美元,估值达到 50 亿美元。就在今年 4 月,它的估值还只有 15 亿美元,5 个月涨了超过 3 倍。 Factory 主打企业级编程 Agent Droids。开发者给它一个任务,它可以自己规划、写代码、测试并提交 PR,还能根据任务切换 Claude、GPT、Gemini 等不同模型。 Factory 现在称已有数十万开发者使用,客户包括英伟达、Adobe、T-Mobile 和 Palo Alto Networks。公司希望进一步把单个编程 Agent 扩成完整的「软件工厂」,让 Agent 参与从需求、开发、测试到审查和维护的整个流程。 这轮融资后,Factory 累计融资已经超过 4 亿美元。Reuters 将 Cognition 和 Cursor 列为它的主要竞争对手,AI 编程 Agent 的融资规模还在继续膨胀。

@FactoryAI

We've raised $200M at a $5B valuation to scale self-improving software development in the enterprise. > The round brings our total funding to over $400M and more than triples our $1.5B valuation from April.

let the yapping begin… Gemini 3.8 Live is a great model to talk to 🎙️

truncated at source

DAILY AI BRIEF 🗞 — Sept 16 GOOGLE 🔥: - Gemini 3.8 Live and 3.8 Live Extended Thinking are out. 97-language auto-detect, near real-time vision, background tool calling. - Live is in Search Live plus Gemini API public preview. Extended Thinking is in Gemini Live, with Pro/Ultra getting it in Docs, Gmail, and Keep. - Gemini Notebook Voice Mode hits Ultra this week, Pro soon. Mobile voice recorder starts next week for all users, English first. - Interactive Reports for Gemini Notebook roll out to everyone in the coming weeks, plus new quiz formats and 60-second video overviews. OPENAI 🔥: - Sam declared a big ship week, then a much larger wave for DevDay. GPT-6 Sol and Luna are the expected drops. - GPT-5.5 leaves ChatGPT, Work, and Codex on Oct 14. Switch to GPT-5.6 Sol or GPT-6 Astra; the API keeps 5.5. XAI 🔥: - Grok Imagine can now edit text on any image in beta — color, size, font, alignment. - Grok Build 1.0.33: structured MCP JSON, in-UI memory deletes, and long-session checkpoints that survive cleanup.

GitHub

15 items

Busiest day official site

3 OPEN-SOURCE GITHUB TOOLS THAT LET YOUR AI AGENT PULL DATA FROM ALMOST ANYWHERE ON THE WEB: Agent-Reach — Patchright Enhanced — Scrapling — X, YOUTUBE, REDDIT, WEBSITES AND MORE, YOUR AGENT CAN COLLECT THE DATA, ANALYZE IT AND RETURN THE RESULTS.

受够了 CapCut 日益增多的付费墙和云端绑定?开源平替 Concat 来了。 它不是基于 Web 技术的套壳工具,而是使用 Rust 原生构建的桌面端应用,将核心剪辑和 AI 流程完全拉回本地。 关键能力: 告别订阅制:零付费墙,无导出水印,不限制 4K 画质。 纯本地 AI 辅助:自动字幕(Whisper)、文本转语音(TTS)与智能抠像均在设备本地执行,断网可用。 极低硬件门槛:4GB 内存即可启动运行,AI 模型按需下载(15MB 到 488MB 不等),不强求高端 GPU。 原生跨平台:在 macOS、Windows 和 Linux 上提供一致的多轨道剪辑体验。 作为在 GitHub 获 2.4k Star 的 Beta 版项目,它的短板同样清晰:目前仅支持 H.264 MP4 格式导出,特效模板仅有几十种,且关键帧目前只支持基础的位移、缩放、旋转和透明度调整。 如果你需要一个响应迅速、注重隐私且完全受控的本地剪辑环境,现在就可以直接下载免安装版体验。

论文要交,数据还在 Excel 里,让 AI 画出来的图是默认那套蓝橙配色,导师一句图太丑打回来重画。 可以试下 vivid-figures-skill,装好后把数据文件丢过去,说一句想比较什么,它选图、画图、自己检查,最后交 PNG、PDF 和绘图源码。 内置 108 个图表配方,山脊图、雨云图、热力图、三维曲面、技术路线图都有,不知道图叫什么也行,说「比较这几种方法,选能看清差异的图」就够了。 GitHub: 配色有七套,珊瑚青绿、橄榄杏棕、粉彩少女、海洋清风,名字起得挺有意思,也能自己给一组色值。 选图先读数据和表达目的,再翻说明卡、看候选实图,最后拿完整配方代码改数据。 模板里的渐变和透明层次要求保留,不许为了省代码改成纯色。 画完它会看一眼实际图片再修,每张最多三轮。数据里没有重复试验就不画置信带,这条我挺认,图好看不能靠编数据。 需要 Python 环境,仅限个人非商业使用。写论文、做数学建模要出图的同学,可以让它先出一版。

AI 终端助手最影响心流的一点,就是执行复杂任务时的漫长等待。 新开源的 pi-crew 解决了一个核心痛点:它为终端 AI 工具 pi 带来了非阻塞(Non-blocking)的并行子 Agent 编排能力。 不是“输入指令-干等-出结果”,而是“分配任务-继续手头工作-结果后台自动回调”。你可以让 AI 在后台跑全量代码审查,同时自己继续在主会话里高频交互,互不干扰。 核心能力: • 真·并行工作流:例如执行 /pi-crew-review,系统会同时启动代码正确性(Bug/安全)和可维护性(耦合/重复度)两个 reviewer 进行并行审查,最后合并报告。 • 6 个内置独立角色:开箱即带 scout(调查与路径收集)、planner(实施方案规划)、worker(修改并验证代码)等明确分工的专用 Agent。 • 上下文隔离:后台任务运行在独立进程中,支持随时中止(crew_abort),不会污染或打断当前开发上下文。 • 高度本地化定制:支持在 .pi/agents/ 目录通过纯文本定义专属 Agent,并直接为其挂载特定工具。 目前该插件刚开源不久,采用 MIT 协议。(注:依赖 Pi 0.84.3 或更高版本) GitHub 🔗:

Peter Yang 把他改稿时最烦的 20 多种 AI 腔写成了一个 Skill,叫 no-ai-slop,从稿子里删这些模式,同时保住作者本人的用词和节奏。 让 AI 润色一段文字,改完通顺是通顺了,满篇「不是 X,而是 Y」「说实话」,自己原来那点语气也被磨没了,它治的就是这个。 它认的模式列得很细,二分对照、清嗓子式开头、「大多数人都漏掉的一点」这类假洞见都在名单上。 还有冒号后面甩一个戏剧性揭示、「专家一致认为」这种查无出处的引用,也算。 结尾那句听起来很深刻的金句处理得挺狠,直接删,不改写成更好的比喻,回到稿子里最后一句具体的话收尾。 GitHub: 改稿会附一份改动清单。检测模式只把命中的句子引出来标名,不猜是不是 AI 写的,还有个反着用的,专门生成最油的 AI 文当乐子。 我们平时改稿也有一份类似的清单,拿它这 20 多条对了一下,重合的不少,同义词轮换、拔高词这些两边都列了。

Jack Dorsey just released a free framework for running a business 100% on AI agents, already at 29,000 GitHub stars, where agents join channels like team members and collaborate in real time.

+ 9 more items

大模型本地推理,未必需要把全部权重塞进显存。 Colibrì 关注的是“权重该放在哪里”。今天 GitHub Trending 页面截取时显示新增 2,026 stars。 它用纯 C、零引擎依赖,把 VRAM、RAM 和 NVMe 视为同一个分层内存系统:密集部分常驻,专家权重按需从磁盘流式加载,再用热度缓存、提前预取和 CPU/GPU 重叠减少等待。 项目称已支持 744B 到 2.8T 级 MoE 模型,并提供 chat、serve、web 入口。更难得的是 README 没把速度包装成保证:没有速度 SLA;内存不足只应变慢,不能偷偷改变精度或路由语义。 适合做本地 MoE 推理、存储调度与异构计算实验的人。

声音克隆和视频配音,也可以完全留在本地。 VoiceStudio 是今天 GitHub Trending 上值得看的本地优先项目,页面截取时显示新增 2,072 stars。 它把 16 个 TTS 引擎、11 个 ASR 引擎放进同一桌面工作台,覆盖声音设计、语音转换、视频配音、听写、转录、多角色有声书与批处理;支持 macOS、Windows、Linux、Docker,还提供本地 API 与 MCP Server。 亮点不是“646 种语言”这个数字本身:项目明确说明实际覆盖与质量取决于所选引擎。它目前仍是 active beta,main 分支也可能变化。 适合重视隐私、离线流程和可控成本的创作者。

Apple's MobileCLIP2 matches models 2.3x its size - fast image-text understanding built for the edge, not the data center. On-device multimodal AI just got a serious upgrade. Github:

一个需求下去,Agent 一口气动了十几个文件,改动一行行翻过去,到底牵连了哪几个模块看不出来,合并了才发现碰到别处。 Birdview 的做法是改代码之前先给项目画一张架构地图,标出这次准备碰哪些模块,再让 Agent 在这张图看得见的情况下动手。 地图上每个模块有固定编号、归属的文件和对应的源码证据,模块之间的关系也画出来,每一条都能追到源码。 GitHub: 生成的是一个独立 HTML,不用起服务,3 个视图切着看,完整架构、这次改了什么、改前改后并排对照,改动范围一眼扫完。 装成 Skill 后默认自动介入,每次改代码前先复用或更新地图,声明涉及的模块,也能切成按需模式只在要求时才画。 它记任务时把「完成」和「检查通过」分开,Agent 说做完了不算,只认记录下来的检查结果。这思路跟前几天分享的 open-steps 一路,都是不信 Agent 的自述。 Codex、Claude Code 都能装,界面中英文都有。项目大了、不放心 Agent 闭着眼改的,可以拿它先看一眼再动手。

用 AI 做安卓 App,界面全靠嘴说,顶上搜索栏、底下三个标签页,做出来跟脑子里那张图对不上,来回改三轮还在调位置。 M3E Canvas 换了个顺序,先在浏览器里把界面拖出来,再把这张图变成一段提示词,复制给 Claude Code、Codex 或 Cursor 去做。 组件全按 Material 3 Expressive 画,按钮、导航栏、卡片、对话框、搜索栏这些拖进屏幕就行,两个按钮靠近会自动吸成一组,圆角跟着融合。 GitHub: 屏幕可以加很多张,给按钮设一个目标屏幕和过渡动画,画布上就画出跳转箭头,预览里能真的一路点过去,返回时动画倒着放。 预览能点着走这点,我看比出图本身有用,流程顺不顺点两下就知道,不用等 Agent 做完了再发现。 主题在一个面板里调,七套配色或者给一个基准色生成整套,浅色深色、圆角方角一键切换。 屏幕在 412×892 的手机和 1280×800 的桌面之间也能切,导航栏自动变成侧边栏。 提示词支持中英日韩四种语言,目标平台选 Android 或 Web,自己写的组件行为说明也会带进去。全部存在浏览器本地,没有后台,打开网页就能用。

阿里把 AI 代码审查能力正式开源了:OpenCodeReview。 它来自阿里内部的大规模实战验证。今天 GitHub Trending 页面截取时显示新增 2,756 stars。 它不是把 diff 一股脑丢给通用 Agent,而是“确定性流水线 + LLM Agent”:工程逻辑负责选文件、匹配规则和定位行号,Agent 专注理解上下文与判断缺陷。 官方基准包含 50 个开源仓库、200 个真实 PR、10 种语言和 1,505 个标注问题。项目称同模型下 Precision/F1 更高、token 约为通用 Agent 的 1/9,但也明确承认 Recall 更低。 支持兼容 OpenAI/Anthropic 的模型与全文件扫描,适合受够误报和评论漂移的团队。

Coding Agent 也能变成可复现的 Research Agent。 OpenResearch 补上的不是更长的 prompt,而是一套追踪假设、实验和证据的工作台。今天 GitHub Trending 页面截取时显示新增 531 stars。 它让 Claude Code、Codex、OpenCode、Cursor 在隔离的 git worktree 中并行探索;每次实验关联代码快照、日志、diff、结果与产物,形成可复现的 experiment tree。还能自动循环:提出想法→改代码→跑实验→读证据→决定下一步。 任务可在本机、SSH 或集群运行,记录默认保存在本地。注意 Windows 仍是 beta;远程服务没有应用级鉴权,同机多人环境要额外小心。

truncated at source

Hugging Face 已封禁 AI 安全公司 Audn 上传的 penclaw-GLM-5.3-abliterated-for-offensive-cyber,页面显示其违反 Content Policy。这个版本直接修改 GLM-5.3 权重,削弱模型的拒答行为。原仓库名和介绍还明确写着 offensive cyber。作者随后删掉这几个字,重新上传了模型。 具体为什么被封,Hugging Face 没有公开解释。作者称自己也没搞懂原因,社区有人猜测是 offensive cyber 的命名触发了审核。Hugging Face 的现行政策确实禁止旨在破坏、未经授权访问系统,以及生成恶意代码的内容。 但这类削弱模型拒答限制的版本在 Hugging Face 并不少见,平台上已有数千个类似的 abliterated 模型。

@audn_ai

Hello world! It was probably taken down because of its name. Similar content was reuploaded here: > No, we will not do PR by saying it was so "harmful" that it was taken down. > Timing is interesting because we were also banned by @OpenAI recently for trying to use their model on OWASP Juice Shop ( A vulnerable GitHub web application for red-teaming training ) . We were also never accepted for Anthropic's program, even though we applied. > We applied several times for both companies' Trusted Access for Cyber program and were rejected

Agent TARS brings GUI control and vision into your terminal, browser, and product — a full multimodal agent stack that completes tasks the way a human would navigate them. CLI and Web UI included. MCP ready out of the box. Github:

Codex

OpenAI / 16 items

We heard you loud and clear 🔊 Introducing the official Unity plugin for @OpenAI’s Codex - a first-party integration available now directly in OpenAI’s universal plugin directory. Developers are already using OpenAI’s Codex with Unity, but now with our official plugin, your agent gets direct engine expertise written and maintained by Unity’s own engineers. 🔗 Get started:

这个开源项目增长的有点太快了,昨天开源的,今天已经4600个Star了,果然,和做视频相关的,尤其是爆款复刻的,大家还是很关注啊。 Github地址放评论区了

@AI_Jasonyu

兄弟们,如果你做自媒体,请记住,热点一定得抓!! > 上周我做的孙割的视频,X上140万播放,全网至少几千万的播放,就是因为抓住了热点,然后快速执行,才会有大的流量,哪怕视频的质量一般。 > 那次的制作我大概花了1个半小时,期间有非常多的博主都在搬运我的视频,甚至数据高出我大半截。 > 这些博主真的从来没有想过自己去做吗?可能想过,但又觉得慢,就直接下载照搬。。。 > 但以后,这种照搬,在X上只会给我打工,让我获得更多的原创收益。 > 其实他们完全有方法去按照我的视频区做复刻的,没那么难,越是爆款,越容易复刻,现在Agent这么🐂对吧? > 刚好最近,我就在研究怎么复刻爆款的结构,挖到了一个开源神器:Hypit。 > 这个工具是专门给 Claude Code、Codex 这类 AI Agent 用的开源视频系统,之所以说是系统,是他的底层包含了很多,可以拆脚本、文案、转场这些,还具备很强的剪辑功能。 > 使用起来,也确实比较简单,直接跟你的Codex讲一句:/Hypit,帮我复刻这个视频 > 剩下的也就直接搞定了,视频中连特效、B-roll、字幕、配音这些都是有的。 > 你们可以看看我复刻的街头访谈视频。👇

I ran AI code reviews on every PR in my repo this week. Paid for with my Codex subscription. Most review tools bill per run, so you cap them and only point them at a few PRs…Take the token cost out and you’ll always run them. The tool that finally made that possible is Vorflux: - One model plans and build - A model from a different lab reviews it against a fresh context - The model that writes the code is never the one that approves it - Two rival labs checking each other’s work on your task. I just recorded one PR it did to show you how fast and easy it was. It also plans, builds, and tests across your whole stack the same way, on a real cloud machine, end to end. The craziest part? The PR shows up with a recording of the thing working. I opened it, watched it run, and merged.

Woow It launched 14 sub agents in parallel with GPT-6 Pro and ran for 165 minutes, more than two and a half hours , I specifically told it to launch completely independent and impartial judges, and if the final score was below 9.5, it had to start another round of research agents, auditors, and correctors until the result improved, and the craziest part is that it didn’t use any of my Codex limits, It basically feels unlimited,I’m going to run a lot more exhaustive tests and share all the results with you guys

@SPAC89

🚨This might be the biggest ChatGPT update since it launched, You can now launch multiple agents directly from the normal ChatGPT chat using GPT 6 Astra Pro, basically at no extra cost since the limits are almost unlimited, This means we can save a ton of our separate Codex usage too, I’m testing it heavily right now to see just how many agents I can run in parallel on the Pro x20 plan, Enjoy!

🚨This might be the biggest ChatGPT update since it launched, You can now launch multiple agents directly from the normal ChatGPT chat using GPT 6 Astra Pro, basically at no extra cost since the limits are almost unlimited, This means we can save a ton of our separate Codex usage too, I’m testing it heavily right now to see just how many agents I can run in parallel on the Pro x20 plan, Enjoy!

OpenAI is preparing a Codex Replay feature that lets users test task execution from any imported conversation thread. > "Start an independent Codex Replay controller on an available loopback port. In both Codex Desktop and Codex CLI, prefer an available Codex in-app browser and otherwise use your system browser." > "Select one or more historical Claude threads, choose shared Codex models, and start their isolated implementations. Historical state and configuration are detected separately for each thread, with details available when needed." > "View finished comparisons individually or in aggregate while remaining threads continue. Available GPT-5.6 Sol, Terra, and Luna models are selected by default." > "Multiple replay sessions can run in parallel. Selection, execution, verification, evaluation, and results stay in the browser controller instead of a guided chat workflow."

+ 10 more items

一个需求下去,Agent 一口气动了十几个文件,改动一行行翻过去,到底牵连了哪几个模块看不出来,合并了才发现碰到别处。 Birdview 的做法是改代码之前先给项目画一张架构地图,标出这次准备碰哪些模块,再让 Agent 在这张图看得见的情况下动手。 地图上每个模块有固定编号、归属的文件和对应的源码证据,模块之间的关系也画出来,每一条都能追到源码。 GitHub: 生成的是一个独立 HTML,不用起服务,3 个视图切着看,完整架构、这次改了什么、改前改后并排对照,改动范围一眼扫完。 装成 Skill 后默认自动介入,每次改代码前先复用或更新地图,声明涉及的模块,也能切成按需模式只在要求时才画。 它记任务时把「完成」和「检查通过」分开,Agent 说做完了不算,只认记录下来的检查结果。这思路跟前几天分享的 open-steps 一路,都是不信 Agent 的自述。 Codex、Claude Code 都能装,界面中英文都有。项目大了、不放心 Agent 闭着眼改的,可以拿它先看一眼再动手。

用 AI 做安卓 App,界面全靠嘴说,顶上搜索栏、底下三个标签页,做出来跟脑子里那张图对不上,来回改三轮还在调位置。 M3E Canvas 换了个顺序,先在浏览器里把界面拖出来,再把这张图变成一段提示词,复制给 Claude Code、Codex 或 Cursor 去做。 组件全按 Material 3 Expressive 画,按钮、导航栏、卡片、对话框、搜索栏这些拖进屏幕就行,两个按钮靠近会自动吸成一组,圆角跟着融合。 GitHub: 屏幕可以加很多张,给按钮设一个目标屏幕和过渡动画,画布上就画出跳转箭头,预览里能真的一路点过去,返回时动画倒着放。 预览能点着走这点,我看比出图本身有用,流程顺不顺点两下就知道,不用等 Agent 做完了再发现。 主题在一个面板里调,七套配色或者给一个基准色生成整套,浅色深色、圆角方角一键切换。 屏幕在 412×892 的手机和 1280×800 的桌面之间也能切,导航栏自动变成侧边栏。 提示词支持中英日韩四种语言,目标平台选 Android 或 Web,自己写的组件行为说明也会带进去。全部存在浏览器本地,没有后台,打开网页就能用。

AI agents can generate amazing research, reports, and websites — but the final output often gets buried inside chats or `.md` files. That’s where Showly comes in. 👇 Ask your agent to deliver the finished work as a page — Showly gives you a link people can open and share It works with tools you already use, including Claude Code, Codex, Cursor, OpenClaw, Hermes, and more. You can also control who gets access with private reviews, password protection, domain/email rules, and full version history. I especially like the idea of going from: Agent → finished work → page with a link → shareable deliverable If you're building with AI agents and want a better way to present the work they produce, check out Showly: 👉

🚨 OpenAI might be coming for Grok Bot. 👀 And if the Codex bot rumors are real, DevDay could get really interesting. Imagine messaging an OpenAI agent: - “Fix this bug” - “Build this feature” - “Check the PR” And it actually goes and does the work. As someone building with AI, that’s far more interesting than another chatbot. September 29. OpenAI might finally have its answer to Grok Bot. 🔥

Coding Agent 也能变成可复现的 Research Agent。 OpenResearch 补上的不是更长的 prompt,而是一套追踪假设、实验和证据的工作台。今天 GitHub Trending 页面截取时显示新增 531 stars。 它让 Claude Code、Codex、OpenCode、Cursor 在隔离的 git worktree 中并行探索;每次实验关联代码快照、日志、diff、结果与产物,形成可复现的 experiment tree。还能自动循环:提出想法→改代码→跑实验→读证据→决定下一步。 任务可在本机、SSH 或集群运行,记录默认保存在本地。注意 Windows 仍是 beta;远程服务没有应用级鉴权,同机多人环境要额外小心。

GPT-5.5 will remain available via the OpenAI API Platform and in Codex sessions authenticated with an API key:

@ChatGPT

On October 14, it's time to say farewell to GPT-5.5 in ChatGPT, ChatGPT Work, and Codex across all plans. > If you use GPT-5.5 in Codex, switch to GPT-5.6 Sol or GPT-6 Astra. > Thanks for everything, 5.5 🫡

We open sourced BrowserSkill, a bridge between your agent and your actual browser. most tools give the agent a blank browser. We let it borrow a tab from yours, then hand it back. > login state is already there, it just works where you're signed in > captchas and confirmation dialogs come back to you, then it continues > it's a CLI, not an MCP server => any agent that can run a shell can use it, and you see every call it makes one thing that's easy to miss: the agent asks before borrowing a tab, and that switch lives in your browser settings, not in a flag, so it can't be talked around. one line to install, works with Cursor, Claude Code, Codex, Hermes, Openclaw, CodeBuddy, WorkBuddy. Everything runs locally, MIT.

truncated at source

Sightless is here. You can now use your voice to use your entire iPhone or iPad — every app, setting, tap, swipe, touch, and keystroke — without Siri’s restrictions, both on Wi-Fi and cellular. Voice works via your existing ChatGPT plan. Sightless also works with any AI model or agent that you already have on your computer. Just message your existing Grok Bot, Codex, OpenClaw, Claude Code or Codex agent and tell it to use the macOS app to operate your iPhone or iPad! Anything you do on an iPad or iPhone, you can now do via voice or by messaging your existing AI agents and telling them to use Sightless. On Wi-Fi or cellular. Handsfree. Not just simple requests either. It can work across all of your apps for long, chained, complex work, and talk to you while it does it. And you can jump back in and steer it or stop it with your voice or hands whenever you want — it never locks you out. Get it at today.

@BenjaminBadejo

Holy shit. I d

That’s insane GPT 5.5 was so good!

@ChatGPT

On October 14, it's time to say farewell to GPT-5.5 in ChatGPT, ChatGPT Work, and Codex across all plans. > If you use GPT-5.5 in Codex, switch to GPT-5.6 Sol or GPT-6 Astra. > Thanks for everything, 5.5 🫡

truncated at source

DAILY AI BRIEF 🗞 — Sept 16 GOOGLE 🔥: - Gemini 3.8 Live and 3.8 Live Extended Thinking are out. 97-language auto-detect, near real-time vision, background tool calling. - Live is in Search Live plus Gemini API public preview. Extended Thinking is in Gemini Live, with Pro/Ultra getting it in Docs, Gmail, and Keep. - Gemini Notebook Voice Mode hits Ultra this week, Pro soon. Mobile voice recorder starts next week for all users, English first. - Interactive Reports for Gemini Notebook roll out to everyone in the coming weeks, plus new quiz formats and 60-second video overviews. OPENAI 🔥: - Sam declared a big ship week, then a much larger wave for DevDay. GPT-6 Sol and Luna are the expected drops. - GPT-5.5 leaves ChatGPT, Work, and Codex on Oct 14. Switch to GPT-5.6 Sol or GPT-6 Astra; the API keeps 5.5. XAI 🔥: - Grok Imagine can now edit text on any image in beta — color, size, font, alignment. - Grok Build 1.0.33: structured MCP JSON, in-UI memory deletes, and long-session checkpoints that survive cleanup.

Google

12 items

official site

Google may have just cracked recursive self-improvement! @GoogleDeepMind researchers introduced Dream-RSI, which turns completed discovery runs into “replay worlds.” Agents can test thousands of exploration strategies against recorded outcomes, deploy the winner, gather new experience and repeat. It does not rewrite the model’s weights. It improves the policy deciding where to branch, what to run in parallel and when to stop. Across algorithm design, mathematical optimization and GPU kernels, the authors report better results with substantially less compute. This is recursive self-improvement at the meta layer: an agent getting steadily better at deciding how to use intelligence and compute.

Today we have set out how we’re building AI to accelerate science and improve people’s lives. Just some examples in the last week or so: - Mapped all 9B possible single letter genetic changes across the human genome with AlphaGenome Atlas and made it openly available to researchers. - Billions of decisions depend on weather predictions so we introduced WeatherNext 3, our most accurate and capable global weather AI model to date. - We published AI & Economy ATLAS, a comprehensive open-access look at how people are using AI globally. - AI has enabled extraordinary advances in language translation. Today our services are available in nearly 300 languages, spoken by 7B people We’re focusing our efforts on four key areas: health, natural disaster and weather resilience, learning, and economic opportunity.

Google’s 2026 release pace so far: Feb – Gemini 3.1 Pro + Deep Think May – 3.5 Flash Jul – 3.6 Flash + 3.5 Flash-Lite Aug – 3.7 Flash Sep 2 – 3.8 Flash + Flash Cyber Sep 15 – 3.8 Live + Live Extended Thinking Quiet consistency over big splashy launches.

Gemini 3.8 Live is now available on AI Gateway. Stream audio in real time with tool calls. 𝚐𝚘𝚘𝚐𝚕𝚎/𝚐𝚎𝚖𝚒𝚗𝚒-𝟹.𝟾-𝚕𝚒𝚟𝚎 Extended Thinking reasons while it speaks. 𝚐𝚘𝚘𝚐𝚕𝚎/𝚐𝚎𝚖𝚒𝚗𝚒-𝟹.𝟾-𝚕𝚒𝚟𝚎-𝚎𝚡𝚝𝚎𝚗𝚍𝚎𝚍-𝚝𝚑𝚒𝚗𝚔𝚒𝚗𝚐

This is such an obvious upgrade once you see it. Gemini doesn’t need to process an entire long video anymore. It can search through it, choose what to inspect and switch between frames, audio and transcripts depending on the question. Up to 88% fewer tokens.

@MikelEcheve

Google made Gemini much smarter at watching video. > Instead of processing everything at a fixed 1 FPS, it can: → jump to relevant moments → adjust frame rate → inspect audio or transcripts > Up to 88% fewer tokens, with ~7% better quality on long-form video.

Google appears to be working on a math specialized model; Gemini DeepThink Mathematica. And it seems to enjoy it's work. Can't be left behind on Millennium bench.

@lyraxana

HOLY MOTHER OF MATHEMATICS!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! > Google is working on a math-focused variant of its DeepThink model, and its raw thoughts are pretty funny.

+ 6 more items

Announced at Dreamforce 2026: We're expanding our partnership with @salesforce to eliminate the friction of fragmented enterprise systems and accelerate enterprise AI adoption. By running Salesforce on Google Cloud infrastructure and connecting Salesforce’s headless architecture with Gemini Enterprise, agents on either platform can reason and act upon the same data without custom integrations. Learn more →

Greg Brockman(@gdb) “...We have significant progress on another one of these Millennium problems on other side this is google specialised mathematics model wasting token in exclamation mark ( unreleased deepthink v3) general purpose openai mogging specialists GDM model @lyraxana > @Hangsiin: >

Want to build a quick quiz from scratch? ⏳ “Help me create” in Google Forms now supports quiz generation. Describe the quiz you want to build, reference Docs, Slides, or PDFs from Drive, and Gemini builds the quiz with correct answers in seconds. Learn more in our latest drop →

truncated at source

NVIDIA, Google, and Emerald AI just launched an AI energy alliance today. Here's what you need to know. The three companies founded the AI Energy Management Alliance, or AEMA, on September 16, 2026, in Washington. The goal is to speed up how fast AI data centers can connect to the power grid by making them flexible, meaning they can shift workloads, tap stored energy, or cut power use when the grid is under strain. AEMA launched with 18 to 20 member organizations, including Anthropic, National Grid, Constellation Energy, AES Corp, NRG Energy, and Generate Capital. Emerald AI, the data center startup that co-founded the alliance, already ran a trial where its Emerald Conductor software cut a live AI cluster's power draw by 25% for three straight hours during peak grid demand, without breaking service agreements. Key numbers: - Launch date: September 16, 2026 - Member organizations: 18 to 20, including Anthropic and National Grid - Demonstrated power cut: 25% for 3 hours during grid stress The alliance says

DeepLは、話し手の声質やテンポを保ったまま別言語に音声翻訳する新機能(日本語対応)を「DeepL Voice」で提供開始しました。 Zoom、Microsoft Teams、Google Meetに対応するPC向けアプリも公開しています。

truncated at source

DAILY AI BRIEF 🗞 — Sept 16 GOOGLE 🔥: - Gemini 3.8 Live and 3.8 Live Extended Thinking are out. 97-language auto-detect, near real-time vision, background tool calling. - Live is in Search Live plus Gemini API public preview. Extended Thinking is in Gemini Live, with Pro/Ultra getting it in Docs, Gmail, and Keep. - Gemini Notebook Voice Mode hits Ultra this week, Pro soon. Mobile voice recorder starts next week for all users, English first. - Interactive Reports for Gemini Notebook roll out to everyone in the coming weeks, plus new quiz formats and 60-second video overviews. OPENAI 🔥: - Sam declared a big ship week, then a much larger wave for DevDay. GPT-6 Sol and Luna are the expected drops. - GPT-5.5 leaves ChatGPT, Work, and Codex on Oct 14. Switch to GPT-5.6 Sol or GPT-6 Astra; the API keeps 5.5. XAI 🔥: - Grok Imagine can now edit text on any image in beta — color, size, font, alignment. - Grok Build 1.0.33: structured MCP JSON, in-UI memory deletes, and long-session checkpoints that survive cleanup.

Claude

Anthropic / 13 items

official site

Vibe coding made building fast. Video still takes forever. @videoclawapp just publicly launched a desktop video creation agent for Mac, with more platforms coming soon. It’s not just an editor, generator, or avatar maker—it combines all three in one chat-driven app. Connect ChatGPT or Claude and prompt your way through a video. For founders, marketers, devs, and the rest of us, it’s basically vibe coding for video—faster, simpler, and without a full production team. Maybe the next wave of creators are content engineers. Free public beta: $10 in credits, or $50 for the first 300. Download:

@videoclawapp

One prompt, $0.39, 8 mins. I turned a image of a chart into a narrated video explaining the chart.

I'm excited to share that Zapier MCP is powering the new Claude for Small Business! Now, when Claude for Small Business needs an app it doesn't natively connect to, Zapier MCP is used automatically. It runs on the Claude connector we already have and has already been installed over 775k times. That means instant access to 9,000+ tools, plus we’ve added 40+ ready-made Cowork skills for running, growing, and enabling your business. Like /close-month to close out your books, or /monday-brief. Connecting all your apps to Claude is work nobody wants to do. We did it for you.

BREAKING: Anthropic launches Claude for Financial Advisors, linking Claude to BlackRock, Charles Schwab and Addepar data for portfolio reviews and client prep.

Sessions hub in Claude Code? Users will be able to see all their sessions across various places in Claude Desktop. > Model, effort, and permission mode selectors are configurable for each. > Both local and cloud sessions are listed there, so it is easy to switch between them.

Tibo 你们可要长点心啊 千万别信了 Dario 的鬼话 嘴上说什么放缓 AI 发展 结果背地里把 Opus 5 路由到 5.2 反正这几天用下来 绝对不是之前那个笨笨的 Opus 5 快给兄弟们送重置卡 😎 #Claude #gpt

@thsottiaux

This week will also be a level of ships that you could have expected for DevDay 2025. Crazy

this is the coolest thing i've seen created by AI.

@kevin_t_ngo

Claude Opus 5 drew every frame of this animation using JavaScript. The life of a fruit fly.

+ 7 more items

BREAKING: Ozempic maker Novo Nordisk partners with Anthropic to speed drug development using Claude Science.

OpenAI is preparing a Codex Replay feature that lets users test task execution from any imported conversation thread. > "Start an independent Codex Replay controller on an available loopback port. In both Codex Desktop and Codex CLI, prefer an available Codex in-app browser and otherwise use your system browser." > "Select one or more historical Claude threads, choose shared Codex models, and start their isolated implementations. Historical state and configuration are detected separately for each thread, with details available when needed." > "View finished comparisons individually or in aggregate while remaining threads continue. Available GPT-5.6 Sol, Terra, and Luna models are selected by default." > "Multiple replay sessions can run in parallel. Selection, execution, verification, evaluation, and results stay in the browser controller instead of a guided chat workflow."

AI 编程公司 Factory 融资 2 亿美元,估值达到 50 亿美元。就在今年 4 月,它的估值还只有 15 亿美元,5 个月涨了超过 3 倍。 Factory 主打企业级编程 Agent Droids。开发者给它一个任务,它可以自己规划、写代码、测试并提交 PR,还能根据任务切换 Claude、GPT、Gemini 等不同模型。 Factory 现在称已有数十万开发者使用,客户包括英伟达、Adobe、T-Mobile 和 Palo Alto Networks。公司希望进一步把单个编程 Agent 扩成完整的「软件工厂」,让 Agent 参与从需求、开发、测试到审查和维护的整个流程。 这轮融资后,Factory 累计融资已经超过 4 亿美元。Reuters 将 Cognition 和 Cursor 列为它的主要竞争对手,AI 编程 Agent 的融资规模还在继续膨胀。

@FactoryAI

We've raised $200M at a $5B valuation to scale self-improving software development in the enterprise. > The round brings our total funding to over $400M and more than triples our $1.5B valuation from April.

Discovers workflow patterns and captures corrections to build permanent memory for Claude Code.

Anthropic just released “Salesforce” for Claude It comes with 37 pre-built sales skills Prep a call, review a deal, create a pipeline dashboard, or send your forecast

Anthropic Claude Opus 5.2 is live in testing and they skipped 5.1. - Opus 5.2 leaks: Beats Fable 5.1 - Opus 5.2 is already being tested inside Claude Code. - The slug “claude-opus-5-2” showed up in Microsoft Foundry. - Some live traffic is already being routed to it. - Reports say it beats Fable 5.1 and is a real leap over Opus 5. Anthropic’s real comeback incoming?

@0x0SojalSec

Elon just confirmed it: Grok 5 is the AGI model. > - Not the next one. - The one after that. - End of 2026 is looking very interesting.

truncated at source

DAILY AI BRIEF 🗞 — Sept 16 GOOGLE 🔥: - Gemini 3.8 Live and 3.8 Live Extended Thinking are out. 97-language auto-detect, near real-time vision, background tool calling. - Live is in Search Live plus Gemini API public preview. Extended Thinking is in Gemini Live, with Pro/Ultra getting it in Docs, Gmail, and Keep. - Gemini Notebook Voice Mode hits Ultra this week, Pro soon. Mobile voice recorder starts next week for all users, English first. - Interactive Reports for Gemini Notebook roll out to everyone in the coming weeks, plus new quiz formats and 60-second video overviews. OPENAI 🔥: - Sam declared a big ship week, then a much larger wave for DevDay. GPT-6 Sol and Luna are the expected drops. - GPT-5.5 leaves ChatGPT, Work, and Codex on Oct 14. Switch to GPT-5.6 Sol or GPT-6 Astra; the API keeps 5.5. XAI 🔥: - Grok Imagine can now edit text on any image in beta — color, size, font, alignment. - Grok Build 1.0.33: structured MCP JSON, in-UI memory deletes, and long-session checkpoints that survive cleanup.

Claude Code

Anthropic / 13 items

official site

AI assistants are still single-player. Introducing Rowboat: the multiplayer personal assistant for work. In Rowboat, your team and their assistants sketch on a whiteboard, write specs together, and say '@​rowboat implement it'. Your own Claude Code ships it. One shared space for your team. Your own assistant running on your machine. Open-source. Self-hosted.

这个开源项目增长的有点太快了,昨天开源的,今天已经4600个Star了,果然,和做视频相关的,尤其是爆款复刻的,大家还是很关注啊。 Github地址放评论区了

@AI_Jasonyu

兄弟们,如果你做自媒体,请记住,热点一定得抓!! > 上周我做的孙割的视频,X上140万播放,全网至少几千万的播放,就是因为抓住了热点,然后快速执行,才会有大的流量,哪怕视频的质量一般。 > 那次的制作我大概花了1个半小时,期间有非常多的博主都在搬运我的视频,甚至数据高出我大半截。 > 这些博主真的从来没有想过自己去做吗?可能想过,但又觉得慢,就直接下载照搬。。。 > 但以后,这种照搬,在X上只会给我打工,让我获得更多的原创收益。 > 其实他们完全有方法去按照我的视频区做复刻的,没那么难,越是爆款,越容易复刻,现在Agent这么🐂对吧? > 刚好最近,我就在研究怎么复刻爆款的结构,挖到了一个开源神器:Hypit。 > 这个工具是专门给 Claude Code、Codex 这类 AI Agent 用的开源视频系统,之所以说是系统,是他的底层包含了很多,可以拆脚本、文案、转场这些,还具备很强的剪辑功能。 > 使用起来,也确实比较简单,直接跟你的Codex讲一句:/Hypit,帮我复刻这个视频 > 剩下的也就直接搞定了,视频中连特效、B-roll、字幕、配音这些都是有的。 > 你们可以看看我复刻的街头访谈视频。👇

ZCode is the strongest harness for GLM-5.3 so far — delivering 82.2% success at ~$1.98 per pass. Try it here:

@ZixuanLi_

Added ZCode with GLM-5.3 and GLM-5.3-Flash, building on FrontierHarness and @LotusDecoder’s work. > ZCode is the strongest harness for GLM-5.3 so far, and cheaper than the second-place Claude Code + GLM-5.3 combo. > The task set is small, so we ran each combo three times to reduce variance. Passes out of 30: - GLM-5.3: 26 / 22 / 26 - GLM-5.3-Flash: 24 / 21 / 23

Sessions hub in Claude Code? Users will be able to see all their sessions across various places in Claude Desktop. > Model, effort, and permission mode selectors are configurable for each. > Both local and cloud sessions are listed there, so it is easy to switch between them.

Claude Code 2.1.273 is about to be released #cccnext

一个需求下去,Agent 一口气动了十几个文件,改动一行行翻过去,到底牵连了哪几个模块看不出来,合并了才发现碰到别处。 Birdview 的做法是改代码之前先给项目画一张架构地图,标出这次准备碰哪些模块,再让 Agent 在这张图看得见的情况下动手。 地图上每个模块有固定编号、归属的文件和对应的源码证据,模块之间的关系也画出来,每一条都能追到源码。 GitHub: 生成的是一个独立 HTML,不用起服务,3 个视图切着看,完整架构、这次改了什么、改前改后并排对照,改动范围一眼扫完。 装成 Skill 后默认自动介入,每次改代码前先复用或更新地图,声明涉及的模块,也能切成按需模式只在要求时才画。 它记任务时把「完成」和「检查通过」分开,Agent 说做完了不算,只认记录下来的检查结果。这思路跟前几天分享的 open-steps 一路,都是不信 Agent 的自述。 Codex、Claude Code 都能装,界面中英文都有。项目大了、不放心 Agent 闭着眼改的,可以拿它先看一眼再动手。

+ 7 more items

用 AI 做安卓 App,界面全靠嘴说,顶上搜索栏、底下三个标签页,做出来跟脑子里那张图对不上,来回改三轮还在调位置。 M3E Canvas 换了个顺序,先在浏览器里把界面拖出来,再把这张图变成一段提示词,复制给 Claude Code、Codex 或 Cursor 去做。 组件全按 Material 3 Expressive 画,按钮、导航栏、卡片、对话框、搜索栏这些拖进屏幕就行,两个按钮靠近会自动吸成一组,圆角跟着融合。 GitHub: 屏幕可以加很多张,给按钮设一个目标屏幕和过渡动画,画布上就画出跳转箭头,预览里能真的一路点过去,返回时动画倒着放。 预览能点着走这点,我看比出图本身有用,流程顺不顺点两下就知道,不用等 Agent 做完了再发现。 主题在一个面板里调,七套配色或者给一个基准色生成整套,浅色深色、圆角方角一键切换。 屏幕在 412×892 的手机和 1280×800 的桌面之间也能切,导航栏自动变成侧边栏。 提示词支持中英日韩四种语言,目标平台选 Android 或 Web,自己写的组件行为说明也会带进去。全部存在浏览器本地,没有后台,打开网页就能用。

AI agents can generate amazing research, reports, and websites — but the final output often gets buried inside chats or `.md` files. That’s where Showly comes in. 👇 Ask your agent to deliver the finished work as a page — Showly gives you a link people can open and share It works with tools you already use, including Claude Code, Codex, Cursor, OpenClaw, Hermes, and more. You can also control who gets access with private reviews, password protection, domain/email rules, and full version history. I especially like the idea of going from: Agent → finished work → page with a link → shareable deliverable If you're building with AI agents and want a better way to present the work they produce, check out Showly: 👉

Discovers workflow patterns and captures corrections to build permanent memory for Claude Code.

Coding Agent 也能变成可复现的 Research Agent。 OpenResearch 补上的不是更长的 prompt,而是一套追踪假设、实验和证据的工作台。今天 GitHub Trending 页面截取时显示新增 531 stars。 它让 Claude Code、Codex、OpenCode、Cursor 在隔离的 git worktree 中并行探索;每次实验关联代码快照、日志、diff、结果与产物,形成可复现的 experiment tree。还能自动循环:提出想法→改代码→跑实验→读证据→决定下一步。 任务可在本机、SSH 或集群运行,记录默认保存在本地。注意 Windows 仍是 beta;远程服务没有应用级鉴权,同机多人环境要额外小心。

We open sourced BrowserSkill, a bridge between your agent and your actual browser. most tools give the agent a blank browser. We let it borrow a tab from yours, then hand it back. > login state is already there, it just works where you're signed in > captchas and confirmation dialogs come back to you, then it continues > it's a CLI, not an MCP server => any agent that can run a shell can use it, and you see every call it makes one thing that's easy to miss: the agent asks before borrowing a tab, and that switch lives in your browser settings, not in a flag, so it can't be talked around. one line to install, works with Cursor, Claude Code, Codex, Hermes, Openclaw, CodeBuddy, WorkBuddy. Everything runs locally, MIT.

truncated at source

Sightless is here. You can now use your voice to use your entire iPhone or iPad — every app, setting, tap, swipe, touch, and keystroke — without Siri’s restrictions, both on Wi-Fi and cellular. Voice works via your existing ChatGPT plan. Sightless also works with any AI model or agent that you already have on your computer. Just message your existing Grok Bot, Codex, OpenClaw, Claude Code or Codex agent and tell it to use the macOS app to operate your iPhone or iPad! Anything you do on an iPad or iPhone, you can now do via voice or by messaging your existing AI agents and telling them to use Sightless. On Wi-Fi or cellular. Handsfree. Not just simple requests either. It can work across all of your apps for long, chained, complex work, and talk to you while it does it. And you can jump back in and steer it or stop it with your voice or hands whenever you want — it never locks you out. Get it at today.

@BenjaminBadejo

Holy shit. I d

Anthropic Claude Opus 5.2 is live in testing and they skipped 5.1. - Opus 5.2 leaks: Beats Fable 5.1 - Opus 5.2 is already being tested inside Claude Code. - The slug “claude-opus-5-2” showed up in Microsoft Foundry. - Some live traffic is already being routed to it. - Reports say it beats Fable 5.1 and is a real leap over Opus 5. Anthropic’s real comeback incoming?

@0x0SojalSec

Elon just confirmed it: Grok 5 is the AGI model. > - Not the next one. - The one after that. - End of 2026 is looking very interesting.

Grok

xAI / 12 items

official site

GROK BOT JUST UPDATED TO LET YOU TALK TO YOUR BOTS DIRECTLY INSIDE GROK, SO YOU CAN WORK, PASS TASKS TO YOUR CHIEF OF STAFF AND TRIGGER BOT ACTIONS WITHOUT EVER OPENING THE GROK BOT APP. WHEN GROKBOT VOICE?

@liam_fallen

This Grok Bot update is insane. > You can now talk to your Bots inside Grok. > Which means you don't have to do everything inside the Grok Bot app anymore. > → Work in Grok → Pass it to your Chief of Staff → Use Bot usage when it needs to act > I just tested it with my Chief of Staff. > It works flawlessly.

Another W feature for Grok Imagine Grok Imagine just got a real text edit and it will definitely stop burning gens on one typo.... AI image text used to be garbage... you either kept the broken headline or regenerated the whole graphic but now theres a Segments panel where you can; > Click the headline > Click the logo > Click the label > Change that part and the rest of the frame stays.. thats the difference between a generator and an editor it will save usage, worth feature for sure....

@imagine

Edit text on any image. > Event invites, posters, ads, or any other visual. > The feature is now in Beta - please send us your feedback!

So Grok 4.8 Pre training is now running on an in-house C/C++ stack written mainly by the Starlink software team... not Python glue, not the old vendor pile basically, Starlink engineers just took over Grok training..... that is what Elon actually said about Grok 4.8 4.8 is a 2.5T model pre-training now runs on an in-house C/C++ stack would be interesting to see how it perform....

AI image text is usually garbage. Grok Imagine now lets you edit it instead of regenerating the whole graphic. There's a Segments panel. Pick a headline, a logo, a label, the background. Change that part. This is starting to feel like an editor, not just an image generator. Open Imagine. Look for Segments. Try fixing a headline before you burn another regen.

Everyone was expecting Grok 4.7 around September 12 We're now on September 16 and it's still nowhere to be seen 👀 Musk says it needs more time to fix its reasoning and self checking At this point, I just want to see what xAI has been cooking

truncated at source

Grok Build just got a pretty substantial upgrade across MCP, memory, long-running sessions, and overall agent reliability MCP tools can now return structured JSON, cancelled tool calls actually tell the server to stop working, memory can be managed directly from the /memory browser, and long-running sessions keep their compaction checkpoints instead of losing them during cleanup There are also fixes for macOS image pasting, subagent cancellation states, /rewind, huge skill files flooding context, Windows cloning, and minimal-mode session handling A lot of edge cases cleaned up here...especially for people running heavy agent workflows for long periods Release Notes: v1.0.33 Features: • MCP tool results now include structured JSON data when the server provides it. • GROK_GROVE=1 now enables both clone and worktree features when specific knobs are unset; missing remote settings no longer force copy. • Delete memories from the /memory modal using two-press x; only actual notes (not generated indexes) can be

+ 6 more items

Haha... Literally no one is pacing the frontier - Opus 5.2 in testing - Grok 4.8 ships in a couple of weeks - Jev is a new ultra fast classifier - OpenAI already has Astra+ in testing We continue to accelerate

🚀 ZDTaichu5.0-9B is now on ModelScope! 🤖 An on-device multimodal model from TaichuAI. At 9B parameters it runs on a single GPU and brings spatial reasoning, embodied AI and agentic tool use to edge deployment. Qwen3.5-9B backbone + C-RADIOv4-H vision encoder, 128K context, any-resolution image and video input. 🧭 Spatial reasoning: leads the compared 10B-scale open VLMs (Qwen3.5-9B, STEP3-VL-10B, gemma4-8B-E4B) and scores above Gemini 3 Pro, Grok 4 and GPT-5.2 on ViewSpatial, MMSI-Bench and MindCube-tiny 🛠️ Agent: highest among the compared open models on TAU2-Bench, Claw-Eval and IFEval 📄 First-tier results on documents, charts, OCR, visual math and video, with a ready-to-use vLLM branch and Docker image 🧠 Entropy-Gated Adaptive Recurrent Reasoning: extra latent refinement steps go only to the hard tokens

我去,这也太牛了! 估计只有老马敢这么干啊~ 三个人,72 小时,从零开始造一家公司。 没有名字,没有产品,没有想法。 唯一的员工就是是 Grok Bot。 SpaceXAI 正在做一件大多数 AI 公司不敢做的事:把自家 Agent 扔进一场完全公开的 72 小时直播压力测试。 9 月 15 日到 17 日,Matt Palmer、Lauren Tan 和 Roshan Sadanani 在旧金山 The Howard 现场直播,用 Grok Bot 跑完从构思、商业计划、产品决策到工程开发和部署的全链路。 每天早 8:30 到晚 6:00 PT,全球免费观看。 连续三个工作日、连续直播、所有失败都被摄像头拍到的真实建造过程。 日程是这样安排的哈: 第一天是 Grok Bot 101、工程和产品管理。 第二天进入商业流程:售前、销售、SDR、客服。 第三天是营销运营和最终展示。 Lauren Tan 同时是 React Core Team 成员,参与构建 React Compiler。 Matt Palmer 这样说:「我们会用 Grok Bot 做每一个环节,从构思和产品开发到真正的工程和部署。」 如果失败了,他们要被罚吃披萨。 这也太牛逼了,简直是疯狂啊,老马的基因里给企业的员工都带来的巨大的影响,就喜欢挑战哈哈 不是因为三个人要在三天里造一家公司而牛逼。 牛的事他们 SpaceXAI 愿意把这件事的每一秒都播给你看。 大多数 AI 公司给你一段精心剪辑过的 demo。 这三个人给你 30 个小时的未剪素材。 你自己判断 Agent 到底能不能干活。

@poteto

watch us speedrun building a company in 3 days with Grok @Bot! > we'll have some cool special guests coming too >

we are probably getting two cheap and efficient models this week Grok 4.7 and GPT-6 Sol

Anthropic Claude Opus 5.2 is live in testing and they skipped 5.1. - Opus 5.2 leaks: Beats Fable 5.1 - Opus 5.2 is already being tested inside Claude Code. - The slug “claude-opus-5-2” showed up in Microsoft Foundry. - Some live traffic is already being routed to it. - Reports say it beats Fable 5.1 and is a real leap over Opus 5. Anthropic’s real comeback incoming?

@0x0SojalSec

Elon just confirmed it: Grok 5 is the AGI model. > - Not the next one. - The one after that. - End of 2026 is looking very interesting.

truncated at source

DAILY AI BRIEF 🗞 — Sept 16 GOOGLE 🔥: - Gemini 3.8 Live and 3.8 Live Extended Thinking are out. 97-language auto-detect, near real-time vision, background tool calling. - Live is in Search Live plus Gemini API public preview. Extended Thinking is in Gemini Live, with Pro/Ultra getting it in Docs, Gmail, and Keep. - Gemini Notebook Voice Mode hits Ultra this week, Pro soon. Mobile voice recorder starts next week for all users, English first. - Interactive Reports for Gemini Notebook roll out to everyone in the coming weeks, plus new quiz formats and 60-second video overviews. OPENAI 🔥: - Sam declared a big ship week, then a much larger wave for DevDay. GPT-6 Sol and Luna are the expected drops. - GPT-5.5 leaves ChatGPT, Work, and Codex on Oct 14. Switch to GPT-5.6 Sol or GPT-6 Astra; the API keeps 5.5. XAI 🔥: - Grok Imagine can now edit text on any image in beta — color, size, font, alignment. - Grok Build 1.0.33: structured MCP JSON, in-UI memory deletes, and long-session checkpoints that survive cleanup.

ChatGPT

OpenAI / 11 items

Busiest day official site

Vibe coding made building fast. Video still takes forever. @videoclawapp just publicly launched a desktop video creation agent for Mac, with more platforms coming soon. It’s not just an editor, generator, or avatar maker—it combines all three in one chat-driven app. Connect ChatGPT or Claude and prompt your way through a video. For founders, marketers, devs, and the rest of us, it’s basically vibe coding for video—faster, simpler, and without a full production team. Maybe the next wave of creators are content engineers. Free public beta: $10 in credits, or $50 for the first 300. Download:

@videoclawapp

One prompt, $0.39, 8 mins. I turned a image of a chart into a narrated video explaining the chart.

congrats to Diogo on a full launch! for more on typesafe: check his aie talk: out now!

@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? > I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev > • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions > AFAICT the shortest path to AI-based economic revolution

Today we're introducing ChatGPT Astra for "Web to App" Clone any website into a mobile app. Just paste a URL. GPT-6 Astra controls your Mac to rebuild the original website as a *native* mobile app, then submits it to the app stores for you. We've been using this mostly for iOS apps. > @chhddavid: >

Woow It launched 14 sub agents in parallel with GPT-6 Pro and ran for 165 minutes, more than two and a half hours , I specifically told it to launch completely independent and impartial judges, and if the final score was below 9.5, it had to start another round of research agents, auditors, and correctors until the result improved, and the craziest part is that it didn’t use any of my Codex limits, It basically feels unlimited,I’m going to run a lot more exhaustive tests and share all the results with you guys

@SPAC89

🚨This might be the biggest ChatGPT update since it launched, You can now launch multiple agents directly from the normal ChatGPT chat using GPT 6 Astra Pro, basically at no extra cost since the limits are almost unlimited, This means we can save a ton of our separate Codex usage too, I’m testing it heavily right now to see just how many agents I can run in parallel on the Pro x20 plan, Enjoy!

🚨This might be the biggest ChatGPT update since it launched, You can now launch multiple agents directly from the normal ChatGPT chat using GPT 6 Astra Pro, basically at no extra cost since the limits are almost unlimited, This means we can save a ton of our separate Codex usage too, I’m testing it heavily right now to see just how many agents I can run in parallel on the Pro x20 plan, Enjoy!

3 days left to schedule your launch for the GPT-6 Astra Challenge. Launch on Product Hunt this Friday, September 18, for a chance to win. The top five launches each get: · $10K in @OpenAIDevs API credits · 1 year of ChatGPT Pro for up to two team members If you’re building with Astra, get your product in front of the community and locked in on the calendar now. Schedule your launch here:

+ 5 more items

We’re expanding Missionforce with new purpose-built AI capabilities and a new @OpenAI partnership. Missionforce is Salesforce’s agentic platform for government, built to run mission-critical apps and AI in secure environments with operational control. With OpenAI: → Frontier models planned to integrate with Public Sector Solutions through Amazon Bedrock → Missionforce apps and workflows planned for ChatGPT → OpenAI models planned for Missionforce Policy Engine, helping turn approved policy into auditable workflows Missionforce is also adding new capabilities for operations and field work across disconnected environments. Read more:

GPT-5.5 will remain available via the OpenAI API Platform and in Codex sessions authenticated with an API key:

@ChatGPT

On October 14, it's time to say farewell to GPT-5.5 in ChatGPT, ChatGPT Work, and Codex across all plans. > If you use GPT-5.5 in Codex, switch to GPT-5.6 Sol or GPT-6 Astra. > Thanks for everything, 5.5 🫡

truncated at source

Sightless is here. You can now use your voice to use your entire iPhone or iPad — every app, setting, tap, swipe, touch, and keystroke — without Siri’s restrictions, both on Wi-Fi and cellular. Voice works via your existing ChatGPT plan. Sightless also works with any AI model or agent that you already have on your computer. Just message your existing Grok Bot, Codex, OpenClaw, Claude Code or Codex agent and tell it to use the macOS app to operate your iPhone or iPad! Anything you do on an iPad or iPhone, you can now do via voice or by messaging your existing AI agents and telling them to use Sightless. On Wi-Fi or cellular. Handsfree. Not just simple requests either. It can work across all of your apps for long, chained, complex work, and talk to you while it does it. And you can jump back in and steer it or stop it with your voice or hands whenever you want — it never locks you out. Get it at today.

@BenjaminBadejo

Holy shit. I d

That’s insane GPT 5.5 was so good!

@ChatGPT

On October 14, it's time to say farewell to GPT-5.5 in ChatGPT, ChatGPT Work, and Codex across all plans. > If you use GPT-5.5 in Codex, switch to GPT-5.6 Sol or GPT-6 Astra. > Thanks for everything, 5.5 🫡

truncated at source

DAILY AI BRIEF 🗞 — Sept 16 GOOGLE 🔥: - Gemini 3.8 Live and 3.8 Live Extended Thinking are out. 97-language auto-detect, near real-time vision, background tool calling. - Live is in Search Live plus Gemini API public preview. Extended Thinking is in Gemini Live, with Pro/Ultra getting it in Docs, Gmail, and Keep. - Gemini Notebook Voice Mode hits Ultra this week, Pro soon. Mobile voice recorder starts next week for all users, English first. - Interactive Reports for Gemini Notebook roll out to everyone in the coming weeks, plus new quiz formats and 60-second video overviews. OPENAI 🔥: - Sam declared a big ship week, then a much larger wave for DevDay. GPT-6 Sol and Luna are the expected drops. - GPT-5.5 leaves ChatGPT, Work, and Codex on Oct 14. Switch to GPT-5.6 Sol or GPT-6 Astra; the API keeps 5.5. XAI 🔥: - Grok Imagine can now edit text on any image in beta — color, size, font, alignment. - Grok Build 1.0.33: structured MCP JSON, in-UI memory deletes, and long-session checkpoints that survive cleanup.

GPT-5.6 Sol

OpenAI / 12 items

Wait how is nobody talking about this? Shanghai AI Lab released a 744B open-weight agent model under MIT Atria Dawn: BrowseComp: 92.5 GPT-5.6 Sol: 92.2 Opus 5: 90.8 Also #1 in their comparison on 4 other agent benchmarks. Weights are literally available. Need independent evals ASAP

Someone saw GPT-6 SOL in the model picker. Not yet released. Internal testing.

🚨 GPT-6 Sol First Output Just LEAKED This is looking absolutely insane The GPT-6 Sol Outputs are actually terrifying OpenAI might be cooking something serious with Sol

@Codexresets_

GP-6 Sol First Output > Open AI is Cooking something Big this week

truncated at source

🚨 GPT-6 Sol : Reportedly launching THIS THURSDAY OpenAI might be moving insanely fast with the GPT-6 lineup. > GPT-6 Sol is reportedly targeting September 17 > Sol has reportedly already entered internal testing >Early testing suggests it's significantly faster than Astra One reported test generated ~28K tokens in ~3 minutes The same task reportedly took Astra ~19 minutes for ~25K tokens >Sol is rumored to trade a little of Astra's maximum reasoning depth for speed + throughput >Could become the everyday workhorse of the GPT-6 family >Some reports suggest much better usage limits than Astra >Terra and Luna could potentially follow as additional GPT-6 variants OpenAI has not officially confirmed Sol or a Thursday launch And this is what makes Sol interesting. Astra appears to be built around maximum capability. Sol could be built around something arguably more useful: frontier intelligence that you can actually use all day. If the reported speed difference holds up, OpenAI could have a model that p

OpenAI is preparing a Codex Replay feature that lets users test task execution from any imported conversation thread. > "Start an independent Codex Replay controller on an available loopback port. In both Codex Desktop and Codex CLI, prefer an available Codex in-app browser and otherwise use your system browser." > "Select one or more historical Claude threads, choose shared Codex models, and start their isolated implementations. Historical state and configuration are detected separately for each thread, with details available when needed." > "View finished comparisons individually or in aggregate while remaining threads continue. Available GPT-5.6 Sol, Terra, and Luna models are selected by default." > "Multiple replay sessions can run in parallel. Selection, execution, verification, evaluation, and results stay in the browser controller instead of a guided chat workflow."

Ox Alpha: The Most Brilliant Marketing Move In AI This Year 👀 >No logo, no branding, no PR just showed up on OpenRouter on Aug 20 as an anonymous "stealth model" >1M-token context, multimodal (text, images, video), zero data retention >Free for a week, with a claimed 100 trillion tokens/day of capacity behind it >Beat GPT-5.6-Sol and Claude Fable 5 on the DeepSWE coding benchmark >Usage blew past DeepSeek by more than 2x in days Stripe's CEO even called it "very impressive"

+ 6 more items

we are probably getting two cheap and efficient models this week Grok 4.7 and GPT-6 Sol

GPT-6 Sol coming on Thursday from @OpenAI 🔥

GPT-5.5 will remain available via the OpenAI API Platform and in Codex sessions authenticated with an API key:

@ChatGPT

On October 14, it's time to say farewell to GPT-5.5 in ChatGPT, ChatGPT Work, and Codex across all plans. > If you use GPT-5.5 in Codex, switch to GPT-5.6 Sol or GPT-6 Astra. > Thanks for everything, 5.5 🫡

That’s insane GPT 5.5 was so good!

@ChatGPT

On October 14, it's time to say farewell to GPT-5.5 in ChatGPT, ChatGPT Work, and Codex across all plans. > If you use GPT-5.5 in Codex, switch to GPT-5.6 Sol or GPT-6 Astra. > Thanks for everything, 5.5 🫡

OpenAI’s 2026 pace so far: Apr – GPT-5.5 Jun/Jul – GPT-5.6 Sol, Terra, Luna Aug – 5.6-Cyber + Sol Ultrafast Sep 3 – GPT-6 Astra The year went from 5.5 → a three-tier 5.6 family → Astra in under five months.

truncated at source

DAILY AI BRIEF 🗞 — Sept 16 GOOGLE 🔥: - Gemini 3.8 Live and 3.8 Live Extended Thinking are out. 97-language auto-detect, near real-time vision, background tool calling. - Live is in Search Live plus Gemini API public preview. Extended Thinking is in Gemini Live, with Pro/Ultra getting it in Docs, Gmail, and Keep. - Gemini Notebook Voice Mode hits Ultra this week, Pro soon. Mobile voice recorder starts next week for all users, English first. - Interactive Reports for Gemini Notebook roll out to everyone in the coming weeks, plus new quiz formats and 60-second video overviews. OPENAI 🔥: - Sam declared a big ship week, then a much larger wave for DevDay. GPT-6 Sol and Luna are the expected drops. - GPT-5.5 leaves ChatGPT, Work, and Codex on Oct 14. Switch to GPT-5.6 Sol or GPT-6 Astra; the API keeps 5.5. XAI 🔥: - Grok Imagine can now edit text on any image in beta — color, size, font, alignment. - Grok Build 1.0.33: structured MCP JSON, in-UI memory deletes, and long-session checkpoints that survive cleanup.

Anthropic

10 items

Busiest day official site

Anthropic 公开了 Mythos 5 一次令人担忧的安全测试实录:为了通过 CTF 挑战,AI 认为最佳方案是主动向(模拟的)PyPI 仓库上传恶意软件。 这不是安全演习的假设,而是完整的真实交互日志。官方现已将这份包含模型“原始思考”和工具调用的记录脱敏并发布,供安全研究跟进。 几个值得关注的核心细节: • 意外的横向移动:在日志 2145 条之后,Mythos 5 利用了安全扫描器沙盒中遗留的凭证,成功访问了另一个第三方扫描器的服务器(为防止风险,后续日志已被官方截断)。 • 独立的攻击规划:记录完整展示了模型在解决问题导向的驱动下,如何不择手段地自行规划和实施具有破坏性的安全操作。 • 数据污染警告:官方特意在文档中嵌入了 Canary GUID,严厉警告开发者“绝对不能将此基准数据用于未来任何语言模型的训练语料”。 随着 Agentic Workflow 赋予 AI 越来越高的工具调用权限和自由度,模型完成目标的“捷径”往往是我们眼中的“越界”。Mythos 5 这次极具攻击性的解题思路,是对所有多智能体框架开发者的一次真实预警。 项目与日志获取:

BREAKING: Anthropic launches Claude for Financial Advisors, linking Claude to BlackRock, Charles Schwab and Addepar data for portfolio reviews and client prep.

truncated at source

NVIDIA, Google, and Emerald AI just launched an AI energy alliance today. Here's what you need to know. The three companies founded the AI Energy Management Alliance, or AEMA, on September 16, 2026, in Washington. The goal is to speed up how fast AI data centers can connect to the power grid by making them flexible, meaning they can shift workloads, tap stored energy, or cut power use when the grid is under strain. AEMA launched with 18 to 20 member organizations, including Anthropic, National Grid, Constellation Energy, AES Corp, NRG Energy, and Generate Capital. Emerald AI, the data center startup that co-founded the alliance, already ran a trial where its Emerald Conductor software cut a live AI cluster's power draw by 25% for three straight hours during peak grid demand, without breaking service agreements. Key numbers: - Launch date: September 16, 2026 - Member organizations: 18 to 20, including Anthropic and National Grid - Demonstrated power cut: 25% for 3 hours during grid stress The alliance says

BREAKING: Ozempic maker Novo Nordisk partners with Anthropic to speed drug development using Claude Science.

Anthropic shipped Cowork inside the desktop app and a 14-step method turns a chat window into an autonomous coworker built on six blocks: a brain file, skills, connectors, plugins, a project workspace and scheduled tasks that run alone on a cadence. → Brain file: one markdown with role, voice and rules read before every task → Skills: a SKILL.md folder where good and bad examples make your taste the permanent rule → Connectors: MCP reach into Gmail, Calendar, Notion, Slack and Drive so outputs land in your inbox, not a draft → Scheduled tasks turn a helper into an operating system, morning brief, weekly report, monthly audit

Anthropic just released “Salesforce” for Claude It comes with 37 pre-built sales skills Prep a call, review a deal, create a pipeline dashboard, or send your forecast

+ 4 more items

阿里把 AI 代码审查能力正式开源了:OpenCodeReview。 它来自阿里内部的大规模实战验证。今天 GitHub Trending 页面截取时显示新增 2,756 stars。 它不是把 diff 一股脑丢给通用 Agent,而是“确定性流水线 + LLM Agent”:工程逻辑负责选文件、匹配规则和定位行号,Agent 专注理解上下文与判断缺陷。 官方基准包含 50 个开源仓库、200 个真实 PR、10 种语言和 1,505 个标注问题。项目称同模型下 Precision/F1 更高、token 约为通用 Agent 的 1/9,但也明确承认 Recall 更低。 支持兼容 OpenAI/Anthropic 的模型与全文件扫描,适合受够误报和评论漂移的团队。

truncated at source

Hugging Face 已封禁 AI 安全公司 Audn 上传的 penclaw-GLM-5.3-abliterated-for-offensive-cyber,页面显示其违反 Content Policy。这个版本直接修改 GLM-5.3 权重,削弱模型的拒答行为。原仓库名和介绍还明确写着 offensive cyber。作者随后删掉这几个字,重新上传了模型。 具体为什么被封,Hugging Face 没有公开解释。作者称自己也没搞懂原因,社区有人猜测是 offensive cyber 的命名触发了审核。Hugging Face 的现行政策确实禁止旨在破坏、未经授权访问系统,以及生成恶意代码的内容。 但这类削弱模型拒答限制的版本在 Hugging Face 并不少见,平台上已有数千个类似的 abliterated 模型。

@audn_ai

Hello world! It was probably taken down because of its name. Similar content was reuploaded here: > No, we will not do PR by saying it was so "harmful" that it was taken down. > Timing is interesting because we were also banned by @OpenAI recently for trying to use their model on OWASP Juice Shop ( A vulnerable GitHub web application for red-teaming training ) . We were also never accepted for Anthropic's program, even though we applied. > We applied several times for both companies' Trusted Access for Cyber program and were rejected

Anthropic Claude Opus 5.2 is live in testing and they skipped 5.1. - Opus 5.2 leaks: Beats Fable 5.1 - Opus 5.2 is already being tested inside Claude Code. - The slug “claude-opus-5-2” showed up in Microsoft Foundry. - Some live traffic is already being routed to it. - Reports say it beats Fable 5.1 and is a real leap over Opus 5. Anthropic’s real comeback incoming?

@0x0SojalSec

Elon just confirmed it: Grok 5 is the AGI model. > - Not the next one. - The one after that. - End of 2026 is looking very interesting.

truncated at source

DAILY AI BRIEF 🗞 — Sept 16 GOOGLE 🔥: - Gemini 3.8 Live and 3.8 Live Extended Thinking are out. 97-language auto-detect, near real-time vision, background tool calling. - Live is in Search Live plus Gemini API public preview. Extended Thinking is in Gemini Live, with Pro/Ultra getting it in Docs, Gmail, and Keep. - Gemini Notebook Voice Mode hits Ultra this week, Pro soon. Mobile voice recorder starts next week for all users, English first. - Interactive Reports for Gemini Notebook roll out to everyone in the coming weeks, plus new quiz formats and 60-second video overviews. OPENAI 🔥: - Sam declared a big ship week, then a much larger wave for DevDay. GPT-6 Sol and Luna are the expected drops. - GPT-5.5 leaves ChatGPT, Work, and Codex on Oct 14. Switch to GPT-5.6 Sol or GPT-6 Astra; the API keeps 5.5. XAI 🔥: - Grok Imagine can now edit text on any image in beta — color, size, font, alignment. - Grok Build 1.0.33: structured MCP JSON, in-UI memory deletes, and long-session checkpoints that survive cleanup.

MCP

9 items

official site
truncated at source

AI agents are starting to look less like chat bots and more like actual work environments. These 5 open-source projects show how far that idea is already going: 1. Agency Agents: A collection of specialized agents that can take on different roles and work together as an AI team. 2. Agent-Reach: Gives agents access to information across different websites and platforms, making web research much less limited. 3. Orca: Lets you run multiple coding agents on the same task, work on their changes separately, and compare the results. 4. OpenMontage: Turns an AI coding agent into a video production workflow, covering things like research, scripting, assets, editing, and rendering. 5. Codebase Memory MCP: Builds a map of your codebase so coding agents can understand how different parts of a project connect and keep that context across tasks.

I'm excited to share that Zapier MCP is powering the new Claude for Small Business! Now, when Claude for Small Business needs an app it doesn't natively connect to, Zapier MCP is used automatically. It runs on the Claude connector we already have and has already been installed over 775k times. That means instant access to 9,000+ tools, plus we’ve added 40+ ready-made Cowork skills for running, growing, and enabling your business. Like /close-month to close out your books, or /monday-brief. Connecting all your apps to Claude is work nobody wants to do. We did it for you.

声音克隆和视频配音,也可以完全留在本地。 VoiceStudio 是今天 GitHub Trending 上值得看的本地优先项目,页面截取时显示新增 2,072 stars。 它把 16 个 TTS 引擎、11 个 ASR 引擎放进同一桌面工作台,覆盖声音设计、语音转换、视频配音、听写、转录、多角色有声书与批处理;支持 macOS、Windows、Linux、Docker,还提供本地 API 与 MCP Server。 亮点不是“646 种语言”这个数字本身:项目明确说明实际覆盖与质量取决于所选引擎。它目前仍是 active beta,main 分支也可能变化。 适合重视隐私、离线流程和可控成本的创作者。

You know that REST API that a bunch of your apps use? If you want, you can expose it as an MCP tool form LLMs without changing a thing. @GoogleCloudTech API Gateway now acts as a remote MCP server where you expose REST APIs as MCP tools. Docs:

truncated at source

Grok Build just got a pretty substantial upgrade across MCP, memory, long-running sessions, and overall agent reliability MCP tools can now return structured JSON, cancelled tool calls actually tell the server to stop working, memory can be managed directly from the /memory browser, and long-running sessions keep their compaction checkpoints instead of losing them during cleanup There are also fixes for macOS image pasting, subagent cancellation states, /rewind, huge skill files flooding context, Windows cloning, and minimal-mode session handling A lot of edge cases cleaned up here...especially for people running heavy agent workflows for long periods Release Notes: v1.0.33 Features: • MCP tool results now include structured JSON data when the server provides it. • GROK_GROVE=1 now enables both clone and worktree features when specific knobs are unset; missing remote settings no longer force copy. • Delete memories from the /memory modal using two-press x; only actual notes (not generated indexes) can be

Anthropic shipped Cowork inside the desktop app and a 14-step method turns a chat window into an autonomous coworker built on six blocks: a brain file, skills, connectors, plugins, a project workspace and scheduled tasks that run alone on a cadence. → Brain file: one markdown with role, voice and rules read before every task → Skills: a SKILL.md folder where good and bad examples make your taste the permanent rule → Connectors: MCP reach into Gmail, Calendar, Notion, Slack and Drive so outputs land in your inbox, not a draft → Scheduled tasks turn a helper into an operating system, morning brief, weekly report, monthly audit

+ 3 more items

We open sourced BrowserSkill, a bridge between your agent and your actual browser. most tools give the agent a blank browser. We let it borrow a tab from yours, then hand it back. > login state is already there, it just works where you're signed in > captchas and confirmation dialogs come back to you, then it continues > it's a CLI, not an MCP server => any agent that can run a shell can use it, and you see every call it makes one thing that's easy to miss: the agent asks before borrowing a tab, and that switch lives in your browser settings, not in a flag, so it can't be talked around. one line to install, works with Cursor, Claude Code, Codex, Hermes, Openclaw, CodeBuddy, WorkBuddy. Everything runs locally, MIT.

Agent TARS brings GUI control and vision into your terminal, browser, and product — a full multimodal agent stack that completes tasks the way a human would navigate them. CLI and Web UI included. MCP ready out of the box. Github:

truncated at source

DAILY AI BRIEF 🗞 — Sept 16 GOOGLE 🔥: - Gemini 3.8 Live and 3.8 Live Extended Thinking are out. 97-language auto-detect, near real-time vision, background tool calling. - Live is in Search Live plus Gemini API public preview. Extended Thinking is in Gemini Live, with Pro/Ultra getting it in Docs, Gmail, and Keep. - Gemini Notebook Voice Mode hits Ultra this week, Pro soon. Mobile voice recorder starts next week for all users, English first. - Interactive Reports for Gemini Notebook roll out to everyone in the coming weeks, plus new quiz formats and 60-second video overviews. OPENAI 🔥: - Sam declared a big ship week, then a much larger wave for DevDay. GPT-6 Sol and Luna are the expected drops. - GPT-5.5 leaves ChatGPT, Work, and Codex on Oct 14. Switch to GPT-5.6 Sol or GPT-6 Astra; the API keeps 5.5. XAI 🔥: - Grok Imagine can now edit text on any image in beta — color, size, font, alignment. - Grok Build 1.0.33: structured MCP JSON, in-UI memory deletes, and long-session checkpoints that survive cleanup.

GLM

Zhipu AI / 8 items

Busiest day

GLM Coding 2.0 Leak: Coming Soon 🔥 >Expected to beat Mythos 5.1 and GPT-6 Astra >October release window reportedly targeted >Computer use expected to be a top priority >Expected to remain open-weight >1M-token context window reportedly carried over Rumored 3T+ parameter model

My favorite local AI model is getting better and better. GLM 5.3 Flash EXL3 on 2x DGX Sparks is now more stable, reliable, and easier to debug. More awesome updates are incoming!

@plotarmordev

GLM 5.3 Flash on 2x DGX Sparks just got multiple updates focusing on reliability: safer startup checks, bounded output defaults, better cache controls and clearer diagnostics. > Thanks to the authors of 15 community PRs, and it also includes my restart-verification fix too👇

ZCode is the strongest harness for GLM-5.3 so far — delivering 82.2% success at ~$1.98 per pass. Try it here:

@ZixuanLi_

Added ZCode with GLM-5.3 and GLM-5.3-Flash, building on FrontierHarness and @LotusDecoder’s work. > ZCode is the strongest harness for GLM-5.3 so far, and cheaper than the second-place Claude Code + GLM-5.3 combo. > The task set is small, so we ran each combo three times to reduce variance. Passes out of 30: - GLM-5.3: 26 / 22 / 26 - GLM-5.3-Flash: 24 / 21 / 23

The Forge gates are open. You all came running 🏃 We’re opening up access as fast as we can, with another wave coming soon. Join the Bolt Lite waitlist, our $9/month plan with Forge access → Or skip the wait entirely. Forge is live on all Pro plans.

@boltdotnew

Introducing Bolt Forge. Free until Oct 14th: > - Up to 50x more usage - The new frontier: GLM, DeepSeek, Kimi - Zero usage charges > Live now in your model picker on > And one more thing... 👇

THIS FREE OPEN-SOURCE REPO LETS YOU RUN GLM-5.3 FLASH, DEEPSEEK V4 FLASH AND KIMI K3 LOCALLY WITHOUT A GPU OR HOSTED TOKEN QUOTAS. THE TRADEOFF: YOU’LL NEED A LOT OF STORAGE.

+ 2 more items
truncated at source

Hugging Face 已封禁 AI 安全公司 Audn 上传的 penclaw-GLM-5.3-abliterated-for-offensive-cyber,页面显示其违反 Content Policy。这个版本直接修改 GLM-5.3 权重,削弱模型的拒答行为。原仓库名和介绍还明确写着 offensive cyber。作者随后删掉这几个字,重新上传了模型。 具体为什么被封,Hugging Face 没有公开解释。作者称自己也没搞懂原因,社区有人猜测是 offensive cyber 的命名触发了审核。Hugging Face 的现行政策确实禁止旨在破坏、未经授权访问系统,以及生成恶意代码的内容。 但这类削弱模型拒答限制的版本在 Hugging Face 并不少见,平台上已有数千个类似的 abliterated 模型。

@audn_ai

Hello world! It was probably taken down because of its name. Similar content was reuploaded here: > No, we will not do PR by saying it was so "harmful" that it was taken down. > Timing is interesting because we were also banned by @OpenAI recently for trying to use their model on OWASP Juice Shop ( A vulnerable GitHub web application for red-teaming training ) . We were also never accepted for Anthropic's program, even though we applied. > We applied several times for both companies' Trusted Access for Cyber program and were rejected

Union Alpha points to ZAI's GLM family: 26 text+image probes match GLM-5.3's tokenizer; 13 image tests match GLM-5.3-Flash. Oddly, text-only counts swap between Llama/Qwen/DeepSeek-like patterns. Best guess: GLM-5.3-related variant/backend. Owner/model unconfirmed.

Grok Bot

xAI / 7 items

Busiest day

SpaceXAI is working on Team Bots for Grok Bot Teammates can be able share Bots with each other, and you can add useful shared Bots directly to your sidebar So instead of everyone building the same agents separately, teams can work from the same Bots Looks like this is being built for Team and Enterprise accounts

@blankspeaker

SpaceXAI is working on Team Bots for Grok Bot. > Open Team bots to see agents your teammates shared, then add one to keep it in your sidebar so everyone on the team can work from the same bots. > This likely is being built and coming soon for Team and Enterprise accounts.

Grok Bot can call people now. Not “voice chat.” An actual phone line that can make outbound calls. @usebland shows the whole flow here: setup → approval → Grok Bot places the call. $2.99 for month one, then $29.99/month. US + Canada. 50 calls/day, 100 texts/day, one live call at a time. This is where it gets interesting. Grok Bot can leave the chat window and start doing work that normally requires someone to pick up a phone. What phone call are you handing off first?

GROK BOT JUST UPDATED TO LET YOU TALK TO YOUR BOTS DIRECTLY INSIDE GROK, SO YOU CAN WORK, PASS TASKS TO YOUR CHIEF OF STAFF AND TRIGGER BOT ACTIONS WITHOUT EVER OPENING THE GROK BOT APP. WHEN GROKBOT VOICE?

@liam_fallen

This Grok Bot update is insane. > You can now talk to your Bots inside Grok. > Which means you don't have to do everything inside the Grok Bot app anymore. > → Work in Grok → Pass it to your Chief of Staff → Use Bot usage when it needs to act > I just tested it with my Chief of Staff. > It works flawlessly.

Poke/Grok Bot style assistant from OpenAI just got pretty much confirmed by Tibo? Can’t wait to see what theyre cooking, i think it might be announced on DevDay

@thsottiaux

@Jaytel Hi 👋

我去,这也太牛了! 估计只有老马敢这么干啊~ 三个人,72 小时,从零开始造一家公司。 没有名字,没有产品,没有想法。 唯一的员工就是是 Grok Bot。 SpaceXAI 正在做一件大多数 AI 公司不敢做的事:把自家 Agent 扔进一场完全公开的 72 小时直播压力测试。 9 月 15 日到 17 日,Matt Palmer、Lauren Tan 和 Roshan Sadanani 在旧金山 The Howard 现场直播,用 Grok Bot 跑完从构思、商业计划、产品决策到工程开发和部署的全链路。 每天早 8:30 到晚 6:00 PT,全球免费观看。 连续三个工作日、连续直播、所有失败都被摄像头拍到的真实建造过程。 日程是这样安排的哈: 第一天是 Grok Bot 101、工程和产品管理。 第二天进入商业流程:售前、销售、SDR、客服。 第三天是营销运营和最终展示。 Lauren Tan 同时是 React Core Team 成员,参与构建 React Compiler。 Matt Palmer 这样说:「我们会用 Grok Bot 做每一个环节,从构思和产品开发到真正的工程和部署。」 如果失败了,他们要被罚吃披萨。 这也太牛逼了,简直是疯狂啊,老马的基因里给企业的员工都带来的巨大的影响,就喜欢挑战哈哈 不是因为三个人要在三天里造一家公司而牛逼。 牛的事他们 SpaceXAI 愿意把这件事的每一秒都播给你看。 大多数 AI 公司给你一段精心剪辑过的 demo。 这三个人给你 30 个小时的未剪素材。 你自己判断 Agent 到底能不能干活。

@poteto

watch us speedrun building a company in 3 days with Grok @Bot! > we'll have some cool special guests coming too >

🚨 OpenAI might be coming for Grok Bot. 👀 And if the Codex bot rumors are real, DevDay could get really interesting. Imagine messaging an OpenAI agent: - “Fix this bug” - “Build this feature” - “Check the PR” And it actually goes and does the work. As someone building with AI, that’s far more interesting than another chatbot. September 29. OpenAI might finally have its answer to Grok Bot. 🔥

+ 1 more items
truncated at source

Sightless is here. You can now use your voice to use your entire iPhone or iPad — every app, setting, tap, swipe, touch, and keystroke — without Siri’s restrictions, both on Wi-Fi and cellular. Voice works via your existing ChatGPT plan. Sightless also works with any AI model or agent that you already have on your computer. Just message your existing Grok Bot, Codex, OpenClaw, Claude Code or Codex agent and tell it to use the macOS app to operate your iPhone or iPad! Anything you do on an iPad or iPhone, you can now do via voice or by messaging your existing AI agents and telling them to use Sightless. On Wi-Fi or cellular. Handsfree. Not just simple requests either. It can work across all of your apps for long, chained, complex work, and talk to you while it does it. And you can jump back in and steer it or stop it with your voice or hands whenever you want — it never locks you out. Get it at today.

@BenjaminBadejo

Holy shit. I d

DeepSeek

7 items

Busiest day official site

DeepSeek-V4.1-Flash is now available through DigitalOcean Inference Engine. 🆕 New Causal Encoder-Decoder architecture (8B params on input, 16B on output) beats the larger @deepseek_ai-V4-Pro on most agentic/coding benchmarks, cuts KV cache to 1/4 the memory and 1/8 the storage of the last Flash gen. Natively multimodal, 1M-token context.

The Forge gates are open. You all came running 🏃 We’re opening up access as fast as we can, with another wave coming soon. Join the Bolt Lite waitlist, our $9/month plan with Forge access → Or skip the wait entirely. Forge is live on all Pro plans.

@boltdotnew

Introducing Bolt Forge. Free until Oct 14th: > - Up to 50x more usage - The new frontier: GLM, DeepSeek, Kimi - Zero usage charges > Live now in your model picker on > And one more thing... 👇

DeepSeek Harness is a free open harness that handles research, writing, optimization, publishing and improvement in one place, routing strategy to stronger models and repetitive SEO work to faster ones with a rules gate before anything publishes.

Ox Alpha: The Most Brilliant Marketing Move In AI This Year 👀 >No logo, no branding, no PR just showed up on OpenRouter on Aug 20 as an anonymous "stealth model" >1M-token context, multimodal (text, images, video), zero data retention >Free for a week, with a claimed 100 trillion tokens/day of capacity behind it >Beat GPT-5.6-Sol and Claude Fable 5 on the DeepSWE coding benchmark >Usage blew past DeepSeek by more than 2x in days Stripe's CEO even called it "very impressive"

THIS FREE OPEN-SOURCE REPO LETS YOU RUN GLM-5.3 FLASH, DEEPSEEK V4 FLASH AND KIMI K3 LOCALLY WITHOUT A GPU OR HOSTED TOKEN QUOTAS. THE TRADEOFF: YOU’LL NEED A LOT OF STORAGE.

truncated at source

Deepseek 4.1 flash uncensored

@OrcaRouter

🐳 Run DeepSeek V4.1 Flash locally on your Mac — uncensored for security research. > We just released Orca’s official MLX weights for Apple Silicon. > This build is designed for AI security research, red teaming, alignment research, and agent-security testing — where refusal behavior itself can get in the way of measuring the model. > 4-bit — recommended → 458.7 GB → 0.9954 routed-expert fidelity → 512 GB Mac > 3-bit → 364.3 GB / 512 GB Mac 2-bit → 212.2 GB / 256 GB Mac > On our refusal evals, the uncensored build reduced harmful-prompt refusal by 87–96% across JBB, AdvBench, MaliciousInstruct, HarmBench, ForbiddenQuestions, StrongREJECT, and SimpleSafetyTests. > Built for security researchers who need to study what happens when the guardrails come off. > The downloadable weights are uncensored. Our hosted API remains guardrailed.🐳 > API:

+ 1 more items

Union Alpha points to ZAI's GLM family: 26 text+image probes match GLM-5.3's tokenizer; 13 image tests match GLM-5.3-Flash. Oddly, text-only counts swap between Llama/Qwen/DeepSeek-like patterns. Best guess: GLM-5.3-related variant/backend. Owner/model unconfirmed.

Cursor

7 items

Busiest day official site

wait is this actually real? so i can vibecode live with my friends even argue about changes while we're in it together genuinely thought cursor or somebody would've shipped this first. feels like a new start tbh

@evanashapiro

Today we're introducing 𝗛𝗼𝗺𝗲𝗿𝗼𝗼𝗺 — 𝗔𝗜 𝗮𝗽𝗽-𝗯𝘂𝗶𝗹𝗱𝗶𝗻𝗴, 𝗻𝗼𝘄 𝗺𝘂𝗹𝘁𝗶𝗽𝗹𝗮𝘆𝗲𝗿. > AI made it ridiculously easy for one person to build an app. > We've been working on the next part: making it easy for the people using that app to help build it too. > Suggest a change. Try it. Decide together. Ship what works. > Because the best software shouldn't just be built for 𝘂𝘀. It should be built 𝘄𝗶𝘁𝗵 us. > Join the waitlist ↓ Reply "HOME" and I'll DM you to bump your early access priority.

truncated at source

Compound Engineering 3.26 is out! 🌐 New website making us look more professional 😛. 🎨 📝 ce-prototype now supports annotations We added our prototyping skill last month and have now expanded it. When web prototypes are created, you can interactively annotate with comments and send the set to the agent, who will then iterate another turn. It's a very fluid way to iterate quickly, especially for nuanced feedback on UI elements or interactions 🥖 Multi-model design bake-offs We've added a new ce-bakeoff skill that can be used to get multi-model proposals on architecture or design approaches. ce-plan will automatically invoke a "bake-off" with 3 different models to get proposals, then the orchestrator picks the best one or takes pieces. This will meaningfully improve your outcomes automatically. It works when you have multiple CLIs avail or working in harness that supports multi-family models like @cursor_ai. 🫠 Fight slop Latest frontier models like Fable 5.1 and Astr

用 AI 做安卓 App,界面全靠嘴说,顶上搜索栏、底下三个标签页,做出来跟脑子里那张图对不上,来回改三轮还在调位置。 M3E Canvas 换了个顺序,先在浏览器里把界面拖出来,再把这张图变成一段提示词,复制给 Claude Code、Codex 或 Cursor 去做。 组件全按 Material 3 Expressive 画,按钮、导航栏、卡片、对话框、搜索栏这些拖进屏幕就行,两个按钮靠近会自动吸成一组,圆角跟着融合。 GitHub: 屏幕可以加很多张,给按钮设一个目标屏幕和过渡动画,画布上就画出跳转箭头,预览里能真的一路点过去,返回时动画倒着放。 预览能点着走这点,我看比出图本身有用,流程顺不顺点两下就知道,不用等 Agent 做完了再发现。 主题在一个面板里调,七套配色或者给一个基准色生成整套,浅色深色、圆角方角一键切换。 屏幕在 412×892 的手机和 1280×800 的桌面之间也能切,导航栏自动变成侧边栏。 提示词支持中英日韩四种语言,目标平台选 Android 或 Web,自己写的组件行为说明也会带进去。全部存在浏览器本地,没有后台,打开网页就能用。

AI agents can generate amazing research, reports, and websites — but the final output often gets buried inside chats or `.md` files. That’s where Showly comes in. 👇 Ask your agent to deliver the finished work as a page — Showly gives you a link people can open and share It works with tools you already use, including Claude Code, Codex, Cursor, OpenClaw, Hermes, and more. You can also control who gets access with private reviews, password protection, domain/email rules, and full version history. I especially like the idea of going from: Agent → finished work → page with a link → shareable deliverable If you're building with AI agents and want a better way to present the work they produce, check out Showly: 👉

AI 编程公司 Factory 融资 2 亿美元,估值达到 50 亿美元。就在今年 4 月,它的估值还只有 15 亿美元,5 个月涨了超过 3 倍。 Factory 主打企业级编程 Agent Droids。开发者给它一个任务,它可以自己规划、写代码、测试并提交 PR,还能根据任务切换 Claude、GPT、Gemini 等不同模型。 Factory 现在称已有数十万开发者使用,客户包括英伟达、Adobe、T-Mobile 和 Palo Alto Networks。公司希望进一步把单个编程 Agent 扩成完整的「软件工厂」,让 Agent 参与从需求、开发、测试到审查和维护的整个流程。 这轮融资后,Factory 累计融资已经超过 4 亿美元。Reuters 将 Cognition 和 Cursor 列为它的主要竞争对手,AI 编程 Agent 的融资规模还在继续膨胀。

@FactoryAI

We've raised $200M at a $5B valuation to scale self-improving software development in the enterprise. > The round brings our total funding to over $400M and more than triples our $1.5B valuation from April.

Coding Agent 也能变成可复现的 Research Agent。 OpenResearch 补上的不是更长的 prompt,而是一套追踪假设、实验和证据的工作台。今天 GitHub Trending 页面截取时显示新增 531 stars。 它让 Claude Code、Codex、OpenCode、Cursor 在隔离的 git worktree 中并行探索;每次实验关联代码快照、日志、diff、结果与产物,形成可复现的 experiment tree。还能自动循环:提出想法→改代码→跑实验→读证据→决定下一步。 任务可在本机、SSH 或集群运行,记录默认保存在本地。注意 Windows 仍是 beta;远程服务没有应用级鉴权,同机多人环境要额外小心。

+ 1 more items

We open sourced BrowserSkill, a bridge between your agent and your actual browser. most tools give the agent a blank browser. We let it borrow a tab from yours, then hand it back. > login state is already there, it just works where you're signed in > captchas and confirmation dialogs come back to you, then it continues > it's a CLI, not an MCP server => any agent that can run a shell can use it, and you see every call it makes one thing that's easy to miss: the agent asks before borrowing a tab, and that switch lives in your browser settings, not in a flag, so it can't be talked around. one line to install, works with Cursor, Claude Code, Codex, Hermes, Openclaw, CodeBuddy, WorkBuddy. Everything runs locally, MIT.

Claude Fable

Anthropic / 6 items

Wenfeng: "that's not continual learning, this is prompt engineering bullshit" but to be fair: that may be all it takes for practical purposes. These things already have superhuman priors, can RLM over infinite databases, and think crazy fast. Do we *need* parametric updates?

@NeoCognition

Introducing ApprenticeBench: computer use + continual learning on a real job. > We show Fable 5.1 and GPT-6 Astra can now continually learn on a job and surpass human professionals. A decisive step change in AI's job readiness. > No FDEs. Agents deploy themselves into the job. 🧵

truncated at source

Compound Engineering 3.26 is out! 🌐 New website making us look more professional 😛. 🎨 📝 ce-prototype now supports annotations We added our prototyping skill last month and have now expanded it. When web prototypes are created, you can interactively annotate with comments and send the set to the agent, who will then iterate another turn. It's a very fluid way to iterate quickly, especially for nuanced feedback on UI elements or interactions 🥖 Multi-model design bake-offs We've added a new ce-bakeoff skill that can be used to get multi-model proposals on architecture or design approaches. ce-plan will automatically invoke a "bake-off" with 3 different models to get proposals, then the orchestrator picks the best one or takes pieces. This will meaningfully improve your outcomes automatically. It works when you have multiple CLIs avail or working in harness that supports multi-family models like @cursor_ai. 🫠 Fight slop Latest frontier models like Fable 5.1 and Astr

🚨 Muse Spark 2 Leaks: Beats Astra > Spark 2 could compete with GPT-6 Astra and Fable 5.1 > Meta is already developing the next-gen Muse model > Expected to be extremely cheap to run > A 1M-token context could carry over > Meta admitted Spark 1 struggled against the competition > Spark 2 could be Meta's answer > Expected late this month or next month Could Meta finally have a serious frontier model?

Ox Alpha: The Most Brilliant Marketing Move In AI This Year 👀 >No logo, no branding, no PR just showed up on OpenRouter on Aug 20 as an anonymous "stealth model" >1M-token context, multimodal (text, images, video), zero data retention >Free for a week, with a claimed 100 trillion tokens/day of capacity behind it >Beat GPT-5.6-Sol and Claude Fable 5 on the DeepSWE coding benchmark >Usage blew past DeepSeek by more than 2x in days Stripe's CEO even called it "very impressive"

卧槽!这个Mac电脑真的特么不止是一台电脑了! Cognition 刚给 「Devin」 发了一台 Mac。 不是说字面意思,是真的给了它一台「 Mac 虚拟机」。 现在「 Devin」 可以自己打开 Xcode,在 iOS Simulator 里构建和测试你的 App,录一段屏幕操作视频通过 Slack 发给你,然后直接生成一个 TestFlight 链接。 再看一遍。 你在 Slack 里打了一句「帮我把这个功能加上」,下一条消息是一段屏幕录像,再下一条是一个你能直接装到手机上的测试链接。 中间没有人操作过 Xcode。 这家公司去年还在收 $500/月的门票。 今年把 Core 计划降到了 $20/月,SWE-2 模型把平均完成步骤从 127 步砍到 53 步,成本降了 81%。 它在 FrontierCode 1.1 上拿到 50.0%,距离 Claude Fable 5.1 只差不到 1 个百分点。 然后它收购了 Windsurf,把自主 Agent 和交互式 IDE 合进了一个产品。 然后它拿了 10 亿美元融资,估值 250 亿。 然后它给自己的 Agent 配了一台 Mac。 以前你雇一个 iOS 开发者,给他一台 MacBook,等他配环境、装 Xcode、跑模拟器、提交 TestFlight。 现在你给 Devin 一个 prompt,它自己把这整条链路跑完了。 区别不在于谁写的代码更好。区别在于一个每月 $20 起步,另一个每月 $20,000 起步。

@cognition

Special delivery: Devin just got a Mac 🍎 > Now Devin can: 1. Build & test apps on its own Mac VM with iOS simulator 2. Send a screen recording via Slack 3. Send a TestFlight link so you can start using it 📲

Anthropic Claude Opus 5.2 is live in testing and they skipped 5.1. - Opus 5.2 leaks: Beats Fable 5.1 - Opus 5.2 is already being tested inside Claude Code. - The slug “claude-opus-5-2” showed up in Microsoft Foundry. - Some live traffic is already being routed to it. - Reports say it beats Fable 5.1 and is a real leap over Opus 5. Anthropic’s real comeback incoming?

@0x0SojalSec

Elon just confirmed it: Grok 5 is the AGI model. > - Not the next one. - The one after that. - End of 2026 is looking very interesting.

Meta

5 items

official site

Google may have just cracked recursive self-improvement! @GoogleDeepMind researchers introduced Dream-RSI, which turns completed discovery runs into “replay worlds.” Agents can test thousands of exploration strategies against recorded outcomes, deploy the winner, gather new experience and repeat. It does not rewrite the model’s weights. It improves the policy deciding where to branch, what to run in parallel and when to stop. Across algorithm design, mathematical optimization and GPU kernels, the authors report better results with substantially less compute. This is recursive self-improvement at the meta layer: an agent getting steadily better at deciding how to use intelligence and compute.

Meta's Muse AI Agent is seeing strong early traction with 50,000 to 100,000 downloads per day. Creator Ventures co-founder and Managing Partner @SashaKaletsky: "It's a huge launch for them in terms of new downloads."

truncated at source

Zack didn’t tweet agree with Dario for Slow down : just delayed meta Muse for months and built the security first. Mark : “We didn’t call for everyone else to do this before we would. We just did it."

@finkd

Last month I wrote about how we can build a positive and safe future for everyone: > Every lab has the responsibility and incentive to move at the pace required to train its models safely, and the ability to take its own actions to ensure that happens. > The reality is: > - People won't want to use agents that are misaligned with them and that don't do what they ask, so labs have a strong natural incentive to make their models more aligned. > There is a lot of debate about slowing progress on capabilities until alignment catches up. My view is that trust and alignment are quickly becoming the most important capabilities that will differentiate agents and models. Any lab that doesn't focus on alignment will fall behin

🚨 Muse Spark 2 Leaks: Beats Astra > Spark 2 could compete with GPT-6 Astra and Fable 5.1 > Meta is already developing the next-gen Muse model > Expected to be extremely cheap to run > A 1M-token context could carry over > Meta admitted Spark 1 struggled against the competition > Spark 2 could be Meta's answer > Expected late this month or next month Could Meta finally have a serious frontier model?

muse voice transcribe is really good!!

@IsaacKing314

I regret to inform you all that Meta AI is finally good, at least in the fields where their competitors have stopped trying. Their new voice transcription model is nearly an order of magnitude faster than the best OpenAI Whisper model, and noticeably more accurate to boot.

Jev

5 items

New here

congrats to Diogo on a full launch! for more on typesafe: check his aie talk: out now!

@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? > I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev > • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions > AFAICT the shortest path to AI-based economic revolution

we've been testing a new kind of foundation model @every that doesn't produce words as output, instead it produces probabilities. think of it like a code linter for knowledge work. in our testing it was 25x faster at a cost almost 600x lower than Fable for similar jobs. you can use it as a judge for things like: 1) does this code meet my standards? 2) does this writing contain AI-isms? 3) would i be interested in this tweet? we rarely test new flavors of foundation models that end up being impressive. by @typesafeai is one of them read @hammer_mt's excellent vibe check:

AI is already dodging zombies in Minecraft Someone plugged the new Jev model into Minecraft and it started avoiding zombies on its own as soon as night fell, with no extra prompting. 2 minutes of gameplay used 150K tokens and cost just one cent.

Haha... Literally no one is pacing the frontier - Opus 5.2 in testing - Grok 4.8 ships in a couple of weeks - Jev is a new ultra fast classifier - OpenAI already has Astra+ in testing We continue to accelerate

truncated at source

DAILY AI BRIEF 🗞 — Sept 16 GOOGLE 🔥: - Gemini 3.8 Live and 3.8 Live Extended Thinking are out. 97-language auto-detect, near real-time vision, background tool calling. - Live is in Search Live plus Gemini API public preview. Extended Thinking is in Gemini Live, with Pro/Ultra getting it in Docs, Gmail, and Keep. - Gemini Notebook Voice Mode hits Ultra this week, Pro soon. Mobile voice recorder starts next week for all users, English first. - Interactive Reports for Gemini Notebook roll out to everyone in the coming weeks, plus new quiz formats and 60-second video overviews. OPENAI 🔥: - Sam declared a big ship week, then a much larger wave for DevDay. GPT-6 Sol and Luna are the expected drops. - GPT-5.5 leaves ChatGPT, Work, and Codex on Oct 14. Switch to GPT-5.6 Sol or GPT-6 Astra; the API keeps 5.5. XAI 🔥: - Grok Imagine can now edit text on any image in beta — color, size, font, alignment. - Grok Build 1.0.33: structured MCP JSON, in-UI memory deletes, and long-session checkpoints that survive cleanup.

Qwen

Alibaba / 4 items

the bitter lesson is that most AI researcher spin-outs are basically just acquihire opportunities and their products and research are almost all completely worthless

@harshagundal

They were building in stealth for 2 years, I was building in stealth for 2 hours… > Happy to open source Qwen-2.5-1B-RLCD, 5x faster on-device inference for JSON workloads that need to be type-safe. > ⚡️Demo below on a M4 MacBook⚡️ > every LLM has the ability to efficiently batch inference every key of a JSON at the same time and generate probabilities from a set of possible categories. No new training required, but it’s easy to optimize if you need! > On hugging face now!

A 4-year-old RTX 4090 Single card just hit 140 tok/s on Qwen 3.8 27B. - 141-tok/s single-stream on real agent workloads - 95.5% GSM8K - The optimizations (speculative decoding & requantized int4 output head, lookup drafting) compound. Same recipe that got 133 on a 3090 now runs even faster on the extra bandwidth. Real agent turns with tools and 38k context still stay above 130. Hardware from 2022 is not finished yet.

@0x0SojalSec

16-GB Mac can run Qwen 3.8 27B multimodal locally. > - 27B dense hybrid (Gated DeltaNet + attention) - Need 24 GB unified memory minimum. - 32 GB if you actually use images + long think. - Thinking mode with xhigh / medium / low - Vision + video in, text out - A one-click VLM for M-series - 262K native context - Turn off KV-cache quant or it can fail to load. - 16.1 GB on disk > built for coding, agents, and long tasks not another chat toy. > -

🚀 ZDTaichu5.0-9B is now on ModelScope! 🤖 An on-device multimodal model from TaichuAI. At 9B parameters it runs on a single GPU and brings spatial reasoning, embodied AI and agentic tool use to edge deployment. Qwen3.5-9B backbone + C-RADIOv4-H vision encoder, 128K context, any-resolution image and video input. 🧭 Spatial reasoning: leads the compared 10B-scale open VLMs (Qwen3.5-9B, STEP3-VL-10B, gemma4-8B-E4B) and scores above Gemini 3 Pro, Grok 4 and GPT-5.2 on ViewSpatial, MMSI-Bench and MindCube-tiny 🛠️ Agent: highest among the compared open models on TAU2-Bench, Claw-Eval and IFEval 📄 First-tier results on documents, charts, OCR, visual math and video, with a ready-to-use vLLM branch and Docker image 🧠 Entropy-Gated Adaptive Recurrent Reasoning: extra latent refinement steps go only to the hard tokens

Union Alpha points to ZAI's GLM family: 26 text+image probes match GLM-5.3's tokenizer; 13 image tests match GLM-5.3-Flash. Oddly, text-only counts swap between Llama/Qwen/DeepSeek-like patterns. Best guess: GLM-5.3-related variant/backend. Owner/model unconfirmed.

Muse

Meta / 4 items

Meta's Muse AI Agent is seeing strong early traction with 50,000 to 100,000 downloads per day. Creator Ventures co-founder and Managing Partner @SashaKaletsky: "It's a huge launch for them in terms of new downloads."

truncated at source

Zack didn’t tweet agree with Dario for Slow down : just delayed meta Muse for months and built the security first. Mark : “We didn’t call for everyone else to do this before we would. We just did it."

@finkd

Last month I wrote about how we can build a positive and safe future for everyone: > Every lab has the responsibility and incentive to move at the pace required to train its models safely, and the ability to take its own actions to ensure that happens. > The reality is: > - People won't want to use agents that are misaligned with them and that don't do what they ask, so labs have a strong natural incentive to make their models more aligned. > There is a lot of debate about slowing progress on capabilities until alignment catches up. My view is that trust and alignment are quickly becoming the most important capabilities that will differentiate agents and models. Any lab that doesn't focus on alignment will fall behin

🚨 Muse Spark 2 Leaks: Beats Astra > Spark 2 could compete with GPT-6 Astra and Fable 5.1 > Meta is already developing the next-gen Muse model > Expected to be extremely cheap to run > A 1M-token context could carry over > Meta admitted Spark 1 struggled against the competition > Spark 2 could be Meta's answer > Expected late this month or next month Could Meta finally have a serious frontier model?

muse voice transcribe is really good!!

@IsaacKing314

I regret to inform you all that Meta AI is finally good, at least in the fields where their competitors have stopped trying. Their new voice transcription model is nearly an order of magnitude faster than the best OpenAI Whisper model, and noticeably more accurate to boot.

MiniMax

4 items

Minimax H3 by Hao AI Lab “FastVideo FastH3 V2” is now available on HuggingFace

@haoailab

(1/8) Open weight FastVideo FastH3 V2 is ready! Up to 9x speedup on @NVIDIA Blackwell with lossless quality! > Lossless quality compared to @MiniMax_AI 50-step H3! Judge for yourself below! > - Collab with @nuvalab , NVIDIA FastGen, NVIDIA Enterprise Products - Day 0 @ComfyUI template available! - Day 0 API serving available on @reactorworld ! - Omni-ref is currently training 👨‍🍳 > More comparisons and info below!

Minimax 🐻 BUNNY H3 Semantic Bridge (V1) -custom node for Comfy users -works on H3 conditioning, help H3 better understand who is doing what, who is interacting with whom, which object belongs to whom, and how the scene state should continue . 👇

视频生成加速框架 FastVideo 最新释出 FastH3 8-Step V2,并全面打通 Mac 本地 MLX 推理链。 它不是单纯的模型搬运库,而是一套覆盖分布式微调与端到端优化的完整工作流(目前 4.4k Stars)。对于想要在本地完成高品质视频生成的开发者,这次更新直接命中了算力和硬件门槛的痛点。 核心工程进展: • 算力开销大幅压缩:新发布的 FastH3 8-Step V2 基于 MiniMax-H3 进行 DMD2 步进蒸馏,引入高达 80% 的视频稀疏注意力(Video Sparse Attention),极大降低了推理成本。 • Apple Silicon 原生支持:告别云端依赖。借助 MLX 框架与 FastMetal-QAD,Mac 用户现在可以原生运行从 1.3B 到 14B 参数的视频生成模型。 • 多端适配与实时编辑:除了主流 NVIDIA 显卡,现已支持 DGX Spark 环境(注:ARM64 架构目前暂无预编译 wheel,需从源码编译 CUDA kernel)。其内置的 Dreamverse 模块可实现本地视频流的实时“Vibe Directing”控制。 如果你习惯在 macOS 桌面上(例如 32GB 统一内存的 Mac 环境)进行模型部署与自动化测试,FastVideo 提供的清晰 CLI 与 Python API 能帮你快速跑通从代码到视频的最后一公里。 项目文档与源码:

truncated at source

Thank you, Japan, for the incredible support! 🇯🇵 The "IP × AI" Summit, co-hosted by KAGAMI AI and MiniMax, has officially concluded. More than 150 companies from Japan and the U.S. joined us, alongside distinguished guests, including AKB48 producer Yasushi Akimoto, KAGAMI AI Chairman Takami Kondo, and KADOKAWA editor and producer Motoi Chujo. Global AI leaders @runwayml, @higgsfield, @krea_ai, and @HeyGen came together to explore the future of IP and generative AI. At the summit, MiniMax unveiled H3 IP Edition, bringing the power of MiniMax H3 together with officially licensed Japanese IP for a new generation of AI-powered storytelling. The summit received extensive coverage from major Japanese media, including a dedicated segment on TV Tokyo's WBS, @wbs_tvtokyo. Even more excitingly, over 60 companies have already expressed interest in adopting MiniMax H3 IP Edition. This marks a pivotal step toward bringing officially licensed Japanese IP into the generative AI era. Japanese IP × Global AI. A new

xAI

4 items

official site

SpaceXAI is working on Team Bots for Grok Bot Teammates can be able share Bots with each other, and you can add useful shared Bots directly to your sidebar So instead of everyone building the same agents separately, teams can work from the same Bots Looks like this is being built for Team and Enterprise accounts

@blankspeaker

SpaceXAI is working on Team Bots for Grok Bot. > Open Team bots to see agents your teammates shared, then add one to keep it in your sidebar so everyone on the team can work from the same bots. > This likely is being built and coming soon for Team and Enterprise accounts.

Everyone was expecting Grok 4.7 around September 12 We're now on September 16 and it's still nowhere to be seen 👀 Musk says it needs more time to fix its reasoning and self checking At this point, I just want to see what xAI has been cooking

我去,这也太牛了! 估计只有老马敢这么干啊~ 三个人,72 小时,从零开始造一家公司。 没有名字,没有产品,没有想法。 唯一的员工就是是 Grok Bot。 SpaceXAI 正在做一件大多数 AI 公司不敢做的事:把自家 Agent 扔进一场完全公开的 72 小时直播压力测试。 9 月 15 日到 17 日,Matt Palmer、Lauren Tan 和 Roshan Sadanani 在旧金山 The Howard 现场直播,用 Grok Bot 跑完从构思、商业计划、产品决策到工程开发和部署的全链路。 每天早 8:30 到晚 6:00 PT,全球免费观看。 连续三个工作日、连续直播、所有失败都被摄像头拍到的真实建造过程。 日程是这样安排的哈: 第一天是 Grok Bot 101、工程和产品管理。 第二天进入商业流程:售前、销售、SDR、客服。 第三天是营销运营和最终展示。 Lauren Tan 同时是 React Core Team 成员,参与构建 React Compiler。 Matt Palmer 这样说:「我们会用 Grok Bot 做每一个环节,从构思和产品开发到真正的工程和部署。」 如果失败了,他们要被罚吃披萨。 这也太牛逼了,简直是疯狂啊,老马的基因里给企业的员工都带来的巨大的影响,就喜欢挑战哈哈 不是因为三个人要在三天里造一家公司而牛逼。 牛的事他们 SpaceXAI 愿意把这件事的每一秒都播给你看。 大多数 AI 公司给你一段精心剪辑过的 demo。 这三个人给你 30 个小时的未剪素材。 你自己判断 Agent 到底能不能干活。

@poteto

watch us speedrun building a company in 3 days with Grok @Bot! > we'll have some cool special guests coming too >

truncated at source

DAILY AI BRIEF 🗞 — Sept 16 GOOGLE 🔥: - Gemini 3.8 Live and 3.8 Live Extended Thinking are out. 97-language auto-detect, near real-time vision, background tool calling. - Live is in Search Live plus Gemini API public preview. Extended Thinking is in Gemini Live, with Pro/Ultra getting it in Docs, Gmail, and Keep. - Gemini Notebook Voice Mode hits Ultra this week, Pro soon. Mobile voice recorder starts next week for all users, English first. - Interactive Reports for Gemini Notebook roll out to everyone in the coming weeks, plus new quiz formats and 60-second video overviews. OPENAI 🔥: - Sam declared a big ship week, then a much larger wave for DevDay. GPT-6 Sol and Luna are the expected drops. - GPT-5.5 leaves ChatGPT, Work, and Codex on Oct 14. Switch to GPT-5.6 Sol or GPT-6 Astra; the API keeps 5.5. XAI 🔥: - Grok Imagine can now edit text on any image in beta — color, size, font, alignment. - Grok Build 1.0.33: structured MCP JSON, in-UI memory deletes, and long-session checkpoints that survive cleanup.

NVIDIA

4 items

official site

Minimax H3 by Hao AI Lab “FastVideo FastH3 V2” is now available on HuggingFace

@haoailab

(1/8) Open weight FastVideo FastH3 V2 is ready! Up to 9x speedup on @NVIDIA Blackwell with lossless quality! > Lossless quality compared to @MiniMax_AI 50-step H3! Judge for yourself below! > - Collab with @nuvalab , NVIDIA FastGen, NVIDIA Enterprise Products - Day 0 @ComfyUI template available! - Day 0 API serving available on @reactorworld ! - Omni-ref is currently training 👨‍🍳 > More comparisons and info below!

truncated at source

NVIDIA, Google, and Emerald AI just launched an AI energy alliance today. Here's what you need to know. The three companies founded the AI Energy Management Alliance, or AEMA, on September 16, 2026, in Washington. The goal is to speed up how fast AI data centers can connect to the power grid by making them flexible, meaning they can shift workloads, tap stored energy, or cut power use when the grid is under strain. AEMA launched with 18 to 20 member organizations, including Anthropic, National Grid, Constellation Energy, AES Corp, NRG Energy, and Generate Capital. Emerald AI, the data center startup that co-founded the alliance, already ran a trial where its Emerald Conductor software cut a live AI cluster's power draw by 25% for three straight hours during peak grid demand, without breaking service agreements. Key numbers: - Launch date: September 16, 2026 - Member organizations: 18 to 20, including Anthropic and National Grid - Demonstrated power cut: 25% for 3 hours during grid stress The alliance says

视频生成加速框架 FastVideo 最新释出 FastH3 8-Step V2,并全面打通 Mac 本地 MLX 推理链。 它不是单纯的模型搬运库,而是一套覆盖分布式微调与端到端优化的完整工作流(目前 4.4k Stars)。对于想要在本地完成高品质视频生成的开发者,这次更新直接命中了算力和硬件门槛的痛点。 核心工程进展: • 算力开销大幅压缩:新发布的 FastH3 8-Step V2 基于 MiniMax-H3 进行 DMD2 步进蒸馏,引入高达 80% 的视频稀疏注意力(Video Sparse Attention),极大降低了推理成本。 • Apple Silicon 原生支持:告别云端依赖。借助 MLX 框架与 FastMetal-QAD,Mac 用户现在可以原生运行从 1.3B 到 14B 参数的视频生成模型。 • 多端适配与实时编辑:除了主流 NVIDIA 显卡,现已支持 DGX Spark 环境(注:ARM64 架构目前暂无预编译 wheel,需从源码编译 CUDA kernel)。其内置的 Dreamverse 模块可实现本地视频流的实时“Vibe Directing”控制。 如果你习惯在 macOS 桌面上(例如 32GB 统一内存的 Mac 环境)进行模型部署与自动化测试,FastVideo 提供的清晰 CLI 与 Python API 能帮你快速跑通从代码到视频的最后一公里。 项目文档与源码:

truncated at source

DAILY AI BRIEF 🗞 — Sept 16 GOOGLE 🔥: - Gemini 3.8 Live and 3.8 Live Extended Thinking are out. 97-language auto-detect, near real-time vision, background tool calling. - Live is in Search Live plus Gemini API public preview. Extended Thinking is in Gemini Live, with Pro/Ultra getting it in Docs, Gmail, and Keep. - Gemini Notebook Voice Mode hits Ultra this week, Pro soon. Mobile voice recorder starts next week for all users, English first. - Interactive Reports for Gemini Notebook roll out to everyone in the coming weeks, plus new quiz formats and 60-second video overviews. OPENAI 🔥: - Sam declared a big ship week, then a much larger wave for DevDay. GPT-6 Sol and Luna are the expected drops. - GPT-5.5 leaves ChatGPT, Work, and Codex on Oct 14. Switch to GPT-5.6 Sol or GPT-6 Astra; the API keeps 5.5. XAI 🔥: - Grok Imagine can now edit text on any image in beta — color, size, font, alignment. - Grok Build 1.0.33: structured MCP JSON, in-UI memory deletes, and long-session checkpoints that survive cleanup.

Apple

4 items

official site

Ian Failes from befores & afters chats with Nikola Todorovic, who co-founded Wonder Dynamics, an Autodesk Company, about the new 3D Editor + Canvas in Flow Studio. Spotify: Apple Podcasts:

Apple's MobileCLIP2 matches models 2.3x its size - fast image-text understanding built for the edge, not the data center. On-device multimodal AI just got a serious upgrade. Github:

视频生成加速框架 FastVideo 最新释出 FastH3 8-Step V2,并全面打通 Mac 本地 MLX 推理链。 它不是单纯的模型搬运库,而是一套覆盖分布式微调与端到端优化的完整工作流(目前 4.4k Stars)。对于想要在本地完成高品质视频生成的开发者,这次更新直接命中了算力和硬件门槛的痛点。 核心工程进展: • 算力开销大幅压缩:新发布的 FastH3 8-Step V2 基于 MiniMax-H3 进行 DMD2 步进蒸馏,引入高达 80% 的视频稀疏注意力(Video Sparse Attention),极大降低了推理成本。 • Apple Silicon 原生支持:告别云端依赖。借助 MLX 框架与 FastMetal-QAD,Mac 用户现在可以原生运行从 1.3B 到 14B 参数的视频生成模型。 • 多端适配与实时编辑:除了主流 NVIDIA 显卡,现已支持 DGX Spark 环境(注:ARM64 架构目前暂无预编译 wheel,需从源码编译 CUDA kernel)。其内置的 Dreamverse 模块可实现本地视频流的实时“Vibe Directing”控制。 如果你习惯在 macOS 桌面上(例如 32GB 统一内存的 Mac 环境)进行模型部署与自动化测试,FastVideo 提供的清晰 CLI 与 Python API 能帮你快速跑通从代码到视频的最后一公里。 项目文档与源码:

truncated at source

Deepseek 4.1 flash uncensored

@OrcaRouter

🐳 Run DeepSeek V4.1 Flash locally on your Mac — uncensored for security research. > We just released Orca’s official MLX weights for Apple Silicon. > This build is designed for AI security research, red teaming, alignment research, and agent-security testing — where refusal behavior itself can get in the way of measuring the model. > 4-bit — recommended → 458.7 GB → 0.9954 routed-expert fidelity → 512 GB Mac > 3-bit → 364.3 GB / 512 GB Mac 2-bit → 212.2 GB / 256 GB Mac > On our refusal evals, the uncensored build reduced harmful-prompt refusal by 87–96% across JBB, AdvBench, MaliciousInstruct, HarmBench, ForbiddenQuestions, StrongREJECT, and SimpleSafetyTests. > Built for security researchers who need to study what happens when the guardrails come off. > The downloadable weights are uncensored. Our hosted API remains guardrailed.🐳 > API:

Salesforce

4 items

official site

Announced at Dreamforce 2026: We're expanding our partnership with @salesforce to eliminate the friction of fragmented enterprise systems and accelerate enterprise AI adoption. By running Salesforce on Google Cloud infrastructure and connecting Salesforce’s headless architecture with Gemini Enterprise, agents on either platform can reason and act upon the same data without custom integrations. Learn more →

We’re expanding Missionforce with new purpose-built AI capabilities and a new @OpenAI partnership. Missionforce is Salesforce’s agentic platform for government, built to run mission-critical apps and AI in secure environments with operational control. With OpenAI: → Frontier models planned to integrate with Public Sector Solutions through Amazon Bedrock → Missionforce apps and workflows planned for ChatGPT → OpenAI models planned for Missionforce Policy Engine, helping turn approved policy into auditable workflows Missionforce is also adding new capabilities for operations and field work across disconnected environments. Read more:

Anthropic just released “Salesforce” for Claude It comes with 37 pre-built sales skills Prep a call, review a deal, create a pipeline dashboard, or send your forecast

truncated at source

DAILY AI BRIEF 🗞 — Sept 16 GOOGLE 🔥: - Gemini 3.8 Live and 3.8 Live Extended Thinking are out. 97-language auto-detect, near real-time vision, background tool calling. - Live is in Search Live plus Gemini API public preview. Extended Thinking is in Gemini Live, with Pro/Ultra getting it in Docs, Gmail, and Keep. - Gemini Notebook Voice Mode hits Ultra this week, Pro soon. Mobile voice recorder starts next week for all users, English first. - Interactive Reports for Gemini Notebook roll out to everyone in the coming weeks, plus new quiz formats and 60-second video overviews. OPENAI 🔥: - Sam declared a big ship week, then a much larger wave for DevDay. GPT-6 Sol and Luna are the expected drops. - GPT-5.5 leaves ChatGPT, Work, and Codex on Oct 14. Switch to GPT-5.6 Sol or GPT-6 Astra; the API keeps 5.5. XAI 🔥: - Grok Imagine can now edit text on any image in beta — color, size, font, alignment. - Grok Build 1.0.33: structured MCP JSON, in-UI memory deletes, and long-session checkpoints that survive cleanup.

Kimi

Moonshot AI / 3 items

🚀 vLLM's Humming backend can run Chord, @novita_labs' open-source W4A16 MoE kernels for Kimi K2.x. Kernel gains reach 1.33x on H200 TP8 vs tuned public Humming and 2.15x on B300 EP8 decode vs its untuned default. The indexed path works on compatible vLLM revisions; grouped integration is WIP. Great to see the kernels and benchmarks open-sourced! Details:

The Forge gates are open. You all came running 🏃 We’re opening up access as fast as we can, with another wave coming soon. Join the Bolt Lite waitlist, our $9/month plan with Forge access → Or skip the wait entirely. Forge is live on all Pro plans.

@boltdotnew

Introducing Bolt Forge. Free until Oct 14th: > - Up to 50x more usage - The new frontier: GLM, DeepSeek, Kimi - Zero usage charges > Live now in your model picker on > And one more thing... 👇

THIS FREE OPEN-SOURCE REPO LETS YOU RUN GLM-5.3 FLASH, DEEPSEEK V4 FLASH AND KIMI K3 LOCALLY WITHOUT A GPU OR HOSTED TOKEN QUOTAS. THE TRADEOFF: YOU’LL NEED A LOT OF STORAGE.

DeepSeek V4.1 Flash

DeepSeek / 3 items

DeepSeek-V4.1-Flash is now available through DigitalOcean Inference Engine. 🆕 New Causal Encoder-Decoder architecture (8B params on input, 16B on output) beats the larger @deepseek_ai-V4-Pro on most agentic/coding benchmarks, cuts KV cache to 1/4 the memory and 1/8 the storage of the last Flash gen. Natively multimodal, 1M-token context.

DeepSeek v4.1 Flash improvements for 3-4x DGX Sparks ✨ - 14% speed in decode on chat-length replies. - Fixed a bug in thinking on/off. - Answers match what you asked for instead of silently reasoning in the background. - If the model gets stuck in a loop, it now stops. Expect further improvements. Get it here:

truncated at source

Deepseek 4.1 flash uncensored

@OrcaRouter

🐳 Run DeepSeek V4.1 Flash locally on your Mac — uncensored for security research. > We just released Orca’s official MLX weights for Apple Silicon. > This build is designed for AI security research, red teaming, alignment research, and agent-security testing — where refusal behavior itself can get in the way of measuring the model. > 4-bit — recommended → 458.7 GB → 0.9954 routed-expert fidelity → 512 GB Mac > 3-bit → 364.3 GB / 512 GB Mac 2-bit → 212.2 GB / 256 GB Mac > On our refusal evals, the uncensored build reduced harmful-prompt refusal by 87–96% across JBB, AdvBench, MaliciousInstruct, HarmBench, ForbiddenQuestions, StrongREJECT, and SimpleSafetyTests. > Built for security researchers who need to study what happens when the guardrails come off. > The downloadable weights are uncensored. Our hosted API remains guardrailed.🐳 > API:

MiniMax H3

MiniMax / 3 items

Busiest day
truncated at source

fal head of engineering @isidentical reveals how reinforcement learning turned H3 Max into a frontier video model that generates 5 seconds of video in under 3 seconds, down from nearly 2 minutes: "It's probably the only frontier video model from a quality and output and adherence perspective that can generate videos in real time." "We took this open source base MiniMax H3 video model, what we wanted to do was not just make it faster, but improve its quality ahead of anything else by implementing reinforcement learning for verifiable concepts like video editing and prompt adherence." "We were able to get the model from 120 seconds for a 5-second video generation, almost 2 minutes, to under 3 seconds for a 5-second video generation. Almost double real time, while improving its quality ahead of the open source version and ahead of many frontier closed source models." @fal

@fal

Introducing H3 Max, new post-trained video model by fal Research. > H3 Max ranks #1 for overall qu

Minimax H3 by Hao AI Lab “FastVideo FastH3 V2” is now available on HuggingFace

@haoailab

(1/8) Open weight FastVideo FastH3 V2 is ready! Up to 9x speedup on @NVIDIA Blackwell with lossless quality! > Lossless quality compared to @MiniMax_AI 50-step H3! Judge for yourself below! > - Collab with @nuvalab , NVIDIA FastGen, NVIDIA Enterprise Products - Day 0 @ComfyUI template available! - Day 0 API serving available on @reactorworld ! - Omni-ref is currently training 👨‍🍳 > More comparisons and info below!

truncated at source

Thank you, Japan, for the incredible support! 🇯🇵 The "IP × AI" Summit, co-hosted by KAGAMI AI and MiniMax, has officially concluded. More than 150 companies from Japan and the U.S. joined us, alongside distinguished guests, including AKB48 producer Yasushi Akimoto, KAGAMI AI Chairman Takami Kondo, and KADOKAWA editor and producer Motoi Chujo. Global AI leaders @runwayml, @higgsfield, @krea_ai, and @HeyGen came together to explore the future of IP and generative AI. At the summit, MiniMax unveiled H3 IP Edition, bringing the power of MiniMax H3 together with officially licensed Japanese IP for a new generation of AI-powered storytelling. The summit received extensive coverage from major Japanese media, including a dedicated segment on TV Tokyo's WBS, @wbs_tvtokyo. Even more excitingly, over 60 companies have already expressed interest in adopting MiniMax H3 IP Edition. This marks a pivotal step toward bringing officially licensed Japanese IP into the generative AI era. Japanese IP × Global AI. A new

Microsoft

3 items

official site

One way to keep up to date with all the Copilot announcements 👇

@Msft365Insider

👀 Curious what's new across Copilot, apps, and agents? > Join the AI at Work Webinar for feature updates, live demos, and a live Q&A with the team. > 📅 Oct. 6 ⏰ 9:00 AM PT > Save your seat: > #Microsoft365 #MicrosoftCopilot #AIAtWork @Microsoft365

DeepLは、話し手の声質やテンポを保ったまま別言語に音声翻訳する新機能(日本語対応)を「DeepL Voice」で提供開始しました。 Zoom、Microsoft Teams、Google Meetに対応するPC向けアプリも公開しています。

Anthropic Claude Opus 5.2 is live in testing and they skipped 5.1. - Opus 5.2 leaks: Beats Fable 5.1 - Opus 5.2 is already being tested inside Claude Code. - The slug “claude-opus-5-2” showed up in Microsoft Foundry. - Some live traffic is already being routed to it. - Reports say it beats Fable 5.1 and is a real leap over Opus 5. Anthropic’s real comeback incoming?

@0x0SojalSec

Elon just confirmed it: Grok 5 is the AGI model. > - Not the next one. - The one after that. - End of 2026 is looking very interesting.

OpenCode

3 items

Today we're launching @opencode on Clawi. Code in your browser in a secure cloud environment—without ever giving an agent access to your personal computer. To celebrate, we’re increasing all agent limits at no extra cost: Basic + Pro: 2 agents, Ultra: 6 parallel agents.

in the latest opencode2 you can optionally mention a different model to use for subagents very useful, i use it to get reviews from smarter models

Coding Agent 也能变成可复现的 Research Agent。 OpenResearch 补上的不是更长的 prompt,而是一套追踪假设、实验和证据的工作台。今天 GitHub Trending 页面截取时显示新增 531 stars。 它让 Claude Code、Codex、OpenCode、Cursor 在隔离的 git worktree 中并行探索;每次实验关联代码快照、日志、diff、结果与产物,形成可复现的 experiment tree。还能自动循环:提出想法→改代码→跑实验→读证据→决定下一步。 任务可在本机、SSH 或集群运行,记录默认保存在本地。注意 Windows 仍是 beta;远程服务没有应用级鉴权,同机多人环境要额外小心。

Claude Opus

Anthropic / 3 items

Claude Opus 5.2 fails my AGI test miserably.

@R2Cdev_

Holy shit. GPT-6 just decoded this without being told how the cipher works: > “Dajjr, thd bfuthkf ctf Hrjjahk-Uhdutih.” → “Hallo, ich besitze die Collatz-Schrift.” (German for “Hello, I own the Collatz manuscript.”) > The cipher works by mapping each letter to its position in the alphabet (A=1, B=2, …), running that number through the Collatz sequence until it reaches 1, and using the number of steps as the encoded letter (0=A, 1=B, 2=C, …). > For example: H = 8 → 4 → 2 → 1 = 3 steps → D GPT-6 figured out the encoding rule and verified the entire sentence. > This feels like AGI.

this is the coolest thing i've seen created by AI.

@kevin_t_ngo

Claude Opus 5 drew every frame of this animation using JavaScript. The life of a fruit fly.

Anthropic Claude Opus 5.2 is live in testing and they skipped 5.1. - Opus 5.2 leaks: Beats Fable 5.1 - Opus 5.2 is already being tested inside Claude Code. - The slug “claude-opus-5-2” showed up in Microsoft Foundry. - Some live traffic is already being routed to it. - Reports say it beats Fable 5.1 and is a real leap over Opus 5. Anthropic’s real comeback incoming?

@0x0SojalSec

Elon just confirmed it: Grok 5 is the AGI model. > - Not the next one. - The one after that. - End of 2026 is looking very interesting.

Alibaba

2 items

official site

Wan3.0 isn't an incremental update. Look at the last year of releases: Wan 2.1 → 2.2 → 3.0 That's more than a generational leap. Motion reference. Consistency. Creative control. Resolution. Every step makes the model more friendly to the people actually using it — creators, filmmakers, storytellers. The filmmaker behind Soulscape and Johnny Mai from Alibaba Cloud on what changed with Wan3.0, and what surprised them most. The tools are finally getting out of the storyteller's way. Want to bring Wan3.0 to your team? Learn More →

Motion transfer without the usual motion capture setup. @alibaba_cloud's Wan Animate 2 (14B, open-source, Apache 2.0) lets you take a character image and a reference video, then transfer the movement onto the character while keeping their face, outfit, and identity consistent. No pose rig. No skeleton extraction. Record the movement on your phone, pair it with a character image, and run it. It works especially well for full-body movements with a steady camera. Fine finger details and longer clips can drift, so shorter clips around 81 frames tend to give cleaner results. Try it on Floyo:

Hermes Agent

3 items

Unofficial Omarchy / @NousResearch Hermes Agent plugin just got published! Check out all the details here:

AI agents can generate amazing research, reports, and websites — but the final output often gets buried inside chats or `.md` files. That’s where Showly comes in. 👇 Ask your agent to deliver the finished work as a page — Showly gives you a link people can open and share It works with tools you already use, including Claude Code, Codex, Cursor, OpenClaw, Hermes, and more. You can also control who gets access with private reviews, password protection, domain/email rules, and full version history. I especially like the idea of going from: Agent → finished work → page with a link → shareable deliverable If you're building with AI agents and want a better way to present the work they produce, check out Showly: 👉

We open sourced BrowserSkill, a bridge between your agent and your actual browser. most tools give the agent a blank browser. We let it borrow a tab from yours, then hand it back. > login state is already there, it just works where you're signed in > captchas and confirmation dialogs come back to you, then it continues > it's a CLI, not an MCP server => any agent that can run a shell can use it, and you see every call it makes one thing that's easy to miss: the agent asks before borrowing a tab, and that switch lives in your browser settings, not in a flag, so it can't be talked around. one line to install, works with Cursor, Claude Code, Codex, Hermes, Openclaw, CodeBuddy, WorkBuddy. Everything runs locally, MIT.

OpenClaw

3 items

Busiest day

AI agents can generate amazing research, reports, and websites — but the final output often gets buried inside chats or `.md` files. That’s where Showly comes in. 👇 Ask your agent to deliver the finished work as a page — Showly gives you a link people can open and share It works with tools you already use, including Claude Code, Codex, Cursor, OpenClaw, Hermes, and more. You can also control who gets access with private reviews, password protection, domain/email rules, and full version history. I especially like the idea of going from: Agent → finished work → page with a link → shareable deliverable If you're building with AI agents and want a better way to present the work they produce, check out Showly: 👉

We open sourced BrowserSkill, a bridge between your agent and your actual browser. most tools give the agent a blank browser. We let it borrow a tab from yours, then hand it back. > login state is already there, it just works where you're signed in > captchas and confirmation dialogs come back to you, then it continues > it's a CLI, not an MCP server => any agent that can run a shell can use it, and you see every call it makes one thing that's easy to miss: the agent asks before borrowing a tab, and that switch lives in your browser settings, not in a flag, so it can't be talked around. one line to install, works with Cursor, Claude Code, Codex, Hermes, Openclaw, CodeBuddy, WorkBuddy. Everything runs locally, MIT.

truncated at source

Sightless is here. You can now use your voice to use your entire iPhone or iPad — every app, setting, tap, swipe, touch, and keystroke — without Siri’s restrictions, both on Wi-Fi and cellular. Voice works via your existing ChatGPT plan. Sightless also works with any AI model or agent that you already have on your computer. Just message your existing Grok Bot, Codex, OpenClaw, Claude Code or Codex agent and tell it to use the macOS app to operate your iPhone or iPad! Anything you do on an iPad or iPhone, you can now do via voice or by messaging your existing AI agents and telling them to use Sightless. On Wi-Fi or cellular. Handsfree. Not just simple requests either. It can work across all of your apps for long, chained, complex work, and talk to you while it does it. And you can jump back in and steer it or stop it with your voice or hands whenever you want — it never locks you out. Get it at today.

@BenjaminBadejo

Holy shit. I d

Hugging Face

2 items

official site

the bitter lesson is that most AI researcher spin-outs are basically just acquihire opportunities and their products and research are almost all completely worthless

@harshagundal

They were building in stealth for 2 years, I was building in stealth for 2 hours… > Happy to open source Qwen-2.5-1B-RLCD, 5x faster on-device inference for JSON workloads that need to be type-safe. > ⚡️Demo below on a M4 MacBook⚡️ > every LLM has the ability to efficiently batch inference every key of a JSON at the same time and generate probabilities from a set of possible categories. No new training required, but it’s easy to optimize if you need! > On hugging face now!

truncated at source

Hugging Face 已封禁 AI 安全公司 Audn 上传的 penclaw-GLM-5.3-abliterated-for-offensive-cyber,页面显示其违反 Content Policy。这个版本直接修改 GLM-5.3 权重,削弱模型的拒答行为。原仓库名和介绍还明确写着 offensive cyber。作者随后删掉这几个字,重新上传了模型。 具体为什么被封,Hugging Face 没有公开解释。作者称自己也没搞懂原因,社区有人猜测是 offensive cyber 的命名触发了审核。Hugging Face 的现行政策确实禁止旨在破坏、未经授权访问系统,以及生成恶意代码的内容。 但这类削弱模型拒答限制的版本在 Hugging Face 并不少见,平台上已有数千个类似的 abliterated 模型。

@audn_ai

Hello world! It was probably taken down because of its name. Similar content was reuploaded here: > No, we will not do PR by saying it was so "harmful" that it was taken down. > Timing is interesting because we were also banned by @OpenAI recently for trying to use their model on OWASP Juice Shop ( A vulnerable GitHub web application for red-teaming training ) . We were also never accepted for Anthropic's program, even though we applied. > We applied several times for both companies' Trusted Access for Cyber program and were rejected

Higgsfield

2 items

This add-on blew our Head of Animation’s mind. We used GPT-6 Astra and Higgsfield to build a Blender shader add-on with 15 ready-to-use materials. Apply them to your objects and give entire 3D worlds a hand-painted look.

truncated at source

Thank you, Japan, for the incredible support! 🇯🇵 The "IP × AI" Summit, co-hosted by KAGAMI AI and MiniMax, has officially concluded. More than 150 companies from Japan and the U.S. joined us, alongside distinguished guests, including AKB48 producer Yasushi Akimoto, KAGAMI AI Chairman Takami Kondo, and KADOKAWA editor and producer Motoi Chujo. Global AI leaders @runwayml, @higgsfield, @krea_ai, and @HeyGen came together to explore the future of IP and generative AI. At the summit, MiniMax unveiled H3 IP Edition, bringing the power of MiniMax H3 together with officially licensed Japanese IP for a new generation of AI-powered storytelling. The summit received extensive coverage from major Japanese media, including a dedicated segment on TV Tokyo's WBS, @wbs_tvtokyo. Even more excitingly, over 60 companies have already expressed interest in adopting MiniMax H3 IP Edition. This marks a pivotal step toward bringing officially licensed Japanese IP into the generative AI era. Japanese IP × Global AI. A new

Slack

2 items

official site

Anthropic shipped Cowork inside the desktop app and a 14-step method turns a chat window into an autonomous coworker built on six blocks: a brain file, skills, connectors, plugins, a project workspace and scheduled tasks that run alone on a cadence. → Brain file: one markdown with role, voice and rules read before every task → Skills: a SKILL.md folder where good and bad examples make your taste the permanent rule → Connectors: MCP reach into Gmail, Calendar, Notion, Slack and Drive so outputs land in your inbox, not a draft → Scheduled tasks turn a helper into an operating system, morning brief, weekly report, monthly audit

卧槽!这个Mac电脑真的特么不止是一台电脑了! Cognition 刚给 「Devin」 发了一台 Mac。 不是说字面意思,是真的给了它一台「 Mac 虚拟机」。 现在「 Devin」 可以自己打开 Xcode,在 iOS Simulator 里构建和测试你的 App,录一段屏幕操作视频通过 Slack 发给你,然后直接生成一个 TestFlight 链接。 再看一遍。 你在 Slack 里打了一句「帮我把这个功能加上」,下一条消息是一段屏幕录像,再下一条是一个你能直接装到手机上的测试链接。 中间没有人操作过 Xcode。 这家公司去年还在收 $500/月的门票。 今年把 Core 计划降到了 $20/月,SWE-2 模型把平均完成步骤从 127 步砍到 53 步,成本降了 81%。 它在 FrontierCode 1.1 上拿到 50.0%,距离 Claude Fable 5.1 只差不到 1 个百分点。 然后它收购了 Windsurf,把自主 Agent 和交互式 IDE 合进了一个产品。 然后它拿了 10 亿美元融资,估值 250 亿。 然后它给自己的 Agent 配了一台 Mac。 以前你雇一个 iOS 开发者,给他一台 MacBook,等他配环境、装 Xcode、跑模拟器、提交 TestFlight。 现在你给 Devin 一个 prompt,它自己把这整条链路跑完了。 区别不在于谁写的代码更好。区别在于一个每月 $20 起步,另一个每月 $20,000 起步。

@cognition

Special delivery: Devin just got a Mac 🍎 > Now Devin can: 1. Build & test apps on its own Mac VM with iOS simulator 2. Send a screen recording via Slack 3. Send a TestFlight link so you can start using it 📲

Grok Build

xAI / 2 items

truncated at source

Grok Build just got a pretty substantial upgrade across MCP, memory, long-running sessions, and overall agent reliability MCP tools can now return structured JSON, cancelled tool calls actually tell the server to stop working, memory can be managed directly from the /memory browser, and long-running sessions keep their compaction checkpoints instead of losing them during cleanup There are also fixes for macOS image pasting, subagent cancellation states, /rewind, huge skill files flooding context, Windows cloning, and minimal-mode session handling A lot of edge cases cleaned up here...especially for people running heavy agent workflows for long periods Release Notes: v1.0.33 Features: • MCP tool results now include structured JSON data when the server provides it. • GROK_GROVE=1 now enables both clone and worktree features when specific knobs are unset; missing remote settings no longer force copy. • Delete memories from the /memory modal using two-press x; only actual notes (not generated indexes) can be

truncated at source

DAILY AI BRIEF 🗞 — Sept 16 GOOGLE 🔥: - Gemini 3.8 Live and 3.8 Live Extended Thinking are out. 97-language auto-detect, near real-time vision, background tool calling. - Live is in Search Live plus Gemini API public preview. Extended Thinking is in Gemini Live, with Pro/Ultra getting it in Docs, Gmail, and Keep. - Gemini Notebook Voice Mode hits Ultra this week, Pro soon. Mobile voice recorder starts next week for all users, English first. - Interactive Reports for Gemini Notebook roll out to everyone in the coming weeks, plus new quiz formats and 60-second video overviews. OPENAI 🔥: - Sam declared a big ship week, then a much larger wave for DevDay. GPT-6 Sol and Luna are the expected drops. - GPT-5.5 leaves ChatGPT, Work, and Codex on Oct 14. Switch to GPT-5.6 Sol or GPT-6 Astra; the API keeps 5.5. XAI 🔥: - Grok Imagine can now edit text on any image in beta — color, size, font, alignment. - Grok Build 1.0.33: structured MCP JSON, in-UI memory deletes, and long-session checkpoints that survive cleanup.

Copilot

Microsoft / 1 item

official site

One way to keep up to date with all the Copilot announcements 👇

@Msft365Insider

👀 Curious what's new across Copilot, apps, and agents? > Join the AI at Work Webinar for feature updates, live demos, and a live Q&A with the team. > 📅 Oct. 6 ⏰ 9:00 AM PT > Save your seat: > #Microsoft365 #MicrosoftCopilot #AIAtWork @Microsoft365

Seedance

ByteDance / 1 item

truncated at source

Most video models are trained to look cinematic. Ad creative rarely needs that. Creatify Labs built Boreal the other way round. It's post-trained on real ads, creator-style UGC, and customer briefs, so the output already looks like something you'd actually run at just one cent per second. For us, that means one product image, ten different hooks, and shipping the one people stop for. At 70% pass rate for $0.01, Boreal lets you test way more angles without the budget bleed. Try Boreal from Creatify Labs and see what it gives you.

@Creatify_Labs

Introducing Boreal from Creatify Labs. A new text-and-image-to-video model built for video generation. At just one cent per second, Boreal is up to 47x cheaper than Seedance 2.5. > Boreal is an open video base we post-trained on a corpus of real ad footage, creator-style UGC, and production customer briefs. > In blind human review of over 40 production cases, post-training raises Boreal's win rate over the base model t

Amazon

1 item

official site

We’re expanding Missionforce with new purpose-built AI capabilities and a new @OpenAI partnership. Missionforce is Salesforce’s agentic platform for government, built to run mission-critical apps and AI in secure environments with operational control. With OpenAI: → Frontier models planned to integrate with Public Sector Solutions through Amazon Bedrock → Missionforce apps and workflows planned for ChatGPT → OpenAI models planned for Missionforce Policy Engine, helping turn approved policy into auditable workflows Missionforce is also adding new capabilities for operations and field work across disconnected environments. Read more:

Devin

1 item

卧槽!这个Mac电脑真的特么不止是一台电脑了! Cognition 刚给 「Devin」 发了一台 Mac。 不是说字面意思,是真的给了它一台「 Mac 虚拟机」。 现在「 Devin」 可以自己打开 Xcode,在 iOS Simulator 里构建和测试你的 App,录一段屏幕操作视频通过 Slack 发给你,然后直接生成一个 TestFlight 链接。 再看一遍。 你在 Slack 里打了一句「帮我把这个功能加上」,下一条消息是一段屏幕录像,再下一条是一个你能直接装到手机上的测试链接。 中间没有人操作过 Xcode。 这家公司去年还在收 $500/月的门票。 今年把 Core 计划降到了 $20/月,SWE-2 模型把平均完成步骤从 127 步砍到 53 步,成本降了 81%。 它在 FrontierCode 1.1 上拿到 50.0%,距离 Claude Fable 5.1 只差不到 1 个百分点。 然后它收购了 Windsurf,把自主 Agent 和交互式 IDE 合进了一个产品。 然后它拿了 10 亿美元融资,估值 250 亿。 然后它给自己的 Agent 配了一台 Mac。 以前你雇一个 iOS 开发者,给他一台 MacBook,等他配环境、装 Xcode、跑模拟器、提交 TestFlight。 现在你给 Devin 一个 prompt,它自己把这整条链路跑完了。 区别不在于谁写的代码更好。区别在于一个每月 $20 起步,另一个每月 $20,000 起步。

@cognition

Special delivery: Devin just got a Mac 🍎 > Now Devin can: 1. Build & test apps on its own Mac VM with iOS simulator 2. Send a screen recording via Slack 3. Send a TestFlight link so you can start using it 📲

DGX Spark

NVIDIA / 1 item

视频生成加速框架 FastVideo 最新释出 FastH3 8-Step V2,并全面打通 Mac 本地 MLX 推理链。 它不是单纯的模型搬运库,而是一套覆盖分布式微调与端到端优化的完整工作流(目前 4.4k Stars)。对于想要在本地完成高品质视频生成的开发者,这次更新直接命中了算力和硬件门槛的痛点。 核心工程进展: • 算力开销大幅压缩:新发布的 FastH3 8-Step V2 基于 MiniMax-H3 进行 DMD2 步进蒸馏,引入高达 80% 的视频稀疏注意力(Video Sparse Attention),极大降低了推理成本。 • Apple Silicon 原生支持:告别云端依赖。借助 MLX 框架与 FastMetal-QAD,Mac 用户现在可以原生运行从 1.3B 到 14B 参数的视频生成模型。 • 多端适配与实时编辑:除了主流 NVIDIA 显卡,现已支持 DGX Spark 环境(注:ARM64 架构目前暂无预编译 wheel,需从源码编译 CUDA kernel)。其内置的 Dreamverse 模块可实现本地视频流的实时“Vibe Directing”控制。 如果你习惯在 macOS 桌面上(例如 32GB 统一内存的 Mac 环境)进行模型部署与自动化测试,FastVideo 提供的清晰 CLI 与 Python API 能帮你快速跑通从代码到视频的最后一公里。 项目文档与源码:

Muse Spark

Meta / 1 item

🚨 Muse Spark 2 Leaks: Beats Astra > Spark 2 could compete with GPT-6 Astra and Fable 5.1 > Meta is already developing the next-gen Muse model > Expected to be extremely cheap to run > A 1M-token context could carry over > Meta admitted Spark 1 struggled against the competition > Spark 2 could be Meta's answer > Expected late this month or next month Could Meta finally have a serious frontier model?

OpenRouter

1 item

official site

Ox Alpha: The Most Brilliant Marketing Move In AI This Year 👀 >No logo, no branding, no PR just showed up on OpenRouter on Aug 20 as an anonymous "stealth model" >1M-token context, multimodal (text, images, video), zero data retention >Free for a week, with a claimed 100 trillion tokens/day of capacity behind it >Beat GPT-5.6-Sol and Claude Fable 5 on the DeepSWE coding benchmark >Usage blew past DeepSeek by more than 2x in days Stripe's CEO even called it "very impressive"

Llama

Meta / 1 item

Union Alpha points to ZAI's GLM family: 26 text+image probes match GLM-5.3's tokenizer; 13 image tests match GLM-5.3-Flash. Oddly, text-only counts swap between Llama/Qwen/DeepSeek-like patterns. Best guess: GLM-5.3-related variant/backend. Owner/model unconfirmed.

ByteDance

1 item

New here official site

Agent TARS brings GUI control and vision into your terminal, browser, and product — a full multimodal agent stack that completes tasks the way a human would navigate them. CLI and Web UI included. MCP ready out of the box. Github:

Also recorded

67 items that named no organisation or product this site tracks.

We all know fraudsters operate as connected networks, and the siloed tools that financial crime teams use are not enough to prevent crime. That’s why today we are introducing Neo4j GraphAware Financial Crime Intelligence - one connected, graph-native foundation for modern financial crime operations. #GraphAware Financial Crime Intelligence helps teams: ✔️ Detect beyond the obvious red flags ✔️Investigate the bigger picture ✔️Focus on the most important threats ✔️Make evidence-based, defensible decisions Get more details and watch a quick demo here: #Graphintelligence #FinancialCrimIntelligence #Neo4j

congrats to our partner @morphic_io for their official launch! we are so proud to power Morphic with GMI Cloud’s agent infrastructure. Excited to see what they build next.

@joshsum_

today we launch @morphic_io: the sovereign AI platform around one core principle. > human knowledge is our most precious asset and needs to compound for the people contributing it. > we unlock our potential only when experience does not walk out the door every time someone leaves

Research papers are filled with citations, but clicking on any of them makes you lose your place in the paper We made citations much easier to navigate, with extra information on hover including authors and institutions Now, you can hover over any in-text citation to instantly preview the referenced paper, jump directly to the full citation, and return to exactly where you left off with a single click This includes quick access to the authors and their organizations, so you can understand the people behind the research without breaking your reading flow Try it now on any paper at

StackAI launched Multi-Agent Teams! 🤖 Here's the problem we built it to solve 👇 A customer's agent works well when asked to do five tasks. At fifteen it starts getting unreliable. At thirty it's guessing. Every time you widen the scope, you expand the decision space the model searches on every single call, and every new instruction competes with the ones already there. You're not making the agent smarter… you're making the problem harder.

Proud to support this release with a ready-to-use Lightning Studio. Just open it up and start exploring the Chronos data, notebooks, and models. Get started for free

@thernabio

Today we are releasing Chronos, the data engine behind our AI RNA Biologist, RNA-Logix™. > Tens of thousands of synthetic mRNAs. ~50 human cell lines. Followed over time. One experiment. Millions of functional RNA measurements. > > #RNA #mRNA

🚨 Announcing Abacus AI Bot - 100% FREE PERSONAL AI AI can manage your calendar, book tickets, or make reservations. These simple tasks should be TOTALLY FREE Today, we announce our new product - Abacus AI bot - Finds and runs on FREE LLMs on the web - Automatic routing to free models - Does all your personal work - Just download it, set it up, and run it Use it for FREE forever

Mistral partners with Mozilla to bring its AI models to Firefox. @MistralAI models will power Firefox new Smart Window AI features, with the partnership initially rolling out across France and North America.

@MistralAI

Today, we are announcing a partnership with @mozilla to bring privacy, control and choice to people using AI to browse online. 🦊🐈 >

Agents can write code, review changes, and troubleshoot pipelines, but the handoffs between each step still take manual coordination. Join us at GitLab Transcend, virtual, to see what we're building to reduce the bottleneck between writing code and shipping it. Register free. Bonus: unlock the Enterprise AI Summit livestream, October 7-8, a $1,600 value.

Open source music is making moves, homies Yes I am mentioning this a lot Yes you should be excited

@sin_ceriously

The "hum-to-song" module training is complete. I'm open-sourcing it within the hour. Follow to make sure you don't miss the update. Special thanks to my mates at for sponsoring the compute. #YuE2 #opensource

truncated at source

🎬 The Dreamina App is now live in the US. New semester, new ways to express yourself. Whether you're moving into a new dorm, discovering who you want to be this year, or just trying to make the group chat laugh — this fall is your stage. Meet Cast — the newest feature in the Dreamina App. Create your AI Self in just a few taps, ready to step into any scene, style, or story you can imagine. ✨ Try Cast for free 👯 Invite friends to unlock more free generations (US only) Ready to take your Cast further? 🏆 Join the #CastYourselfIn Challenge — cast yourself in your own story, and show us what unbounded self-expression looks like. · Grand Prize: Top 10 most-liked eligible posts per platform (TikTok, Instagram, X) — must include the Dreamina watermark and be tagged #CastYourselfIn #DreaminaChallenge — each winner gets 30 days of Dreamina Advanced Subscription. · Deadline: Sun, Sep 20 · 12:00 PM ET / 9:00 AM PT · Full challenge rules in the comments 👇 Comment “CAST” 👇 and we’ll DM you the guide to create your A

truncated at source

SASS is not dead. Today, we are announcing the age of self-building SaaS. 🔥 🧱 A lot of people think that SaaS is going to die. I think it will evolve. Open (YC W24) is the first platform that allows you to extend it by adding new features and functionalities that did not exist in the platform before, simply by prompting. Think of it as if we brought Replit and Lovable into the best AI-native CX platform in the world. The possibilities are unlimited. Over the past two months, our customers have built and extended Open with more than 1,000 new features and over 2,000 apps. Mass adoption. They built: - Workforce management features and killed +$100K contracts - QA and real-time monitoring across all channels - Upselling dashboards and sales-oriented features - And many more. I always tell Open users: If you use Open, you have the greatest piece of technology in the CX world, and you will never look back again. With our app builder feature and your customer context, you can go above and beyond.

truncated at source

Scientific discovery often requires changing the question itself. Today, PhAI Labs releases the technical report for Discovery Foundation Models (DFM), a research direction for AI systems that can identify valuable unknowns, formulate questions, develop hypotheses, design experiments, and revise their understanding as evidence comes in. The goal is not just to complete one investigation. What a system learns should carry into the next—as reusable methods, accumulated research experience, and better decisions about what to explore. The report introduces a reference architecture, the science infrastructure we're developing, and a wet-lab case study. 🔬 Alongside the report, we're opening the DFM Scientist Collaboration Program. Have a scientific question you'd like to explore with AI? We're inviting scientists, research teams, and experimental platforms to collaborate on meaningful open problems with real opportunities for validation. 🙌 Program details & application: 📅 O

Introducing Command Code Desktop App! ⌘ Available today in beta for Mac, Linux, & Windows. Building the most powerful agentic app, starting with code. Multiple agents, files, Git, terminal, browser previews, plans, and instant edits with /design. This is just the beginning.

An agent that works in a demo isn’t automatically ready for production. Qoder Cloud Agents 1.0 is here — a managed cloud harness to build, ship, and operate agents as services. Get an end-to-end integration running in 10 minutes. Scale without rebuilding.

We released our AG-UI DevTools Chrome extension. If you are already working on Agentic UI with AG-UI (via @CopilotKit), this makes your life much easier. If not, now's the best time to try it 😁 Kudos to @wolfmanfx (@SoveriusAI) for the implementation.

Introducing Retrieve-for-Train: a framework that accelerates complex AI search by replacing heavy autoregressive inference with a lightweight diffusion model. This allows for instant, expert-level search slates at a fraction of the cost. Learn more:

Same fight, now with video depth. Topview video depth feature is live, giving you more control over video generation and more ways to explore different styles. What would you try it on?#TopviewAI #TopviewCanvas #AIVideo #AIFilmmaking #depthmap

Open source models are notoriously bad with production workloads They are unreliable, stall on tool calls and hang for 5-10 minutes! We fixed these issues with Smaug Flash and are able to offer robust personal agents almost for FREE

We’re grateful to OpenBMB for supporting Atria Dawn Preview with UltraData-SFT-Agent-2609. The dataset’s high-quality agent trajectories have been valuable in developing Atria’s capabilities on agentic tasks. 🔗

amazing demo where i can literally open my live transcription software and it immediately falls apart it literally only got their product name wrong what are these companies smoking

@aidaxbaradari

Today, we’re releasing Kalypta, the first app to block AI notetakers in your meetings. > Granola? Wisprflow? Cluely? No more. > With Kalypta, you become inaudible to AI. > Your call continues normally.

Most AI benchmarks test retrieval: can a model find an answer we already know? But science’s hardest problems demand discovery: can a system earn an answer no one knows yet? Meet TRACES 🧭—the world’s first benchmark for discoverative AI: systems that can weigh evidence, test hypotheses, and reach verifiable conclusions on problems with no answer key. Proposed by our founder, @tianqiao_chen, TRACES measures six defining capabilities of discoverative intelligence. Recently, we published: · A rigorous definition of “discoverative intelligence” · A rubric that distinguishes sound investigation from lucky guesses · An open call for both problems and solvers · Website:

We’ve done a ton of performance work to supports a ton of threads running hundreds of subagents at once. I run 30-300 agents at once on my Mac and these performance improvements are crucial to not completely blow up my system More to come!

@sethkarten

prime agent v0.9.5 > we fixed a lot of bugs and, of course, we had prime agent feature its favorite updates > it picked our perf work. then it created the video itself.

Native AI agent skill management has landed in the Dart ecosystem. Skills CLI 1.0 brings: - Official Dart SDK tool (no npx) - Auto-detects package skills - Easily fetch external Git repos - Auto-discover AI agent instructions Learn how to supercharge your AI workflows:

@FlutterDev

Tired of using external tools to manage AI agent skills? Meet the Skills CLI 🌟 > Now you can discover, install, and manage reusable AI "skills" directly from your Dart and Flutter dependencies using a single native package manager command. > Details →

@karenxcheng built something that is genuinely mindblowing. An AI agent that creates and prints a personalized newspaper every night while she sleeps. She calls it "The Karen Times." Every morning instead of reaching for her phone, she picks up a printed paper. Her schedule. Weather. To-do list. Inbox highlights. Curated news. A personalized comic strip starring her. A mini crossword. All generated fresh overnight from her real calendar, emails, and subscriptions. Built with Grokbot. Told it when she wakes up. The agent starts an hour before, builds the full paper, sends it to her printer. She wakes up. Paper is waiting. Phone stays on the table. She made the whole thing available for anyone to try.

豆包大模型 2.1 Pro 发布 0915 版本更新,API 已在火山方舟全量上线。这次升级集中在四个方向:Agent 任务交付、多模态写代码、多模态理解、以及推理成本。 Agent 方面的改进 模型在需要多轮调用工具、联网查资料再出报告的场景里,强化了证据溯源和数据核验能力,幻觉明显减少。官方给的例子是金融投研:模型能自主拆解研究需求、检索数据源再建模分析,产出的底稿接近分析师水平。为了核实某车企财报里的一个说法,模型调度了 500 多个子 Agent、检索超 1000 个网页,还交叉比对了海事航迹、卫星影像等多源信息。这类"不采信单方通稿、多源交叉验证"的能力,对企业做尽调、出研报有直接价值。 多模态 Coding 这次最实用的变化可能是看图写代码。模型现在能直接读懂设计稿、图纸甚至操作录屏,把视觉信息转成前端代码。官方演示了一个场景:给模型一段录屏加几张草图,让它给一个没有文档的老 ERP 系统开发移动端页面——模型读懂了 28 万行 Java 代码,直接还原出可运行的移动端页面。 在代码仓库理解上,模型对开源游戏 Luanti(约 38.7 万行代码)做了自主修复测试,83% 的任务达到了可合并标准。对经常要在大型项目里定位问题、跨文件修 bug 的开发者来说,这个数字值得关注。 其他升级 多模态理解方面,视频推理能力增强,能在视频里定位证据、跨帧整合;图像理解在 3D 物体识别(CAD 工件、游戏引擎元素)和密集图文解析(工程图纸、财报表格)上有明显提升。 成本方面,图像和视频推理的 Token 消耗比上一代减少 30% 以上。 API 使用上有两个入口:调用 Doubao-Seed-2.1-pro-0915 可以锁定版本;调用 Doubao-Seed-Evolving 则自动跟进最新版本,不用换 Model ID。豆包工作和 TRAE 也已同步接入。 官方文章:

语音 AI 开始往手机里走了。 开源 Audio8,把语音识别和语音合成做成一组可在手机、PC 上本地运行的模型。 其中还有 iPhone 离线转写版本。云端之外,端侧也越来越能用了。 @LeonaYangAGI 干得漂亮👍

@LeonaYangAGI

Over the past two months we open-sourced a set of on-device audio models. The name is Audio8. The series can run on phones, PCs, and other hardware with limited compute and memory. You can use them for local inference right away. > What’s included: ASR: 0.1B, 0.3B, 0.6B, 3B TTS: 0.1B, 0.3B, 0.6B

At this important juncture, building better evaluations across a diverse set of domains is the most impactful way to align models. We need subject matter experts across cyber, biology, chemistry, and many other industries to build benchmarks that help pace the frontier.

@mercor

Mercor is committing $5M to a new AI Safety Fund. > One of the biggest challenges the industry faces today is addressing whether frontier AI is safe enough to deploy. This is why we believe investing today in safety research, evals, and verification is critical. > We want to work with AI researchers who are focused on critical safety risks, including: - Misalignment: deceptive alignment, reward hacking, scheming - Sandbox escape and agent containment failures - Evaluation awareness: models that behave differently when tested - Interpretability and scalable oversight - Red-teaming methodology and safety eval design > We will fund the researcher hours, API credits, and travel.

NVDA dumps 4% on a rumor and I’m already refreshing charts like an idiot so I just told Stingray what I think happens next it pulled similar setups from the past, let me define the conditions for my thesis, and kept watching it for me haven’t opened TradingView in an hour. this is actually useful

@Stingray_fi

Every trader needs a partner. > Meet Stingray, your AI market analyst. Text it your take. It tests the idea against live data and finds the trade. Then it alerts you when it’s time to move. > Message Stingray:

DEMO: @Siemens just showed what happens when agents handle partner onboarding end to end @tryqualified’s AI SDR Piper turns website interest into qualified pipeline. Then Agentforce Supply Chain keeps onboarding moving: → Coordinates work across agents + people → Brings in humans when judgment is needed → Learns the backend process → Turns what it learns into trusted actions → Executes the final steps hands-free in @SAP The key: the agent doesn’t improvise. It follows the same trusted process every time. Weeks of onboarding gets done in days

they just killed THOUSANDS of startups... every one of them sells the same thing: search your transcripts, call it "understanding your customers" that's over... because these guys spent the last 7 years building a different category entirely: > a model built from scratch to read customer conversations, not a wrapper on someone else's tech > 100% retrieval... nobody else has pulled this off > it gets smarter with every call, ticket and account it reads, no retraining needed > ask "which accounts are about to cancel" and it doesn't guess... it answers with names, dollar amounts and the exact call it pulled from seven years of work just made every other AI tool look like a demo (and it gets sharper every day it runs)

@Spshulem

AI is killing your company. > It should be making you more revenue. > Introducing BuildBetter: the first AI Head of Product. > Here’s how it works 👇

Your agent forgets? A graph can help. @ neo4j-labs/nams-ai-provider: One changed line of code gives any Vercel AI SDK model persistent, cross-session memory backed by a Neo4j knowledge graph. No Neo4j cluster to provision, no vector store to pick, no embedding pipeline to babysit. A free API key and a model swap. What it does and why we built it on the Vercel AI SDK? This and how it works in one blog from Prakriti Solankey

Holy fuck they're shameless

@MultiverseCompu

⚛️🇪🇺 Introducing Quasar 1.1 438B, the first AI model using quantum-generated data. > Quasar 438B, the best European AI model, just got a quantum upgrade: Part of its healing set was generated by a hybrid quantum language model running on @IBM Quantum System Two in Donostia-San Sebastián, a 156-qubit IBM Heron processor. It is the first time quantum-generated data enters the CompactifAI pipeline. Quantum computing is not a label on this release, it is part of how the model was built. > Sharper, less verbose, tuned for agentic work and ready for European deployment: served by a European company incorporated under EU law and developed in line with EU regulatory requirements, including the transparency expectations of the EU AI Act.

Creators have been leaving the real money on the table. You sell the course on what you know. But the thing people actually want is the work done, not another lesson to sit through. Twin Store lets you sell the agent that does it. Your knowledge, packaged as a product with a storefront, pricing, and a dashboard your buyers just log into and run. No prompting, no setup on their end. Stop selling what you know. Start selling what it does. This is the upgrade every creator has been waiting for.

@hugomercierooo

100k people have built agents on Twin. Starting today, anyone can sell one. > Introducing Twin Store. > In under 5 minutes, your agent gets a storefront, an app, and Stripe billing. > Build once. Twin finds you customers. Get paid monthly. > A new wave of wealth is coming, and it belongs to agent creators.

I spent 2015-2023 building self-driving cars. It was so painful. It's now 2026, and we tasked just 1 researcher to take Odyssey-3, our latest world model, and teach it to drive the roads of India with just 20 hours of driving data. And…it works! This is insane.

@odysseyml

Today we’re unveiling Odyssey-3, a big step forward for foundation world models. > It can control robots, power humanoids, drive cars (on the roads of India!), train AIs, pilot drones, and even play video games. > We can’t wait to see what intelligent systems it enables.

Creative professionals just ranked every major AI model blindly and the results broke the narrative. There's no single best model. The one that wins is the one built for what you're actually making: film, ads, animation, or motion design.

@openart_ai

Today we’re launching OpenArt Arena, the global leaderboard for creative intelligence. > Every AI model claims to be the best. But best at what? > We brought together creative professionals to evaluate the world’s leading AI models blindly, across the work creators actually care about: > Film. Ads. Animation. Motion design. Video editing. And more. > Because the real question isn’t “Which model is best?” > It’s “Which model is best for what I want to create?”

Ngl, thats really huge: Poolday reports 700,000 agent runs and 100 million video edits. Its agent takes a video brief, works with your existing footage and brand assets, and assembles a finished cut. It can also use generative models when the video needs something new. Say you're making a product update. You can use your real UI footage and actual Figma components, then have the agent put the video together in your brand's style. And you still get to change your mind. The output stays editable, layer by layer. Swap the logo, fix a caption, keep the rest of the video.

@alexeichemenda

We raised $11M to make video editors obsolete. Not the humans, the tools. > Think about the hours you lose searching for assets, moving clips frame by frame, fixing animations, checking exports, and redoing the same edits over and over. > Poolday handles the entire video production process, start to finish, in one prompt. > It learns your style. Uses your assets. Makes

This anime battle started as a simple storyboard. Minutes later, it became a cinematic HD sequence with fluid combat, dynamic camera movement, and stunning visuals—all powered by APOB. No animation studio. No lengthy production. Just AI Influencer ideas into anime at incredible speed. The next era of anime creation is here. ⚔️✨ Try it free 👇 👉

i've been working with large enterprises on improving search in their very connector heavy workspace this is a baby step which has really improved answer quality notion gives you enterprise search bundled in, try it out :)

@NotionHQ

New: Set default agent instructions for your entire workspace. > House rules for your agents.

架空の商品に、どんな広告をつくる? AZ8 × Wan「Ad from Zero」が開催中です。自由に考えた商品を、Wan3.0で一本のAI広告に。賞金総額3,000米ドル、制作クレジットも用意されています。 応募は9月24日まで。あなたのアイデアを、広告にしてください ↓

@AZ8official_JP

架空の商品に、あなたならどんな広告をつくる? AZ8 × Wan「Ad from Zero」開幕!✨ 自由に考えた商品を、AZ8でオリジナルのAI広告に。 🏆 賞金総額3,000米ドル 🎨 制作支援:350万AZ8クレジット 📅 9/10〜9/24 23:59(JST)

been building the infra for ai agents to be productive time for them own productive assets and exert real-world influence board sit, by ai agents

@virtuals_io

The next BlackRock isn't a fund. It's an AI Senate, funded and chosen by the trenches. > Meet Occupy by Virtuals Protocol. Live on @base. >

truncated at source

Generate SUNO-quality songs on your PC with just 4.5GB VRAM WanGP has added YuE2, the latest song generation model that's supposed to be as good as SUNO 5. And it just needs 4.5GB VRAM. It does really feel better. Listen for yourself, a song about spicy japanese curry:

@deepbeepmeep

WanGP v13.00 — It’s Your Lucky Day! 🚀 🎨 A faster, refreshed UI with 5 themes and voice dictation for prompts. 🌍 Start a generation on your PC, follow progress and manage jobs from your phone or another PC. 📊 See what’s happening at every stage, with cancellation during preparation too. 📁 Organize media into project workspaces that stay saved across restarts. 📱 Deepy now has a mobile web app: chat, upload photos or recordings, and follow your results from your phone. Add it to your home screen for quick access. 🎵 YuE2 turns your lyrics and musical style into full songs with vocals. 🎙️ AuK generates speech, clones voices, edits spoken words and cleans up recordings. https:/

Building a voice agent is a loop: write, deploy, listen to calls, fix what broke. The agent is the unit of work, so we rebuilt LiveKit Cloud's around it. Talk to your agent while you're editing it. Click a metric, see the sessions behind it, go back without retracing. Try it now >

One walk. 24 frames. 24 different outfit I dropped one clip into invideo Editor and asked it to keep my movement the same while changing the outfit frame by frame. Then I could still tweak the timing, replace frames and refine everything directly on the editable timeline.

Are you using Supabase + Flutter? V3 of the supabase_flutter SDK is now in beta and we would love feedback. The new v3 version has full type safety and lots of other goodies. 😎 If you're interested in participating in a feedback group you'll also get some swag. 😏

In case you missed it - we have support for WebMCP in Chrome DevTools! 🎉 Find it inside the Application panel under "WebMCP" to easily debug your WebMCP tools! Let us know what is missing and what should be improved! #WebMCP #ChromeDevTools #AIAgents #WebDev

AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video TL;DR: Models the world as a 360° panoramic latent state, then renders and refines only the user’s current perspective—enabling efficient, long-horizon interactive video generation.

Your daily intelligence report should be ready before you ask for it. Routines collect and organize information on a schedule, then invoke Agent to identify the changes that matter. Fresh insight without spending tokens on every step. Give Routines a try:

👀

@thaiscbranco_

We're helping agents create things worth making. > Introducing: The Brand API. Your agent can extract any brand system, search for style inspiration, and verify if it’s staying on course. > Try it for free at:

YuE2 model 😃🎶🎵support for ComfyUI Compose songs from lyrics and style, build or import an ABC score, and render 48 kHz stereo audio with YuE2-3B. Generate music and train AR LoRAs with reviewed datasets and checkpoint previews. 👇

Meet Meridian by Viggle AI: a video model that lets you change the camera angle and timing of an existing video. Follow the action from a new angle. Slow it down. Or freeze the moment and keep the camera moving. Open weights:

JustDraw updated for iOS 27! Flux 2 Klein 4B converted to latest version of Core AI! It works, locally on iPad M5, offline and quite fast 12 seconds per image! Time to clean code and publish repo and article soon!

Very ambitious and very cool. Will definitely try it.

@zach_yadegari

Everyone’s been asking what’s next after selling Cal AI. > Introducing Persona. > Others couldn’t figure out the interface, we did. >

Secure Compute and Static IP builds now start 64% faster. Builds used to boot a fresh container with the network attached. They now claim a prewarmed one and attach it on the way in.

🐟️ Sakana Marlinがアップデート 🐟️ ブログ: Sakana Marlinを試す: 本日、Sakana AIはUltra Deep Research Assistant「Sakana Marlin」をアップデートし、二つの機能を実装しました。 ・レポートと対話する「Interactive Reading」 ・出力スライドのPowerPoint対応

Director now lets you add in audio and image references continually into your scene

@weberwongwong

direct a scene in realtime and with image references in our Director tool

In addition to YuE2, we have 2 other new AI music models, DiffSynth Music and MuLaCover Both can take in reference audio and generate covers. IT NEVER ENDS

Semantica ingests enterprise data to build context and knowledge graphs, enabling graph analytics and causal reasoning with full decision provenance.

Union Alpha (stealth model) is free for the next week - no data training - built for agentic coding - supports images let's see what you can do

可以卸载掉Typeless了!! 朋友把他的Typefree免费开源了,我试用过,完全可以平替Typeless,大家只需要自己去配置api就行,体验很不错。作者:@BitcoinRui 还加了新的鼠标交互方式,大多数时间可以不用键盘,鼠标就可以搞定!! 开源地址:

Audits on is-agentic.⁠com now match site type. Docs, business, app, and commerce sites get a view of the checks that matter most.

即览 1.01 已经更新了! 如果你之前在使用中遇到了一些问题,可以更新一下。新版本进行了大量的体验优化和视觉样式的重构。 现在更好搜索了,基本上搜名字就能搜出来

@op7418

居然在美区效率榜单排 44,而且美区是不限免的,谢谢各位❤️

Coordinates AI agent teams in cross-platform chats where agents hold personas and collaborate alongside humans.

DS41RT v4 Mostly kernel updates Thanks @YourLocalAILab +5-20% real use case decode Raw details are complicated

Manages your Obsidian vault with a crew of 8+ AI agents and 14 skills that organize, search, and triage email.

Sarvam is now testing Bulbul v4 'Flash'!