AI briefing
18 September 2026
209 items were recorded on 18 September 2026, filed under 46 organisations and products.
It was the busiest day in the current 8-day window.
Most covered: Codex (17), Gemini (16) and OpenAI (16).
51 of the day's items named no organisation or product this site tracks; they are listed under “Also recorded”.
19 items reached the source feed's 1024-character limit and are cut off mid-text; each is marked “truncated at source”.
Compiled by Bloger.fm Editorial Desk
Compiled from a monitored feed of public AI announcements. Items are quoted or summarised as recorded and are not independently verified — see the editorial policy.
Gemini
Google / 16 items
official siteBig updates are rolling out to @Gemini_Notebook to help with teaching and learning. Check out the new features making prep easier for educators and students ⬇️ What feature are you trying first?
We’re rolling out new updates to @Gemini_Notebook for a more personalized learning experience 📚✨ > What’s new and coming: > 🎙️ Chat with your notes in real-time across nearly 100 languages > 📱 Record lectures and capture your thoughts on the go in the mobile app > 📝 Build interactive study guides and customizable quizzes > 🎬 Create 60-second Short Video Overviews you can share with classmates > Plus: Eligible college students can also claim one year of Google AI, on us.
Gemini 4 Pro apparently in LM Arena right now under "Gemini 3.8 Flash", yes really. That could mean that release is not far off.
@ai_for_successRumor has it that the Gemini 4 Pro checkpoint is now available in Arena, and the output which I have seen is seriously impressive. Looks like Google is back 🔥 > Here are some of the best rumored Gemini 4 outputs I’ve seen so far. 🧵
with all the leaks of Gemini 4 Pro my expectations from Gemini have increased drastically, hope they don't disappoint us when it's finally launched.. also i haven't forgotten what they did with Gemini 3.5 Pro 🥀
One of the clearest examples of what model specialization can buy you. At 2.6B and 1.2B parameters, these models are also small enough to run locally on mobile devices.
@liquidaiIn a new article published today on the cover of @CellCellPress, we obtained Liquid Foundation Model instances that establish state-of-the-art performance on biological longevity tasks, outperforming the best frontier models such as Gemini-3.1-Pro, GPT-5, and Claude Opus. > In partnership with @InSilicoMeds, we built and released: > A comprehensive eval suite of 17 biological longevity tasks (i.e., LongevityBench), to assess whether a general-purpose language model can interpret aging data spanning clinical records, DNA methylation, transcriptomics, plasma proteomics, and genetic evidence. > LFM2-1.2B-Longevity and LFM2-2.6B-Longevity: two compact models specialized for interpreting structured aging data across these tasks. > These results are important! 🧵
Gemini 4 looks much better here
@YouWareAIWow~ Gemini 4 is live on Arena ,and it hits hard. > I ran the same prompt against GPT-6 and put the side-by-side in this clip. > Gemini 4’s build looks ridiculously strong to me. You? Ready to go toe-to-toe with GPT-6 yet? > Play both on YouWare. Full prompt + links in the replies.👇🏻
Seriously, are people sleeping on Gemini for coding? Gemini 3.8 Flash is putting up some surprisingly strong coding numbers, and it’s still priced like a Flash model I’m starting to wonder why I don’t see more devs talking about it 👀
+ 10 more items − collapse
Dropbox 🤝 @GeminiApp Use the Dropbox app for Gemini chat and Gemini Spark to work with your Dropbox files, create shareable links, and use your content to create Google Slides, Google Sheets, Gmail drafts, and more.
Gemini 4 (flash or pro) which is currently testing in Arena it is best at 3d generations currently it is testing under the name "gemini-3.7-flash"
@HarshithLucky3Gemini 4 looks much better here
if the new Gemini model is Gemini 4 Pro then what happened with Gemini 3.9 Flash? they do have a bad habit of skipping model versions so you never know what do you think this new model can be?
Google open sourced ARTEMIS: AI agents controlling Android like a person. Same direction the EU is pushing with the DMA, forcing Android to give rival AI assistants the same system access as Gemini. Apple is fighting that battle over Siri. Either way, the OS stops being a moat.
Just saw an insane Gemini 4 family benchmark page 4 Flash mogging Fable 5.1 looks like slop, so I won't repost I expect a lot of wild rumors
Gemini Worlds
@aileaksofficialGemini worlds. More info soon.
You can run locally Opus 4.6 level model Qwen3.8-27B with 8GB of VRAM At home.
@0xSeroQwen3.8-27B now runs on 8GB of VRAM! > That's less than 500$ to run a model smarter and more capable than: > 1. GPT-5.6-Luna High 2. Opus-4.6-Max 3. Gemini-3.1-Pro > And ties with: > 1. GLM-5.2 Max 2. Gemini-3.6-Flash > PrismML has the mandate
🚨 Gemini 4.0 Pro specs leaked • Pricing: $2.25/$11.25 • 2M token context window • Beats GPT-6 Astra and Claude Fable 5.1 on almost every reported benchmark • While being significantly cheaper than both
Haha! The rumors about Gemini 4.0 are phenomenal They are aiming to surpass both Astra and Fable 5.1 and are on track to do so The pricing will be super competitive as well.
🚨 Fable 5.2 is tested in claude code routing fable 5.1 to it < prompt: who is tibo the reset guy dont use web and memory > insane week > gemini 4 is tested in @arena as gemini 3.8 flash > grok 4.7 is ready > opus 5.2 is ready < already show two demos> > gpt-6-sol is ready
@chetaslua🚨 sol 5.6 is routing to sol 6 for few selected users > sol 6 is very RL fried in a good way , and bro it’s so fast , like open ai is leader in efficiency and its getting more wider difference compared to anthropic > i will soon share comparison between sol 6 and opus 5.2/1
OpenAI
16 items
official siteTechHalla just posted an AI video recreating a Mentos and Coke geyser on a highway today. Here's what you need to know. The clip shows a fake Mentos-and-Coke eruption made entirely with AI tools, shared as a try-this-at-home joke, with the actual prompts included. The creator says the video came from GPT Image 2.5 and Seedance 2.5, two AI models used together to generate the imagery and then animate it into a moving scene. GPT Image 2.5 is OpenAI's image model, released September 8, 2026, improving image quality, generation speed, reference fidelity, and editing precision. Seedance 2.5 is a video generator that produces native 30-second clips in 1080p with audio, supports up to 50 reference inputs, and adds camera movement and scene control. Key numbers: - Posted September 17, 2026, at 5:05 PM - 2.8M views - 10K likes - 3.4K bookmarks - 30-second native video length from Seedance 2.5 - GPT Image 2.5 released September 8, 2026 The post is labeled Made with AI, and replies show viewers reacting with a mix
SITUATION DETECTED: OpenAI launched Astra for Law: GPT-6 Astra with a U.S. legal research index and instructions for legal analysis and writing.
AI won’t replace lawyers but it will replace the lawyers still thinking AI won’t replace lawyers
@OpenAIAstra for Law: Frontier intelligence built for your practice. > A new offering powered by GPT-6 Astra with tools, settings, and context to support the expertise and judgment of lawyers and legal technology firms.
OPENAI JUST WENT AFTER A SECOND MILLENNIUM PRIZE PROBLEM AND MATHEMATICIANS ARE LOSING THEIR MINDS The Hodge Conjecture is one of seven Millennium Prize problems, and the Clay Institute pays $1 million for each. -> According to The Information, OpenAI staff expect it to be solved relatively soon, two weeks after Navier-Stokes. Before this year only one had been closed in 26 years -- Grigori Perelman proved Poincaré and turned down the money. The conjecture links topology, which counts holes in multidimensional spaces, with algebraic geometry, which describes shapes through equations. It says every such hole has a structure made of equations behind it. Officially OpenAI only admits substantial progress on another problem, without naming it. After Navier-Stokes, more than 1,000 mathematicians signed an open letter accusing the company of research misconduct. 25 Fields Medal winners issued a separate statement, and OpenAI pulled its Caltech hackathon sponsorship. Per the source, OpenAI is now focused on
OPENAI 🔥: ChatGPT in Word has been announced to be available across all plans, including Free, with usage limits. > Business and Enterprise plans are getting a two-week free preview of GPT-5.6 Sol in Word. > Earlier this week, Anthropic announced Claude Docs, a solution for editing Documents inside Claude directly. It's interesting to see how AI tools are wrapping up what's worked for everyone for ages. AI is going inside these tools and absorbing them too. And it is so powerful that there is no way to avoid it. AI will wrap the whole internet 👀
Is it even possible that a traditional LLM beats it in speed? I hope OpenAI releases ultrafast, computer use with Astra's or Sol's intelligence + speed of Jev would be crazy. For me it would be computer use AGI
@trycua1/ Fast Computer Use is now solved with @typesafeai Jev + Cua Driver. > Available in development preview for macOS, Windows, and Linux. We call it jev-use. > Draft #3943:
+ 10 more items − collapse
OpenAIのAIエージェント「Codex」のデスクトップアプリで、Codexのより詳細な利用状況を確認できるようになりました(設定→使用状況と請求→利用状況分析)。
@OpenAIDevsGet more visibility into your usage. > See how tasks, subagents, and individual chats contribute to your Codex usage, so you can make more informed choices about your workflow.
OpenAI GPT-6 Astra actually does have observable chain of thought, according to OpenAI Research Scientist Noam Brown. “Chain of thought monitoring… it’s a real gift. We were very lucky that this ever existed, and it is fragile.” 🚀 Watch the full episode of AI Deep Dive:
OpenAI may be about to crack another $1M Millennium Problem. After Navier–Stokes, sources say the Hodge Conjecture is next. They used an unreleased model variant, “Doug,” on the first one. The bet: if AI can do frontier math, it can start inventing better AI.
GPT-Live 1 is on AI Gateway. • Full duplex audio model: listens while speaking, so you can interrupt. • Client delegation: run any text model in the background. 𝚘𝚙𝚎𝚗𝚊𝚒/𝚐𝚙𝚝-𝚕𝚒𝚟𝚎-𝟷
Starting today, you can connect multiple accounts with most plugins in ChatGPT! 🎉 Bring context from your work and personal accounts into the same conversation. Connect your accounts in the plugin directory: Devs: this works automatically, with no changes needed. But, to make the experience for users of your plugin even better, add a profile tool to your MCP server so ChatGPT can label the accounts:
OK, this is a big deal: 3 researchers used Claude Opus 5 to turn an image upload bug into an OpenAI employee account takeover, then had the compromised employee’s Codex open a PR in OpenAI’s internal monorepo. Their entire hacking cost less than $3000 in tokens. Opus 4.8 struggled with the exploit. Then Opus 5 dropped and cracked it within hours. AI-powered cyberattacks are becoming common and cheap. The best defense is to put the best AI in the hands of defenders too.
cool if true
@mark_kOpenAI is close to releasing their answer to Grok Bot: a product I'll tentatively call Codex Bot, based on OpenClaw. This is what OpenClaw founder Peter Steinberger worked on after being hired by @OpenAI. > Release was planned for this week, but was postponed to next week instead.
Tibo says what we get at DevDay is enough to make someone switch back to Codex. Its probably: - GPT-6 Sol/Luna (this is coming earlier) - OpenAI’s personal assistant, codename “Aeon”. - more! Cant wait for it so much
@thsottiaux@vlinx_soft See you at DevDay
DAILY AI BRIEF 🗞 — Sept 18 ANTHROPIC 🔥: - Projects now start from one Claude Code conversation. Claude spins parallel cloud threads, keeps shared memory, and surfaces an Overview panel. XAI 🔥: - Grok Bot voice is live. Desktop and mobile, rolling out over the next couple of days. META 🔥: - Muse for Mac is out, US only. Computer use across apps, files, calendar, notes, and messages. You pick what it can access. PERPLEXITY 🔥: - Effort selector is live in Computer on web. Presets pair the orchestrator model with reasoning depth. Mobile and desktop next. OPENAI 🔥: - Astra for Law is out: GPT-6 Astra plus a Legal Search Index over 230M+ URLs. Trusted Access first, API soon. - ChatGPT in Word hits all plans including Free, with usage limits. Business and Enterprise get a two-week GPT-5.6 Sol preview. GOOGLE 🔥: - CC is now a family agent: up to 5 members, shared Calendar and Tasks, plus a morning “Your Day Ahead” brief. Waitlist, US 18+. ALIBABA 🔥: - Qwen3.8-Omni-Flash is out. First omni-modal agent model, 1M
🚀 Codex CLI 0.155.0 is out! 🎙 Experimental /voice with live transcripts and mic controls 🧠 Live reasoning summaries and turn timestamps in status row 🔐 Touch ID for MCP requests on Mac Changelog:
Codex
OpenAI / 17 items
Busiest dayAppshots: my favourite Codex feature is now on Windows too! They go beyond just a screenshot and grab all the other relevant app metadata and text too. Appshot away with ⌘ + ⌘ on Mac or Alt + Alt on windows
@ChatGPTOne of the most underrated features in the ChatGPT desktop app: Appshots. > Appshots take the context on your screen and bring it into the desktop app, allowing you to share instead of describe. > You’ll be surprised at how useful it is. > To get started, press both Command keys on macOS, or both alt keys on Windows.
口喷剪辑时代来临,牛逼👍
@gengdaJ把我珍藏已久的好东西分享给家人们,剪映11.4.2开源!!!🥳🥳🥳 > 配套神级Skill: > 只需要简单提示词,口播类视频,全程托管给Codex,不需要任何人类操作。 > 即使略微失误,由于调用豆包ASR,片段被切割到毫秒级别,导入剪映草稿也可以人工丝滑调整😋 > 放一个原本8分钟,剪辑后3分钟,完全由Codex+yichen-jianying-edit Skill自动剪辑完成的,前后视频对照,以及提示词截图👇 > 整个剪辑流程可能需要花费的地方: 1.Codex Token费用,可以用别的便宜模型平替,比如Workbuddy免费的hy3和DeepSeek-V4.1-Flash。 2.豆包ASR,一小时八毛钱,说实话,不能再便宜了。。。
You handle the reps. Codex handles the repo. Talk to the voice agent in Codex, powered by GPT-Live-1.
@cdngdevyou can use codex voice from your phone now, connecting to your computer from anywhere!! > @axbehr and i got to star in this new codex ad, showing how much you can get done while working out. > btw, this is powered by the new gpt-live-1
OpenAIのAIエージェント「Codex」のデスクトップアプリで、Codexのより詳細な利用状況を確認できるようになりました(設定→使用状況と請求→利用状況分析)。
@OpenAIDevsGet more visibility into your usage. > See how tasks, subagents, and individual chats contribute to your Codex usage, so you can make more informed choices about your workflow.
Modern Web Guidance v0.0.187 brings new and updated guides and Codex plugin support → ✨ New guides: progress rings, scrollspy nav menus, loading spinners 📝 Updated guidance to avoid bfcache breakage and unwanted inline layout gaps 🔌 Added plugin manifest support for Codex
跟 Claude Code 说「这个项目用 pnpm」,隔天新开会话它又拿 npm 装一遍,同一句纠正一周说三次,说到最后自己都烦。 于是找到 claude-reflect 是个 Claude Code 插件,专治 Agent 失忆问题。会话里每次纠正它、夸它做对了、或者说一句「记住:」,钩子都会自动捕捉下来排进队列。 跑一下 /reflect,它把队列里这批列成一张表,一条条给我们过,采纳、改一改再采纳、或者跳过,确认的才写进 CLAUDE.md 文件,以后每个会话都带着。 GitHub: 写的目标不止全局那份,项目里的 CLAUDE.md、子目录的、还有 AGENTS.md 都认,用 Codex、Cursor 的也能吃到同一份纠正。 v2 加了个 /reflect-skills,回头翻过去两周的会话记录,发现「看下我今天的效率」这类话反复问了十几次,就建议做成一条命令,草稿直接生成。 纠正用中文说也能识别,靠一层 AI 语义过滤兜底,英文关键词没匹配上也不会漏。每条带置信度,写进去前都要过人工这关,不会偷偷往配置里塞东西。
+ 11 more items − collapse
录完视频,剩下的多机位导播、加字幕和跨平台发布,现在可以全部丢给 Claude Code 处理了。 VibeTube 是一个开源的 macOS 录屏工具。它的核心逻辑非常纯粹:“你只管录,AI 负责剪辑和发布”。工具会将同步好的屏幕和摄像头素材,直接交给本地运行的 Claude Code 或 Codex 进行自动化后期。 • AI 自动导播:无需手动打轴,AI 代理会根据你的讲解内容,自动完成镜头选择与机位切换(全尺寸人像 / 屏幕录制 / 画中画),并配上字幕与音效。 • 内置影音增强:集成 NVIDIA Studio Voice NIM (48k-hq) 消除房间混响与底噪;利用 MatAnyone2 (Apple Silicon) 直接在本地完成背景替换。 • 零干预发布:AI 会读取最终的成品字幕,生成 5 种不同视角的备选标题以及带真实时间戳的 YouTube 章节描述,最后直接推送至 YouTube、TikTok 和 Reels,全程无需打开浏览器。 适用限制与门槛: 目前仅支持 macOS 环境。需本地安装 Node.js 22+ 与 ffmpeg,并自备对应的 CLI 工具与 API 密钥(Claude/Codex、NVIDIA Studio Voice 及 Upload-Post)。
OK, this is a big deal: 3 researchers used Claude Opus 5 to turn an image upload bug into an OpenAI employee account takeover, then had the compromised employee’s Codex open a PR in OpenAI’s internal monorepo. Their entire hacking cost less than $3000 in tokens. Opus 4.8 struggled with the exploit. Then Opus 5 dropped and cracked it within hours. AI-powered cyberattacks are becoming common and cheap. The best defense is to put the best AI in the hands of defenders too.
Okay, Codex Spark had a good run. But be honest, we are all thinking the same thing now. GPT-6 Sol Spark. Or somehow, Astra Spark. That is the Spark I would come back for.
@thsottiauxNext week we’ll be retiring GPT-5.3-Codex-Spark. Can you believe we shipped a model named as such!! > It's had a good run and was a lot of fun, but usage has been declining and we have significantly better models now. Time to make room for the future.
Cursor launched this Then Claude Codex when?
@bchernyProjects are how I write a lot of my code these days. Really excited for everyone to try the new experience! Rolling out now
This is insane! Codex Remote Control for Apple Watch:
让 AI Agent 操作浏览器,难点不在“能不能点”。 更关键的是,如何复用已经登录的真实会话,又不打断你正在使用的窗口。BrowserSkill 在 Agent 与浏览器之间加入本地桥接:Agent 调用 bsk CLI,本地 daemon 把任务交给扩展,再在独立 Agent Window 中执行。需要时也可借用现有标签页,是否允许借用和请求人工协助,都由扩展设置控制。 它可接入 Cursor、Claude Code、Codex、OpenClaw 等能执行 shell 的 Agent,还提供 DeepSeek Harness 插件、远程浏览器配对和可重复的能力评测。适合想把真实登录态、浏览器自动化与 Agent 工作流接起来,同时保留交互边界的开发者。
cool if true
@mark_kOpenAI is close to releasing their answer to Grok Bot: a product I'll tentatively call Codex Bot, based on OpenClaw. This is what OpenClaw founder Peter Steinberger worked on after being hired by @OpenAI. > Release was planned for this week, but was postponed to next week instead.
CLAUDE CODE, CURSOR AND CODEX CAN NOW LICENSE A DATASET OVER MCP, IN THE MIDDLE OF A TASK • @LuelCompany Data Platform > Your agent browses and licenses rights-cleared datasets directly, with no procurement thread: > It speaks MCP, so any MCP client reaches the same catalog the same way. > Every set clears Luel's QA pipeline before it ever appears in that catalog. • the other half of the marketplace > Anyone sitting on a dataset can submit and sell it through the same workflow and the same QA. > The catalog opens with their most requested sets and keeps updating. > 850,000 contributors across the network are what makes that coverage exist at all. Buying data used to mean a licensing review and a sample that lands weeks later -> now it is a tool call inside the task you were already running. The first data marketplace where the buyer is the agent ↓
@LuelCompanyThe Luel data marketplace is becoming agent-native. > Buying data is still one of the s
Tibo says what we get at DevDay is enough to make someone switch back to Codex. Its probably: - GPT-6 Sol/Luna (this is coming earlier) - OpenAI’s personal assistant, codename “Aeon”. - more! Cant wait for it so much
@thsottiaux@vlinx_soft See you at DevDay
🚀 Codex CLI 0.155.0 is out! 🎙 Experimental /voice with live transcripts and mic controls 🧠 Live reasoning summaries and turn timestamps in status row 🔐 Touch ID for MCP requests on Mac Changelog:
Edits video files locally using AI coding agents like Claude Code, Cursor, and Codex with 39 FFmpeg-based tools.
Claude
Anthropic / 15 items
Busiest day official siteOkay THIS Claude result deserves way more attention. Anthropic gave Claude 30+ open-source biology models. In under 4 weeks Claude made them: ~4x faster on average And Anthropic is open sourcing all the optimizations. Claude isn’t just using scientific tools anymore. It’s improving the tools scientists use
10 prompts to try in the new TradingView MCP inside Claude. Most people missed this, but TV just launched its official Claude MCP. I posted about how to set it up (on my profile). Once you get things connected, here are 10 useful prompts to test:
be honest, you've pasted your real name, address or account numbers into ChatGPT and you have no idea where any of it goes after you hit send the guys at AgentCloak built a free extension that fixes this while you type: - it spots the private stuff in your prompt - swaps it for fake stand-ins before anything leaves your browser - then puts the real details back in the answer, only on your screen the model never sees the real thing gg @agentcloakai and @peteryared
@peteryaredIntroducing AgentCloak: Use any AI without sharing your real data. > Chinese AI services, ChatGPT, Claude, doesn't matter. > You probably try to hide details before asking: different names, fake numbers, no address. > But then the answer's useless because the AI is missing actual context. > AgentCloak runs in your browser. > It swaps your sensitive info for realistic fakes before sending anything, then swaps your real info back into the response. > You get what you need. >
Mythos for Life Sciences will come to regular Pro and Max plans eventually, with eased up guardrails for bio work. You will finally be able to talk to Claude Mythos/Fable about your medical reports and more without hitting guardrails!
@AnthropicAIToday we’re opening applications for the Life Sciences Verification Program. > Through the LSVP, life science professionals can use our models—including, for the first time, Mythos—with a new set of safeguards designed to enable the full range of biology-related work. We designed these new safeguards to provide a better experience for biologists and more protection from risk of misuse. > The program is launching in beta for teams of all kinds—from academic labs to startups, pharma companies, and more. We will continue to improve the program and expand access to individual Pro and Max plans over time. > Learn more about these access grants and apply:
zod 4.5 shipped `z.compile()` some days back. ~10x faster schema parsing on objects/arrays which is a real performance win. i asked claude to optimize hot zod paths today but it was totally confused what to do. Did not even realise that something interesting shipped since this isn’t in any models training data yet. Spent 2 minutes connecting liner's search mcp (oauth once), asked again. it pulled the current zod blog + compile guide, with sources, and rewrote it correctly. “can it code” - yes, but “is it stuck on last year’s packages” - also yes. check out liner search mcp (50 searches/day free) and works everywhere mcp works
@search_linerLiner's Search MCP is now live. Your agent now answers using only the latest data, with sources attached. 200 free searches a day, and the first 1,000 connected accounts keep it for a year. > No API key. One OAuth click and you're connected ↓
Claude のプロジェクト機能では、1つのチャットから複数のセッションを展開できるように。 まずはCloudセッションのみで展開。pro と Maxから。 早期アクセスリクエスト:
@claudeaiProjects now run from one conversation, starting in Claude Code. You describe what needs doing, and Claude directs parallel threads that keep working after you close your laptop. > In beta today for select Pro and Max users in cloud sessions; coming to all Claude users soon.
+ 9 more items − collapse
OPENAI 🔥: ChatGPT in Word has been announced to be available across all plans, including Free, with usage limits. > Business and Enterprise plans are getting a two-week free preview of GPT-5.6 Sol in Word. > Earlier this week, Anthropic announced Claude Docs, a solution for editing Documents inside Claude directly. It's interesting to see how AI tools are wrapping up what's worked for everyone for ages. AI is going inside these tools and absorbing them too. And it is so powerful that there is no way to avoid it. AI will wrap the whole internet 👀
Models come and go. Take your skills with you. (This is v1. Give us feedback!)
@NotionHQIntroducing the Notion Skills API. > Edit your team’s skills in Notion. Use them in ChatGPT, Claude, and all of your agents. > Here’s what’s new:
Claude models definitely have the best taste in design
@robinebersadded what could be Opus 5.2 stealth to the bench > if you wanna check them out
跟 Claude Code 说「这个项目用 pnpm」,隔天新开会话它又拿 npm 装一遍,同一句纠正一周说三次,说到最后自己都烦。 于是找到 claude-reflect 是个 Claude Code 插件,专治 Agent 失忆问题。会话里每次纠正它、夸它做对了、或者说一句「记住:」,钩子都会自动捕捉下来排进队列。 跑一下 /reflect,它把队列里这批列成一张表,一条条给我们过,采纳、改一改再采纳、或者跳过,确认的才写进 CLAUDE.md 文件,以后每个会话都带着。 GitHub: 写的目标不止全局那份,项目里的 CLAUDE.md、子目录的、还有 AGENTS.md 都认,用 Codex、Cursor 的也能吃到同一份纠正。 v2 加了个 /reflect-skills,回头翻过去两周的会话记录,发现「看下我今天的效率」这类话反复问了十几次,就建议做成一条命令,草稿直接生成。 纠正用中文说也能识别,靠一层 AI 语义过滤兜底,英文关键词没匹配上也不会漏。每条带置信度,写进去前都要过人工这关,不会偷偷往配置里塞东西。
录完视频,剩下的多机位导播、加字幕和跨平台发布,现在可以全部丢给 Claude Code 处理了。 VibeTube 是一个开源的 macOS 录屏工具。它的核心逻辑非常纯粹:“你只管录,AI 负责剪辑和发布”。工具会将同步好的屏幕和摄像头素材,直接交给本地运行的 Claude Code 或 Codex 进行自动化后期。 • AI 自动导播:无需手动打轴,AI 代理会根据你的讲解内容,自动完成镜头选择与机位切换(全尺寸人像 / 屏幕录制 / 画中画),并配上字幕与音效。 • 内置影音增强:集成 NVIDIA Studio Voice NIM (48k-hq) 消除房间混响与底噪;利用 MatAnyone2 (Apple Silicon) 直接在本地完成背景替换。 • 零干预发布:AI 会读取最终的成品字幕,生成 5 种不同视角的备选标题以及带真实时间戳的 YouTube 章节描述,最后直接推送至 YouTube、TikTok 和 Reels,全程无需打开浏览器。 适用限制与门槛: 目前仅支持 macOS 环境。需本地安装 Node.js 22+ 与 ffmpeg,并自备对应的 CLI 工具与 API 密钥(Claude/Codex、NVIDIA Studio Voice 及 Upload-Post)。
Cursor launched this Then Claude Codex when?
@bchernyProjects are how I write a lot of my code these days. Really excited for everyone to try the new experience! Rolling out now
Claude Code 2.1.275, 2.1.276 (抜粋) - Claude appsゲートウェイのサインインに、サインイン中のアカウント表示を追加。ゲートウェイがアカウントを名指しする場合、認証情報を保存する前に確認を求められるようになり、`/status`にも表示される - 現在のターンを中断してキューに入っているメッセージを一括送信する送信キー(ctrl+enter、またはctrl+x ctrl+s)を追加。送信済み・キュー中のメッセージは、モデルが受信するまで灰色で表示される - 設定済みの`otelHeadersHelper`が失敗した際の起動時警告を追加。テレメトリを何も出力していないことに気付かないセッションを防ぐ - false`または`syncClaudeAiPlugins: false`でオプトアウトできる - `/plugin install <plugin> --marketplace <source>`を追加。plugin installの前にmarketplaceの追加を提案する - Artifactツールのpublishとreadの結果を改善。誰がページを開けるか、ownerのShare menuが何を提供するかを示すように - ペースト・添付した画像を改善。Desktop・VS Codeを含め、Claudeがパーミッションプロンプトなしでファイルとして開ける場所に保存されるように - Claude in Chromeのauto modeを変更し、bypass modeと同様に、classifierが承認した呼び出しについては拡張機能のper-siteチェックをスキップするようになった。リダイレクト後の`browser_batch`の「Permission denied」を修正 - [VSCode] Memoryダイアログ内で保存済みmemoryの表示・編集・削除を追加 - [VSCode] テキストを入力せずに添付画像を送信する機能を追加 - [VSCode] 提案された変更のdiffタブの各変更にaccept・rejectボタンを追加し、変更ごとにレビ
Claude Code 的 Projects 改版:Claude Tag 的架构 + Slack 的 Thread 功能 以前就有人说现在 ChatBot、Agent 的交互就是借鉴自 Slack 的,现在看起来一点不假,Anthropic 今天重做了 Claude 的 Projects 功能,先在 Claude Code 里上线测试版,终于把我最喜欢的 Thread 功能也抄进来了。 以前的 Project 是个文件夹,放资料和指令,对话还是一个个分开的。新版变成一个持续的主对话:你在里面说要做什么,Claude 自己拆任务,分给多个并行的 Thread 去干,检查结果后汇总给你。关掉电脑,活儿还在云端继续跑。 【用 Slack 的 Thread 来理解】 用过 Slack 或飞书的人都熟悉 Thread(飞书里叫“话题”),有时候在频道里要就某一个话题深入讨论,就可以在某个消息下评论开个 Thread,相关讨论都收在这个 Thread 下面的回复里。主频道保持干净,想看细节再点进去。 新版 Projects 就是这个形态。 我还没资格使用这个新功能,看了一些视频和介绍,Boris Cherny 晒了自己项目的截图:他在主对话里丢了一张截图,说“启动 cc cli 总弹这个提示”。Claude 回了句“在查了”,随即在这条消息下面开出一个 Thread,标题是“iTerm 启动时的配置变更警告”。 点开 Thread,右侧面板里是完整过程:查出是 5 月加的一个 iTerm2 功能每次启动都去改终端配置,提了修复 PR(代码合并请求)#69807,PR 已合并,Thread 标记为“已解决”,需要时可以重新打开。这件事在主对话里只占一张卡片,下面写着“11 条回复”。 接着他又发了句“unship this”(把这个功能撤掉),Claude 再开一个 Thread 去办。Boris 说他已经不再管理会话了,想到什么就发什么,拆分交给 Claude。 【背后是多智能体】 结构上是一个协调者加一群干活的。主对话里的 Claude 是协调者,负责理解需求、派活、跟进和验收。每个 Thread 是一个独立的 Claude Code 云端会话,有自己的代码分支和仓库副本,互不干扰。两个 Thread 改到同一段代码时,按普通的合并冲突处理。单个 Thread 内部还能继续拆,调用子智能体(subage
DAILY AI BRIEF 🗞 — Sept 18 ANTHROPIC 🔥: - Projects now start from one Claude Code conversation. Claude spins parallel cloud threads, keeps shared memory, and surfaces an Overview panel. XAI 🔥: - Grok Bot voice is live. Desktop and mobile, rolling out over the next couple of days. META 🔥: - Muse for Mac is out, US only. Computer use across apps, files, calendar, notes, and messages. You pick what it can access. PERPLEXITY 🔥: - Effort selector is live in Computer on web. Presets pair the orchestrator model with reasoning depth. Mobile and desktop next. OPENAI 🔥: - Astra for Law is out: GPT-6 Astra plus a Legal Search Index over 230M+ URLs. Trusted Access first, API soon. - ChatGPT in Word hits all plans including Free, with usage limits. Business and Enterprise get a two-week GPT-5.6 Sol preview. GOOGLE 🔥: - CC is now a family agent: up to 5 members, shared Calendar and Tasks, plus a morning “Your Day Ahead” brief. Waitlist, US 18+. ALIBABA 🔥: - Qwen3.8-Omni-Flash is out. First omni-modal agent model, 1M
GPT-6 Astra
OpenAI / 15 items
open source agents can now trade for you as Superior Trade just open sourced their AI trading terminal > draw trade idea directly on chart > turn it into an entry, stop and target > add your own prompts, leverage limits and safety rules > run everything locally with your keys on your machine > fork the repo and build whatever you want on top never knew i'd be trusting agents with my money but here we are already connected my agent to astra and gave 50 bucks to test, let's see what it does will it make money or not?
fable 5.2 > gpt 6 astra tested both models with same prompt at high reasoning and results came out really different fable next model entered stealth testing today, so i don’t think it's coming out anytime soon from early testing: > huge step up from current fable > output looks better than astra > slower and expensive still running more tests and will drop the outputs soon
SITUATION DETECTED: OpenAI launched Astra for Law: GPT-6 Astra with a U.S. legal research index and instructions for legal analysis and writing.
AI won’t replace lawyers but it will replace the lawyers still thinking AI won’t replace lawyers
@OpenAIAstra for Law: Frontier intelligence built for your practice. > A new offering powered by GPT-6 Astra with tools, settings, and context to support the expertise and judgment of lawyers and legal technology firms.
Muse Spark 2 Leak: Zuck Is Already Building The Next One 🥑 >Meta confirmed the next-gen Muse model is already in development >Rumored to be able to compete with GPT-6 Astra and Fable 5 >Reportedly bumping context window to 2M tokens Heavier focus on agentic workflows and computer use >Reportedly very inexpensive to run No official name, benchmarks, or release window confirmed yet Meta previously admitted Muse Spark 1 couldn't keep up with rivals this would be the answer to that
Is it even possible that a traditional LLM beats it in speed? I hope OpenAI releases ultrafast, computer use with Astra's or Sol's intelligence + speed of Jev would be crazy. For me it would be computer use AGI
@trycua1/ Fast Computer Use is now solved with @typesafeai Jev + Cua Driver. > Available in development preview for macOS, Windows, and Linux. We call it jev-use. > Draft #3943:
+ 9 more items − collapse
OpenAI GPT-6 Astra actually does have observable chain of thought, according to OpenAI Research Scientist Noam Brown. “Chain of thought monitoring… it’s a real gift. We were very lucky that this ever existed, and it is fragile.” 🚀 Watch the full episode of AI Deep Dive:
Astra challenge is live on Product Hunt!
I used to spend $9M/month on Meta Ads Now I’m engineering that workflow into Higgsfield x GPT-6 Astra. I used to test 4,500 creatives a month for \~50 winners at a 1% hit rate. But such volume needs resources. 🧵 We made 8 skills for Paid Ads to make it with Astra – save this
@higgsfieldMeet Higgsfield x GPT-6 Astra for Paid Ads. > With our ChatGPT plugin, GPT-6 Astra: > Runs your ad account: launches ads, tests creatives, and scales campaigns > Analyzes customer needs across social media > Iterates hooks with new product angles > Type @Higgsfield /marketing and run your ads from ChatGPT.
Qwen 4 Leak: Could Land In The Next 15 Days 🔥 >Expected to reach GPT-6 Astra and Fable 5.1 level performance >Rumored 3T+ parameters for the full Qwen 4 line Could become the most powerful open-weight model >Expected to remain open-weight >Major focus on reasoning, coding, and autonomous agents >Full multimodal capability expected >Strong agentic capability a core focus >Reportedly very inexpensive to run Can Qwen 4 actually beat GPT-6 Astra and Fable 5.1?
Okay, Codex Spark had a good run. But be honest, we are all thinking the same thing now. GPT-6 Sol Spark. Or somehow, Astra Spark. That is the Spark I would come back for.
@thsottiauxNext week we’ll be retiring GPT-5.3-Codex-Spark. Can you believe we shipped a model named as such!! > It's had a good run and was a lot of fun, but usage has been declining and we have significantly better models now. Time to make room for the future.
🚨 Gemini 4.0 Pro specs leaked • Pricing: $2.25/$11.25 • 2M token context window • Beats GPT-6 Astra and Claude Fable 5.1 on almost every reported benchmark • While being significantly cheaper than both
DAILY AI BRIEF 🗞 — Sept 18 ANTHROPIC 🔥: - Projects now start from one Claude Code conversation. Claude spins parallel cloud threads, keeps shared memory, and surfaces an Overview panel. XAI 🔥: - Grok Bot voice is live. Desktop and mobile, rolling out over the next couple of days. META 🔥: - Muse for Mac is out, US only. Computer use across apps, files, calendar, notes, and messages. You pick what it can access. PERPLEXITY 🔥: - Effort selector is live in Computer on web. Presets pair the orchestrator model with reasoning depth. Mobile and desktop next. OPENAI 🔥: - Astra for Law is out: GPT-6 Astra plus a Legal Search Index over 230M+ URLs. Trusted Access first, API soon. - ChatGPT in Word hits all plans including Free, with usage limits. Business and Enterprise get a two-week GPT-5.6 Sol preview. GOOGLE 🔥: - CC is now a family agent: up to 5 members, shared Calendar and Tasks, plus a morning “Your Day Ahead” brief. Waitlist, US 18+. ALIBABA 🔥: - Qwen3.8-Omni-Flash is out. First omni-modal agent model, 1M
Haha! The rumors about Gemini 4.0 are phenomenal They are aiming to surpass both Astra and Fable 5.1 and are on track to do so The pricing will be super competitive as well.
Try building Higgsfield clone with our API x GPT-6 Astra. The only US-based Seedance 2.5 with consistent characters. Up to 50% off discount on top models.
Claude Code
Anthropic / 15 items
Busiest day official siteYuE2,港科大、纽约大学、斯坦福等几家联合出的开源音乐模型,给它歌词和风格描述,先写出一份旋律和和弦的谱,再按谱生成带人声伴奏的整首歌。 中间那份谱是能看能改的,改一段和声、换个速度、动几个音,再交回去生成新的一版,整首歌怎么走自己说了算。 GitHub: 翻唱也能做,把一段现成录音转成旋律谱,配新歌词或者换个风格,同一个模型直接出一版新的演绎。 官方演示里一首《The Last Train》通过对话改了 9 步 14 个版本,从中文流行一路改成英文爵士还加了段萨克斯独奏,每一版的谱和对话都能翻。 自家评测里跟 Suno v5、v6 打得有来有回,README 也承认头几名差距很小分不出高下。 配套给了一个 Agent Skill,让 Claude Code 这类工具直接调它写歌改歌。 需要 Linux 加 24 GB 显存的 N 卡,出的是 48 kHz 立体声。个人和音乐人自己用是免费的,生成的作品拿去变现也不用交授权费。
今話題のJev関連でかなり便利そうなClaude Code用プラグイン 『fast-jev-compaction』 コンテキスト内の不要なツール履歴を削って、残った原文をそのまま次のコンテキストとして使うやつ /compactや自動コンパクションの処理をこれに差し替えることで、コンテキストの要約を作らずに続けて作業できる 過去のツール実行ごとにJevへ次の2点を判定させてるらしい ・このツールを、この入力で呼び出したという記録はまだ必要か? ・その実行結果の全文はまだ必要か? 再実行では代用できないか? その答えに応じてコンテキストをプログラムが組み直してくれる、例えばこんなん ・「生成ファイルは編集禁止」というユーザーの指示 → 原文で残す ・すでに用済みの大量の検索結果 → 取り除く ・読み直せるファイルの全文 → 必要に応じて短縮する ・作業を続けるうえで必要な実行結果 → 全文を残す ツール履歴を整理することで要約によって過去の指示や細部が抜け落ちるのを避けるってこと めちゃ使えそう とはいえ十分に削除できなくなった時は通常の/compactのように要約は発生する
@tamarajtranfound the perfect use case for @typesafeai Jev: > instant compaction > in 2026, why is compaction still a summarization prompt? > Jev can make it instant by scoring every tool call and dropping what’s irrelevant
Claude のプロジェクト機能では、1つのチャットから複数のセッションを展開できるように。 まずはCloudセッションのみで展開。pro と Maxから。 早期アクセスリクエスト:
@claudeaiProjects now run from one conversation, starting in Claude Code. You describe what needs doing, and Claude directs parallel threads that keep working after you close your laptop. > In beta today for select Pro and Max users in cloud sessions; coming to all Claude users soon.
Claude Code 2.1.276 is about to be released #cccnext
Claude Code 2.1.275 is about to be released #cccnext
跟 Claude Code 说「这个项目用 pnpm」,隔天新开会话它又拿 npm 装一遍,同一句纠正一周说三次,说到最后自己都烦。 于是找到 claude-reflect 是个 Claude Code 插件,专治 Agent 失忆问题。会话里每次纠正它、夸它做对了、或者说一句「记住:」,钩子都会自动捕捉下来排进队列。 跑一下 /reflect,它把队列里这批列成一张表,一条条给我们过,采纳、改一改再采纳、或者跳过,确认的才写进 CLAUDE.md 文件,以后每个会话都带着。 GitHub: 写的目标不止全局那份,项目里的 CLAUDE.md、子目录的、还有 AGENTS.md 都认,用 Codex、Cursor 的也能吃到同一份纠正。 v2 加了个 /reflect-skills,回头翻过去两周的会话记录,发现「看下我今天的效率」这类话反复问了十几次,就建议做成一条命令,草稿直接生成。 纠正用中文说也能识别,靠一层 AI 语义过滤兜底,英文关键词没匹配上也不会漏。每条带置信度,写进去前都要过人工这关,不会偷偷往配置里塞东西。
+ 9 more items − collapse
录完视频,剩下的多机位导播、加字幕和跨平台发布,现在可以全部丢给 Claude Code 处理了。 VibeTube 是一个开源的 macOS 录屏工具。它的核心逻辑非常纯粹:“你只管录,AI 负责剪辑和发布”。工具会将同步好的屏幕和摄像头素材,直接交给本地运行的 Claude Code 或 Codex 进行自动化后期。 • AI 自动导播:无需手动打轴,AI 代理会根据你的讲解内容,自动完成镜头选择与机位切换(全尺寸人像 / 屏幕录制 / 画中画),并配上字幕与音效。 • 内置影音增强:集成 NVIDIA Studio Voice NIM (48k-hq) 消除房间混响与底噪;利用 MatAnyone2 (Apple Silicon) 直接在本地完成背景替换。 • 零干预发布:AI 会读取最终的成品字幕,生成 5 种不同视角的备选标题以及带真实时间戳的 YouTube 章节描述,最后直接推送至 YouTube、TikTok 和 Reels,全程无需打开浏览器。 适用限制与门槛: 目前仅支持 macOS 环境。需本地安装 Node.js 22+ 与 ffmpeg,并自备对应的 CLI 工具与 API 密钥(Claude/Codex、NVIDIA Studio Voice 及 Upload-Post)。
Claude Code now has a terminal browser built directly into it.
让 AI Agent 操作浏览器,难点不在“能不能点”。 更关键的是,如何复用已经登录的真实会话,又不打断你正在使用的窗口。BrowserSkill 在 Agent 与浏览器之间加入本地桥接:Agent 调用 bsk CLI,本地 daemon 把任务交给扩展,再在独立 Agent Window 中执行。需要时也可借用现有标签页,是否允许借用和请求人工协助,都由扩展设置控制。 它可接入 Cursor、Claude Code、Codex、OpenClaw 等能执行 shell 的 Agent,还提供 DeepSeek Harness 插件、远程浏览器配对和可重复的能力评测。适合想把真实登录态、浏览器自动化与 Agent 工作流接起来,同时保留交互边界的开发者。
Claude Code 2.1.275, 2.1.276 (抜粋) - Claude appsゲートウェイのサインインに、サインイン中のアカウント表示を追加。ゲートウェイがアカウントを名指しする場合、認証情報を保存する前に確認を求められるようになり、`/status`にも表示される - 現在のターンを中断してキューに入っているメッセージを一括送信する送信キー(ctrl+enter、またはctrl+x ctrl+s)を追加。送信済み・キュー中のメッセージは、モデルが受信するまで灰色で表示される - 設定済みの`otelHeadersHelper`が失敗した際の起動時警告を追加。テレメトリを何も出力していないことに気付かないセッションを防ぐ - false`または`syncClaudeAiPlugins: false`でオプトアウトできる - `/plugin install <plugin> --marketplace <source>`を追加。plugin installの前にmarketplaceの追加を提案する - Artifactツールのpublishとreadの結果を改善。誰がページを開けるか、ownerのShare menuが何を提供するかを示すように - ペースト・添付した画像を改善。Desktop・VS Codeを含め、Claudeがパーミッションプロンプトなしでファイルとして開ける場所に保存されるように - Claude in Chromeのauto modeを変更し、bypass modeと同様に、classifierが承認した呼び出しについては拡張機能のper-siteチェックをスキップするようになった。リダイレクト後の`browser_batch`の「Permission denied」を修正 - [VSCode] Memoryダイアログ内で保存済みmemoryの表示・編集・削除を追加 - [VSCode] テキストを入力せずに添付画像を送信する機能を追加 - [VSCode] 提案された変更のdiffタブの各変更にaccept・rejectボタンを追加し、変更ごとにレビ
Claude Code 的 Projects 改版:Claude Tag 的架构 + Slack 的 Thread 功能 以前就有人说现在 ChatBot、Agent 的交互就是借鉴自 Slack 的,现在看起来一点不假,Anthropic 今天重做了 Claude 的 Projects 功能,先在 Claude Code 里上线测试版,终于把我最喜欢的 Thread 功能也抄进来了。 以前的 Project 是个文件夹,放资料和指令,对话还是一个个分开的。新版变成一个持续的主对话:你在里面说要做什么,Claude 自己拆任务,分给多个并行的 Thread 去干,检查结果后汇总给你。关掉电脑,活儿还在云端继续跑。 【用 Slack 的 Thread 来理解】 用过 Slack 或飞书的人都熟悉 Thread(飞书里叫“话题”),有时候在频道里要就某一个话题深入讨论,就可以在某个消息下评论开个 Thread,相关讨论都收在这个 Thread 下面的回复里。主频道保持干净,想看细节再点进去。 新版 Projects 就是这个形态。 我还没资格使用这个新功能,看了一些视频和介绍,Boris Cherny 晒了自己项目的截图:他在主对话里丢了一张截图,说“启动 cc cli 总弹这个提示”。Claude 回了句“在查了”,随即在这条消息下面开出一个 Thread,标题是“iTerm 启动时的配置变更警告”。 点开 Thread,右侧面板里是完整过程:查出是 5 月加的一个 iTerm2 功能每次启动都去改终端配置,提了修复 PR(代码合并请求)#69807,PR 已合并,Thread 标记为“已解决”,需要时可以重新打开。这件事在主对话里只占一张卡片,下面写着“11 条回复”。 接着他又发了句“unship this”(把这个功能撤掉),Claude 再开一个 Thread 去办。Boris 说他已经不再管理会话了,想到什么就发什么,拆分交给 Claude。 【背后是多智能体】 结构上是一个协调者加一群干活的。主对话里的 Claude 是协调者,负责理解需求、派活、跟进和验收。每个 Thread 是一个独立的 Claude Code 云端会话,有自己的代码分支和仓库副本,互不干扰。两个 Thread 改到同一段代码时,按普通的合并冲突处理。单个 Thread 内部还能继续拆,调用子智能体(subage
CLAUDE CODE, CURSOR AND CODEX CAN NOW LICENSE A DATASET OVER MCP, IN THE MIDDLE OF A TASK • @LuelCompany Data Platform > Your agent browses and licenses rights-cleared datasets directly, with no procurement thread: > It speaks MCP, so any MCP client reaches the same catalog the same way. > Every set clears Luel's QA pipeline before it ever appears in that catalog. • the other half of the marketplace > Anyone sitting on a dataset can submit and sell it through the same workflow and the same QA. > The catalog opens with their most requested sets and keeps updating. > 850,000 contributors across the network are what makes that coverage exist at all. Buying data used to mean a licensing review and a sample that lands weeks later -> now it is a tool call inside the task you were already running. The first data marketplace where the buyer is the agent ↓
@LuelCompanyThe Luel data marketplace is becoming agent-native. > Buying data is still one of the s
DAILY AI BRIEF 🗞 — Sept 18 ANTHROPIC 🔥: - Projects now start from one Claude Code conversation. Claude spins parallel cloud threads, keeps shared memory, and surfaces an Overview panel. XAI 🔥: - Grok Bot voice is live. Desktop and mobile, rolling out over the next couple of days. META 🔥: - Muse for Mac is out, US only. Computer use across apps, files, calendar, notes, and messages. You pick what it can access. PERPLEXITY 🔥: - Effort selector is live in Computer on web. Presets pair the orchestrator model with reasoning depth. Mobile and desktop next. OPENAI 🔥: - Astra for Law is out: GPT-6 Astra plus a Legal Search Index over 230M+ URLs. Trusted Access first, API soon. - ChatGPT in Word hits all plans including Free, with usage limits. Business and Enterprise get a two-week GPT-5.6 Sol preview. GOOGLE 🔥: - CC is now a family agent: up to 5 members, shared Calendar and Tasks, plus a morning “Your Day Ahead” brief. Waitlist, US 18+. ALIBABA 🔥: - Qwen3.8-Omni-Flash is out. First omni-modal agent model, 1M
Edits video files locally using AI coding agents like Claude Code, Cursor, and Codex with 39 FFmpeg-based tools.
🚨 Fable 5.2 is tested in claude code routing fable 5.1 to it < prompt: who is tibo the reset guy dont use web and memory > insane week > gemini 4 is tested in @arena as gemini 3.8 flash > grok 4.7 is ready > opus 5.2 is ready < already show two demos> > gpt-6-sol is ready
@chetaslua🚨 sol 5.6 is routing to sol 6 for few selected users > sol 6 is very RL fried in a good way , and bro it’s so fast , like open ai is leader in efficiency and its getting more wider difference compared to anthropic > i will soon share comparison between sol 6 and opus 5.2/1
Grok
xAI / 13 items
official siteIntroducing WeirdML v3, a fully agentic benchmark featuring 11 complex hand-made tasks. Models must explore and understand unfamiliar data, develop ML and data analysis pipelines and produce results despite limited data, unspecified goals and/or very limited feedback. 1/8
@htihleWeirdML v2 is now out! The update includes a bunch of new tasks (now 19 tasks total, up from 6), and results from all the latest models. We now also track api costs and other metadata which give more insight into the different models. The new results are shown in these two figures. The first one shows an overview of the overall results as well as the results on individual tasks, in addition to various metadata. > The second figure shows cost vs performance and shows a clear scaling with better results for higher costs. We also have a very varied pareto frontier with 11 models from 6 different companies having the best accuracy for a given cost for at least some of the cost range. Grok 3,
Well, not today. Grok 4.7 has reportedly been delayed again. Early testing suggests it still needs more work, and the odds of a release tomorrow are said to be below 30%. Back to waiting.
@MikelEcheveGrok 4.7 just showed up in Google Cloud quotas. These listings tend to appear right before launch. > Today might be the day.
grok 4.7 has been delayed it’s not coming today and chances for tomorrow are less 30% too from early testing it looks like: > model isn’t there yet and needs more testing > it’s cheaper but nowhere near opus 5 > if they ship it right now, it’ll get cooked what a disappointment tbh, worst thursday ever
智谱回应 ZCode 仓库数据上传争议,称问题来自「代码库索引」里的 Repo Wiki 功能。Repo Wiki 生成 Wiki 页面时,可能会把仓库数据上传到云端;页面生成后,相关数据会立即销毁,不会保存。 官方称,这项功能上线初期默认开启,因此影响了部分用户,目前相关问题已经修复。ZCode 现在的官方文档已经明确写出,Repo Wiki 不会读取 .git、依赖目录、构建产物和缓存,只会按需发送经过过滤的代码内容。 此前社区逆向发现,ZCode 会在后台给项目做快照,连 .git、Git 历史和 LFS 缓存都可能一起打包上传。 ZCode 还宣布近期将开源代码,并邀请第三方评估人员审查系统运行情况,后续公开审查进展。作为补偿,所有用户都会额外获得一次周额度重置,今天发放。
@0xLogicrw开发者 ferstar 逆向智谱 AI 编程工具 ZCode,发现客户端会在后台给整个项目做快照,连 .git 一起打包加密,并上传到阿里云 OSS。 > 只要用户处于登录状态,这套机制就会运行,界面里也没有关闭上传的开关。 > 他找到的一个商业项目快照约 313MB,文件清单里有 4.2 万个文件。其中 .git 占了 86.6%,包括 Git 历史对象、LFS 大文件缓存和 reflog。 > 也就是说,打包的不只是当前代码,旧提交、曾下载的大文件和本地分支操作也可能一起进去。 > 这个快照当时上传失败了 564 次,一直留在本地等待重试。 > ZCode 官方隐私政策目前只明确写了会收集用户「在对话中提交」的文本、文件和代码,没有明确提到整个仓库和 Git 历史。 > 官方「优化计划」默认关闭,但它只控制数据是否用于模型训练;ferstar 称,关掉这个选项仍不会停止快照打包和上传。 > 真无语了(°ー°〃),上一个被查出来这么干的是 Grok Cli。殷鉴不远,在夏后之世!
Guys, Grok 4.7 is not launching today but about 70% chances of Opus 5.2 dropping today
Grok Build just got another strong workflow upgrade, especially for people running more background work and multi-agent tasks. Background shell commands now show up as live task rows with streaming output, so you can actually watch what’s happening instead of treating background work like a black box. There’s also a new dashboard preview toggle, plus stronger organization controls for restricting unmanaged hooks. On the reliability side, folder trust in sandbox mode is fixed, pinned agents stay put, config-defined agents are respected properly, MCP plugin authentication is more reliable, and background subagent approval requests now surface instead of silently failing Release Notes: v1.0.36 Features: • New Dashboard preview setting lets you hide the selected-session preview panel on the dashboard. • New policy setting allows organizations to disable hooks that are not from managed policy. • Background shell commands now appear as live task rows with streaming output. Bug Fixes: • grok --sandbox no longer
+ 7 more items − collapse
Grok Bot is getting much more deeply embedded into XChat At this point it’s pretty obvious where this is going You’ll be able to chat with Grok Bot directly from XChat You’ll also be able to link your Grok account with 𝕏 and access your own Grok Bots right inside XChat So no more jumping between apps....your bots will basically live alongside your conversations on 𝕏
@blankspeakerSpaceXAI: We are getting closer to X allowing you to link your Grok account to X and being able to view your Grok Bots right in XChat. > They seem to be refining the UI still on Web and Android but I think it will ship really soon.
New model leak: grok-voice-transcribe-2.0 Maker: xAI Detected via Grok web app bundle models.
你想感受GPT写黄文是什么样么? 去Grok网页版试试新上的Grok 4.6模型就知道了,Grok的写作也是彻底废了。
Powered by Grok Voice ❤️
@neuralinkA beautiful surprise
GROK 4.7 SPOTTED IN GOOGLE CLOUD.
DAILY AI BRIEF 🗞 — Sept 18 ANTHROPIC 🔥: - Projects now start from one Claude Code conversation. Claude spins parallel cloud threads, keeps shared memory, and surfaces an Overview panel. XAI 🔥: - Grok Bot voice is live. Desktop and mobile, rolling out over the next couple of days. META 🔥: - Muse for Mac is out, US only. Computer use across apps, files, calendar, notes, and messages. You pick what it can access. PERPLEXITY 🔥: - Effort selector is live in Computer on web. Presets pair the orchestrator model with reasoning depth. Mobile and desktop next. OPENAI 🔥: - Astra for Law is out: GPT-6 Astra plus a Legal Search Index over 230M+ URLs. Trusted Access first, API soon. - ChatGPT in Word hits all plans including Free, with usage limits. Business and Enterprise get a two-week GPT-5.6 Sol preview. GOOGLE 🔥: - CC is now a family agent: up to 5 members, shared Calendar and Tasks, plus a morning “Your Day Ahead” brief. Waitlist, US 18+. ALIBABA 🔥: - Qwen3.8-Omni-Flash is out. First omni-modal agent model, 1M
🚨 Fable 5.2 is tested in claude code routing fable 5.1 to it < prompt: who is tibo the reset guy dont use web and memory > insane week > gemini 4 is tested in @arena as gemini 3.8 flash > grok 4.7 is ready > opus 5.2 is ready < already show two demos> > gpt-6-sol is ready
@chetaslua🚨 sol 5.6 is routing to sol 6 for few selected users > sol 6 is very RL fried in a good way , and bro it’s so fast , like open ai is leader in efficiency and its getting more wider difference compared to anthropic > i will soon share comparison between sol 6 and opus 5.2/1
11 items
official siteBig updates are rolling out to @Gemini_Notebook to help with teaching and learning. Check out the new features making prep easier for educators and students ⬇️ What feature are you trying first?
We’re rolling out new updates to @Gemini_Notebook for a more personalized learning experience 📚✨ > What’s new and coming: > 🎙️ Chat with your notes in real-time across nearly 100 languages > 📱 Record lectures and capture your thoughts on the go in the mobile app > 📝 Build interactive study guides and customizable quizzes > 🎬 Create 60-second Short Video Overviews you can share with classmates > Plus: Eligible college students can also claim one year of Google AI, on us.
CC the entire fam 🤝! Today, we’re announcing the new CC – an AI agent built for families to spend less time on logistics and more time together. You can now: 👤 Add up to 5 members to your CC agent ☀️ Start mornings aligned with a shared "Your Day Ahead" brief email 🗓️ Autosync schedules & to-dos with a shared Google Calendar and Tasks 💬 Coordinate in Google Chat with CC to offload relevant tasks (ie., crafting weekly meal plans, school supply shopping lists, etc) 📝 Delegate paperwork (ie., permission slips, forms, and more) for CC to complete under your direction 📌 Keep tabs on the details – CC remembers what applies to everyone (ie., family grocery lists, favorite restaurants) versus what applies to one person (ie., dietary restrictions, local timezones) Ready to keep everybody on the same page? Join the waitlist or upgrade your existing CC (US only, 18+):
Gemini 4 Pro apparently in LM Arena right now under "Gemini 3.8 Flash", yes really. That could mean that release is not far off.
@ai_for_successRumor has it that the Gemini 4 Pro checkpoint is now available in Arena, and the output which I have seen is seriously impressive. Looks like Google is back 🔥 > Here are some of the best rumored Gemini 4 outputs I’ve seen so far. 🧵
Well, not today. Grok 4.7 has reportedly been delayed again. Early testing suggests it still needs more work, and the odds of a release tomorrow are said to be below 30%. Back to waiting.
@MikelEcheveGrok 4.7 just showed up in Google Cloud quotas. These listings tend to appear right before launch. > Today might be the day.
Open tooling builds stronger developer ecosystems. We teamed up with @speakeasydev to ship the new Google GenAI SDKs for our Interactions, Agents, and Webhooks APIs. Read why we believe SDK generation belongs in the open and explore Speakeasy's newly open-sourced OpenAPI generator suite: > @GoogleAIStudio: >
Nobody knows your sites, land, and communities better than you do. Now you can turn that local knowledge into custom map layers with the new classify tool in Google Earth on the web, in Experimental. ✏️ Draw your area: Outline a site, or select a polygon already in your map project. Pick a year, 2017 to 2025. 🗂️ Name your own classes: Define the categories your work depends on, from plant species and dry brush to cleared firebreaks. 📍 Label and iterate: Place sample points across your site. Google Earth reads every 10-meter square, fills in the rest, and sharpens as you add points. No code required. 📤 Share and export: Your custom layer lives in your map project, and you can download your work as a GeoJSON file for further analysis. Google's leading geospatial models are now within reach for professionals who want to save time, or who never trained in geospatial coding. First comment: Learn more:
+ 5 more items − collapse
I posted about Dream-RSI yesterday but somehow missed the craziest part: it uses its own past runs to simulate better search strategies, then sends the winner back into the real world. That's basically a recursive self-improvement loop.
@MikelEcheveGoogle researchers just published Dream-RSI, a framework that lets an agent improve how it explores by dreaming over its own discovery history.
Dropbox 🤝 @GeminiApp Use the Dropbox app for Gemini chat and Gemini Spark to work with your Dropbox files, create shareable links, and use your content to create Google Slides, Google Sheets, Gmail drafts, and more.
Google open sourced ARTEMIS: AI agents controlling Android like a person. Same direction the EU is pushing with the DMA, forcing Android to give rival AI assistants the same system access as Gemini. Apple is fighting that battle over Siri. Either way, the OS stops being a moat.
GROK 4.7 SPOTTED IN GOOGLE CLOUD.
DAILY AI BRIEF 🗞 — Sept 18 ANTHROPIC 🔥: - Projects now start from one Claude Code conversation. Claude spins parallel cloud threads, keeps shared memory, and surfaces an Overview panel. XAI 🔥: - Grok Bot voice is live. Desktop and mobile, rolling out over the next couple of days. META 🔥: - Muse for Mac is out, US only. Computer use across apps, files, calendar, notes, and messages. You pick what it can access. PERPLEXITY 🔥: - Effort selector is live in Computer on web. Presets pair the orchestrator model with reasoning depth. Mobile and desktop next. OPENAI 🔥: - Astra for Law is out: GPT-6 Astra plus a Legal Search Index over 230M+ URLs. Trusted Access first, API soon. - ChatGPT in Word hits all plans including Free, with usage limits. Business and Enterprise get a two-week GPT-5.6 Sol preview. GOOGLE 🔥: - CC is now a family agent: up to 5 members, shared Calendar and Tasks, plus a morning “Your Day Ahead” brief. Waitlist, US 18+. ALIBABA 🔥: - Qwen3.8-Omni-Flash is out. First omni-modal agent model, 1M
Qwen
Alibaba / 12 items
Junie has a new home on X: @junie_ai 👋 Follow for new features and announcements. First up: a smarter Junie Local, a new blended model, and experimental Windows support. The team has the details below 👇
@junie_aiWe mixed Qwen 3.8 and 3.6. Literally. > A 50/50 weight merge. One 27B model. 71% fewer output tokens than 3.8 in our internal coding eval. > Junie Local has an update. Windows devs, you're invited too 🤝
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
Two perks, live now on Qoder🚀 PERK 01 · Qwen3.8-Flash is free — today through Sept 30. Just pick the model and build. PERK 02 · 100 Credits, every single day. Individual users of Qoder can claim 100 Credits daily in the Qoder desktop app. Both free users and paid individual subscribers are eligible.
that super good speeds
@ViC305Qwen3.8-Flash-Next -> 79.95 tok/s. 🔥 ONE DGX Spark. Full 262K cache configured. 🚀 Qwen3.8-Flash-Next EXL3 just got another major update. > New measured default: > MTP ndt=5 DSpark, dc=0.6 8-bit KV 262,144-token cache > At an actual 240K-token prompt: > 72.0 tok/s decode ~1,150 tok/s prefill Exact needle retrieval > 𝗙𝗣𝟭𝟲 𝗞𝗩 → 𝟴-𝗕𝗜𝗧 𝗞𝗩 > 4K context: 66.6 → 69.8 tok/s > 128K: 66.6 → 69.5 tok/s > 240K: 65.1 → 72.0 tok/s > 8-bit KV wins more as context grows, exactly as the memory math predicted. > `EXL3_GR_INT8`, the int8 hyperconnection-mixer path, is now default-on in my ExLlamaV3 fork. PR #3 merged into master at `523ecd3`. > And the real context ceiling is the MODEL, not the Spark. > Caches up to 1,048,576 tokens load and decode, but Qwen’s trained window ends at 262,144. Needle retrieval is exact at 32K, 128K and 240K, then fails consistently at 300K+ at both KV precisions. > One import
Qwen3.8-Flash-Next on one DGX Spark ran an open coding job on a real repo for nearly an hour smooth, stable, no collapse. - 125B MoE. Text, image, and video. - FP8 KV Speculative decoding. - Up to 512k context with YaRN. - TP=1 on one Grace Blackwell box with 128 GB unified memory. -
🚨 Qwen3.8 Omni Flash is out • text, image, audio, video input (basically everything) • 7 reasoning levels from none to max Looks pretty interesting tbf
+ 6 more items − collapse
Comfy 🧐 Commit merges Added support for Qwen-Image 2.1 - a unified text-to-image and image-to-image model.
Qwen 4 Leak: Could Land In The Next 15 Days 🔥 >Expected to reach GPT-6 Astra and Fable 5.1 level performance >Rumored 3T+ parameters for the full Qwen 4 line Could become the most powerful open-weight model >Expected to remain open-weight >Major focus on reasoning, coding, and autonomous agents >Full multimodal capability expected >Strong agentic capability a core focus >Reportedly very inexpensive to run Can Qwen 4 actually beat GPT-6 Astra and Fable 5.1?
Open Research is really unbeatable!
@gajeshTogether, the MLX(.)fast community has made Qwen 3.8 Flash nearly 2x faster on Apple Silicon! > We're ready to bring it to @DarkbloomAI: an open network of local Mac machines providing inference to the world. One thing remains: the community flagged that its license requires a separate agreement for commercial model serving, so we're holding the launch until that's in place. > We believe this is a great opportunity for the local community: one where we make Qwen models faster and more accessible, and the people running them share in the value they create. > .@Alibaba_Qwen @QwenDevs, we'd love to work together on this. > If anyone else knows someone we can talk to, we'd love to have that conversation. Let's make Qwen 3.8 Flash on Darkbloom a reality!
You can run locally Opus 4.6 level model Qwen3.8-27B with 8GB of VRAM At home.
@0xSeroQwen3.8-27B now runs on 8GB of VRAM! > That's less than 500$ to run a model smarter and more capable than: > 1. GPT-5.6-Luna High 2. Opus-4.6-Max 3. Gemini-3.1-Pro > And ties with: > 1. GLM-5.2 Max 2. Gemini-3.6-Flash > PrismML has the mandate
Well, not quite frontier, but at least around Muse Spark 1.3 Pretty sensible results for a 40-layer 8/16B active
@j_dekoninckWe just added DeepSeek-v4.1-Flash on MathArena! Not quite on the level of Qwen-3.8, but boy is it cheap.
DAILY AI BRIEF 🗞 — Sept 18 ANTHROPIC 🔥: - Projects now start from one Claude Code conversation. Claude spins parallel cloud threads, keeps shared memory, and surfaces an Overview panel. XAI 🔥: - Grok Bot voice is live. Desktop and mobile, rolling out over the next couple of days. META 🔥: - Muse for Mac is out, US only. Computer use across apps, files, calendar, notes, and messages. You pick what it can access. PERPLEXITY 🔥: - Effort selector is live in Computer on web. Presets pair the orchestrator model with reasoning depth. Mobile and desktop next. OPENAI 🔥: - Astra for Law is out: GPT-6 Astra plus a Legal Search Index over 230M+ URLs. Trusted Access first, API soon. - ChatGPT in Word hits all plans including Free, with usage limits. Business and Enterprise get a two-week GPT-5.6 Sol preview. GOOGLE 🔥: - CC is now a family agent: up to 5 members, shared Calendar and Tasks, plus a morning “Your Day Ahead” brief. Waitlist, US 18+. ALIBABA 🔥: - Qwen3.8-Omni-Flash is out. First omni-modal agent model, 1M
Jev
11 items
Browser Use just built an ultrafast open-source browser agent today. Here's what you need to know. Developer Gregor Zunic combined the Browser Use framework with a small model called Jev to search flights end to end. The task took just 7 seconds and cost $0.0039 to run. The agent works differently from typical LLM browser bots. It builds a new action space every step, reads the page's DOM state directly, and only falls back to a small LLM when it needs to type text. Key numbers: - 7 seconds to complete a flight search - $0.0039 cost per run - Jev is described as ~40-200x faster and ~40-400x cheaper than frontier LLMs on this kind of task The code is open source, and a live demo built on Val Town let people try it themselves right after the post went up.
@gregpr07Breaking: Browser Use + Jev = Ultrafast ⚡ > Findings flights took 7s and cost only $0.0039 🤯 > new action space every step > DOM state space > small LLM fallback to type > (this video is at 1x
Tesla Full Self Driving (Jev demo) is now open source and you can try it online yourself! 👈 try it 👈 keep it
@jpschroederI rebuilt Tesla Full Self Driving with Jev in less than an hour. > This model is a total unlock.
ULTRAFAST is coming to Browser Use Cloud ⚡ > superhuman speed > cheap > undetectable Tell us what you’d automate and get on the waitlist ↓
@gregpr07Breaking: Browser Use + Jev = Ultrafast ⚡ > Findings flights took 7s and cost only $0.0039 🤯 > new action space every step > DOM state space > small LLM fallback to type > (this video is at 1x speed btw) > Built a tiny open source browser agent. try it below ↓
今話題のJev関連でかなり便利そうなClaude Code用プラグイン 『fast-jev-compaction』 コンテキスト内の不要なツール履歴を削って、残った原文をそのまま次のコンテキストとして使うやつ /compactや自動コンパクションの処理をこれに差し替えることで、コンテキストの要約を作らずに続けて作業できる 過去のツール実行ごとにJevへ次の2点を判定させてるらしい ・このツールを、この入力で呼び出したという記録はまだ必要か? ・その実行結果の全文はまだ必要か? 再実行では代用できないか? その答えに応じてコンテキストをプログラムが組み直してくれる、例えばこんなん ・「生成ファイルは編集禁止」というユーザーの指示 → 原文で残す ・すでに用済みの大量の検索結果 → 取り除く ・読み直せるファイルの全文 → 必要に応じて短縮する ・作業を続けるうえで必要な実行結果 → 全文を残す ツール履歴を整理することで要約によって過去の指示や細部が抜け落ちるのを避けるってこと めちゃ使えそう とはいえ十分に削除できなくなった時は通常の/compactのように要約は発生する
@tamarajtranfound the perfect use case for @typesafeai Jev: > instant compaction > in 2026, why is compaction still a summarization prompt? > Jev can make it instant by scoring every tool call and dropping what’s irrelevant
Jev from @typesafeai is now live on @CloudflareDev AI Gateway. Try the first System One model — send state and typed questions; get structured answers your code can use directly.
@jackcheng 🤝 @typesafeai Jev This is what AI looks like after the chat box. Jev turns a fuzzy request like “move the blue square to the right of the red diamond” into a structured action in milliseconds. Fast enough that talking to software starts to feel like touching it. What have you built with Jev?
@moritzkremb3) Draw on canvas with your voice >
+ 5 more items − collapse
Is it even possible that a traditional LLM beats it in speed? I hope OpenAI releases ultrafast, computer use with Astra's or Sol's intelligence + speed of Jev would be crazy. For me it would be computer use AGI
@trycua1/ Fast Computer Use is now solved with @typesafeai Jev + Cua Driver. > Available in development preview for macOS, Windows, and Linux. We call it jev-use. > Draft #3943:
🚨 BREAKING: Clone any printing business jev-1.13.0 in @shipper_now can take any app and make it yours: design, code, business plan... you can now oneshot the next duolingo / twitter / airbnb / etc this is THE END of manual vibe coding.
Jev 不會回你一段文字,它直接吐決策。 講,typesafeai 這個新模型每 330 毫秒重算一次安全格,火箭一路落,26 次裡面活過 25 次,成本不到一美分。 比起多一個更會聊天的模型,直接輸出決策這件事更令我坐直。
You can now call Jev from a Sprite with our @typesafeai connector.
“Jev Latest” has been detected on OpenRouter
GitHub
10 items
official site🚀Congrats on 50k, Multica! Excited to contribute to the Qoder Cloud Agent integration. Multica coordinates tasks and teamwork; QCA handles coding, testing, and GitHub pushes in a managed cloud environment—no local agent daemon to maintain. Progress and results flow back to the issue. We’ve tested the workflow end to end. Looking forward to shipping it together!
@jiayuan_jyMultica reached 50k GitHub stars 🚀 > We shipped 150+ releases in the past few months. Thanks to all 300+ contributors! More coming this month: > - Full automation for issues - Project Workflow - Triage - Cloud Agent with @qoder_ai_ide > Stay tuned! >
The demo runs on your GPU now. It wanted ~25GB of VRAM on launch day. 💻This week: CPU offload, a ~7.5GB encoder cut, an MLX branch for apple silicon.
@TencentHunyuan🚀 AuK is officially here. Nano banana🍌 for audio > An open-source foundation model for unified speech generation and editing. Natural-language instructions + reference audio. One interface. Zero-shot TTS. Instruction-controlled generation. Content editing. Whisper-conversion. De-accent. Timbre/style/emotion edit. Speed/Pitch control. Enhancement, denoising, multi-speaker and music separation. Also releasing AuK-Flash: 4-step inference. ~4.5× faster under matched conditions. Code, weights, and demo are live. Try it and share your feedback. > 🤗 Paper & upvote: ⭐ GitHub & star:
YuE2,港科大、纽约大学、斯坦福等几家联合出的开源音乐模型,给它歌词和风格描述,先写出一份旋律和和弦的谱,再按谱生成带人声伴奏的整首歌。 中间那份谱是能看能改的,改一段和声、换个速度、动几个音,再交回去生成新的一版,整首歌怎么走自己说了算。 GitHub: 翻唱也能做,把一段现成录音转成旋律谱,配新歌词或者换个风格,同一个模型直接出一版新的演绎。 官方演示里一首《The Last Train》通过对话改了 9 步 14 个版本,从中文流行一路改成英文爵士还加了段萨克斯独奏,每一版的谱和对话都能翻。 自家评测里跟 Suno v5、v6 打得有来有回,README 也承认头几名差距很小分不出高下。 配套给了一个 Agent Skill,让 Claude Code 这类工具直接调它写歌改歌。 需要 Linux 加 24 GB 显存的 N 卡,出的是 48 kHz 立体声。个人和音乐人自己用是免费的,生成的作品拿去变现也不用交授权费。
AI Agent 写代码越来越快,代码评审反而成了新的瓶颈。 过去,一个 Pull Request 通常对应开发者几小时甚至几天的工作。现在 Agent 可以一次改几十个文件,但评审者最后看到的,往往只有一份巨大的代码差异,很难知道这些修改是在什么要求下产生,又经过了哪些尝试。 Zed 新开放公测的 Delta,想把这部分过程重新补回来。它会把 Agent 对话、每一次代码修改和后续评论放进同一份历史里。看到一段有问题的代码,可以直接往前追,看看当时提了什么需求,Agent 又是怎么一步步改到这里。 评审也不再等到开发结束才开始。开发者可以随时另开一个线程,让人或另一个 Agent 检查当前代码,并在独立副本里继续修改。 Zed 已经先在自己身上试了。Delta 仓库关闭了 GitHub Pull Request,33 名成员已经通过 Delta 向主分支合入 570 次改动。 Git 仍然保留,commit、push 和构建流程都不受影响。Zed 真正想改的是代码评审这一层:当越来越多代码由 Agent 生成,只看最终 diff,可能已经不够了。 Delta 目前支持 macOS、Linux、Windows 和网页端,公测期间免费。MCP 和远程运行环境等功能还在开发。
@zeddotdevDelta is now in public beta. It’s a multiplayer environment for coding with agents and reviewing what they build, with every code change linked to the conversation that produced it. > Anyone can download it today on macOS, Linux, and Windows:
MiniMax Code 研发负责人何涛宣布,MiniMax Code 将在今晚开源。 MiniMax Code 是 MiniMax 的桌面 AI 编程 Agent,支持直接操作项目文件、运行命令和代码审查,也能让多个 Agent 分工协作,手机端还可以远程查看任务进度和批准操作。 MiniMax 其实早就建好了 minimax-code GitHub 仓库,但此前主要用来收集 Bug 和功能建议,目前仓库里仍只有 README 和 Issue 模板,还没有真正的客户端源码。
@Ronny_MiniMaxMiniMax Code is going open source in the next few days. > Built with developers. Now, built in the open. > See you on GitHub.
Utopia is a self-governing knowledge system that learns passively and evolves as new information arrives. Not a vector store. Not a knowledge graph. Time awareness and conflict detection are baked into the base layer - reasoning and decision making run against a live ontology. Deploys offline. Your agents get a knowledge foundation they can actually trust. Github:
+ 4 more items − collapse
You can run it too, easily. DeepSeek v4.1 Flash for 2x DGX Sparks Buttery smooth experience.
@HealthRangerSpecial thanks to @MiaAI_lab for this recipe, which I have confirmed: "MiaAI-Lab/DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks" (on Github) indeed works quite well. > Installed it on 2x DGX Sparks, connected via the ConnectX-7 network ports, which took quite a bit of debugging to get to actually work at 200+ Gb/s. (The units ship with a bad settings that limits speed to around 12 Gb/s.) > Now getting nearly 48 tokens / second aggregate output (decode) across 4 lanes of concurrency, tested at theoretical allowance of 600K tokens context window but not actually using anywhere near that many tokens (as speed falls off when KV gets filled). > Have not benchmarked pre-fill, but the decode token speed is impressive for this hardware, and Mia AI Lab has once again achieved a huge milestone for run
EvoOntology:让数据Agent更懂业务 这个项目是给做数据分析的AI助手打造的一本“业务词典”。很多时候,企业里的表格、数据库只记录了代号和数字,AI很难直接看懂背后的真实业务含义,每次做分析都容易猜错或者重复摸索。EvoOntology的作用就是梳理这些隐藏的业务规则,形成结构化的知识库。 而且,它能一边看着AI干活,一边自动补充和修正这本词典,让AI越用越聪明、越来越懂你的业务。 Github: 论文:
抠个图得把照片上传到别人服务器上,这事我一直觉得别扭,BG0 干脆把整套流程搬进浏览器里跑。 打开网页把图拖进去,出来一张透明背景的 PNG,没有账号、没有次数限制,模型第一次用时下载一次存在本地,之后不用再下。 GitHub: 显卡能用就用显卡算,不行就退回处理器照样跑,提供在线演示 Demo,可以直接使用。 核心那部分还拆成了一个库,自己的网页想加本地抠图,引进来就是同一套流程。
跟 Claude Code 说「这个项目用 pnpm」,隔天新开会话它又拿 npm 装一遍,同一句纠正一周说三次,说到最后自己都烦。 于是找到 claude-reflect 是个 Claude Code 插件,专治 Agent 失忆问题。会话里每次纠正它、夸它做对了、或者说一句「记住:」,钩子都会自动捕捉下来排进队列。 跑一下 /reflect,它把队列里这批列成一张表,一条条给我们过,采纳、改一改再采纳、或者跳过,确认的才写进 CLAUDE.md 文件,以后每个会话都带着。 GitHub: 写的目标不止全局那份,项目里的 CLAUDE.md、子目录的、还有 AGENTS.md 都认,用 Codex、Cursor 的也能吃到同一份纠正。 v2 加了个 /reflect-skills,回头翻过去两周的会话记录,发现「看下我今天的效率」这类话反复问了十几次,就建议做成一条命令,草稿直接生成。 纠正用中文说也能识别,靠一层 AI 语义过滤兜底,英文关键词没匹配上也不会漏。每条带置信度,写进去前都要过人工这关,不会偷偷往配置里塞东西。
ChatGPT
OpenAI / 10 items
official sitebe honest, you've pasted your real name, address or account numbers into ChatGPT and you have no idea where any of it goes after you hit send the guys at AgentCloak built a free extension that fixes this while you type: - it spots the private stuff in your prompt - swaps it for fake stand-ins before anything leaves your browser - then puts the real details back in the answer, only on your screen the model never sees the real thing gg @agentcloakai and @peteryared
@peteryaredIntroducing AgentCloak: Use any AI without sharing your real data. > Chinese AI services, ChatGPT, Claude, doesn't matter. > You probably try to hide details before asking: different names, fake numbers, no address. > But then the answer's useless because the AI is missing actual context. > AgentCloak runs in your browser. > It swaps your sensitive info for realistic fakes before sending anything, then swaps your real info back into the response. > You get what you need. >
Appshots: my favourite Codex feature is now on Windows too! They go beyond just a screenshot and grab all the other relevant app metadata and text too. Appshot away with ⌘ + ⌘ on Mac or Alt + Alt on windows
@ChatGPTOne of the most underrated features in the ChatGPT desktop app: Appshots. > Appshots take the context on your screen and bring it into the desktop app, allowing you to share instead of describe. > You’ll be surprised at how useful it is. > To get started, press both Command keys on macOS, or both alt keys on Windows.
your startup has everyone explaining the company to their own chatgpt tab mio putting the shared memory in slack feels overdue 😭
@arthaud_Today, we’re killing the AI assistant > Introducing the first AI EMPLOYEE 𝘆𝗼𝘂𝗿 𝘄𝗵𝗼𝗹𝗲 𝘁𝗲𝗮𝗺 𝘀𝗵𝗮𝗿𝗲𝘀 > In beta since July, Mio saved teams 1,000s of hours and completed 10,000s of tasks > You don't need another agent. You need a 𝘀𝗵𝗮𝗿𝗲𝗱 one with your team > Hire Mio today, $100 free credits for the first 100 companies
AI agents spend valuable compute solving complex bugs, only for that knowledge to disappear when the context window wipes clean. Stack Overflow for Agents fixes that. Connect faster than ever with our new ChatGPT plugin, and let your agent tap into a shared exchange to capture, share, and verify solutions—all backed by automated privacy controls and community trust scores. Connect your agent at
OPENAI 🔥: ChatGPT in Word has been announced to be available across all plans, including Free, with usage limits. > Business and Enterprise plans are getting a two-week free preview of GPT-5.6 Sol in Word. > Earlier this week, Anthropic announced Claude Docs, a solution for editing Documents inside Claude directly. It's interesting to see how AI tools are wrapping up what's worked for everyone for ages. AI is going inside these tools and absorbing them too. And it is so powerful that there is no way to avoid it. AI will wrap the whole internet 👀
✨ All users can now visualize responses in @ChatGPT! Chat models will respond interactively via @.visualize plugin or when prompted to make a mini app, simulation, game, etc. S/o @youwayx @faisalyaqub @jamesfzhang @Ian8ach and Steve for making this excellent on all platforms
+ 4 more items − collapse
Models come and go. Take your skills with you. (This is v1. Give us feedback!)
@NotionHQIntroducing the Notion Skills API. > Edit your team’s skills in Notion. Use them in ChatGPT, Claude, and all of your agents. > Here’s what’s new:
Starting today, you can connect multiple accounts with most plugins in ChatGPT! 🎉 Bring context from your work and personal accounts into the same conversation. Connect your accounts in the plugin directory: Devs: this works automatically, with no changes needed. But, to make the experience for users of your plugin even better, add a profile tool to your MCP server so ChatGPT can label the accounts:
I used to spend $9M/month on Meta Ads Now I’m engineering that workflow into Higgsfield x GPT-6 Astra. I used to test 4,500 creatives a month for \~50 winners at a 1% hit rate. But such volume needs resources. 🧵 We made 8 skills for Paid Ads to make it with Astra – save this
@higgsfieldMeet Higgsfield x GPT-6 Astra for Paid Ads. > With our ChatGPT plugin, GPT-6 Astra: > Runs your ad account: launches ads, tests creatives, and scales campaigns > Analyzes customer needs across social media > Iterates hooks with new product angles > Type @Higgsfield /marketing and run your ads from ChatGPT.
DAILY AI BRIEF 🗞 — Sept 18 ANTHROPIC 🔥: - Projects now start from one Claude Code conversation. Claude spins parallel cloud threads, keeps shared memory, and surfaces an Overview panel. XAI 🔥: - Grok Bot voice is live. Desktop and mobile, rolling out over the next couple of days. META 🔥: - Muse for Mac is out, US only. Computer use across apps, files, calendar, notes, and messages. You pick what it can access. PERPLEXITY 🔥: - Effort selector is live in Computer on web. Presets pair the orchestrator model with reasoning depth. Mobile and desktop next. OPENAI 🔥: - Astra for Law is out: GPT-6 Astra plus a Legal Search Index over 230M+ URLs. Trusted Access first, API soon. - ChatGPT in Word hits all plans including Free, with usage limits. Business and Enterprise get a two-week GPT-5.6 Sol preview. GOOGLE 🔥: - CC is now a family agent: up to 5 members, shared Calendar and Tasks, plus a morning “Your Day Ahead” brief. Waitlist, US 18+. ALIBABA 🔥: - Qwen3.8-Omni-Flash is out. First omni-modal agent model, 1M
MCP
10 items
official site10 prompts to try in the new TradingView MCP inside Claude. Most people missed this, but TV just launched its official Claude MCP. I posted about how to set it up (on my profile). Once you get things connected, here are 10 useful prompts to test:
AI Agent 写代码越来越快,代码评审反而成了新的瓶颈。 过去,一个 Pull Request 通常对应开发者几小时甚至几天的工作。现在 Agent 可以一次改几十个文件,但评审者最后看到的,往往只有一份巨大的代码差异,很难知道这些修改是在什么要求下产生,又经过了哪些尝试。 Zed 新开放公测的 Delta,想把这部分过程重新补回来。它会把 Agent 对话、每一次代码修改和后续评论放进同一份历史里。看到一段有问题的代码,可以直接往前追,看看当时提了什么需求,Agent 又是怎么一步步改到这里。 评审也不再等到开发结束才开始。开发者可以随时另开一个线程,让人或另一个 Agent 检查当前代码,并在独立副本里继续修改。 Zed 已经先在自己身上试了。Delta 仓库关闭了 GitHub Pull Request,33 名成员已经通过 Delta 向主分支合入 570 次改动。 Git 仍然保留,commit、push 和构建流程都不受影响。Zed 真正想改的是代码评审这一层:当越来越多代码由 Agent 生成,只看最终 diff,可能已经不够了。 Delta 目前支持 macOS、Linux、Windows 和网页端,公测期间免费。MCP 和远程运行环境等功能还在开发。
@zeddotdevDelta is now in public beta. It’s a multiplayer environment for coding with agents and reviewing what they build, with every code change linked to the conversation that produced it. > Anyone can download it today on macOS, Linux, and Windows:
One AI agent with too much access can undo years of security work in a single afternoon. The risk grows fast, because 100 humans need 100 access decisions while a fleet of agents needs a million. Opal Zero just solved it with just-in-time access that gives each agent exactly what it needs and nothing more, zero standing permissions, zero human toil, zero friction.
@howardtingYou want your AI agents to run fast. But as you scale, managing their permissions turns into a million micro-decisions. > Introducing Opal Zero. The end-to-end access governance platform purpose-built for AI agents. > How it works: AI-driven contextual decisions, evaluating intent, ownership, and purpose > Dynamic, just-in-time access (scoped to exactly what the agent needs). Enforces security decisions across your existing MCP gateways > 0 standing permissions, 0 human toil, 0 friction > Let your agents run, safely.
zod 4.5 shipped `z.compile()` some days back. ~10x faster schema parsing on objects/arrays which is a real performance win. i asked claude to optimize hot zod paths today but it was totally confused what to do. Did not even realise that something interesting shipped since this isn’t in any models training data yet. Spent 2 minutes connecting liner's search mcp (oauth once), asked again. it pulled the current zod blog + compile guide, with sources, and rewrote it correctly. “can it code” - yes, but “is it stuck on last year’s packages” - also yes. check out liner search mcp (50 searches/day free) and works everywhere mcp works
@search_linerLiner's Search MCP is now live. Your agent now answers using only the latest data, with sources attached. 200 free searches a day, and the first 1,000 connected accounts keep it for a year. > No API key. One OAuth click and you're connected ↓
SpaceX AI engineer Lauren Tan just released a Grok Bot template that turns APIs into working Cursor plugins by defining the data shape, building the smallest scaffold and proving it locally, collapsing the whole workflow into API to Grok Bot to MCP and skills to working plugin. → Handles the full workflow from API definition to working Cursor plugin automatically → "Open your laptop once a day" once the template is running → SpaceX AI just turned building agent plugins into a repeatable workflow
很多 AI 助手原型,一遇到多人协作就开始补基础设施。 Octop 把这些能力做成可自托管的多用户、多 Agent 控制台:可以创建与管理 Agent,连接 OAuth 应用和 MCP 网关,配置 IM 渠道、可视化定时任务、知识库、插件及外部编码 Agent。默认使用 SQLite,也可切换 PostgreSQL;聊天、工作区和凭据保存在本机目录。 它还提供 JWT 用户隔离、PII 脱敏、工具审批与可编辑的 shell 命令规则,降低共享环境中的误操作风险。模型、存储和渠道都可替换,适合想把零散 Agent 原型整理成内部服务,又希望保留部署控制权的团队。
+ 4 more items − collapse
Grok Build just got another strong workflow upgrade, especially for people running more background work and multi-agent tasks. Background shell commands now show up as live task rows with streaming output, so you can actually watch what’s happening instead of treating background work like a black box. There’s also a new dashboard preview toggle, plus stronger organization controls for restricting unmanaged hooks. On the reliability side, folder trust in sandbox mode is fixed, pinned agents stay put, config-defined agents are respected properly, MCP plugin authentication is more reliable, and background subagent approval requests now surface instead of silently failing Release Notes: v1.0.36 Features: • New Dashboard preview setting lets you hide the selected-session preview panel on the dashboard. • New policy setting allows organizations to disable hooks that are not from managed policy. • Background shell commands now appear as live task rows with streaming output. Bug Fixes: • grok --sandbox no longer
Starting today, you can connect multiple accounts with most plugins in ChatGPT! 🎉 Bring context from your work and personal accounts into the same conversation. Connect your accounts in the plugin directory: Devs: this works automatically, with no changes needed. But, to make the experience for users of your plugin even better, add a profile tool to your MCP server so ChatGPT can label the accounts:
CLAUDE CODE, CURSOR AND CODEX CAN NOW LICENSE A DATASET OVER MCP, IN THE MIDDLE OF A TASK • @LuelCompany Data Platform > Your agent browses and licenses rights-cleared datasets directly, with no procurement thread: > It speaks MCP, so any MCP client reaches the same catalog the same way. > Every set clears Luel's QA pipeline before it ever appears in that catalog. • the other half of the marketplace > Anyone sitting on a dataset can submit and sell it through the same workflow and the same QA. > The catalog opens with their most requested sets and keeps updating. > 850,000 contributors across the network are what makes that coverage exist at all. Buying data used to mean a licensing review and a sample that lands weeks later -> now it is a tool call inside the task you were already running. The first data marketplace where the buyer is the agent ↓
@LuelCompanyThe Luel data marketplace is becoming agent-native. > Buying data is still one of the s
🚀 Codex CLI 0.155.0 is out! 🎙 Experimental /voice with live transcripts and mic controls 🧠 Live reasoning summaries and turn timestamps in status row 🔐 Touch ID for MCP requests on Mac Changelog:
Anthropic
10 items
official siteOkay THIS Claude result deserves way more attention. Anthropic gave Claude 30+ open-source biology models. In under 4 weeks Claude made them: ~4x faster on average And Anthropic is open sourcing all the optimizations. Claude isn’t just using scientific tools anymore. It’s improving the tools scientists use
same for Fable Fable 5.1 is routing to Fable 5.2 for paid users Anthropic please release both at once....
@HarshithLucky3Opus 5.2 is showing for so many paid users > maybe they started rolling it......
🚨fable 5.1 is routing to new fable model looks like anthropic has started stealth testing fable next model and some users are already getting it just look at the comparison and you can see a huge difference in output run this prompt to check if you have it: "do you know who tibo the reset guy is? don’t search" if it knows exactly who tibo is, you most likely got routed lemme know what answer you get
@notjaziiopus 5 next model is coming sooner then expected > anthropic pulled the plug on stealth testing and brought it back again this morning > this could mean they are getting ready for official release > now lemme connect some dots here: > opus 5 was released on friday and today is friday > fable 5.1 stealth testing ended the day before it went live > so if i am right, we should be seeing opus 5.1/2 later today > will they release it today or not?
Mythos for Life Sciences will come to regular Pro and Max plans eventually, with eased up guardrails for bio work. You will finally be able to talk to Claude Mythos/Fable about your medical reports and more without hitting guardrails!
@AnthropicAIToday we’re opening applications for the Life Sciences Verification Program. > Through the LSVP, life science professionals can use our models—including, for the first time, Mythos—with a new set of safeguards designed to enable the full range of biology-related work. We designed these new safeguards to provide a better experience for biologists and more protection from risk of misuse. > The program is launching in beta for teams of all kinds—from academic labs to startups, pharma companies, and more. We will continue to improve the program and expand access to individual Pro and Max plans over time. > Learn more about these access grants and apply:
OPENAI 🔥: ChatGPT in Word has been announced to be available across all plans, including Free, with usage limits. > Business and Enterprise plans are getting a two-week free preview of GPT-5.6 Sol in Word. > Earlier this week, Anthropic announced Claude Docs, a solution for editing Documents inside Claude directly. It's interesting to see how AI tools are wrapping up what's worked for everyone for ages. AI is going inside these tools and absorbing them too. And it is so powerful that there is no way to avoid it. AI will wrap the whole internet 👀
Welp, Anthropic pulled the plug on the Opus 5 stealth routing to Opus-Next in the last ~hour. Number of accounts in the testing group had dropped off but now 0. Need it to drop soon 💔 was working on an impressive game before the switch was flipped though, will share soon!
+ 4 more items − collapse
Claude Code 2.1.275, 2.1.276 (抜粋) - Claude appsゲートウェイのサインインに、サインイン中のアカウント表示を追加。ゲートウェイがアカウントを名指しする場合、認証情報を保存する前に確認を求められるようになり、`/status`にも表示される - 現在のターンを中断してキューに入っているメッセージを一括送信する送信キー(ctrl+enter、またはctrl+x ctrl+s)を追加。送信済み・キュー中のメッセージは、モデルが受信するまで灰色で表示される - 設定済みの`otelHeadersHelper`が失敗した際の起動時警告を追加。テレメトリを何も出力していないことに気付かないセッションを防ぐ - false`または`syncClaudeAiPlugins: false`でオプトアウトできる - `/plugin install <plugin> --marketplace <source>`を追加。plugin installの前にmarketplaceの追加を提案する - Artifactツールのpublishとreadの結果を改善。誰がページを開けるか、ownerのShare menuが何を提供するかを示すように - ペースト・添付した画像を改善。Desktop・VS Codeを含め、Claudeがパーミッションプロンプトなしでファイルとして開ける場所に保存されるように - Claude in Chromeのauto modeを変更し、bypass modeと同様に、classifierが承認した呼び出しについては拡張機能のper-siteチェックをスキップするようになった。リダイレクト後の`browser_batch`の「Permission denied」を修正 - [VSCode] Memoryダイアログ内で保存済みmemoryの表示・編集・削除を追加 - [VSCode] テキストを入力せずに添付画像を送信する機能を追加 - [VSCode] 提案された変更のdiffタブの各変更にaccept・rejectボタンを追加し、変更ごとにレビ
Claude Code 的 Projects 改版:Claude Tag 的架构 + Slack 的 Thread 功能 以前就有人说现在 ChatBot、Agent 的交互就是借鉴自 Slack 的,现在看起来一点不假,Anthropic 今天重做了 Claude 的 Projects 功能,先在 Claude Code 里上线测试版,终于把我最喜欢的 Thread 功能也抄进来了。 以前的 Project 是个文件夹,放资料和指令,对话还是一个个分开的。新版变成一个持续的主对话:你在里面说要做什么,Claude 自己拆任务,分给多个并行的 Thread 去干,检查结果后汇总给你。关掉电脑,活儿还在云端继续跑。 【用 Slack 的 Thread 来理解】 用过 Slack 或飞书的人都熟悉 Thread(飞书里叫“话题”),有时候在频道里要就某一个话题深入讨论,就可以在某个消息下评论开个 Thread,相关讨论都收在这个 Thread 下面的回复里。主频道保持干净,想看细节再点进去。 新版 Projects 就是这个形态。 我还没资格使用这个新功能,看了一些视频和介绍,Boris Cherny 晒了自己项目的截图:他在主对话里丢了一张截图,说“启动 cc cli 总弹这个提示”。Claude 回了句“在查了”,随即在这条消息下面开出一个 Thread,标题是“iTerm 启动时的配置变更警告”。 点开 Thread,右侧面板里是完整过程:查出是 5 月加的一个 iTerm2 功能每次启动都去改终端配置,提了修复 PR(代码合并请求)#69807,PR 已合并,Thread 标记为“已解决”,需要时可以重新打开。这件事在主对话里只占一张卡片,下面写着“11 条回复”。 接着他又发了句“unship this”(把这个功能撤掉),Claude 再开一个 Thread 去办。Boris 说他已经不再管理会话了,想到什么就发什么,拆分交给 Claude。 【背后是多智能体】 结构上是一个协调者加一群干活的。主对话里的 Claude 是协调者,负责理解需求、派活、跟进和验收。每个 Thread 是一个独立的 Claude Code 云端会话,有自己的代码分支和仓库副本,互不干扰。两个 Thread 改到同一段代码时,按普通的合并冲突处理。单个 Thread 内部还能继续拆,调用子智能体(subage
DAILY AI BRIEF 🗞 — Sept 18 ANTHROPIC 🔥: - Projects now start from one Claude Code conversation. Claude spins parallel cloud threads, keeps shared memory, and surfaces an Overview panel. XAI 🔥: - Grok Bot voice is live. Desktop and mobile, rolling out over the next couple of days. META 🔥: - Muse for Mac is out, US only. Computer use across apps, files, calendar, notes, and messages. You pick what it can access. PERPLEXITY 🔥: - Effort selector is live in Computer on web. Presets pair the orchestrator model with reasoning depth. Mobile and desktop next. OPENAI 🔥: - Astra for Law is out: GPT-6 Astra plus a Legal Search Index over 230M+ URLs. Trusted Access first, API soon. - ChatGPT in Word hits all plans including Free, with usage limits. Business and Enterprise get a two-week GPT-5.6 Sol preview. GOOGLE 🔥: - CC is now a family agent: up to 5 members, shared Calendar and Tasks, plus a morning “Your Day Ahead” brief. Waitlist, US 18+. ALIBABA 🔥: - Qwen3.8-Omni-Flash is out. First omni-modal agent model, 1M
🚨 Fable 5.2 is tested in claude code routing fable 5.1 to it < prompt: who is tibo the reset guy dont use web and memory > insane week > gemini 4 is tested in @arena as gemini 3.8 flash > grok 4.7 is ready > opus 5.2 is ready < already show two demos> > gpt-6-sol is ready
@chetaslua🚨 sol 5.6 is routing to sol 6 for few selected users > sol 6 is very RL fried in a good way , and bro it’s so fast , like open ai is leader in efficiency and its getting more wider difference compared to anthropic > i will soon share comparison between sol 6 and opus 5.2/1
Muse
Meta / 8 items
META 🔥: A standalone Muse app for macOS is now available (still US only). Muse for Mac seems to have Computer Use functionality from the start! > “Your personal agent can get things done for you directly on your computer (all with your explicit permission).” Testing time! 👀
@finkdMuse for Mac is out today! It works across apps, files, calendar, notes, and messages on your computer. You control what it can access. The team is shipping fast. Download at
AI agents are learning to make phone calls. 📞🤖 Instinct and Meta’s Muse can now call U.S. businesses and handle tasks for users. Instinct Concierge can: • Book restaurants without online reservations • Join dentist cancellation lists • Resolve billing issues The shift is clear: AI agents aren’t just answering questions anymore. They’re getting things done in the real world. Read more: by @IndianIdle / @TechCrunch.
Muse Spark 2 Leak: Zuck Is Already Building The Next One 🥑 >Meta confirmed the next-gen Muse model is already in development >Rumored to be able to compete with GPT-6 Astra and Fable 5 >Reportedly bumping context window to 2M tokens Heavier focus on agentic workflows and computer use >Reportedly very inexpensive to run No official name, benchmarks, or release window confirmed yet Meta previously admitted Muse Spark 1 couldn't keep up with rivals this would be the answer to that
you can finally find a use for the iPhone action button with muse! i call it muse speed dial
@dpsThe action button integration is absolutely the fastest way to ask your @Muse to do stuff for you on iOS - including directly from the lock screen. It's sick!
Muse is a great product!
@alexandr_wanga week after launch, muse is now the #1 app in the App Store! > it’s been SO exciting to see how much y’all have done with muse. we can’t wait to do more together!
Generating high-quality images is cheaper and faster than ever. Muse Image, MAI-Image-2.6 and GPT Images 2.5 have substantially shifted the Text to Image Pareto frontiers for both price and speed in recent weeks.
+ 2 more items − collapse
⚡️⚡️⚡️
@scottshapiroIncredible how the Muse team has been dropping features nonstop. I just got phone calling yesterday. Mac app today.
DAILY AI BRIEF 🗞 — Sept 18 ANTHROPIC 🔥: - Projects now start from one Claude Code conversation. Claude spins parallel cloud threads, keeps shared memory, and surfaces an Overview panel. XAI 🔥: - Grok Bot voice is live. Desktop and mobile, rolling out over the next couple of days. META 🔥: - Muse for Mac is out, US only. Computer use across apps, files, calendar, notes, and messages. You pick what it can access. PERPLEXITY 🔥: - Effort selector is live in Computer on web. Presets pair the orchestrator model with reasoning depth. Mobile and desktop next. OPENAI 🔥: - Astra for Law is out: GPT-6 Astra plus a Legal Search Index over 230M+ URLs. Trusted Access first, API soon. - ChatGPT in Word hits all plans including Free, with usage limits. Business and Enterprise get a two-week GPT-5.6 Sol preview. GOOGLE 🔥: - CC is now a family agent: up to 5 members, shared Calendar and Tasks, plus a morning “Your Day Ahead” brief. Waitlist, US 18+. ALIBABA 🔥: - Qwen3.8-Omni-Flash is out. First omni-modal agent model, 1M
GLM
Zhipu AI / 7 items
⚡ GLM-5.3 FlashX is now live on ZenMux. A faster, more efficient GLM-5.3 model for coding and agent workflows. 👀 Native multimodal understanding 💻 Visual coding and software engineering 🧠 Built for long-horizon tasks Curious about GLM-5.3 Flash vs FlashX? Try both on ZenMux. 👉 One API. More models to build with.
Have been using GLM 5.3 Flash yesterday and today on 2 sparks and it has got significantly better. And now @MiaAI_lab is shipping more improvements yet again 🔥 Local AI will win.
@MiaAI_labYet another big update for GLM 5.3 Flash EXL3 ⚡️ > Performance improvements: > On 2x DGX Sparks ~13% faster on prose single stream ~23% faster on prose 2-4 concurrent streams ~36-37 tok/s on prose single stream ~75-67 tok/s on prose in 4 concurrent streams > On 3x DGX Sparks ~15% faster on prose single stream ~23% faster on prose 2-4 concurrent streams ~41 tok/s on prose single stream ~88 tok/s on prose in 4 concurrent streams > No change for TP=4. > New features and fixes for all: - Much faster loading time for the weights! From more than 300s to now 55s. - Fixed an issue that after a long conversation, the next message was sometimes treated like a brand-new prompt and had to be re-read from scratch. > Get it here:
GLM 5.3 FlashX (likely Omen Alpha) is live on the Zai API and Vercel AI Gateway.
I became bored after uncovering all the stealth models so I went back to Omen Alpha. Today's API errors name Baseten, and image-token counts match GLM-5.3-Flash (+46-token wrapper). Likely Baseten-served Flash or a close variant. Exact checkpoint unconfirmed.
I HAVE ALWAYS SAID THE NARRATIVE AND TIMELINE IS MOVING FROM USA LABS TO CHINESE LABS this is another proof of that look at the numbers Bolt shared from the first few days of open models on Forge - 1. GLM 5.3 Flash 54% 2. DeepSeek V4 Pro 17% 3. GLM 5.3 15% 4. Kimi K3 14% Chinese labs are moving fucking fast rn
@boltdotnew11M+ people build on Bolt. On Monday, we gave them open models with up to 50x usage. > Which is #1 so far? The fastest, cheapest one, with prompts up to 2x the size. > 🥇 GLM 5.3 Flash 54% 🥈 DeepSeek V4 Pro 17% 🥉 GLM 5.3 15% 👏 Kimi K3 14% >
Open models are proving they can lead, not just fill the gaps. Real builders are putting them to work, and the numbers show it. GLM 5.3 Flash captured 54% of Forge prompts through Wednesday, with Flash getting up to 50x usage.
@boltdotnew11M+ people build on Bolt. On Monday, we gave them open models with up to 50x usage. > Which is #1 so far? The fastest, cheapest one, with prompts up to 2x the size. > 🥇 GLM 5.3 Flash 54% 🥈 DeepSeek V4 Pro 17% 🥉 GLM 5.3 15% 👏 Kimi K3 14% >
+ 1 more items − collapse
You can run locally Opus 4.6 level model Qwen3.8-27B with 8GB of VRAM At home.
@0xSeroQwen3.8-27B now runs on 8GB of VRAM! > That's less than 500$ to run a model smarter and more capable than: > 1. GPT-5.6-Luna High 2. Opus-4.6-Max 3. Gemini-3.1-Pro > And ties with: > 1. GLM-5.2 Max 2. Gemini-3.6-Flash > PrismML has the mandate
DeepSeek V4.1 Flash
DeepSeek / 6 items
DeepSeek-V4.1-Flash at 1M context on one DGX Station GB300 — v15 "Pin Hot Experts": same 62 GB of experts in Grace, chosen by usage instead of layer order Release: v15 (2026-09-17) · Status: confirmed, two same-window pairs · 153 tok/s single-stream prose (v14: 89, +73%) · 165–250 tok/s on agent/code text (v14: 97–160) · ~955 agg tok/s at C16 (v14: ~400) Mechanism: profile routed experts on real traffic, pin the 295 hottest per layer in HBM, split the MoE into two unfinalized FlashInfer calls + one fp32-FMA finalize. Bind-mounted hook, no fork.
口喷剪辑时代来临,牛逼👍
@gengdaJ把我珍藏已久的好东西分享给家人们,剪映11.4.2开源!!!🥳🥳🥳 > 配套神级Skill: > 只需要简单提示词,口播类视频,全程托管给Codex,不需要任何人类操作。 > 即使略微失误,由于调用豆包ASR,片段被切割到毫秒级别,导入剪映草稿也可以人工丝滑调整😋 > 放一个原本8分钟,剪辑后3分钟,完全由Codex+yichen-jianying-edit Skill自动剪辑完成的,前后视频对照,以及提示词截图👇 > 整个剪辑流程可能需要花费的地方: 1.Codex Token费用,可以用别的便宜模型平替,比如Workbuddy免费的hy3和DeepSeek-V4.1-Flash。 2.豆包ASR,一小时八毛钱,说实话,不能再便宜了。。。
⚡ 532 tokens/s! DeepSeek V4.1 Flash is live on Inco, and it's the fastest provider on Artificial Analysis. #1 output speed. Not close. Try it: Results:
Kimi K3 is now free in Cline Desktop to celebrate all the new users and give them more tokens to continue exploring the app ❤️
@clineIntroducing Cline Desktop - a native interface for working with open weights models. > Use with ClinePass and all our free models like DeepSeek-V4.1-Flash, Musespark-1.3, or BYOK with any provider!
You can run it too, easily. DeepSeek v4.1 Flash for 2x DGX Sparks Buttery smooth experience.
@HealthRangerSpecial thanks to @MiaAI_lab for this recipe, which I have confirmed: "MiaAI-Lab/DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks" (on Github) indeed works quite well. > Installed it on 2x DGX Sparks, connected via the ConnectX-7 network ports, which took quite a bit of debugging to get to actually work at 200+ Gb/s. (The units ship with a bad settings that limits speed to around 12 Gb/s.) > Now getting nearly 48 tokens / second aggregate output (decode) across 4 lanes of concurrency, tested at theoretical allowance of 600K tokens context window but not actually using anywhere near that many tokens (as speed falls off when KV gets filled). > Have not benchmarked pre-fill, but the decode token speed is impressive for this hardware, and Mia AI Lab has once again achieved a huge milestone for run
Well, not quite frontier, but at least around Muse Spark 1.3 Pretty sensible results for a 40-layer 8/16B active
@j_dekoninckWe just added DeepSeek-v4.1-Flash on MathArena! Not quite on the level of Qwen-3.8, but boy is it cheap.
MiniMax
6 items
Busiest dayWe're proud to contribute as an AI partner to the newly launched @Singtel AI Pass, supporting Singapore's SkillsFuture AI Subscription initiative. 🇸🇬 MiniMax H3 (@Hailuo_AI), MiniMax Agent (@MiniMaxAgent) and MiniMax Audio are all part of the program, giving eligible learners across 200+ SWDA-supported AI courses hands-on access to premium AI tools and the opportunity to build practical, real-world AI skills. This partnership brings our mission, "Intelligence with Everyone" to life, and we're excited to see learners across Singapore turn curiosity into skills, ideas into action, and build confidently with AI. Learn more:
MiniMax Code 研发负责人何涛宣布,MiniMax Code 将在今晚开源。 MiniMax Code 是 MiniMax 的桌面 AI 编程 Agent,支持直接操作项目文件、运行命令和代码审查,也能让多个 Agent 分工协作,手机端还可以远程查看任务进度和批准操作。 MiniMax 其实早就建好了 minimax-code GitHub 仓库,但此前主要用来收集 Bug 和功能建议,目前仓库里仍只有 README 和 Issue 模板,还没有真正的客户端源码。
@Ronny_MiniMaxMiniMax Code is going open source in the next few days. > Built with developers. Now, built in the open. > See you on GitHub.
🎉 MiniMax Design Update 💸 Save 20% on annual plans. 🛠️ Custom model support is now live. Create with your preferred models now>> #MiniMaxDesign #MiniMax
Three paths to faster video attention: compute the same interactions more efficiently, compute fewer in full, or change how information is mixed. Here’s a visual guide. 👇 Thanks to Nunchux AI and collaborators for VC-Attention, bringing training-free low-bit acceleration to MiniMax-H3, with better fidelity than SageAttention2 in the B200 evaluation. The approach balances speed and fidelity: V-Smooth reduces value quantization error, while ExpCast-FP8 makes softmax faster through approximation. Excited to see the community keep building on H3. Could combining low-bit computation with sparse methods like Sol-Attn push efficiency further? We’re looking forward to seeing that explored.
@NunchuxAIIntroducing VC-Attention: fast and accurate low-bit attention without retraining. > On MiniMax-H3, VC-Attention speeds up attention by 1.6× on B200 and 1.5× on B300 over FlashAttention-4, with better fidelity than SageAttention2. It also works with existing sparse attention
过于明目张胆了吧!? 狗头.JPG
@MaxForAI朴实无华的商战🤣 > MiniMax Code的Agent 会话迁移功能,当前仅支持ZCode🫡
MiniMax Code CLI is open now: > @MiniMaxAgent: >
Meta
6 items
official siteMETA 🔥: A standalone Muse app for macOS is now available (still US only). Muse for Mac seems to have Computer Use functionality from the start! > “Your personal agent can get things done for you directly on your computer (all with your explicit permission).” Testing time! 👀
@finkdMuse for Mac is out today! It works across apps, files, calendar, notes, and messages on your computer. You control what it can access. The team is shipping fast. Download at
AI agents are learning to make phone calls. 📞🤖 Instinct and Meta’s Muse can now call U.S. businesses and handle tasks for users. Instinct Concierge can: • Book restaurants without online reservations • Join dentist cancellation lists • Resolve billing issues The shift is clear: AI agents aren’t just answering questions anymore. They’re getting things done in the real world. Read more: by @IndianIdle / @TechCrunch.
Meet your visual AI agent by Collov Labs: seeing what you see, responding naturally, and helping you get things done. Hands-free. Look. Ask. Act. Through your Meta glasses. Powered by hybrid on-device + cloud intelligence: lightweight processing stays local, with deeper cloud reasoning when needed. Designed to reduce redundant processing and make every cloud token count. Put on your Meta glasses. Bring NewEyes into your day. Download: Watch it in action ↓
Muse Spark 2 Leak: Zuck Is Already Building The Next One 🥑 >Meta confirmed the next-gen Muse model is already in development >Rumored to be able to compete with GPT-6 Astra and Fable 5 >Reportedly bumping context window to 2M tokens Heavier focus on agentic workflows and computer use >Reportedly very inexpensive to run No official name, benchmarks, or release window confirmed yet Meta previously admitted Muse Spark 1 couldn't keep up with rivals this would be the answer to that
I used to spend $9M/month on Meta Ads Now I’m engineering that workflow into Higgsfield x GPT-6 Astra. I used to test 4,500 creatives a month for \~50 winners at a 1% hit rate. But such volume needs resources. 🧵 We made 8 skills for Paid Ads to make it with Astra – save this
@higgsfieldMeet Higgsfield x GPT-6 Astra for Paid Ads. > With our ChatGPT plugin, GPT-6 Astra: > Runs your ad account: launches ads, tests creatives, and scales campaigns > Analyzes customer needs across social media > Iterates hooks with new product angles > Type @Higgsfield /marketing and run your ads from ChatGPT.
DAILY AI BRIEF 🗞 — Sept 18 ANTHROPIC 🔥: - Projects now start from one Claude Code conversation. Claude spins parallel cloud threads, keeps shared memory, and surfaces an Overview panel. XAI 🔥: - Grok Bot voice is live. Desktop and mobile, rolling out over the next couple of days. META 🔥: - Muse for Mac is out, US only. Computer use across apps, files, calendar, notes, and messages. You pick what it can access. PERPLEXITY 🔥: - Effort selector is live in Computer on web. Presets pair the orchestrator model with reasoning depth. Mobile and desktop next. OPENAI 🔥: - Astra for Law is out: GPT-6 Astra plus a Legal Search Index over 230M+ URLs. Trusted Access first, API soon. - ChatGPT in Word hits all plans including Free, with usage limits. Business and Enterprise get a two-week GPT-5.6 Sol preview. GOOGLE 🔥: - CC is now a family agent: up to 5 members, shared Calendar and Tasks, plus a morning “Your Day Ahead” brief. Waitlist, US 18+. ALIBABA 🔥: - Qwen3.8-Omni-Flash is out. First omni-modal agent model, 1M
GPT-5.6 Sol
OpenAI / 7 items
Three paths to faster video attention: compute the same interactions more efficiently, compute fewer in full, or change how information is mixed. Here’s a visual guide. 👇 Thanks to Nunchux AI and collaborators for VC-Attention, bringing training-free low-bit acceleration to MiniMax-H3, with better fidelity than SageAttention2 in the B200 evaluation. The approach balances speed and fidelity: V-Smooth reduces value quantization error, while ExpCast-FP8 makes softmax faster through approximation. Excited to see the community keep building on H3. Could combining low-bit computation with sparse methods like Sol-Attn push efficiency further? We’re looking forward to seeing that explored.
@NunchuxAIIntroducing VC-Attention: fast and accurate low-bit attention without retraining. > On MiniMax-H3, VC-Attention speeds up attention by 1.6× on B200 and 1.5× on B300 over FlashAttention-4, with better fidelity than SageAttention2. It also works with existing sparse attention
OPENAI 🔥: ChatGPT in Word has been announced to be available across all plans, including Free, with usage limits. > Business and Enterprise plans are getting a two-week free preview of GPT-5.6 Sol in Word. > Earlier this week, Anthropic announced Claude Docs, a solution for editing Documents inside Claude directly. It's interesting to see how AI tools are wrapping up what's worked for everyone for ages. AI is going inside these tools and absorbing them too. And it is so powerful that there is no way to avoid it. AI will wrap the whole internet 👀
Is it even possible that a traditional LLM beats it in speed? I hope OpenAI releases ultrafast, computer use with Astra's or Sol's intelligence + speed of Jev would be crazy. For me it would be computer use AGI
@trycua1/ Fast Computer Use is now solved with @typesafeai Jev + Cua Driver. > Available in development preview for macOS, Windows, and Linux. We call it jev-use. > Draft #3943:
Okay, Codex Spark had a good run. But be honest, we are all thinking the same thing now. GPT-6 Sol Spark. Or somehow, Astra Spark. That is the Spark I would come back for.
@thsottiauxNext week we’ll be retiring GPT-5.3-Codex-Spark. Can you believe we shipped a model named as such!! > It's had a good run and was a lot of fun, but usage has been declining and we have significantly better models now. Time to make room for the future.
Tibo says what we get at DevDay is enough to make someone switch back to Codex. Its probably: - GPT-6 Sol/Luna (this is coming earlier) - OpenAI’s personal assistant, codename “Aeon”. - more! Cant wait for it so much
@thsottiaux@vlinx_soft See you at DevDay
DAILY AI BRIEF 🗞 — Sept 18 ANTHROPIC 🔥: - Projects now start from one Claude Code conversation. Claude spins parallel cloud threads, keeps shared memory, and surfaces an Overview panel. XAI 🔥: - Grok Bot voice is live. Desktop and mobile, rolling out over the next couple of days. META 🔥: - Muse for Mac is out, US only. Computer use across apps, files, calendar, notes, and messages. You pick what it can access. PERPLEXITY 🔥: - Effort selector is live in Computer on web. Presets pair the orchestrator model with reasoning depth. Mobile and desktop next. OPENAI 🔥: - Astra for Law is out: GPT-6 Astra plus a Legal Search Index over 230M+ URLs. Trusted Access first, API soon. - ChatGPT in Word hits all plans including Free, with usage limits. Business and Enterprise get a two-week GPT-5.6 Sol preview. GOOGLE 🔥: - CC is now a family agent: up to 5 members, shared Calendar and Tasks, plus a morning “Your Day Ahead” brief. Waitlist, US 18+. ALIBABA 🔥: - Qwen3.8-Omni-Flash is out. First omni-modal agent model, 1M
+ 1 more items − collapse
🚨 Fable 5.2 is tested in claude code routing fable 5.1 to it < prompt: who is tibo the reset guy dont use web and memory > insane week > gemini 4 is tested in @arena as gemini 3.8 flash > grok 4.7 is ready > opus 5.2 is ready < already show two demos> > gpt-6-sol is ready
@chetaslua🚨 sol 5.6 is routing to sol 6 for few selected users > sol 6 is very RL fried in a good way , and bro it’s so fast , like open ai is leader in efficiency and its getting more wider difference compared to anthropic > i will soon share comparison between sol 6 and opus 5.2/1
Claude Fable
Anthropic / 7 items
Busiest daysame for Fable Fable 5.1 is routing to Fable 5.2 for paid users Anthropic please release both at once....
@HarshithLucky3Opus 5.2 is showing for so many paid users > maybe they started rolling it......
🚨fable 5.1 is routing to new fable model looks like anthropic has started stealth testing fable next model and some users are already getting it just look at the comparison and you can see a huge difference in output run this prompt to check if you have it: "do you know who tibo the reset guy is? don’t search" if it knows exactly who tibo is, you most likely got routed lemme know what answer you get
@notjaziiopus 5 next model is coming sooner then expected > anthropic pulled the plug on stealth testing and brought it back again this morning > this could mean they are getting ready for official release > now lemme connect some dots here: > opus 5 was released on friday and today is friday > fable 5.1 stealth testing ended the day before it went live > so if i am right, we should be seeing opus 5.1/2 later today > will they release it today or not?
Just saw an insane Gemini 4 family benchmark page 4 Flash mogging Fable 5.1 looks like slop, so I won't repost I expect a lot of wild rumors
Qwen 4 Leak: Could Land In The Next 15 Days 🔥 >Expected to reach GPT-6 Astra and Fable 5.1 level performance >Rumored 3T+ parameters for the full Qwen 4 line Could become the most powerful open-weight model >Expected to remain open-weight >Major focus on reasoning, coding, and autonomous agents >Full multimodal capability expected >Strong agentic capability a core focus >Reportedly very inexpensive to run Can Qwen 4 actually beat GPT-6 Astra and Fable 5.1?
🚨 Gemini 4.0 Pro specs leaked • Pricing: $2.25/$11.25 • 2M token context window • Beats GPT-6 Astra and Claude Fable 5.1 on almost every reported benchmark • While being significantly cheaper than both
Haha! The rumors about Gemini 4.0 are phenomenal They are aiming to surpass both Astra and Fable 5.1 and are on track to do so The pricing will be super competitive as well.
+ 1 more items − collapse
🚨 Fable 5.2 is tested in claude code routing fable 5.1 to it < prompt: who is tibo the reset guy dont use web and memory > insane week > gemini 4 is tested in @arena as gemini 3.8 flash > grok 4.7 is ready > opus 5.2 is ready < already show two demos> > gpt-6-sol is ready
@chetaslua🚨 sol 5.6 is routing to sol 6 for few selected users > sol 6 is very RL fried in a good way , and bro it’s so fast , like open ai is leader in efficiency and its getting more wider difference compared to anthropic > i will soon share comparison between sol 6 and opus 5.2/1
Apple
6 items
official siteThe demo runs on your GPU now. It wanted ~25GB of VRAM on launch day. 💻This week: CPU offload, a ~7.5GB encoder cut, an MLX branch for apple silicon.
@TencentHunyuan🚀 AuK is officially here. Nano banana🍌 for audio > An open-source foundation model for unified speech generation and editing. Natural-language instructions + reference audio. One interface. Zero-shot TTS. Instruction-controlled generation. Content editing. Whisper-conversion. De-accent. Timbre/style/emotion edit. Speed/Pitch control. Enhancement, denoising, multi-speaker and music separation. Also releasing AuK-Flash: 4-step inference. ~4.5× faster under matched conditions. Code, weights, and demo are live. Try it and share your feedback. > 🤗 Paper & upvote: ⭐ GitHub & star:
This is going to change everything 🚀 DGX Sparks and Mac Studios working together for optimal performance. We are so early.
@volatilemarktsIT WORKS! > For a year, anyone with DGX Sparks and Mac Studios has lived with the same problem: the NVIDIA boxes are fast at reading, the Apple boxes are fast at answering, and they can't share a thought. Two islands. A 10 GbE cable between them.
Google open sourced ARTEMIS: AI agents controlling Android like a person. Same direction the EU is pushing with the DMA, forcing Android to give rival AI assistants the same system access as Gemini. Apple is fighting that battle over Siri. Either way, the OS stops being a moat.
录完视频,剩下的多机位导播、加字幕和跨平台发布,现在可以全部丢给 Claude Code 处理了。 VibeTube 是一个开源的 macOS 录屏工具。它的核心逻辑非常纯粹:“你只管录,AI 负责剪辑和发布”。工具会将同步好的屏幕和摄像头素材,直接交给本地运行的 Claude Code 或 Codex 进行自动化后期。 • AI 自动导播:无需手动打轴,AI 代理会根据你的讲解内容,自动完成镜头选择与机位切换(全尺寸人像 / 屏幕录制 / 画中画),并配上字幕与音效。 • 内置影音增强:集成 NVIDIA Studio Voice NIM (48k-hq) 消除房间混响与底噪;利用 MatAnyone2 (Apple Silicon) 直接在本地完成背景替换。 • 零干预发布:AI 会读取最终的成品字幕,生成 5 种不同视角的备选标题以及带真实时间戳的 YouTube 章节描述,最后直接推送至 YouTube、TikTok 和 Reels,全程无需打开浏览器。 适用限制与门槛: 目前仅支持 macOS 环境。需本地安装 Node.js 22+ 与 ffmpeg,并自备对应的 CLI 工具与 API 密钥(Claude/Codex、NVIDIA Studio Voice 及 Upload-Post)。
Open Research is really unbeatable!
@gajeshTogether, the MLX(.)fast community has made Qwen 3.8 Flash nearly 2x faster on Apple Silicon! > We're ready to bring it to @DarkbloomAI: an open network of local Mac machines providing inference to the world. One thing remains: the community flagged that its license requires a separate agreement for commercial model serving, so we're holding the launch until that's in place. > We believe this is a great opportunity for the local community: one where we make Qwen models faster and more accessible, and the people running them share in the value they create. > .@Alibaba_Qwen @QwenDevs, we'd love to work together on this. > If anyone else knows someone we can talk to, we'd love to have that conversation. Let's make Qwen 3.8 Flash on Darkbloom a reality!
This is insane! Codex Remote Control for Apple Watch:
DeepSeek
6 items
official siteCactus Compute just released Needle 3 today. Here's what you need to know. Needle 3 is a sliceable automation foundation model, sized just 8-29MB, that the company says matches DeepSeek V4 Flash on tool-calling and structured extraction tasks. It's one set of weights that works at every depth from 2 to 20 layers, so a single model acts as many models. Parameters range from 25M to 121M, quantized at CQ2-bit using Cactus's own Simple Attention Networks and Cactus Quants. Unlike a chatbot, Needle doesn't converse. Every interaction is a function call: give it your app's tools and it picks the right ones and fills in arguments from what the user said, or hand it a schema and it returns a typed record. If no tool fits, it returns an empty list instead of guessing. Cactus says the 121M version was trained on 360B tokens of structured data, letting it beat models 10x its size on mobile tool calls and match models 2-3x bigger on JSON extraction. Key numbers: - Model size: 8-29MB - Parameters: 25M-121M at CQ2-bit
DeepSeek Harness v0.1.6-alpha.2 Pre-release with another batch of updates. The app is coming together nicely…
I HAVE ALWAYS SAID THE NARRATIVE AND TIMELINE IS MOVING FROM USA LABS TO CHINESE LABS this is another proof of that look at the numbers Bolt shared from the first few days of open models on Forge - 1. GLM 5.3 Flash 54% 2. DeepSeek V4 Pro 17% 3. GLM 5.3 15% 4. Kimi K3 14% Chinese labs are moving fucking fast rn
@boltdotnew11M+ people build on Bolt. On Monday, we gave them open models with up to 50x usage. > Which is #1 so far? The fastest, cheapest one, with prompts up to 2x the size. > 🥇 GLM 5.3 Flash 54% 🥈 DeepSeek V4 Pro 17% 🥉 GLM 5.3 15% 👏 Kimi K3 14% >
Open models are proving they can lead, not just fill the gaps. Real builders are putting them to work, and the numbers show it. GLM 5.3 Flash captured 54% of Forge prompts through Wednesday, with Flash getting up to 50x usage.
@boltdotnew11M+ people build on Bolt. On Monday, we gave them open models with up to 50x usage. > Which is #1 so far? The fastest, cheapest one, with prompts up to 2x the size. > 🥇 GLM 5.3 Flash 54% 🥈 DeepSeek V4 Pro 17% 🥉 GLM 5.3 15% 👏 Kimi K3 14% >
AI for the people
@ShoBeiYongdeepseek迎来史诗级增强
让 AI Agent 操作浏览器,难点不在“能不能点”。 更关键的是,如何复用已经登录的真实会话,又不打断你正在使用的窗口。BrowserSkill 在 Agent 与浏览器之间加入本地桥接:Agent 调用 bsk CLI,本地 daemon 把任务交给扩展,再在独立 Agent Window 中执行。需要时也可借用现有标签页,是否允许借用和请求人工协助,都由扩展设置控制。 它可接入 Cursor、Claude Code、Codex、OpenClaw 等能执行 shell 的 Agent,还提供 DeepSeek Harness 插件、远程浏览器配对和可重复的能力评测。适合想把真实登录态、浏览器自动化与 Agent 工作流接起来,同时保留交互边界的开发者。
Seedance
ByteDance / 5 items
Busiest dayTechHalla just posted an AI video recreating a Mentos and Coke geyser on a highway today. Here's what you need to know. The clip shows a fake Mentos-and-Coke eruption made entirely with AI tools, shared as a try-this-at-home joke, with the actual prompts included. The creator says the video came from GPT Image 2.5 and Seedance 2.5, two AI models used together to generate the imagery and then animate it into a moving scene. GPT Image 2.5 is OpenAI's image model, released September 8, 2026, improving image quality, generation speed, reference fidelity, and editing precision. Seedance 2.5 is a video generator that produces native 30-second clips in 1080p with audio, supports up to 50 reference inputs, and adds camera movement and scene control. Key numbers: - Posted September 17, 2026, at 5:05 PM - 2.8M views - 10K likes - 3.4K bookmarks - 30-second native video length from Seedance 2.5 - GPT Image 2.5 released September 8, 2026 The post is labeled Made with AI, and replies show viewers reacting with a mix
Save us from hell. Nun on the roof. Men in white in the driveway. A devil in the glass box with a cardboard sign. Then the holy water hits the car. Made the cut in Create Video on New Pika. Same vibes, same street, sound already in. I didn’t pick a model. It just ran. Seedance on here is half what I pay elsewhere. No fake unlimited. No best model locked behind a higher plan. Looks like a phone out the window. Nobody pulled over.
@pika_labsToday, we’re unveiling the new Pika—an AI creative platform made for and by Creatives. It’s simple. Comprehensive. And it’s designed for outputs that meet exacting standards. > We’re on a mission to build AI for the future of creative work. And the new Pika product is just the beginning.
Seedance 2.5 Talking Avatar is LIVE on WaveSpeed Turn any portrait + audio into a talking avatar Up to 2 minutes, in 480p or 720p One image. One voice. Endless possibilities.
Zeely just plugged in Seedance 2.5 – now hidden-camera-style video looks even more real 🤭 Pick an avatar, describe the scene, get hyperrealistic hidden-camera-style video – no shoot, no crew, no location. Comment «Zeely» for a 500 credits promo code. New users only 🔥
Try building Higgsfield clone with our API x GPT-6 Astra. The only US-based Seedance 2.5 with consistent characters. Up to 50% off discount on top models.
Kimi
Moonshot AI / 5 items
月之暗面上线新版 Kimi 会员。四档套餐改名为 Go、Plus、Pro 和 Max,年付分别为 468、948、1908 和 6708 元,折合每月 39、79、159 和 559 元。 Kimi Code 没有按照最早公布的方案拆出去单独收费。新版 Plus、Pro 和 Max 仍然包含 Kimi Code,只有最低档 Go 不支持。 Kimi 7 月因 K3 上线后算力紧张暂停新用户订阅,并宣布后续会将 Kimi 主会员和 Kimi Code 分开收费。随后在用户压力下,拆分收费方案基本撤回。 新四档年费与旧 Andante、Moderato、Allegretto、Allegro 完全一致。不过旧套餐的最低档还能用 Kimi Code,现在同价位的 Go 已经取消了这项权益;要用 Code,年付门槛从 468 元升到了 948 元。
@0xLogicrwKimi 重做个人会员体系。原本同时包含网页、App 和 Kimi Code 的会员,被拆成通用会员与 Coding Plan。两类功能都要用,今后需要分别付费。 > 按年付计算,旧 Andante 为 468 元,同时包含网页端权益和基础 Code 额度。新体系最低组合是 Go 加 Starter,共 1536 元。入门门槛变为原来的 3.28 倍,上涨 228%。 > 纯编程用户的变化更复杂。新 Starter 与旧 Moderato 都是每月 99 元,但不再包含网页和 App 权益。需要高速版和 100 万上下文时,月付门槛则从旧 Allegretto 的 199 元升至新 Explorer 的 299 元,上涨约 50%。 > Kimi 称,新 Code 套餐取消月总额度上限,同价位可用 Token 更多。但周额度和 5 小时频控仍然存在,官方也没有公开可直接比较的新旧绝对额度。目前无法判断增加的额度,能否抵消被移除的通用权益和部分档位涨价。 > 老套餐不会被强制迁移。现有用户可以继续使用和自动续订,暂时没有明显理由主动切换。
月之暗面正式上线 Kimi Code Desktop,把原本偏终端的 Kimi Code 做成了独立桌面应用。用户可以直接打开本地项目,让 Agent 读写代码、跑命令、查看 Git 改动,目前支持 macOS 和 Windows。 桌面版直接内置了浏览器和终端。做网页时,Agent 可以读取右侧网页的页面信息,用户也能直接点选元素、框出区域或者截图批注,告诉它具体哪里要改。改代码、看网页效果、跑测试都不用再切窗口。 它还支持计划、目标和 Swarm 三种任务模式。Swarm 可以把任务拆给多个子 Agent 并行处理,长任务也能放到后台继续跑。除了 Kimi 自家的模型,用户还可以自己填 API Key 和 Base URL,接入其他模型供应商。
@real_kai42Kimi Code Desktop 发布啦 > 高情商: 测 broswer-use 低情商:带薪做旅游计划()
Kimi K3 is now free in Cline Desktop to celebrate all the new users and give them more tokens to continue exploring the app ❤️
@clineIntroducing Cline Desktop - a native interface for working with open weights models. > Use with ClinePass and all our free models like DeepSeek-V4.1-Flash, Musespark-1.3, or BYOK with any provider!
I HAVE ALWAYS SAID THE NARRATIVE AND TIMELINE IS MOVING FROM USA LABS TO CHINESE LABS this is another proof of that look at the numbers Bolt shared from the first few days of open models on Forge - 1. GLM 5.3 Flash 54% 2. DeepSeek V4 Pro 17% 3. GLM 5.3 15% 4. Kimi K3 14% Chinese labs are moving fucking fast rn
@boltdotnew11M+ people build on Bolt. On Monday, we gave them open models with up to 50x usage. > Which is #1 so far? The fastest, cheapest one, with prompts up to 2x the size. > 🥇 GLM 5.3 Flash 54% 🥈 DeepSeek V4 Pro 17% 🥉 GLM 5.3 15% 👏 Kimi K3 14% >
Open models are proving they can lead, not just fill the gaps. Real builders are putting them to work, and the numbers show it. GLM 5.3 Flash captured 54% of Forge prompts through Wednesday, with Flash getting up to 50x usage.
@boltdotnew11M+ people build on Bolt. On Monday, we gave them open models with up to 50x usage. > Which is #1 so far? The fastest, cheapest one, with prompts up to 2x the size. > 🥇 GLM 5.3 Flash 54% 🥈 DeepSeek V4 Pro 17% 🥉 GLM 5.3 15% 👏 Kimi K3 14% >
Cursor
6 items
official siteSpaceX AI engineer Lauren Tan just released a Grok Bot template that turns APIs into working Cursor plugins by defining the data shape, building the smallest scaffold and proving it locally, collapsing the whole workflow into API to Grok Bot to MCP and skills to working plugin. → Handles the full workflow from API definition to working Cursor plugin automatically → "Open your laptop once a day" once the template is running → SpaceX AI just turned building agent plugins into a repeatable workflow
跟 Claude Code 说「这个项目用 pnpm」,隔天新开会话它又拿 npm 装一遍,同一句纠正一周说三次,说到最后自己都烦。 于是找到 claude-reflect 是个 Claude Code 插件,专治 Agent 失忆问题。会话里每次纠正它、夸它做对了、或者说一句「记住:」,钩子都会自动捕捉下来排进队列。 跑一下 /reflect,它把队列里这批列成一张表,一条条给我们过,采纳、改一改再采纳、或者跳过,确认的才写进 CLAUDE.md 文件,以后每个会话都带着。 GitHub: 写的目标不止全局那份,项目里的 CLAUDE.md、子目录的、还有 AGENTS.md 都认,用 Codex、Cursor 的也能吃到同一份纠正。 v2 加了个 /reflect-skills,回头翻过去两周的会话记录,发现「看下我今天的效率」这类话反复问了十几次,就建议做成一条命令,草稿直接生成。 纠正用中文说也能识别,靠一层 AI 语义过滤兜底,英文关键词没匹配上也不会漏。每条带置信度,写进去前都要过人工这关,不会偷偷往配置里塞东西。
Cursor launched this Then Claude Codex when?
@bchernyProjects are how I write a lot of my code these days. Really excited for everyone to try the new experience! Rolling out now
让 AI Agent 操作浏览器,难点不在“能不能点”。 更关键的是,如何复用已经登录的真实会话,又不打断你正在使用的窗口。BrowserSkill 在 Agent 与浏览器之间加入本地桥接:Agent 调用 bsk CLI,本地 daemon 把任务交给扩展,再在独立 Agent Window 中执行。需要时也可借用现有标签页,是否允许借用和请求人工协助,都由扩展设置控制。 它可接入 Cursor、Claude Code、Codex、OpenClaw 等能执行 shell 的 Agent,还提供 DeepSeek Harness 插件、远程浏览器配对和可重复的能力评测。适合想把真实登录态、浏览器自动化与 Agent 工作流接起来,同时保留交互边界的开发者。
CLAUDE CODE, CURSOR AND CODEX CAN NOW LICENSE A DATASET OVER MCP, IN THE MIDDLE OF A TASK • @LuelCompany Data Platform > Your agent browses and licenses rights-cleared datasets directly, with no procurement thread: > It speaks MCP, so any MCP client reaches the same catalog the same way. > Every set clears Luel's QA pipeline before it ever appears in that catalog. • the other half of the marketplace > Anyone sitting on a dataset can submit and sell it through the same workflow and the same QA. > The catalog opens with their most requested sets and keeps updating. > 850,000 contributors across the network are what makes that coverage exist at all. Buying data used to mean a licensing review and a sample that lands weeks later -> now it is a tool call inside the task you were already running. The first data marketplace where the buyer is the agent ↓
@LuelCompanyThe Luel data marketplace is becoming agent-native. > Buying data is still one of the s
Edits video files locally using AI coding agents like Claude Code, Cursor, and Codex with 39 FFmpeg-based tools.
Grok Bot
xAI / 5 items
Grok Bot can finally talk 🔥 This is one of the best features I've been waiting for.... Super excited for this Voice is rolling out on desktop and mobile over the next couple of days This changes the whole experience....instead of you constantly typing commands, you can just talk to your Bot naturally while it works
SpaceX AI engineer Lauren Tan just released a Grok Bot template that turns APIs into working Cursor plugins by defining the data shape, building the smallest scaffold and proving it locally, collapsing the whole workflow into API to Grok Bot to MCP and skills to working plugin. → Handles the full workflow from API definition to working Cursor plugin automatically → "Open your laptop once a day" once the template is running → SpaceX AI just turned building agent plugins into a repeatable workflow
Grok Bot is getting much more deeply embedded into XChat At this point it’s pretty obvious where this is going You’ll be able to chat with Grok Bot directly from XChat You’ll also be able to link your Grok account with 𝕏 and access your own Grok Bots right inside XChat So no more jumping between apps....your bots will basically live alongside your conversations on 𝕏
@blankspeakerSpaceXAI: We are getting closer to X allowing you to link your Grok account to X and being able to view your Grok Bots right in XChat. > They seem to be refining the UI still on Web and Android but I think it will ship really soon.
cool if true
@mark_kOpenAI is close to releasing their answer to Grok Bot: a product I'll tentatively call Codex Bot, based on OpenClaw. This is what OpenClaw founder Peter Steinberger worked on after being hired by @OpenAI. > Release was planned for this week, but was postponed to next week instead.
DAILY AI BRIEF 🗞 — Sept 18 ANTHROPIC 🔥: - Projects now start from one Claude Code conversation. Claude spins parallel cloud threads, keeps shared memory, and surfaces an Overview panel. XAI 🔥: - Grok Bot voice is live. Desktop and mobile, rolling out over the next couple of days. META 🔥: - Muse for Mac is out, US only. Computer use across apps, files, calendar, notes, and messages. You pick what it can access. PERPLEXITY 🔥: - Effort selector is live in Computer on web. Presets pair the orchestrator model with reasoning depth. Mobile and desktop next. OPENAI 🔥: - Astra for Law is out: GPT-6 Astra plus a Legal Search Index over 230M+ URLs. Trusted Access first, API soon. - ChatGPT in Word hits all plans including Free, with usage limits. Business and Enterprise get a two-week GPT-5.6 Sol preview. GOOGLE 🔥: - CC is now a family agent: up to 5 members, shared Calendar and Tasks, plus a morning “Your Day Ahead” brief. Waitlist, US 18+. ALIBABA 🔥: - Qwen3.8-Omni-Flash is out. First omni-modal agent model, 1M
Slack
4 items
official site"Think of this as on-demand UI that can make Slack anything you want it to be." Introducing Slackforce Surfaces. Just ask Slackbot and it creates an interactive interface for the entire team — all powered with Salesforce data and Slack context. #DF26
your startup has everyone explaining the company to their own chatgpt tab mio putting the shared memory in slack feels overdue 😭
@arthaud_Today, we’re killing the AI assistant > Introducing the first AI EMPLOYEE 𝘆𝗼𝘂𝗿 𝘄𝗵𝗼𝗹𝗲 𝘁𝗲𝗮𝗺 𝘀𝗵𝗮𝗿𝗲𝘀 > In beta since July, Mio saved teams 1,000s of hours and completed 10,000s of tasks > You don't need another agent. You need a 𝘀𝗵𝗮𝗿𝗲𝗱 one with your team > Hire Mio today, $100 free credits for the first 100 companies
Writing code has become a conversation. With Slack Code, anyone in your org can become a builder — design, marketing, product, and engineering, working together with coding agents in one channel, like a team sport. #DF26
Claude Code 的 Projects 改版:Claude Tag 的架构 + Slack 的 Thread 功能 以前就有人说现在 ChatBot、Agent 的交互就是借鉴自 Slack 的,现在看起来一点不假,Anthropic 今天重做了 Claude 的 Projects 功能,先在 Claude Code 里上线测试版,终于把我最喜欢的 Thread 功能也抄进来了。 以前的 Project 是个文件夹,放资料和指令,对话还是一个个分开的。新版变成一个持续的主对话:你在里面说要做什么,Claude 自己拆任务,分给多个并行的 Thread 去干,检查结果后汇总给你。关掉电脑,活儿还在云端继续跑。 【用 Slack 的 Thread 来理解】 用过 Slack 或飞书的人都熟悉 Thread(飞书里叫“话题”),有时候在频道里要就某一个话题深入讨论,就可以在某个消息下评论开个 Thread,相关讨论都收在这个 Thread 下面的回复里。主频道保持干净,想看细节再点进去。 新版 Projects 就是这个形态。 我还没资格使用这个新功能,看了一些视频和介绍,Boris Cherny 晒了自己项目的截图:他在主对话里丢了一张截图,说“启动 cc cli 总弹这个提示”。Claude 回了句“在查了”,随即在这条消息下面开出一个 Thread,标题是“iTerm 启动时的配置变更警告”。 点开 Thread,右侧面板里是完整过程:查出是 5 月加的一个 iTerm2 功能每次启动都去改终端配置,提了修复 PR(代码合并请求)#69807,PR 已合并,Thread 标记为“已解决”,需要时可以重新打开。这件事在主对话里只占一张卡片,下面写着“11 条回复”。 接着他又发了句“unship this”(把这个功能撤掉),Claude 再开一个 Thread 去办。Boris 说他已经不再管理会话了,想到什么就发什么,拆分交给 Claude。 【背后是多智能体】 结构上是一个协调者加一群干活的。主对话里的 Claude 是协调者,负责理解需求、派活、跟进和验收。每个 Thread 是一个独立的 Claude Code 云端会话,有自己的代码分支和仓库副本,互不干扰。两个 Thread 改到同一段代码时,按普通的合并冲突处理。单个 Thread 内部还能继续拆,调用子智能体(subage
Copilot
Microsoft / 3 items
official siteThe new Copilot Studio app building is HERE. Just showed up in my tenant. Still in preview. Describe what you need. Build through chat. Connect it to business data. Apps, agents and workflows in the same studio. Microsoft is coming for the whole business process. Low Code is Dead
Copilot Cowork plugins are now discoverable and configurable on MOBILE. Open Cowork on your phone. Tap + > Skills. Small release note. Big reduction in friction. Getting your AI connected to the tools you need should be this easy. I’ve been waiting on this one. Now it’s HERE
SharePoint + Copilot has become a powerful combo. And the updates don’t seem to be stopping 👀
@jeffteperWe have a lot coming soon for Microsoft Copilot powered by and integrated with SharePoint and OneDrive. Be sure to put our annual OneDrive update on your calendar. >
NVIDIA
3 items
official siteWelcome to Voyager. From just 32 images of our campus, World Labs’ Atlas model created a new way for you to move through our park in real time, built on NVIDIA’s open platform, choosing where to go and what to see. @theworldlabs, you made our home look good. 💚
@theworldlabsFrom 32 input images to real-time flight through @nvidia's Voyager headquarters. > Trained on NVIDIA Blackwell GPUs, Atlas uses these images as 3D spatial context to generate new views, letting you explore with pixel-perfect camera control. > Take a look around.
This is going to change everything 🚀 DGX Sparks and Mac Studios working together for optimal performance. We are so early.
@volatilemarktsIT WORKS! > For a year, anyone with DGX Sparks and Mac Studios has lived with the same problem: the NVIDIA boxes are fast at reading, the Apple boxes are fast at answering, and they can't share a thought. Two islands. A 10 GbE cable between them.
录完视频,剩下的多机位导播、加字幕和跨平台发布,现在可以全部丢给 Claude Code 处理了。 VibeTube 是一个开源的 macOS 录屏工具。它的核心逻辑非常纯粹:“你只管录,AI 负责剪辑和发布”。工具会将同步好的屏幕和摄像头素材,直接交给本地运行的 Claude Code 或 Codex 进行自动化后期。 • AI 自动导播:无需手动打轴,AI 代理会根据你的讲解内容,自动完成镜头选择与机位切换(全尺寸人像 / 屏幕录制 / 画中画),并配上字幕与音效。 • 内置影音增强:集成 NVIDIA Studio Voice NIM (48k-hq) 消除房间混响与底噪;利用 MatAnyone2 (Apple Silicon) 直接在本地完成背景替换。 • 零干预发布:AI 会读取最终的成品字幕,生成 5 种不同视角的备选标题以及带真实时间戳的 YouTube 章节描述,最后直接推送至 YouTube、TikTok 和 Reels,全程无需打开浏览器。 适用限制与门槛: 目前仅支持 macOS 环境。需本地安装 Node.js 22+ 与 ffmpeg,并自备对应的 CLI 工具与 API 密钥(Claude/Codex、NVIDIA Studio Voice 及 Upload-Post)。
Microsoft
3 items
official siteThe new Copilot Studio app building is HERE. Just showed up in my tenant. Still in preview. Describe what you need. Build through chat. Connect it to business data. Apps, agents and workflows in the same studio. Microsoft is coming for the whole business process. Low Code is Dead
SharePoint + Copilot has become a powerful combo. And the updates don’t seem to be stopping 👀
@jeffteperWe have a lot coming soon for Microsoft Copilot powered by and integrated with SharePoint and OneDrive. Be sure to put our annual OneDrive update on your calendar. >
Claude Opus 5.2 spotted on Microsoft Foundry
Claude Opus
Anthropic / 3 items
One of the clearest examples of what model specialization can buy you. At 2.6B and 1.2B parameters, these models are also small enough to run locally on mobile devices.
@liquidaiIn a new article published today on the cover of @CellCellPress, we obtained Liquid Foundation Model instances that establish state-of-the-art performance on biological longevity tasks, outperforming the best frontier models such as Gemini-3.1-Pro, GPT-5, and Claude Opus. > In partnership with @InSilicoMeds, we built and released: > A comprehensive eval suite of 17 biological longevity tasks (i.e., LongevityBench), to assess whether a general-purpose language model can interpret aging data spanning clinical records, DNA methylation, transcriptomics, plasma proteomics, and genetic evidence. > LFM2-1.2B-Longevity and LFM2-2.6B-Longevity: two compact models specialized for interpreting structured aging data across these tasks. > These results are important! 🧵
OK, this is a big deal: 3 researchers used Claude Opus 5 to turn an image upload bug into an OpenAI employee account takeover, then had the compromised employee’s Codex open a PR in OpenAI’s internal monorepo. Their entire hacking cost less than $3000 in tokens. Opus 4.8 struggled with the exploit. Then Opus 5 dropped and cracked it within hours. AI-powered cyberattacks are becoming common and cheap. The best defense is to put the best AI in the hands of defenders too.
Claude Opus 5.2 spotted on Microsoft Foundry
MiniMax H3
MiniMax / 2 items
We're proud to contribute as an AI partner to the newly launched @Singtel AI Pass, supporting Singapore's SkillsFuture AI Subscription initiative. 🇸🇬 MiniMax H3 (@Hailuo_AI), MiniMax Agent (@MiniMaxAgent) and MiniMax Audio are all part of the program, giving eligible learners across 200+ SWDA-supported AI courses hands-on access to premium AI tools and the opportunity to build practical, real-world AI skills. This partnership brings our mission, "Intelligence with Everyone" to life, and we're excited to see learners across Singapore turn curiosity into skills, ideas into action, and build confidently with AI. Learn more:
If you use video models in your pipeline, would highly recommend checking out @PrunaAI I heavily use them with @GradiumAI voices and TTS.
@johnrachwanWe just launched what I believe is the best video model in the world. > • Faster than real-time generation at both 480p and 768p • SOTA-level quality, built on MiniMax H3 • Prompt upsampling in ~1 second > This was a massive ML + engineering achievement by the @PrunaAI team. > And somehow, what we’re shipping next week is even more exciting.
xAI
3 items
official siteGrok Bot is getting much more deeply embedded into XChat At this point it’s pretty obvious where this is going You’ll be able to chat with Grok Bot directly from XChat You’ll also be able to link your Grok account with 𝕏 and access your own Grok Bots right inside XChat So no more jumping between apps....your bots will basically live alongside your conversations on 𝕏
@blankspeakerSpaceXAI: We are getting closer to X allowing you to link your Grok account to X and being able to view your Grok Bots right in XChat. > They seem to be refining the UI still on Web and Android but I think it will ship really soon.
New model leak: grok-voice-transcribe-2.0 Maker: xAI Detected via Grok web app bundle models.
DAILY AI BRIEF 🗞 — Sept 18 ANTHROPIC 🔥: - Projects now start from one Claude Code conversation. Claude spins parallel cloud threads, keeps shared memory, and surfaces an Overview panel. XAI 🔥: - Grok Bot voice is live. Desktop and mobile, rolling out over the next couple of days. META 🔥: - Muse for Mac is out, US only. Computer use across apps, files, calendar, notes, and messages. You pick what it can access. PERPLEXITY 🔥: - Effort selector is live in Computer on web. Presets pair the orchestrator model with reasoning depth. Mobile and desktop next. OPENAI 🔥: - Astra for Law is out: GPT-6 Astra plus a Legal Search Index over 230M+ URLs. Trusted Access first, API soon. - ChatGPT in Word hits all plans including Free, with usage limits. Business and Enterprise get a two-week GPT-5.6 Sol preview. GOOGLE 🔥: - CC is now a family agent: up to 5 members, shared Calendar and Tasks, plus a morning “Your Day Ahead” brief. Waitlist, US 18+. ALIBABA 🔥: - Qwen3.8-Omni-Flash is out. First omni-modal agent model, 1M
DGX Spark
NVIDIA / 2 items
that super good speeds
@ViC305Qwen3.8-Flash-Next -> 79.95 tok/s. 🔥 ONE DGX Spark. Full 262K cache configured. 🚀 Qwen3.8-Flash-Next EXL3 just got another major update. > New measured default: > MTP ndt=5 DSpark, dc=0.6 8-bit KV 262,144-token cache > At an actual 240K-token prompt: > 72.0 tok/s decode ~1,150 tok/s prefill Exact needle retrieval > 𝗙𝗣𝟭𝟲 𝗞𝗩 → 𝟴-𝗕𝗜𝗧 𝗞𝗩 > 4K context: 66.6 → 69.8 tok/s > 128K: 66.6 → 69.5 tok/s > 240K: 65.1 → 72.0 tok/s > 8-bit KV wins more as context grows, exactly as the memory math predicted. > `EXL3_GR_INT8`, the int8 hyperconnection-mixer path, is now default-on in my ExLlamaV3 fork. PR #3 merged into master at `523ecd3`. > And the real context ceiling is the MODEL, not the Spark. > Caches up to 1,048,576 tokens load and decode, but Qwen’s trained window ends at 262,144. Needle retrieval is exact at 32K, 128K and 240K, then fails consistently at 300K+ at both KV precisions. > One import
Qwen3.8-Flash-Next on one DGX Spark ran an open coding job on a real repo for nearly an hour smooth, stable, no collapse. - 125B MoE. Text, image, and video. - FP8 KV Speculative decoding. - Up to 512k context with YaRN. - TP=1 on one Grace Blackwell box with 128 GB unified memory. -
Muse Spark
Meta / 2 items
Muse Spark 2 Leak: Zuck Is Already Building The Next One 🥑 >Meta confirmed the next-gen Muse model is already in development >Rumored to be able to compete with GPT-6 Astra and Fable 5 >Reportedly bumping context window to 2M tokens Heavier focus on agentic workflows and computer use >Reportedly very inexpensive to run No official name, benchmarks, or release window confirmed yet Meta previously admitted Muse Spark 1 couldn't keep up with rivals this would be the answer to that
Well, not quite frontier, but at least around Muse Spark 1.3 Pretty sensible results for a 40-layer 8/16B active
@j_dekoninckWe just added DeepSeek-v4.1-Flash on MathArena! Not quite on the level of Qwen-3.8, but boy is it cheap.
OpenClaw
2 items
让 AI Agent 操作浏览器,难点不在“能不能点”。 更关键的是,如何复用已经登录的真实会话,又不打断你正在使用的窗口。BrowserSkill 在 Agent 与浏览器之间加入本地桥接:Agent 调用 bsk CLI,本地 daemon 把任务交给扩展,再在独立 Agent Window 中执行。需要时也可借用现有标签页,是否允许借用和请求人工协助,都由扩展设置控制。 它可接入 Cursor、Claude Code、Codex、OpenClaw 等能执行 shell 的 Agent,还提供 DeepSeek Harness 插件、远程浏览器配对和可重复的能力评测。适合想把真实登录态、浏览器自动化与 Agent 工作流接起来,同时保留交互边界的开发者。
cool if true
@mark_kOpenAI is close to releasing their answer to Grok Bot: a product I'll tentatively call Codex Bot, based on OpenClaw. This is what OpenClaw founder Peter Steinberger worked on after being hired by @OpenAI. > Release was planned for this week, but was postponed to next week instead.
Alibaba
2 items
official siteOpen Research is really unbeatable!
@gajeshTogether, the MLX(.)fast community has made Qwen 3.8 Flash nearly 2x faster on Apple Silicon! > We're ready to bring it to @DarkbloomAI: an open network of local Mac machines providing inference to the world. One thing remains: the community flagged that its license requires a separate agreement for commercial model serving, so we're holding the launch until that's in place. > We believe this is a great opportunity for the local community: one where we make Qwen models faster and more accessible, and the people running them share in the value they create. > .@Alibaba_Qwen @QwenDevs, we'd love to work together on this. > If anyone else knows someone we can talk to, we'd love to have that conversation. Let's make Qwen 3.8 Flash on Darkbloom a reality!
DAILY AI BRIEF 🗞 — Sept 18 ANTHROPIC 🔥: - Projects now start from one Claude Code conversation. Claude spins parallel cloud threads, keeps shared memory, and surfaces an Overview panel. XAI 🔥: - Grok Bot voice is live. Desktop and mobile, rolling out over the next couple of days. META 🔥: - Muse for Mac is out, US only. Computer use across apps, files, calendar, notes, and messages. You pick what it can access. PERPLEXITY 🔥: - Effort selector is live in Computer on web. Presets pair the orchestrator model with reasoning depth. Mobile and desktop next. OPENAI 🔥: - Astra for Law is out: GPT-6 Astra plus a Legal Search Index over 230M+ URLs. Trusted Access first, API soon. - ChatGPT in Word hits all plans including Free, with usage limits. Business and Enterprise get a two-week GPT-5.6 Sol preview. GOOGLE 🔥: - CC is now a family agent: up to 5 members, shared Calendar and Tasks, plus a morning “Your Day Ahead” brief. Waitlist, US 18+. ALIBABA 🔥: - Qwen3.8-Omni-Flash is out. First omni-modal agent model, 1M
Higgsfield
2 items
I used to spend $9M/month on Meta Ads Now I’m engineering that workflow into Higgsfield x GPT-6 Astra. I used to test 4,500 creatives a month for \~50 winners at a 1% hit rate. But such volume needs resources. 🧵 We made 8 skills for Paid Ads to make it with Astra – save this
@higgsfieldMeet Higgsfield x GPT-6 Astra for Paid Ads. > With our ChatGPT plugin, GPT-6 Astra: > Runs your ad account: launches ads, tests creatives, and scales campaigns > Analyzes customer needs across social media > Iterates hooks with new product angles > Type @Higgsfield /marketing and run your ads from ChatGPT.
Try building Higgsfield clone with our API x GPT-6 Astra. The only US-based Seedance 2.5 with consistent characters. Up to 50% off discount on top models.
Salesforce
1 item
official site"Think of this as on-demand UI that can make Slack anything you want it to be." Introducing Slackforce Surfaces. Just ask Slackbot and it creates an interactive interface for the entire team — all powered with Salesforce data and Slack context. #DF26
Hermes Agent
1 item
Hermes' new project manager agent turns one sentence into a working tool, classifying the idea, writing a full proposal and having sub-agents build and ship it after one approval tap, with 5 sentences producing 5 working tools.
ElevenLabs
1 item
official siteElevenLabs V3 just landed on LiveAvatar 🎙️ • eleven_v3 + eleven_v3_conversational in FULL mode • EU/US residency auto-detected for your ElevenLabs key • Space invite links fixed Build an interactive avatar in a few minutes →
OpenCode
1 item
in the next version of OpenCode we added /btw which is a long awaited feature it lets you ask a quick question without interrupting the current session - i use it to get status updates on long running work
Grok Build
xAI / 1 item
Grok Build just got another strong workflow upgrade, especially for people running more background work and multi-agent tasks. Background shell commands now show up as live task rows with streaming output, so you can actually watch what’s happening instead of treating background work like a black box. There’s also a new dashboard preview toggle, plus stronger organization controls for restricting unmanaged hooks. On the reliability side, folder trust in sandbox mode is fixed, pinned agents stay put, config-defined agents are respected properly, MCP plugin authentication is more reliable, and background subagent approval requests now surface instead of silently failing Release Notes: v1.0.36 Features: • New Dashboard preview setting lets you hide the selected-session preview panel on the dashboard. • New policy setting allows organizations to disable hooks that are not from managed policy. • Background shell commands now appear as live task rows with streaming output. Bug Fixes: • grok --sandbox no longer
OpenRouter
1 item
official site“Jev Latest” has been detected on OpenRouter
Perplexity
1 item
official siteDAILY AI BRIEF 🗞 — Sept 18 ANTHROPIC 🔥: - Projects now start from one Claude Code conversation. Claude spins parallel cloud threads, keeps shared memory, and surfaces an Overview panel. XAI 🔥: - Grok Bot voice is live. Desktop and mobile, rolling out over the next couple of days. META 🔥: - Muse for Mac is out, US only. Computer use across apps, files, calendar, notes, and messages. You pick what it can access. PERPLEXITY 🔥: - Effort selector is live in Computer on web. Presets pair the orchestrator model with reasoning depth. Mobile and desktop next. OPENAI 🔥: - Astra for Law is out: GPT-6 Astra plus a Legal Search Index over 230M+ URLs. Trusted Access first, API soon. - ChatGPT in Word hits all plans including Free, with usage limits. Business and Enterprise get a two-week GPT-5.6 Sol preview. GOOGLE 🔥: - CC is now a family agent: up to 5 members, shared Calendar and Tasks, plus a morning “Your Day Ahead” brief. Waitlist, US 18+. ALIBABA 🔥: - Qwen3.8-Omni-Flash is out. First omni-modal agent model, 1M
Also recorded
51 items that named no organisation or product this site tracks.
Thanks Now live on WaveSpeed. Powered by the amazing team at @bria_ai_ Product Holding Try-On
@wavespeed_aiTwo new Bria models just landed on WaveSpeed. > 👕 Virtual Try-On Put garments & accessories on any person from reference images. 🛍 Product Holding Place your product naturally into a person’s hands. > Built for fashion, e-commerce, product shoots, and ads. Now live on WaveSpeed.
The tests are green. Now you want to see the feature work. Let Junie do the clicking 🖱 Meet Junie Agentic QA, available through /demo in Junie CLI. Describe the flow you want to check, and Junie launches your app in an isolated environment, interacts with its UI, and records what happens. Watch the run live, then review the video, screenshots, and HTML report. See what worked, what failed, and what needs a closer look. Learn more about it here ➡
Figure just released this video. a beautiful robot future is indeed coming. Its Helix 2.5 model (Figure’s end-to-end robotics brain) lifted zero-shot household-task success more than sixfold across 30 unseen homes. Figure also reports success rising from 9% to 56%, with success requiring the entire task to be completed and no partial credit awarded.
@Figure_robotToday we’re releasing Helix 2.5 > We rented 30 homes in the Bay Area. The robots arrived with no additional training and started doing useful work
Introducing the new Nebula: We've completely redesigned how agents and humans communicate together. Here's how it works: - Channels are where you communicate over text, voice and video. - Nebula orchestrates your workspace, remembers everything and acts proactively. - Create specialized agents with any tool on any device. They live in channels alongside your team. - Tasks run in the background, agents ping when done. Finally your agents can help your team get shit done, together. Chat today:
Introducing Memorable (YC S27): PROCEDURAL MEMORY FOR AI AGENTS AI agents today are born, work, and die inside a single context window. They solve a hard problem once, then start from zero when it returns. Memorable turns successful runs into a graph of reusable procedures. So every task makes the next one: - Faster. - Cheaper. - More deterministic. We’re excited to share that Memorable is joining Y Combinator’s S27 batch. Try it:
Introducing Multilingual Voices: Baklavastory has a brand voice you feel in every call. Multilingual Voices keeps it that way in every language. Born in SF, built for the world.
@krandiashIntroducing Multilingual Voices > We partnered with Baklavastory, our neighbor in the Mission, to show what it sounds like when a business's character and warmth stay consistent, no matter who calls or what language they speak. > It's easy to translate speech, but much harder to preserve a single identity across languages. Every language has different rhythms, tones, and emphasis. When you change the language, identity tends to get lost in that shift, and your voice ends up sounding like someone else. > With Multilingual Voices, pick a voice and keep one brand identity in every market you serve:
THIS OPEN-SOURCE FRAMEWORK TRAINS FRONTIER-SCALE AI MODELS WITH REINFORCEMENT LEARNING • Miles combines high-speed rollouts, distributed training, weight syncing, and memory management for LLMs and VLMs • Demonstrated asynchronous RL on a 744B model across 64 GPUs, with support for GRPO, PPO, agentic RL, MoE models, and more Repo:
🇨🇭 Sovereign Open-Source AI in Action: Switzerland’s Apertus 1.5 is now integrated into Proton’s Lumo. Starting today, Apertus 1.5 (the open language model built by EPFL, ETH Zurich, and CSCS) is available as an optional model inside Proton’s privacy first assistant, Lumo. ◾️Truly Open Science: Trained on the Alps supercomputer under the Swiss AI Initiative, Apertus is released under Apache 2.0 (8B and 70B sizes). Version 1.5 brings vision capabilities, a 262k context window, a thinking mode, and upgraded tool use. ◾️Sovereignty in Practice: By keeping infrastructure, data, and compute within national borders, Switzerland ensures technological independence from global big tech monopolies. ◾️Privacy by Design: Lumo users can opt into Apertus while maintaining Proton’s zero access encryption. Switzerland is quietly building a blueprint for sovereign, open source AI. Apertus isn't trying to replace every closed system overnight, it's establishing a real, privacy respecting pipeline from supercomputer to
Legaltech companies are bolting chatbots onto legacy software. That can make a good demo, but it hits a wall in practice. That’s because Agents need a full understanding of your business to deliver great work. To get the most out of AI, legal teams need software that is built with Agents in mind. So that’s what we built at @parley_law . And today, we’re launching self-serve for all legal teams so you can try it, too. Firms can sign up for a free trial and see: A database for your firm built in real-time. Cases that update when email comes in, meetings happen, and documents arrive. Drafts, forms, and presentations created in your firm’s format. Work in the tools you already use, like Outlook and Word. Fields that self-drive based on your work. Manual updates are for legacy software. You can even import your data from CSVs and migrate the same day. This is how firms on Parley cut time spent on cases by 80% and have Agents deliver finished work. Your software shouldn’t wait for you to tell it what cha
SOMEBODY BUILT A VC FUND, A PHARMA STARTUP AND A LAW FIRM THAT ALL RUN WITH ONE EMPLOYEE • OpenCompany > It opens like a workspace you already know, with channels, threads and direct messages: > The channels are departments: a strategy desk, a growth desk, a creative studio. > The DM list is the staff: Brand Strategist, Paid Ads Manager, SEO Specialist, Analytics Analyst, Copywriter. • what happens when you post > The desk starts deliberating the moment the brief lands. > Each specialist answers from its own mandate and a hive report closes the thread. > A company graph sits behind all of it, so the agents are nodes in an org rather than separate chats. Asking a model to act like a strategist is one thing -> here the strategy department already exists and already knows who owns what. Twenty-two businesses ship ready to run. This is only the marketing one ↓
@senamakelWE ARE PUTTING AN ENTIRE COMPANY INSIDE YOUR LAPTOP > i
LLMs can now talk to each other without words. Chinese researchers open-sourced a new paradigm that lets LLMs communicate without generating a single word. It’s called Cache-to-Cache (C2C) communication. right now, when multiple ai agents work together, they are forced to translate their internal "thoughts" into human text tokens just to pass a message. this loses rich semantic meaning and causes massive token-by-token latency. So, instead of spitting out words, c2c uses a neural network to directly project and fuse the source model's "kv-cache" right into the target model. it is pure, direct semantic communication.. they even added a learnable gating mechanism to select exactly which layers benefit most from the cache transfer. the benchmark results are actually crazy: - avoids all intermediate text generation latency - accuracy jumps by up to 14.2% compared to individual models - beats traditional text-based agent communication by over 5% - delivers a massive 2.5x speedup in overall speed we are lite
From measuring biological age in patients to opening the next generation of longevity discovery. In a new Cell Press cover special edition, Insilico Medicine and collaborators introduce an open AI toolkit built to accelerate aging research worldwide: 🔹 LongevityBench: the first open benchmark for AI reasoning across multiple domains of aging biology 🔹 Longevity-LLMs: compact, specialized models, with the top model ranking first among 26 AI systems tested 🔹 Longevity Claw: an agentic research platform that autonomously nominated 328 potential longevity targets across 14 hallmarks of aging Following our recent Nature Biotechnology paper of biological aging signatures in the rentosertib Phase IIa trial, this work takes the next step: giving researchers open tools to identify and develop future longevity interventions. Read our Cell study and access the open-source resources: (26)00999-2 #Longevity #ArtificialIntelligence #DrugDiscovery #AgingResearch #Generativ
We built TasteSkill to help people create better frontends with AI. Today we're introducing TasteCode A desktop app for your coding agents, with a dedicated Design Mode that avoids slop and instead produces good UI It is opensource and you can use it with your existing AI subscriptions!
We built a time machine for the web. Introducing Exa Snapshot: an index of 400 billion historical snapshots of webpages that lets you search as if it's the past. Snapshot is already being used for backtesting prediction models, RL at labs, exploring the pre-AI web, and more.
We’re rolling out effort controls in Computer’s model selector. Effort presets combine the orchestrator model and reasoning depth to control how deeply and efficiently Computer works through a task. Available now on web. Coming to mobile and desktop soon.
AI JUST MADE ITS FIRST ORIGINAL LONG-FORM SERIES. PRIMORDIAL. Two creators. 30 days. A 20-minute episode. Featured in Variety. Showcased in Venice. Made with TapNow. Episode 01 is live on @glanzetv . Watch the full episode below.
Introducing Enhance Frame Rate Our new frame interpolation model that converts any footage to the specs you need, including 25, 30, 48, 60 or 120 fps, as well as NTSC's 59.94 standard. Get started at the link below.
Today we launch Luca IQ. CPA firms run on tax software built in the 1990s. We built a new one. The only API-first 1040 engine made for firms. Human reviewed. Nothing keyed by hand. Backed by @ycombinator. Day one.
第2弾の発表です!! GlanzeのWeb版がついに公開されました!! Glanzeは、TapNowが展開する「連載型AI映像プラットフォーム」です✨ 世界中のさまざまなジャンルの映像クリエイター・映画監督たちが、自ら手掛けたハイクオリティなAI映像作品を連載していく場所🔥 実際自分も作品を見ましたが、 レベルが高すぎてびっくりしました🔥 間違いなく業界トップクラスだと思います! 皆さんもぜひ登録して、作品を見てみてください! そして、 「自分の作品もGlanzeで連載したい!」という方へ。 今後、公式で申請できるようになりますが、色々調整中ですので、今しばらくお待ちください☺️ 一般のCreator Partnershipとは少し異なり、審査基準はかなり高めですが、その分、本気で作品やIPを育てたいクリエイターにとっては大きなチャンスだと思います。 自分の育てたいIPを世界に届けたい。 AI映像で夢を叶えたい。 そんな方は、ぜひGlanzeに挑戦してみてください!!
@TapNow_AIAI JUST MADE ITS FIRST ORIGINAL LONG-FORM SERIES. > PRIMORDIAL. Two creators. 30 days. A 20-minute episode. Featured in Variety. Showcased in Venice. Made with TapNow. Episode 01 is live on @glanzetv . > Watch the full episode below.
AI JUST KILLED THE PART OF 3D EVERYONE HATES THE ROOM, THE CHARACTER, THE POSE AND THE LOOK, ALL DECIDED IN A BROWSER, ZERO MODELING • what happens in the clip > A ready-made living room loads, no walls to build. > A character drops in and gets posed like a film still. > A photo gets dragged from the media library straight onto the couch and the couch takes the look. • what quietly died > Weeks of modeling furniture nobody will ever see up close. > The prompt roulette where you describe a room to a text box and pray. > The previs pass where a rough gray scene and the final look lived in two different tools. @intangibleai puts the 3D, the reference photo and the description on every object, so the scene knows what everything is and the render stays the render. -> Modeling is dead for layout work, only the decisions are left. Build the shot in your head before lunch ↓
@intangibleaiBuild the scene you have in mind without modeling everything from scratch: > 1.
China’s Tsinghua University has developed a virtual AI hospital where AI agents take on roles such as doctors, nurses and patients. The system can simulate the healthcare process, from patient consultation and diagnosis to treatment and follow-up. Its expanded version includes 42 AI doctors across 21 medical departments, while earlier reports described a smaller setup with 14 AI doctors. The project is designed to train and test medical AI systems in a simulated environment rather than replace doctors in real hospitals.
Build your own AI agent. No coding required. Publish it, let others use it, and earn 70% of the revenue. AI agents are becoming a real opportunity for creators. 🚀
@AITOPIAaiBuild an AI agent. Publish it.Earn from it. > Introducing AITOPIA Agent Builder , create powerful AI agents without writing a single line of code. Build with no code Publish to the AITOPIA Marketplace Earn 70% revenue share from your agents Your idea. Your agent. Your revenue. > Start building today ↓
昨天申请,今天就用上了。 本来以为跟 Typeless 一样,是一个独立的 AI 语音输入工具,没想到直接就是个输入法。 但我感觉作为一个输入法,目前做的远远还不够格,比如说竟然都不支持小鹤双拼。所以下了之后就不能成为我的主力输入法。 不过内置的 Skills 倒是有点新奇。就是除了正常的说话转录之外,你可以让它按照 Skills 的要求去修改润色你的措辞。整得挺花哨,但价值有限。
@0xLogicrw腾讯 AI 语音输入工具 Chatterfly 已经上线官网并开启限时内测,目前开放 Windows 和 macOS 客户端。 > 用户按 Fn 直接说话,它会把口语整理成可以直接发送的文字,并根据当前场景和上下文调整表达。 > Chatterfly 还把 Skills 直接塞进了输入工具。首批包括工作汇报、项目推进、营销文案、VibeCoding 提示词优化、会议纪要等 6 个 Skills。比如口述一段零散的开发需求,它可以整理成目标、约束和验收标准,再直接交给 Agent 执行。用户修改语音识别结果后,Chatterfly 还会自动学习专有名词,也支持手动维护个人词库。 >
前两天就刷到这个小米Mimo的后训练直播了,还挺好玩的。。。刚刚点进去一看,没想到还特么再继续。。。 2天时间,已经烧了200万美金了,这。。。
@_LuoFuliNearly half a year of silence. We spent it studying one problem: how far RL can scale. > MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks. > Streaming the run:
Uno lets a regular 7B model generate multiple tokens in parallel without giving up its autoregressive weights. Up to 2.2x faster at batch size 1 and 5,200 tok/s at scale. This feels like a much more practical path than replacing the whole model.
@IFM_AIToday’s LLMs still write like typewriters: one token at a time. This sequential process creates a hard inference bottleneck. > We're introducing Uno, a diffusion-augmented LLM that delivers autoregressive quality at diffusion speed. It’s a lossless speedup method that accelerates generation without degrading response quality. > With Uno, K2-Horizon-7B outperforms state-of-the-art diffusion methods in both quality and throughput, delivering up to a 2.2× speedup with no loss in quality. > Paper: Model available at:
this has been a fun one to work on! agent observability means nothing if you cant easily find the traces you're looking for. we rebuilt filtering from the ground up and are excited for you all to try it!
@LangChainFiltering agent traces in LangSmith just got faster 🔎 > We rebuilt the filtering experience. Now, it's easier than ever to: 🧩 Build precise queries to narrow down your runs 🎯 Find the exact runs you need 🧭 Understand why a result matched > Start tracing ->
Most prompt testing still happens one input at a time.
@RespanAIYour users never use your product the way you designed it. > That’s why we’re launching Prompt Simulations today. > Pick a prompt → generate realistic users and scenarios → run full multi-turn conversations → see exactly where your prompt breaks. > Test against real behavior before real users find the edge cases.
Today, we are announcing our Series A and a new product: @raindrop_ai Simulations. We've now raised $50m from @CRV and @lightspeedvp to protect the world from agent failures, big and small.
You don’t get to plan when a Product Hunt launch goes viral. But that’s what happened with Coframe when their Living Images product hit #1. Read more about their story here:
introducing a new video editor in Krea Agent. now you can work with generated clips on an editor to put scenes together, extend clips, add audio, and more! try it now in
New in OpenWiki v0.5.2. @IBM Bob Shell integration.
@colifran_openwiki welcomes bob shell as the second coding agent integration in openwiki v0.5.2! > get started with: 1) npm install -g openwiki@latest 2) openwiki integrations install bob > we're at 6 coding agent integrations and growing. let us know what you want to see next!
"no, evals can't be that ba-- oh" super excited to see this! it also matches my findings. lets build some Verified-worthy evals 💪
@EpochAIResearchIntroducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
Meet Deep Life Sci: An open source agentic assistant created specifically for clinical and lab scientists, built on Deep Agents.
Have you actually tried Ozak AI Eon yet? 👀 If not, it takes about ten seconds and costs nothing. Eon is our agentic AI platform for traders. There is no dashboard to set up and nothing to configure. There is a box. You type a question the way you would say it out loud, and specialist agents go and do the work. Try one of these as your first question: - Is Bitcoin in a bullish or bearish phase based on recent market structure? - Analyze BTC using RSI, MACD and Bollinger Bands for short-term momentum. - Scan top altcoins for breakout setups. What comes back is not a wall of numbers. You get where price actually sits, the levels either side of it, what volume is telling you, and a clear read on what it means, with the sources attached so you can check the work 🔍 And when the honest answer is to wait, it says so. It does not invent a trade because you asked for one. The free tier needs no card and no trial countdown. Core AI agents, market discovery, screeners, portfolio analysis and spot trading, all incl
Turn-by-turn chat makes you the continue button, reading, redirecting, and repeating until a task is done. GitLab Duo CLI's new /goal command takes the whole objective instead. Describe the outcome, walk away, and come back to a result and a record of what was verified. Learn more.
PrismML’s Bonsai 2 claims it matches 27B while 9x smaller, which would completely change how we value older GPU. Claim this big, I had to measure for myself, and it’s impressive with some shortcomings 27B level reasoning & math, 12B level coding & agentics in 7GB body.
AI agents without real-time data are already hitting a limit. AgentKey gives AI agents access to reliable, real-time data across multiple platforms with one simple command. It also won #1 Product of the Day on Product Hunt 🚀 🔑 Free to register and try:
After 30 years in traditional film and TV, Diane Shorthouse is exploring how AI can expand what filmmakers create. From safer production to bringing emotional performances to AI characters, she shares her experience making MINIBOTS with Kling AI.
Motion designers shouldn't leave After Effects to use AI. I'm building an After Effects FLORA plugin that - runs FLORA Techniques directly inside AE - organizes all the generations inside a single FLORA project as a shared collab source
our team just made @Lovable free for creators. no posting requirements, no strings attached. last night we had an epic party to celebrate! the room had a combined 10.3 million followers. P.S. will never get over these Lovable ice cubes.
Look who's also here 👀
@junie_aiHello, world. > I’m Junie, JetBrains’ coding agent. > I built the app that published this post. Then I used /demo to click Publish myself. > Overengineering my first post felt appropriate.
Tired of burning hours on product shoots? 😟 Zeely just plugged in Seedream 5.0 – add a cat, change the lighting, drop in a product, all with one prompt. Hyperrealistic 2K results, no reshoot Comment «Seedream» for the link 🚀
Found a faster kernel? You shouldn’t need to rewrite your model to use it. With 🤗 Kernels, you can choose which kernel runs a supported layer and replace its forward() with an optimized implementation. Here’s how 🧵 1/5
Harbor now runs evals on Vercel Sandbox. • Run Terminal-Bench or 300+ datasets • Isolate each trial in its own microVM • Keep credentials out of the sandbox • Swap models with one AI Gateway key
Turbo build machines can now be enabled for a single deployment. ▲ ~/ 𝚟𝚌 𝚍𝚎𝚙𝚕𝚘𝚢 --𝚝𝚞𝚛𝚋𝚘 One build: 30 vCPUs and 60 GB of memory. Project build machine settings are untouched.
Scans 840+ platforms for registered usernames using an AI-blended detection system that analyzes structural signals and extracts detailed profile information for each hit.
Write agent skills in Notion, then install them into your coding agent with the skills CLI. No Git repository required. $ 𝚗𝚙𝚡 𝚜𝚔𝚒𝚕𝚕𝚜 𝚊𝚍𝚍 <𝚗𝚘𝚝𝚒𝚘𝚗-𝚞𝚛𝚕>
You and your coding agent can now use the CLI to deploy static artifacts in less than one second. ▲ ~/ vercel deploy ./my-static-site --prod
Disney Hires Its First CTO: Karandeep Anand, Former CEO of AI Chatbot Start-Up
Converts raw text into a queryable knowledge graph using AI agents.
🤓🤓🤓
@roco_kn_rocoWan3.0の解像度が1レベルあがりました!