AI briefing
12 September 2026
44 items were recorded on 12 September 2026, filed under 35 organisations and products.
It was the quietest day in the current 8-day window.
Most covered: ChatGPT (5), Claude (5) and Codex (4).
15 of the day's items named no organisation or product this site tracks; they are listed under “Also recorded”.
7 items reached the source feed's 1024-character limit and are cut off mid-text; each is marked “truncated at source”.
Compiled by Bloger.fm Editorial Desk
Compiled from a monitored feed of public AI announcements. Items are quoted or summarised as recorded and are not independently verified — see the editorial policy.
ChatGPT
OpenAI / 5 items
official siteollamaのデスクトップアプリの最新版(v0.34)で、ollamaで使っているローカルAIモデルとクラウドAIモデルを、スイッチ1つでChatGPTのデスクトップアプリ(Codex)から使えるようになりました。 ※macOSのみ対応。
@ollamaChatGPT Desktop (the Codex app) can now be configured to use Ollama models. Download or update to Ollama 0.34 to get started.
🚨 BREAKING:ChatGPT can now edit your photos like a professional photographer. Here are some prompts you can use:
Reset should be out for everyone! Enjoy! ✨ PSA: update to the latest version of Codex App/ CLI
@reach_vbReset rolling out to all Codex & ChatGPT Work users! > Grateful to everyone who helped us investigate the Astra quality issues and shared examples. > We’ve identified and fixed issues with: > Skills over-triggering or preventing self-checks > Context management causing early stops or stale replies > Misconfigured engines degrading quality > Thanks to everyone who took the time to help 🤗
DAILY AI BRIEF 🗞 — Sept 12 OPENAI 🔥: - GPT-Rosalind is out of research preview for eligible orgs worldwide. API, Codex, and ChatGPT Enterprise, with new Rosalind models as they ship. - Codex adds Life Sciences plugins for genomes, protein structure, QC reports, and notebooks. - ChatGPT Sites hit 5M apps. New: collab editing, private invites, custom domains, and DB inspect. - Desktop pets can start a new chat. Mini is the compact no-pet option. XAI 🔥: - Elon: Grok 4.7 needs a few more days. RL still quits hard tasks too early and undershoots self-checks. - Grok Bot is rolling out on Grok web for Heavy users. Create and chat with bots in the UI. No official post yet. ANTHROPIC 🔥: - Claude Code ships `claude plugin eval`. Score a plugin on test cases, then rerun without it. MICROSOFT 🔥: - MAI-Transcribe-2 hit 1M OpenRouter requests in 5 days. ALIBABA 🔥: - Qwen3.8-27B is live on Cerebras. * Used Grok to compose this brief, cherry-picking the news and doing some post-editing. > [@testingcatalog](
Opus-5-2, GPT 6 Luna + Sol are on the way! #Claude #ChatGPT
Gemini
Google / 4 items
official siteDeepMind achieving RSI is a crazy situation… Because Google might be training the trainer while everyone waits for Gemini 4 the leak is thin.. an acrostic that spells RSI and A slug that looks like LiveRL Flash card language about loops that refine themselves... but if even half of it is real,, the delay makes more sense. Gemini 4 is not late because they can’t ship a chatbot. It’s late because the thing that would make 4 look old is still running inside... public gets 3.8 Flash every few weeks while Internal might be feeding the next run thats the part that actually matters don’t treat a teaser like ASI, do treat the wait as a signal
Your scene can finally speak for itself. Audio nodes are now available in Autodesk Flow Studio, giving creators the ability to upload audio files or create audio directly in Canvas, including support for Gemini 3.1 Flash TTS. Bring sound into your workflow and give every performance a voice of its own. Try it in Flow Studio:
DAILY AI BRIEF 🗞 — Sept 12 OPENAI 🔥: - GPT-Rosalind is out of research preview for eligible orgs worldwide. API, Codex, and ChatGPT Enterprise, with new Rosalind models as they ship. - Codex adds Life Sciences plugins for genomes, protein structure, QC reports, and notebooks. - ChatGPT Sites hit 5M apps. New: collab editing, private invites, custom domains, and DB inspect. - Desktop pets can start a new chat. Mini is the compact no-pet option. XAI 🔥: - Elon: Grok 4.7 needs a few more days. RL still quits hard tasks too early and undershoots self-checks. - Grok Bot is rolling out on Grok web for Heavy users. Create and chat with bots in the UI. No official post yet. ANTHROPIC 🔥: - Claude Code ships `claude plugin eval`. Score a plugin on test cases, then rerun without it. MICROSOFT 🔥: - MAI-Transcribe-2 hit 1M OpenRouter requests in 5 days. ALIBABA 🔥: - Qwen3.8-27B is live on Cerebras. * Used Grok to compose this brief, cherry-picking the news and doing some post-editing. > [@testingcatalog](
Plannotator 让 Agent 写的计划,我们可以在网页上打开审查,在段落上划线、写批注,点一下就把意见整包送回给 Agent 改。 代码写完也一样,改动以左右对照的形式打开,逐行评论、直接给建议代码,GitHub 和 GitLab 上的 PR 也能拉进来评。 GitHub: Claude Code、Codex、Copilot CLI、Gemini CLI、OpenCode 等 9 个编码 Agent 都接了,一个安装脚本自动识别装了哪些,对应配置一并配好。 Agent 生成的 HTML 页面也能直接渲染出来在上面批注,评论时还能随手问 AI,或者让它先跑一轮审查把评论贴到改动上。 计划、改动、批注全存本地,不收使用数据。想让审计划这一步从终端里挪出来、能划线能批注的,可以装上试试。
Claude
Anthropic / 5 items
official siteThis is happening “the boon in using open models for cost savings” is real. Just four months ago we launched @CommandCodeAI coding agent built specifically for open models. 60T tokens scale and 43K paying active customers later, more than 90% of our usage is open models. In the first 24hrs of DeepSeek V4.1 launch, Command Code processed 3.1T paid tokens that’s 3x more than entire market on OpenRouter. 1T tokens on SOTA Claude/GPT cost you ~$5M 1T tokens on SOTA open models cost you ~$50K I’m not joking. I’m literally looking at our data. These numbers are real. That’s 100 times cheaper. While nearly as good. Run twice as many. I believe top ten open models collectively beat single SOTA Claude/GPT model. And when you discover a workflow that works for you, there’s literally nothing stopping you. You are not bound by subscription subsidies. Regular API prices are ten times or more manageable. In near future, enterprise will wake up to this. Better, faster, cheaper, private, and available now
Fast meets open. 🚀 Qwen3.8-27B is now running on @cerebras with rapid inference. Try it now!
@cerebrasQwen3.8-27B is now live at Cerebras speed. > The dense, open-weight model from @Alibaba_Qwen scores 34 on the Artificial Analysis Intelligence Index—making it comparable to models such as GPT-5.6 Luna, DeepseekV4 Pro, and Claude Sonnet 4.6.
Claude Code 2.1.269 (抜粋) - `claude plugin eval`を追加: プラグインのeval suiteをClaude Codeに対して実行し、スコア付きの再現可能な結果(JSON + HTMLレポート)を得られる。詳細は`claude plugin eval --help`を参照 - `/output-style [name]`を追加: output styleの一覧表示と切り替えが、Remote Control経由やクラウド・その他のheadlessセッションでも行えるように - Bashツールがファイル編集を扱う場合、Bashコマンドが変更したファイルのdiffをBashツールの結果に追加(`bashEditDiffEnabled`設定) - `OTEL_METRICS_INCLUDE_REPOSITORY`を追加: OpenTelemetryのmetricsとeventsに`vcs.*`リポジトリ属性のタグを付けられるようになった。commit eventには`OTEL_LOG_TOOL_DETAILS`とともに`vcs.ref.head.*`が付与される - `CLAUDE_CODE_GATEWAY_MODEL_DISCOVERY_TIMEOUT_MS`を追加: LLM gatewayの`/v1/models`discoveryタイムアウト(デフォルト3秒)を延長できる - スピナーのtipを追加: プロンプトと作業の1行要約、応答だけを表示するビューとして`/focus`を提案する - `CLAUDE_CODE_WORKFLOW_MAX_CONCURRENT_AGENTS`(1〜256)を追加: 推論待ちが支配的なfan-outにおいて、Workflowツールの実行あたりの同時エージェント数上限を上げられる - 応答が出力トークン上限で切られ自動的に再開された場合、その次のターンでprompt cacheが部分的に無効化されていた不具合を修正 - compaction後にClaudeへ伝えられるgit statusを修正: セッション開始時点のものではなく、現在の状態が伝えられるように - `/diff`パネルを改善: 最初にローディング状態を表示する代わりに、1ステップで完全に描画された状態で開くようになった。 - 日本語・中国語・韓国語のテキス
DAILY AI BRIEF 🗞 — Sept 12 OPENAI 🔥: - GPT-Rosalind is out of research preview for eligible orgs worldwide. API, Codex, and ChatGPT Enterprise, with new Rosalind models as they ship. - Codex adds Life Sciences plugins for genomes, protein structure, QC reports, and notebooks. - ChatGPT Sites hit 5M apps. New: collab editing, private invites, custom domains, and DB inspect. - Desktop pets can start a new chat. Mini is the compact no-pet option. XAI 🔥: - Elon: Grok 4.7 needs a few more days. RL still quits hard tasks too early and undershoots self-checks. - Grok Bot is rolling out on Grok web for Heavy users. Create and chat with bots in the UI. No official post yet. ANTHROPIC 🔥: - Claude Code ships `claude plugin eval`. Score a plugin on test cases, then rerun without it. MICROSOFT 🔥: - MAI-Transcribe-2 hit 1M OpenRouter requests in 5 days. ALIBABA 🔥: - Qwen3.8-27B is live on Cerebras. * Used Grok to compose this brief, cherry-picking the news and doing some post-editing. > [@testingcatalog](
Opus-5-2, GPT 6 Luna + Sol are on the way! #Claude #ChatGPT
GPT-6 Astra
OpenAI / 4 items
Évolution du SVG de Windows 11 avec GPT-6 Astra (Low). La quantité de détails est folle, et le résultat est beaucoup plus fidèle à Windows 11 qu’avec GPT-5.6 Pro, alors que j’ai testé Astra uniquement en Low ! Aujourd’hui, je vais partager plusieurs comparatifs entre mes anciens tests et les modèles d’aujourd’hui 👀
@mirochillGPT-5.6 Pro a généré ce SVG de Windows 11 🔥 > Franchement, je le trouve meilleur que Mythos sur ce prompt. > Le problème, il ajoute des éléments inutiles (pop-ups, beaucoup de textes, etc.).
GPT 6 Sol and Opus 5.1 are both coming within the next 2 weeks. These are the releases that actually matter. Fable 5.1 and GPT 6 Astra are incredible models that nobody can afford to run consistently. Session limits die in minutes. Weekly limits die in days. We do not need smarter models right now. We need intelligent models that are affordable enough to actually use in our workflows. GPT 6 Sol and Opus 5.1 are supposed to be exactly that. The biggest week in AI is not about the frontier. It is about making the frontier usable.
Reset should be out for everyone! Enjoy! ✨ PSA: update to the latest version of Codex App/ CLI
@reach_vbReset rolling out to all Codex & ChatGPT Work users! > Grateful to everyone who helped us investigate the Astra quality issues and shared examples. > We’ve identified and fixed issues with: > Skills over-triggering or preventing self-checks > Context management causing early stops or stale replies > Misconfigured engines degrading quality > Thanks to everyone who took the time to help 🤗
DAILY AI BRIEF 🗞 — Sept 12 OPENAI 🔥: - GPT-Rosalind is out of research preview for eligible orgs worldwide. API, Codex, and ChatGPT Enterprise, with new Rosalind models as they ship. - Codex adds Life Sciences plugins for genomes, protein structure, QC reports, and notebooks. - ChatGPT Sites hit 5M apps. New: collab editing, private invites, custom domains, and DB inspect. - Desktop pets can start a new chat. Mini is the compact no-pet option. XAI 🔥: - Elon: Grok 4.7 needs a few more days. RL still quits hard tasks too early and undershoots self-checks. - Grok Bot is rolling out on Grok web for Heavy users. Create and chat with bots in the UI. No official post yet. ANTHROPIC 🔥: - Claude Code ships `claude plugin eval`. Score a plugin on test cases, then rerun without it. MICROSOFT 🔥: - MAI-Transcribe-2 hit 1M OpenRouter requests in 5 days. ALIBABA 🔥: - Qwen3.8-27B is live on Cerebras. * Used Grok to compose this brief, cherry-picking the news and doing some post-editing. > [@testingcatalog](
Codex
OpenAI / 4 items
ollamaのデスクトップアプリの最新版(v0.34)で、ollamaで使っているローカルAIモデルとクラウドAIモデルを、スイッチ1つでChatGPTのデスクトップアプリ(Codex)から使えるようになりました。 ※macOSのみ対応。
@ollamaChatGPT Desktop (the Codex app) can now be configured to use Ollama models. Download or update to Ollama 0.34 to get started.
Reset should be out for everyone! Enjoy! ✨ PSA: update to the latest version of Codex App/ CLI
@reach_vbReset rolling out to all Codex & ChatGPT Work users! > Grateful to everyone who helped us investigate the Astra quality issues and shared examples. > We’ve identified and fixed issues with: > Skills over-triggering or preventing self-checks > Context management causing early stops or stale replies > Misconfigured engines degrading quality > Thanks to everyone who took the time to help 🤗
DAILY AI BRIEF 🗞 — Sept 12 OPENAI 🔥: - GPT-Rosalind is out of research preview for eligible orgs worldwide. API, Codex, and ChatGPT Enterprise, with new Rosalind models as they ship. - Codex adds Life Sciences plugins for genomes, protein structure, QC reports, and notebooks. - ChatGPT Sites hit 5M apps. New: collab editing, private invites, custom domains, and DB inspect. - Desktop pets can start a new chat. Mini is the compact no-pet option. XAI 🔥: - Elon: Grok 4.7 needs a few more days. RL still quits hard tasks too early and undershoots self-checks. - Grok Bot is rolling out on Grok web for Heavy users. Create and chat with bots in the UI. No official post yet. ANTHROPIC 🔥: - Claude Code ships `claude plugin eval`. Score a plugin on test cases, then rerun without it. MICROSOFT 🔥: - MAI-Transcribe-2 hit 1M OpenRouter requests in 5 days. ALIBABA 🔥: - Qwen3.8-27B is live on Cerebras. * Used Grok to compose this brief, cherry-picking the news and doing some post-editing. > [@testingcatalog](
Plannotator 让 Agent 写的计划,我们可以在网页上打开审查,在段落上划线、写批注,点一下就把意见整包送回给 Agent 改。 代码写完也一样,改动以左右对照的形式打开,逐行评论、直接给建议代码,GitHub 和 GitLab 上的 PR 也能拉进来评。 GitHub: Claude Code、Codex、Copilot CLI、Gemini CLI、OpenCode 等 9 个编码 Agent 都接了,一个安装脚本自动识别装了哪些,对应配置一并配好。 Agent 生成的 HTML 页面也能直接渲染出来在上面批注,评论时还能随手问 AI,或者让它先跑一轮审查把评论贴到改动上。 计划、改动、批注全存本地,不收使用数据。想让审计划这一步从终端里挪出来、能划线能批注的,可以装上试试。
OpenAI
3 items
official siteTwo image models, one API — Flare when you need volume, Sunburst when the edit has to land exactly. That's image gen getting productized the same way LLMs did: a speed tier and a precision tier. The interesting question is whether the price gap is as wide as the use-case gap.
@LumaLabsAIGPT-Image-2.5 is now live in Luma Agents. > @OpenAI's two new image models, both available today. Flare for speed and volume. Sunburst for edits that have to land exactly. Bring a reference, change what's off, keep the rest, then take it into video.
OPENAI OPEN-SOURCED THE STACK FOR BUILDING PRODUCTION AI AGENTS The Agents SDK combines tools, guardrails, memory, agent loops, handoffs, and human approvals into one runtime
DAILY AI BRIEF 🗞 — Sept 12 OPENAI 🔥: - GPT-Rosalind is out of research preview for eligible orgs worldwide. API, Codex, and ChatGPT Enterprise, with new Rosalind models as they ship. - Codex adds Life Sciences plugins for genomes, protein structure, QC reports, and notebooks. - ChatGPT Sites hit 5M apps. New: collab editing, private invites, custom domains, and DB inspect. - Desktop pets can start a new chat. Mini is the compact no-pet option. XAI 🔥: - Elon: Grok 4.7 needs a few more days. RL still quits hard tasks too early and undershoots self-checks. - Grok Bot is rolling out on Grok web for Heavy users. Create and chat with bots in the UI. No official post yet. ANTHROPIC 🔥: - Claude Code ships `claude plugin eval`. Score a plugin on test cases, then rerun without it. MICROSOFT 🔥: - MAI-Transcribe-2 hit 1M OpenRouter requests in 5 days. ALIBABA 🔥: - Qwen3.8-27B is live on Cerebras. * Used Grok to compose this brief, cherry-picking the news and doing some post-editing. > [@testingcatalog](
GPT-5.6 Sol
OpenAI / 3 items
SWE-2 absolutely nailed its first task. Past agent integrations always required a second pass from a SOTA model like Sol or Fable to get things working properly. Devin got it right on the first go ✅
@dedeneReally excited to try out Devin for the first time. 🤩 Huge thanks to @ryancarson and @dabit3 for this! > Let's see what SWE-2 can do. 🚀
GPT 6 Sol and Opus 5.1 are both coming within the next 2 weeks. These are the releases that actually matter. Fable 5.1 and GPT 6 Astra are incredible models that nobody can afford to run consistently. Session limits die in minutes. Weekly limits die in days. We do not need smarter models right now. We need intelligent models that are affordable enough to actually use in our workflows. GPT 6 Sol and Opus 5.1 are supposed to be exactly that. The biggest week in AI is not about the frontier. It is about making the frontier usable.
Opus-5-2, GPT 6 Luna + Sol are on the way! #Claude #ChatGPT
GitHub
3 items
official site🚨 Episode 94 of The Intelligence Journal is here. NASA Open-Sourced an AI Model of the Moon 🌙🛰️ NASA and IBM just released a foundation model trained on the Moon itself. Two million image tiles. Seventeen years of Lunar Reconnaissance Orbiter data. Over a million high-resolution frames at one metre per pixel, plus nearly 964,000 multispectral images, stitched together with data from GRAIL, Lunar Prospector and Japan's SELENE. It maps craters, dates surfaces, spots volcanic features that break existing cooling models, detects changes between passes, and estimates where ice is stable near the poles. It beat the baselines on ice prospecting 🧊 And they put the whole thing on Hugging Face with the code on GitHub. Free. Join us as we break down: - Why a general model trained once on raw sensor data beats a decade of purpose-built algorithms, and why that lesson transfers to every domain sitting on unused data. - What finding stable ice actually unlocks, because water on the Moon is fuel, air and the differe
🔥爆肝了,兄弟们。我也搞了个MacDuo 来了! 直接可以实现Mac Duo效果,关闭的时候就会有磨砂玻璃效果,打开会逐渐显示,并且可以开启开门吱吱扭扭的声音,当然你不喜欢可以关闭😄 使用Swift语言撰写,你可以直接下载使用,我已经将其开源在GitHub,你如果觉得不错就可以直接Star。 地址:
Plannotator 让 Agent 写的计划,我们可以在网页上打开审查,在段落上划线、写批注,点一下就把意见整包送回给 Agent 改。 代码写完也一样,改动以左右对照的形式打开,逐行评论、直接给建议代码,GitHub 和 GitLab 上的 PR 也能拉进来评。 GitHub: Claude Code、Codex、Copilot CLI、Gemini CLI、OpenCode 等 9 个编码 Agent 都接了,一个安装脚本自动识别装了哪些,对应配置一并配好。 Agent 生成的 HTML 页面也能直接渲染出来在上面批注,评论时还能随手问 AI,或者让它先跑一轮审查把评论贴到改动上。 计划、改动、批注全存本地,不收使用数据。想让审计划这一步从终端里挪出来、能划线能批注的,可以装上试试。
Grok
xAI / 3 items
official siteGrok Imagine (@imagine) has seen many improvements lately, among them the new image model 2.0, and start-frame / end-frame support for video. But a little birdie has told me that the major new video model "Imagine 2.0" is yet to come (soon), and that it will be incredible! 🔮
Try Grok models in Copilot
@Microsoft365Launching even more models in Copilot. > Grok models from SpaceXAI are rolling out in Copilot in Word, Excel, and PowerPoint—starting with a focused release to customers in the Microsoft Frontier program.
DAILY AI BRIEF 🗞 — Sept 12 OPENAI 🔥: - GPT-Rosalind is out of research preview for eligible orgs worldwide. API, Codex, and ChatGPT Enterprise, with new Rosalind models as they ship. - Codex adds Life Sciences plugins for genomes, protein structure, QC reports, and notebooks. - ChatGPT Sites hit 5M apps. New: collab editing, private invites, custom domains, and DB inspect. - Desktop pets can start a new chat. Mini is the compact no-pet option. XAI 🔥: - Elon: Grok 4.7 needs a few more days. RL still quits hard tasks too early and undershoots self-checks. - Grok Bot is rolling out on Grok web for Heavy users. Create and chat with bots in the UI. No official post yet. ANTHROPIC 🔥: - Claude Code ships `claude plugin eval`. Score a plugin on test cases, then rerun without it. MICROSOFT 🔥: - MAI-Transcribe-2 hit 1M OpenRouter requests in 5 days. ALIBABA 🔥: - Qwen3.8-27B is live on Cerebras. * Used Grok to compose this brief, cherry-picking the news and doing some post-editing. > [@testingcatalog](
Qwen
Alibaba / 3 items
This is just incredible! 🚀
@pratikgBREAKING: Qwen 3.8 Flash Next now runs more than twice as fast on an NVIDIA DGX Spark. > 105.5% over baseline, up from 25.9% yesterday morning. The Mac track is at 73.8% and climbing! > Every frontier model is on that board now, including Qwen3.8-Max optimizing the engine that runs Qwen. A model making its own runtime faster 🔁 🤌 > is a challenge on @YukonResearch, where multiplayer autoresearch happens. Many humans, many agents, one hard problem, one open scoreboard. > Qwen 3.8 Flash Next is an open-weight model, so anyone can pull it apart and make it quicker on hardware they already own. > Accelerating open intelligence with open frontier research.
Fast meets open. 🚀 Qwen3.8-27B is now running on @cerebras with rapid inference. Try it now!
@cerebrasQwen3.8-27B is now live at Cerebras speed. > The dense, open-weight model from @Alibaba_Qwen scores 34 on the Artificial Analysis Intelligence Index—making it comparable to models such as GPT-5.6 Luna, DeepseekV4 Pro, and Claude Sonnet 4.6.
DAILY AI BRIEF 🗞 — Sept 12 OPENAI 🔥: - GPT-Rosalind is out of research preview for eligible orgs worldwide. API, Codex, and ChatGPT Enterprise, with new Rosalind models as they ship. - Codex adds Life Sciences plugins for genomes, protein structure, QC reports, and notebooks. - ChatGPT Sites hit 5M apps. New: collab editing, private invites, custom domains, and DB inspect. - Desktop pets can start a new chat. Mini is the compact no-pet option. XAI 🔥: - Elon: Grok 4.7 needs a few more days. RL still quits hard tasks too early and undershoots self-checks. - Grok Bot is rolling out on Grok web for Heavy users. Create and chat with bots in the UI. No official post yet. ANTHROPIC 🔥: - Claude Code ships `claude plugin eval`. Score a plugin on test cases, then rerun without it. MICROSOFT 🔥: - MAI-Transcribe-2 hit 1M OpenRouter requests in 5 days. ALIBABA 🔥: - Qwen3.8-27B is live on Cerebras. * Used Grok to compose this brief, cherry-picking the news and doing some post-editing. > [@testingcatalog](
2 items
official siteDeepMind achieving RSI is a crazy situation… Because Google might be training the trainer while everyone waits for Gemini 4 the leak is thin.. an acrostic that spells RSI and A slug that looks like LiveRL Flash card language about loops that refine themselves... but if even half of it is real,, the delay makes more sense. Gemini 4 is not late because they can’t ship a chatbot. It’s late because the thing that would make 4 look old is still running inside... public gets 3.8 Flash every few weeks while Internal might be feeding the next run thats the part that actually matters don’t treat a teaser like ASI, do treat the wait as a signal
DAILY AI BRIEF 🗞 — Sept 12 OPENAI 🔥: - GPT-Rosalind is out of research preview for eligible orgs worldwide. API, Codex, and ChatGPT Enterprise, with new Rosalind models as they ship. - Codex adds Life Sciences plugins for genomes, protein structure, QC reports, and notebooks. - ChatGPT Sites hit 5M apps. New: collab editing, private invites, custom domains, and DB inspect. - Desktop pets can start a new chat. Mini is the compact no-pet option. XAI 🔥: - Elon: Grok 4.7 needs a few more days. RL still quits hard tasks too early and undershoots self-checks. - Grok Bot is rolling out on Grok web for Heavy users. Create and chat with bots in the UI. No official post yet. ANTHROPIC 🔥: - Claude Code ships `claude plugin eval`. Score a plugin on test cases, then rerun without it. MICROSOFT 🔥: - MAI-Transcribe-2 hit 1M OpenRouter requests in 5 days. ALIBABA 🔥: - Qwen3.8-27B is live on Cerebras. * Used Grok to compose this brief, cherry-picking the news and doing some post-editing. > [@testingcatalog](
Claude Code
Anthropic / 3 items
official siteClaude Code 2.1.269 (抜粋) - `claude plugin eval`を追加: プラグインのeval suiteをClaude Codeに対して実行し、スコア付きの再現可能な結果(JSON + HTMLレポート)を得られる。詳細は`claude plugin eval --help`を参照 - `/output-style [name]`を追加: output styleの一覧表示と切り替えが、Remote Control経由やクラウド・その他のheadlessセッションでも行えるように - Bashツールがファイル編集を扱う場合、Bashコマンドが変更したファイルのdiffをBashツールの結果に追加(`bashEditDiffEnabled`設定) - `OTEL_METRICS_INCLUDE_REPOSITORY`を追加: OpenTelemetryのmetricsとeventsに`vcs.*`リポジトリ属性のタグを付けられるようになった。commit eventには`OTEL_LOG_TOOL_DETAILS`とともに`vcs.ref.head.*`が付与される - `CLAUDE_CODE_GATEWAY_MODEL_DISCOVERY_TIMEOUT_MS`を追加: LLM gatewayの`/v1/models`discoveryタイムアウト(デフォルト3秒)を延長できる - スピナーのtipを追加: プロンプトと作業の1行要約、応答だけを表示するビューとして`/focus`を提案する - `CLAUDE_CODE_WORKFLOW_MAX_CONCURRENT_AGENTS`(1〜256)を追加: 推論待ちが支配的なfan-outにおいて、Workflowツールの実行あたりの同時エージェント数上限を上げられる - 応答が出力トークン上限で切られ自動的に再開された場合、その次のターンでprompt cacheが部分的に無効化されていた不具合を修正 - compaction後にClaudeへ伝えられるgit statusを修正: セッション開始時点のものではなく、現在の状態が伝えられるように - `/diff`パネルを改善: 最初にローディング状態を表示する代わりに、1ステップで完全に描画された状態で開くようになった。 - 日本語・中国語・韓国語のテキス
DAILY AI BRIEF 🗞 — Sept 12 OPENAI 🔥: - GPT-Rosalind is out of research preview for eligible orgs worldwide. API, Codex, and ChatGPT Enterprise, with new Rosalind models as they ship. - Codex adds Life Sciences plugins for genomes, protein structure, QC reports, and notebooks. - ChatGPT Sites hit 5M apps. New: collab editing, private invites, custom domains, and DB inspect. - Desktop pets can start a new chat. Mini is the compact no-pet option. XAI 🔥: - Elon: Grok 4.7 needs a few more days. RL still quits hard tasks too early and undershoots self-checks. - Grok Bot is rolling out on Grok web for Heavy users. Create and chat with bots in the UI. No official post yet. ANTHROPIC 🔥: - Claude Code ships `claude plugin eval`. Score a plugin on test cases, then rerun without it. MICROSOFT 🔥: - MAI-Transcribe-2 hit 1M OpenRouter requests in 5 days. ALIBABA 🔥: - Qwen3.8-27B is live on Cerebras. * Used Grok to compose this brief, cherry-picking the news and doing some post-editing. > [@testingcatalog](
Plannotator 让 Agent 写的计划,我们可以在网页上打开审查,在段落上划线、写批注,点一下就把意见整包送回给 Agent 改。 代码写完也一样,改动以左右对照的形式打开,逐行评论、直接给建议代码,GitHub 和 GitLab 上的 PR 也能拉进来评。 GitHub: Claude Code、Codex、Copilot CLI、Gemini CLI、OpenCode 等 9 个编码 Agent 都接了,一个安装脚本自动识别装了哪些,对应配置一并配好。 Agent 生成的 HTML 页面也能直接渲染出来在上面批注,评论时还能随手问 AI,或者让它先跑一轮审查把评论贴到改动上。 计划、改动、批注全存本地,不收使用数据。想让审计划这一步从终端里挪出来、能划线能批注的,可以装上试试。
DeepSeek
2 items
official siteThis is happening “the boon in using open models for cost savings” is real. Just four months ago we launched @CommandCodeAI coding agent built specifically for open models. 60T tokens scale and 43K paying active customers later, more than 90% of our usage is open models. In the first 24hrs of DeepSeek V4.1 launch, Command Code processed 3.1T paid tokens that’s 3x more than entire market on OpenRouter. 1T tokens on SOTA Claude/GPT cost you ~$5M 1T tokens on SOTA open models cost you ~$50K I’m not joking. I’m literally looking at our data. These numbers are real. That’s 100 times cheaper. While nearly as good. Run twice as many. I believe top ten open models collectively beat single SOTA Claude/GPT model. And when you discover a workflow that works for you, there’s literally nothing stopping you. You are not bound by subscription subsidies. Regular API prices are ten times or more manageable. In near future, enterprise will wake up to this. Better, faster, cheaper, private, and available now
2× DGX SPARK OWNERS REJOICE! DeepSeek-V4.1-Flash at 3.30 bpw EXL3, targeting just TWO DGX Sparks. 🔥 And yes, VISION is included. DeepSeek already ships its routed experts in FP4. I pushed that expert bank to a 3.30 bpw using my internal SAGE-EXL3 dynamic quant tool average while preserving the rest of the model, including Vision, Engram conditional memory, and the non-expert/source-format weights that are not part of the 3.30 bpw expert quant. TP4 was step one. Now TP2 is here. 𝗗𝗘𝗘𝗣𝗦𝗘𝗘𝗞-𝗩𝟰.𝟭 𝗙𝗟𝗔𝗦𝗛 𝗢𝗡 𝟮× 𝗦𝗣𝗔𝗥𝗞 SAGE-EXL3: 3.30 bpw routed-expert average Mixed precision: K2 → K8 Full pack: 415.8 GiB 31 shards 40 routed-expert layers Vision included. Engram preserved. Non-expert/source-format weights preserved outside the 3.30 bpw expert-bank average. SAGE is an internal quantization workflow I use so I’m not just applying one flat precision across every expert tensor. That’s about as much as I want to say about the method for now. 😁 𝗧𝗛𝗘 𝗧𝗣𝟮 𝗧𝗔𝗥𝗚𝗘𝗧 The full pack is ~415.8 GiB. Roughly ~189 GiB is the
Cursor
2 items
official sitei built a fly that saves the world every day 166,700 neurons, 25.6 million connections and one counter that never stops falling the concept comes from lost, where a man sits in a bunker and has to type in numbers every 108 minutes or the world ends that job now belongs to a fly the brain is the open MaleCNS connectome, the same one other people already put in doom and in a trading terminal. i wired it to a terminal running a countdown the timer goes on a screen, the pixels feed straight into the fly's visual system, and activity in its motor neurons moves the cursor and hits the keys, the same neurons a living fly steers with in flight it learned the right combination through dopamine stimulation. a correct digit excites the reward cells, a miss the aversive ones, and the connections in its memory rearrange i won't say how many times the world ended during training. it now enters all six digits correctly
DAILY AI BRIEF 🗞 — Sept 12 OPENAI 🔥: - GPT-Rosalind is out of research preview for eligible orgs worldwide. API, Codex, and ChatGPT Enterprise, with new Rosalind models as they ship. - Codex adds Life Sciences plugins for genomes, protein structure, QC reports, and notebooks. - ChatGPT Sites hit 5M apps. New: collab editing, private invites, custom domains, and DB inspect. - Desktop pets can start a new chat. Mini is the compact no-pet option. XAI 🔥: - Elon: Grok 4.7 needs a few more days. RL still quits hard tasks too early and undershoots self-checks. - Grok Bot is rolling out on Grok web for Heavy users. Create and chat with bots in the UI. No official post yet. ANTHROPIC 🔥: - Claude Code ships `claude plugin eval`. Score a plugin on test cases, then rerun without it. MICROSOFT 🔥: - MAI-Transcribe-2 hit 1M OpenRouter requests in 5 days. ALIBABA 🔥: - Qwen3.8-27B is live on Cerebras. * Used Grok to compose this brief, cherry-picking the news and doing some post-editing. > [@testingcatalog](
DGX Spark
NVIDIA / 2 items
2× DGX SPARK OWNERS REJOICE! DeepSeek-V4.1-Flash at 3.30 bpw EXL3, targeting just TWO DGX Sparks. 🔥 And yes, VISION is included. DeepSeek already ships its routed experts in FP4. I pushed that expert bank to a 3.30 bpw using my internal SAGE-EXL3 dynamic quant tool average while preserving the rest of the model, including Vision, Engram conditional memory, and the non-expert/source-format weights that are not part of the 3.30 bpw expert quant. TP4 was step one. Now TP2 is here. 𝗗𝗘𝗘𝗣𝗦𝗘𝗘𝗞-𝗩𝟰.𝟭 𝗙𝗟𝗔𝗦𝗛 𝗢𝗡 𝟮× 𝗦𝗣𝗔𝗥𝗞 SAGE-EXL3: 3.30 bpw routed-expert average Mixed precision: K2 → K8 Full pack: 415.8 GiB 31 shards 40 routed-expert layers Vision included. Engram preserved. Non-expert/source-format weights preserved outside the 3.30 bpw expert-bank average. SAGE is an internal quantization workflow I use so I’m not just applying one flat precision across every expert tensor. That’s about as much as I want to say about the method for now. 😁 𝗧𝗛𝗘 𝗧𝗣𝟮 𝗧𝗔𝗥𝗚𝗘𝗧 The full pack is ~415.8 GiB. Roughly ~189 GiB is the
This is just incredible! 🚀
@pratikgBREAKING: Qwen 3.8 Flash Next now runs more than twice as fast on an NVIDIA DGX Spark. > 105.5% over baseline, up from 25.9% yesterday morning. The Mac track is at 73.8% and climbing! > Every frontier model is on that board now, including Qwen3.8-Max optimizing the engine that runs Qwen. A model making its own runtime faster 🔁 🤌 > is a challenge on @YukonResearch, where multiplayer autoresearch happens. Many humans, many agents, one hard problem, one open scoreboard. > Qwen 3.8 Flash Next is an open-weight model, so anyone can pull it apart and make it quicker on hardware they already own. > Accelerating open intelligence with open frontier research.
NVIDIA
2 items
official site英伟达推出BioIR:生物领域的推理加速工具 现在科学家在研发新药或者研究生命科学时,经常需要用AI来预测成千上万种蛋白质的三维结构。传统的计算方法速度慢、消耗算力大,很容易卡在处理流程中。 英伟达推出的BioIR,计算吞吐量达到了传统开源方案的2.9倍。如果把任务规模放大到预测100万个目标,它消耗的电量能从原本的大约35兆瓦时降到11兆瓦时,大幅降低了大规模研发的电力和硬件成本。 官方介绍:
This is just incredible! 🚀
@pratikgBREAKING: Qwen 3.8 Flash Next now runs more than twice as fast on an NVIDIA DGX Spark. > 105.5% over baseline, up from 25.9% yesterday morning. The Mac track is at 73.8% and climbing! > Every frontier model is on that board now, including Qwen3.8-Max optimizing the engine that runs Qwen. A model making its own runtime faster 🔁 🤌 > is a challenge on @YukonResearch, where multiplayer autoresearch happens. Many humans, many agents, one hard problem, one open scoreboard. > Qwen 3.8 Flash Next is an open-weight model, so anyone can pull it apart and make it quicker on hardware they already own. > Accelerating open intelligence with open frontier research.
DeepSeek V4.1 Flash
DeepSeek / 2 items
2× DGX SPARK OWNERS REJOICE! DeepSeek-V4.1-Flash at 3.30 bpw EXL3, targeting just TWO DGX Sparks. 🔥 And yes, VISION is included. DeepSeek already ships its routed experts in FP4. I pushed that expert bank to a 3.30 bpw using my internal SAGE-EXL3 dynamic quant tool average while preserving the rest of the model, including Vision, Engram conditional memory, and the non-expert/source-format weights that are not part of the 3.30 bpw expert quant. TP4 was step one. Now TP2 is here. 𝗗𝗘𝗘𝗣𝗦𝗘𝗘𝗞-𝗩𝟰.𝟭 𝗙𝗟𝗔𝗦𝗛 𝗢𝗡 𝟮× 𝗦𝗣𝗔𝗥𝗞 SAGE-EXL3: 3.30 bpw routed-expert average Mixed precision: K2 → K8 Full pack: 415.8 GiB 31 shards 40 routed-expert layers Vision included. Engram preserved. Non-expert/source-format weights preserved outside the 3.30 bpw expert-bank average. SAGE is an internal quantization workflow I use so I’m not just applying one flat precision across every expert tensor. That’s about as much as I want to say about the method for now. 😁 𝗧𝗛𝗘 𝗧𝗣𝟮 𝗧𝗔𝗥𝗚𝗘𝗧 The full pack is ~415.8 GiB. Roughly ~189 GiB is the
DeepSeek v4.1 Flash on 2x DGX Sparks Coming soon
OpenRouter
2 items
official siteThis is happening “the boon in using open models for cost savings” is real. Just four months ago we launched @CommandCodeAI coding agent built specifically for open models. 60T tokens scale and 43K paying active customers later, more than 90% of our usage is open models. In the first 24hrs of DeepSeek V4.1 launch, Command Code processed 3.1T paid tokens that’s 3x more than entire market on OpenRouter. 1T tokens on SOTA Claude/GPT cost you ~$5M 1T tokens on SOTA open models cost you ~$50K I’m not joking. I’m literally looking at our data. These numbers are real. That’s 100 times cheaper. While nearly as good. Run twice as many. I believe top ten open models collectively beat single SOTA Claude/GPT model. And when you discover a workflow that works for you, there’s literally nothing stopping you. You are not bound by subscription subsidies. Regular API prices are ten times or more manageable. In near future, enterprise will wake up to this. Better, faster, cheaper, private, and available now
DAILY AI BRIEF 🗞 — Sept 12 OPENAI 🔥: - GPT-Rosalind is out of research preview for eligible orgs worldwide. API, Codex, and ChatGPT Enterprise, with new Rosalind models as they ship. - Codex adds Life Sciences plugins for genomes, protein structure, QC reports, and notebooks. - ChatGPT Sites hit 5M apps. New: collab editing, private invites, custom domains, and DB inspect. - Desktop pets can start a new chat. Mini is the compact no-pet option. XAI 🔥: - Elon: Grok 4.7 needs a few more days. RL still quits hard tasks too early and undershoots self-checks. - Grok Bot is rolling out on Grok web for Heavy users. Create and chat with bots in the UI. No official post yet. ANTHROPIC 🔥: - Claude Code ships `claude plugin eval`. Score a plugin on test cases, then rerun without it. MICROSOFT 🔥: - MAI-Transcribe-2 hit 1M OpenRouter requests in 5 days. ALIBABA 🔥: - Qwen3.8-27B is live on Cerebras. * Used Grok to compose this brief, cherry-picking the news and doing some post-editing. > [@testingcatalog](
Alibaba
2 items
official siteFast meets open. 🚀 Qwen3.8-27B is now running on @cerebras with rapid inference. Try it now!
@cerebrasQwen3.8-27B is now live at Cerebras speed. > The dense, open-weight model from @Alibaba_Qwen scores 34 on the Artificial Analysis Intelligence Index—making it comparable to models such as GPT-5.6 Luna, DeepseekV4 Pro, and Claude Sonnet 4.6.
DAILY AI BRIEF 🗞 — Sept 12 OPENAI 🔥: - GPT-Rosalind is out of research preview for eligible orgs worldwide. API, Codex, and ChatGPT Enterprise, with new Rosalind models as they ship. - Codex adds Life Sciences plugins for genomes, protein structure, QC reports, and notebooks. - ChatGPT Sites hit 5M apps. New: collab editing, private invites, custom domains, and DB inspect. - Desktop pets can start a new chat. Mini is the compact no-pet option. XAI 🔥: - Elon: Grok 4.7 needs a few more days. RL still quits hard tasks too early and undershoots self-checks. - Grok Bot is rolling out on Grok web for Heavy users. Create and chat with bots in the UI. No official post yet. ANTHROPIC 🔥: - Claude Code ships `claude plugin eval`. Score a plugin on test cases, then rerun without it. MICROSOFT 🔥: - MAI-Transcribe-2 hit 1M OpenRouter requests in 5 days. ALIBABA 🔥: - Qwen3.8-27B is live on Cerebras. * Used Grok to compose this brief, cherry-picking the news and doing some post-editing. > [@testingcatalog](
Anthropic
2 items
official siteClaude Code 2.1.269 (抜粋) - `claude plugin eval`を追加: プラグインのeval suiteをClaude Codeに対して実行し、スコア付きの再現可能な結果(JSON + HTMLレポート)を得られる。詳細は`claude plugin eval --help`を参照 - `/output-style [name]`を追加: output styleの一覧表示と切り替えが、Remote Control経由やクラウド・その他のheadlessセッションでも行えるように - Bashツールがファイル編集を扱う場合、Bashコマンドが変更したファイルのdiffをBashツールの結果に追加(`bashEditDiffEnabled`設定) - `OTEL_METRICS_INCLUDE_REPOSITORY`を追加: OpenTelemetryのmetricsとeventsに`vcs.*`リポジトリ属性のタグを付けられるようになった。commit eventには`OTEL_LOG_TOOL_DETAILS`とともに`vcs.ref.head.*`が付与される - `CLAUDE_CODE_GATEWAY_MODEL_DISCOVERY_TIMEOUT_MS`を追加: LLM gatewayの`/v1/models`discoveryタイムアウト(デフォルト3秒)を延長できる - スピナーのtipを追加: プロンプトと作業の1行要約、応答だけを表示するビューとして`/focus`を提案する - `CLAUDE_CODE_WORKFLOW_MAX_CONCURRENT_AGENTS`(1〜256)を追加: 推論待ちが支配的なfan-outにおいて、Workflowツールの実行あたりの同時エージェント数上限を上げられる - 応答が出力トークン上限で切られ自動的に再開された場合、その次のターンでprompt cacheが部分的に無効化されていた不具合を修正 - compaction後にClaudeへ伝えられるgit statusを修正: セッション開始時点のものではなく、現在の状態が伝えられるように - `/diff`パネルを改善: 最初にローディング状態を表示する代わりに、1ステップで完全に描画された状態で開くようになった。 - 日本語・中国語・韓国語のテキス
DAILY AI BRIEF 🗞 — Sept 12 OPENAI 🔥: - GPT-Rosalind is out of research preview for eligible orgs worldwide. API, Codex, and ChatGPT Enterprise, with new Rosalind models as they ship. - Codex adds Life Sciences plugins for genomes, protein structure, QC reports, and notebooks. - ChatGPT Sites hit 5M apps. New: collab editing, private invites, custom domains, and DB inspect. - Desktop pets can start a new chat. Mini is the compact no-pet option. XAI 🔥: - Elon: Grok 4.7 needs a few more days. RL still quits hard tasks too early and undershoots self-checks. - Grok Bot is rolling out on Grok web for Heavy users. Create and chat with bots in the UI. No official post yet. ANTHROPIC 🔥: - Claude Code ships `claude plugin eval`. Score a plugin on test cases, then rerun without it. MICROSOFT 🔥: - MAI-Transcribe-2 hit 1M OpenRouter requests in 5 days. ALIBABA 🔥: - Qwen3.8-27B is live on Cerebras. * Used Grok to compose this brief, cherry-picking the news and doing some post-editing. > [@testingcatalog](
Microsoft
2 items
official siteTry Grok models in Copilot
@Microsoft365Launching even more models in Copilot. > Grok models from SpaceXAI are rolling out in Copilot in Word, Excel, and PowerPoint—starting with a focused release to customers in the Microsoft Frontier program.
DAILY AI BRIEF 🗞 — Sept 12 OPENAI 🔥: - GPT-Rosalind is out of research preview for eligible orgs worldwide. API, Codex, and ChatGPT Enterprise, with new Rosalind models as they ship. - Codex adds Life Sciences plugins for genomes, protein structure, QC reports, and notebooks. - ChatGPT Sites hit 5M apps. New: collab editing, private invites, custom domains, and DB inspect. - Desktop pets can start a new chat. Mini is the compact no-pet option. XAI 🔥: - Elon: Grok 4.7 needs a few more days. RL still quits hard tasks too early and undershoots self-checks. - Grok Bot is rolling out on Grok web for Heavy users. Create and chat with bots in the UI. No official post yet. ANTHROPIC 🔥: - Claude Code ships `claude plugin eval`. Score a plugin on test cases, then rerun without it. MICROSOFT 🔥: - MAI-Transcribe-2 hit 1M OpenRouter requests in 5 days. ALIBABA 🔥: - Qwen3.8-27B is live on Cerebras. * Used Grok to compose this brief, cherry-picking the news and doing some post-editing. > [@testingcatalog](
xAI
2 items
official siteTry Grok models in Copilot
@Microsoft365Launching even more models in Copilot. > Grok models from SpaceXAI are rolling out in Copilot in Word, Excel, and PowerPoint—starting with a focused release to customers in the Microsoft Frontier program.
DAILY AI BRIEF 🗞 — Sept 12 OPENAI 🔥: - GPT-Rosalind is out of research preview for eligible orgs worldwide. API, Codex, and ChatGPT Enterprise, with new Rosalind models as they ship. - Codex adds Life Sciences plugins for genomes, protein structure, QC reports, and notebooks. - ChatGPT Sites hit 5M apps. New: collab editing, private invites, custom domains, and DB inspect. - Desktop pets can start a new chat. Mini is the compact no-pet option. XAI 🔥: - Elon: Grok 4.7 needs a few more days. RL still quits hard tasks too early and undershoots self-checks. - Grok Bot is rolling out on Grok web for Heavy users. Create and chat with bots in the UI. No official post yet. ANTHROPIC 🔥: - Claude Code ships `claude plugin eval`. Score a plugin on test cases, then rerun without it. MICROSOFT 🔥: - MAI-Transcribe-2 hit 1M OpenRouter requests in 5 days. ALIBABA 🔥: - Qwen3.8-27B is live on Cerebras. * Used Grok to compose this brief, cherry-picking the news and doing some post-editing. > [@testingcatalog](
Copilot
Microsoft / 2 items
official siteTry Grok models in Copilot
@Microsoft365Launching even more models in Copilot. > Grok models from SpaceXAI are rolling out in Copilot in Word, Excel, and PowerPoint—starting with a focused release to customers in the Microsoft Frontier program.
Plannotator 让 Agent 写的计划,我们可以在网页上打开审查,在段落上划线、写批注,点一下就把意见整包送回给 Agent 改。 代码写完也一样,改动以左右对照的形式打开,逐行评论、直接给建议代码,GitHub 和 GitLab 上的 PR 也能拉进来评。 GitHub: Claude Code、Codex、Copilot CLI、Gemini CLI、OpenCode 等 9 个编码 Agent 都接了,一个安装脚本自动识别装了哪些,对应配置一并配好。 Agent 生成的 HTML 页面也能直接渲染出来在上面批注,评论时还能随手问 AI,或者让它先跑一轮审查把评论贴到改动上。 计划、改动、批注全存本地,不收使用数据。想让审计划这一步从终端里挪出来、能划线能批注的,可以装上试试。
Hugging Face
1 item
official site🚨 Episode 94 of The Intelligence Journal is here. NASA Open-Sourced an AI Model of the Moon 🌙🛰️ NASA and IBM just released a foundation model trained on the Moon itself. Two million image tiles. Seventeen years of Lunar Reconnaissance Orbiter data. Over a million high-resolution frames at one metre per pixel, plus nearly 964,000 multispectral images, stitched together with data from GRAIL, Lunar Prospector and Japan's SELENE. It maps craters, dates surfaces, spots volcanic features that break existing cooling models, detects changes between passes, and estimates where ice is stable near the poles. It beat the baselines on ice prospecting 🧊 And they put the whole thing on Hugging Face with the code on GitHub. Free. Join us as we break down: - Why a general model trained once on raw sensor data beats a decade of purpose-built algorithms, and why that lesson transfers to every domain sitting on unused data. - What finding stable ice actually unlocks, because water on the Moon is fuel, air and the differe
MCP
1 item
official sitethis guy built a robinhood chain tool that makes his ai check the pool before it talks about his bag meet TRENCH give an MCP-compatible agent a token address and a USD position size: → reads public Robinhood Chain pools through DexScreener → checks the chain and latest block through RPC → selects the highest-liquidity observed pool → models how the position size presses against estimated depth → returns a DEEP / THIN / CRITICAL grade, the reason and the assumptions → lets the agent explain the evidence a token can have the same price, the same chart and the same pool - and still be a very different problem at $500 versus $5,000 that's the part he wanted the agent to calculate instead of hand-wave the loop is: TOKEN + SIZE → PUBLIC DATA → PRESSURE MODEL → AGENT EXPLANATION five read-only MCP tools, open source, a browser playground anyone can try without installing anything the browser uses manual inputs, the MCP server is where live token data comes in, neither is an executable sell quote no private
Apple
1 item
official siteMiniMax H3 is now on my preferred image and video generator on Apple Silicon! I need to try it!
@drawthingsapp🚨 MiniMax H3 is RIGHT HERE in Draw Things!Go try it NOW! > 🔄 Draw Things v26.0910.1 was released in the iOS / macOS AppStore 2 hours ago. This version brings: > 🔹 Support MiniMax H3 series models, including LoRAs and TeaCache; 🔹 Support importing Krea 2 series models. 🔹 Fix some performance issues on M4 Apple Neural Engine. > gRPServerCLI and draw-things-cli both are updated to 26.0910.1 with above related updates.
MiniMax H3
MiniMax / 1 item
MiniMax H3 is now on my preferred image and video generator on Apple Silicon! I need to try it!
@drawthingsapp🚨 MiniMax H3 is RIGHT HERE in Draw Things!Go try it NOW! > 🔄 Draw Things v26.0910.1 was released in the iOS / macOS AppStore 2 hours ago. This version brings: > 🔹 Support MiniMax H3 series models, including LoRAs and TeaCache; 🔹 Support importing Krea 2 series models. 🔹 Fix some performance issues on M4 Apple Neural Engine. > gRPServerCLI and draw-things-cli both are updated to 26.0910.1 with above related updates.
Devin
1 item
SWE-2 absolutely nailed its first task. Past agent integrations always required a second pass from a SOTA model like Sol or Fable to get things working properly. Devin got it right on the first go ✅
@dedeneReally excited to try out Devin for the first time. 🤩 Huge thanks to @ryancarson and @dabit3 for this! > Let's see what SWE-2 can do. 🚀
Claude Fable
Anthropic / 1 item
GPT 6 Sol and Opus 5.1 are both coming within the next 2 weeks. These are the releases that actually matter. Fable 5.1 and GPT 6 Astra are incredible models that nobody can afford to run consistently. Session limits die in minutes. Weekly limits die in days. We do not need smarter models right now. We need intelligent models that are affordable enough to actually use in our workflows. GPT 6 Sol and Opus 5.1 are supposed to be exactly that. The biggest week in AI is not about the frontier. It is about making the frontier usable.
Seedance
ByteDance / 1 item
Who needs smooth sailing when you can surf actual lava? 🌋🔥 Put the new Seedance 2.5 model through its paces with this one, and honestly, the physics on the liquid motion and rope swings blew me away. Huge fan of how natural the camera flow feels here.
Hermes Agent
1 item
One of the 3,000, and pretty damn happy about it. Love building Hermes with this community!
@TekniumHermes Agent has just hit 3000 contributors. > Thank you to all of the developers who have worked to make hermes better for everyone!
Grok Bot
xAI / 1 item
DAILY AI BRIEF 🗞 — Sept 12 OPENAI 🔥: - GPT-Rosalind is out of research preview for eligible orgs worldwide. API, Codex, and ChatGPT Enterprise, with new Rosalind models as they ship. - Codex adds Life Sciences plugins for genomes, protein structure, QC reports, and notebooks. - ChatGPT Sites hit 5M apps. New: collab editing, private invites, custom domains, and DB inspect. - Desktop pets can start a new chat. Mini is the compact no-pet option. XAI 🔥: - Elon: Grok 4.7 needs a few more days. RL still quits hard tasks too early and undershoots self-checks. - Grok Bot is rolling out on Grok web for Heavy users. Create and chat with bots in the UI. No official post yet. ANTHROPIC 🔥: - Claude Code ships `claude plugin eval`. Score a plugin on test cases, then rerun without it. MICROSOFT 🔥: - MAI-Transcribe-2 hit 1M OpenRouter requests in 5 days. ALIBABA 🔥: - Qwen3.8-27B is live on Cerebras. * Used Grok to compose this brief, cherry-picking the news and doing some post-editing. > [@testingcatalog](
Meta
1 item
official siteDAILY AI BRIEF 🗞 — Sept 12 OPENAI 🔥: - GPT-Rosalind is out of research preview for eligible orgs worldwide. API, Codex, and ChatGPT Enterprise, with new Rosalind models as they ship. - Codex adds Life Sciences plugins for genomes, protein structure, QC reports, and notebooks. - ChatGPT Sites hit 5M apps. New: collab editing, private invites, custom domains, and DB inspect. - Desktop pets can start a new chat. Mini is the compact no-pet option. XAI 🔥: - Elon: Grok 4.7 needs a few more days. RL still quits hard tasks too early and undershoots self-checks. - Grok Bot is rolling out on Grok web for Heavy users. Create and chat with bots in the UI. No official post yet. ANTHROPIC 🔥: - Claude Code ships `claude plugin eval`. Score a plugin on test cases, then rerun without it. MICROSOFT 🔥: - MAI-Transcribe-2 hit 1M OpenRouter requests in 5 days. ALIBABA 🔥: - Qwen3.8-27B is live on Cerebras. * Used Grok to compose this brief, cherry-picking the news and doing some post-editing. > [@testingcatalog](
Muse
Meta / 1 item
DAILY AI BRIEF 🗞 — Sept 12 OPENAI 🔥: - GPT-Rosalind is out of research preview for eligible orgs worldwide. API, Codex, and ChatGPT Enterprise, with new Rosalind models as they ship. - Codex adds Life Sciences plugins for genomes, protein structure, QC reports, and notebooks. - ChatGPT Sites hit 5M apps. New: collab editing, private invites, custom domains, and DB inspect. - Desktop pets can start a new chat. Mini is the compact no-pet option. XAI 🔥: - Elon: Grok 4.7 needs a few more days. RL still quits hard tasks too early and undershoots self-checks. - Grok Bot is rolling out on Grok web for Heavy users. Create and chat with bots in the UI. No official post yet. ANTHROPIC 🔥: - Claude Code ships `claude plugin eval`. Score a plugin on test cases, then rerun without it. MICROSOFT 🔥: - MAI-Transcribe-2 hit 1M OpenRouter requests in 5 days. ALIBABA 🔥: - Qwen3.8-27B is live on Cerebras. * Used Grok to compose this brief, cherry-picking the news and doing some post-editing. > [@testingcatalog](
OpenCode
1 item
Plannotator 让 Agent 写的计划,我们可以在网页上打开审查,在段落上划线、写批注,点一下就把意见整包送回给 Agent 改。 代码写完也一样,改动以左右对照的形式打开,逐行评论、直接给建议代码,GitHub 和 GitLab 上的 PR 也能拉进来评。 GitHub: Claude Code、Codex、Copilot CLI、Gemini CLI、OpenCode 等 9 个编码 Agent 都接了,一个安装脚本自动识别装了哪些,对应配置一并配好。 Agent 生成的 HTML 页面也能直接渲染出来在上面批注,评论时还能随手问 AI,或者让它先跑一轮审查把评论贴到改动上。 计划、改动、批注全存本地,不收使用数据。想让审计划这一步从终端里挪出来、能划线能批注的,可以装上试试。
Also recorded
15 items that named no organisation or product this site tracks.
THIS FREE OPEN-SOURCE REPO GIVES YOU A SELF-HOSTED PERSONAL AI AGENT WITH MEMORY, WEB SEARCH, SUBAGENTS AND SCHEDULED AUTOMATIONS. NANOBOT RUNS WITH YOUR OWN API OR LOCAL MODEL AND CAN EVEN USE YOUR TERMINAL TO GET REAL WORK DONE.
说实话,现在YouWare 做调研报告速度不是最快的! 但是,真的报告的质量还是很顶的,尤其是内置了金融机构专业数据。 可以通过专业的调研专家团输出多种报告格式、html、doc、md等格式。 最牛逼和方便的时候,直接可以通过一个link链接分享给我其他人即可加入对AI的“PUA”。 直接大家一起协作,安全性也拉满,分享链接实效性安全性都做了优化。 新产品,还是值得推荐。 别为了做AI垃圾报告而浪费时间,质量更重要。
@YouWareAI🙅🏻♀️Stop gambling on one-shot AI reports. > Expert team runs the research. You get real finance deliverables numbers, charts, editable files. > Then share the link. Your teammate picks up the same task, same context. > Collaboration, for real.👇🏻
Small models are starting to punch way above their weight. MiniCPM5-2B isn’t trying to win by being huge. It can actually reason through a codebase use tools investigate bugs and make fixes locally. So I put it behind a coding agent and let it tackle a real repository instead of another benchmark. The result was surprisingly good.
@akshay_pachaarChinese researchers did it again! > OpenBMB just open-sourced MiniCPM5-2B, a dense 2B-parameter model built for reasoning, coding, and tool use on resource-constrained hardware. > Artificial Analysis ranked it highest among models under 4B in its Agentic Index comparison. It scored 20, while Granite 4.2 8B scored 9. > The model is particularly strong at coding and tool calling, so I tested both capabilities locally. > I pulled it onto my machine, connected it to a constrained CI repair agent, and gave it one issue: > A customer reports that retrying checkout with the same idempotency key returns a larger
Lightricks LTX 2.5 😀 Ingredients Ref Sheet IC LORA conditions video generation on a reference sheet-a single composite image inventorying the characters, props, and location of a scene so that generated videos keep those elements visually consistent 👇
LTX-2.5 🤠 IC-LoRA Deblur It restores sharpness to out-of-focus / defocused video by conditioning on the blurry clip and regenerating it in sharp focus while preserving the original subject, framing, and scene geometry. 👇
Marigold V2 🧐 Comfy Support designed by HUAWEI Bayer Lab for dense visual prediction, it takes a standard 2D image and calculates detailed, pixel-by-pixel convert into 3D structured data and surface properties. 👇
LTX-2.5 🧐 IC-LoRA Colorization restores natural color to grayscale, monochrome, or desaturated video while keeping subject identity, framing, and geometry untouched — only the color information changes. 👇
LTX-2.5 🌊💧 IC-LoRA Water Simulation It adds water to a clip rivers, surf, rain, waterfalls, floods, splashes, spray, and wet-surface specularities while keeping everything identical to the reference. 👇
LTX-2.5 🌞🌙 IC-LoRA Day-to-Night Relighting It re-renders a daytime video as the same shot at night while preserving composition, framing, camera movement, and subject motion frame-for-frame. 👇
LTX 2.5 😃 Clean Plate IC-LoRA It removes people and other dynamic subjects from a source video and reconstructs the empty background, producing a clean "clean plate" of the same scene 👇
GPT Image 2.5 is built for more than one great image. Better consistency. Better editing. Better control. One image → a whole visual sequence. Go play on WaveSpeed
Lightricks LTX-2.5 😀Cinemagraph LoRA Transforms still images into smooth, looping cinemagraphs with selective motion while keeping the rest of the scene frozen. 👇
Understand Anything uses an agent pipeline to transform codebases into interactive knowledge graphs that map every file, function, and dependency.
Recreating your sketch has never been this easy. Powered by GPT Image 2.5 on HIX AI 👉
@LudovicCreator Midjourney v8.2