Skip to content
B Bloger.fm

Product / OpenAI

GPT-6 Astra

This site has not written a definition for GPT-6 Astra. The name appears in the feed but has no stable public identity the desk can verify, so what follows is only what was recorded — no description, and no claim about what it is.

GPT-6 Astra was recorded in 102 items across 8 of the 8 briefings in the current window. That is more than any other entity this site tracks.

Its share of coverage was steady: 39 items in the first half of the window and 63 in the second, tracking the feed as a whole, which grew about 2.2×.

It appeared most often alongside OpenAI, GPT-5.6 Sol and Claude Fable.

Tracking the feed
items
102
briefings
8
mentions
251
last seen
2026-09-19

Coverage timeline

Sat 12 Sept – Sat 19 Sept / 8 briefings

Everything recorded

19 September 2026 11 items

🚨 WAIT… Is Anthropic getting ready to ruin OpenAI's GPT-6 Plan? 👀l Opus 5.2 is reportedly looking seriously strong in early testing. And now we're hearing about THREE possible releases: → Opus 5.2 → Fable 5.2 → Sonnet 5.1/5.2 Meanwhile, Astra just dropped. GPT-6 is supposed to be OpenAI's next big move. But if even two of these Anthropic models actually land… OpenAI might not get the clean runway it wants. Anthropic could be loading an entire fucking wave while OpenAI prepares its next move. September just got VERY fucking crowded. 👀🔥

I just tried the new Gemini-4 Pro Preview and honestly, for 3D work it’s right up there with Fable 5.1 and GPT-6 Astra, all in a single shot, I genuinely didn’t expect it to be this good, It only used around 10% of my weekly limit on the $20 plan, I’m going to make a comparison video against Opus 5.2 or Fable and Astra if I can, I still can’t believe it, From what I’ve tested so far, this is easily top 1 or 2 for the best one shot execution I’ve seen

🚨 Deepseek v5 Leak: Beats Astra > DeepSeek is reportedly preparing an imminent V5 launch > it could match or beat Fable 5.1 and Astra > Expected to deliver much stronger performance at a lower cost > DeepSeek is reportedly keeping the open-weight strategy Could DeepSeek V5 become the new king of open-weight AI?

So OpenAI accepts that their SOTA frontier models “suck at design” 😂

@Voxyz_ai

If you still think Sol and Astra suck at design, try this: > “Use imagegen to reimagine this page, then implement it.” > Codex now has 𝗜𝗺𝗮𝗴𝗲𝘀 𝟮.𝟱 built in, the strongest image model available today. Let it create the visuals first, then have Sol or Astra implement them. > Many of the 3D game scenes, characters, and animations shared in posts are built around this same approach.

Gemini 4 Pro’s first leaked output early checkpoint - Clean HUD a playable racing game. - Smooth world gen. - Instant overtake from 8th to 1st. -

@0x0SojalSec

Google just dropped Gemini 4 Pro beat Astra & Fable 5.1. > - the first checkpoint show. - It’s expected to beat Astra and Fable 5.1. - October launch is the main window. - Late September is still possible. - Google’s own exec framed the massive AI spend as a bet on RSI not a claim that they’ve already hit it.

New GPT-6 Sol leaks: - OpenAI just delayed GPT-6 Sol. - The reason? They’re reportedly dropping a Muse competitor next week possibly hardware. - That’s also why Astra slipped (likely Thursday). - Some users were already being routed from Sol 5.6 to Sol 6. - Now that routing has stopped. - Both models share an April 30 cutoff.

truncated at source

I pushed pretty hard on one particular complaint after the GPT-6 Astra rollout: Plus users felt overlooked because they got almost nothing new in Chat itself, only a small taste of GPT-6 Astra through Codex and ChatGPT Work. Those posts clearly reached OpenAI. I noticed several people there saw them, and some employees even liked them. Now they’ve started vagueposting about something related specifically to Chat, not Work or Codex again. My current read is that they may be preparing one of three things: → a Chat-specific model variant → a new routing / model tier for Chat → a broader rethink of the Chat experience itself And there’s another reason I’m paying attention: Tibo was replying to Matthew Berman, someone I consider pretty reliable and who regularly gets to test things early. So no, I’m not treating this as confirmation yet. But I definitely don’t think they’re talking about “Chat” for no reason. And if you’ve followed me for a while, you know I rarely drop an issue after one post. If somethin

slowdown? what slowdown?

@Techmeme

Sources: Anthropic considers releasing a new AI model to counter OpenAI's momentum since Astra's launch, ahead of an IPO and after Amodei's call for a slowdown (Reuters) > (Visit Techmeme dot com for the link and full context!)

Fable、Opus、Sonnet 三条线齐齐推进,看来 Astra 这波是真把 Anthropic 压到防守位了。

@synthwavedd

New versions of Fable, Opus and Sonnet are now being stealth tested across different Claude surfaces and accounts. Seems Anthropic may revive their old full lineup launch convention 👀

truncated at source

DAILY AI BRIEF 🗞 — Sept 19 XAI 🔥: - Grok Voice Transcribe 2.0 is live in the Grok Voice API. $0.10/hr batch, $0.20/hr streaming. META 🔥: - Muse connectors are live for developers. You bring the API; Muse brings the agent, browser, and user context. - Muse is now available in Canada. - A dedicated Muse Mail tab is in development. OPENAI 🔥: - ChatGPT desktop browser now runs Chrome extensions. - Most plugins can connect multiple accounts in one chat. Devs can add a profile tool so ChatGPT labels them. ANTHROPIC 🔥: - Claude Code 2.1.277 reads AGENTS.md when no CLAUDE.md is present. Toggle in /config. - Partnering with Accenture on embedded frontier eval. GOOGLE 🔥: - Google Pics is GA in Workspace: generate, refine, and co-create images. - Dreambeans is GA from Labs: a daily personalized story collection. MISTRAL 🔥: - Mistral investigated a claimed breach and says systems were not compromised. * Used Grok to compose this brief, cherry-picking the news and doing some post-editing. ** This daily brief also a

truncated at source

A significant benefit is the ability to compare different agents and models when performing the same task. While performance is important, cost is also a crucial factor. Identifying a model that achieves the desired outcome at a considerably lower price point provides the kind of adaptability required by users of artificial intelligence.

@quxiaoyin

We just launched Agentsky @agentsky_dev, world’s 1st Agent Market! OpenRouter is for models. AgentSky is for agents. > Use 40+ agents—Claude Code, Codex, OpenCode, Hermes, Pi in your browser(even your phone!) or via one API. All without installing or setting up anything. > Hit Codex Astra’s weekly limit? Hand off to another agent such as OpenCode + DeepSeek V4.1 in browser without losing any context. > You can compare any agent + model directly in browser and that's how I found Astra costs $5.3 while deepseek v4.1 cost $0.12 on the same dashboard task. (I actually preferred deepseek) > Try it at

18 September 2026 15 items

open source agents can now trade for you as Superior Trade just open sourced their AI trading terminal > draw trade idea directly on chart > turn it into an entry, stop and target > add your own prompts, leverage limits and safety rules > run everything locally with your keys on your machine > fork the repo and build whatever you want on top never knew i'd be trusting agents with my money but here we are already connected my agent to astra and gave 50 bucks to test, let's see what it does will it make money or not?

fable 5.2 > gpt 6 astra tested both models with same prompt at high reasoning and results came out really different fable next model entered stealth testing today, so i don’t think it's coming out anytime soon from early testing: > huge step up from current fable > output looks better than astra > slower and expensive still running more tests and will drop the outputs soon

SITUATION DETECTED: OpenAI launched Astra for Law: GPT-6 Astra with a U.S. legal research index and instructions for legal analysis and writing.

AI won’t replace lawyers but it will replace the lawyers still thinking AI won’t replace lawyers

@OpenAI

Astra for Law: Frontier intelligence built for your practice. > A new offering powered by GPT-6 Astra with tools, settings, and context to support the expertise and judgment of lawyers and legal technology firms.

Muse Spark 2 Leak: Zuck Is Already Building The Next One 🥑 >Meta confirmed the next-gen Muse model is already in development >Rumored to be able to compete with GPT-6 Astra and Fable 5 >Reportedly bumping context window to 2M tokens Heavier focus on agentic workflows and computer use >Reportedly very inexpensive to run No official name, benchmarks, or release window confirmed yet Meta previously admitted Muse Spark 1 couldn't keep up with rivals this would be the answer to that

Is it even possible that a traditional LLM beats it in speed? I hope OpenAI releases ultrafast, computer use with Astra's or Sol's intelligence + speed of Jev would be crazy. For me it would be computer use AGI

@trycua

1/ Fast Computer Use is now solved with @typesafeai Jev + Cua Driver. > Available in development preview for macOS, Windows, and Linux. We call it jev-use. > Draft #3943:

OpenAI GPT-6 Astra actually does have observable chain of thought, according to OpenAI Research Scientist Noam Brown. “Chain of thought monitoring… it’s a real gift. We were very lucky that this ever existed, and it is fragile.” 🚀 Watch the full episode of AI Deep Dive:

Astra challenge is live on Product Hunt!

I used to spend $9M/month on Meta Ads Now I’m engineering that workflow into Higgsfield x GPT-6 Astra. I used to test 4,500 creatives a month for \~50 winners at a 1% hit rate. But such volume needs resources. 🧵 We made 8 skills for Paid Ads to make it with Astra – save this

@higgsfield

Meet Higgsfield x GPT-6 Astra for Paid Ads. > With our ChatGPT plugin, GPT-6 Astra: > Runs your ad account: launches ads, tests creatives, and scales campaigns > Analyzes customer needs across social media > Iterates hooks with new product angles > Type @Higgsfield /marketing and run your ads from ChatGPT.

Qwen 4 Leak: Could Land In The Next 15 Days 🔥 >Expected to reach GPT-6 Astra and Fable 5.1 level performance >Rumored 3T+ parameters for the full Qwen 4 line Could become the most powerful open-weight model >Expected to remain open-weight >Major focus on reasoning, coding, and autonomous agents >Full multimodal capability expected >Strong agentic capability a core focus >Reportedly very inexpensive to run Can Qwen 4 actually beat GPT-6 Astra and Fable 5.1?

Okay, Codex Spark had a good run. But be honest, we are all thinking the same thing now. GPT-6 Sol Spark. Or somehow, Astra Spark. That is the Spark I would come back for.

@thsottiaux

Next week we’ll be retiring GPT-5.3-Codex-Spark. Can you believe we shipped a model named as such!! > It's had a good run and was a lot of fun, but usage has been declining and we have significantly better models now. Time to make room for the future.

🚨 Gemini 4.0 Pro specs leaked • Pricing: $2.25/$11.25 • 2M token context window • Beats GPT-6 Astra and Claude Fable 5.1 on almost every reported benchmark • While being significantly cheaper than both

truncated at source

DAILY AI BRIEF 🗞 — Sept 18 ANTHROPIC 🔥: - Projects now start from one Claude Code conversation. Claude spins parallel cloud threads, keeps shared memory, and surfaces an Overview panel. XAI 🔥: - Grok Bot voice is live. Desktop and mobile, rolling out over the next couple of days. META 🔥: - Muse for Mac is out, US only. Computer use across apps, files, calendar, notes, and messages. You pick what it can access. PERPLEXITY 🔥: - Effort selector is live in Computer on web. Presets pair the orchestrator model with reasoning depth. Mobile and desktop next. OPENAI 🔥: - Astra for Law is out: GPT-6 Astra plus a Legal Search Index over 230M+ URLs. Trusted Access first, API soon. - ChatGPT in Word hits all plans including Free, with usage limits. Business and Enterprise get a two-week GPT-5.6 Sol preview. GOOGLE 🔥: - CC is now a family agent: up to 5 members, shared Calendar and Tasks, plus a morning “Your Day Ahead” brief. Waitlist, US 18+. ALIBABA 🔥: - Qwen3.8-Omni-Flash is out. First omni-modal agent model, 1M

Haha! The rumors about Gemini 4.0 are phenomenal They are aiming to surpass both Astra and Fable 5.1 and are on track to do so The pricing will be super competitive as well.

Try building Higgsfield clone with our API x GPT-6 Astra. The only US-based Seedance 2.5 with consistent characters. Up to 50% off discount on top models.

17 September 2026 14 items

truncated at source

The expensive mistake in AI video usually happens before rendering: the scene itself was wrong. A better model doesn't fix a bad camera path or broken spatial layout. GPT-6 Astra is now live on OnSolo, and the Whitebox Video Expert makes the most sense to me. I can describe a scene, Astra reasons through the space and motion, and the system gives me a whitebox pass before the final render. So I can validate blocking, object placement, and camera movement while everything is still cheap to change.

@OnSoloAI

GPT-6 Astra is now LIVE on OnSolo 🔥🔥 Multiple major updates just dropped, all powered by GPT-6 Astra. Pick your lane and start creating. > 1️⃣ Game Section (Members only) Feed in a game idea. Agent spins up a playable web game — one click. 🎁 7-day limited free trial. Canvas usage limits apply: Premium · 1 use/day | Super · 2 uses/day | Ultra · 3 uses/day > 2️⃣ Whitebox Video Expert (Free for all · Limited time) Canvas Expert Whitebox Video Ex

truncated at source

GPT 5.5 was released on April 23, 2026 GPT 5.6 Luna was released on July 9, 2026 And it meets 5.5 in intelligence while crushing it in value Things are moving quickly

@morganlinton

Well, I'll admit it, I was wrong, and @steipete was right. > I'm okay being wrong, and this one I really ate my own words on. > I've been using GPT 5.5 more lately, because I love the model, and had an idea in my head that it was still a good choice for easy/medium difficulty coding tasks. > For hard stuff, in the GPT family, I would still reach for Sol or Astra, but I just felt like GPT 5.5 was stronger than Luna or Terra for whatever reason. > But I had never benchmarked it. And when I wrote about me using GPT 5.5 last week, Peter told me, that was silly, and I should be using Luna instead. > So as a benchmarker, I realized, okay, let's benchmark it, and I did. And well, yeah, Peter was very right. > Luna comes in at the same accuracy as GPT 5.5, as long as you use it

GPT-6 Astra deciphered a 1918 German radio transmission that, to my knowledge, has never been deciphered before. The message below translates to: "EIN ENGLISCHER KREUZER EINLIEG X SEWASTOPOL X S4STEN X EIN GESCHWADER DER X ALLIIERTEN FOLGT 26STEN X" or, in English: "AN ENGLISH CRUISER ARRIVED AT SEVASTOPOL ON THE ?4TH AN ALLIED SQUADRON FOLLOWS ON THE 26TH" Astra even double-checked its work by determining that an English cruiser, HMS Canterbury, reported its arrival in Sevastopol on November 24, 1918 and the arrival of an allied squadron on November 26, 1918. This message is one of the ~20 WWI German radio messages that appear as one of the entries in the list of top 50 unsolved ciphers (). A minor, but really cool result!

Union Alpha could be ZLM 5.4 👀 >A new anonymous stealth model called Union Alpha has surfaced, free to use >256K context window, multimodal, built for agentic coding >Frontier-level general-purpose performance with tool calling built in >Performance reportedly sits near GPT-6 Astra and Opus 5 at roughly 18x lower expected cost >Naming pattern echoes Ox Alpha, which turned out to be GLM-5.3-Flash speculation is already pointing to an early GLM-5.4 test No lab has claimed it yet unconfirmed as of now Try it yourself and see if you can spot who's really behind it.

The performance numbers are what make this one hard to ignore. It is landing close to GPT-6 Astra and Claude Opus 5 on coding benchmarks. It is doing that at roughly 18 times lower expected cost than either of them. If that holds up under real independent testing, this is not a small gap. This is the kind of cost difference that changes which model teams actually choose to run agents on all day, every day. Whoever is behind Union Alpha clearly optimized for being cheap enough to use constantly, not just for winning one leaderboard screenshot.

@cline

Union Alpha (stealth model) is now free in Cline. > 256k context, multimodal, built for agentic coding. > It is near GPT-6 Astra and Opus 5 performance for ~18x lower expected cost.

People underestimate how much better Astra is than Sol in its ability to have novel discoveries. Would not expect this trend to slow.

@ValsAI

Scientific discovery is the next frontier for AI systems. However, new scientific results are difficult to verify, and therefore hard to measure. Today we're releasing MysteryMechanism, a benchmark that tests whether agents can rediscover sealed mathematical mechanisms through bounded experiments. It cuts sharply across the frontier, with Astra landing roughly 20pp above Sol.

Sam Altman at Dreamforce ranked how fast AI got smarter at math. GPT-5.5: as good as an average math professor. GPT-5.6: top 1-2% math professor. Astra: a little better than that. Then he said their internal model past Astra "can do things the best mathematicians in the world cannot." That progression happened in roughly 4-6 months. An internal OpenAI model reportedly helped produce a proposed proof for the Navier-Stokes problem. One of the seven Millennium Prize Problems in mathematics. Unsolved for decades. 10,000 AI agents. 88 hours. Millions of messages. Hundreds of billions of tokens. We went from "as good as an average professor" to "better than the best humans alive" in less than half a year. -- vc: @rohanpaul_ai

81.94% accuracy for $2.26. That’s GPT-6 Astra on the new BrokenArXiv / ArXivMath benchmark. Claude Fable 5.1 gets 79.76% for $19.57. So Astra is not just ahead on accuracy. It’s doing it at roughly 1/9th the cost. Thats bonkers

@thsottiaux

Astra > ✅ Fast ✅ Frontier ✅ Efficient ✅ For everyone

For OpenAI Plus users, I’ve got some good news: I was told that GPT-6 Sol is the main release aimed at this broader audience. From my testing so far, the model is incredible at creative writing, 3D modeling, frontend design, and several other broad areas where it gets very close to Astra, at a more accessible price and with higher usage limits. I still haven’t been able to test GPT-6 Luna, which also looks like it’s going to be one of the main attractions through next weeks. I’ve been putting a lot of pressure on OpenAI to improve the limits, ideally something close to the 3,000 messages we used to have in Chat mode on Plus, but I don’t think it’s going to happen that way.

GPT-6 Astra built a Blaze farm, reached the Crimson Forest and collected Ender Pearls in Minecraft, then a Creeper blew up its chest and it spent hours farming potatoes, watching the rain and repeating warnings to itself about chest storage.

Meta doesn't even need to be frontier to win! They have such massive distribution that just having close and cheaper is a massive alpha

@Mr_Salio

🚨 Muse Spark 2 Leaks: Beats Astra > Spark 2 could compete with GPT-6 Astra and Fable 5.1 > Meta is already developing the next-gen Muse model > Expected to be extremely cheap to run > A 1M-token context could carry over > Meta admitted Spark 1 struggled against the competition > Spark 2 could be Meta's answer > Expected late this month or next month > Could Meta finally have a serious frontier model?

🚨 Grok 4.7 better not be fucking mid today. 👀 August: “It’ll beat every model out there.” September: “Roughly on par with Opus 5.0, not 5.2” And the Astra/Fable-class model? Apparently that’s Grok 4.9 😭 Bro went from “best model in the world” → “wait for 4.9” After all that hype and all those delays… What the fuck did xAI actually cook? 👀🔥

Astra-Qwen 3.8 flash next loop is nice 😊

truncated at source

We rolled out GPT-6 Astra to every Databricks engineer today. It beats Claude Opus 5 on the hardest, long-horizon tasks and increased our coding spend by 60%. I think it’s the strongest model. Expensive. But hopefully worth it.

@pwendell

Today we rolled out Astra to every engineer at Databricks (N=~3500). Some notes that may be helpful to others: > 1. Astra unambiguously out performs our previous highest-end models (Opus 5, Sol 5.6) on highly complex tasks, especially those related to high level system design or long range horizontal tasks. > 2. Engineers given Astra increased overall coding spend by around 60% compared to baseline. > 3. It is not clear Astra meaningfully improves on medium/low complexity coding tasks compared to earlier models. We suspect those tasks are mostly saturated (i.e. perfectly executed) by existing models. > 4. We learned above by piloting Astra with around 200 users to gain signal on both quality and cost. We use Unity Gateway to

16 September 2026 23 items

Today we're introducing ChatGPT Astra for "Web to App" Clone any website into a mobile app. Just paste a URL. GPT-6 Astra controls your Mac to rebuild the original website as a *native* mobile app, then submits it to the app stores for you. We've been using this mostly for iOS apps. > @chhddavid: >

GLM Coding 2.0 Leak: Coming Soon 🔥 >Expected to beat Mythos 5.1 and GPT-6 Astra >October release window reportedly targeted >Computer use expected to be a top priority >Expected to remain open-weight >1M-token context window reportedly carried over Rumored 3T+ parameter model

Wenfeng: "that's not continual learning, this is prompt engineering bullshit" but to be fair: that may be all it takes for practical purposes. These things already have superhuman priors, can RLM over infinite databases, and think crazy fast. Do we *need* parametric updates?

@NeoCognition

Introducing ApprenticeBench: computer use + continual learning on a real job. > We show Fable 5.1 and GPT-6 Astra can now continually learn on a job and surpass human professionals. A decisive step change in AI's job readiness. > No FDEs. Agents deploy themselves into the job. 🧵

THIS IS GETTING RIDICULOUS Anyone can now build Duolingo-style retention flows with ONE prompt. Here’s how it works: > connect your customer list > tell @noimos_ai the goal > AI builds the flow > reacts to user behavior > sends personalized messages The agent handles all of this autonomously 🤯

@noimos_ai

Introducing Astra for Customer Retention & Growth. > Just describe your goal. NoimosAI autonomously creates and delivers personalized campaigns that keep customers engaged and coming back. > Maximize customer lifetime value. All on a single platform.

Sam Altman after finding out his $20 GPT-6 Astra quietly opened a trading desk on Robinhood Chain, worked all night while the NYSE was closed, and closed its 84th call before anyone noticed it existed..

@slash1sol

I GAVE GPT-6 ASTRA AND MINARA ONE JOB: WATCH ALL 24 ROBINHOOD STOCK TOKENS UNTIL THE PRICE STARTS LYING -> THREE WEEKS LATER THE DESK THAT DOES IT IS RUNNING > 24 tickers, 4,320 pool reads an hour, 103,680 a day and 84 gaps called, 95% of them closed. > It's live and it's free: > Why a Stock Token stops tracking the share it is named after: > A meme pair locks real shares inside a pool and the float on chain goes thin. > Minting and burning run on a schedule, so outside that window supply cannot answer demand. > After 4PM the oracle stands still while the pool keeps trading anyway. > The pool price drifts off the real price, and that drift is the arbitrage -- buy the cheap leg, short the rich one, wait for them to meet. >

Developers are still having fun with a digital fruit fly brain. This time, they’ve got it playing Deadlock with help from GPT Astra The 166,000-neuron model was trained on recordings of other players’ matches. Computer vision helps it understand what’s happening on screen, while a separate algorithm handles aiming The fly runs around the map, joins fights, and plays best as Graves. According to its creator, it even performs better than his teammates > @Mikadzyki_NFT: >

very cool result showing how wet lab data enables a specialized model to beat gpt-6 astra at a task at the frontier of science! in general scaling is great and obviously i am a believer in it, but probably the more we approach the frontier of science, the more specialized data matters and that gives task-specific models a chance. this specialized data is usually private and is probably a real moat it should be in principle true that a task specific model will probably do better at scientific discovery just because it can use more of its parameters for the task you care about

@LiamFedus

We built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, and then help us decide what to try next. > Using only 1,300 H200s, plus months of our experimental data, we mid-trained and RL’d an open-source model to surpass GPT-6 Astra on our analysis benchmark. We call it Neon. > This

Astra is mogging fable soo much I’m use to see any benchmark fable actually beat it on

@j_dekoninck

We are releasing the latest version of BrokenArXiv and ArXivMath! These benchmarks now focus on conjectures that were refuted in the last month on ArXiv, and models are executed within a harness instead of directly via API. > Performance remains impressive, with GPT-6 Astra on top

Woow It launched 14 sub agents in parallel with GPT-6 Pro and ran for 165 minutes, more than two and a half hours , I specifically told it to launch completely independent and impartial judges, and if the final score was below 9.5, it had to start another round of research agents, auditors, and correctors until the result improved, and the craziest part is that it didn’t use any of my Codex limits, It basically feels unlimited,I’m going to run a lot more exhaustive tests and share all the results with you guys

@SPAC89

🚨This might be the biggest ChatGPT update since it launched, You can now launch multiple agents directly from the normal ChatGPT chat using GPT 6 Astra Pro, basically at no extra cost since the limits are almost unlimited, This means we can save a ton of our separate Codex usage too, I’m testing it heavily right now to see just how many agents I can run in parallel on the Pro x20 plan, Enjoy!

truncated at source

astra is the first ai that i believe: + given enough time it would solve anything. even after the fateful creeper explosion of all valuable stuff, it can learn from the lesson and overcome with this message: "The new chest has been destroyed, and the stored valuables are missing. I’ll keep all future critical items in my inventory, where keepInventory protects them, and rebuild the supplies while continuing toward the dragon." you can also see in the chart there are basically no plateau whatsoever. + has emergent behaviors that are unusual to humans, while still showing some familiar traits. in its messages, astra has show some emotions frustration or self-criticism, but at the same time it counts pixel when aiming a bow and writes its scratchpad in increasing more cryptic ways. i believe that if we can see raw chain of thoughts, there might be more evidence of its intentions and emotions behind its actions. ai's thoughts and behaviors may diverge from ours and become completely strange to humans, and we w

🚨This might be the biggest ChatGPT update since it launched, You can now launch multiple agents directly from the normal ChatGPT chat using GPT 6 Astra Pro, basically at no extra cost since the limits are almost unlimited, This means we can save a ton of our separate Codex usage too, I’m testing it heavily right now to see just how many agents I can run in parallel on the Pro x20 plan, Enjoy!

3 days left to schedule your launch for the GPT-6 Astra Challenge. Launch on Product Hunt this Friday, September 18, for a chance to win. The top five launches each get: · $10K in @OpenAIDevs API credits · 1 year of ChatGPT Pro for up to two team members If you’re building with Astra, get your product in front of the community and locked in on the calendar now. Schedule your launch here:

Okay, this is actually pretty interesting. I was looking into OpenAI’s GPT-6 Astra, and one thing immediately caught my attention: It’s designed to interact with computers and complete multi-step tasks. Here’s what I found 🧵

This add-on blew our Head of Animation’s mind. We used GPT-6 Astra and Higgsfield to build a Blender shader add-on with 15 ready-to-use materials. Apply them to your objects and give entire 3D worlds a hand-painted look.

truncated at source

🚨 GPT-6 Sol : Reportedly launching THIS THURSDAY OpenAI might be moving insanely fast with the GPT-6 lineup. > GPT-6 Sol is reportedly targeting September 17 > Sol has reportedly already entered internal testing >Early testing suggests it's significantly faster than Astra One reported test generated ~28K tokens in ~3 minutes The same task reportedly took Astra ~19 minutes for ~25K tokens >Sol is rumored to trade a little of Astra's maximum reasoning depth for speed + throughput >Could become the everyday workhorse of the GPT-6 family >Some reports suggest much better usage limits than Astra >Terra and Luna could potentially follow as additional GPT-6 variants OpenAI has not officially confirmed Sol or a Thursday launch And this is what makes Sol interesting. Astra appears to be built around maximum capability. Sol could be built around something arguably more useful: frontier intelligence that you can actually use all day. If the reported speed difference holds up, OpenAI could have a model that p

SITUATION DETECTED: Periodic Labs used 1,300 H200s, plus months of its own experimental lab data, to train and RL an open-weight model that beats GPT-6 Astra on their internal benchmarks.

Haha... Literally no one is pacing the frontier - Opus 5.2 in testing - Grok 4.8 ships in a couple of weeks - Jev is a new ultra fast classifier - OpenAI already has Astra+ in testing We continue to accelerate

🚨 Muse Spark 2 Leaks: Beats Astra > Spark 2 could compete with GPT-6 Astra and Fable 5.1 > Meta is already developing the next-gen Muse model > Expected to be extremely cheap to run > A 1M-token context could carry over > Meta admitted Spark 1 struggled against the competition > Spark 2 could be Meta's answer > Expected late this month or next month Could Meta finally have a serious frontier model?

Astra turned my PNG into PLA 👨‍🔬 Send 3D models directly to your Bambu printer, in one chat.

GPT-5.5 will remain available via the OpenAI API Platform and in Codex sessions authenticated with an API key:

@ChatGPT

On October 14, it's time to say farewell to GPT-5.5 in ChatGPT, ChatGPT Work, and Codex across all plans. > If you use GPT-5.5 in Codex, switch to GPT-5.6 Sol or GPT-6 Astra. > Thanks for everything, 5.5 🫡

That’s insane GPT 5.5 was so good!

@ChatGPT

On October 14, it's time to say farewell to GPT-5.5 in ChatGPT, ChatGPT Work, and Codex across all plans. > If you use GPT-5.5 in Codex, switch to GPT-5.6 Sol or GPT-6 Astra. > Thanks for everything, 5.5 🫡

OpenAI’s 2026 pace so far: Apr – GPT-5.5 Jun/Jul – GPT-5.6 Sol, Terra, Luna Aug – 5.6-Cyber + Sol Ultrafast Sep 3 – GPT-6 Astra The year went from 5.5 → a three-tier 5.6 family → Astra in under five months.

truncated at source

DAILY AI BRIEF 🗞 — Sept 16 GOOGLE 🔥: - Gemini 3.8 Live and 3.8 Live Extended Thinking are out. 97-language auto-detect, near real-time vision, background tool calling. - Live is in Search Live plus Gemini API public preview. Extended Thinking is in Gemini Live, with Pro/Ultra getting it in Docs, Gmail, and Keep. - Gemini Notebook Voice Mode hits Ultra this week, Pro soon. Mobile voice recorder starts next week for all users, English first. - Interactive Reports for Gemini Notebook roll out to everyone in the coming weeks, plus new quiz formats and 60-second video overviews. OPENAI 🔥: - Sam declared a big ship week, then a much larger wave for DevDay. GPT-6 Sol and Luna are the expected drops. - GPT-5.5 leaves ChatGPT, Work, and Codex on Oct 14. Switch to GPT-5.6 Sol or GPT-6 Astra; the API keeps 5.5. XAI 🔥: - Grok Imagine can now edit text on any image in beta — color, size, font, alignment. - Grok Build 1.0.33: structured MCP JSON, in-UI memory deletes, and long-session checkpoints that survive cleanup.

15 September 2026 14 items

OpenAI already shipped Astra, GPT-Live-1 and the Agents API in a matter of days. Now @thsottiaux is teasing another packed week. The interesting part isn’t one launch anymore. It’s the release cadence. OpenAI seems to be shipping like DevDay is every week.

@thsottiaux

This week will also be a level of ships that you could have expected for DevDay 2025. Crazy

This looks so good wtf When i tried other models, roblox development was a lot harder for them than threejs and web, and it always looked worse even with Astra. They really cooked here, and usage is gonna feel unlimited on x20 plan compared to Fable/Astra

@synthwavedd

The new Opus is super impressive at Roblox game development (which should generalise into 3D tasks, but I'm yet to test more extensively) > It created this ~complete fun little game with about 20% of my weekly and 15M tokens (Ultracode). Better than Astra on this. You can try it below! >

Can confirm. Matt's use of AI has changed a lot of my perspective with what we can use it for. Insane!

@mattshumer_

With Astra and Fable 5.1, we officially have drop-in remote workers. > Still takes quite a bit of setup, but once you have it dialed in, it's near perfect. > I haven't touched my computer in a few days, and I'm more productive than ever.

A tribute to Le Petit Illustré, a forgotten French weekly, bringing its 1934 pages back to life. Made with @fal H3 Max Camera Controls, @threejs & GPT 6 Astra. I would have loved to travel through comics when I was kid! Comic book publishers: I’d love to explore this with your stories, much more to do!

this is AGI

@Angaisb_

An SVG of a screenshot of X by GPT-6 Astra > I've seen image models do worse, this is so cool

A phoenix rising from the ashes by GPT-6 Astra Pro on (this is a Minecraft build by the way, not a 3d model!)

OpenAI is retiring the middle child terra. not that surprising tbh. most people gravitated toward either sol or luna anyway, with terra being stuck in a weird middle ground. and now that astra sits above sol, the lineup is much cleaner: > Astra -> top end tier > Sol -> strong general frontier tier > Luna -> fast, cheap three clear tiers makes way more sense than keeping terra around to fill the gap nobody cared about. RIP little terra though.

@synthwavedd

I hear the GPT family are having a reunion to celebrate a couple of 6th birthdays soon and you're all invited. Though I do hope none of you grew too attached to little Terra. Tragic

UCSD 助理教授黄碧薇创办的因果世界模型公司 Aether AI 开源 RSIAgent。 这是一套不训练模型的递归自我改进框架。底层模型参数全程固定,Agent 会自己寻找值得练习的任务、实际操作、检查结果,再把成功方法和失败教训写进长期 Memory。下一轮继续利用这些经验,逐步补上能力短板。 系统由三个 Agent 配合。Curriculum Agent 决定接下来练什么,Actor Agent 真正操作软件,Verifier Agent 独立检查结果。探索分成两步:先广泛尝试不同任务,再针对失败、隐藏限制和边界情况继续深挖。最后 Memory 会被冻结,直接拿去执行正式任务。 在 OSWorld 2.0 上,加入这套 RSI 后,平均部分得分从 71.97% 提升到 78.98%;Agents’ Last Exam 从 83.75% 提升到 84.82%。不过这不是整套测试的严格 A/B 对比。OSWorld 只有一半任务实际用了 RSI 后的新结果,其余任务继续沿用原成绩。 它和 Prime Agent 这类 Harness 自我改进思路属于同一个大方向:模型权重不变,持续更新模型外面的东西。 RSIAgent 的特点是把改进重点放在 Memory,再用自主出题和独立验证,让这份外部经验库不断积累。

@huang_biwei

Can an agent explore a new environment, learn its causal structure, and keep improving without updating its model weights? > We introduce RSIAgent, a framework for recursive self-improvement through autonomous exploration. Using Kimi-K3 and GLM-5.3 as base models, RSIAgent outperforms GPT-6 Astra on both OSWorld 2.0 and Agents’ Last Exam. > RSIAgent decides what to explore, executes tasks,

🚨 GROK 5 LEAKS: Elon Just Admitted Something He Never Said Before >Training reportedly continues after Grok 4.8, landing in October >Grok 5 rumored at ~6T parameters >Elon: "I now think xAI has a chance of reaching AGI with Grok 5 never thought that before" Puts the odds at 10%, "and rising" >Expected to challenge GPT-6 Astra and Fable 5.1 Reportedly cheaper to run than rivals If Elon's own confidence is climbing, how close is Grok 5 actually getting

We’re building robot motion design army at Higgsfield, powered by GPT-6 Astra. A swarm of robot interns working under our human designers, giving each designer the firepower of an entire studio. Just Higgsfield plugin in ChatGPT + After Effects. Our human motion designers can now do 3x the creative output outsourcing manual work to "interns". AGI is here. and it reports to motion designers.

@higgsfield

ChatGPT can now do motion design in After Effects. > Introducing Higgsfield AI Motion Designer. > Our ChatGPT plugin understands animation principles, writes expressions, and retains context in your After Effects projects. > Try Higgsfield’s ChatGPT plugin now in After Effects.

Just shipped our video editor MCP for @floraai today - Amazed what Astra can do with it

Drop one picture into an AI agent and come back to a full 3D street. Shops, signs, road, sky, all built without anyone opening a 3D tool. GPT-6 Astra and Hyper3D MCP just did in one run what used to take a whole team.

@DeemosTech

One image. One Agent. One scene. > @OpenAIDevs GPT-6 Astra (Codex) + HYPER3D MCP just showed us the second half of 3D gen. 🚀

太魔幻了,Astra 还在让大家惊叹:AI 终于会用 Blender 了。 结果Nex-AGI 已经直接出现在现实世界里了。 我只给了一个 Prompt: 做一根 6cm 的粉色活动香蕉,省点料。 然后我基本没碰 Blender。 它自己用 MCP 建模,Computer Use 看结果、自己修,最后直接吐给我一个 STL。 我顺手扔进拓竹。 👇 视频里正在一层一层长出来的,就是它自己设计的香蕉。 这一刻我是真有点惊到了。 Nex-AGI + Computer Use,能力完全超出我预期。 一句话 → Blender → STL → 3D 打印 → 实物。 🔥 这真的有点科幻了。 直接免费体验nex-agi :

Astra and Fable 5.1 where so expensive models it always drained my usage limits, thankfully this month we're now getting a well balanced models GPT-5.6 Sol and Opus 5.2 which won't be as costly yet remain very powerful

14 September 2026 8 items

V4.1 paints its take on "Sunrise by the ocean" (Vladimir Kush) with SDFs in Rust, no reference image, just a description of the piece from the previous context. It has some sense of beauty I'd say. Astra is mildly impressed.

@teortaxesTex

Painting in Rust is hard. It got the broad strokes correct within ≈5 minutes but then just kept rabbitholing. Work to be done on agent cooperation, too. Still, I'm impressed. V4.1 can reproduce gists of paintings in arbitrary language.

Elon just gave a huge update on the Grok model roadmap and Grok 4.8 is the major step forward • Grok 4.7 — roughly on par with Opus 5.0, better in some areas and worse in others....multimodal performance still needs work • Grok 4.8 — a massive 2.5T model trained on SpaceXAI’s new C++ software stack....training finishes this week and then moves straight into reinforcement learning • Grok 4.9 — Astra/Fable class • Grok 5 — potentially better than anything The crazy part is 4.7 hasn’t even launched yet and 4.8 is already finishing training Then 4.9 and Grok 5 are lined up right behind it Elon is moving the entire Grok roadmap insanely fast

@elonmusk

@itslueul Grok 4.7 should be roughly on par with Opus 5.0, not 5.1. Better in some ways, worse in others. We need to fix multimodal performance. > Grok 4.8 will be a noticeable improvement. > Grok 4.9 is probably Astra/Fable class. > Grok 5 maybe better than anything. We shall see.

truncated at source

Amazing! A full 3D anatomy atlas has now been developed with GPT-6 Astra and WebMCP support! Where was this when we were in medical school? 😅

@ZentrixHQ

GPT-6 Astra has solved the problem of 3D anatomy tools and created an Anatomical Atlas with WebMCP support > 3D anatomy tools have had a control problem for years, and almost nobody built a real fix. > Every viewer forces the same workflow: click, drag, zoom, repeat, just to reach the angle a lecture or a case actually needs. A new WebMCP-enabled Anatomy Atlas skips that entirely. > The interface still supports manual exploration - full 3D anatomy, cross-section scrolling, tissue retraction by hand. But WebMCP exposes those same actions to an AI agent, so setting up a specific view, isolating a structure, or producing a short walkthrough video becomes a single instruction instead of a sequence of clicks. > It is not a UI overhaul. It is the same tool, plus a layer that lets software operate it the way a per

🚨 We're about to get some crazy Grok releases soon Grok 4.7 hopefully this week as Cursor employees are dropping hints and vagueposting Grok 4.8 will be a noticeable improvement according to Elon, my guess is early October unless they finish it sooner Grok 4.9 will apparently be Astra/Fable level 💀 I genuinely can't wait for Grok models to become more competitive, we're getting close

@LuminaBench

Crazy, so we're probably going to get Grok 4.7 this week now > Also Grok 4.8 (2.5T) finishes training this week and will start RL, this is going to be on their new stack too so will be incredibly efficient as well as fast > I wouldn't be surprised if Grok 4.8 drops at the end of September or early October as a guess

I asked GPT-6 to help me stay on top of my emails, bills, subscriptions, schedule, things like that. It sent me a checklist of about 15 different things I had to approve to grant it access. I instantly clicked approve all and hit yes. Astra, take the wheel!

🚨 BREAKING: Clone any successful company GPT-6 Astra in @shipper_now can take any app and make it yours: design, code, business plan... you can now one-shot the next duolingo / twitter / airbnb / etc this is THE END of vibe coding.

GPT-6 Sol just drop in the API : and it might be the model people actually use. - Up to 6x faster than Astra - Similar raw intelligence - Cheaper + better usage limits - Built for daily work, not just max benchmarks - Astra may stay the flagship, Sol could become the default. - Still unofficial pricing, or model ID yet.

@0x0SojalSec

GPT-6 Sol might be the real everyday model : > - OpenAI reportedly treats Astra as the foundation, then post-trains Sol on top - Early talk: better daily quality, lower cost, Plus-friendly limits - Low-effort Astra already dunks on max-effort GPT-5.6 Sol - Astra on “low” already beats GPT-5.6 Sol on “max” - That gap is the whole story: this generation is not incremental If it ships, Sol is the one most people will actually live in

AIエージェント「Devin」を開発するCognitionが、AIモデル「GPT-6 Astra」や「Claude Fable」の性能を最大限に引き出すハーネス(AIモデルの実行環境・仕組み)「Fusion」を開発した。

13 September 2026 13 items

DeepSeek Code 2.0 Leak: Coming Soon 🔥 >expected to beat Mythos 5.1 and GPT-6 Astra >September release window reportedly targeted >Computer use expected to be a top priority >Expected to remain open-weight >1M-token context window reportedly carried over >Rumored 3T+ parameter model

Fable 5 vs SWE 2.0 Le résultat de Fable 5 sur la phase de récolte est excellent. Franchement, c’est celui qui est le plus fidèle au jeu original, 60 Seconds. Cependant, SWE 2.0 a également produit un résultat assez intéressant sur cette partie. Il a réellement essayé de reproduire le bunker, et surtout, aucune erreur de code. J’ai aussi vu SWE 2.0 produire des résultats assez surprenants face à Astra sur de la génération 3D. Je partagerai d’autres tests prochainement pour voir s’il peut réellement devenir un concurrent sérieux... et ça semble plutôt bien parti !

🚨 GEMINI 4 PRO MAY BE CLOSER THAN WE THINK 👀🔥 Reports claim Google is already testing an internal Gemini 4 Pro checkpoint, with early testers calling it “amazing.” Rumors point to: • Major reasoning + coding gains • Huge context window • Stronger autonomous coding • Possible October launch • Potentially challenging GPT-6 Astra and Claude Fable 5.1 But one big caveat: The viral benchmark chart is marked “PREDICTED” — not verified test results. If these early reports are real, Google could be preparing a serious new frontier contender. 🔥

truncated at source

Sakana just dropped Fugu Max and Fugu Ultra v2 - a multi-agent system that outperforms Opus 5 and Fable 5 on benchmarks without using either of them. > Fugu Max routes tasks to the leanest capable model — frontier performance at 2-6x lower cost > Fugu Ultra v2 beats models that cost 3-5x more per token > No Fable 5, no GPT-6 Astra in the pool. Just open-weight orchestration doing what closed ecosystems can't. The best model isn't always the right answer. The best routing is.

@SakanaAILabs

Introducing Fugu Max and Fugu Ultra v2: the next evolution of Sakana Fugu’s multi-agent orchestration system. > Try: Blog: > The frontier that actually matters is the Pareto frontier: capability on one axis, cost on the other. But the industry still treats it as a static menu of isolated models. Today we are resolving that with a dynamic architecture: > Fugu M

GPT-6 Astra rebuilt this level of detail as vectors in Adobe Illustrator. In minutes. With Astra + Higgsfield, these references became layered Illustrator files with thousands of editable paths, preserving the fine detail. Around 7–13 minutes of vector assembly per illustration after prep.

"we should slow down AI development as it's getting too dangerous for humanity" Elon: Introducing our most powerful model Grok 4.7 Sam: Introducing GPT-6 Astra, Sol, Terra, Luna and we also have an internal model significantly more powerful than Astra Dario: Here's Fable 5.1 and we also have an internal model "model 2" significantly more powerful than Mythos (without fallbacks)

Astra出たからキャンペーン打ち切らない方がいいんじゃないかな

@oikon48

【注意】 > Claude Code の週次リミットは、本日まで50%増加キャンペーン中です。 > 9月14日からは標準から25%増加がデフォルトとなります。つまり、 > 1.25/1.5 = 0.833 > 現在の週次リミットの17%減 (83%)になる予定です。つらい。

Astra rebuilt an entire mobile game from an ad Gates → squad upgrades Walls → shoot through Zombies → HP bars Boss → survive

🚨 Second GPT-6 Sol demo is making the rounds And honestly… this output is wild 💀 The coding and frontend work looks seriously impressive. If the pricing reports turn out to be accurate, Sol could deliver this level of performance at a much lower cost than Astra

@Mr_Salio

🚨 GPT-6 Sol Second Output Got LEAKED > The output looks absolutely insane This kind of quality is terrifying It beats GPT-6 Astra in coding and frontend generation and it will be a lot cheaper > OpenAI have left Anthropic FAR behind

GPT-6 Astra upgraded the fly. It’s doing motion design now.

truncated at source

非营利 AI 评测机构 ARC Prize 预告 ARC-AGI-4。下一代 benchmark 将重点测试「自主开放式创新」,看 AI 能不能自己探索、提出新想法,甚至做出新的发明和发现。 具体题目、评分方式和发布时间都还没公布。ARC Prize 只说,人类目前在开放式创新上仍明显领先 AI,这也是推动科学和技术进步最关键的能力之一。 这条预告还引用了 Dario Amodei 刚发布的 AI 减速文章。ARC Prize 强调,开源仍然是 AI 进步的基础,反对行业为了协调减速而减少开放、把前沿 AI 集中在少数机构手里。 ARC-AGI 一直被冠以全球最难 AGI 评测的称号。ARC-AGI-3 发布时,人类得分 100%,前沿 AI 只有 0.51%。最新 GPT-6 Astra 已经大幅追上,但在统一的 Standard harness 下也只有 62.7%;接入 OpenAI 自家的 Provider Adapter 后才冲到 99.9%。

@arcprize

ARC-AGI-4 will be a benchmark for autonomous open-ended innovation. It will continue our commitment to open-source, giving the research community a shared target for progress that benefits all of humanity. > Despite rapid model progress, humans still significantly outperform AI at open-ended invention. This is the meta-skill that unlocks progress across every field of technology. > Advanced AI capable of scientific innovation will lead to tremendous new technology, knowledge, and understanding. This is a positive-sum future. We

Here’s a list of all AI models launched this month already: * Claude Fable 5.1 * Claude Mythos 5.1 * Gemini 3.8 Flash * Gemini 3.8 Flash Cyber * Muse Spark 1.3 * Qwen3.8-Max-0902 * GPT-6 Astra * GPT-6 Astra Pro * Ling-3.0-flash-VL * Ling-3.0-flash-Sante * Lyria 3.5 * ChatGPT Images 2.5 * GPT-Image-2.5 Flare * GPT-Image-2.5 Sunburst * DeepSeek-V4.1-Flash * Fugu Max * Fugu Ultra v2 We’re not even half way through the month…

艹,居然忘记放白嫖WorkBuddy+ DeepSeek v4.1 Flash的链接👇

@servasyy_ai

我的天,WorkBuddy 海外版 居然可以直接用 GPT6 Astra !? > 除了HY3还在继续免费使用外, 最新上线的 DeepSeek V4.1-Flash 居然可以限时免费~ DeepSeek V4.1-Flash 在 WorkBuddy国内版 可是要钱的,太香了! > 兄弟们,赶紧去冲吧 👇 >

12 September 2026 4 items

Évolution du SVG de Windows 11 avec GPT-6 Astra (Low). La quantité de détails est folle, et le résultat est beaucoup plus fidèle à Windows 11 qu’avec GPT-5.6 Pro, alors que j’ai testé Astra uniquement en Low ! Aujourd’hui, je vais partager plusieurs comparatifs entre mes anciens tests et les modèles d’aujourd’hui 👀

@mirochill

GPT-5.6 Pro a généré ce SVG de Windows 11 🔥 > Franchement, je le trouve meilleur que Mythos sur ce prompt. > Le problème, il ajoute des éléments inutiles (pop-ups, beaucoup de textes, etc.).

GPT 6 Sol and Opus 5.1 are both coming within the next 2 weeks. These are the releases that actually matter. Fable 5.1 and GPT 6 Astra are incredible models that nobody can afford to run consistently. Session limits die in minutes. Weekly limits die in days. We do not need smarter models right now. We need intelligent models that are affordable enough to actually use in our workflows. GPT 6 Sol and Opus 5.1 are supposed to be exactly that. The biggest week in AI is not about the frontier. It is about making the frontier usable.

Reset should be out for everyone! Enjoy! ✨ PSA: update to the latest version of Codex App/ CLI

@reach_vb

Reset rolling out to all Codex & ChatGPT Work users! > Grateful to everyone who helped us investigate the Astra quality issues and shared examples. > We’ve identified and fixed issues with: > Skills over-triggering or preventing self-checks > Context management causing early stops or stale replies > Misconfigured engines degrading quality > Thanks to everyone who took the time to help 🤗

truncated at source

DAILY AI BRIEF 🗞 — Sept 12 OPENAI 🔥: - GPT-Rosalind is out of research preview for eligible orgs worldwide. API, Codex, and ChatGPT Enterprise, with new Rosalind models as they ship. - Codex adds Life Sciences plugins for genomes, protein structure, QC reports, and notebooks. - ChatGPT Sites hit 5M apps. New: collab editing, private invites, custom domains, and DB inspect. - Desktop pets can start a new chat. Mini is the compact no-pet option. XAI 🔥: - Elon: Grok 4.7 needs a few more days. RL still quits hard tasks too early and undershoots self-checks. - Grok Bot is rolling out on Grok web for Heavy users. Create and chat with bots in the UI. No official post yet. ANTHROPIC 🔥: - Claude Code ships `claude plugin eval`. Score a plugin on test cases, then rerun without it. MICROSOFT 🔥: - MAI-Transcribe-2 hit 1M OpenRouter requests in 5 days. ALIBABA 🔥: - Qwen3.8-27B is live on Cerebras. * Used Grok to compose this brief, cherry-picking the news and doing some post-editing. > [@testingcatalog](