← Back to blog

OpenAI Says GPT-6 Astra Just Entered the 'AGI Era.' Here's What Actually Changed.

⭐ Featured

OpenAI Says GPT-6 Astra Just Entered the 'AGI Era.' Here's What Actually Changed.

OpenAI shipped a new flagship model today called GPT-6 Astra, and the framing wasn't subtle. President Greg Brockman closed the announcement with "welcome to the AGI era." Sam Altman called it OpenAI's best model yet for computer use, professional work, science, coding, and cybersecurity. That's a lot of superlatives for one release β€” so let's look at what's actually in the box.

The Benchmarks, Briefly

Astra's numbers are genuinely strong. OpenAI is reporting 98% on FrontierMath Tier 4, 100% on ExploitBench, and a saturating 99.9% on ARC-AGI-3 β€” a benchmark specifically designed to resist memorization and reward genuine reasoning. It's also claiming SOTA on TerminalBench-4.0, plus leading scores on Terminal-Bench Science and HealthBench Pro.

The ARC-AGI-3 number is the one people are actually arguing about. FranΓ§ois Chollet, who built the benchmark, said Astra represents "a step-function change" for interactive reasoning β€” but he was careful to note the 99.9% figure used a custom continuous-dialogue harness with heavy compaction, at roughly $360 per game. With the standard evaluation harness, Astra scores 66%. Both numbers are real; they're just measuring different things. The benchmark was released six months ago, and Chollet says he expected it to take frontier models about a year to saturate. It took half that.

The First "Critical" Model

The detail that matters more than any leaderboard: OpenAI's own safety overview classifies Astra as the first model to reach "Critical" capability under its Preparedness Framework for cybersecurity β€” meaning it can discover unknown vulnerabilities and develop working exploits without step-by-step human guidance. That's not a hypothetical concern; it's why the rollout is staged. Astra is going first to vetted organizations in OpenAI's Daybreak cybersecurity program, and only over the following days to everyone else on Plus, Pro, Business, Enterprise, the API, and AWS.

OpenAI paired the release with a $1 billion commitment called Daybreak for Frontline Defenders β€” subsidized access, training, and support aimed at water utilities, power grids, local governments, community banks, and open-source maintainers who don't have enterprise security budgets. Reading between the lines: if you're shipping a model that can find zero-days on its own, you'd better also be arming the people who have to defend against them.

Not everyone is convinced the caution is enough. TechCrunch's coverage flagged what it called Astra's "opaque recurrence" β€” a term for internal reasoning behavior that's harder to inspect than prior models β€” as the source of the "controversial" label in its own headline.

Where This Leaves Claude

On pricing, Astra lands at $10 per million input tokens and $50 per million output β€” the same tier as Claude Fable 5 and 5.1. On Artificial Analysis's Coding Agent Index, Astra scored 67, putting it roughly even with Claude Opus 5 and Fable 5, at under half of Fable 5's cost and about 70% better token efficiency than OpenAI's own GPT-5.6 Sol. Several commentators framed today's release as directly aimed at Fable 5.1's two-day run at the top of the leaderboards. Whether that holds depends on real-world agentic use, not a single index β€” but the gap between frontier labs on coding-agent quality is clearly narrowing, not widening.

What This Means If You Use OpenClaw

The headline claim in every one of these releases is basically the same: the model is smarter, and increasingly, it can act on its own β€” write code, use a computer, chain tool calls, find its own way to a goal. That's exactly the territory OpenClaw already lives in.

The model race matters less to you as an OpenClaw user than it sounds like it should. OpenClaw isn't betting on a single underlying model β€” it's the agent layer on top: the memory, the tool connections, the task history that persists whether the model underneath is from OpenAI, Anthropic, or whoever ships the next "AGI era" headline. When a new frontier model lands, what actually improves for you is how well your agent reasons and acts β€” not what you have to rebuild to get there.

That's the practical upside of using an agent platform instead of chasing model releases yourself: the capability jumps keep coming, and you don't have to re-architect anything to benefit from them.

Start your free trial β†’