The Big Story
OpenAI just seized the code generation throne. GPT-6 Astra Max has overtaken Claude Fable 5.1 Max to claim the #1 spot in code generation with a massive ELO of 1797 — blowing past every model on the leaderboard by a wide margin. Meanwhile, xAI quietly slipped a new video model into the text-to-video top 10.
What Changed Today
- 🏆 Code Generation — New #1: GPT-6 Astra Max overtook Claude Fable 5.1 Max to become the top-ranked code generation model
- 🆕 Code Generation — New Entry: GPT-6 Astra Max entered the top 10 with an ELO of 1797
- 🆕 Text-to-Video — New Entry: Grok Imagine Video 1.5 Agent entered the top 10 with an ELO of 1491
Code Generation: GPT-6 Astra Max Takes the Crown
GPT-6 Astra Max from OpenAI has rocketed to #1 in code generation with an ELO of 1797 — a staggering 236 points above the previous reference leader (Claude Opus 4.6 at 1561) documented in our last full category update. This isn’t an incremental improvement; it’s a generational leap. The “Astra” line appears to represent OpenAI’s most capable reasoning-heavy architecture yet, and the code generation benchmarks reflect that with a score that dwarfs everything else on the board.
For developers building coding agents, autocomplete pipelines, or code review tools, this is the new model to beat. Use the API model ID gpt-6-astra-max via OpenAI’s API. Pricing hasn’t been confirmed in our reference data yet, so watch OpenAI’s pricing page closely — given the “max” tier designation, expect this to sit at the premium end of their lineup. If you’ve been locked into Anthropic’s Claude stack for code tasks, now is the time to run head-to-head evals on your own workloads.
FREE GUIDE
Stop Writing Design Specs by Hand
Get the free visual guide: how AI tools generate GAMP 5 documentation directly from your PLC and DCS exports. Used by Life Sciences engineers who are done doing it manually.
No spam. Unsubscribe anytime.
Text-to-Video: xAI’s Grok Imagine Video Gets an Agentic Upgrade
Grok Imagine Video 1.5 Agent from xAI has entered the text-to-video top 10 with an ELO of 1491 — and that score is notable because it significantly exceeds the current category leader, Veo 3.1 Audio 1080p (ELO 1392), as documented in our last reference update. The “Agent” suffix suggests this model is designed for programmatic, multi-step video generation workflows rather than one-shot prompting, which could make it particularly compelling for developers building automated content pipelines.
The existing xAI entry, Grok Imagine Video 720p, already sat at #6 with an ELO of 1357, so this 1.5 Agent variant represents a major quality jump from xAI. Developers in the text-to-video space should keep an eye on xAI’s API documentation for the model ID and pricing — given the “agent” framing, expect tool-use and multi-turn capabilities that could differentiate it from Google’s Veo and OpenAI’s Sora offerings.
Current Leaders at a Glance
| Category | #1 Model | Provider | Score (ELO) |
|---|---|---|---|
| Code Generation | GPT-6 Astra Max 🆕 | OpenAI | 1797 |
| Text-to-Video | Veo 3.1 Audio 1080p | 1392* |
*Note: Grok Imagine Video 1.5 Agent entered the top 10 at ELO 1491, which exceeds the documented leader score. The official leaderboard position may be updating — we’ll confirm the new #1 in tomorrow’s report once vote counts stabilize.
So What?
If you’re building anything that touches code generation — AI coding assistants, automated PR review, agentic dev tools — you need to benchmark gpt-6-astra-max against your current stack immediately. An ELO of 1797 isn’t a marginal win; it suggests a qualitative shift in what’s possible. For video builders, the emergence of Grok Imagine Video 1.5 Agent with its agent-oriented architecture signals that text-to-video is moving beyond “generate a clip” toward programmable, multi-step video workflows — exactly what production pipelines need. The practical move today: spin up eval runs against GPT-6 Astra Max for code, add xAI’s new video model to your shortlist for video generation, and don’t assume last month’s model choices are still optimal. The leaderboard just shifted hard, and your product should shift with it.

