The Devin agent benchmark gives the most honest read on GPT-6 Astra so far: it shows no real gap over Claude Fable 5. Roughly the same level of LLM, just cheaper.
That pattern is familiar. Kimi K3 and Qwen 3.8 Max already delivered near-top Claude results at lower prices, and Anthropic never cut prices in response — Dario's line was always "бенчи-бенчами, но у меня LLM для реальной разработки."
What differs is Altman made passing benchmarks his stated priority, which invites scrutiny. WebDev Arena is the moment of truth: if the claims are honest, GPT-6 Astra should beat Claude Fable 5.1 by at least 100 Elo — the margin Claude itself showed on genuinely breakthrough releases. A smaller gap or a loss would make the case for extreme benchmark overfitting.
The Devin result is one data point, and 100 Elo is an arbitrary threshold. The real test will be how GPT-6 Astra performs in actual development workflows, not just one arena.


















