Top 10 LLM Models by DeepSWE Score
DeepSWE measures average pass@1 on real software engineering tasks - can the model actually fix the bug, not describe how it would. Current standings:
• GPT-5.6 Sol (max) - 69
• GPT-5.6 Sol (xHigh) - 67
• GPT-5.6 Terra (max) - 67
• Claude Fable 5 (max) - 66
• GPT-5.6 Sol (High) - 65
• GPT-5.5 (xHigh) - 64
• GPT-5.6 Sol (medium) - 64
Kimi K3 - 64
GPT-5.6 Luna (max) - 63
Claude Opus 5 (max) - 63
Seven of the top ten are OpenAI. But the more useful read is the spread: six points separate first place from tenth, and GPT-5.6 Sol at medium reasoning scores the same 64 as Kimi K3 and GPT-5.5 at xHigh.
Kimi K3 at 64 is the number worth watching. It's the only non-US model in the top ten and it lands inside the same band as models costing several times more.
Meanwhile the leader's 69 carries an asterisk - Artificial Analysis flagged a higher rate of content safety filtering on that endpoint since benchmarking.
Data source 🔗 Artificial Analysis
Follow @top7ico to be on top of Tech, AI & Crypto

August 27, 2026 844 15