Latest news

Announcements, benchmark releases, and the work behind both.

Research & Engineering•2 min read

FrontierSWE v2

34 tasks, 13 frontier models, 20-hour budgets. Our hardest coding benchmark yet.

Read more →
Leaderboard
#ModelScore
1
Claude Fable 5.1
56.3%
2
GPT-5.6
32.2%
3
GLM-5.3
30.2%
4
Kimi K3
25.9%
5
Grok 4.6
25.3%
6
Gemini 3.7 Flash
20.3%
7
Qwen3.8-Max
15.8%
8
DeepSeek V4 Flash Exp
14.8%
9
Muse Spark 1.2
12.0%
10
Inkling
4.1%