Latest news
Announcements, benchmark releases, and the work behind both.
Research & Engineering•2 min read
FrontierSWE v2
34 tasks, 13 frontier models, 20-hour budgets. Our hardest coding benchmark yet.
| # | Model | Score |
|---|---|---|
| 1 | Claude Fable 5.1 | 56.3% |
| 2 | GPT-5.6 | 32.2% |
| 3 | GLM-5.3 | 30.2% |
| 4 | Kimi K3 | 25.9% |
| 5 | Grok 4.6 | 25.3% |
| 6 | Gemini 3.7 Flash | 20.3% |
| 7 | Qwen3.8-Max | 15.8% |
| 8 | DeepSeek V4 Flash Exp | 14.8% |
| 9 | Muse Spark 1.2 | 12.0% |
| 10 | Inkling | 4.1% |
Research & Engineering
Post-training infrastructure at Proximal
How we run agentic RL across Modal GPUs and our own Kubernetes sandbox fleet.
Read more →Research & Engineering5 min read
Robustly Evaluating Cyber Capabilities
Our investigation into the CyberGym benchmark and reliable cyber evaluations.
Read more →Company2 min read
New Frontiers
Announcing our Seed Fundraise and Revenue Milestones
Read more →Research & Engineering9 min read
FrontierSWE
Our ultra-long-horizon coding benchmark. Frontier models clear only a fraction of its tasks.
Read more →Research & Engineering7 min read
Our Problems
An overview of the problems we’re working on at Proximal.
Read more →Company4 min read
Announcing Proximal
We believe data is becoming one of the central research problems in AI, and no one is working on it the right way. Proximal is a research lab for data.
Read more →Post-training infrastructure at Proximal
Research & Engineering
Robustly Evaluating Cyber Capabilities
Research & Engineering•5 min read
New Frontiers
Company•2 min read
FrontierSWE v2
Research & Engineering•2 min read
FrontierSWE
Research & Engineering•9 min read
Our Problems
Research & Engineering•7 min read
Announcing Proximal
Company•4 min read