Coding
FrontierSWE v2
Long-horizon engineering: difficult software, research and visual tasks within 20 hours
What it measures
FrontierSWE v2, developed by Proximal Labs, tests whether coding agents can sustain autonomous work on complex projects. Its 34 tasks cover system implementation, performance optimisation, scientific computing, visual reasoning and AI research, with five trials per task and up to 20 hours per trial.
Think of it as a day-long engineering challenge: the agent must build, test, diagnose problems and iterate, rather than merely explain how to write the code. Examples include reimplementing Git in Zig, building a PostgreSQL-compatible service on SQLite, recreating a video in Remotion, and training a racing bot that drives from camera images alone.
Models use the Proximus harness in Linux environments, sometimes with GPUs. The harness supports image viewing, context compaction and saved candidate submissions to encourage further improvements within the remaining time.
Automated verifiers grade the final work, so these are model-plus-harness results, not measurements of the same model in Codex or Claude Code.
Reading the scores
each task earns a score from 0 to 1; scores are averaged over five trials and aggregated across tasks as mean@5, displayed as a percentage. A score of 65% does not mean 65% of tasks were fully solved. The worst@5–best@5 range aggregates each task’s worst and best trial scores; it is not a 95% confidence interval. Cost in USD and time in hours are per-trial averages, not the cost of a full benchmark evaluation. This page includes all 14 models from the official leaderboard’s All view, consistently using its scores, costs and times. Its analysis charts contain slightly different aggregates, which are not mixed here. V2 adds 21 tasks, retires four older tasks and changes scoring, so results are not directly comparable with v1 rankings. Task content is public, but the standalone runner and public prebuilt images are still marked as forthcoming.
Model performance
GPT-6 AstraOpenAIAgent · Proximus | Proximus | 65.5%55.1–73.0%worst–best@5 | $1,029.65 | 12.1 h |
|---|---|---|---|---|
Claude Fable 5.1AnthropicAgent · Proximus | Proximus | 56.3%144.5–66.7%worst–best@5 | $138.55 | 11.6 h |
Claude Opus 5AnthropicAgent · Proximus | Proximus | 52.0%39.1–60.7%worst–best@5 | $196.71 | 16.1 h |
Claude Fable 5AnthropicAgent · Proximus | Proximus | 47.0%33.9–61.8%worst–best@5 | $301.28 | 15.8 h |
GPT-5.6OpenAIAgent · Proximus | Proximus | 32.2%22.2–44.2%worst–best@5 | $179.64 | 8.6 h |
GLM-5.3Z AIAgent · Proximus | Proximus | 30.2%17.7–40.6%worst–best@5 | $97.22 | 17.0 h |
Kimi K3Moonshot AIAgent · Proximus | Proximus | 25.9%13.9–37.6%worst–best@5 | $109.71 | 18.4 h |
Grok 4.6xAIAgent · Proximus | Proximus | 25.3%12.9–37.3%worst–best@5 | $243.43 | 13.8 h |
Gemini 3.7 FlashGoogleAgent · Proximus | Proximus | 20.3%10.3–31.2%worst–best@5 | $34.14 | 7.9 h |
Gemini 3.8 FlashGoogleAgent · Proximus | Proximus | 19.6%9.5–31.6%worst–best@5 | $38.75 | 7.2 h |
Qwen3.8-MaxAlibabaAgent · Proximus | Proximus | 15.8%8.7–24.3%worst–best@5 | $55.14 | 18.5 h |
DeepSeek V4 Flash Vision ExpDeepSeekAgent · Proximus | Proximus | 14.8%7.0–26.5%worst–best@5 | $8.57 | 14.8 h |
Muse Spark 1.2MetaAgent · Proximus | Proximus | 12.0%6.8–18.4%worst–best@5 | $27.81 | 4.6 h |
InklingThinking MachinesAgent · Proximus | Proximus | 4.1%0.1–10.8%worst–best@5 | $9.15 | 1.2 h |