Benchmarks

Coding

FrontierSWE v2

Long-horizon engineering: difficult software, research and visual tasks within 20 hours

What it measures

FrontierSWE v2, developed by Proximal Labs, tests whether coding agents can sustain autonomous work on complex projects. Its 34 tasks cover system implementation, performance optimisation, scientific computing, visual reasoning and AI research, with five trials per task and up to 20 hours per trial.

Think of it as a day-long engineering challenge: the agent must build, test, diagnose problems and iterate, rather than merely explain how to write the code. Examples include reimplementing Git in Zig, building a PostgreSQL-compatible service on SQLite, recreating a video in Remotion, and training a racing bot that drives from camera images alone.

Models use the Proximus harness in Linux environments, sometimes with GPUs. The harness supports image viewing, context compaction and saved candidate submissions to encourage further improvements within the remaining time.

Automated verifiers grade the final work, so these are model-plus-harness results, not measurements of the same model in Codex or Claude Code.

Reading the scores

each task earns a score from 0 to 1; scores are averaged over five trials and aggregated across tasks as mean@5, displayed as a percentage. A score of 65% does not mean 65% of tasks were fully solved. The worst@5–best@5 range aggregates each task’s worst and best trial scores; it is not a 95% confidence interval. Cost in USD and time in hours are per-trial averages, not the cost of a full benchmark evaluation. This page includes all 14 models from the official leaderboard’s All view, consistently using its scores, costs and times. Its analysis charts contain slightly different aggregates, which are not mixed here. V2 adds 21 tasks, retires four older tasks and changes scoring, so results are not directly comparable with v1 rankings. Task content is public, but the standalone runner and public prebuilt images are still marked as forthcoming.

SourceProximal Labs · FrontierSWE

Official methodology and task examples

Model performance

Checked 14 results

FrontierSWE v2 — Model performance. Select a column heading to sort.
GPT-6 AstraOpenAIAgent · Proximus
Proximus65.5%55.1–73.0%worst–best@5$1,029.6512.1 h
Claude Fable 5.1AnthropicAgent · Proximus
Proximus56.3%144.5–66.7%worst–best@5$138.5511.6 h
Claude Opus 5AnthropicAgent · Proximus
Proximus52.0%39.1–60.7%worst–best@5$196.7116.1 h
Claude Fable 5AnthropicAgent · Proximus
Proximus47.0%33.9–61.8%worst–best@5$301.2815.8 h
GPT-5.6OpenAIAgent · Proximus
Proximus32.2%22.2–44.2%worst–best@5$179.648.6 h
GLM-5.3Z AIAgent · Proximus
Proximus30.2%17.7–40.6%worst–best@5$97.2217.0 h
Kimi K3Moonshot AIAgent · Proximus
Proximus25.9%13.9–37.6%worst–best@5$109.7118.4 h
Grok 4.6xAIAgent · Proximus
Proximus25.3%12.9–37.3%worst–best@5$243.4313.8 h
Gemini 3.7 FlashGoogleAgent · Proximus
Proximus20.3%10.3–31.2%worst–best@5$34.147.9 h
Gemini 3.8 FlashGoogleAgent · Proximus
Proximus19.6%9.5–31.6%worst–best@5$38.757.2 h
Qwen3.8-MaxAlibabaAgent · Proximus
Proximus15.8%8.7–24.3%worst–best@5$55.1418.5 h
DeepSeek V4 Flash Vision ExpDeepSeekAgent · Proximus
Proximus14.8%7.0–26.5%worst–best@5$8.5714.8 h
Muse Spark 1.2MetaAgent · Proximus
Proximus12.0%6.8–18.4%worst–best@5$27.814.6 h
InklingThinking MachinesAgent · Proximus
Proximus4.1%0.1–10.8%worst–best@5$9.151.2 h
  1. The official v2 report states that Fable 5.1 uses Opus 5 as a fallback on tasks blocked by content filters. This result should not be read as every task being completed exclusively by Fable 5.1.source

⌘KSearch the knowledge base

RECENTLY ADDED

Loading the knowledge index…