Benchmarks

DeepSWE v1.1

Software engineering: completing long-horizon, multi-file tasks in real repositories

What it measures

DeepSWE is a software engineering benchmark from Datacurve that tests whether coding agents can understand an existing project, plan changes and deliver working software. Its 113 original tasks span 91 repositories and five languages: TypeScript, JavaScript, Python, Go and Rust.

Tasks include bug fixes, new features, API extensions and complex behavior changes, such as adding configuration-file support to a CLI or implementing XML diff and patch operations. Authors design tasks from scratch rather than adapting existing commits or pull requests, aiming to reduce exposure to reference solutions; this does not mean the underlying repositories were never in training data.

The official leaderboard uses mini-swe-agent throughout and grades solutions with behavioral tests. In v1.1, only the committed patch is evaluated in a separate container, and Git history after the starting commit is hidden to reduce opportunities to bypass evaluation.

It is useful for comparing complex engineering ability, while real-world results still depend on the agent scaffold, reasoning effort and task domain.

Reading the scores

Pass@1 is the task pass rate for a single attempt; higher is better. Repeated runs estimate variability, not success after multiple retries. This page includes the 21 models shown in the official v1.1 default Best view, retaining the best-scoring effort setting for each displayed model and naming that setting explicitly. Official values are converted to percentages and rounded to one decimal place. ± is the half-width of a 95% confidence interval estimated from variation across four whole-benchmark runs, in percentage points, rather than a task-sampling interval. Some configurations have fewer than 452 valid attempts, so use the official denominator convention and do not treat small gaps as stable leads. This snapshot follows the official September 3, 2026 update and does not mix v1 scores.

SourceDatacurve · DeepSWE

Model performance

Checked 21 results

DeepSWE v1.1 — Model performance. Select a column heading to sort.
GPT-6 Astra (xhigh)OpenAIAgent · mini-swe-agent
mini-swe-agent74.1%± 2.9
Gemini 3.8 Flash (high)GoogleAgent · mini-swe-agent
mini-swe-agent73.8%± 1.4
Claude Opus 5 (max)AnthropicAgent · mini-swe-agent
mini-swe-agent73.6%± 3.9
GPT-5.6 Sol (max)OpenAIAgent · mini-swe-agent
mini-swe-agent72.7%± 2.8
Claude Fable 5 (xhigh)AnthropicAgent · mini-swe-agent
mini-swe-agent69.9%1± 3.2
GLM-5.3 (max)Z AIAgent · mini-swe-agent
mini-swe-agent69.0%± 3
Kimi K3 (max)Moonshot AIAgent · mini-swe-agent
mini-swe-agent68.5%± 4.5
Grok 4.6 (medium)xAIAgent · mini-swe-agent
mini-swe-agent67.5%± 2.3
GPT-5.6 Luna (max)OpenAIAgent · mini-swe-agent
mini-swe-agent67.2%± 4
GPT-5.5 (xhigh)OpenAIAgent · mini-swe-agent
mini-swe-agent67.0%± 6.5
Gemini 3.7 Flash (medium)GoogleAgent · mini-swe-agent
mini-swe-agent65.5%± 3.1
GLM-5.3 Flash (max)Z AIAgent · mini-swe-agent
mini-swe-agent63.4%± 4.4
DeepSeek V4 Pro (max)DeepSeekAgent · mini-swe-agent
mini-swe-agent62.8%± 6.3
Claude Opus 4.8 (max)AnthropicAgent · mini-swe-agent
mini-swe-agent59.0%± 1.8
Qwen3.8 Max (xhigh)AlibabaAgent · mini-swe-agent
mini-swe-agent57.5%± 2.7
Muse Spark 1.2 (xhigh)MetaAgent · mini-swe-agent
mini-swe-agent54.9%± 2.1
Claude Sonnet 5 (max)AnthropicAgent · mini-swe-agent
mini-swe-agent53.8%± 4.2
DeepSeek V4 Flash (max)DeepSeekAgent · mini-swe-agent
mini-swe-agent53.3%± 3.6
Gemini 3.6 Flash (high)GoogleAgent · mini-swe-agent
mini-swe-agent46.7%± 3.7
GLM-5.2 (max)Z AIAgent · mini-swe-agent
mini-swe-agent43.8%± 1.7
Gemini 3.5 Flash (high)GoogleAgent · mini-swe-agent
mini-swe-agent36.1%± 4
  1. The official note reports 73 incomplete trials out of 2,260 across all Claude Fable 5 effort settings due to interrupted access; pass rates use completed trials. The xhigh setting shown here has 452 attempts. The all-effort missing-trial count does not apply to this row alone.source

⌘KSearch the knowledge base

RECENTLY ADDED

Loading the knowledge index…