DeepSWE v1.1
Software engineering: completing long-horizon, multi-file tasks in real repositories
What it measures
DeepSWE is a software engineering benchmark from Datacurve that tests whether coding agents can understand an existing project, plan changes and deliver working software. Its 113 original tasks span 91 repositories and five languages: TypeScript, JavaScript, Python, Go and Rust.
Tasks include bug fixes, new features, API extensions and complex behavior changes, such as adding configuration-file support to a CLI or implementing XML diff and patch operations. Authors design tasks from scratch rather than adapting existing commits or pull requests, aiming to reduce exposure to reference solutions; this does not mean the underlying repositories were never in training data.
The official leaderboard uses mini-swe-agent throughout and grades solutions with behavioral tests. In v1.1, only the committed patch is evaluated in a separate container, and Git history after the starting commit is hidden to reduce opportunities to bypass evaluation.
It is useful for comparing complex engineering ability, while real-world results still depend on the agent scaffold, reasoning effort and task domain.
Reading the scores
Pass@1 is the task pass rate for a single attempt; higher is better. Repeated runs estimate variability, not success after multiple retries. This page includes the 21 models shown in the official v1.1 default Best view, retaining the best-scoring effort setting for each displayed model and naming that setting explicitly. Official values are converted to percentages and rounded to one decimal place. ± is the half-width of a 95% confidence interval estimated from variation across four whole-benchmark runs, in percentage points, rather than a task-sampling interval. Some configurations have fewer than 452 valid attempts, so use the official denominator convention and do not treat small gaps as stable leads. This snapshot follows the official September 3, 2026 update and does not mix v1 scores.
SourceDatacurve · DeepSWE
Model performance
GPT-6 Astra (xhigh)OpenAIAgent · mini-swe-agent | mini-swe-agent | 74.1%± 2.9 |
|---|---|---|
Gemini 3.8 Flash (high)GoogleAgent · mini-swe-agent | mini-swe-agent | 73.8%± 1.4 |
Claude Opus 5 (max)AnthropicAgent · mini-swe-agent | mini-swe-agent | 73.6%± 3.9 |
GPT-5.6 Sol (max)OpenAIAgent · mini-swe-agent | mini-swe-agent | 72.7%± 2.8 |
Claude Fable 5 (xhigh)AnthropicAgent · mini-swe-agent | mini-swe-agent | 69.9%1± 3.2 |
GLM-5.3 (max)Z AIAgent · mini-swe-agent | mini-swe-agent | 69.0%± 3 |
Kimi K3 (max)Moonshot AIAgent · mini-swe-agent | mini-swe-agent | 68.5%± 4.5 |
Grok 4.6 (medium)xAIAgent · mini-swe-agent | mini-swe-agent | 67.5%± 2.3 |
GPT-5.6 Luna (max)OpenAIAgent · mini-swe-agent | mini-swe-agent | 67.2%± 4 |
GPT-5.5 (xhigh)OpenAIAgent · mini-swe-agent | mini-swe-agent | 67.0%± 6.5 |
Gemini 3.7 Flash (medium)GoogleAgent · mini-swe-agent | mini-swe-agent | 65.5%± 3.1 |
GLM-5.3 Flash (max)Z AIAgent · mini-swe-agent | mini-swe-agent | 63.4%± 4.4 |
DeepSeek V4 Pro (max)DeepSeekAgent · mini-swe-agent | mini-swe-agent | 62.8%± 6.3 |
Claude Opus 4.8 (max)AnthropicAgent · mini-swe-agent | mini-swe-agent | 59.0%± 1.8 |
Qwen3.8 Max (xhigh)AlibabaAgent · mini-swe-agent | mini-swe-agent | 57.5%± 2.7 |
Muse Spark 1.2 (xhigh)MetaAgent · mini-swe-agent | mini-swe-agent | 54.9%± 2.1 |
Claude Sonnet 5 (max)AnthropicAgent · mini-swe-agent | mini-swe-agent | 53.8%± 4.2 |
DeepSeek V4 Flash (max)DeepSeekAgent · mini-swe-agent | mini-swe-agent | 53.3%± 3.6 |
Gemini 3.6 Flash (high)GoogleAgent · mini-swe-agent | mini-swe-agent | 46.7%± 3.7 |
GLM-5.2 (max)Z AIAgent · mini-swe-agent | mini-swe-agent | 43.8%± 1.7 |
Gemini 3.5 Flash (high)GoogleAgent · mini-swe-agent | mini-swe-agent | 36.1%± 4 |