Terminal-Bench 4.0
Agentic ability: completing real tasks in a terminal, unattended
What it measures
The model is dropped into a container with nothing but a shell and a one-paragraph task—build a project, fix a bug, process a dataset, configure a service—and then left alone. It has to run commands, read the output and adjust until the job is done; a pre-written test script decides pass or fail.
The score is the share of tasks solved. It measures getting things done, not knowing things.
Results belong to a model + agent-scaffold pairing: the same model run under Claude Code, Codex or mini-SWE-agent scores differently, which is why the agent is shown next to each number.
Reading the scores
percentage of tasks solved with a 95% confidence interval; when the official leaderboard and a vendor’s own report differ, both are footnoted.
Sourcetbench.ai
Model performance
| GPT-6 Astra (max)OpenAIAgent · Codex | Codex | 58.2%± 2.8 |
|---|---|---|
| Claude Fable 5.1 (max)AnthropicAgent · Claude Code | Claude Code | 57.9%1± 3.8 |
| Claude Opus 5 (max)AnthropicAgent · Claude Code | Claude Code | 51.8%± 3.4 |
| Claude Fable 5AnthropicAgent · Claude Code | Claude Code | 44.5%2± 3.8 |
| GLM-5.3 (max)Z AIAgent · Claude Code | Claude Code | 41.8%± 3.2 |
| GPT-5.6 Sol (max)OpenAIAgent · Codex | Codex | 37.3%± 3.8 |