Benchmarks

Terminal-Bench 4.0

Agentic ability: completing real tasks in a terminal, unattended

What it measures

The model is dropped into a container with nothing but a shell and a one-paragraph task—build a project, fix a bug, process a dataset, configure a service—and then left alone. It has to run commands, read the output and adjust until the job is done; a pre-written test script decides pass or fail.

The score is the share of tasks solved. It measures getting things done, not knowing things.

Results belong to a model + agent-scaffold pairing: the same model run under Claude Code, Codex or mini-SWE-agent scores differently, which is why the agent is shown next to each number.

Reading the scores

percentage of tasks solved with a 95% confidence interval; when the official leaderboard and a vendor’s own report differ, both are footnoted.

Sourcetbench.ai

Model performance

Checked 6 results

Terminal-Bench 4.0 — Model performance. Select a column heading to sort.
GPT-6 Astra (max)OpenAIAgent · CodexCodex58.2%± 2.8
Claude Fable 5.1 (max)AnthropicAgent · Claude CodeClaude Code57.9%1± 3.8
Claude Opus 5 (max)AnthropicAgent · Claude CodeClaude Code51.8%± 3.4
Claude Fable 5AnthropicAgent · Claude CodeClaude Code44.5%2± 3.8
GLM-5.3 (max)Z AIAgent · Claude CodeClaude Code41.8%± 3.2
GPT-5.6 Sol (max)OpenAIAgent · CodexCodex37.3%± 3.8
  1. Anthropic system card self-reports 55.8%.source
  2. Terminal-Bench run at max reasoning effort.

⌘KSearch the knowledge base

RECENTLY ADDED

Loading the knowledge index…