Benchmarks

General intelligence

Vals Index

GDP-weighted agent performance across finance, coding and legal work

What it measures

Vals Index is Vals AI’s composite measure of agent performance on real-world work, currently at version 2. It covers finance, coding and legal tasks: financial research and Excel modeling; terminal work, app building and code migration; and legal research and long-horizon legal work products.

Think of it as a report card with differently weighted subjects: average each sector’s benchmark scores first, then combine the three sector averages. Finance has the greatest weight, so models with similar coding results can have different overall scores because of financial research and modeling.

Finance comprises Finance Agent v2 and the Excel Modeling Benchmark; coding comprises Terminal-Bench 2.1, Vibe Code Bench and Code Migration; legal comprises Legal Research Bench and Harvey’s Legal Agent Benchmark. The weights use the sectors’ approximate shares of U.S.

GDP: finance and insurance 8.0, information 5.6, and legal services 1.2, totaling 14.8; normalized weights are approximately 54.1%, 37.8% and 8.1%.

Reading the scores

Higher scores mean stronger overall performance under this specific task mix and weighting. The formula is (8 × Finance + 5.6 × Coding + 1.2 × Legal) ÷ 14.8. For example, sector scores of 60, 80 and 40 yield approximately 65.9%. This is neither a simple success rate across all tasks nor a percentage of jobs automated or GDP generated. This snapshot includes all 57 model results from the v2 leaderboard marked updated September 11, 2026, retaining published reasoning settings. Components and weights change between versions, so their scores should not be compared directly. Code Migration uses a fixed subset of 50 CLI migration tasks plus all 10 COBOL tasks, weighted 75% / 25%; its full standalone leaderboard score cannot simply be substituted. We retain the published composite scores instead of recomputing them from other benchmarks on this site, and do not present source standard errors as 95% confidence intervals.

SourceVals AI

Explore the finance component: Finance Agent v2

Model performance

Checked 57 results

Vals Index — Model performance. Select a column heading to sort.
Claude Fable 5.1 (max)Anthropic
68.8%1
Claude Opus 5 (max)Anthropic
67.2%2
GPT-6 Astra (max)OpenAI
66.6%3
Claude Fable 5 (max)Anthropic
66.0%4
Muse Spark 1.3 Max (max)Meta
64.5%
GPT-5.6 Sol (max)OpenAI
63.7%
Gemini 3.8 Flash (high)Google
62.3%
Claude Opus 4.8 (max)Anthropic
60.9%
Muse Spark 1.3 (xhigh)Meta
60.3%
GPT-5.6 Luna (max)OpenAI
59.9%5
Claude Sonnet 5 (max)Anthropic
59.6%
GPT-5.6 Terra (max)OpenAI
59.6%6
Gemini 3.7 Flash (high)Google
59.3%
Grok 4.6 (high)xAI
59.2%
DeepSeek V4.1 Flash (high)DeepSeek
57.9%
Kimi K3Moonshot AI
57.8%
GPT-5.5 (xhigh)OpenAI
57.4%
Muse Spark 1.2 (xhigh)Meta
57.1%
GLM-5.3 (max)Z AI
57.0%
Claude Opus 4.7 (high)Anthropic
56.1%
Gemini 3.6 Flash (high)Google
55.4%
Muse Spark 1.1 (xhigh)Meta
54.8%
DeepSeek V4 Flash 0731 (high)DeepSeek
53.6%
GLM-5.2Z AI
53.1%
Gemini 3.5 Flash (high)Google
53.1%
DeepSeek V4 Pro 0813 (max)DeepSeek
52.4%
Qwen 3.8 MaxAlibaba
51.8%
Grok 4.5 (high)xAI
51.5%
Claude Sonnet 4.6 (max)Anthropic
50.6%
Qwen 3.8 27B (xhigh)Alibaba
48.5%
GLM-5.3 Flash (max)Z AI
47.2%
Qwen 3.7 MaxAlibaba
44.8%
Kimi K2.6Moonshot AI
43.5%
DeepSeek V4 Pro (max)DeepSeek
42.9%
MiniMax-M3MiniMax
42.7%
Gemini 3.1 Pro Preview (02/26) (high)Google
41.9%
MiMo V2.5 ProXiaomi
41.0%
MiMo V2.5Xiaomi
39.9%
GPT 5.4 Mini (xhigh)OpenAI
39.6%
Qwen 3.7 PlusAlibaba
38.6%
Gemini 3.5 Flash Lite (high)Google
36.7%
Inkling (0.99)Thinking Machines
34.1%
GPT 5.4 Nano (high)OpenAI
33.0%
Qwen 3.6 PlusAlibaba
32.0%
Inkling Small (0.99)Thinking Machines
31.9%
Gemini 3 Flash (12/25) (high)Google
29.4%
Nemotron 3 UltraNVIDIA
27.4%
Kimi K2.5Moonshot AI
26.3%
MiniMax-M2.7MiniMax
24.6%
Grok 4.3 (high)xAI
24.3%
Claude Haiku 4.5 (Thinking)Anthropic
22.9%
Ling 3.0 FlashAnt Group
21.7%
Mistral Medium 3.5 (high)Mistral
17.9%
Grok 4.20 (Reasoning)xAI
17.6%
Gemini 3.1 Flash Lite Preview (high)Google
15.5%
Mercury 2.5 (high)Inception
13.0%
Nemotron 3.5 LightningNVIDIA
11.5%
  1. The default published score includes fallback-assisted results. Vals AI flags 35 fallbacks across 1,257 task results (2.78%); this is not a score obtained entirely by one model without fallback assistance.source
  2. The default published score includes fallback-assisted results. Vals AI flags 9 fallbacks across 2,157 task results (0.42%); this is not a score obtained entirely by one model without fallback assistance.source
  3. Vals AI reports 31 provider refusals across 1,257 task results (2.47%), scored as failures.source
  4. The default published score includes fallback-assisted results. Vals AI flags 84 fallbacks across 2,156 task results (3.90%); this is not a score obtained entirely by one model without fallback assistance.source
  5. Vals AI reports 1 provider refusals across 2,157 task results (0.05%), scored as failures.source
  6. Vals AI reports 31 provider refusals across 1,257 task results (2.47%), scored as failures.source

⌘KSearch the knowledge base

RECENTLY ADDED

Loading the knowledge index…