General intelligence
Vals Index
GDP-weighted agent performance across finance, coding and legal work
What it measures
Vals Index is Vals AI’s composite measure of agent performance on real-world work, currently at version 2. It covers finance, coding and legal tasks: financial research and Excel modeling; terminal work, app building and code migration; and legal research and long-horizon legal work products.
Think of it as a report card with differently weighted subjects: average each sector’s benchmark scores first, then combine the three sector averages. Finance has the greatest weight, so models with similar coding results can have different overall scores because of financial research and modeling.
Finance comprises Finance Agent v2 and the Excel Modeling Benchmark; coding comprises Terminal-Bench 2.1, Vibe Code Bench and Code Migration; legal comprises Legal Research Bench and Harvey’s Legal Agent Benchmark. The weights use the sectors’ approximate shares of U.S.
GDP: finance and insurance 8.0, information 5.6, and legal services 1.2, totaling 14.8; normalized weights are approximately 54.1%, 37.8% and 8.1%.
Reading the scores
Higher scores mean stronger overall performance under this specific task mix and weighting. The formula is (8 × Finance + 5.6 × Coding + 1.2 × Legal) ÷ 14.8. For example, sector scores of 60, 80 and 40 yield approximately 65.9%. This is neither a simple success rate across all tasks nor a percentage of jobs automated or GDP generated. This snapshot includes all 57 model results from the v2 leaderboard marked updated September 11, 2026, retaining published reasoning settings. Components and weights change between versions, so their scores should not be compared directly. Code Migration uses a fixed subset of 50 CLI migration tasks plus all 10 COBOL tasks, weighted 75% / 25%; its full standalone leaderboard score cannot simply be substituted. We retain the published composite scores instead of recomputing them from other benchmarks on this site, and do not present source standard errors as 95% confidence intervals.
SourceVals AI
Model performance
Claude Fable 5.1 (max)Anthropic | 68.8%1 |
|---|---|
Claude Opus 5 (max)Anthropic | 67.2%2 |
GPT-6 Astra (max)OpenAI | 66.6%3 |
Claude Fable 5 (max)Anthropic | 66.0%4 |
Muse Spark 1.3 Max (max)Meta | 64.5% |
GPT-5.6 Sol (max)OpenAI | 63.7% |
Gemini 3.8 Flash (high)Google | 62.3% |
Claude Opus 4.8 (max)Anthropic | 60.9% |
Muse Spark 1.3 (xhigh)Meta | 60.3% |
GPT-5.6 Luna (max)OpenAI | 59.9%5 |
Claude Sonnet 5 (max)Anthropic | 59.6% |
GPT-5.6 Terra (max)OpenAI | 59.6%6 |
Gemini 3.7 Flash (high)Google | 59.3% |
Grok 4.6 (high)xAI | 59.2% |
DeepSeek V4.1 Flash (high)DeepSeek | 57.9% |
Kimi K3Moonshot AI | 57.8% |
GPT-5.5 (xhigh)OpenAI | 57.4% |
Muse Spark 1.2 (xhigh)Meta | 57.1% |
GLM-5.3 (max)Z AI | 57.0% |
Claude Opus 4.7 (high)Anthropic | 56.1% |
Gemini 3.6 Flash (high)Google | 55.4% |
Muse Spark 1.1 (xhigh)Meta | 54.8% |
DeepSeek V4 Flash 0731 (high)DeepSeek | 53.6% |
GLM-5.2Z AI | 53.1% |
Gemini 3.5 Flash (high)Google | 53.1% |
DeepSeek V4 Pro 0813 (max)DeepSeek | 52.4% |
Qwen 3.8 MaxAlibaba | 51.8% |
Grok 4.5 (high)xAI | 51.5% |
Claude Sonnet 4.6 (max)Anthropic | 50.6% |
Qwen 3.8 27B (xhigh)Alibaba | 48.5% |
GLM-5.3 Flash (max)Z AI | 47.2% |
Qwen 3.7 MaxAlibaba | 44.8% |
Kimi K2.6Moonshot AI | 43.5% |
DeepSeek V4 Pro (max)DeepSeek | 42.9% |
MiniMax-M3MiniMax | 42.7% |
Gemini 3.1 Pro Preview (02/26) (high)Google | 41.9% |
MiMo V2.5 ProXiaomi | 41.0% |
MiMo V2.5Xiaomi | 39.9% |
GPT 5.4 Mini (xhigh)OpenAI | 39.6% |
Qwen 3.7 PlusAlibaba | 38.6% |
Gemini 3.5 Flash Lite (high)Google | 36.7% |
Inkling (0.99)Thinking Machines | 34.1% |
GPT 5.4 Nano (high)OpenAI | 33.0% |
Qwen 3.6 PlusAlibaba | 32.0% |
Inkling Small (0.99)Thinking Machines | 31.9% |
Gemini 3 Flash (12/25) (high)Google | 29.4% |
Nemotron 3 UltraNVIDIA | 27.4% |
Kimi K2.5Moonshot AI | 26.3% |
MiniMax-M2.7MiniMax | 24.6% |
Grok 4.3 (high)xAI | 24.3% |
Claude Haiku 4.5 (Thinking)Anthropic | 22.9% |
Ling 3.0 FlashAnt Group | 21.7% |
Mistral Medium 3.5 (high)Mistral | 17.9% |
Grok 4.20 (Reasoning)xAI | 17.6% |
Gemini 3.1 Flash Lite Preview (high)Google | 15.5% |
Mercury 2.5 (high)Inception | 13.0% |
Nemotron 3.5 LightningNVIDIA | 11.5% |