Checking session…
What do model benchmarks actually measure? Start with the task, then explore the scores.
General intelligence: knowledge, reasoning, math and coding rolled into one number
Agentic ability: completing real tasks in a terminal, unattended
⌘KSearch the knowledge base
Loading the knowledge index…