Benchmarks
Scores, leaderboards, and test methods for the frontier benchmarks worth tracking
Agentic coding
Terminal-Bench 4.0
A 66-task benchmark of long-horizon work that an AI agent finishes inside a real Linux terminal sandbox, graded pass or fail by each task's own tests.
Current record65.2%Claude Opus 5.5
- models tested
- 15
- points from first to last
- 36.4
- points ahead of second
- +1.0
Top score by release dateVals AI independent run, Oct 1, 2026
29 benchmarks
Computer useAgents' Last Exam (ALE)A UC Berkeley RDI benchmark that tests whether an AI agent can finish long professional tasks in real software, with deterministic checkers grading the deliverable file.LeaderClaude Opus 5.538.2%ReasoningARC-AGI-2 and ARC-AGI-3The ARC Prize Foundation's skill-acquisition benchmarks, which measure how efficiently a model learns a skill it has never seen, through static grid puzzles and interactive games.LeaderGPT-6 Astra95.0%MathematicsArXivMathMathArena's monthly benchmark of research-level math questions drawn from new arXiv papers, each with one exact final answer that a script or judge model checks.LeaderClaude Opus 5.586.0%Knowledge workAutomationBenchZapier's benchmark for AI agents that run business workflows across simulated SaaS apps, graded on the final state of every system with no judge model in the loop.LeaderGemini 4 Argon51.3%Knowledge workBenchCADA benchmark that asks a model to write executable CadQuery code rebuilding a real mechanical part from four rendered views, scored by voxel IoU against the reference.LeaderClaude Sonnet 5.50.963Long context and multimodalChartography and LVBenchTwo multimodal benchmarks. Chartography is Surge AI's 100-task test of professional chart reading, and LVBench asks questions about videos that run past an hour.LeaderGemini 4 Argon71.6%Agentic codingCursorBenchCursor's private coding-agent benchmark, built from real Cursor sessions, that scores ambiguous multi-file tasks on correctness, code quality, efficiency, and interaction behavior.LeaderClaude Opus 5.557.8%Agentic codingDeepSWEDatacurve's coding-agent benchmark of 113 original, long-horizon engineering tasks across 91 open-source repositories, graded by hand-written functional verifiers.LeaderGPT-6 Astra74.1%Knowledge workDRACO (Deep Research Accuracy, Completeness, and Objectivity)A 100-task benchmark for deep research agents, built from real Perplexity Deep Research queries, where an LLM judge grades each report against expert-written rubrics.LeaderClaude Fable 5 + GPT-5.5, fused panel69Agentic codingFrontierCodeCognition's coding benchmark that asks whether an agent's patch to a real open-source repository is one the project's maintainers would merge, graded on six rubric axes.LeaderClaude Opus 5.554.6%MathematicsFrontierMath Tier 4The hardest tier of Epoch AI's private FrontierMath set, research-level problems with one exact answer that a Python function returns and code checks.LeaderGPT-6.1 Sol100.0%Agentic codingFrontierSWE v2Proximal's ultra-long-horizon coding benchmark of 34 open-ended engineering projects, each with a 20-hour budget per trial and a deterministic partial-credit verifier.LeaderGPT-6 Astra65.5%Knowledge workGDPval-AAArtificial Analysis's agentic version of OpenAI's GDPval, which runs 220 public tasks through the Stirrup harness and ranks models by pairwise judge preferences.LeaderGrok 4.71715Long context and multimodalGraphWalksAn OpenAI long-context benchmark that gives a model a directed graph as an edge list and asks for the exact set of nodes from a breadth-first search or a parents lookup.LeaderGemini 4 Argon99.7%Knowledge workHarvey's Legal Agent Benchmark (LAB)Harvey's benchmark that gives an AI agent a partner-style instruction and a closed folder of matter documents, then grades the deliverable against expert rubrics.LeaderMuse Spark 1.225.4%HealthHealthBench Professional and HealthBench HardOpenAI's physician-rubric benchmarks for clinician chat tasks. Professional has 525 tasks, and Hard is a subset of the 2025 HealthBench, with a model grading each response.LeaderGPT-6 Astra64.7%ReasoningHumanity's Last Exam (HLE)A 2,500-question closed-ended academic exam from the Center for AI Safety and Scale AI, with one checkable reference answer per question and an LLM judge for grading.LeaderGPT-6 Astra54.8%ScienceInternal Research DebuggingOpenAI's private eval of whether a model can find and fix 41 real bugs from OpenAI research experiments, plus 6 alignment-auditing tasks, scored by rubric.LeaderGPT-6 Astra78.0%Agentic codingKernelGen 1P and NanoGPTTwo private OpenAI evals that test whether a model can speed up AI research by optimizing kernels and by cutting the training time of a small language model.LeaderGPT-6 Astra66.7%HealthMentalHealthBenchAn open OpenAI benchmark that grades a model's next reply in 1,215 synthetic mental health conversations against rubrics written by clinicians.LeaderGPT-6 Astra57.3%Long context and multimodalOpenAI MRCR v2 (Multi-Round Co-reference Resolution)A long-context benchmark that hides identical requests in a synthetic conversation of up to 1M tokens and asks the model to reproduce the answer to one of them exactly.LeaderGPT-5.263.9%Computer useOSWorld 2.0 and 2.1A benchmark of 108 long-horizon desktop tasks on a live Ubuntu virtual machine, where an agent works through screenshots, mouse, and keyboard and is graded on weighted checkpoints.LeaderClaude Opus 534.7%Agentic codingProgramBenchA benchmark that asks an AI agent to rebuild a complete program, in any language, from a compiled binary and its documentation, graded by hidden behavioral tests.LeaderClaude Opus 5.518.5%Computer useScreenSpot-ProA GUI grounding benchmark that tests whether a model can find and click the one correct element on full-resolution screenshots of professional desktop software.LeaderIndeed-UI-8B73.4%Security and operationsSEC-Bench ProA vulnerability-discovery benchmark that asks a coding agent to find a disclosed bug in V8, SpiderMonkey, or the Linux kernel and prove it with a crashing proof-of-concept.LeaderGPT-5.558.4%Security and operationsSRE-Bench (Software Reverse Engineering)A benchmark that tests whether AI agents can work out what private, often hardened compiled binaries do when they have no source code.LeaderGPT-6 Astra56.9%Agentic codingSWE-benchAn execution-based benchmark that measures whether AI models and agents can resolve real software engineering issues from GitHub repositories, judged by running unit tests.LeaderClaude Opus 4.580.0%Agentic codingTerminal-Bench 4.0A 66-task benchmark of long-horizon work that an AI agent finishes inside a real Linux terminal sandbox, graded pass or fail by each task's own tests.LeaderClaude Opus 5.565.2%ScienceTerminal-Bench Science (TB-Science) 0.1A 70-task benchmark of real scientific research workflows that an agent completes in a containerized terminal, graded pass or fail by each task's verifier.LeaderGPT-6 Astra62.9%
0%9 models · spread 24 pts100%
0%11 models · spread 28 pts100%
0%7 models · spread 47 pts100%
0%10 models · spread 17 pts100%
04 models · spread 0.0641
0%42 models · spread 63 pts100%
0%63 models · spread 42 pts100%
0%31 models · spread 62 pts100%
43.112 models · spread 25.969
0%14 models · spread 10 pts100%
0%17 models · spread 44 pts100%
0%10 models · spread 61 pts100%
160019 models · spread 115 pts1715
0%4 models · spread 9.1 pts100%
0%17 models · spread 23 pts100%
0%9 models · spread 17 pts100%
0%10 models · spread 20 pts100%
0%4 models · spread 14 pts100%
0%4 models · spread 29 pts100%
0%17 models · spread 28 pts100%
0%3 models · spread 19 pts100%
0%11 models · spread 31 pts100%
0%14 models · spread 18 pts100%
0%15 models · spread 73 pts100%
0%6 models · spread 58 pts100%
0%8 models · spread 57 pts100%
0%13 models · spread 78 pts100%
0%15 models · spread 36 pts100%
0%38 models · spread 63 pts100%
Historical
Older benchmarks with the scores and dates they were published with
- AlpacaEvalChat and preference
- BBHard (BIG-Bench Hard)Reasoning
- BigCodeBenchCoding
- GAIA (General AI Assistants)Computer use
- GPQA (Graduate-Level Google-Proof Q&A)Science
- GSM8K (Grade School Math 8K)Mathematics
- HellaSwagReasoning
- HumanEvalCoding
- LMSYS Chatbot Arena LeaderboardChat and preference
- MATH (Mathematics Assessment of Textual Heuristics)Mathematics
- MixEvalKnowledge
- MMLU (Massive Multi-task Language Understanding)Knowledge
- MMLU-ProKnowledge
- MMMU (Massive Multi-discipline Multimodal Understanding and Reasoning)Long context and multimodal
- MT-Bench (Multi-turn Benchmark)Chat and preference
- MTEB (Massive Text Embedding Benchmark)Embeddings
- MuSR (Multistep Soft Reasoning)Reasoning
- Needle In A HaystackLong context and multimodal
- RewardBenchChat and preference
- SuperGLUEKnowledge
- TruthfulQAKnowledge
A benchmark score isn't your score
Leaderboards rank models on someone else's tasks. Your prompts, your data, and your idea of a good answer can reorder that list. Klu runs the same models on your evals, gives you your own ranking, and reruns it every time a new model ships.
Terminal-Bench 4.0Your use case
- 1Claude Opus 5.565.2%
- 2Claude Sonnet 5.564.1%
- 3GPT-6 Astra59.6%
- 4Claude Fable 5.158.1%
- 5Gemini 4 Argon57.6%
- 1Claude Fable 5.1▲ 3
- 2Claude Opus 5.5
- 3GPT-6 Astra
- 4Claude Sonnet 5.5
- 5Gemini 4 Argon
Find your best model
Klu turns your use case into evals and reruns them every time a new model ships