Benchmarks

Scores, leaderboards, and test methods for the frontier benchmarks worth tracking

Agentic coding

Terminal-Bench 4.0

A 66-task benchmark of long-horizon work that an AI agent finishes inside a real Linux terminal sandbox, graded pass or fail by each task's own tests.

Current record65.2%Claude Opus 5.5
models tested
15
points from first to last
36.4
points ahead of second
+1.0
Read the breakdown
Top score by release dateVals AI independent run, Oct 1, 2026
40%60%JulAugSepOct
29 benchmarks
Computer useAgents' Last Exam (ALE)A UC Berkeley RDI benchmark that tests whether an AI agent can finish long professional tasks in real software, with deterministic checkers grading the deliverable file.
LeaderClaude Opus 5.538.2%
ReasoningARC-AGI-2 and ARC-AGI-3The ARC Prize Foundation's skill-acquisition benchmarks, which measure how efficiently a model learns a skill it has never seen, through static grid puzzles and interactive games.
LeaderGPT-6 Astra95.0%
MathematicsArXivMathMathArena's monthly benchmark of research-level math questions drawn from new arXiv papers, each with one exact final answer that a script or judge model checks.
LeaderClaude Opus 5.586.0%
Knowledge workAutomationBenchZapier's benchmark for AI agents that run business workflows across simulated SaaS apps, graded on the final state of every system with no judge model in the loop.
LeaderGemini 4 Argon51.3%
Knowledge workBenchCADA benchmark that asks a model to write executable CadQuery code rebuilding a real mechanical part from four rendered views, scored by voxel IoU against the reference.
LeaderClaude Sonnet 5.50.963
Long context and multimodalChartography and LVBenchTwo multimodal benchmarks. Chartography is Surge AI's 100-task test of professional chart reading, and LVBench asks questions about videos that run past an hour.
LeaderGemini 4 Argon71.6%
Agentic codingCursorBenchCursor's private coding-agent benchmark, built from real Cursor sessions, that scores ambiguous multi-file tasks on correctness, code quality, efficiency, and interaction behavior.
LeaderClaude Opus 5.557.8%
Agentic codingDeepSWEDatacurve's coding-agent benchmark of 113 original, long-horizon engineering tasks across 91 open-source repositories, graded by hand-written functional verifiers.
LeaderGPT-6 Astra74.1%
Knowledge workDRACO (Deep Research Accuracy, Completeness, and Objectivity)A 100-task benchmark for deep research agents, built from real Perplexity Deep Research queries, where an LLM judge grades each report against expert-written rubrics.
LeaderClaude Fable 5 + GPT-5.5, fused panel69
Agentic codingFrontierCodeCognition's coding benchmark that asks whether an agent's patch to a real open-source repository is one the project's maintainers would merge, graded on six rubric axes.
LeaderClaude Opus 5.554.6%
MathematicsFrontierMath Tier 4The hardest tier of Epoch AI's private FrontierMath set, research-level problems with one exact answer that a Python function returns and code checks.
LeaderGPT-6.1 Sol100.0%
Agentic codingFrontierSWE v2Proximal's ultra-long-horizon coding benchmark of 34 open-ended engineering projects, each with a 20-hour budget per trial and a deterministic partial-credit verifier.
LeaderGPT-6 Astra65.5%
Knowledge workGDPval-AAArtificial Analysis's agentic version of OpenAI's GDPval, which runs 220 public tasks through the Stirrup harness and ranks models by pairwise judge preferences.
LeaderGrok 4.71715
Long context and multimodalGraphWalksAn OpenAI long-context benchmark that gives a model a directed graph as an edge list and asks for the exact set of nodes from a breadth-first search or a parents lookup.
LeaderGemini 4 Argon99.7%
Knowledge workHarvey's Legal Agent Benchmark (LAB)Harvey's benchmark that gives an AI agent a partner-style instruction and a closed folder of matter documents, then grades the deliverable against expert rubrics.
LeaderMuse Spark 1.225.4%
HealthHealthBench Professional and HealthBench HardOpenAI's physician-rubric benchmarks for clinician chat tasks. Professional has 525 tasks, and Hard is a subset of the 2025 HealthBench, with a model grading each response.
LeaderGPT-6 Astra64.7%
ReasoningHumanity's Last Exam (HLE)A 2,500-question closed-ended academic exam from the Center for AI Safety and Scale AI, with one checkable reference answer per question and an LLM judge for grading.
LeaderGPT-6 Astra54.8%
ScienceInternal Research DebuggingOpenAI's private eval of whether a model can find and fix 41 real bugs from OpenAI research experiments, plus 6 alignment-auditing tasks, scored by rubric.
LeaderGPT-6 Astra78.0%
Agentic codingKernelGen 1P and NanoGPTTwo private OpenAI evals that test whether a model can speed up AI research by optimizing kernels and by cutting the training time of a small language model.
LeaderGPT-6 Astra66.7%
HealthMentalHealthBenchAn open OpenAI benchmark that grades a model's next reply in 1,215 synthetic mental health conversations against rubrics written by clinicians.
LeaderGPT-6 Astra57.3%
Long context and multimodalOpenAI MRCR v2 (Multi-Round Co-reference Resolution)A long-context benchmark that hides identical requests in a synthetic conversation of up to 1M tokens and asks the model to reproduce the answer to one of them exactly.
LeaderGPT-5.263.9%
Computer useOSWorld 2.0 and 2.1A benchmark of 108 long-horizon desktop tasks on a live Ubuntu virtual machine, where an agent works through screenshots, mouse, and keyboard and is graded on weighted checkpoints.
LeaderClaude Opus 534.7%
Agentic codingProgramBenchA benchmark that asks an AI agent to rebuild a complete program, in any language, from a compiled binary and its documentation, graded by hidden behavioral tests.
LeaderClaude Opus 5.518.5%
Computer useScreenSpot-ProA GUI grounding benchmark that tests whether a model can find and click the one correct element on full-resolution screenshots of professional desktop software.
LeaderIndeed-UI-8B73.4%
Security and operationsSEC-Bench ProA vulnerability-discovery benchmark that asks a coding agent to find a disclosed bug in V8, SpiderMonkey, or the Linux kernel and prove it with a crashing proof-of-concept.
LeaderGPT-5.558.4%
Security and operationsSRE-Bench (Software Reverse Engineering)A benchmark that tests whether AI agents can work out what private, often hardened compiled binaries do when they have no source code.
LeaderGPT-6 Astra56.9%
Agentic codingSWE-benchAn execution-based benchmark that measures whether AI models and agents can resolve real software engineering issues from GitHub repositories, judged by running unit tests.
LeaderClaude Opus 4.580.0%
Agentic codingTerminal-Bench 4.0A 66-task benchmark of long-horizon work that an AI agent finishes inside a real Linux terminal sandbox, graded pass or fail by each task's own tests.
LeaderClaude Opus 5.565.2%
ScienceTerminal-Bench Science (TB-Science) 0.1A 70-task benchmark of real scientific research workflows that an agent completes in a containerized terminal, graded pass or fail by each task's verifier.
LeaderGPT-6 Astra62.9%

Historical

Older benchmarks with the scores and dates they were published with

A benchmark score isn't your score

Leaderboards rank models on someone else's tasks. Your prompts, your data, and your idea of a good answer can reorder that list. Klu runs the same models on your evals, gives you your own ranking, and reruns it every time a new model ships.

Find your best model

Klu turns your use case into evals and reruns them every time a new model ships

Talk to sales