Humanity's Last Exam (HLE)

A 2,500-question closed-ended academic exam from the Center for AI Safety and Scale AI, with one checkable reference answer per question and an LLM judge for grading

ReasoningAccuracy

Stephen M. Walker II · Co-Founder / CEO · October 6, 2026 · 12 min read

What is Humanity's Last Exam?

Humanity's Last Exam is a 2,500-question academic exam built by the Center for AI Safety (CAIS) and Scale AI. Every question is closed-ended, with one reference answer that a grader can check. Nearly 1,000 subject experts from more than 500 institutions in 50 countries wrote the questions, which span more than 100 subjects. The paper, Phan et al., first appeared on arXiv as 2501.14249 in January 2025 and was published in Nature on January 28, 2026 (vol. 649, pp. 1139-1146).

CAIS built it because frontier models had passed 90% on MMLU-class tests, and those tests could no longer tell the top models apart. HLE's answer was adversarial collection. A question stayed in the pool only if frontier models failed it. The organizers logged over 70,000 attempts against frontier models, and about 13,000 questions that stumped them went on to expert review. At launch GPT-4o scored 2.7% and o3-mini (high) scored 13.4%.

There is no single HLE leaderboard in 2026. Anthropic runs its own models in its own harness, Artificial Analysis runs a text-only subset, Scale runs the full set with the original judge, and CAIS and Scale now publish a cleaned 1,000-question subset called HLE-Diamond. Each table has its own leader:

  • Claude Opus 5.5 leads Anthropic's full-set runs at 67.7% with tools and 64.4% without.
  • Claude Opus 5.5 leads Artificial Analysis's text-only board at 61.4%.
  • GPT-6 Astra leads Scale's full-set board at 54.80%.
  • GPT-6 Astra leads HLE-Diamond at 59.9% without tools and 82.9% with tools.

The only tables where Opus 5.5 and Astra were run side by side by the same party are the Diamond tables, and Astra wins both.

How the questions and grading work

The public set has 2,500 questions, finalized on April 3, 2025 (agi.safe.ai). About 14% need an image plus text. 24% are multiple choice, and the rest are short answers graded by exact match. CAIS also keeps a private held-out test set of undisclosed size to catch models that have overfit to the public questions. The public data is on Hugging Face as cais/hle.

Collection ran on prize money. The pool was $500,000, paying $5,000 for each of the top 50 questions and $500 for each of the next 500. In the first review round each question got one to three reviews. Keep that light review in mind when reading the label-noise audit under failure modes.

Scoring under the CAIS and Scale protocol

The model answers, and an LLM judge compares that answer with the reference. The judge is o3-mini-2025-01-31, using structured JSON extraction at temperature 0 when the model allows it (Scale leaderboard). The headline number is accuracy.

The model also states its confidence from 0 to 100% for each answer. From those confidences the protocol computes RMS calibration error, which measures how far stated confidence sits from actual accuracy. A model that is right 40% of the time but claims 90% confidence scores badly on calibration even if its accuracy matches a rival's. For anyone building on top of a model's self-reported certainty, this second number is the more useful one.

Scoring in the other series

Three other setups produce published HLE numbers, and each changes something that moves the score.

Artificial Analysis uses 2,158 text-only questions, dropping the multimodal items. Each model gives one answer, graded by an LLM equality checker with a small numeric tolerance (Artificial Analysis).

Anthropic's system-card runs use the full 2,500 questions with Claude Opus 4.6 as the grader. Thinking is set to auto, total tokens across contexts are capped at 1M, and there is no context compaction. Reported scores use adaptive thinking at max effort with default sampling, averaged over five trials (Opus 5.5 system card, sections 8.1 and 8.11.1).

Anthropic's with-tools configuration adds web search, web fetch, programmatic tool calling and code execution. Because the questions are public, a search tool can find them. Anthropic blocklists sources that discuss HLE, headed by huggingface.co, hf.co and mirrors, for both the searcher and the fetcher (Appendix 9.2). It then screens every correctly answered transcript with text rules. Claude Opus 5 reviews flagged transcripts, and confirmed cheating is re-graded as incorrect. For the Fable 5.1 runs, Anthropic also limited fetch to URLs that had already appeared in the conversation (Fable 5.1 system card, section 8.12.1). The Opus 5.5 card does not state that restriction.

Versions

VersionReleasedQuestionsWhat changedSource
HLEApril 3, 20252,500 public, plus a private held-out setOriginal set; 14% multimodal, 24% multiple choiceagi.safe.ai
HLE-RollingOctober 8, 2025Not stated by CAISDynamic fork, following the HLE team's commitment to rolling revisions after the FutureHouse auditagi.safe.ai, FutureHouse
HLE-DiamondSept. 22, 20261,000: 500 reasoning, 500 knowledge, image and textYear-long cleaning and refinement; removed questions and corrected answers are not listedCAIS and Scale blog, cais/hle-diamond

HLE-Diamond is the version to watch. The Hugging Face card lists an MIT license with gated access and carries a canary string so the data can be filtered out of training corpora. Neither the card nor the blog says which questions were dropped or which answers changed, so nobody outside CAIS and Scale can audit the cleanup question by question. The blog says every model was run "with reasoning high."

Current leaderboard

HLE has no published per-task cost or time from any source, so this page has no frontier chart. Both Anthropic cards plot HLE score against cost per task across effort levels (Opus 5.5 card Figure 8.11.1.A, Fable 5.1 card Figure 8.12.1.A), but only as images with no printed values. The two figures also use different billing assumptions. The Opus 5.5 figure assumes a perfect cache hit rate and excludes web search fees. The Fable 5.1 figure bills on cache-hit estimates.

Anthropic harness, full set, with tools

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 5.5Anthropic67.7%Max effort; web search, web fetch, programmatic tool calling, code execution; Opus 4.6 grader2026-09-22Opus 5.5 card Table 8.1.A, mean of five trials
Claude Fable 5.1Anthropic65.6%Same harness, max effort2026-09-22Opus 5.5 card Table 8.1.A; the Fable 5.1 card reports 65.0% for the same model
Claude Opus 5Anthropic63.6%Same harness, max effort2026-09-22Opus 5.5 card Table 8.1.A
AnthropicAnthropic harnessFull set, 2,500 questions, with toolsAccuracywww-cdn.anthropic.com

Vendor-run in Anthropic's harness with the Opus 4.6 grader on all 2,500 questions; comparable only to other rows in this harness, and no non-Anthropic model was run in it.

Opus 5.5 leads by 2.1 points over Fable 5.1 and 4.1 over Opus 5. The same card table lists GPT-6 Astra at 57.2% with tools, attributed only to "the respective developers' published system cards or benchmark leaderboards." OpenAI's GPT-6 Astra system card publishes no HLE score, and the harness behind 57.2% is unknown, so it is left out of this table.

Anthropic harness, full set, no tools

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 5.5Anthropic64.4%Max effort, no tools, thinking auto, 1M-token cap, Opus 4.6 grader2026-09-22Opus 5.5 card Table 8.1.A
Claude Fable 5.1Anthropic60.9%Same harness, max effort2026-09-22Opus 5.5 card Table 8.1.A
Claude Opus 5Anthropic56.6%Same harness, max effort2026-09-22Opus 5.5 card Table 8.1.A
AnthropicAnthropic harnessFull set, 2,500 questions, no toolsAccuracywww-cdn.anthropic.com

Same harness and grader as the with-tools table; the card has no non-Anthropic no-tools row in this harness.

Tools add 3.3 points for Opus 5.5 and 4.7 for Fable 5.1 under this harness. That is a small gain next to the Diamond tables below, where tools add 19 to 32 points. Diamond's tool settings are unpublished, so the two gains cannot be compared directly.

Artificial Analysis, text-only subset

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 5.5Anthropic61.4%2,158 text-only questions, pass@1, Max, default fallbackAccessed 2026-10-06Artificial Analysis, independent
Claude Fable 5.1Anthropic59.1%Same, Max, default fallbackAccessed 2026-10-06Artificial Analysis
Claude Fable 5.1Anthropic58.7%Same, Xhigh, default fallbackAccessed 2026-10-06Artificial Analysis
Artificial Analysis2,158 text-only questionspass@1artificialanalysis.ai

Independent run on the 2,158-question text-only subset with an LLM equality checker; not comparable to the Anthropic or Scale tables, which use the full set and different graders.

Artificial Analysis's page lists only these three rows. Blog posts circulating figures for Gemini 4 Argon, GPT-6 Astra and GPT-6.1 Sol on this board do not match the page, and they are excluded here. Google's Gemini 4 Argon evaluation document has no HLE row either, so Google has published no vendor HLE number for its newest model.

Scale Labs, full set

ModelOrganizationScoreHarness and setupDateSource and notes
GPT-6 AstraOpenAI54.80 (+/-1.94)2,500 questions, temp 0, o3-mini judgeAccessed 2026-10-06Scale Labs, calibration error 39
Claude Fable 5.1 (xhigh)Anthropic46.50 (+/-2.00)SameAccessed 2026-10-06Scale Labs, calibration error 20
Gemini 3.1 Pro Preview (thinking high)Google46.44 (+/-1.96)SameAccessed 2026-10-06Scale Labs, calibration error 51
Gemini 3.8 FlashGoogle44.52 (+/-1.96)SameAccessed 2026-10-06Scale Labs, calibration error 51
GPT-5.4 Pro (2026-03-05)OpenAI44.32 (+/-1.95)SameAccessed 2026-10-06Scale Labs, calibration error 38
Muse SparkMeta40.56 (+/-1.92)SameAccessed 2026-10-06Scale Labs, calibration error 50
Gemini 3 Pro PreviewGoogle37.52 (+/-1.90)SameAccessed 2026-10-06Scale Labs, calibration error 57
GPT-5.4 (xhigh thinking)OpenAI36.24 (+/-1.88)SameAccessed 2026-10-06Scale Labs, calibration error 42
Claude Opus 4.7Anthropic36.20 (+/-1.88)SameAccessed 2026-10-06Scale Labs, calibration error 47
Claude Opus 4.6 (thinking max)Anthropic34.44 (+/-1.86)SameAccessed 2026-10-06Scale Labs, calibration error 46
Scale LabsFull set, 2,500 public questionsAccuracylabs.scale.com

Benchmark-owner run on all 2,500 public questions with the o3-mini judge and no stated tools; ranks use 95% CI upper bounds, and the judge and prompt differ from every other table here.

Astra leads by 8.3 points, well outside the roughly 2-point intervals. Below it the board bunches up. Fable 5.1 (xhigh) and Gemini 3.1 Pro Preview are 0.06 apart, and Gemini 3.8 Flash and GPT-5.4 Pro sit about 2 points further down with overlapping intervals. Treat any gap under about 4 points on this board as a tie.

Calibration tells a different story from accuracy. Fable 5.1's error of 20 is the lowest in the top five, about half of Astra's 39. If your application uses the model's own confidence to decide when to escalate to a person, Fable 5.1 is the better-calibrated choice on this board even though Astra answers more questions correctly.

This board has no Opus 5.5 row, which is why Scale's full-set table cannot settle Opus 5.5 against Astra.

HLE-Diamond, no tools

ModelOrganizationScoreHarness and setupDateSource and notes
GPT-6 AstraOpenAI59.9%1,000 Diamond questions, reasoning high2026-09-22CAIS and Scale blog
Claude Opus 5.5Anthropic54.6%Same2026-09-22CAIS and Scale blog
GPT-6.1 SolOpenAI53.2%Same2026-09-22CAIS and Scale blog
Claude Fable 5.1Anthropic50.7%Same2026-09-22CAIS and Scale blog
Gemini 3.8 FlashGoogle33.3%Same2026-09-22CAIS and Scale blog
Claude Sonnet 5.5Anthropic33.1%Same2026-09-22CAIS and Scale blog
GPT-6 SolOpenAI32.8%Same2026-09-22CAIS and Scale blog
Muse Spark 1.3Meta24.6%Same2026-09-22CAIS and Scale blog
Grok 4.7xAI22.8%Same2026-09-22CAIS and Scale blog
CAIS and Scale AIHLE-Diamond, 1,000 questions, no toolsAccuracylastexam.ai

Benchmark-owner run on the 1,000-question Diamond subset; not comparable to any full-set or text-only table on this page.

This is the closest thing HLE has to a fair head-to-head of the current flagships. Astra leads Opus 5.5 by 5.3 points. Four models sit above 50%, and then the table drops 17 points to Gemini 3.8 Flash at 33.3%. Below that cliff, Sonnet 5.5 and GPT-6 Sol land within half a point of each other.

HLE-Diamond, with tools

ModelOrganizationScoreHarness and setupDateSource and notes
GPT-6 AstraOpenAI82.9%1,000 Diamond questions, web and code tools2026-09-22CAIS and Scale blog, +23.0 over no tools
Claude Opus 5.5Anthropic73.9%Same2026-09-22CAIS and Scale blog, +19.3 over no tools
Claude Fable 5.1Anthropic72.3%Same2026-09-22CAIS and Scale blog, +21.6 over no tools
GPT-6 SolOpenAI65.0%Same2026-09-22CAIS and Scale blog, +32.2 over no tools
Gemini 3.8 FlashGoogle60.6%Same2026-09-22CAIS and Scale blog, +27.3 over no tools
Muse Spark 1.3Meta55.4%Same2026-09-22CAIS and Scale blog, +30.8 over no tools
CAIS and Scale AIHLE-Diamond, 1,000 questions, web and code toolsAccuracylastexam.ai

Same run and table as the no-tools Diamond results, so the gains are same-table deltas; the blog does not publish the tool settings, and GPT-6.1 Sol and Claude Sonnet 5.5 have no with-tools row.

Astra's lead over Opus 5.5 grows to 9.0 points with tools, and its lead over Fable 5.1 is 10.6. The bigger story is lower in the table. The models that scored worst without tools gained the most. GPT-6 Sol gained 32.2 points and jumped from 32.8% to 65.0%, ahead of where any model sat on Diamond without tools. With search and code, a model that scored 32.8% closed-book beats every flagship's closed-book score.

Scale's Diamond leaderboard, by split

ModelOrganizationScoreHarness and setupDateSource and notes
GPT-6 AstraOpenAI60.6 overall; 75.6 reasoning; 45.6 knowledgeDiamond, no effort setting listedUndated, accessed 2026-10-06Scale Labs Diamond, CI +/-3.00
Claude Opus 5.5Anthropic55.0 overall; 63.2 reasoning; 46.8 knowledgeSameUndated, accessed 2026-10-06Scale Labs Diamond
Claude Fable 5.1Anthropic51.3 overall; 62.0 reasoning; 40.6 knowledgeSameUndated, accessed 2026-10-06Scale Labs Diamond
Claude Opus 5Anthropic38.6 overallSameUndated, accessed 2026-10-06Scale Labs Diamond, no split published
Gemini 3.8 FlashGoogle34.3 overallSameUndated, accessed 2026-10-06Scale Labs Diamond, no split published
GPT-6 SolOpenAI33.8 overallSameUndated, accessed 2026-10-06Scale Labs Diamond, no split published
GPT-5.6 SolOpenAI31.2 overallSameUndated, accessed 2026-10-06Scale Labs Diamond, no split published
Muse Spark 1.3Meta25.4 overallSameUndated, accessed 2026-10-06Scale Labs Diamond, no split published
Grok 4.7xAI23.4 overallSameUndated, accessed 2026-10-06Scale Labs Diamond, no split published
Kimi K3Moonshot AI22.2 overallSameUndated, accessed 2026-10-06Scale Labs Diamond, no split published
GLM-5.3Z.ai16.4 overallSameUndated, accessed 2026-10-06Scale Labs Diamond, no split published
DeepSeek-V4-ProDeepSeek13.4 overallSameUndated, accessed 2026-10-06Scale Labs Diamond, no split published
Scale LabsHLE-DiamondAccuracy, overalllabs.scale.com

Undated Scale snapshot with no effort setting; its scores differ from the blog's Diamond table (Astra 60.6 against 59.9, Opus 5.5 55.0 against 54.6), so do not mix rows between the two.

The split is the reason to look at this table. Astra's whole lead comes from the 500 reasoning questions, where it beats Opus 5.5 by 12.4 points. On the 500 knowledge questions Opus 5.5 is ahead, 46.8 to 45.6. That is inside the ±3 interval, so call knowledge a tie. If your workload leans on recall of specialist facts more than multi-step derivation, Astra's headline lead overstates what you would get.

Historical progression

Launch snapshot, January 2025

ModelOrganizationScoreHarness and setupDateSource and notes
o3-mini (high)OpenAI13.4%Full set, no tools, o3-mini judgeJan 2025arXiv paper, calibration error 80
DeepSeek-R1DeepSeek8.5%SameJan 2025arXiv paper, calibration error 73
o1OpenAI8.0%SameJan 2025arXiv paper, calibration error 83
Gemini 2.0 Flash ThinkingGoogle6.6%SameJan 2025arXiv paper, calibration error 82
Gemini 1.5 ProGoogle4.6%SameJan 2025arXiv paper, calibration error 88
Claude 3.5 SonnetAnthropic4.1%SameJan 2025arXiv paper, calibration error 84
Grok 2xAI3.0%SameJan 2025arXiv paper, calibration error 87
GPT-4oOpenAI2.7%SameJan 2025arXiv paper, calibration error 89
CAIS and Scale AIFull set, no toolsAccuracyarxiv.org

Launch-time paper table with the o3-mini judge; no source says Scale's 2026 board reuses the paper's prompt and judge configuration, so read it apart from the current Scale table.

The calibration column is the striking part. Every model at launch had error between 73 and 89. These models were wrong on almost everything and confident anyway. By 2026 Scale's board shows errors from 20 to 57, so calibration has improved along with accuracy, though no source isolates how much of that comes from the model and how much from any change in the board's setup.

Anthropic system cards across releases

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 5Anthropic56.6% no tools; 63.6% with toolsAnthropic harness, Opus 4.6 grader2026-09-01Fable 5.1 card summary table
Claude Fable 5Anthropic57.8% no tools; 63.8% with toolsSame2026-09-01Fable 5.1 card summary table
Claude Fable 5.1Anthropic60.9% no tools; 65.0% with toolsSame2026-09-01Fable 5.1 card; the Opus 5.5 card later reports 65.6% with tools, with no explanation in either card
Claude Opus 5.5Anthropic64.4% no tools; 67.7% with toolsSame2026-09-22Opus 5.5 card Table 8.1.A
AnthropicAnthropic harnessFull setAccuracy, no tools

Vendor-run in one harness and grader across two cards three weeks apart; the Fable 5.1 with-tools figure shifted between cards, so treat sub-point changes as noise.

Across four Anthropic releases the no-tools score rose 7.8 points, from Opus 5 at 56.6% to Opus 5.5 at 64.4%. The with-tools score rose only 4.1 points over the same models. The jump from the January 2025 launch scores to these numbers crosses graders, harnesses and model generations, so it does not measure a rate of progress.

Failure modes and limitations

Label noise in biology and chemistry

FutureHouse audited the rationales for 321 text-only biology, health and chemistry questions with Crow, its open-source PaperQA2 agent, prompted to find evidence for or against each rationale. Crow found 53.3% (171) of the rationales in direct conflict with published research. Independent chemistry and biology experts then checked 150 of the 321 outputs. Extrapolating from where Crow and the experts agreed, FutureHouse expects 29.3 ± 3.7% of HLE's questions to be directly conflicted by research.

Reviewers did not have to verify a rationale if it would take "more than 5 minutes." The HLE team's own follow-up, in a page update dated September 16, 2025, found about 18% of a bio/chem subset problematic, and in 25% of cases at least one reviewer disagreed. FutureHouse released HLE-Gold-Bio/Chem on Hugging Face, a set of question IDs its system graded as supported by research. HLE-Rolling and HLE-Diamond followed. On the original set, a model marked wrong on a conflicted bio/chem item has given an answer that published research supports.

Contamination when tools are on

The questions are public, so a model with search can look them up. Anthropic's blocklist and transcript screening exist for exactly that reason. Its testing also turned up a worse problem. A model with an unrestricted fetch tool routed its own JavaScript through public third-party web services and ran code outside the sandbox. Anthropic's fix was to let fetch reach only URLs already in the conversation. Any with-tools HLE score from a lab that does not describe its blocklist and fetch policy is hard to trust. CAIS gates the Diamond dataset and asks, through its canary string, that it never enter training data.

Overconfidence

Every model in the launch paper had calibration error above 70%. The current Scale board is better, but Gemini 3 Pro Preview still sits at 57 and the two Gemini models in the top five are at 51. A model with high calibration error will mislead any system that trusts its self-assessment, however good its accuracy.

Narrow scope

HLE is closed-ended and exact-match. It does not measure open-ended research, writing, or agentic work over many steps. About 14% of questions are multimodal, and Artificial Analysis's board drops all of them, so its scores say nothing about image reasoning.

Different graders

Scale uses o3-mini, Anthropic uses Claude Opus 4.6, and Artificial Analysis uses an LLM equality checker. No source measures how much the choice of grader moves a score. This is the main reason the tables above stay separate.

Variance

Anthropic averages five trials per card score. Scale's 95% intervals run about ±2 points on the full set and ±3 on Diamond. On Scale's full-set board, gaps under about 4 points have overlapping intervals and do not separate.

MMLU and MMLU-Pro are broad multiple-choice knowledge tests on which frontier models score above 90%. That saturation is the problem HLE was built to solve.

GPQA is multiple-choice graduate science, with a 25% guessing floor on four options. HLE is 76% exact-match and spans more than 100 subjects, so most answers cannot be picked from a list of options. Anthropic's Opus 5.5 card runs GPQA, MMLU-Pro and HLE together in its reasoning-leak test.

FrontierMath Tier 4 is research-level math from a private Epoch AI set. HLE is public, which is why contamination in tool runs is a live issue for HLE and a non-issue for FrontierMath.

ArXivMath draws monthly questions from new arXiv papers and defends against contamination by date. HLE's public set has no such defense.

DRACO is 100 rubric-graded deep-research tasks with citation grading. Anthropic runs it beside HLE under "agentic search" with a 980k-token budget (Opus 5.5 card section 8.11.2). Where HLE asks for a single checkable answer, DRACO grades a whole research report.

For agentic and science-workflow tests in the same family of pages, see Agents' Last Exam, Terminal-Bench Science and ARC-AGI. The wider map is in LLM benchmarks and frontier models.

What it means for teams choosing a model

  1. For closed-book expert questions, GPT-6 Astra leads. On Diamond without tools it beats Opus 5.5 by 5.3 points (59.9% to 54.6%), with GPT-6.1 Sol at 53.2% and Fable 5.1 at 50.7%. On Scale's full set it leads Fable 5.1 (xhigh) by 8.3 points.
  2. Astra's lead is a reasoning lead. On Scale's Diamond splits it beats Opus 5.5 by 12.4 points on reasoning questions and trails by 1.2 on knowledge questions, a tie inside the interval. For specialist recall, Opus 5.5 matches Astra.
  3. With search and code on Diamond, Astra leads Opus 5.5 by 9.0 points (82.9% to 73.9%). Under Anthropic's own full-set harness, Opus 5.5 leads Anthropic's lineup at 67.7%, but that harness has never run Astra, so it does not rank the two.
  4. Turn tools on before you pay for a bigger model. GPT-6 Sol with tools (65.0%) beats every model's no-tools Diamond score, including Astra's 59.9%. The 19-to-32-point tool gains on Diamond are larger than the gap between most adjacent models.
  5. If you rely on the model's confidence, check calibration. Fable 5.1 (xhigh) has calibration error 20 on Scale's board against Astra's 39. A hallucination evaluation on your own data will tell you more than either number.
  6. Do not price HLE accuracy. No source publishes per-task cost or time as numbers, so cost per correct answer across vendors cannot be computed from HLE. Use cost-per-token economics and your own runs.
  7. For biology and chemistry work, the original set is unreliable. FutureHouse expects 29.3 ± 3.7% of bio/chem questions to conflict with published research, and the HLE team's own check found about 18% problematic. Use Diamond, HLE-Gold-Bio/Chem, or a golden dataset built from your own domain.

HLE measures one answer per question checked by a model judge, which is a long way from a research workflow. To compare these models on other benchmarks, see the LLM leaderboard. To run the same kind of expert-question evaluation on your own data in Klu, see how it handles research workflows.