FrontierMath Tier 4

The hardest tier of Epoch AI's private FrontierMath set, research-level problems with one exact answer that a Python function returns and code checks

MathematicsAccuracy

Stephen M. Walker II · Co-Founder / CEO · October 6, 2026 · 13 min read

What is FrontierMath Tier 4?

FrontierMath Tier 4 is the hardest part of FrontierMath, Epoch AI's private set of original mathematics problems written and vetted by expert mathematicians. Epoch released FrontierMath in November 2024 with Tiers 1-3 (Glazer et al., arXiv:2411.04872). Epoch describes those tiers as "undergraduate problems through exploratory problems suitable for an advanced graduate student." Tier 4, added in July 2025, is "research-level mathematics" (FrontierMath overview).

Every problem has one exact answer that a program can check and a model cannot guess, and the problems are unpublished. That combination is why Tier 4 is worth tracking. There are no answer choices to eliminate, no public test set to memorize, and no judge model deciding whether an argument sounds right. The model either returns the right object or it scores zero.

The current version is Tier 4 v2, released June 12, 2026 with 43 problems. Epoch scores models on the 41 that are private. On that set GPT-6.1 Sol at max effort scores 100.0%, GPT-6 Astra 97.6% and Claude Opus 5.5 95.0%. The best models from any other developer, Qwen3.8-Max and Meta's Muse Spark 1.3, tie at 46.3%.

Every score on this page is from Epoch's own runs on its own scaffold, taken from Epoch's benchmark data download (CC-BY 4.0) on October 6, 2026. No lab-published Tier 4 score was retrievable on that date, and Anthropic's Claude Opus 5.5 release post does not mention FrontierMath. OpenAI funded FrontierMath and has exclusive access to part of it. That conflict gets its own entry under failure modes below.

How the problems and harness work

The v2 FrontierMath set has 338 problems: 295 in Tiers 1-3 and 43 in Tier 4. Twelve are public, ten from Tiers 1-3 and two from Tier 4. Topics include computation-heavy fields such as number theory and real analysis and abstract ones such as algebraic geometry and category theory. It is computational mathematics in the literal sense, because the final step is always a program.

The model does not write a proof. It submits a Python function called answer() that takes no arguments and returns the answer, usually an integer or a sympy object. Epoch's scaffold gives the model three things:

  • free-form reasoning in its own messages
  • a stateless python tool that returns stdout only, with a 30-second limit per call
  • a submit_answer tool

The submitted answer() function also has a 30-second runtime limit, and there is no proof assistant. A brute-force search that needs ten minutes of compute fails, so the model has to cut the problem down by reasoning before the final code runs.

Each run has a hard cap of 1,000,000 tokens, input plus output. If a run crosses it, the run ends on the spot. At 660,000 tokens the scaffold forces the model to call submit_answer on its next message. Epoch publishes no step limit and no wall-clock limit.

Epoch also ran some models through their web apps, for models that had no API at the time. Those runs used a simple prompt, allowed web search and code execution, and were graded by hand (FrontierMath Tier 4: Battle Royale, October 13, 2025). Epoch allows web search there because the problems are not public. It is a different protocol from the scaffold, and its results sit in their own table in the history section.

Scoring

Scoring is binary. A problem earns 1 point if answer() returns the correct value and 0 if the value is wrong or nothing was submitted. There is no partial credit and no proof grading. Epoch designs answers to be "guessproof," meaning large numbers or structured objects with under a 1% chance of a lucky guess (benchmark design).

The private v2 set has 41 problems, so one problem is worth 2.44 points. That number decides how to read the leaderboard. Epoch's standard errors for most rows run 4 to 8 points, which means two configurations within about five points of each other, two problems, are not distinguishable on this data.

Several rows carry values that are not multiples of 1/41, such as 90.0%, 95.0%, 72.5% and 75.6%. Epoch does not state the denominator for those rows. Opus 5.5's 95.0% is one of them.

Versions

VersionReleasedTier 4 problemsPrivate problems scoredOne problem is worthPage
v1July 2025 (hub version FrontierMath-Tier-4-2025-07-01-Private, 2025-07-11)50482.08 pointsfrontiermath-tier-4
v22026-06-1243412.44 pointsfrontiermath-tier-4-v2

The v2 changelog on Epoch's page lists the fixes. In Tier 4, Epoch "corrected 12 problems" and "removed 7 problems," taking the set from 50 to 43. In Tiers 1-3 it "corrected 123 problems" and "removed 5 problems." Epoch says the update addressed "errors in 42% of problems" across the dataset.

Treat v1 and v2 as two benchmarks. The same model at the same effort on the same scaffold scores 3.9 to 24.9 points higher on v2 than on v1, so any jump from a v1 number to a v2 number mixes model progress with the corrections. v1 is still published, and its results are in the history section.

Current leaderboard (Epoch v2, October 2026)

All rows are Epoch-run on the 41 private v2 problems with the scaffold described above: Python tool, submit_answer, 1,000,000-token cap. The setting is the effort or token-budget suffix on Epoch's model-version string. The date is when Epoch started the run. Standard errors (SE) are Epoch's. Epoch publishes no cost, token count or time per task for Tier 4 v2, so this page has no cost or time chart.

Frontier models

ModelOrganizationScoreHarness and setupDateSource and notes
GPT-6.1 SolOpenAI100.0%Epoch scaffold, max effort2026-09-29Epoch v2, SE 0.0, no log link
GPT-6 AstraOpenAI97.6%Epoch scaffold, high, xhigh and max (all identical)2026-08-30Epoch v2, SE 2.4, no log link
Claude Opus 5.5Anthropic95.0%Epoch scaffold, max effort2026-09-22Epoch v2, SE 4.1, no log link
Claude Fable 5Anthropic90.2%Epoch scaffold, max effort2026-06-09Epoch v2, SE 4.6, log linked
GPT-6 SolOpenAI90.0%Epoch scaffold, max effort2026-09-22Epoch v2, SE 5.2, no log link
Claude Fable 5.1Anthropic87.8%Epoch scaffold, max effort2026-09-01Epoch v2, SE 5.2, no log link
GPT-5.6 SolOpenAI82.9%Epoch scaffold, max effort2026-07-09Epoch v2, SE 5.9, log linked
Claude Sonnet 5.5Anthropic80.5%Epoch scaffold, max effort2026-09-29Epoch v2, SE 6.3, no log link
GPT-5.6 SolOpenAI80.5%Epoch scaffold, promax setting2026-07-09Epoch v2, SE 6.3, log linked
GPT-5.5 ProOpenAI78.0%Epoch scaffold, xhigh effort2026-06-12Epoch v2, SE 6.5, no log link
Claude Opus 5Anthropic73.2%Epoch scaffold, max effort2026-07-24Epoch v2, SE 7.0, log linked
GPT-5.5OpenAI72.5%Epoch scaffold, xhigh effort2026-06-11Epoch v2, SE 7.2, no log link
GPT-5.6 TerraOpenAI70.7%Epoch scaffold, max effort2026-07-09Epoch v2, SE 7.2, log linked
GPT-5.6 LunaOpenAI61.0%Epoch scaffold, max effort2026-07-09Epoch v2, SE 7.7, log linked
GPT-5.4 ProOpenAI58.5%Epoch scaffold, xhigh effort2026-06-13Epoch v2, SE 7.8, log linked
GPT-6 LunaOpenAI56.1%Epoch scaffold, max effort2026-09-22Epoch v2, SE 7.8, no log link
Claude Opus 4.8Anthropic56.1%Epoch scaffold, max effort2026-06-10Epoch v2, SE 7.8, log linked
Epoch AIEpoch scaffoldTier 4 v2, 41 private problemsAccuracyepoch.ai

Comparable to every other v2 row on this page from Epoch's scaffold; not comparable to v1 scores, web-app runs or lab-reported numbers.

GPT-6.1 Sol leads at 100.0%, 2.4 points ahead of GPT-6 Astra and 5.0 ahead of Claude Opus 5.5. With standard errors of 0.0, 2.4 and 4.1, Epoch's data does not separate those three. What it does separate is developers. Opus 5.5, the lowest of the three, sits 48.7 points above the best bare model from anyone other than OpenAI or Anthropic. None of the three top rows has a log viewer link, so nobody outside Epoch can inspect those runs from the published data.

Model size inside a family moves the score more than anything else in the table. GPT-6 Sol at 90.0% beats GPT-6 Luna at 56.1% by 33.9 points, and GPT-5.6 Sol beats GPT-5.6 Luna by 21.9. Claude Sonnet 5.5 at 80.5% trails Opus 5.5 by 14.5 points.

GPT-6 Astra by reasoning effort

EffortScoreChange from previous step
none82.9%(baseline)
low87.8%+4.9 points
medium97.56%+9.8 points
high97.6%0
xhigh97.6%0
max97.6%0

One Epoch run date (2026-08-30), one scaffold, one model; comparable only within this table and to the v2 leaderboard above.

Going from no reasoning to medium effort adds 14.6 points, six problems. Everything above medium adds nothing on this run. It is the cleanest test-time compute curve in the Tier 4 data, and it flattens at medium.

Other developers

ModelOrganizationScoreHarness and setupDateSource and notes
Qwen3.8-MaxAlibaba46.3%Epoch scaffold, xhigh effort2026-08-04Epoch v2, SE 7.9, log linked
Muse Spark 1.3Meta AI46.3% (max); 41.5% (xhigh)Epoch scaffold, max and xhighOn or after 2026-08-25Epoch v2, SE 7.9 / 7.8, no log link
Kimi K3Moonshot39.0%Epoch scaffold, max effortBefore 2026-08-25Epoch v2, SE 7.7, log linked
Gemini 3.7 FlashGoogle DeepMind36.6%Epoch scaffold, high effortBefore 2026-08-25Epoch v2, SE 7.6, log linked
Qwen3.8-Max (0902)Alibaba34.1%Epoch scaffold, xhigh effortOn or after 2026-08-25Epoch v2, SE 7.5, no log link
Qwen3.7-MaxAlibaba34.1%Epoch scaffold, default settingBefore 2026-08-25Epoch v2, SE 7.5, log linked
Grok 4.6xAI31.7%Epoch scaffold, xhigh effortBefore 2026-08-25Epoch v2, SE 7.4, log linked
GLM-5.3Z.ai29.3%Epoch scaffold, max effortOn or after 2026-08-25Epoch v2, SE 7.2, no log link
GLM-5.2Z.ai29.3%Epoch scaffold, max effortBefore 2026-08-25Epoch v2, SE 7.2, log linked
Gemini 3.1 Pro PreviewGoogle DeepMind26.8%Epoch scaffold, default settingBefore 2026-08-25Epoch v2, SE 7.0, log linked
DeepSeek V4 Pro (0813)DeepSeek26.8%Epoch scaffold, max effortBefore 2026-08-25Epoch v2, SE 7.0, log linked
Gemini 3.8 FlashGoogle DeepMind22.0%Epoch scaffold, high effortOn or after 2026-08-25Epoch v2, SE 6.5, no log link
Grok 4.7xAI17.1%Epoch scaffold, xhigh effortOn or after 2026-08-25Epoch v2, SE 5.9, no log link
Epoch AIEpoch scaffoldTier 4 v2, 41 private problemsAccuracyepoch.ai

Same version, problem set and scaffold as the frontier table, so these rows compare directly with it; not comparable to v1, web-app runs or lab-reported numbers.

Where the date is a range, it comes from the log column, because Epoch's CSV leaves every run started on or after 2026-08-25 unlinked and every row here with a link started earlier. No model in this table passes 47%. Qwen3.8-Max and Muse Spark 1.3 tie at 46.3%, which is 48.7 points behind Opus 5.5 and 53.7 behind GPT-6.1 Sol. Google's best bare model is Gemini 3.7 Flash at 36.6%, and Gemini 3.1 Pro Preview scores 26.8%. Epoch's v2 table has no Gemini 4 and no Google Pro-class model newer than Gemini 3.1 Pro Preview, so this table does not describe Google's current flagship. Two later versions score lower than their predecessors on the same set, Qwen3.8-Max (0902) at 34.1% against Qwen3.8-Max at 46.3%, and Gemini 3.8 Flash at 22.0% against Gemini 3.7 Flash at 36.6%.

Agent system

ModelOrganizationScoreHarness and setupDateSource and notes
AI co-mathematician (gdm-ai-co-mathematician)Google DeepMind75.6%Research agent system; effort, tools and token budget not published2026-05-08 release, entered 2026-06-12Epoch v2, SE 6.7, no log link
Epoch AITier 4 v2, 41 private problemsAccuracyepoch.ai

A research system rather than a bare model, so it is not ranked against the two tables above; same v2 problem set.

The co-mathematician's 75.6% is Google DeepMind's highest Tier 4 number, but Epoch lists no effort setting, scaffold description or log for it. It shows what a Google system can reach with its own agent around the model. It does not tell you what a Gemini API call scores.

Historical progression

The history lives on v1, the 50-problem set Epoch launched in July 2025 and scored on 48 private problems, so one problem is worth 2.08 points. Every v1 score predates the v2 corrections. Dates are the start of Epoch's run.

v1 API runs on Epoch's scaffold

Epoch run dateModel (setting)v1 score
2025-07-01o4-mini (high)6.2%
2025-07-01o3-mini (high)4.2%
2025-07-01Claude Opus 4 (27K thinking)4.2%
2025-07-01o3 (high)2.1%
2025-07-03Gemini 2.5 Pro4.2%
2025-08-11Grok 4 (0709)2.1%
2025-10-09GPT-5 Pro (high)14.6%
2025-11-21Gemini 3 Pro Preview18.8%
2025-12-14GPT-5.2 (xhigh)18.8%
2025-12-16DeepSeek V3.22.1%
2026-02-12Claude Opus 4.6 (max)22.9%
2026-02-19Gemini 3.1 Pro Preview16.7%
2026-03-06GPT-5.4 (xhigh)27.1%
2026-04-23GPT-5.5 Pro pre-release (xhigh)39.6%
2026-06-08Claude Opus 4.8 (max)31.2%
Epoch AIEpoch scaffoldTier 4 v1, 48 private problemsAccuracy

v1 problem set and Epoch API scaffold with settings that differ by row; not comparable to any v2 table.

v1 web-app runs, graded by hand

Epoch run dateModelv1 score
2025-08-01Gemini 2.5 Deep Think10.4%
2025-10-09Grok 4 Heavy2.1%
2025-12-24GPT-5.2 Pro31.3%
2026-03-06GPT-5.4 Pro37.5%
Epoch AIWeb app, graded by handTier 4 v1, 48 private problemsAccuracy

v1 problem set with web search, code execution and a simple prompt; not comparable to the v1 API table or to v2.

v1 agent system

Epoch run dateSystemOrganizationv1 score
2026-05-08AI co-mathematician (gdm-ai-co-mathematician)Google DeepMind47.9%
Epoch AITier 4 v1, 48 private problemsAccuracy

v1 problem set, agent system; not comparable to bare-model rows or to v2.

Epoch's October 2025 Battle Royale post adds context to the early rows. Before the GPT-5 Pro runs, models had solved 8 of the 50 Tier 4 problems at least once, across all attempts. GPT-5 Pro added a ninth, putting the share ever solved at 19%. Five of the eight problems GPT-5 Pro solved at least once were in Epoch's 20-problem holdout. Epoch called GPT-5 Pro's one-problem lead, 13%, over Gemini 2.5 Deep Think "not statistically significant." The post reports 6 of 48 for GPT-5 Pro across web app and API, while the hub lists 14.6%.

Same model on v1 and v2

Model (setting)v1 scorev2 scoreDifference
GPT-5.2 (xhigh)18.8%31.7%+12.9
Claude Opus 4.6 (max)22.9%26.8%+3.9
Gemini 3.1 Pro Preview16.7%26.8%+10.1
GPT-5.4 (xhigh)27.1%49.0%+21.9
Claude Opus 4.8 (max)31.2%56.1%+24.9
Epoch AIEpoch scaffoldTier 4 v1 and v2Accuracy, v2

Same model, effort and Epoch scaffold on both versions; the difference measures the problem-set change, not model progress.

All five models score higher on v2, by 3.9 to 24.9 points. The co-mathematician moved from 47.9% to 75.6%. When a chart online shows a model "jumping" from a v1 number to a v2 number, part or all of that jump is the corrected problem set.

v2 progression by developer

DeveloperModel (setting)Model release datev2 score
OpenAIGPT-5.2 (xhigh)2025-12-1131.7%
OpenAIGPT-5.4 (xhigh)2026-03-0549.0%
OpenAIGPT-5.5 (xhigh)2026-04-2372.5%
OpenAIGPT-5.6 Sol (max)2026-07-0982.9%
OpenAIGPT-6 Astra2026-09-0397.6%
OpenAIGPT-6.1 Sol (max)2026-09-29100.0%
AnthropicClaude Opus 4.6 (max)2026-02-0526.8%
AnthropicClaude Opus 4.7 (max)2026-04-1631.7%
AnthropicClaude Opus 4.8 (max)2026-05-2856.1%
AnthropicClaude Fable 5 (max)2026-06-0990.2%
AnthropicClaude Fable 5.1 (max)2026-09-0187.8%
AnthropicClaude Opus 5.5 (max)2026-09-2295.0%
Epoch AIEpoch scaffoldTier 4 v2, 41 private problemsAccuracy

All v2 on Epoch's scaffold, so rows compare directly; release dates are from Epoch's CSV, not run dates.

This is the progression you can trust, because every row is the same problem set run the same way. OpenAI went from 31.7% to 100.0% across releases under ten months apart. Anthropic went from 26.8% to 95.0% over a slightly shorter span, with the biggest single step from Opus 4.8 to Fable 5, 34.1 points between releases 12 days apart. Fable 5.1 scores 2.4 points below Fable 5, one problem, which is inside both models' standard errors and is not a regression you can measure on this set.

Documented failure modes and limitations

Final answer only. Epoch checks the value answer() returns and does not grade the reasoning. A correct answer reached by a wrong argument scores 1. Tier 4 measures whether a model can land on the right object, not whether its mathematics would survive review.

Errors in the problem set. The v2 update corrected 12 Tier 4 problems and removed 7, and Epoch says it addressed errors in 42% of problems across the whole dataset. Every v1 Tier 4 score was earned on the uncorrected set.

OpenAI funding and access. "FrontierMath was developed with funding from OpenAI, who has exclusive access to a subset of the benchmark," according to Epoch's v2 page. Under Epoch's conflict-of-interest statement, OpenAI commissioned and owns the 300 original problems; Epoch can publish evaluations of any model but cannot publish questions without OpenAI's written permission. Epoch also wrote, "we recognize we have not communicated clearly enough about the relationship between FrontierMath and OpenAI, leading to questions and concerns among contributors, researchers, and the public." For v1 Tier 4, OpenAI had access to 28 of the 48 evaluated problems and their solutions, and Epoch held out the other 20. Epoch has not published that split for v2. OpenAI models hold the top two v2 rows, and without the v2 split the published data cannot show whether access played any part. The one data point Epoch did publish runs the other way for v1: five of the eight problems GPT-5 Pro solved at least once came from the holdout.

No logs for the newest runs. 27 of the 70 v2 rows have no log viewer link in Epoch's CSV, and that includes every run started on or after August 25, 2026: all GPT-6 and GPT-6.1 rows, Claude Fable 5.1, Opus 5.5 and Sonnet 5.5. The top of the current leaderboard is the part nobody outside Epoch can audit.

Token cap. Runs end at 1,000,000 tokens and are forced to submit at 660,000. A model that needs a longer chain scores 0 on that problem. Epoch publishes no per-model token usage, so you cannot tell which misses were cap hits.

Contamination. The problems are private and the answers are built to be unguessable. The FrontierMath paper names unpublished problems and automated verification as its contamination defenses. Epoch's web-app runs allow web search, so published literature can help on those runs even though the problems themselves are not online.

Two protocols. Scaffold runs get a Python tool and no proof assistant. Web-app runs get web search, a simple prompt and hand grading. This page keeps them in separate tables because they are different tests.

Saturation. The top three rows of the frontier table score 95.0% or higher, and nine rows score 80.5% or higher. Tier 4 no longer separates the top OpenAI and Anthropic models. Epoch's newer FrontierMath Erdős set, released September 3, 2026, is nowhere near saturated. GPT-6 Astra (max) scores 2.9% there, and Claude Fable 5.1 and GPT-5.5 score 0.0% on the Epoch hub.

FrontierMath Tiers 1-3

ModelTier 4 v2 (41 private)Tiers 1-3 v2 (285 private)
GPT-6.1 Sol (max)100.0%93.7%
GPT-6 Astra (max)97.6%93.7%
Claude Opus 5.5 (max)95.0%91.2%
Claude Fable 5.1 (max)87.8%90.2%
GPT-6 Sol (max)90.0%89.8%
GPT-5.6 Sol (max)82.9%89.1%
Claude Sonnet 5.5 (max)80.5%88.8%
Claude Opus 5 (max)73.2%85.6%
GPT-5.6 Luna (max)61.0%82.1%
Epoch AIEpoch scaffoldTier 4 v2, 41 private problemsAccuracy, Tier 4 v2

Same Epoch scaffold and models, different problem sets; compare ranks across columns, not raw scores.

Tiers 1-3 has 295 problems, 285 of them private, and GPT-6.1 Sol's 93.7% is 267 of 285. Tier 4 sits below Tiers 1-3 for every model through GPT-5.6 Sol, as you would expect from the harder tier. For GPT-6.1 Sol, Astra, Opus 5.5 and GPT-6 Sol it sits above. Tiers 1-3 ties GPT-6.1 Sol with Astra, while Tier 4 puts 2.4 points between them, one problem. Fable 5.1 and GPT-6 Sol land within half a point of each other on both sets.

Other math and reasoning benchmarks

MATH and GSM8K are public problem sets with short derivations. Tier 4 answers are exact Python objects that cannot be guessed, and the problems take researchers hours to days. A model that scores near the top of MATH tells you little about Tier 4.

GPQA is multiple-choice graduate science. Tier 4 has no choices, so the model has to produce the exact answer instead of ruling out three wrong ones.

Humanity's Last Exam covers expert questions across many fields with a public question set. Epoch's hub lists GPT-6 Astra at 54.8% and Claude Fable 5.1 (xhigh) at 46.5% there, from an external-source table that Epoch did not run. FrontierMath is math-only, private and checked by running code, so the two measure different things even when the same models top both.

For the broader map of where Tier 4 fits, see LLM benchmarks and frontier models.

What it means for teams choosing a model

  1. The top three are a tie on this data. GPT-6.1 Sol (max) at 100.0%, GPT-6 Astra at 97.6% and Claude Opus 5.5 at 95.0% sit inside each other's standard errors. Pick among them on cost, latency and your own evals. None of the three runs has a published log.
  2. Medium effort is enough for Astra. Astra scores 97.56% at medium and 97.6% at high, xhigh and max on Epoch's run. Paying for max buys nothing on Tier 4. Dropping to none costs 14.6 points, six problems.
  3. Size within a family matters more than generation. GPT-6 Luna at 56.1% scores below GPT-5.6 Sol at 82.9%, an older model. Sonnet 5.5 trails Opus 5.5 by 14.5 points. If hard math is the workload, buy the largest tier.
  4. Outside OpenAI and Anthropic, the ceiling is 46.3%. Qwen3.8-Max and Muse Spark 1.3 are the best open or non-US-lab options at 46.3%. Google's best bare model in Epoch's table is Gemini 3.7 Flash at 36.6%. The co-mathematician's 75.6% is a research system, not something you call through an API.
  5. Do not rank inside five points. Fable 5 at 90.2%, GPT-6 Sol at 90.0% and Fable 5.1 at 87.8% are the same score for decision purposes, and Tiers 1-3 says the same thing about Fable 5.1 and GPT-6 Sol.
  6. Use only v2 numbers. v1 scores for the same model and effort run 3.9 to 24.9 points lower. Claude Opus 4.8 scores 31.2% on v1 and 56.1% on v2. Mixing versions in a comparison will mislead you.

Tier 4 measures exact-answer research math with a Python tool and a 30-second runtime. It does not measure proof writing, explaining mathematics to a reader, or general coding. If your workload is one of those, test on that instead. Its saturation at the top also means the benchmark's value for choosing among OpenAI and Anthropic flagships is mostly spent. The gap it shows between those two developers and everyone else is still the clearest signal in the data.

To compare these models across other benchmarks, see the LLM leaderboard. To run the same kind of exact-answer evaluation on your own problems in Klu, see how it handles research workflows.