What is ArXivMath?
ArXivMath is a benchmark of research-level mathematics that MathArena rebuilds every month from new arXiv papers. MathArena comes from SRI Lab at ETH Zurich and INSAIT. The launch post is by Jasper Dekoninck, Tim Gehrunger and Martin Vechev, and it went up on February 4, 2026 with 40 problems drawn from December 2025 and January 2026 papers.
Each release uses papers posted in the previous month, so every question post-dates the training data of the models being tested. Each question has one exact answer, which is a number, a formula or a specific mathematical object. A script or an LLM judge can grade that without a human reading a proof. The model is asked for the answer, not for a proof.
MathArena built it because the existing options all had a hole. Research-level math sets were either private (FrontierMath, MathScienceBench, IMProofBench, EternalMath) or static (BrokenMath, HLE). Contest sets such as AIME and HMMT were close to saturated. ArXivMath publishes every question and every model output, and it gets a fresh set each month.
The newest release on MathArena's leaderboard as of October 6, 2026 is August 2026: 57 questions, published September 15 with a rebuilt pipeline, a new harness and a new grader (August update). GPT-6.1 Sol at max effort tops it at 96.49%. GPT-6 Sol (max) scores 92.98%, GPT-6 Astra (max) 88.60%, Claude Fable 5.1 at high effort 87.72% and Claude Opus 5.5 85.96%. Every model from outside OpenAI and Anthropic scores under 58%.
How the questions are built
The original pipeline ran from the December 2025 release through July 2026. Every month about 4,000 arXiv math papers went through it. GPT-5.2 wrote candidate questions from each abstract. Automated GPT-5.2 filters then removed weak ones, including any whose answer could be guessed from prior work. Human reviewers made the final cut, which left about 40 questions a month. Reviewers kept only questions that were self-contained, had an answer the parser could check, and had an answer other than 0.
Release sizes grew over the year: 17 questions for December 2025, 23 for January 2026, 32 for February, 30 for March, 40 each for April and May, and 48 for June. Anthropic's system cards count June as 49 problems, while MathArena's own June table has 48 problem columns (competitions page, June dataset).
The August release changed how questions are made. GPT-6 Astra now generates them from the paper's full TeX source instead of the abstract, followed by automated verification and human review. The pipeline gives priority to results that refute or resolve an earlier conjecture. Of the 57 August questions, 25 come from refuted or resolved conjectures, and each is worded so that answering with the conjectured value is wrong. A model that recalls the folklore answer instead of working the problem loses the point. MathArena says the September release admits only conjecture-based questions.
How the harness and scoring work
Original releases
Through July 2026, models answered with no tools and no web search. MathArena's platform paper says web search risks contamination, since the source paper is online. MathArena did test a Semantic Scholar literature-search tool with date cutoffs on Gemini-3-Flash, DeepSeek-V3.2 and GPT-5.2 (low). Scores moved by under 3 points, so the tool was dropped. Each model ran with its recommended hyperparameters.
Grading used a rule-based LaTeX parser built on SymPy. The parser wrongly rejected about 1% of responses, so a Gemini-3-Flash judge reviewed every answer the parser marked incorrect or could not parse, and a human checked every answer that judge accepted.
August 2026 release
August turned ArXivMath into an agentic, tool-using benchmark. Each model runs inside its vendor's own coding harness: Antigravity CLI for Gemini, Codex for GPT-6 Astra and Claude Code for Fable 5.1. The harness runs in Docker with no internet and an empty workspace, plus the Python scientific stack and SageMath. Each attempt is capped at 12 hours and $100 of model cost. MathArena ran Gemini-3.8-Flash through Antigravity CLI, OpenCode, Kimi Code and Qwen Code and got essentially identical accuracy and cost, so it treats the choice of harness as neutral.
The prompt asks the model for a best guess when it cannot solve the problem. Anthropic describes MathArena's setting as two attempts per problem graded by an LLM judge (Opus 5.5 and Sonnet 5.5 system cards, Section 8.9).
The SymPy script is gone. Gemini-3.8-Flash grades every answer, and MathArena reports "essentially perfect accuracy in our extensive human validation." When MathArena swapped in GPT-6 Astra as judge, it "found 100% agreement with Gemini-3.8-Flash" on ArXivMath. The 97.5% agreement figure in the same post is for BrokenArXiv, the companion benchmark, not for ArXivMath.
Two details of the August runs affect how to read the table.
Fable 5.1 ran at high effort, not max. MathArena's reason: "this model regularly spends over 512,000 output tokens without making a single tool call, frequently triggering errors in Claude Code. For this reason, we also set its reasoning effort to high instead of max."
The time limit arrived partway through. MathArena added the 12-hour and $100 limits "after running several models, when Fable 5.1 revealed the need for them." The limits come with a short prompt addition that tells the model about the time limit. In preliminary comparisons on Fable 5.1 and Gemini-3.8-Flash, the instruction raised output-token usage by roughly 20-30% "with similar performance." Fable 5.1 received it in roughly half of its ArXivMath runs. MathArena names no other model that got the instruction on ArXivMath, and it does not say which other August rows ran under the limits. It plans to apply the limits to every model in future releases.
What the leaderboard reports
The live leaderboard shows accuracy with a 95% confidence interval, cost in USD for one model run on one problem, average output tokens, average input tokens, average retries and average time per problem. With 57 questions, one question is worth about 1.75 points, and the published intervals run from ±4.78 to ±12.82 points. Gaps of three or four points between neighbors sit inside those intervals.
Versions
| Release | Questions | Published | Question source | Tools | Grader |
|---|---|---|---|---|---|
| Launch, Dec 2025 + Jan 2026 | 40 | 2026-02-04 | GPT-5.2 from abstracts, human review | None | SymPy parser, Gemini-3-Flash fallback, human check |
| Platform paper set, Jan-Apr 2026 | 103 | arXiv 2605.00674 | Same pipeline | None | Same |
| Monthly, Feb to Jul 2026 | 30 to 48 through June | Monthly | Same pipeline | None | Same |
| August 2026 | 57, 25 conjecture-based | 2026-09-15 | GPT-6 Astra from full TeX, human review | Vendor harness, Python stack, SageMath, no internet | Gemini-3.8-Flash judge |
The August release is a different benchmark in practice. It changed the question generator, gave models code execution and replaced the grader. No score from before August belongs on the same axis as an August score.
Current leaderboard (August 2026 release)
Every MathArena row below is MathArena's own run on the 57 August questions with the best-guess prompt and the Gemini-3.8-Flash judge, read from the leaderboard table on October 6, 2026. The table carries no per-row evaluation date. Twelve of the 13 rows carry MathArena's flag for a model released after the competition release; Kimi K3 (Think) is the exception. Cost, tokens and time come from the same table.
This page has no cost or time chart. The one table with named harnesses has four rows that differ on effort and on whether the time limit applied, the Sol table has two rows, and the remaining MathArena rows have no published harness or effort. Anthropic publishes its cost data only as log-scale images.
OpenAI Sol models
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| GPT-6.1 Sol | OpenAI | 96.49% (±4.78) | Max effort; harness and limits not published | Model released 2026-09-29 | MathArena; $0.93 per problem, 63,909 output tokens, 23m 55s per problem |
| GPT-6 Sol | OpenAI | 92.98% (±6.63) | Max effort; harness and limits not published | No eval date published | MathArena; $4.82 per problem, 157,043 output tokens, 40m 38s; 93.0% also quoted in the Sonnet 5.5 system card |
Same effort, questions and grader for both rows; MathArena publishes no harness or limit setting for either, so they do not line up row for row with the named-harness table.
Vendor-native harness named
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | 88.60% (±5.83) | Codex, max effort; limits not stated | No eval date published | MathArena; $3.63 per problem, 27,180 output tokens, time not published; 88.6% also in Opus 5.5 card |
| Claude Fable 5.1 | Anthropic | 87.72% (±6.03) | Claude Code, high effort; time-limit prompt in about half of runs | No eval date published | MathArena; $12.77 per problem, 146,134 output tokens, 64m 26s; 87.7% also in Opus 5.5 card |
| GPT-6 Astra | OpenAI | 85.09% (±6.54) | Codex, low effort; limits not stated | No eval date published | MathArena; $1.22 per problem, 12,664 output tokens, 21m 22s |
| Claude Fable 5.1 | Anthropic | 78.95% (±10.58) | Claude Code, low effort; time-limit prompt in about half of runs | No eval date published | MathArena; $9.15 per problem, 72,681 output tokens, 78m 16s |
Same questions and grader; the rows are not matched on effort or time limits, and each Fable row blends runs with and without the time-limit prompt because MathArena publishes no split.
At the top of this table, Astra (max) leads Fable 5.1 (high) by 0.88 points, half a question, while costing $3.63 per problem against $12.77. Fable spends 146,134 output tokens per problem to Astra's 27,180. MathArena says Fable's cost "far exceeds the estimated 20% increase due to the time-limit instruction," so the prompt addition does not explain the gap.
No harness or effort published
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 5.5 | Anthropic | 85.96% (±9.02) | Label "high"; harness not published | Results listed 2026-09-24 | MathArena; $1.68 per problem, 45,039 output tokens, 34m 53s |
| Qwen3.8-Max | Qwen | 57.89% (±12.82) | No label; harness not published; open weights | No eval date published | MathArena; $5.76 per problem, 297,226 output tokens, 241m 57s |
| Grok 4.7 | xAI | 46.49% (±9.78) | Label "xhigh"; harness not published | No eval date published | MathArena; $7.95 per problem, 178,253 output tokens, 91m 41s |
| Muse Spark 1.3 | Meta AI | 45.61% (±9.91) | No label; harness not published | No eval date published | MathArena; $1.15 per problem, 129,636 output tokens, 36m 32s |
| DeepSeek-V4.1-Flash | DeepSeek | 43.86% (±9.11) | Label "Max"; harness not published; open weights | No eval date published | MathArena; $0.43 per problem, 466,786 output tokens, 153m 47s |
| Gemini 3.8 Flash | 40.35% (±9.05) | Antigravity CLI; no effort label | No eval date published | MathArena; $0.92 per problem, 95,160 output tokens, time not published | |
| Kimi K3 | Moonshot AI | 38.60% (±12.64) | Label "Think"; harness not published; open weights | No eval date published | MathArena; $5.77 per problem, 332,043 output tokens, 240m 54s; the only row not flagged as released after the competition |
Same questions and grader as the tables above; the effort labels are each vendor's own and no harness is published except Gemini's, so these are reported scores, not a ranking against the other MathArena tables.
Anthropic internal runs (vendor-reported)
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 5.5 | Anthropic | 96.9% with tools, 91.2% without | Anthropic harness, max effort, four attempts averaged, code sandbox | 2026-09-22 | Opus 5.5 system card, Section 8.9 |
| Claude Sonnet 5.5 | Anthropic | 95.2% with tools, 86.8% without | Same; pre-release serving of the launch checkpoint | 2026-09-28 | Sonnet 5.5 system card, Section 8.9 |
| Claude Fable 5.1 | Anthropic | 92.1% with tools, 82.9% without | Same | 2026-09-22 | Opus 5.5 and Sonnet 5.5 cards quote the same value |
| Claude Opus 5 | Anthropic | 90.4% with tools, 78.1% without | Same | 2026-09-22 | Opus 5.5 and Sonnet 5.5 cards quote the same value |
Anthropic-run with its own prompt, judge and harness on the same 57 questions, with no safeguards classifiers; Anthropic says MathArena's numbers "are closest to our 'with tools' results but are not directly comparable."
"With tools" here means a code-execution sandbox with no internet, the closest match to MathArena's August setup. Every Claude model in both places scores higher in Anthropic's table: Opus 5.5 at 96.9% against 85.96% on MathArena, Fable 5.1 at 92.1% against 87.72%. Anthropic runs at max effort and averages four attempts, and MathArena ran Fable at high. Use Anthropic's table to compare Claude models with each other, and MathArena's to compare across vendors.
June 2026 release
MathArena's live June table is not reproduced here. Its rows mix run protocols, with average times from 0m 20s to 27m 56s and 21 of 29 rows flagged as released after the competition, and MathArena publishes no per-row protocol. Three sourced statements cover June:
- MathArena's own run on the pre-August no-tool pipeline: "GPT-6 Astra essentially saturated both ArXivMath (94%) and BrokenArXiv (96%)." MathArena adds that with coding tools Astra "would have solved both questions it missed on ArXivMath June."
- MathArena values quoted in the Fable 5.1 and Mythos 5.1 system card, Section 8.10, with SymPy-era grading: GPT-5.6 Sol (max) 86.73%, Gemini 3.1 Pro Preview 65.99%.
- Anthropic's internal run in the same card, four runs per problem on 49 problems: Claude Mythos 5.1 at max effort scores 91.33% without tools and 93.88% with tools. Mythos 5.1 is not in the August results.
The saturation in June is why MathArena rebuilt the benchmark for August.
Historical progression
Each row is one release, one run owner and one protocol. Rows with different protocols are not a trend line, even for the same model family.
| Release | Questions | Run by | Protocol | Result | Source |
|---|---|---|---|---|---|
| Launch, Dec 2025 + Jan 2026 | 40 | MathArena | No tools, SymPy with Gemini-3-Flash fallback | GPT-5.2 60%, Gemini-3-Pro 56.5% | Launch post, 2026-02-04 |
| Platform paper set, Jan-Apr 2026 | 103 | MathArena | Same | GPT-5.5 74.1%; other models 45 to 67%; $0.08 to $3.16 per model | Platform paper |
| Mar + Apr 2026 | 71 | MathArena | Same | GPT-5.5 (xhigh) 71.48%, Gemini 3.1 Pro Preview 64.79% | MathArena values quoted in Opus 4.8 system card, Section 8.8, footnote 38 |
| Mar + Apr 2026 | 71 | Anthropic (vendor-reported) | Extended thinking, Anthropic's protocol | Claude Opus 4.8 71.82% | Opus 4.8 system card, Section 8.8 |
| June 2026 | 48 on MathArena, 49 per Anthropic | MathArena | Pre-August no-tool pipeline | GPT-6 Astra 94% | August update |
| June 2026 | 48 or 49 | MathArena, quoted by Anthropic | SymPy-era grading | GPT-5.6 Sol (max) 86.73%, Gemini 3.1 Pro Preview 65.99% | Fable 5.1 and Mythos 5.1 system card, Section 8.10 |
| June 2026 | 49 | Anthropic (vendor-reported) | Four runs per problem | Claude Mythos 5.1 (max) 91.33% without tools, 93.88% with tools | Fable 5.1 and Mythos 5.1 system card, Section 8.10 |
| August 2026 | 57 | MathArena | New pipeline, vendor harness with code, Gemini-3.8-Flash judge | GPT-6.1 Sol (max) 96.49% | Leaderboard, read 2026-10-06 |
| August 2026 | 57 | Anthropic (vendor-reported) | Four attempts, max effort, code sandbox | Claude Opus 5.5 96.9% with tools | Opus 5.5 system card, Section 8.9 |
Historical rows use older question sets, older pipelines and, before August, no tools; they are not comparable to the August tables or to each other across releases.
The pattern worth taking from this table is about the benchmark, not the models. By June, MathArena itself called the no-tool pipeline essentially saturated, with GPT-6 Astra at 94%. MathArena answered by changing what it measures. The August release asks harder questions, built around refuted conjectures, and gives models a sandbox. GPT-6.1 Sol's 96.49% on that release shows the top of the board is crowded again.
Failure modes and limitations
A narrow slice of research. The benchmark checks final answers, not proofs, exposition or problem choice. MathArena warns that its results "should not be interpreted as evidence that models can autonomously write 60% of recent mathematical papers." Anthropic calls ArXivMath "a narrow proxy for actual research ability" (Opus 4.8 system card, Section 8.8).
Leakage from prior work. A PhD mathematician reviewed about 20 launch questions and found about 30% answerable from earlier work the paper cited. Those were almost all among the easiest items, and MathArena kept them in the set.
Training-data overlap. Anthropic states that Mythos 5.1's training data "may overlap" the June 2026 abstract period and "some contamination cannot be ruled out" (Fable 5.1 and Mythos 5.1 system card, Section 8.10). The monthly refresh is the main defense against this, and here a vendor says the defense was not airtight.
Papers written with AI help. MathArena dropped its "AI usage" paper filter in August because it assumes most papers now involve some AI use. Its own assessment: "we are building benchmarks from questions that may have been answered using the models themselves, so the model used most in practice may have an advantage," and "there is no good way around this issue."
Wrong source papers. arXiv papers are not all peer reviewed, so some reference answers could be wrong. MathArena judges this risk small, because questions that models fail show answers spread across many values rather than one consistent wrong answer.
Invented citations. On one graph-theory question about feedback vertex sets of digraphs, the correct answer is 3/7. Of all attempts, 89.6% answered 2/5 and 10.4% answered 3/7, and models cited theorems that do not exist. This is the hallucination pattern a research team would face in practice: a confident wrong answer backed by a fake reference.
Harness failures. Some August runs ended with no final answer. Fable 5.1 once signaled completion too early on BrokenArXiv, and Qwen3.8-Max once returned an API error. MathArena did not rerun those attempts. Fable 5.1's habit of passing 512,000 output tokens without a tool call "frequently" triggered errors in Claude Code.
Calibration. In every question GPT-6 Astra missed in August, the model stated its uncertainty and said it could not prove its answer. For a team, that is useful behavior: the misses announce themselves.
Saturation. MathArena also built 58 physics questions (ArXivPhys) and 46 computer science questions (ArXivCS). GPT-6 Astra "solved all but three" across the two sets, so MathArena did not adopt either as a benchmark.
A mixed August set. Only 25 of the 57 August questions use the conjecture-based anti-recall design. The other 32 do not.
Missing setup data. MathArena does not publish a per-model harness or effort for Qwen3.8-Max, Muse Spark 1.3, Grok 4.7, DeepSeek-V4.1-Flash, Kimi K3, GPT-6 Sol, GPT-6.1 Sol or Opus 5.5. It publishes no average time for Astra (max) or Gemini 3.8 Flash. It does not say which August rows ran under the 12-hour and $100 limits.
How ArXivMath compares to related benchmarks
Humanity's Last Exam is a static set of expert-written questions. Once published, it can be memorized. ArXivMath swaps in new questions each month, and MathArena names HLE as one of the static benchmarks it set out to replace.
FrontierMath Tier 4 is private, so outsiders cannot audit a score. ArXivMath publishes every question and model output through its Hugging Face datasets and outputs links on the competitions page.
BrokenArXiv is MathArena's companion benchmark, with 56 August questions scored on a 0 to 3 scale. The model receives a false claim taken from a refuted conjecture and earns credit for flagging it as false. On August, Astra scores 81.94%, Fable 5.1 79.76%, Qwen3.8-Max 69.64% and Kimi K3 61.90%. BrokenArXiv tests whether a model catches a false premise. ArXivMath tests whether it can derive the answer. A research assistant needs both.
MATH, AIME and HMMT use contest problems with known solution paths, and MathArena states that models score near-perfect on them. A strong contest score says little about how a model handles a result posted to arXiv last month. For graduate-level science Q&A, see GPQA.
Two other research benchmarks on this site are Terminal-Bench Science and DRACO. For the wider map, see LLM benchmarks, LLM evaluation and frontier models.
What it means for teams choosing a model
- GPT-6.1 Sol is the default for hard math on MathArena's numbers. It scores 96.49% at $0.93 per problem, against 92.98% at $4.82 for GPT-6 Sol at the same effort. It costs 19% as much and scores 3.51 points higher. The intervals (±4.78 and ±6.63) overlap, so the cost gap is the firmer finding. It also uses 63,909 output tokens per problem to GPT-6 Sol's 157,043.
- Budget Fable 5.1 at its own measured cost. GPT-6 Astra (max) costs $3.63 per problem and Fable 5.1 (high) costs $12.77, 3.5 times as much, for scores within a point of each other. If you plan a Fable workload from Astra's numbers, you will underestimate spend.
- Astra at low effort is the cheap option. Low scores 85.09% at $1.22 against 88.60% at $3.63 for max, 3.51 points for roughly three times the cost, and the intervals overlap. Moving Astra from low to max test-time compute costs $2.41 more per problem for those 3.51 points.
- Among Claude models, Opus 5.5 leads on Anthropic's own protocol. It scores 96.9% with tools and 91.2% without, ahead of Sonnet 5.5 by 1.7 points with tools and 4.4 points without. Fable 5.1 (92.1 / 82.9) and Opus 5 (90.4 / 78.1) follow.
- Give Claude models a code sandbox. In Anthropic's runs, the sandbox adds 5.7 points for Opus 5.5, 8.4 for Sonnet 5.5, 9.2 for Fable 5.1 and 12.3 for Opus 5. In this data, the lower a model's no-tools score, the bigger its gain.
- Open-weight models are reported, not ranked. Qwen3.8-Max scores 57.89%, DeepSeek-V4.1-Flash 43.86% and Kimi K3 38.60%, but MathArena publishes no harness or effort for these rows. Qwen3.8-Max and Kimi K3 also average over 240 minutes per problem, the longest times on the board. If latency matters, that alone rules them out for interactive math work.
ArXivMath tells you which model gets the exact answer to a fresh research question with a sandbox and a generous time budget. It does not tell you which model writes a correct proof or explains its reasoning to a colleague. If that is your workload, test it directly.
The LLM leaderboard puts these models side by side across other benchmarks. To run the same kind of exact-answer evaluation on your own problem set in Klu, see how it handles research workflows.