What is OpenAI MRCR v2?
OpenAI MRCR, short for Multi-Round Co-reference Resolution, tests whether a model can tell identical requests apart inside a long conversation. OpenAI published the dataset as openai/mrcr on Hugging Face on April 12, 2025, under an MIT license.
Each prompt is a synthetic chat between a user and an assistant. Somewhere in it, the user asks for the same thing several times, for example "write a poem about tapirs." Those identical requests sit among distractor requests on other topics. At the end, the model gets an instruction such as "return the 2nd poem about tapirs," plus a random alphanumeric hash it has to put at the start of its answer.
The 8-needle setting is the hardest one. Eight copies of the same request produce eight candidate answers that all look right. The model has to pick the correct one by position and then reproduce it word for word.
A needle-in-a-haystack test hides one distinctive fact in filler, and a model can pass by matching a single string. MRCR gives every needle the same shape, so string matching gets the model nowhere. It has to keep count across the whole context window and then copy a long passage back accurately. That makes MRCR one of the few public tests that checks whether a model actually uses a 1M-token window instead of merely accepting one.
The current picture is thin. The newest primary-sourced scores come from Anthropic's Claude Sonnet 4.6 system card, published February 17, 2026, where Claude Opus 4.6 scores 93.0% at 256K tokens and 76.0% at 1M at max effort. No primary document I could read gives an MRCR score for GPT-6 Astra, Claude Opus 5.5, Claude Fable 5.1, Claude Mythos 5.1 or Gemini 4.
How the tasks and harness work
The dataset has 2,400 rows. They split across three needle counts, 2, 4 and 8, and eight length bins that double from 4,096 tokens up to 1,048,576. Every combination of needle count and bin has 100 samples, so a single bin at 8 needles is a 100-problem test.
Vendors label bins by their upper bound. In Anthropic's card, the "256K" column covers prompts above 128K and up to 256K tokens, and the "1M" column covers prompts above 524K and up to 1,024K. Every table on this page reports scores per bin, so a model's 256K and 1M numbers are two separate measurements, not one score at two settings.
There are no tools. The model receives the prompt and returns text in a single call, so no time or step limit applies. The variable that changes between reports is reasoning effort, and each lab sets its own. Anthropic reports its models at a 64k extended thinking budget and at max effort with adaptive thinking. The Gemini rows in Anthropic's table ran at high thinking and GPT-5.2 at xhigh. No lab shares a common prompt harness with another, which is why the leaderboard below splits into separate tables.
Google describes its own MRCR-V2 in Table 11 of the Gemini 2.5 technical report. Google's version moved from 4 needles in MRCR-V1 to 8 and added a style parameter, so a request reads "write a poem about penguins in an archaic style." The model now has to match topic and style, not topic alone. Google reports two variants, an average over contexts up to 128K and a pointwise score at exactly 1M.
The 1M bin runs into real API limits. Anthropic's card says tokenizer differences push some 1M-bin problems past the 1,000,000-token Claude API window, so Anthropic ran internal configurations beyond that limit and published a fitting subset for Sonnet 4.6 only. OpenAI's GPT-6 Astra model page lists a 1,050,000-token window and a 922,000-token maximum input, which sits below the bin's 1,048,576-token ceiling. OpenAI has not published how samples above 922K tokens are handled.
How scoring works
Scoring has two steps. First, the response must begin with the required hash. If it doesn't, the sample scores 0 whatever the rest says. Second, the grader strips the hash and compares the remaining text with the ground-truth answer using Python's difflib SequenceMatcher ratio, which returns a number between 0 and 1. No LLM judge is involved. Anthropic's card calls the averaged result "Mean Match Ratio."
That design has two effects on how to read a score. A 76% result does not mean the model got 76% of samples exactly right. It means the average similarity between its answers and the ground truth was 0.76, with partial credit for near misses. And a model that drops the hash gets nothing even when it found the right poem, so formatting discipline and retrieval land in the same number.
Versions and the December 2025 fix
OpenAI released the dataset on April 12, 2025, and shipped a bugfix on December 5, 2025. The dataset card says "~10% of datapoints contain too many target needles, and ~5% contain incorrect ground truth." Corrected samples carry a date_added field. Anthropic's card states that its runs used the OpenAI version "with the v2 fix introduced on December 5, 2025." Any score on the OpenAI dataset published before that date ran on the unfixed data.
Google's MRCR-V2 is a separate implementation. The Gemini 2.5 report does not say whether it used OpenAI's published dataset or whether Google reran its scores after the fix. Both lines trace back to Vodrahalli et al., Michelangelo, published in 2024 and cited by both Anthropic and Google.
Current leaderboard
Every score below comes from Anthropic's February 2026 card or Google's Gemini 2.5 report. OpenAI's API model pages for GPT-6 Astra, GPT-5.6 Sol and GPT-6.1 Sol publish context limits and prices but no MRCR score, and the GPT-6 Astra system card has none, so OpenAI's current models are absent. Rows marked independent come from contextarena.ai as reproduced in Anthropic's card. Everything else is vendor-reported.
No source publishes cost or time per MRCR sample for any model, so this page has no cost or latency chart.
Anthropic internal runs, max effort
Anthropic card Table 2.16.A, average of 5 trials at default sampling. Comparable only within this table, not with the 64k-thinking runs or the contextarena.ai rows.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 4.6 | Anthropic | 93.0% at 256K; 76.0% at 1M | Adaptive thinking, max effort; Anthropic's protocol on the fixed dataset; 5-trial average | 2026-02-17 | Sonnet 4.6 system card, Table 2.16.A. The 1M run went beyond the public API window, and the card gives no fitting-subset score |
| Claude Sonnet 4.6 | Anthropic | 90.3% at 256K; 65.8% at 1M | Same as above | 2026-02-17 | Table 2.16.A. Footnote 15 says the 1M result "is not reproducible via the public API." The 29-problem subset that fits scores 77.8% |
Opus 4.6 leads Sonnet 4.6 by 2.7 points at 256K and by 10.2 points at 1M, so the gap between the two models widens as prompts get longer. Sonnet's 77.8% subset score is higher than its full 1M score, but it covers 29 problems that fit the window, not the 100-problem bin. Both 1M numbers include problems beyond the public API window.
Anthropic internal runs, 64k extended thinking
Same card, bins and 5-trial average as the max-effort table, with a fixed 64k thinking budget instead. Do not mix these rows with the max-effort rows.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 4.6 | Anthropic | 91.9% at 256K; 78.3% at 1M | 64k extended thinking; Anthropic's protocol on the fixed dataset; 5-trial average | 2026-02-17 | Sonnet 4.6 system card, Table 2.16.A. Same 1M window overflow as the max-effort run; no fitting-subset score |
| Claude Sonnet 4.6 | Anthropic | 90.6% at 256K; 65.1% at 1M | Same as above | 2026-02-17 | Table 2.16.A. The 1M result is not reproducible via the public API. The 54-problem subset that fits scores 71.3% |
| Claude Sonnet 4.5 | Anthropic | 10.8% at 256K; 18.5% at 1M | Same as above | 2026-02-17 | Table 2.16.A. The card gives no subset figure for this 1M cell |
The jump from Sonnet 4.5 to Sonnet 4.6 is the largest move on this page, from 10.8% to 90.6% at 256K under one protocol. Opus 4.6 at the 64k budget scores 78.3% at 1M, 2.3 points above Opus 4.6 at max effort. The max setting did not beat a fixed 64k budget for Opus on the longest bin, so "highest effort" is not a safe default for this workload.
Independent runs on contextarena.ai
Scores from contextarena.ai as reproduced in Anthropic's card, footnote 11. This is a different harness from Anthropic's internal runs, and the card does not state sampling or trial count, so compare these rows only with each other.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| GPT-5.2 | OpenAI | 63.9% at 256K; no 1M score | xhigh thinking; contextarena.ai harness | Quoted 2026-02-17 | Sonnet 4.6 system card, Table 2.16.A and footnotes 11 and 13. Independent run. No 1M score with a 400K window. The card also prints OpenAI's self-reported 70.0%, with effort and bin unstated |
| Gemini 3 Flash | 58.5% at 256K; 32.6% at 1M | High thinking; contextarena.ai harness | Quoted 2026-02-17 | Table 2.16.A, footnote 11. Independent run | |
| Gemini 3 Pro | 45.4% at 256K; 24.5% at 1M | High thinking; contextarena.ai harness | Quoted 2026-02-17 | Table 2.16.A, footnote 11. Independent run |
GPT-5.2 leads this table at 256K, 5.4 points ahead of Gemini 3 Flash. The more surprising result is inside Google's lineup. Flash beats Pro on both bins under the same harness and thinking setting, by 13.1 points at 256K and 8.1 points at 1M. If your workload is long-context retrieval on Gemini 3, the cheaper model is the stronger one in this run.
Google's MRCR-V2 runs from the Gemini 2.5 report
Google's own implementation, Table 4 of arXiv 2507.06261. The up-to-128K score is an average across lengths and the 1M score is pointwise. The report does not state the dataset source or whether it reran after the December 2025 fix, so these rows are not comparable to any other table.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Gemini 2.5 Pro | 58.0% up to 128K; 16.4% at 1M | Google MRCR-V2; no setting stated | 2025-07-07, v6 2025-12-19 | Gemini 2.5 report, Table 4. Vendor run | |
| OpenAI o3 | OpenAI | 57.1% up to 128K; no 1M score | Google MRCR-V2; column header "o3 high" | 2025-07-07, v6 2025-12-19 | Table 4, Google's comparison table |
| Claude 4 Sonnet | Anthropic | 39.1% up to 128K; no 1M score | Google MRCR-V2; no setting stated | 2025-07-07, v6 2025-12-19 | Table 4, Google's comparison table |
| OpenAI o4-mini | OpenAI | 36.3% up to 128K; no 1M score | Google MRCR-V2; column header "o4-mini high" | 2025-07-07, v6 2025-12-19 | Table 4, Google's comparison table |
| Grok 3 Beta | xAI | 34.0% up to 128K; no 1M score | Google MRCR-V2; no setting stated | 2025-07-07, v6 2025-12-19 | Table 4, Google's comparison table |
Gemini 2.5 Pro edges o3 by 0.9 points on the up-to-128K average and is the only model in the table with a 1M score. Its own score falls from 58.0% to 16.4%, a 41.6-point drop between the two variants in the same report. Table 4 also lists Claude 4 Opus at 16.1% up to 128K, but its caption marks that run as having no thinking and API refusals, so I left it out of the ranking. The report's DeepSeek R1 0528 column has no MRCR-V2 entry.
Historical progression
Each row below uses its own protocol. Read the scores down a single row's protocol, never across rows.
| Date | Milestone | Figure | Protocol | Source |
|---|---|---|---|---|
| 2024 | Michelangelo paper introduces the MRCR family | No scores used here | Vodrahalli et al. | arXiv 2409.12640 |
| 2025-04-12 | OpenAI publishes openai/mrcr | 2,400 rows at 2, 4 and 8 needles | OpenAI dataset, original release | Hugging Face card |
| 2025-07-07 | Gemini 2.5 technical report | Gemini 2.5 Pro 58.0% up to 128K, 16.4% at 1M | Google MRCR-V2, 8 needles with style key | arXiv 2507.06261, Tables 4 and 11 |
| 2025-12-05 | OpenAI dataset bugfix | About 10% of rows had too many needles, about 5% wrong ground truth | OpenAI dataset, fixed | Hugging Face card |
| 2026-02-17 | Anthropic Sonnet 4.6 system card | 1M bin at 64k thinking: Sonnet 4.5 18.5%, Sonnet 4.6 65.1%, Opus 4.6 78.3% | Anthropic internal, fixed OpenAI dataset | Sonnet 4.6 system card, Table 2.16.A |
The only multi-model series under one protocol is Anthropic's. At 64k thinking, the 256K bin goes from Sonnet 4.5 at 10.8% to Sonnet 4.6 at 90.6% to Opus 4.6 at 91.9%. The 1M bin goes from 18.5% to 65.1% to 78.3%, with the window-overflow caveat on every 1M cell. None of the sources I read give release dates for these three models, so this is a progression by model generation, not by calendar.
Gemini 2.5 Pro's 16.4% at 1M and Opus 4.6's 78.3% at 1M are not a measure of progress. They come from different implementations, and one predates the dataset fix. No primary source covers Gemini 3 versions over time or any Claude, Gemini or OpenAI MRCR score newer than February 17, 2026.
Failure modes and limitations
A missing hash scores zero. The grader gives 0 to any response that doesn't begin with the required hash. A model that retrieves the right passage but opens with a preamble fails the sample, so MRCR counts instruction-following slips as retrieval failures.
Partial credit inflates the headline. SequenceMatcher ratio rewards near matches. A score is mean similarity, not the share of answers reproduced exactly, and a model that grabs the wrong poem can still collect some credit for overlapping text.
Data bugs before December 5, 2025. About 10% of rows held too many target needles and about 5% held wrong ground truth. Any OpenAI-dataset score published before the fix carries that noise.
The 1M bin overflows real windows. Tokenizer differences push some 1M problems past the 1,000,000-token Claude API limit, and Anthropic's headline 1M numbers come from internal runs it could not reproduce through the public API. On the OpenAI side, Astra caps input at 922,000 tokens against a bin ceiling of 1,048,576.
Labs report the same model differently. Anthropic's card prints GPT-5.2 at 256K as 63.9% from contextarena.ai at xhigh and 70.0% as self-reported by OpenAI. That is a 6.1-point gap for one model, and the card states neither the effort nor the bin behind the 70.0%.
Effort settings move the result. Opus 4.6 scores 78.3% at 1M with 64k thinking and 76.0% at max effort. Sonnet 4.6 at 256K scores 90.6% and 90.3% across the same two settings. Any report without an effort setting is incomplete.
Little independent verification. The only independent runs in the primary sources cover Gemini 3 Pro, Gemini 3 Flash and GPT-5.2, via contextarena.ai as quoted by Anthropic. Every other score is vendor-reported.
Public data, no contamination audit. The dataset is public under MIT and the prompts are synthetic. No primary source reports a contamination audit for MRCR.
Narrow domain. Every needle is a writing request, such as a poem or a blog post, inside a synthetic chat. MRCR says nothing direct about retrieval from contracts, codebases or spreadsheets.
How it compares to related benchmarks
Needle in a haystack is the baseline MRCR makes harder. One distinctive fact in filler can be found by matching one string. MRCR's identical needles force the model to track order and reproduce text. Even the leader in Anthropic's runs, Opus 4.6, drops from 93.0% at 256K to 76.0% at 1M on MRCR at max effort.
GraphWalks is the closest sibling. Each prompt is a directed graph of hexadecimal-hash nodes, and the model runs a breadth-first search or a parents lookup across it. Anthropic runs it next to MRCR in the same card, at 256K and 1M with 100 problems each, scored by F1 over a 5-trial average. Footnote 16 says half the 1M problems exceed the public API limit. GraphWalks tests multi-hop reasoning over the context. MRCR tests ordered retrieval and exact reproduction. The two rank the same pair of models in opposite order.
| Model | MRCR v2 8-needle, 1M, max effort | GraphWalks BFS, 1M, max effort | Source |
|---|---|---|---|
| Claude Opus 4.6 | 76.0% | 38.7% | Sonnet 4.6 system card, Tables 2.16.A and 2.16.B |
| Claude Sonnet 4.6 | 65.8% | 73.8% | Sonnet 4.6 system card, Tables 2.16.A and 2.16.B |
Opus 4.6 leads by 10.2 points on MRCR and trails by 35.1 points on GraphWalks BFS. At 64k thinking the BFS 1M scores are Sonnet 4.6 68.4%, Opus 4.6 41.2% and Sonnet 4.5 25.6%. The pattern holds at shorter lengths on the 256K BFS subset, where Sonnet 4.6 scores 72.8% at 64k and 74.5% at max against Opus 4.6's 61.5% and 61.1%. On the 256K parents subset the models converge, with Sonnet 4.6 at 96.9% and 97.9%, Opus 4.6 at 95.1% and 95.4%, and Sonnet 4.5 at 81.0%. A team choosing a long-context model needs both numbers.
LOFT, from Lee et al. in 2024, tests hard retrieval over long corpora and appears in the same Gemini 2.5 report. Gemini 2.5 Pro scores 87.0% up to 128K and 69.8% at 1M on LOFT, against 58.0% and 16.4% on MRCR-V2. That is a 53.4-point gap at 1M on the same model in the same report. Finding relevant content in a corpus and reproducing the n-th of eight lookalikes are different skills.
MRCR also anchors discussions of context window expansion. A 1M-token window is a spec-sheet number. Astra's 1,050,000-token window with a 922,000-token input cap shows that even the spec has fine print, and MRCR is how you measure what the model does with the space. When the answer is "not enough," retrieval-augmented generation is the alternative to putting 500K tokens in the prompt. The effort results above connect to test-time compute. Max effort did not beat a 64k thinking budget for Opus 4.6 at 1M. For the general picture, see LLM benchmarks and LLM evaluation.
What it means for teams choosing a model
You have no MRCR number for today's flagships. No primary-sourced MRCR score exists for GPT-6 Astra, GPT-5.6 Sol, GPT-6.1 Sol, Claude Opus 5.5, Claude Fable 5.1, Claude Mythos 5.1 or Gemini 4. If you are choosing among current frontier models for long-context work, you need your own MRCR-style test at your real prompt length.
Among the models with sourced scores, Opus 4.6 leads on ordered retrieval. In Anthropic's max-effort runs, Opus 4.6 scores 93.0% at 256K and 76.0% at 1M, ahead of Sonnet 4.6 by 2.7 and 10.2 points. Both 1M numbers include problems above the public API's 1M-token limit, and only Sonnet 4.6 has a fitting-subset score, 77.8% on 29 problems.
Match the benchmark to the job. If your workload is pulling back the exact text of one item among many similar ones, such as the third version of a clause or a particular reply in a long thread, MRCR is the right signal and Opus 4.6 is ahead. If it is chaining facts across a long input, GraphWalks is the right signal and Sonnet 4.6 is ahead by 35.1 points at 1M.
Long prompts carry a surcharge. OpenAI lists GPT-6 Astra at $10 per 1M input tokens and $50 per 1M output, with 2x input and 1.5x output pricing above 272K input tokens. A single 524,288-token prompt costs at least $10.49 before any output, and a prompt at the 922,000-token cap costs at least $18.44. GPT-5.6 Sol lists $4 and $20, and GPT-6.1 Sol $2 and $10, with the same surcharge. Test at real length on a small sample before you commit a pipeline to it.
Don't rank across harnesses. Gemini 3 Flash scores 58.5% at 256K and 32.6% at 1M on contextarena.ai at high thinking, ahead of Gemini 3 Pro on both bins. Opus 4.6's max-effort numbers are 34.5 and 43.4 points higher, but they come from Anthropic's internal harness, so those gaps are not a like-for-like ranking.
The LLM leaderboard tracks the broader set of scores. To run the same kind of test on your own long documents and threads, you can build it in Klu for a knowledge assistant workflow at your real prompt lengths.