OpenAI MRCR v2 (Multi-Round Co-reference Resolution)

A long-context benchmark that hides identical requests in a synthetic conversation of up to 1M tokens and asks the model to reproduce the answer to one of them exactly

Long context and multimodal

Stephen M. Walker II · Co-Founder / CEO · October 6, 2026 · 12 min read

What is OpenAI MRCR v2?

OpenAI MRCR, short for Multi-Round Co-reference Resolution, tests whether a model can tell identical requests apart inside a long conversation. OpenAI published the dataset as openai/mrcr on Hugging Face on April 12, 2025, under an MIT license.

Each prompt is a synthetic chat between a user and an assistant. Somewhere in it, the user asks for the same thing several times, for example "write a poem about tapirs." Those identical requests sit among distractor requests on other topics. At the end, the model gets an instruction such as "return the 2nd poem about tapirs," plus a random alphanumeric hash it has to put at the start of its answer.

The 8-needle setting is the hardest one. Eight copies of the same request produce eight candidate answers that all look right. The model has to pick the correct one by position and then reproduce it word for word.

A needle-in-a-haystack test hides one distinctive fact in filler, and a model can pass by matching a single string. MRCR gives every needle the same shape, so string matching gets the model nowhere. It has to keep count across the whole context window and then copy a long passage back accurately. That makes MRCR one of the few public tests that checks whether a model actually uses a 1M-token window instead of merely accepting one.

The current picture is thin. The newest primary-sourced scores come from Anthropic's Claude Sonnet 4.6 system card, published February 17, 2026, where Claude Opus 4.6 scores 93.0% at 256K tokens and 76.0% at 1M at max effort. No primary document I could read gives an MRCR score for GPT-6 Astra, Claude Opus 5.5, Claude Fable 5.1, Claude Mythos 5.1 or Gemini 4.

How the tasks and harness work

The dataset has 2,400 rows. They split across three needle counts, 2, 4 and 8, and eight length bins that double from 4,096 tokens up to 1,048,576. Every combination of needle count and bin has 100 samples, so a single bin at 8 needles is a 100-problem test.

Vendors label bins by their upper bound. In Anthropic's card, the "256K" column covers prompts above 128K and up to 256K tokens, and the "1M" column covers prompts above 524K and up to 1,024K. Every table on this page reports scores per bin, so a model's 256K and 1M numbers are two separate measurements, not one score at two settings.

There are no tools. The model receives the prompt and returns text in a single call, so no time or step limit applies. The variable that changes between reports is reasoning effort, and each lab sets its own. Anthropic reports its models at a 64k extended thinking budget and at max effort with adaptive thinking. The Gemini rows in Anthropic's table ran at high thinking and GPT-5.2 at xhigh. No lab shares a common prompt harness with another, which is why the leaderboard below splits into separate tables.

Google describes its own MRCR-V2 in Table 11 of the Gemini 2.5 technical report. Google's version moved from 4 needles in MRCR-V1 to 8 and added a style parameter, so a request reads "write a poem about penguins in an archaic style." The model now has to match topic and style, not topic alone. Google reports two variants, an average over contexts up to 128K and a pointwise score at exactly 1M.

The 1M bin runs into real API limits. Anthropic's card says tokenizer differences push some 1M-bin problems past the 1,000,000-token Claude API window, so Anthropic ran internal configurations beyond that limit and published a fitting subset for Sonnet 4.6 only. OpenAI's GPT-6 Astra model page lists a 1,050,000-token window and a 922,000-token maximum input, which sits below the bin's 1,048,576-token ceiling. OpenAI has not published how samples above 922K tokens are handled.

How scoring works

Scoring has two steps. First, the response must begin with the required hash. If it doesn't, the sample scores 0 whatever the rest says. Second, the grader strips the hash and compares the remaining text with the ground-truth answer using Python's difflib SequenceMatcher ratio, which returns a number between 0 and 1. No LLM judge is involved. Anthropic's card calls the averaged result "Mean Match Ratio."

That design has two effects on how to read a score. A 76% result does not mean the model got 76% of samples exactly right. It means the average similarity between its answers and the ground truth was 0.76, with partial credit for near misses. And a model that drops the hash gets nothing even when it found the right poem, so formatting discipline and retrieval land in the same number.

Versions and the December 2025 fix

OpenAI released the dataset on April 12, 2025, and shipped a bugfix on December 5, 2025. The dataset card says "~10% of datapoints contain too many target needles, and ~5% contain incorrect ground truth." Corrected samples carry a date_added field. Anthropic's card states that its runs used the OpenAI version "with the v2 fix introduced on December 5, 2025." Any score on the OpenAI dataset published before that date ran on the unfixed data.

Google's MRCR-V2 is a separate implementation. The Gemini 2.5 report does not say whether it used OpenAI's published dataset or whether Google reran its scores after the fix. Both lines trace back to Vodrahalli et al., Michelangelo, published in 2024 and cited by both Anthropic and Google.

Current leaderboard

Every score below comes from Anthropic's February 2026 card or Google's Gemini 2.5 report. OpenAI's API model pages for GPT-6 Astra, GPT-5.6 Sol and GPT-6.1 Sol publish context limits and prices but no MRCR score, and the GPT-6 Astra system card has none, so OpenAI's current models are absent. Rows marked independent come from contextarena.ai as reproduced in Anthropic's card. Everything else is vendor-reported.

No source publishes cost or time per MRCR sample for any model, so this page has no cost or latency chart.

Anthropic internal runs, max effort

Anthropic card Table 2.16.A, average of 5 trials at default sampling. Comparable only within this table, not with the 64k-thinking runs or the contextarena.ai rows.

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 4.6Anthropic93.0% at 256K; 76.0% at 1MAdaptive thinking, max effort; Anthropic's protocol on the fixed dataset; 5-trial average2026-02-17Sonnet 4.6 system card, Table 2.16.A. The 1M run went beyond the public API window, and the card gives no fitting-subset score
Claude Sonnet 4.6Anthropic90.3% at 256K; 65.8% at 1MSame as above2026-02-17Table 2.16.A. Footnote 15 says the 1M result "is not reproducible via the public API." The 29-problem subset that fits scores 77.8%
Anthropic256K and 1M binsSimilarity to ground truth, 5-trial averagewww-cdn.anthropic.com

Opus 4.6 leads Sonnet 4.6 by 2.7 points at 256K and by 10.2 points at 1M, so the gap between the two models widens as prompts get longer. Sonnet's 77.8% subset score is higher than its full 1M score, but it covers 29 problems that fit the window, not the 100-problem bin. Both 1M numbers include problems beyond the public API window.

Anthropic internal runs, 64k extended thinking

Same card, bins and 5-trial average as the max-effort table, with a fixed 64k thinking budget instead. Do not mix these rows with the max-effort rows.

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 4.6Anthropic91.9% at 256K; 78.3% at 1M64k extended thinking; Anthropic's protocol on the fixed dataset; 5-trial average2026-02-17Sonnet 4.6 system card, Table 2.16.A. Same 1M window overflow as the max-effort run; no fitting-subset score
Claude Sonnet 4.6Anthropic90.6% at 256K; 65.1% at 1MSame as above2026-02-17Table 2.16.A. The 1M result is not reproducible via the public API. The 54-problem subset that fits scores 71.3%
Claude Sonnet 4.5Anthropic10.8% at 256K; 18.5% at 1MSame as above2026-02-17Table 2.16.A. The card gives no subset figure for this 1M cell
Anthropic256K and 1M binsSimilarity to ground truth, 5-trial averagewww-cdn.anthropic.com

The jump from Sonnet 4.5 to Sonnet 4.6 is the largest move on this page, from 10.8% to 90.6% at 256K under one protocol. Opus 4.6 at the 64k budget scores 78.3% at 1M, 2.3 points above Opus 4.6 at max effort. The max setting did not beat a fixed 64k budget for Opus on the longest bin, so "highest effort" is not a safe default for this workload.

Independent runs on contextarena.ai

Scores from contextarena.ai as reproduced in Anthropic's card, footnote 11. This is a different harness from Anthropic's internal runs, and the card does not state sampling or trial count, so compare these rows only with each other.

ModelOrganizationScoreHarness and setupDateSource and notes
GPT-5.2OpenAI63.9% at 256K; no 1M scorexhigh thinking; contextarena.ai harnessQuoted 2026-02-17Sonnet 4.6 system card, Table 2.16.A and footnotes 11 and 13. Independent run. No 1M score with a 400K window. The card also prints OpenAI's self-reported 70.0%, with effort and bin unstated
Gemini 3 FlashGoogle58.5% at 256K; 32.6% at 1MHigh thinking; contextarena.ai harnessQuoted 2026-02-17Table 2.16.A, footnote 11. Independent run
Gemini 3 ProGoogle45.4% at 256K; 24.5% at 1MHigh thinking; contextarena.ai harnessQuoted 2026-02-17Table 2.16.A, footnote 11. Independent run
contextarena.aicontextarena.ai256K and 1M bins

GPT-5.2 leads this table at 256K, 5.4 points ahead of Gemini 3 Flash. The more surprising result is inside Google's lineup. Flash beats Pro on both bins under the same harness and thinking setting, by 13.1 points at 256K and 8.1 points at 1M. If your workload is long-context retrieval on Gemini 3, the cheaper model is the stronger one in this run.

Google's MRCR-V2 runs from the Gemini 2.5 report

Google's own implementation, Table 4 of arXiv 2507.06261. The up-to-128K score is an average across lengths and the 1M score is pointwise. The report does not state the dataset source or whether it reran after the December 2025 fix, so these rows are not comparable to any other table.

ModelOrganizationScoreHarness and setupDateSource and notes
Gemini 2.5 ProGoogle58.0% up to 128K; 16.4% at 1MGoogle MRCR-V2; no setting stated2025-07-07, v6 2025-12-19Gemini 2.5 report, Table 4. Vendor run
OpenAI o3OpenAI57.1% up to 128K; no 1M scoreGoogle MRCR-V2; column header "o3 high"2025-07-07, v6 2025-12-19Table 4, Google's comparison table
Claude 4 SonnetAnthropic39.1% up to 128K; no 1M scoreGoogle MRCR-V2; no setting stated2025-07-07, v6 2025-12-19Table 4, Google's comparison table
OpenAI o4-miniOpenAI36.3% up to 128K; no 1M scoreGoogle MRCR-V2; column header "o4-mini high"2025-07-07, v6 2025-12-19Table 4, Google's comparison table
Grok 3 BetaxAI34.0% up to 128K; no 1M scoreGoogle MRCR-V2; no setting stated2025-07-07, v6 2025-12-19Table 4, Google's comparison table
GoogleGoogle MRCR-V2Up to 128K and 1Marxiv.org

Gemini 2.5 Pro edges o3 by 0.9 points on the up-to-128K average and is the only model in the table with a 1M score. Its own score falls from 58.0% to 16.4%, a 41.6-point drop between the two variants in the same report. Table 4 also lists Claude 4 Opus at 16.1% up to 128K, but its caption marks that run as having no thinking and API refusals, so I left it out of the ranking. The report's DeepSeek R1 0528 column has no MRCR-V2 entry.

Historical progression

Each row below uses its own protocol. Read the scores down a single row's protocol, never across rows.

DateMilestoneFigureProtocolSource
2024Michelangelo paper introduces the MRCR familyNo scores used hereVodrahalli et al.arXiv 2409.12640
2025-04-12OpenAI publishes openai/mrcr2,400 rows at 2, 4 and 8 needlesOpenAI dataset, original releaseHugging Face card
2025-07-07Gemini 2.5 technical reportGemini 2.5 Pro 58.0% up to 128K, 16.4% at 1MGoogle MRCR-V2, 8 needles with style keyarXiv 2507.06261, Tables 4 and 11
2025-12-05OpenAI dataset bugfixAbout 10% of rows had too many needles, about 5% wrong ground truthOpenAI dataset, fixedHugging Face card
2026-02-17Anthropic Sonnet 4.6 system card1M bin at 64k thinking: Sonnet 4.5 18.5%, Sonnet 4.6 65.1%, Opus 4.6 78.3%Anthropic internal, fixed OpenAI datasetSonnet 4.6 system card, Table 2.16.A

The only multi-model series under one protocol is Anthropic's. At 64k thinking, the 256K bin goes from Sonnet 4.5 at 10.8% to Sonnet 4.6 at 90.6% to Opus 4.6 at 91.9%. The 1M bin goes from 18.5% to 65.1% to 78.3%, with the window-overflow caveat on every 1M cell. None of the sources I read give release dates for these three models, so this is a progression by model generation, not by calendar.

Gemini 2.5 Pro's 16.4% at 1M and Opus 4.6's 78.3% at 1M are not a measure of progress. They come from different implementations, and one predates the dataset fix. No primary source covers Gemini 3 versions over time or any Claude, Gemini or OpenAI MRCR score newer than February 17, 2026.

Failure modes and limitations

A missing hash scores zero. The grader gives 0 to any response that doesn't begin with the required hash. A model that retrieves the right passage but opens with a preamble fails the sample, so MRCR counts instruction-following slips as retrieval failures.

Partial credit inflates the headline. SequenceMatcher ratio rewards near matches. A score is mean similarity, not the share of answers reproduced exactly, and a model that grabs the wrong poem can still collect some credit for overlapping text.

Data bugs before December 5, 2025. About 10% of rows held too many target needles and about 5% held wrong ground truth. Any OpenAI-dataset score published before the fix carries that noise.

The 1M bin overflows real windows. Tokenizer differences push some 1M problems past the 1,000,000-token Claude API limit, and Anthropic's headline 1M numbers come from internal runs it could not reproduce through the public API. On the OpenAI side, Astra caps input at 922,000 tokens against a bin ceiling of 1,048,576.

Labs report the same model differently. Anthropic's card prints GPT-5.2 at 256K as 63.9% from contextarena.ai at xhigh and 70.0% as self-reported by OpenAI. That is a 6.1-point gap for one model, and the card states neither the effort nor the bin behind the 70.0%.

Effort settings move the result. Opus 4.6 scores 78.3% at 1M with 64k thinking and 76.0% at max effort. Sonnet 4.6 at 256K scores 90.6% and 90.3% across the same two settings. Any report without an effort setting is incomplete.

Little independent verification. The only independent runs in the primary sources cover Gemini 3 Pro, Gemini 3 Flash and GPT-5.2, via contextarena.ai as quoted by Anthropic. Every other score is vendor-reported.

Public data, no contamination audit. The dataset is public under MIT and the prompts are synthetic. No primary source reports a contamination audit for MRCR.

Narrow domain. Every needle is a writing request, such as a poem or a blog post, inside a synthetic chat. MRCR says nothing direct about retrieval from contracts, codebases or spreadsheets.

Needle in a haystack is the baseline MRCR makes harder. One distinctive fact in filler can be found by matching one string. MRCR's identical needles force the model to track order and reproduce text. Even the leader in Anthropic's runs, Opus 4.6, drops from 93.0% at 256K to 76.0% at 1M on MRCR at max effort.

GraphWalks is the closest sibling. Each prompt is a directed graph of hexadecimal-hash nodes, and the model runs a breadth-first search or a parents lookup across it. Anthropic runs it next to MRCR in the same card, at 256K and 1M with 100 problems each, scored by F1 over a 5-trial average. Footnote 16 says half the 1M problems exceed the public API limit. GraphWalks tests multi-hop reasoning over the context. MRCR tests ordered retrieval and exact reproduction. The two rank the same pair of models in opposite order.

ModelMRCR v2 8-needle, 1M, max effortGraphWalks BFS, 1M, max effortSource
Claude Opus 4.676.0%38.7%Sonnet 4.6 system card, Tables 2.16.A and 2.16.B
Claude Sonnet 4.665.8%73.8%Sonnet 4.6 system card, Tables 2.16.A and 2.16.B
Anthropic8-needle, 1MMRCR v2 8-needle, 1M, max effort

Opus 4.6 leads by 10.2 points on MRCR and trails by 35.1 points on GraphWalks BFS. At 64k thinking the BFS 1M scores are Sonnet 4.6 68.4%, Opus 4.6 41.2% and Sonnet 4.5 25.6%. The pattern holds at shorter lengths on the 256K BFS subset, where Sonnet 4.6 scores 72.8% at 64k and 74.5% at max against Opus 4.6's 61.5% and 61.1%. On the 256K parents subset the models converge, with Sonnet 4.6 at 96.9% and 97.9%, Opus 4.6 at 95.1% and 95.4%, and Sonnet 4.5 at 81.0%. A team choosing a long-context model needs both numbers.

LOFT, from Lee et al. in 2024, tests hard retrieval over long corpora and appears in the same Gemini 2.5 report. Gemini 2.5 Pro scores 87.0% up to 128K and 69.8% at 1M on LOFT, against 58.0% and 16.4% on MRCR-V2. That is a 53.4-point gap at 1M on the same model in the same report. Finding relevant content in a corpus and reproducing the n-th of eight lookalikes are different skills.

MRCR also anchors discussions of context window expansion. A 1M-token window is a spec-sheet number. Astra's 1,050,000-token window with a 922,000-token input cap shows that even the spec has fine print, and MRCR is how you measure what the model does with the space. When the answer is "not enough," retrieval-augmented generation is the alternative to putting 500K tokens in the prompt. The effort results above connect to test-time compute. Max effort did not beat a 64k thinking budget for Opus 4.6 at 1M. For the general picture, see LLM benchmarks and LLM evaluation.

What it means for teams choosing a model

You have no MRCR number for today's flagships. No primary-sourced MRCR score exists for GPT-6 Astra, GPT-5.6 Sol, GPT-6.1 Sol, Claude Opus 5.5, Claude Fable 5.1, Claude Mythos 5.1 or Gemini 4. If you are choosing among current frontier models for long-context work, you need your own MRCR-style test at your real prompt length.

Among the models with sourced scores, Opus 4.6 leads on ordered retrieval. In Anthropic's max-effort runs, Opus 4.6 scores 93.0% at 256K and 76.0% at 1M, ahead of Sonnet 4.6 by 2.7 and 10.2 points. Both 1M numbers include problems above the public API's 1M-token limit, and only Sonnet 4.6 has a fitting-subset score, 77.8% on 29 problems.

Match the benchmark to the job. If your workload is pulling back the exact text of one item among many similar ones, such as the third version of a clause or a particular reply in a long thread, MRCR is the right signal and Opus 4.6 is ahead. If it is chaining facts across a long input, GraphWalks is the right signal and Sonnet 4.6 is ahead by 35.1 points at 1M.

Long prompts carry a surcharge. OpenAI lists GPT-6 Astra at $10 per 1M input tokens and $50 per 1M output, with 2x input and 1.5x output pricing above 272K input tokens. A single 524,288-token prompt costs at least $10.49 before any output, and a prompt at the 922,000-token cap costs at least $18.44. GPT-5.6 Sol lists $4 and $20, and GPT-6.1 Sol $2 and $10, with the same surcharge. Test at real length on a small sample before you commit a pipeline to it.

Don't rank across harnesses. Gemini 3 Flash scores 58.5% at 256K and 32.6% at 1M on contextarena.ai at high thinking, ahead of Gemini 3 Pro on both bins. Opus 4.6's max-effort numbers are 34.5 and 43.4 points higher, but they come from Anthropic's internal harness, so those gaps are not a like-for-like ranking.

The LLM leaderboard tracks the broader set of scores. To run the same kind of test on your own long documents and threads, you can build it in Klu for a knowledge assistant workflow at your real prompt lengths.