What is GraphWalks?
GraphWalks is a long-context reasoning benchmark that OpenAI published as the public dataset openai/graphwalks under an MIT license. It arrived with GPT-4.1 in April 2025, in OpenAI's Introducing GPT-4.1 in the API post, which is the origin Anthropic's system cards cite.
Each prompt is a directed graph written out as a list of edges between random hexadecimal hashes, followed by one operation. The model either runs a breadth-first search from a start node to a stated depth, or it answers a "parents" query by listing every node with an edge into a target node. It has to return the exact set of nodes.
That design is what separates GraphWalks from needle-in-a-haystack tests. A needle test hides one fact in filler and asks the model to find it. GraphWalks has no filler. The edges that matter for an answer are scattered across the whole prompt, and a BFS answer at depth three depends on chaining hops through edges that sit hundreds of thousands of characters apart. A model that retrieves well but can't hold a frontier of visited nodes in its head fails.
The dataset has 1,150 problems, 550 BFS and 600 parents, according to the Hugging Face dataset viewer statistics. Prompts run from 2,638 to 1,748,226 characters, with a median of 110,175. Scores are F1 over the returned node set. GraphWalks sits alongside other long-context tests in the broader set of LLM benchmarks, and it measures computation over a full context window rather than lookup.
How a task works
The prompt opens with three worked examples, then the graph, then the operation. The dataset README spells out the instructions the model gets. Every node has degree at least one. For BFS, return only the nodes at the stated depth and leave out the start node. For parents, return only direct predecessors and leave out the node itself.
The answer goes on the final line in a fixed format, Final Answer: [a, b, c]. A response without that line parses to an empty list.
There are no tools, no sandbox and no agent loop. It is one prompt in and one answer out, so the score reflects what the model does with its context and its reasoning budget, not what a scaffold adds. No source sets a time or step limit.
The answers vary a lot in size. The viewer reports answer sets from 0 to 6,943 nodes, with a median of 3. A model that returns three nodes when the truth is three hundred gets punished on recall, and one that dumps half the graph gets punished on precision.
The data ships as two parquet files. graphwalks_128k_and_shorter.parquet is 22 MB and graphwalks_256k_to_1mil.parquet is 297 MB, per the file listing. The viewer's length histogram puts 650 rows below 177K characters, 100 between 177K and 352K, 200 between 352K and 526K, and 200 between 1.57M and 1.75M, with nothing in between. That works out to 650 rows in the short file and 500 in the long file. With the README's count of 400 parents samples in the short file, the long file holds 200 parents rows and 300 BFS rows.
One detail for anyone writing their own grader. The README calls the answer column answer, but the parquet schema the viewer serves names it answer_nodes. Read the schema, not the README, when you load the files.
How scoring works
The grader extracts the node list from the Final Answer: line and computes precision, recall and F1 against the ground-truth set. If the line fails to parse, all three are 0.
The current README grader has a branch for empty sets. When the ground truth is empty and the model also returns an empty list, recall, precision and F1 are all 1.0. That branch was not always there, and the README's changelog doesn't record when it was added.
Anthropic's Claude Opus 4.6 system card, section 2.18.2, is the record of the earlier version. It reproduces the original suggested scoring code, which scored 0 whenever ground truth was empty, even when the model correctly predicted an empty set. Anthropic reports that ground truth was empty "in many cases" and published its fix, which scores 1.0. The current public grader now matches that fix. Anthropic's later cards describe it as correcting "an ambiguity in the published F1 metric".
A score computed with the original grader penalizes correct empty answers, and a score computed with the current one doesn't. They are not the same metric.
Versions, bins and the February 2026 patch
There is one public dataset. Reported results differ by operation and length bin, not by dataset version.
The README changelog lists the initial publication on April 12, 2025 and a bugfix on February 27, 2026. The fix corrected 24 of the 400 parents samples in the 128k-and-shorter file, whose ground truth wrongly included the root node. It also changed the BFS prompt to ask only for nodes at exactly the stated depth. The old wording, "reachable at that depth", implied that revisited nodes counted. OpenAI credits the Opus 4.6 system card for finding the parents issue, and the viewer's date_added field marks exactly 24 rows with the patch date.
Anthropic made three changes of its own before the patch, described in the Opus 4.6 card and restated in the Opus 4.8 card and the Claude Fable 5 and Claude Mythos 5 card:
- Empty prediction against empty ground truth scores 1.0 instead of 0.
- The 24 parents problems with self-loops had the target node removed from ground truth. These are the same 24 rows the README patch later fixed. They live in the short file, so Anthropic's reported 256K and 1M results are unaffected.
- The BFS prompt asks for nodes "exactly at depth N (not up to N)" because the original wording was ambiguous and "some models made different assumptions".
Labs slice the long file differently. Anthropic runs 100 problems at 256K tokens and 100 at 1,024K tokens for each of BFS and parents. Google, in its Gemini 4 Argon evaluation, uses an "Up to 128k" subset of 650 items, which matches the short file, and a "256 to 1M" subset of 200 problems. Both labs draw their long-context problems from the same 500-row file, and neither says how it picked them.
The 1M variants also can't be run through the public API. The Opus 4.6 card says Anthropic used "an internal setting to support the full prompt + thinking + output", and the Opus 4.8 and Mythos 5 cards state that "1M context subset results are not reproducible via the public API, as the problems exceed its 1M token limit."
Current leaderboard
There is no independent GraphWalks leaderboard. Every number below comes from a lab that also built one of the models in its table. Google ran the first two tables, Anthropic ran the next two, and the GPT-5.5 row is OpenAI's own figure as quoted by Anthropic. No source publishes cost, tokens or minutes per task, so this page has no cost or time chart.
Google run, BFS up to 128k tokens
Google computed GraphWalks for all four models itself, on the 650-item short subset, with results dated "as of October, 2026". Comparable within this table and to the 256k-to-1M table below; not comparable to Anthropic's tables, which use different bins and Anthropic's prompt and scoring changes.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Gemini 4 Argon | 99.7% | Gemini API, highest thinking setting, single attempt | Oct 2026 | Argon evaluation PDF, page 5 | |
| GPT-6 Astra | OpenAI | 98.7% | Maximum or best available reasoning, per Google's methodology; no per-model effort | Oct 2026 | Argon PDF, page 5. Self-computed by Google |
| Claude Fable 5.1 | Anthropic | 91.4% | Maximum or best available reasoning; no per-model effort | Oct 2026 | Argon PDF, page 5. Self-computed by Google |
| Claude Opus 5.5 | Anthropic | 90.6% | Maximum or best available reasoning; no per-model effort | Oct 2026 | Argon PDF, page 5. Self-computed by Google |
Argon and Astra are 1.0 point apart, which is close to saturated. The Anthropic models trail Astra by 7.3 and 8.1 points on this run.
Google run, BFS 256k to 1M tokens
Same lab, same four models, on Google's 200-problem long subset. Comparable to the table above; not comparable to Anthropic's 256K or 1M columns, because Google's bin spans both lengths and its problem selection is not published.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Gemini 4 Argon | 84.2% | Gemini API, highest thinking setting, single attempt | Oct 2026 | Argon evaluation PDF, page 5. Highlighted as best in row | |
| GPT-6 Astra | OpenAI | 71.8% | Maximum or best available reasoning; no per-model effort | Oct 2026 | Argon PDF, page 5. Self-computed by Google |
| Claude Opus 5.5 | Anthropic | 66.8% | Maximum or best available reasoning; no per-model effort | Oct 2026 | Argon PDF, page 5. Self-computed by Google |
| Claude Fable 5.1 | Anthropic | 65.0% | Maximum or best available reasoning; no per-model effort | Oct 2026 | Argon PDF, page 5. Self-computed by Google |
Past 256k tokens the field spreads out. Argon leads Astra by 12.4 points, Opus 5.5 by 17.4 and Fable 5.1 by 19.2. Astra leads Opus 5.5 by 5.0. Opus 5.5 and Fable 5.1 swap places relative to the short bin, with Opus 5.5 ahead by 1.8.
Google's methodology says Gemini runs use "the highest thinking settings" and that it averages "over multiple trials for smaller benchmarks" without giving a count. For the three non-Gemini models it reports "maximum thinking/reasoning settings available" or, failing that, "best available reasoning results", and doesn't say which applied to which model. Google also doesn't say whether it used the patched data, the current grader, or Anthropic's exact-depth prompt.
Anthropic run, BFS
Anthropic ran its own models with the exact-depth BFS prompt and the empty-set F1 fix, averaged over five trials with default sampling. Comparable across rows in this table; not comparable to Google's tables, to the Opus 4.6 card's February run, or to the GPT-5.5 row.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Mythos 5 | Anthropic | 91.1% at 256K; 79.4% at 1M | Anthropic prompt and F1 fix, 5-trial average; 1M run on an internal setting | 2026-06-09 | Fable 5 and Mythos 5 card, Table 8.13.A |
| Claude Opus 4.8 | Anthropic | 85.9% at 256K; 68.1% at 1M | Adaptive thinking, max effort; same prompt, F1 fix and 5-trial average | 2026-05-28 | Opus 4.8 card, Table 8.9.A; same values in the Mythos 5 card |
| Claude Mythos Preview | Anthropic | 85.7% at 256K; 74.3% at 1M | Anthropic prompt and F1 fix, 5-trial average; 1M run on an internal setting | 2026-06-09 | Mythos 5 card, Table 8.13.A |
| Claude Opus 4.7 | Anthropic | 76.9% at 256K; 40.3% at 1M | Anthropic prompt and F1 fix, 5-trial average; 1M run on an internal setting | 2026-05-28 | Opus 4.8 card, Table 8.9.A |
| Claude Opus 4.6 | Anthropic | 61.1% at 256K; 16.3% at 1M | Anthropic prompt and F1 fix, 5-trial average; 1M run on an internal setting | 2026-05-28 | Opus 4.8 card, Table 8.9.A. The Opus 4.6 card reports different 1M numbers; see history below |
Mythos 5 leads at both lengths. At 1M it is 11.3 points ahead of Opus 4.8, and it loses 11.7 points going from 256K to 1M where Opus 4.8 loses 17.8.
Anthropic run, parents
Same harness and cards as the BFS table. Comparable across rows here and to the BFS table as a measure of the same models on a different operation; not comparable to Google's tables or the Opus 4.6 card's February run.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Mythos 5 | Anthropic | 99.96% at 256K; 97.5% at 1M | Anthropic prompt and F1 fix, 5-trial average; 1M run on an internal setting | 2026-06-09 | Mythos 5 card, Table 8.13.A. Four of five 256K runs scored 99.95, each missing one node on one shared problem; one scored 100.0 |
| Claude Mythos Preview | Anthropic | 99.9% at 256K; 95.5% at 1M | Anthropic prompt and F1 fix, 5-trial average; 1M run on an internal setting | 2026-06-09 | Mythos 5 card, Table 8.13.A |
| Claude Opus 4.8 | Anthropic | 99.3% at 256K; 83.3% at 1M | Adaptive thinking, max effort; same prompt, F1 fix and 5-trial average | 2026-05-28 | Opus 4.8 card, Table 8.9.A; same values in the Mythos 5 card |
| Claude Opus 4.6 | Anthropic | 95.4% at 256K; 48.6% at 1M | Anthropic prompt and F1 fix, 5-trial average; 1M run on an internal setting | 2026-05-28 | Opus 4.8 card, Table 8.9.A |
| Claude Opus 4.7 | Anthropic | 93.6% at 256K; 56.6% at 1M | Anthropic prompt and F1 fix, 5-trial average; 1M run on an internal setting | 2026-05-28 | Opus 4.8 card, Table 8.9.A |
Parents at 256K is solved for the top three. The 1M column still separates them, with Mythos 5 at 97.5 and Opus 4.8 at 83.3.
OpenAI-reported GPT-5.5
OpenAI ran this itself at xhigh thinking, and Anthropic prints the numbers beside its own models in the Opus 4.8 and Mythos 5 cards without rerunning them. Not comparable to either Anthropic table, since neither card states the prompt or grader OpenAI used.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| GPT-5.5 | OpenAI | BFS 73.7% at 256K, 45.4% at 1M; parents 90.1% at 256K, 58.5% at 1M | xhigh thinking; prompt and grader not stated | Quoted 2026-05-28 | Introducing GPT-5.5, as quoted in Opus 4.8 card Table 8.9.A and Mythos 5 card Table 8.13.A. The OpenAI page was not retrievable directly |
The newest models have gaps. Anthropic's Opus 5.5 card and Fable 5.1 and Mythos 5.1 card report no GraphWalks result. OpenAI's GPT-6 Astra card and GPT-6.1 Sol addendum don't either. Google's run is the only place the current Anthropic and OpenAI models appear.
Historical progression
The Opus 4.6 card, published in February 2026, ran GraphWalks before the README patch. Its "1M" rows cover the full 1M variant, where footnote 15 says "half the problems exceed its 1M token limit". Its "256K subset" rows filter to problems the public API can handle. The Opus 4.6 card's columns for Gemini 3 Pro, Gemini 3 Flash and GPT-5.2 show a dash on every GraphWalks row.
This table is Anthropic-run with Anthropic's prompt and scoring changes, averaged over five trials. Comparable across its own rows; not comparable to the May and June cards, which define "1M" differently.
| Model | Setting | BFS 1M | BFS 256K subset | Parents 1M | Parents 256K subset |
|---|---|---|---|---|---|
| Claude Opus 4.6 | Max effort, adaptive | 38.7% | 61.1% | 72.0% | 95.4% |
| Claude Opus 4.6 | 64k extended thinking | 41.2% | 61.5% | 71.1% | 95.1% |
| Claude Sonnet 4.5 | 64k extended thinking | 25.6% | 44.9% | 50.2% | 81.0% |
Source: Claude Opus 4.6 system card, Table 2.18.A, February 2026.
Within that card, at 64k thinking, Opus 4.6 gained 16.6 points over Sonnet 4.5 on the BFS 256K subset, 15.6 on BFS 1M, 20.9 on parents 1M and 14.1 on the parents 256K subset.
The bigger jump came over the next four months on the harness Anthropic used in May and June. On BFS 256K, Opus 4.6 scored 61.1, Opus 4.7 76.9, Opus 4.8 85.9 and Mythos 5 91.1, a 30.0-point gain from February to June 2026. BFS 1M went from 16.3 to 40.3 to 68.1 to 79.4, a 63.1-point gain. Parents 1M went from 48.6 to 56.6 to 83.3 to 97.5. Parents 256K moved from 95.4 to 99.96 and has no room left.
The two Anthropic harnesses don't agree on Opus 4.6. The February card reports BFS 1M at 38.7 and parents 1M at 72.0 at max effort. The Opus 4.8 card restates the same model at 16.3 and 48.6. The 256K numbers match exactly, at 61.1 and 95.4, so the gap is in the 1M bin. The Opus 4.8 and Mythos 5 cards say they are "separating out" a 256K subset and a 1M subset, but neither explains how its "1M subset" maps to the February card's full 1M variant. Anthropic has not published that mapping, so treat the February 1M numbers as a separate series.
Before 2026 there are no scores on this page. GraphWalks launched with GPT-4.1 in April 2025, but OpenAI's launch post was not retrievable for this page, so its original results are not included.
Documented failure modes
The length cliff. Every model loses BFS points as prompts grow. In Google's run, going from up to 128k to 256k-1M costs Argon 15.5 points, Opus 5.5 23.8, Fable 5.1 26.4 and Astra 26.9. In Anthropic's run, going from 256K to 1M costs Mythos 5 11.7, Mythos Preview 11.4, Opus 4.8 17.8, Opus 4.7 36.6 and Opus 4.6 44.8. GPT-5.5, on OpenAI's own run, loses 28.3. The cliff has shrunk a lot from Opus 4.6 to Mythos 5, but no model is flat.
BFS is harder than parents. A parents query is a single-hop lookup. Scan the edge list for every edge pointing at the target. BFS has to chain hops, track which nodes are already visited and stop at the right depth. At 256K, Mythos 5 scores 99.96 on parents and 91.1 on BFS, and Opus 4.8 scores 99.3 and 85.9. Parents also degrades less at 1M. Mythos 5 drops 2.5 points on parents against 11.7 on BFS. Opus 4.8 drops 16.0, Opus 4.7 37.0 and GPT-5.5 31.6 on parents.
Prompt ambiguity. The original BFS prompt didn't say whether depth N meant exactly N or up to N, and models answered both ways, per the Opus 4.6 card. A model that returns every node within N hops gets dinged on precision for a perfectly reasonable reading. Anthropic's prompt and the February 2026 README patch both fix this, but any result computed on the original wording carries that noise.
Bad ground truth. In 24 parents problems in the short file, the ground truth wrongly included the target node itself. The Opus 4.6 card traces this to self-loops. The README patch fixed them on February 27, 2026. Google's up-to-128k subset includes those rows, and Google doesn't say whether it used the patched data.
The empty-set scoring bug. The original grader scored a correct empty answer as 0, and empty ground truth came up "in many cases", per Anthropic. The current README grader scores it 1.0. Google doesn't say which grader it used for its run.
Format failures. A response that omits the Final Answer: [...] line parses to an empty list. On a problem with a non-empty answer, that scores 0, however good the reasoning above it.
1M runs need special access. The 1M variants exceed the public API's 1M-token limit, and Anthropic ran them on an internal setting. You can't reproduce Anthropic's 1M numbers yourself.
No error bars and no independent runner. Anthropic averages five trials and Google states pass@1 with an unstated trial count. No source reports confidence intervals. Google ran its own table and Anthropic ran its own tables, and no third party publishes GraphWalks results.
Contamination is untested. The dataset is public. No source reports a contamination check or a held-out set for GraphWalks.
How GraphWalks compares to related benchmarks
OpenAI's MRCR v2 8-needle test is the closest relative. It asks a model to retrieve a specific item among near-identical ones in a long multi-turn conversation, scored by mean match ratio. Anthropic reports both together in the Opus 4.6 card, sections 2.18.1 and 2.18.2. The difference is the kind of work. MRCR tests ordinal retrieval, picking out one item among near-identical ones. GraphWalks tests computation over everything in the prompt.
Needle in a haystack is the simplest long-context test, one fact dropped into filler text. It checks that a model can find something, not that it can use what it finds. GraphWalks has no filler, and every edge can change the answer.
ProgramBench is an agentic long-context coding task, where a model rebuilds a program from a binary and documentation. Anthropic's Opus 5.5 card uses ProgramBench as its long-context evaluation in section 8.10 and reports no GraphWalks. If you want a single-turn reasoning signal, GraphWalks is cleaner. If you want an agent working across a large codebase with tools, ProgramBench is closer to the job.
The task itself is textbook graph traversal, which makes failures easy to read. A wrong answer is either missing nodes, which shows up as low recall, or extra nodes, which shows up as low precision, and you can check which by hand. The reasoning budget is the other lever. The Opus 4.6 card ran two test-time compute settings, 64k extended thinking and max-effort adaptive thinking, and they land within 2.5 points of each other on BFS 1M and within 0.9 on every other column. On that card, the setting moved the score far less than the model generation did.
What GraphWalks means for teams choosing a model
Below 128K tokens, Argon and Astra are tied in practice. On Google's run, Argon scores 99.7 and Astra 98.7. Fable 5.1 and Opus 5.5 trail Astra by 7.3 and 8.1 points at 91.4 and 90.6. If your relational data fits in 128K tokens and you use Argon or Astra, pick between them on price and other evals.
From 256K to 1M tokens, Gemini 4 Argon leads clearly. It is 12.4 points ahead of Astra at 84.2 against 71.8, and 17.4 ahead of Opus 5.5. Astra beats Opus 5.5 by 5.0 and Fable 5.1 by 6.8. That ranking comes from Google running its own model against competitors, at settings it describes only as maximum or best available for the non-Gemini models. Nobody else has published a comparable run, so it is the best evidence available and it is not neutral.
Every model loses 15 to 27 points of BFS accuracy once prompts pass 128K. If you are feeding dependency graphs, org charts, citation chains or permission trees into one long prompt and asking multi-hop questions, keep the prompt under 128K. Above that, chunk the graph or run the traversal as an external graph query and give the model the result.
Lookups are solved, chained traversal is not. Parents at 256K is near ceiling for Anthropic's recent models, at 99.96 for Mythos 5 and 99.3 for Opus 4.8. BFS at 1M is 79.4 and 68.1. A model that answers "who reports to X" reliably can still miss nodes on "who is three levels below X". Test the multi-hop version of your question, not the single-hop one.
Anthropic's own numbers show fast progress, but stop at Mythos 5. BFS 256K rose 30.0 points from Opus 4.6 to Mythos 5 between February and June 2026, and BFS 1M rose from 16.3 to 79.4. Anthropic has not published GraphWalks for Opus 5.5 or Fable 5.1, and OpenAI has not published it for Astra or Sol. Google's run is the only view of those models, and it puts Opus 5.5 at 66.8 on the long bin. If long-context traversal is central to your product, ask your vendor for the number at your context length, or run it yourself.
Check the grader before you compare. A score from the original empty-set grader, the original BFS prompt or the unpatched parents rows is not comparable to one from the current README. When you run GraphWalks as part of your own LLM evaluation, use the current grader and note which subset you drew from the long file. Pin the date of any frontier model number you rely on.
You can compare these models across other benchmarks on the Klu LLM leaderboard. To run the same kind of multi-hop check on your own relational data, see how data teams set up evals in Klu on the data analysis page.