DRACO (Deep Research Accuracy, Completeness, and Objectivity)

A 100-task benchmark for deep research agents, built from real Perplexity Deep Research queries, where an LLM judge grades each report against expert-written rubrics

Knowledge workScore, Gemini 3.1 Pro Preview judge

Stephen M. Walker II · Co-Founder / CEO · October 6, 2026 · 15 min read

What is DRACO?

DRACO, short for Deep Research Accuracy, Completeness, and Objectivity, is a benchmark for deep research agents. Each of its 100 tasks is an open-ended research question. The system browses, analyzes sources and writes a cited report, and an LLM judge grades that report against a hand-built rubric of about 39 weighted criteria. Perplexity AI built it with Harvard. The ten authors include Joey Zhong, Hao Zhang, Denis Yarats and Jerry Ma, and the paper went up on arXiv on 12 February 2026. The tasks and full rubrics are on Hugging Face under the MIT license.

What sets DRACO apart is where the questions come from. Perplexity sampled 1,000 hard English queries from real Deep Research traffic in September and October 2025 and cut them down to 100. Earlier deep research benchmarks either ask closed-ended questions a program can check, like GAIA, Humanity's Last Exam and BrowseComp, or use tasks that people or models wrote from scratch. The paper's Table 1 compares 13 earlier open-ended benchmarks, and none of them draws tasks from a production deep research system.

The rubrics punish mistakes. Criteria can carry negative weights, down to -500 for dangerous medical advice, so a polished report with wrong claims scores low. That makes DRACO a test of hallucination under research conditions as much as a test of coverage.

As of October 2026, every frontier score comes from Anthropic's system cards. In the latest card that compares the top Claude models, Claude Opus 5 leads at 88.3 at max effort, 0.6 points ahead of Claude Fable 5.1 and 0.9 ahead of Claude Opus 5.5. OpenAI and Google DeepMind publish no DRACO score. The GPT-6 Astra System Card, the GPT-6.1 Sol addendum and the Gemini 4 evaluation methodology page contain no DRACO entry.

How the tasks are built

The 100 tasks span 10 domains and draw on information sources from 40 countries. Finance makes up 20% of the set, Shopping and Product Comparison 16%, and Academic 12%. The rest are Technology, General Knowledge, UX Design, Law, Medicine, Needle in a Haystack and Personalized Assistant. Every input is a single-turn, text-only English query, and every output is a written report.

Perplexity used user unhappiness as its difficulty signal. A query counted as hard if the user showed negative sentiment or gave the earlier response a thumbs-down. The team rewrote each query to remove personal information and ambiguity, filtered the set, and had experts from Perplexity and The LLM Data Company verify it.

Rubrics took more work than the tasks. 26 domain experts wrote them, including doctors, attorneys, financial analysts, software engineers and designers. One expert drafted each rubric with LLM help, which took 45 to 60 minutes and at least 6 model turns, and a second expert reviewed it. Then came a saturation test. Perplexity Deep Research ran on the task, and if it scored above 90% the task went back for rework. About 45% of tasks were sent back. A final review by an in-house domain expert and an AI expert returned about 10% of rubrics.

That saturation step tuned difficulty against Perplexity's own product. Keep it in mind when you read the paper's leaderboard, where that product ranks first.

How scoring works

The 100 rubrics hold 3,934 criteria, an average of 39.3 per task, and 415 of them are negative. Each criterion belongs to one of four axes.

AxisCriteria per taskShare of criteriaWeight range
Factual Accuracy20.552%-500 to +20
Breadth and Depth8.622%-100 to +10
Presentation Quality5.614%-50 to +20
Citation Quality4.812%-150 to +10

Medical harm penalties run from -50 to -500. Penalties outside medicine are usually -10 to -25. Half the criteria check facts, and the biggest penalties sit there, so factual errors cost far more than a dull layout.

The judge marks each criterion MET or UNMET. A met positive criterion adds its weight, an unmet one adds nothing, and a met negative criterion subtracts its weight. The task score is the raw total divided by the sum of positive weights, clamped between 0 and 1 and multiplied by 100. The benchmark score is the mean across all 100 tasks, averaged over 5 independent judge runs. The paper also reports a pass rate, the unweighted share of criteria met.

The original judge was Gemini-3-Pro at low thinking and temperature 0.2. The paper reran everything with GPT-5.2 at no reasoning and temperature 0, and with Sonnet-4.5 with reasoning off and temperature 0. Changing the judge moved absolute scores by 10 to 25 points while the ranking of systems held. Anthropic repeats that 10 to 25 point figure in all six system cards that report DRACO. A DRACO number without its judge is close to meaningless.

Two harnesses: the paper and Anthropic's cards

The paper tests products as black boxes. It ran Perplexity Deep Research on Opus 4.6 and on Opus 4.5, Gemini Deep Research, and OpenAI Deep Research with o3 and with o4-mini. It also ran Claude Opus 4.5 and Opus 4.6 as standard models with built-in web search and code execution, because Anthropic offers no dedicated research API. The paper says plainly that these products use different retrieval and browsing stacks, so its comparison is between systems, not models. It does not disclose the reasoning effort or tool settings for the OpenAI and Google products.

Anthropic's cards use Anthropic's own agent harness without the paper's Appendix C.4 system prompt. The model gets web search, web fetch, programmatic tool calling and code execution, a 980k-token task budget inside a 1M-token limit, and adaptive thinking at five effort levels, from low through medium, high and xhigh to max. The 1M figure is Anthropic's limit, not a DRACO rule. DRACO sets no context window length.

Anthropic swapped the judge to Claude Opus 4.6, with 5 grading runs and the mean reported. The cards give the reason as "Gemini 3 Pro, which is no longer available." Anthropic also changed how it pulls the report out of the transcript:

  • The Opus 4.8 card of 28 May 2026 and the Fable 5 / Mythos 5 card of 9 June 2026 grade only the text the model wraps in <result> tags.
  • From the Opus 5 card of 24 July 2026 onward, the model writes its final report to a file in its execution environment, and only that file is graded.

The reason was a real failure. At high effort, models sometimes left out the tags, and complete reports were scored as empty. Anthropic says control runs showed the file method "recovers those losses without otherwise affecting scores." Compaction changed too. The Opus 4.8 card compacted context at 200k tokens. The Mythos 5 card dropped compaction, "given that it does not significantly help for this task."

Versions

DRACO has one task set, the 100 tasks released in February 2026. There is no v2. What changed is the grading protocol, and those changes matter more than any version label would:

  1. Gemini-3-Pro judge, products as shipped, in the February 2026 paper.
  2. Opus 4.6 judge with <result>-tag extraction in Anthropic's May and June 2026 cards.
  3. Opus 4.6 judge with file extraction in Anthropic's cards from July 2026 onward.
  4. Gemini 3.1 Pro Preview judge with OpenRouter's tools in OpenRouter's June 2026 run.

Scores from different protocols do not belong in the same ranking.

Current leaderboard

Every Anthropic table below is vendor-reported, with Anthropic grading its own models using a Claude Opus 4.6 judge. Each table is one card and one run. Anthropic publishes these scores only as labels on figures in the system card PDFs, with no score table. Only two values, Mythos 5 at 86.4 and Opus 4.8 at 80.4, also appear in card text. The Sonnet 5.5 and Fable 5.1 figure labels were rechecked against the rendered pages on 6 October 2026.

Opus 5.5 card, 22 September 2026

This is the most recent card that puts the three strongest Claude models side by side. Opus 5 leads at 88.3, Fable 5.1 follows at 87.7, and Opus 5.5 sits at 87.4. All three land within one point of each other at max effort.

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 5Anthropic88.3Anthropic agent harness, max effort, file-graded report, 980k budget2026-09-22Opus 5.5 card Fig 8.11.2.A, p. 187. Low / med / high / xhigh 83.3 / 85.6 / 87.3 / 87.4
Claude Fable 5.1Anthropic87.7Anthropic agent harness, max effort, file-graded report, 980k budget2026-09-22Same figure. Low / med / high / xhigh 84.2 / 85.7 / 86.5 / 86.9
Claude Opus 5.5Anthropic87.4Anthropic agent harness, max effort, file-graded report, 980k budget2026-09-22Same figure. Low / med / high / xhigh 72.5 / 83.9 / 85.0 / 86.7
AnthropicAnthropic agent harness100 tasks, file-graded reportScore, Claude Opus 4.6 judgewww-cdn.anthropic.com

Comparable within this table only. It is one Anthropic run with an Opus 4.6 judge, and gaps against other cards are not model changes.

Sonnet 5.5 card, 28 September 2026

Sonnet 5.5 reaches 87.0 at max effort, 0.4 points under Opus 5.5 in the same run, and 6.2 points above Sonnet 5.

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 5.5Anthropic87.4Anthropic agent harness, max effort, file-graded report, 980k budget2026-09-28Sonnet 5.5 card Fig 8.11.2.A, p. 121. Low / med / high / xhigh 72.5 / 83.9 / 85.0 / 86.7
Claude Sonnet 5.5Anthropic87.0Anthropic agent harness, max effort, file-graded report, 980k budget2026-09-28Same figure. Low / med / high / xhigh 64.6 / 71.2 / 81.0 / 85.7
Claude Sonnet 5Anthropic80.8Anthropic agent harness, max effort, file-graded report, 980k budget2026-09-28Same figure. Low / med / high / xhigh 64.8 / 70.0 / 75.0 / 77.9. Sonnet 5 scored 84.3 in the Opus 5 card, and Anthropic gives no cause for the gap
AnthropicAnthropic agent harness100 tasks, file-graded reportScore, Claude Opus 4.6 judgewww-cdn.anthropic.com

Comparable within this table only. Same judge and extraction as the Opus 5.5 card, but a separate run.

Fable 5.1 card, 1 September 2026

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Fable 5.1Anthropic87.7Anthropic agent harness, max effort, file-graded report, 1M-token limit2026-09-01Fable 5.1 / Mythos 5.1 card Fig 8.12.2.A, p. 179. Low / med / high / xhigh 84.6 / 85.3 / 85.4 / 87.1
Claude Opus 5Anthropic87.6Anthropic agent harness, max effort, file-graded report, 1M-token limit2026-09-01Same figure. Low / med / high / xhigh 82.6 / 85.0 / 87.5 / 87.9. Opus 5 scores higher at xhigh than at max in this run
Claude Fable 5Anthropic86.0Anthropic agent harness, max effort, file-graded report, 1M-token limit2026-09-01Same figure. Low / med / high / xhigh 77.1 / 80.1 / 80.7 / 83.4
AnthropicAnthropic agent harness100 tasks, file-graded reportScore, Claude Opus 4.6 judgewww-cdn.anthropic.com

Comparable within this table only. One run, Opus 4.6 judge, file extraction.

Opus 5 card, 24 July 2026

This was the first card to grade a report file instead of tagged text.

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 5Anthropic88.6Anthropic agent harness, max effort, file-graded report, 1M-token limit2026-07-24Opus 5 card Fig 8.10.4.A, p. 166. Low / med / high / xhigh 83.2 / 86.8 / 87.0 / 88.1
Claude Mythos 5Anthropic87.7Anthropic agent harness, max effort, file-graded report, 1M-token limit2026-07-24Same figure. Low / med / high / xhigh 78.3 / 81.0 / 83.5 / 86.3
Claude Sonnet 5Anthropic84.3Anthropic agent harness, max effort, file-graded report, 1M-token limit2026-07-24Same figure. Low / med / high / xhigh 66.8 / 72.7 / 76.4 / 81.6
Claude Opus 4.8Anthropic80.6Anthropic agent harness, max effort, file-graded report, 1M-token limit2026-07-24Same figure. Low / med / high / xhigh 70.8 / 74.6 / 76.3 / 79.4
AnthropicAnthropic agent harness100 tasks, file-graded reportScore, Claude Opus 4.6 judgewww-cdn.anthropic.com

Comparable within this table only. Opus 5 appears at 88.6 here, 87.6 in the Fable 5.1 card and 88.3 in the Opus 5.5 card, and Anthropic does not say whether checkpoint, budget or harness changed between those runs.

Original paper results and latency, 12 February 2026

The paper's own leaderboard tests shipped products. Perplexity Deep Research on Opus 4.6 leads at 70.5, 11.5 points ahead of Gemini Deep Research and 10.7 ahead of plain Claude Opus 4.6 with web search and code execution.

ModelOrganizationScoreHarness and setupDateSource and notes
Perplexity Deep Research (Opus 4.6)Perplexity70.5 ± 0.3Product as shipped, Gemini-3-Pro judge2026-02-12Paper Tables 8, 9, 15. Pass rate 72.8. GPT-5.2 judge 50.4, Sonnet-4.5 judge 75.5
Perplexity Deep Research (Opus 4.5)Perplexity67.2 ± 0.3Product as shipped, Gemini-3-Pro judge2026-02-12Pass rate 70.9. GPT-5.2 judge 43.3, Sonnet-4.5 judge 70.3
Claude Opus 4.6 + web search + code execAnthropic59.8 ± 0.3Standard model with built-in tools, Gemini-3-Pro judge2026-02-12Pass rate 63.1. GPT-5.2 judge 42.7, Sonnet-4.5 judge 70.1
Gemini Deep ResearchGoogle59.0 ± 0.4Product as shipped, settings undisclosed, Gemini-3-Pro judge2026-02-12Pass rate 62.7. GPT-5.2 judge 37.8, Sonnet-4.5 judge 61.4
OpenAI Deep Research (o3)OpenAI52.1 ± 0.2Product as shipped, settings undisclosed, Gemini-3-Pro judge2026-02-12Pass rate 56.9. GPT-5.2 judge 31.7, Sonnet-4.5 judge 49.4
Claude Opus 4.5 + web search + code execAnthropic46.7 ± 0.3Standard model with built-in tools, Gemini-3-Pro judge2026-02-12Pass rate 50.2. GPT-5.2 judge 31.9, Sonnet-4.5 judge 58.7
OpenAI Deep Research (o4-mini)OpenAI41.9 ± 0.4Product as shipped, settings undisclosed, Gemini-3-Pro judge2026-02-12Pass rate 48.0. GPT-5.2 judge 25.3, Sonnet-4.5 judge 41.7
Perplexity100 tasks, systems as shippedScore, Gemini-3-Pro judgearxiv.org

Comparable within this table only. Perplexity wrote the benchmark, ran this table and tuned task difficulty against its own product, and the Gemini-3-Pro judge puts these scores on a different scale from every Anthropic card.

The leading system's profile is uneven. Perplexity Deep Research on Opus 4.6 scores 90.2 on Law, 82.8 on Academic and 80.5 on Medicine, but 63.8 on Personalized Assistant and 58.1 on Needle in a Haystack. By axis it gets 90.3 on Presentation and 64.6 on Citation, with Factual Accuracy at 67.9 and Breadth and Depth at 66.0. Writing a clean report is the easy part. Getting the facts and citations right is where every system loses points.

The paper also timed each system, which gives the only score-versus-time data DRACO has. Anthropic publishes cost per task for its 2026 runs only as an unlabeled log-scale figure, so there are no dollar values to chart.

DRACO score vs. time per taskOriginal paper, February 2026, Gemini-3-Pro judge, systems as shipped
40%50%60%70%5 min10 min20 min
FrontierAnthropicOtherBehind the frontier
Average time per task, log scale
Source: Zhong et al., arXiv 2602.11685, Table 8 for scores and Table 10 for average latency per task, converted from seconds to minutes. DR means Deep Research. All seven systems come from the paper's own run, and no 2026 frontier model appears.

Perplexity Deep Research on Opus 4.6 is both the most accurate system and one of the fastest, at 4.1 minutes per task. Gemini Deep Research takes 9.9 minutes for 59.0. OpenAI Deep Research with o3 is the slowest at 30.1 minutes and scores 52.1. The two plain Claude runs finish in about 3 minutes. The fast systems did not read less. Perplexity on Opus 4.6 averaged 778,711 input tokens per task and plain Opus 4.6 averaged 691,338, against 315,548 for Gemini and 44,587 for o3.

Anthropic's only 2026 timing figure is a range. In the Opus 5.5 card, single-task latencies on DRACO ran from 13 to 71 minutes, with no effort level stated, so that range cannot go on this chart.

OpenRouter fusion run, June 2026

OpenRouter's run is the only third-party DRACO result with 2026 models. It tests "fusion," where a panel of models answers in parallel and a synthesizer model merges the answers. Claude Opus 4.8 was the synthesizer for every fused panel.

ModelOrganizationScoreHarness and setupDateSource and notes
Fable 5 + GPT-5.5, fused panelAnthropic, OpenAI69.0OpenRouter web search, web fetch and bash; Opus 4.8 synthesizer2026-06-123.7 above solo Fable 5 and 9.0 above solo GPT-5.5
Opus 4.8 + GPT-5.5 + Gemini 3.1 Pro, fused panelAnthropic, OpenAI, Google68.3Same tools; Opus 4.8 synthesizer2026-06-12Page updated 2026-07-13
Opus 4.8 + GPT-5.5, fused panelAnthropic, OpenAI67.6Same tools; Opus 4.8 synthesizer2026-06-12
Opus 4.8 + Opus 4.8, fused panelAnthropic65.5Same tools; Opus 4.8 synthesizer2026-06-12
Claude Fable 5, soloAnthropic65.3Same tools, no synthesis2026-06-12Scored on 93 of 100 tasks. Content filters blocked 7, with no fallback
Budget panel: Gemini 3 Flash, Kimi K2.6, DeepSeek V4 ProGoogle, Moonshot AI, DeepSeek64.7Same tools; Opus 4.8 synthesizer2026-06-125.9 above solo Opus 4.8, its own synthesizer
DeepSeek V4 Pro, soloDeepSeek60.3Same tools, no synthesis2026-06-12
GPT-5.5, soloOpenAI60.0Same tools, no synthesis2026-06-12
Claude Opus 4.8, soloAnthropic58.8Same tools, no synthesis2026-06-12
Kimi K2.6, soloMoonshot AI53.7Same tools, no synthesis2026-06-12
Gemini 3.1 Pro, soloGoogle45.4Same tools, no synthesis2026-06-12
Gemini 3 Flash, soloGoogle43.1Same tools, no synthesis2026-06-12
OpenRouterOpenRouter web search, web fetch and bash100 tasksScore, Gemini 3.1 Pro Preview judgeopenrouter.ai

Comparable within this table only. The Gemini 3.1 Pro Preview judge and Exa-backed OpenRouter tools differ from every other table, and each fused score includes an Opus 4.8 synthesis step.

The steel.dev aggregator lists MiniMax M3 at 73.23% under an Opus 4.6 judge and MiniMax's internal harness. No primary source with that number exists, so it stays out of the tables.

Historical results

DRACO's history is a series of separate runs under different protocols. Each row below is one run. Read gains within a row, never across rows.

DateRun and protocolJudgeScores at max effort or as shipped
2026-02-12Paper, products as shippedGemini-3-ProOpus 4.5 46.7 to Opus 4.6 59.8 with the same tools. Perplexity DR 67.2 on Opus 4.5, 70.5 on Opus 4.6
2026-05-28Opus 4.8 card, <result> tags, compaction at 200kClaude Opus 4.6Sonnet 4.6 75.8, Opus 4.6 76.5, Opus 4.7 77.7, Opus 4.8 80.4, Mythos Preview 83.7
2026-06-09Fable 5 / Mythos 5 card, <result> tags, no compactionClaude Opus 4.6Opus 4.8 80.6, Mythos Preview 83.6, Mythos 5 86.4
2026-07-24Opus 5 card, file extractionClaude Opus 4.6Opus 4.8 80.6, Sonnet 5 84.3, Mythos 5 87.7, Opus 5 88.6
2026-09-01Fable 5.1 card, file extractionClaude Opus 4.6Fable 5 86.0, Opus 5 87.6, Fable 5.1 87.7
2026-09-22Opus 5.5 card, file extractionClaude Opus 4.6Opus 5.5 87.4, Fable 5.1 87.7, Opus 5 88.3
2026-09-28Sonnet 5.5 card, file extractionClaude Opus 4.6Sonnet 5 80.8, Sonnet 5.5 87.0, Opus 5.5 87.4

Mythos 5 scores 86.4 in the tag-graded card and 87.7 in the file-graded card. That is a change of protocol, not of model.

The two tag-graded cards are kept in full below. They predate the file-extraction fix, so a dropped <result> tag at high effort counted as an empty report.

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Mythos 5Anthropic86.4Anthropic agent harness, max effort, <result> tags, 980k budget, no compaction2026-06-09Fable 5 / Mythos 5 card Fig 8.14.4.A, p. 269; score also in text, p. 268. Low / med / high / xhigh 76.7 / 80.6 / 82.8 / 84.9
Claude Mythos PreviewAnthropic83.6Same2026-06-09Same figure. Low / med / high 79.2 / 81.3 / 83.5, no xhigh point
Claude Opus 4.8Anthropic80.6Same2026-06-09Same figure. Low / med / high / xhigh 70.2 / 73.3 / 75.1 / 79.0
AnthropicAnthropic agent harness100 tasks, <result> tags, no compactionScore, Claude Opus 4.6 judgewww-cdn.anthropic.com

Comparable within this table only. Tag extraction, no compaction, Opus 4.6 judge.

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Mythos PreviewAnthropic83.7Anthropic agent harness, max effort, <result> tags, compaction at 200k2026-05-28Opus 4.8 card Fig 8.10.4.A, p. 208. Pass rate 74.7%
Claude Opus 4.8Anthropic80.4Same2026-05-28Score also in card text. Fig 8.10.4.C low / med / high / xhigh 70.5 / 75.0 / 75.7 / 78.9. Pass rate 71.5%
Claude Opus 4.7Anthropic77.7Same2026-05-2877.6 at max in the effort figure. Low / med / high / xhigh 72.3 / 75.0 / 77.3 / 78.8. Pass rate 69.4%
Claude Opus 4.6Anthropic76.5Same2026-05-28Effort figure low / med / high 61.4 / 69.9 / 74.9, no xhigh point, 75.0 at max. Pass rate 68.3%
Claude Sonnet 4.6Anthropic75.8Same2026-05-28Pass rate 68.2%
AnthropicAnthropic agent harness100 tasks, <result> tags, compaction at 200kScore, Claude Opus 4.6 judgewww-cdn.anthropic.com

Comparable within this table only. Tag extraction, compaction at 200k tokens, Opus 4.6 judge. The effort figure plots average tokens per task on its x axis, with ticks from 20k to 200k, not cost. Anthropic calls Opus 4.8 "a strict improvement over Opus 4.7 at max effort level," and Opus 4.7 is the one model here whose max-effort score sits below its xhigh score.

Failure modes and limitations

Judge dependence. Swapping the judge moves absolute scores by 10 to 25 points. In the paper, the top system scores 70.5 under Gemini-3-Pro, 50.4 under GPT-5.2 and 75.5 under Sonnet-4.5. The order of the top four held across all three judges, which is the reassuring half. The paper also says the judge may "not perfectly align with human expert preferences across all domains."

Self-grading. Anthropic grades Claude reports with a Claude judge. The cards report no check against a non-Claude judge for any 2026 model. The one 2026 run with a non-Claude judge is OpenRouter's, and it covers Fable 5 and fused panels, not Opus 5 or Opus 5.5.

Lost reports. Under the <result>-tag protocol, models at high effort sometimes skipped the tags and full reports scored as empty. Every card before the Opus 5 card of 24 July 2026 uses that protocol, so its high-effort numbers carry this risk.

Unexplained card-to-card drift. Opus 5 at max scores 88.6, 87.6 and 88.3 across three cards. Sonnet 5 scores 84.3 and then 80.8. Anthropic does not say whether the checkpoint, budget or harness changed, so these gaps cannot serve as a noise estimate. The 3.5-point Sonnet 5 swing is larger than the whole spread between the top three models.

System, not model. A DRACO score mixes the model with retrieval and orchestration. The paper's clearest example is Perplexity Deep Research on Opus 4.6 beating plain Opus 4.6 with web search and code execution by 10.7 points, 70.5 to 59.8, on the same model family.

Narrow input format. Tasks are static, single-turn, text-only and English-only. The agent cannot ask a clarifying question or read an image. The paper also warns that rewriting the queries "risks over-specifying tasks and dampening the natural variability of user queries," and names the human effort behind rubric writing as a scaling limit.

Content filters. In OpenRouter's run, filters blocked 7 of 100 tasks for Fable 5, which was scored on the remaining 93.

Contamination. All 100 tasks and full rubrics are public under MIT, and the set has not changed since February 2026. Neither the paper nor the dataset card reports a contamination analysis or a canary string. Nothing published rules out training exposure for models released since.

Missing data. OpenAI and Google DeepMind have not published DRACO scores for GPT-6 or Gemini 4. Anthropic gives cost only as an unlabeled log-scale figure with no token counts.

GAIA asks closed-ended questions with one checkable answer, and the DRACO paper names it among the benchmarks it set out to complement. DRACO has no single right answer. It grades a whole report against roughly 39 criteria.

Humanity's Last Exam also uses expert questions, but scores them by matching a short answer. DRACO measures long-form synthesis with citations and charges for errors.

BrowseComp, which Anthropic reports beside DRACO, "tests an agent's ability to find hard-to-locate information on the open web," in the Opus 5 card's words. It checks whether the agent finds one fact. DRACO checks the written deliverable.

WANDR is Perplexity's wide-search benchmark. Its 500 tasks ask the model to list tens to hundreds of entities with facts and citations. The Sonnet 5.5 card calls it the wide-search counterpart to DRACO and says Anthropic ran a modified version. DRACO is the deep side, with 100 tasks and one long rubric-graded report each.

The paper's Table 1 lists other open-ended deep research benchmarks, including DeepResearch Bench, ResearchRubrics, LiveResearchBench, DeepResearch Bench II and ReportBench. None uses production tasks.

Law and Medicine are two of DRACO's ten domains. Klu's pages on the Harvey legal agent benchmark and HealthBench cover those fields on their own, and GDPval-AA and SWE-bench cover other kinds of agent work. For the general method, see LLM evaluation, retrieval-augmented generation and frontier models.

What it means for teams choosing a model

The top three are a tie for practical purposes. Opus 5 leads at 88.3 in the Opus 5.5 card, 0.6 ahead of Fable 5.1 at 87.7 and 0.9 ahead of Opus 5.5 at 87.4. Those gaps are smaller than the drift the same model shows across Anthropic's cards. Price, latency and policy behavior should decide between them. Sonnet 5.5 at 87.0 sits 0.4 under Opus 5.5 in its own card's run.

Effort matters more than model choice below xhigh. This is the most useful finding in the cards. Opus 5.5 climbs from 72.5 at low to 83.9 at medium, a 10.8-point jump, then gains only 2.4 more from high to max. At low effort it trails every other configuration in its figure by 10.8 points or more. Do not run Opus 5.5 at low for research reports. Anthropic publishes no dollar or token cost per effort level, so you will need to price the tradeoff on your own tasks. For background on effort levels, see test-time compute.

Sonnet 5.5 only catches Opus 5.5 at the top efforts. In the Sonnet 5.5 card it scores 81.0 at high against Opus 5.5's 85.0, a 4.0-point gap. At xhigh the gap shrinks to 1.0, and at max to 0.4. At low and medium Sonnet 5.5 trails by 7.9 and 12.7 points. If you pick Sonnet 5.5 over Opus 5.5, run it at xhigh or max.

The scaffold moves scores as much as the model. Perplexity's agent beats plain Opus 4.6 by 10.7 points in the paper. In OpenRouter's run, a fused Fable 5 and GPT-5.5 panel beats solo Fable 5 by 3.7 points under the same judge and tools, and a budget panel of three cheaper models beats solo Opus 4.8 by 5.9. Opus 4.8 synthesized every fused panel there, so part of the gain is that synthesis step. Before you swap models, check whether your orchestration has more headroom.

Run it yourself with a fixed judge. Every published 2026 number is Anthropic grading Claude with Claude. No non-Claude-judge run exists for Opus 5 or Opus 5.5, and OpenAI and Google have published nothing. The 100 tasks are public. Run them with your own tools and one judge for every candidate, since the judge alone moves scores 10 to 25 points. Budget for wall-clock time too. Anthropic reports 13 to 71 minutes per task on Opus 5.5, and in the same card a five-agent team given half the single-agent latency budget matched the single-agent score at about 2.8 times the speed.

DRACO is graded against Perplexity's queries and rubrics. Your research workload has its own questions and its own costly mistakes. You can build rubric-graded evals like this one on your own research tasks in Klu and see results next to the public numbers on the LLM leaderboard. The research industry page covers research workflows in Klu.