Chartography and LVBench

Two multimodal benchmarks. Chartography is Surge AI's 100-task test of professional chart reading, and LVBench asks questions about videos that run past an hour

Long context and multimodalMean pass@1

Stephen M. Walker II · Co-Founder / CEO · October 6, 2026 · 15 min read

What are Chartography and LVBench?

This page covers two multimodal benchmarks. Chartography tests whether a model can read the charts professionals make decisions from. LVBench tests question answering over videos that run past an hour. Chartography gets most of the space because it has the 2026 results.

Chartography is a 100-task benchmark built by Surge AI. Suhaas Garre, Chris Mutty, Sushant Mehta and Edwin Chen published the paper on arXiv on 11 August 2026 as arXiv:2608.10677, under CC BY 4.0. Domain professionals wrote every task around a chart type that shows up in real work: Kaplan-Meier survival curves, candlestick charts, contour maps, wind roses, Sankey diagrams, Bode plots and 3D surface plots. To answer, the model has to estimate an unlabeled value, separate overlapping traces, or apply a convention of the field that the question does not spell out. Each task passed three independent expert reviews and a difficulty screen against frontier models.

Surge built it because chart benchmarks had stopped telling models apart. Models that scored 80 to 90% on earlier chart sets scored below 50% on Chartography. In the paper's snapshot of 16 July 2026 the best configuration, GPT-5.6 Sol at max reasoning, reached 45.0%. The earlier sets test bar, line and pie charts and academic figures. Chartography tests the formats an engineer, clinician or analyst reads before signing off on something.

LVBench is a long-video question-answering benchmark from Weihan Wang and colleagues at Zhipu AI and Tsinghua, first released in June 2024 as arXiv:2406.08035 and revised to v3 on 9 August 2025. It has 103 videos totaling 117 hours, an average of 4,101 seconds each, and 1,549 four-option multiple-choice questions. Most video benchmarks before it used clips under a minute. In the v1 paper the best model, Gemini 1.5 Pro, scored 33.1% against 94.4% for humans.

How Chartography works

Each task is one user message holding the prompt and the chart image inline. There is no system prompt, no few-shot example and no image normalization, so the model gets the image exactly as the task author supplied it.

The 100 tasks spread across 12 domains:

DomainTasks
STEM general19
Manufacturing / supply chain14
General engineering11
Mechanical engineering10
Finance / investing10
Healthcare10
Electrical engineering9
Chemistry6
Civil / environmental4
Geosciences4
Biology2
Physics1

Engineering and manufacturing dominate. Biology has two tasks and physics one, so a strong Chartography score says little about lab science figures.

Surge's protocol forbids tools. The model cannot run code, crop, zoom, browse or pull external data. It writes a free-form answer, and the score is mean pass@1. Surge runs the benchmark on Inspect AI, with the dataset on Hugging Face. The repo README sets the official leaderboard configuration as --epochs 10 with -T judge_model=google/gemini-3.5-flash and an optional --reasoning-effort. The paper itself ran every task 20 times, 2,000 evaluations per configuration. That split between 20 trials and 10 epochs is why the paper's numbers and the live leaderboard sit in separate tables below.

A typical task asks for the 50th-percentile height of a four-year-old girl from a growth chart, rounded to 0.5 cm. Any answer from 100 to 101 cm passes. The range is the point. An expert reading the same chart lands somewhere in that band, and Chartography grades the model against that tolerance.

Scoring and the judge

An LLM judge, Gemini 3.5 Flash, receives the question, the golden answer and the model's response. It never sees the chart. Its only job is to decide whether the response is equivalent to the golden answer.

The answer key mixes three kinds of answer. 51 answers carry numeric ranges calibrated by experts, 25 are exact numbers and 24 are categorical. 21 tasks have more than one range. Multi-part answers are all-or-nothing, so getting three of four values right scores zero. Calling a question unanswerable counts as wrong unless the golden answer allows it, which closes off the easy escape of refusing hard reads.

Gemini 3.5 Flash is also a model under test. The paper says so directly: "The judge is itself an evaluated configuration; Gemini 3.5 Flash's leaderboard row is therefore a self-judging condition." The paper's named mitigation is that the judge never sees the chart. Gemini 3.5 Flash does not appear among the 42 rows on the live page. It does appear in the paper's July snapshot at 35.9%.

Reasoning effort

The paper ran 12 models at both default and elevated reasoning. Elevated means OpenAI "max" or "xHigh", xAI "high" and Anthropic "adaptive/max". 11 of the 12 scored higher at the elevated setting. The median gain was 4.5 points, the largest 15.5 and the smallest -0.6. That is a modest return for the extra tokens, and the failure-mode section explains why.

Anthropic's with-tools protocol

Anthropic runs its own version for Claude system cards. Claude models use adaptive thinking at max effort, with and without tools, averaged over five runs. With tools, the model gets a container holding the image file and standard libraries, plus an image-cropping tool. Anthropic uses Gemini 3.5 Flash as judge to match Surge. Earlier Anthropic cards used Claude Sonnet 4.6 as judge for Claude models only, and those runs scored slightly lower.

How LVBench works

LVBench asks four-option multiple-choice questions about long videos, so random guessing scores 25%. The questions cover six capabilities: temporal grounding, summarization, reasoning, entity recognition, event understanding and key information retrieval. People annotated the questions by hand. The authors then used two LLMs, GLM-4 among them, to drop any question a model could answer without watching the video.

The videos come from YouTube, and the repo ships download scripts keyed to video IDs rather than the files. The license is CC-BY-NC-SA-4.0, academic use only. The metric is overall accuracy across all 1,549 questions.

The protocol changed between versions. v1 gave Gemini 1.5 Pro 3,600 input frames, with its public interface capping video at one hour. v3 sets 1 FPS for models that take long video natively, downsampling only when a video exceeds the model's maximum input, and gives models without native long-video support a fixed 32 or 96 sampled frames. Frame budget decides how much of each video the model sees, the video version of a context window limit.

Versions

Chartography has one public version. The paper is a frozen snapshot dated 16 July 2026, with 30 configurations across 18 models, 20 trials per task and 95% confidence intervals of about plus or minus one point. The live leaderboard is a later, separate snapshot with 42 rows. Surge publishes no changelog. On 6 October 2026 the repo had no releases and a single commit, dated 16 July 2026. The README states that all 100 tasks defeated at least one of two frontier models during curation and that no frontier model scored above 45% at release.

LVBench has v1, June 2024, which reports Gemini 1.5 Pro as the best model, and v3, 9 August 2025, which adds Gemini 2.5 Pro, Seed1.5-VL, MR.Video, GPT-4.1 and others under the 1 FPS rule.

Current leaderboard

Chartography results come in three groups that do not merge into one ranking: Surge's live leaderboard, Anthropic's no-tools runs and Anthropic's with-tools runs. LVBench results come from the authors' August 2025 runs, split by input budget.

Chartography, Surge live leaderboard, no tools

Same dataset and protocol on every row, run by the benchmark owner. Not comparable to the paper snapshot, which used 20 trials, or to any Anthropic table.

Surge labels scores mean pass@1. The judge and the 10-epoch run count come from the repo README, not the page, and the page publishes no confidence intervals. The page prints every rank as "1", so rows here follow score order. Surge prints no organization column, so this one names a lab only where the model family identifies it.

ModelOrganizationScoreHarness and setupDateSource and notes
Gemini 4 ArgonGoogle71.6%No tools, High2026-10-06Surge leaderboard. Leads GPT-6 Astra by 0.6.
GPT-6 AstraOpenAI71.0%No tools, Max2026-10-06Surge leaderboard. 4.7 ahead of Opus 5.5.
Claude Opus 5.5Anthropic66.3%No tools, Adaptive/Max2026-10-06Surge leaderboard. 20.1 ahead of Fable 5.1.
GPT-6 SolOpenAI53.6%No tools, Max2026-10-06Surge leaderboard. 17.4 behind Astra.
Claude Fable 5.1Anthropic46.2%No tools, Adaptive/Max2026-10-06Surge leaderboard
GPT-5.6 SolOpenAI45.0%No tools, Max2026-10-06Surge leaderboard. Same score as in the July paper snapshot.
Gemini 3.8 FlashGoogle40.9%No tools, High2026-10-06Surge leaderboard
Gemini 3.7 FlashGoogle40.4%No tools, High2026-10-06Surge leaderboard
Claude Fable 5Anthropic34.8%No tools, Adaptive/Max2026-10-06Surge leaderboard
Gemini 3.6 FlashGoogle34.0%No tools, Medium2026-10-06Surge leaderboard
GPT-5.6 TerraOpenAI34.0%No tools, Max2026-10-06Surge leaderboard
Muse Spark 1.2Not listed32.1%No tools, xHigh2026-10-06Surge leaderboard
GPT-5.6 LunaOpenAI31.4%No tools, Max2026-10-06Surge leaderboard
GPT-5.5OpenAI31.0%No tools, xHigh2026-10-06Surge leaderboard
GPT-5.4OpenAI29.6%No tools, xHigh2026-10-06Surge leaderboard
GPT-6 LunaOpenAI29.1%No tools, Max2026-10-06Surge leaderboard
Qwen 3.8 MaxNot listed29.1%No tools, xHigh2026-10-06Surge leaderboard
Muse Spark 1.3Not listed27.6%No tools, xHigh2026-10-06Surge leaderboard
Claude Opus 5Anthropic27.3%No tools, Adaptive/Max2026-10-06Surge leaderboard. 39.0 behind Opus 5.5.
Kimi K3Not listed26.6%No tools, Max2026-10-06Surge leaderboard
Gemini 3.1 ProGoogle26.1%No tools, High2026-10-06Surge leaderboard
Muse Spark 1.1Not listed24.4%No tools, xHigh2026-10-06Surge leaderboard
Grok 4.6Not listed20.5%No tools, xHigh2026-10-06Surge leaderboard
Qwen 3.8 FlashNot listed17.4%No tools, xHigh2026-10-06Surge leaderboard
Muse Glimmer 30BNot listed17.2%No tools, xHigh2026-10-06Surge leaderboard
Grok 4.5Not listed17.0%No tools, High2026-10-06Surge leaderboard
Claude Sonnet 5Anthropic16.6%No tools, Adaptive/Max2026-10-06Surge leaderboard
Claude Opus 4.7Anthropic16.5%No tools, Adaptive/Max2026-10-06Surge leaderboard
GLM 5.3 FlashNot listed16.3%No tools, Max2026-10-06Surge leaderboard
Claude Opus 4.8Anthropic15.9%No tools, Adaptive/Max2026-10-06Surge leaderboard
Qwen 3.5 PlusNot listed15.9%No tools, Thinking on2026-10-06Surge leaderboard
Qwen 3.7 PlusNot listed15.8%No tools, Thinking on2026-10-06Surge leaderboard
Grok 4.7Not listed14.7%No tools, xHigh2026-10-06Surge leaderboard
Grok 4.3Not listed12.9%No tools, High2026-10-06Surge leaderboard
Kimi K2.5Not listed12.6%No tools, Thinking on2026-10-06Surge leaderboard
Kimi K2.6Not listed12.2%No tools, Thinking on2026-10-06Surge leaderboard
DeepSeek V4.1 FlashDeepSeek11.6%No tools, Max2026-10-06Surge leaderboard
DeepSeek V4 Flash Vision (experimental)DeepSeek11.2%No tools, Max2026-10-06Surge leaderboard
Gemini 3.5 Flash-LiteGoogle10.2%No tools, Minimal2026-10-06Surge leaderboard
InklingNot listed10.0%No tools, High2026-10-06Surge leaderboard
Inkling SmallNot listed9.1%No tools, xHigh2026-10-06Surge leaderboard
Mistral Large 3Mistral9.0%No tools, none listed2026-10-06Surge leaderboard. No reasoning setting listed.
Surge AIInspect AI, no toolsChartography, 100 tasks, 10 epochsMean pass@1surgehq.ai

Three models have broken away. Gemini 4 Argon leads at 71.6% at high reasoning, with GPT-6 Astra 0.6 points behind at max. Surge publishes no confidence interval for these 10-epoch rows, so the order between those two rests on a 0.6-point gap with no published error bar. Claude Opus 5.5 is third at 66.3%, 5.3 points behind Argon and 4.7 behind Astra. Then the floor drops. GPT-6 Sol, the next model, trails Argon by 18.0 and Astra by 17.4, and Opus 5.5 sits 20.1 points clear of Claude Fable 5.1.

Chartography, Anthropic harness, no tools

Same card, judge and five-run mean on every row, following Surge's no-tools rules. Not comparable to the with-tools table below or to Surge's 10-epoch leaderboard, which is a different run set.

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 5.5Anthropic64.4%No tools, adaptive thinking, max effort, five runs, Gemini 3.5 Flash judge2026-09-22Opus 5.5 system card, section 8.13.1, p. 200. Vendor-reported.
Claude Fable 5.1Anthropic44.8%Same as above2026-09-22Opus 5.5 card, p. 200. Vendor-reported.
Claude Opus 5Anthropic29.8%Same as above2026-09-22Opus 5.5 card, p. 200. Vendor-reported.
AnthropicAnthropic harness, no toolsChartographyFive-run mean, Gemini 3.5 Flash judgeanthropic.com

Chartography, Anthropic harness, with tools

Same card, judge and five-run mean on every row. Surge's protocol forbids tools, so these scores compare to nothing else on this page.

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 5.5Anthropic89.0%Container with image file and standard libraries, crop tool, adaptive thinking, max effort, five runs2026-09-22Opus 5.5 card, p. 200. Vendor-reported. 24.6 above its no-tools score.
Claude Fable 5.1Anthropic88.4%Same as above2026-09-22Opus 5.5 card, p. 200. Vendor-reported. 43.6 above its no-tools score.
Claude Opus 5Anthropic83.4%Same as above2026-09-22Opus 5.5 card, p. 200. Vendor-reported. 53.6 above its no-tools score.
AnthropicAnthropic harness, with toolsChartographyFive-run mean, Gemini 3.5 Flash judgeanthropic.com

The card plots 95% CI whiskers but states no CI numbers in its text, and it does not say whether the 0.6-point gap between Opus 5.5 and Fable 5.1 with tools is significant. Opus 5.5 leads Opus 5 by 5.6. Sonnet 5 also appears in the card's figure, but only as image bars, so it is left out.

No with-tools score exists for any OpenAI or Google model. Google's Gemini 4 Argon methodology says its Chartography figure is "without tools and taken from the official Surge public leaderboard". OpenAI's GPT-6 Astra system card has no Chartography score at all.

Anthropic makes two cost claims without publishing the numbers behind them. The Opus 5 card says using the models' coding ability to manipulate, analyze and crop images "can be significantly more cost-effective than simply enabling adaptive thinking". The Opus 5.5 card says that with tools Opus 5.5 "Pareto-dominates prior Claude models on the score-cost frontier". Anthropic's figures plot score against dollars per task without a data table, so this page has no cost chart.

LVBench, native long-video models

Authors' own runs from arXiv v3. Same protocol on every row: 1 FPS, downsampled only when the video exceeds the model's maximum input. Not comparable to the sampled-frame table below. Humans score 94.4%.

ModelOrganizationScoreHarness and setupDateSource and notes
Gemini-2.5-ProGoogle67.4%1 FPS, native long video2025-08-09LVBench v3, Table 2. 27.0 behind humans.
Seed1.5-VL-ThinkingNot listed64.6%1 FPS, native long video2025-08-09LVBench v3, Table 2
Seed1.5-VLNot listed64.0%1 FPS, native long video2025-08-09LVBench v3, Table 2
MR.Video (Gemini-2.0-Flash based)Not listed60.8%1 FPS, native long video2025-08-09LVBench v3, Table 2
Gemini-2.5-FlashGoogle56.7%1 FPS, native long video2025-08-09LVBench v3, Table 2
AdaReTaKe (Qwen2.5-72B based)Not listed53.3%1 FPS, native long video2025-08-09LVBench v3, Table 2
Gemini-2.0-FlashGoogle48.6%1 FPS, native long video2025-08-09LVBench v3, Table 2
Qwen2.5-VL-72BNot listed44.0%1 FPS, native long video2025-08-09LVBench v3, Table 2
LVBench authorsLVBench v3, 1 FPSOverall accuracyarxiv.org

Gemini 2.5 Pro leads Seed1.5-VL-Thinking by 2.8 points. The per-category breakdown in Table 3 covers only Gemini-2.5-Pro and Seed1.5-VL, and it shows the overall lead is not uniform. Gemini scores 71.3 on Sports and 54.3 on Documentary. Seed1.5-VL scores 63.3 on Sports and 67.5 on Documentary, ahead of Gemini there by a wide margin. For documentary footage, the overall ranking points at the wrong model.

LVBench, sampled-frame models

Authors' own runs from arXiv v3. These models got a fixed 32 or 96 sampled frames, a different input budget from the native table, so do not rank them against it.

ModelOrganizationScoreHarness and setupDateSource and notes
GPT-4.1OpenAI60.1%32 or 96 sampled frames2025-08-09LVBench v3, Table 2. No per-capability breakdown.
GPT-4o-20241120OpenAI48.9%32 or 96 sampled frames2025-08-09LVBench v3, Table 2
GLM4V-PlusNot listed48.7%32 or 96 sampled frames2025-08-09LVBench v3, Table 2
VideoLLaMA3-7BNot listed45.3%32 or 96 sampled frames2025-08-09LVBench v3, Table 2
LVBench authorsLVBench v3, 32 or 96 sampled framesOverall accuracyarxiv.org

The paper states the frame count per class of model, not per model, so which of these got 32 frames and which got 96 is not published.

LVBench in 2026

Google's Gemini 4 Argon evaluation PDF, dated "as of October, 2026", reports LVBench, but its score cells are an embedded image with no text layer, and no other text source carries them. The PDF's text footnote does survive: "LVBench results are self computed without tools. We use 1FPS for Gemini, 800 frames for GPT-6 Astra and 300 frames for fable 5.1 and 600 for opus 5.5 due to API limitations."

That footnote means each vendor's model saw a different amount of each video. At the 4,101-second average length, 1 FPS works out to about 4,100 frames per video, a figure derived here rather than published by Google. GPT-6 Astra saw 800, Claude Opus 5.5 saw 600 and Claude Fable 5.1 saw 300. Whatever the image says, it is not a same-input comparison.

Historical progression

Chartography paper snapshot, 16 July 2026

Surge's own runs with 20 trials per task, no tools and the Gemini 3.5 Flash judge. An earlier snapshot than the live leaderboard with a different run count, so use it for confidence intervals and the reasoning-effort gain, not for ranking against live rows.

ModelOrganizationScoreHarness and setupDateSource and notes
GPT-5.6 SolOpenAI45.0%No tools, max, 20 trials2026-07-16Chartography paper. 95% CI +/-1.1.
GPT-5.6 SolOpenAI39.5%No tools, default, 20 trials2026-07-16Paper. CI +/-1.2.
Gemini 3.5 FlashGoogle35.9%No tools, default, 20 trials2026-07-16Paper. CI +/-1.3. Self-judging row.
Claude Fable 5Anthropic34.8%No tools, adaptive/max, 20 trials2026-07-16Paper. CI +/-1.0.
Claude Fable 5Anthropic29.5%No tools, default, 20 trials2026-07-16Paper. CI +/-1.1.
GPT-5.6 TerraOpenAI34.0%No tools, max, 20 trials2026-07-16Paper. CI +/-1.2.
GPT-5.5OpenAI31.0%No tools, xHigh, 20 trials2026-07-16Paper. CI +/-1.2.
GPT-5.5OpenAI28.5%No tools, default, 20 trials2026-07-16Paper. CI +/-1.3.
Surge AINo tools, 20 trialsChartography, 100 tasks, 20 trialsMean pass@1arxiv.org

The paper's lowest rows include Mistral Large 3 at 9.0%. Across the July and October snapshots, the top score went from 45.0% to 71.6%. That 26.6-point gap is between two snapshots with different run counts, not a same-protocol measurement. GPT-5.6 Sol scores 45.0 in both, which is the one direct check the two snapshots allow.

On the live page, Fable 5.1 is 11.4 points above Fable 5 and GPT-6 Astra is 26.0 above GPT-5.6 Sol.

Anthropic system cards over time

Each card is its own run set and judge, so compare models only within one table. The Opus 5.5 card says earlier cards "showed slightly lower scores due to using Claude Sonnet 4.6 as a judge only for Claude models", which makes the July and early September cards Sonnet 4.6-judged.

DateCardJudgeModelNo toolsWith tools
2026-07-24Opus 5 card, 8.12.1Claude Sonnet 4.6Claude Opus 4.817.075.0
2026-07-24Opus 5 cardClaude Sonnet 4.6Claude Opus 529.683.0
2026-07-24Opus 5 cardClaude Sonnet 4.6Claude Mythos 536.085.2
2026-09-01Fable 5.1 & Mythos 5.1 card, 8.14.1Claude Sonnet 4.6Claude Opus 529.683.0
2026-09-01Fable 5.1 & Mythos 5.1 cardClaude Sonnet 4.6Claude Fable 536.684.2
2026-09-01Fable 5.1 & Mythos 5.1 cardClaude Sonnet 4.6Claude Fable 5.142.686.2
2026-09-22Opus 5.5 card, 8.13.1Gemini 3.5 FlashClaude Opus 529.883.4
2026-09-22Opus 5.5 cardGemini 3.5 FlashClaude Fable 5.144.888.4
2026-09-22Opus 5.5 cardGemini 3.5 FlashClaude Opus 5.564.489.0
AnthropicChartographyNo tools score

The with-tools column has been crowded at the top since July. Every Claude model from Opus 4.8 on scores 75 or above with tools, and the no-tools column is where releases separate.

LVBench, June 2024 to August 2025

v1 baseline, 3,600 input frames. An earlier protocol and version, which v3 Table 2 reprints unchanged next to its 1 FPS rows.

ModelOrganizationScoreHarness and setupDateSource and notes
Gemini-1.5-ProGoogle33.1%3,600 frames; public interface capped video at 1 hour2024-06LVBench v1, Table 1. Humans 94.4%.
LVBench authorsLVBench v1, 3,600 framesOverall accuracyarxiv.org

In June 2024 the leader was Gemini 1.5 Pro at 33.1%. In August 2025 it was Gemini 2.5 Pro at 67.4%, with Seed1.5-VL-Thinking at 64.6%, both at 1 FPS. The protocols differ, so the two scores do not measure an improvement.

Documented failure modes

The paper and Surge's blog post of 4 October 2026 name five recurring errors:

  1. Missed visual elements. Thin, faint or overlapping features drop out of the model's reading.
  2. Right feature, wrong value. The model finds the correct curve and reads its peak "two gridlines too low".
  3. Complex geometry. 3D surface plots score lowest. Models ignore the projected gridlines.
  4. Unstated domain conventions. A task expects the reader to use the curve's lowest plotted value when extrapolating, and the model does something else.
  5. Error compounding. One misread value inside otherwise correct reasoning pushes the final answer outside the accepted range.

All five are perception errors, and more thinking does not fix them. Some max-reasoning configurations produce 20,000 to 30,000+ tokens without matching accuracy gains. From trace review, the paper reports that "longer reasoning often preserves an incorrect visual observation rather than correcting it." The model misreads the chart early and then reasons carefully from the wrong number. That explains the 4.5-point median reasoning gain, and why tools work. A crop tool lets the model look again at the region it misread. Opus 5.5 goes from 64.4% to 89.0% with tools on Anthropic's harness.

Chartography limitations

  • Small task count. 100 tasks. The paper's CIs of +/-1.0 to +/-1.3 come from repeated trials on the same 100 tasks, so they measure run noise, not task-sampling noise. They say nothing about how scores would shift on a different 100 charts.
  • Self-judging. Gemini 3.5 Flash judges every row and is itself on the paper's leaderboard. Anthropic's cards before 22 September judged Claude models with Claude Sonnet 4.6 instead.
  • No contamination check. The paper publishes no contamination test or hold-out set. It says tasks were "screened for difficulty against frontier models". The dataset is public on Hugging Face and the expert answer ranges sit in the GitHub repo.
  • Run-count split. The paper used 20 trials, the repo's leaderboard config uses 10 epochs, and the live page states neither.
  • No version trail. Nothing documents what changed between the July snapshot and the live page.

LVBench limitations

  • Guessing floor. Four options mean a 25% floor, so a 44% score is closer to chance than it looks.
  • Link rot. Videos are YouTube IDs and disappear over time, which changes the test set for anyone re-running it. The license is non-commercial, academic use only.
  • No contamination analysis. Videos and questions have been public since June 2024.
  • Mixed protocols in one table. v3 Table 2 lists Gemini-1.5-Pro at 33.1 from the v1 run next to v3-protocol rows.
  • Unequal frame budgets in 2026 results. Google's October 2026 evaluation feeds each vendor a different number of frames, and its scores have no text source.

The paper's own comparison table puts Chartography next to three established chart sets. ChartQAPro has 1,341 charts and 1,948 questions on bar, line and pie charts and dashboards, with a top score of 72%. ChartMuseum has 928 charts and 1,162 questions on standard types and infographics, top 86%. CharXiv has 2,323 academic figures and 11,615 questions, top 93%. Chartography is far smaller than any of them and much harder. In the July snapshot no model passed 50%. ParseBench, also in the paper's table, covers 568 pages of finance and insurance PDFs with 4,864 checks and a top score of 65%. It measures document parsing. Chartography hands over one chart image and asks one decision-shaped question.

BenchCAD sits next to Chartography in the Multimodal section of Anthropic's cards. It grades CadQuery code generated from multi-view renders of mechanical parts, 17,900 programs across 106 part families. Chartography grades a number or a category.

OSWorld 2.0 scores computer use from screenshots on task success. Chartography scores one answer about one static image. Google's Argon table lists both. For broad multimodal reasoning across academic subjects, see MMMU. Sibling pages in this set cover ScreenSpot-Pro, Humanity's Last Exam, MRCR v2 and GraphWalks.

LVBench's paper contrasts it with Video-MME, EgoSchema and MLVU, which use shorter videos. LVBench averages 4,101 seconds across 103 videos.

Background reading: LLM evaluation, LLM benchmarks, frontier models and computer vision.

What Chartography and LVBench mean for teams choosing a model

Without tools, pick Gemini 4 Argon or GPT-6 Astra. They lead Surge's leaderboard at 71.6% and 71.0%, ahead of Claude Opus 5.5 at 66.3% by 5.3 and 4.7 points and ahead of GPT-6 Sol at 53.6% by 18.0 and 17.4. The 0.6 points between Argon and Astra carries no published confidence interval, so treat them as a pair. No cost-per-task or latency figure exists for either on Chartography, so measure both on your own charts.

If your pipeline can run code and crop images, give the model tools. That is the largest lever anyone has documented on this benchmark. Opus 5.5 goes from 64.4% to 89.0% on Anthropic's harness, a 24.6-point gain. Fable 5.1 gains 43.6 and Opus 5 gains 53.6. No same-protocol comparison of that 89.0% to Gemini or GPT exists, because OpenAI and Google have published no with-tools score.

Among Claude models with tools, Opus 5.5 and Fable 5.1 are close. Opus 5.5 scores 89.0% and Fable 5.1 88.4%, 0.6 apart with no published CI. Opus 5.5 leads Opus 5 by 5.6. Anthropic says Opus 5.5 Pareto-dominates prior Claude models on score and cost with tools, but per-effort scores and costs are not published as numbers, so test effort settings on your own charts.

Do not buy reasoning effort to fix chart reading. Without tools, max reasoning over default gives a median 4.5 points, ranging from 15.5 down to -0.6, and some configurations burn 20,000 to 30,000+ tokens for it. That spend is test-time compute aimed at the wrong problem. The model needs a second look at the image, not more thinking about its first look.

Ignore CharXiv-style scores for professional charts. Models at 80 to 90% on earlier chart sets scored under 50% on Chartography in July. Build a golden dataset from your own engineering, clinical or finance charts, with acceptable ranges set by the people who read them, the way Chartography does. Grade with a judge that sees only the question, the key and the answer.

For long video, the confirmed leader is still Gemini 2.5 Pro from August 2025. It scored 67.4% among native long-video models at 1 FPS, 2.8 points ahead of Seed1.5-VL-Thinking at 64.6% and 27.0 behind humans at 94.4%. No confirmed 2026 LVBench score exists, because Google's October figures sit in an image with no text layer. Google's own footnote confirms each vendor got a different frame budget. Run your own clips at the frame budget you can afford before buying on any vendor's number.

To run the same kind of eval on your own charts or footage, build the test set in Klu and compare models on the Klu LLM leaderboard. For chart-heavy reporting work, start from the data analysis page.