GDPval-AA

Artificial Analysis's agentic version of OpenAI's GDPval, which runs 220 public tasks through the Stirrup harness and ranks models by pairwise judge preferences

Knowledge workElo

Stephen M. Walker II · Co-Founder / CEO · October 6, 2026 · 15 min read

What is GDPval-AA v2.1?

GDPval-AA is Artificial Analysis's agentic, model-versus-model version of OpenAI's GDPval. OpenAI introduced GDPval in a paper submitted October 5, 2025. Its tasks come from the representative work of industry professionals who average 14 years of experience, and they cover the majority of U.S. Bureau of Labor Statistics work activities for 44 occupations across the top 9 sectors by contribution to U.S. GDP. Earlier LLM benchmarks scored multiple choice, string matches or test suites. GDPval grades the finished work product, whether a document, spreadsheet, presentation or diagram, by blinded expert comparison.

OpenAI open-sourced a 220-task gold subset along with an automated grading service. Artificial Analysis (AA) takes those 220 tasks, runs any model on them inside its open-source Stirrup agent harness with shell access and web browsing, has frontier-model judges rank the anonymized deliverables against each other, and publishes Elo ratings with 95% confidence intervals. The current version is v2.1, per the AA changelog entry of September 19, 2026.

On the live board, retrieved October 6, 2026 with 282 models listed, Claude Opus 5.5 at Max effort leads at 1866 Elo. Eight of the top ten configurations are Anthropic's. GDPval-AA v2.1 carries a 10% weight in the AA Intelligence Index v4.3.2, down from 20% in v4.1.

Despite the name, nothing here is weighted by GDP or dollars. The score is a blind pairwise preference between deliverables, and every task counts the same.

How the tasks, harness and scoring work

The full OpenAI GDPval set has 1,320 tasks across 44 occupations and nine sectors, including finance and insurance, government, healthcare, manufacturing and information. Epoch AI describes the occupations as "the five highest-earning predominantly digital occupations" in each sector. In OpenAI's own grading, "a domain expert sees only the task and two unlabeled deliverables (the model's output and a human expert's), and ranks them without knowing which is which." Each task ends in a win, tie or loss against the human baseline, and OpenAI's leaderboard reports a win rate and a wins-plus-ties rate.

GDPval-AA keeps the 220 public tasks and changes the rest. A run has two stages.

  1. Task submission. The model works in an agentic loop through Stirrup, which AA calls "our open source reference agent harness", with a shell and a web browser. It produces files: documents, slides, diagrams and spreadsheets.
  2. Pairwise grading. An LLM judge sees two anonymized submissions for the same task and picks a winner. Since v2 the judges come from "a rotating panel of frontier-model judges."

AA fits Elo ratings from those pairwise results. Since v2.1 the fit uses a Crowd-BT model and pins DeepSeek V4.1 Flash at Max effort to 1600, so every other score is a distance from that one model. The board prints a 95% interval per row, shown as a half-width such as ±26.

AA states that "all evaluations are conducted independently by Artificial Analysis." For two Anthropic launches the numbers are AA-measured but vendor-published, because AA says it "supported Anthropic to evaluate Claude Opus 5 ahead of release" (AA Opus 5 article).

The index takes GDPval-AA scores "frozen at the time of a model's addition and normalized as clamp((Elo - 500) / 2000)". In Index v4.3.2 the Agents category is 30% of the total: AA-Briefcase v1.1 at 15%, GDPval-AA v2.1 at 10% and AutomationBench-AA at 5%. Coding and Scientific Reasoning are 20% each, and General is 30%.

AA has not published the wall-clock limit per task, a token cap, the sandbox network policy, the judge identities, the judge-panel size or any per-occupation scores.

Versions

VersionDateWhat changedScale
v1Not stated100-turn limitNot stated
v2June 15, 2026, with Index v4.1"Re-baselines ELO to human performance at 1000, introduces a rotating panel of frontier-model judges, and raises the turn limit from 100 to 250"; 20% of Index v4.1, "the highest weighted evaluation"Human expert = 1000
v2.1September 19, 2026"GDPval-AA Elo is now anchored to DeepSeek V4.1 Flash (max) at 1600, with updates to fitting methodology for GDPval-AA and AA-Briefcase to improve Elo stability"; same 220 tasks; Crowd-BT fittingDeepSeek V4.1 Flash (Max) = 1600

Sources: AA Index v4.1 article, AA changelog, AA methodology.

The v2 and v2.1 numbers sit on different scales. A v2 score of 1800 and a v2.1 score of 1800 mean different things, and v2.1 has no human-expert reference point at all.

How to read the run labels

Some rows on the board carry a label inside the model name. Opus 5.5, Sonnet 5.5 and Fable 5.1 rows read "(Max, Default Fallback)" or similar. The Fable 5 row reads "(Max, Opus 4.8 Fallback)". Every other row has no label. The page carries no legend or footnote that defines these strings.

AA's Opus 5 article says Opus 5 has "support for server-side fallback as with Fable 5. Intelligence Index evaluations were run with Opus 4.8 fallback enabled." Anthropic's Opus 5.5 launch page adds: "Claude Opus 5.5 was evaluated with its production safeguards enabled. When they intervened, cybersecurity tasks were completed by Claude Opus 4.8, and biology and frontier LLM development tasks were completed by Claude Opus 5. This likely reduces Claude Opus 5.5's performance on these benchmarks." On AutomationBench-AA, the same page data shows Opus 5.5 at Xhigh with 236 of 12,973 LLM turns and 21 of 657 tasks served by a fallback model.

Three things are missing:

  • AA does not say which model "Default Fallback" names.
  • AA publishes no count of GDPval-AA turns or tasks served by a fallback model.
  • The Opus 5 rows carry no label, yet AA's Opus 5 article says its Index runs had Opus 4.8 fallback enabled. A missing label does not prove a no-fallback run.

The tables below split rows by label, because a fallback-labeled row includes turns completed by a different model. A gap inside one table compares runs with the same label. A gap across tables is a position difference on one board.

GDPval-AA leaderboard (October 2026)

Every row below comes from the AA live board, retrieved October 6, 2026. All are AA-run on GDPval-AA v2.1 with Stirrup, shell and web, the same 220 tasks and the same rating scale. Rank is position by Elo among all 282 models.

Default Fallback rows

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 5.5 (Max)Anthropic1866 ±26Stirrup, Max, Default FallbackOct 6, 2026AA board, rank 1
Claude Sonnet 5.5 (Max)Anthropic1839 ±24Stirrup, Max, Default FallbackOct 6, 2026AA board, rank 2
Claude Opus 5.5 (Xhigh)Anthropic1837 ±25Stirrup, Xhigh, Default FallbackOct 6, 2026AA board, rank 3
Claude Fable 5.1 (Max)Anthropic1758 ±20Stirrup, Max, Default FallbackOct 6, 2026AA board, rank 4
Claude Fable 5.1 (Xhigh)Anthropic1735 ±21Stirrup, Xhigh, Default FallbackOct 6, 2026AA board, rank 5
Claude Sonnet 5.5 (Xhigh)Anthropic1731 ±26Stirrup, Xhigh, Default FallbackOct 6, 2026AA board, rank 6
Claude Opus 5.5 (High)Anthropic1707 ±23Stirrup, High, Default FallbackOct 6, 2026AA board, rank 10
Claude Fable 5.1 (High)Anthropic1635 ±20Stirrup, High, Default FallbackOct 6, 2026AA board, rank 18
Claude Opus 5.5 (Medium)Anthropic1586 ±22Stirrup, Medium, Default FallbackOct 6, 2026AA board, rank 34
Artificial AnalysisStirrupGDPval-AA v2.1, 220 tasksEloartificialanalysis.ai

Comparable to each other on one v2.1 scale; these rows include turns served by an unnamed fallback model, so gaps to unlabeled rows are board positions, not same-condition gaps.

Opus 5.5 Max leads Sonnet 5.5 Max by 27 and Opus 5.5 Xhigh by 29. Both gaps fall inside the ±24 to ±26 intervals, so the top three are not separable on this benchmark. Opus 5.5 Max leads Fable 5.1 Max by 108, which is well clear of the noise.

GDPval-AA v2.1 Elo vs. cost per task, Default Fallback rowsAA live board, October 6, 2026, Stirrup harness, 220 tasks, all Anthropic
175018001850$2$5$10
FrontierAnthropicBehind the frontier
Cost per task, log scale, lower to the right
Elo from the Artificial Analysis GDPval-AA board, retrieved October 6, 2026. AA does not print cost per task; each cost is derived from the token counts and prices in the board's page data, averaged over 220 tasks. Costs exclude cache-write premiums and fallback-model pricing, so they are lower bounds.

Unlabeled rows, ranks 8 to 30

ModelOrganizationScoreHarness and setupDateSource and notes
Grok 4.7 (Xhigh)SpaceXAI1715 ±23Stirrup, XhighOct 6, 2026AA board, rank 8
Grok 4.7 (High)SpaceXAI1710 ±20Stirrup, HighOct 6, 2026AA board, rank 9
MiMo-V2.6-ProXiaomi1686 ±23Stirrup, no effort listedOct 6, 2026AA board, rank 12
Muse Spark 1.3 (Max)Meta1684 ±20Stirrup, MaxOct 6, 2026AA board, rank 13
Qwen3.8 Max (0902)Alibaba1671 ±21Stirrup, no effort listedOct 6, 2026AA board, rank 14
GLM-5.3 (Max)Z.ai1653 ±20Stirrup, MaxOct 6, 2026AA board, rank 15
GLM 5.3 FlashZ.ai1647 ±18Stirrup, no effort listedOct 6, 2026AA board, rank 16
Grok 4.6 (Xhigh)SpaceXAI1647 ±20Stirrup, XhighOct 6, 2026AA board, rank 17
Qwen3.8-Flash-NextAlibaba1633 ±20Stirrup, no effort listedOct 6, 2026AA board, rank 19
Muse Spark 1.3 (Xhigh)Meta1630 ±20Stirrup, XhighOct 6, 2026AA board, rank 20
Gemini 4 Argon (High)Google1627 ±22Stirrup, HighOct 6, 2026AA board, rank 21
Ling 3.1 FlashInclusionAI1622 ±18Stirrup, no effort listedOct 6, 2026AA board, rank 22
Grok 4.6 (High)SpaceXAI1622 ±18Stirrup, HighOct 6, 2026AA board, rank 23
Grok 4.6 (Medium)SpaceXAI1617 ±20Stirrup, MediumOct 6, 2026AA board, rank 24
Qwen3.8 2.4T A95BAlibaba1613 ±20Stirrup, no effort listedOct 6, 2026AA board, rank 26
GPT-5.6 Sol (Max)OpenAI1611 ±17Stirrup, MaxOct 6, 2026AA board, rank 27; best OpenAI row
MiMo-V2.6-FlashXiaomi1611 ±23Stirrup, no effort listedOct 6, 2026AA board, rank 28
Qwen3.8 MaxAlibaba1608 ±19Stirrup, no effort listedOct 6, 2026AA board, rank 29
DeepSeek V4.1 Flash (Max)DeepSeek1600 (anchor)Stirrup, MaxOct 6, 2026AA board, rank 30; scale anchor, fixed at 1600
Artificial AnalysisStirrupGDPval-AA v2.1, 220 tasksEloartificialanalysis.ai

Comparable to each other on one v2.1 scale with no fallback label; not comparable to v2 scores, where human experts sat at 1000.

Grok 4.7 Xhigh at 1715 is the best non-Anthropic configuration. It leads Gemini 4 Argon High by 88 and GPT-5.6 Sol Max by 104. MiMo-V2.6-Pro is the surprise of this group, 29 points behind Grok 4.7 Xhigh at a small fraction of the cost.

GDPval-AA v2.1 Elo vs. cost per task, unlabeled rowsAA live board, October 6, 2026, Stirrup harness, 220 tasks, no fallback label
160016501700$0.2$0.5$1$2$5
FrontierOtherxAIBehind the frontier
Cost per task, log scale, lower to the right
Elo from the Artificial Analysis GDPval-AA board, retrieved October 6, 2026. AA does not print cost per task; each cost is derived from the token counts and list prices in the board's page data, averaged over 220 tasks. GPT-5.6 Sol Max is not plotted because the page data has no token counts for it.

Claude Opus 5

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 5 (Max)Anthropic1724 ±20Stirrup, Max, no labelOct 6, 2026AA board, rank 7
Claude Opus 5 (Xhigh)Anthropic1692 ±20Stirrup, Xhigh, no labelOct 6, 2026AA board, rank 11
Claude Opus 5 (High)Anthropic1596 ±20Stirrup, High, no labelOct 6, 2026AA board, rank 31
Artificial AnalysisStirrupGDPval-AA v2.1, 220 tasksEloartificialanalysis.ai

Same v2.1 scale as the tables above; AA's Opus 5 article says its Index runs had Opus 4.8 fallback enabled, so these are not no-fallback runs despite the missing label.

On the board, Opus 5.5 Max sits 142 points above Opus 5 Max. The two carry different labels, so read that as a board position, not a controlled same-condition gap.

Opus 4.8 Fallback row

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Fable 5 (Max)Anthropic1613 ±17Stirrup, Max, Opus 4.8 FallbackOct 6, 2026AA board, rank 25; June 2026 release
Artificial AnalysisStirrupGDPval-AA v2.1, 220 tasksEloartificialanalysis.ai

Same v2.1 scale; the only row that names its fallback model, Opus 4.8.

OpenAI flagships and other rows below rank 30

These are the OpenAI flagship rows below rank 30 plus Step 5 Preview and Kimi K3 Max, the two other rows in this range with turn data in the page payload.

ModelOrganizationScoreHarness and setupDateSource and notes
Step 5 PreviewStepFun1587 ±26Stirrup, no effort listedOct 6, 2026AA board, rank 33; 63.7 turns, derived $0.60 per task
GPT-6.1 Sol (Max)OpenAI1575 ±20Stirrup, MaxOct 6, 2026AA board, rank 35; 24.9 turns, derived $0.82 per task
GPT-5.6 Sol (Xhigh)OpenAI1572 ±18Stirrup, XhighOct 6, 2026AA board, rank 36; no turn or token data in payload
GPT-6 Astra (Max)OpenAI1542 ±25Stirrup, MaxOct 6, 2026AA board, rank 40; 24.2 turns, derived $4.21 per task
Kimi K3 (Max)Kimi1537 ±20Stirrup, MaxOct 6, 2026AA board, rank 41; 37.5 turns, derived $1.97 per task
GPT-6 Sol (Max)OpenAI1510 ±24Stirrup, MaxOct 6, 2026AA board, rank 44; 39.9 turns, derived $1.19 per task
Artificial AnalysisStirrupGDPval-AA v2.1, 220 tasksEloartificialanalysis.ai

Same v2.1 scale and no fallback label, so directly comparable to the unlabeled ranks 8 to 30 table.

OpenAI does poorly here. Its best row, GPT-5.6 Sol Max at 1611, ranks 27th. GPT-6.1 Sol Max leads GPT-6 Astra Max by 33, inside the intervals, and leads GPT-6 Sol Max by 65. Astra Max, OpenAI's newest flagship, ranks 40th of 282.

Effort ladders

Values are Elo with the 95% half-width, from the same board. Anthropic 5.5 and Fable 5.1 rows carry "Default Fallback"; Opus 5, OpenAI and Grok rows carry no label.

ModelLowMediumHighXhighMaxMax minus Low
Claude Opus 5.51235 ±221586 ±221707 ±231837 ±251866 ±26631
Claude Sonnet 5.51179 ±231324 ±231551 ±241731 ±261839 ±24660
Claude Fable 5.11469 ±201549 ±201635 ±201735 ±211758 ±20289
Claude Opus 51304 ±221492 ±201596 ±201692 ±201724 ±20420
GPT-6 Astra1366 ±261468 ±241485 ±261517 ±261542 ±25176
GPT-6.1 Sol1297 ±191433 ±191486 ±191510 ±181575 ±20278
GPT-6 Sol1204 ±241350 ±251396 ±231457 ±241510 ±24306
GPT-5.6 Sol1305 ±181422 ±181505 ±181572 ±181611 ±17306
Artificial AnalysisStirrupGDPval-AA v2.1, 220 tasksElo by effort settingartificialanalysis.ai

Grok 4.7 has Low 1589 ±24, High 1710 ±20 and Xhigh 1715 ±23, with no Medium or Max row. Grok 4.6 has Low 1409, Medium 1617, High 1622 and Xhigh 1647.

Within each row the label is constant, so the steps are same-condition comparisons; across rows the fallback split above still applies.

The Claude ladders are steep. Sonnet 5.5 at Low scores 1179, below every OpenAI Max row, and at Max it is second on the board. Astra's ladder is the flattest at 176 points, which says extra test-time compute buys Astra little on this task set. On the old v2 scale AA reported that Opus 5's "effort levels span 407 Elo points, with output token usage ranging around 8x from low to max effort"; the same ladder spans 420 on v2.1.

Anthropic's launch page claims: "At default effort (medium), Opus 5.5 beats GPT-6 Astra at max effort for about a fifth of the cost per task." The board agrees on the order, with Opus 5.5 Medium at 1586 against Astra Max at 1542. That 44-point lead is inside the ±22 and ±25 intervals. The page data has no token counts for Opus 5.5 Medium, so the "fifth of the cost" figure is Anthropic's alone.

Cost per task and turns

AA's page has a cost chart but does not print cost per task as a number. Every cost on this page is derived from the board's embedded page data. Turn counts sit at gdpvalBreakdown.avgTurns, and token totals over the 220 tasks sit at canonicalEvalTokenCounts.gdpval, which holds input, answer, reasoning and cacheableInput, the cache-hit part of input. Prices are USD per million tokens. Only 30 rows carry this breakdown.

The formula, per task in USD, is [(input - cacheableInput) x input price + cacheableInput x cache-hit price + (answer + reasoning) x output price] / 1,000,000 / 220. For Opus 5.5 Max that is (3,793,360,879 - 3,726,668,094) x $4 + 3,726,668,094 x $0.20 + (14,463,522 + 26,714,602) x $20, divided by 1,000,000 and by 220, which gives $8.34.

Anthropic costs are lower bounds. The payload has no GDPval cache-write token count, so cache-write premiums are missing, and fallback-model pricing is not included. AA also footnotes that refused tasks are estimated from other models' token use.

Default Fallback rows (all Anthropic)

LabelEloDerived cost / taskAvg turnsOutput tokens / taskFrontier
Opus 5.5 (Max)1866$8.3485.7187kFrontier
Sonnet 5.5 (Max)1839$8.4892.3285k
Opus 5.5 (Xhigh)1837$3.9358.992kFrontier
Fable 5.1 (Max)1758$8.8560.4106k
Fable 5.1 (Xhigh)1735$6.5450.779k
Sonnet 5.5 (Xhigh)1731$2.2743.693kFrontier
Artificial AnalysisStirrupGDPval-AA v2.1, 220 tasksEloartificialanalysis.ai

The cost frontier here runs Sonnet 5.5 Xhigh ($2.27, 1731), Opus 5.5 Xhigh ($3.93, 1837), Opus 5.5 Max ($8.34, 1866). Sonnet 5.5 Max and Fable 5.1 Max both cost more than Opus 5.5 Xhigh and score no higher. Opus 5.5 Xhigh costs 47% of Opus 5.5 Max.

Unlabeled rows

LabelEloDerived cost / taskAvg turnsOutput tokens / taskFrontier
Grok 4.7 (Xhigh)1715$4.5054.4136kFrontier
Grok 4.7 (High)1710$3.6848.8118kFrontier
MiMo-V2.6-Pro1686$0.1546.088kFrontier
Muse Spark 1.3 (Max)1684$1.6474.585k
Qwen3.8 Max (0902)1671$3.5479.4122k
GLM-5.3 (Max)1653$2.2270.094k
GLM 5.3 Flash1647$0.3080.9126k
Gemini 4 Argon (High)1627$1.2835.475k
DeepSeek V4.1 Flash (Max)1600$0.2572.1116k
GPT-6.1 Sol (Max)1575$0.8224.946k
Artificial AnalysisStirrupGDPval-AA v2.1, 220 tasksEloartificialanalysis.ai

Only three of these ten sit on the cost frontier: MiMo-V2.6-Pro ($0.15, 1686), Grok 4.7 High ($3.68, 1710) and Grok 4.7 Xhigh ($4.50, 1715). Paying 30 times more for Grok 4.7 Xhigh than MiMo-V2.6-Pro buys 29 Elo.

Two reference points sit outside both groups. Claude Opus 5 Max scores 1724 at $6.44 with 60.5 turns and 97k output tokens per task. GPT-6 Astra Max scores 1542 at $4.21 with 24.2 turns and 40k output tokens.

A few readings follow from these numbers. Opus 5.5 Max costs more than Opus 5 Max, $8.34 against $6.44, because it takes 85.7 turns against 60.5 and emits 187k output tokens against 97k. GPT-6.1 Sol Max costs one fifth of Astra Max and scores 33 higher, inside the intervals.

Low turn counts go with low scores. Among the 30 rows with turn data, five average 40 turns or fewer per task. They are Astra at 24.2, GPT-6.1 Sol at 24.9, Gemini 4 Argon at 35.4, Kimi K3 at 37.5 and GPT-6 Sol at 39.9, and all five score 1627 or lower. The top Claude configurations take 60 to 86. AA reports the turn gap and gives no cause for it.

On the full Intelligence Index, not GDPval alone, AA's GPT-6 Astra article puts Astra Max at 53, level with Fable 5.1 at Max with fallback, for $3.26 against $7.63 per Index task.

Historical results

Anthropic launch snapshot (September 22, 2026)

Anthropic's Opus 5.5 launch page published GDPval-AA v2.1 scores that AA measured. Anthropic states that all its results use "adaptive thinking at max effort" unless noted. It does not state the effort for the GPT rows.

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 5.5Anthropic1846Stirrup, adaptive thinking at max effortSep 22, 2026Anthropic launch page; live board 1866 (Max), +20
Claude Fable 5.1Anthropic1735Stirrup, adaptive thinking at max effortSep 22, 2026Anthropic launch page; live board 1758 (Max); 1735 is the board's Xhigh value
Claude Opus 5Anthropic1708Stirrup, adaptive thinking at max effortSep 22, 2026Anthropic launch page; live board 1724 (Max), +16
GPT-5.6 SolOpenAI1588Stirrup, effort not statedSep 22, 2026Anthropic launch page; live board 1611 (Max), +23
GPT-6 AstraOpenAI1542Stirrup, effort not statedSep 22, 2026Anthropic launch page; live board 1542 (Max), unchanged
AnthropicStirrupGDPval-AA v2.1Eloanthropic.com

Same v2.1 scale as the live board, but an earlier fit; AA refit the ratings after launch, so rows moved by 0 to 23 points.

The launch page says: "On GDPval-AA v2.1, a test of real-world work across 44 occupations, Opus 5.5 scores 1846 Elo, ahead of Fable 5.1 and Opus 5." The page also has an Elo-versus-cost chart, with Elo from 1200 to 1800 against "Estimated cost per task (USD, log scale)" from $0.20 to $10, but the point values are not in the page text.

GDPval-AA v2 results (June to July 2026)

DateModelOrganizationElo (v2)Source and notes
Jun 15, 2026Claude Fable 5 (with fallback)Anthropic1818AA Index v4.1 article
Jun 15, 2026Claude Opus 4.8Anthropic1638Same
Jun 15, 2026GPT-5.5 (xhigh)OpenAI1531Same
Jul 24, 2026Claude Opus 5 (max)Anthropic1861AA Opus 5 article; +114 over Fable 5, ">100 points ahead of Claude Fable 5 and GPT-5.6 Sol (max)"
Artificial AnalysisGDPval-AA v2Elo (v2)

v2 scale anchored at human experts = 1000; not comparable to any v2.1 table on this page.

AA's September 9, 2026 Astra article gives no absolute v2 score for GPT-6 Astra. It says: "Compared to GPT-5.6 Sol, we observe a drop of ~45 Elo points in GDPval-AA v2 - a benchmark we adapted from OpenAI's dataset measuring economically valuable tasks across 44 occupations. We observed GPT-6 Astra using significantly fewer turns than other models in GDPval tasks - 24 per task at max effort, compared to 45 for GPT-5.6 Sol and 60 turns per task for Claude Fable 5.1 and Claude Opus 5." On v2.1 the board points the same way, with Astra Max 69 below GPT-5.6 Sol Max, outside the intervals.

Timeline

  • October 5, 2025. OpenAI's GDPval paper reports that frontier performance "is improving roughly linearly over time", that the best models are "approaching industry experts in deliverable quality", and that "increased reasoning effort, increased task context, and increased scaffolding improves model performance."
  • June 15, 2026. GDPval-AA v2 ships with Index v4.1. Fable 5 with fallback leads at 1818.
  • July 24, 2026. Claude Opus 5 at max reaches 1861 on v2. The same article puts Opus 5 at 1720 on AA-Briefcase, 146 over Fable 5. Across the whole Intelligence Index, Opus 5 at max costs $2.03 per task against $2.75 for Fable 5 with fallback, $1.80 for Opus 4.8 at max and $1.53 for Sonnet 5 at max.
  • September 7, 2026. Index v4.3 moves Terminal-Bench to v4.0 and adds AutomationBench-AA.
  • September 14, 2026. AA-Briefcase is announced in Capability Indices v1.1.
  • September 19, 2026. GDPval-AA v2.1 changes the anchor and fitting method.
  • September 22, 2026. Anthropic's launch page shows Opus 5.5 at 1846 against Opus 5 at 1708 on the v2.1 snapshot, a 138-point gap.
  • October 6, 2026. The live board in the tables above, now including GPT-6.1 Sol, released September 29, at 1575.

Failure modes and limitations

A relative scale. Elo ranks models against each other around one anchor. On v2.1 that anchor is DeepSeek V4.1 Flash at 1600; on v2 it was human experts at 1000. A GDPval-AA score means nothing without a version and a date.

Version changes move the whole scale. Opus 5 Max is 1861 on v2 and 1724 on v2.1. The model did not get worse. The scale moved.

Judge dependence. Scores come from LLM judges picking winners. v2 switched to a rotating panel of frontier-model judges, and AA has not published who the judges are or why it made the change. See LLM-as-a-judge for the general problem.

Precision. The 95% intervals run ±17 to ±26 on the top 30 rows. The Xhigh-to-Max step falls inside the intervals for Opus 5.5 (+29) and Fable 5.1 (+23), and outside them for Sonnet 5.5 (+108). For Opus 5.5 and Fable 5.1, the board cannot tell Max from Xhigh. For Sonnet 5.5, it can.

Fallback runs. Rows labeled "Default Fallback" or "Opus 4.8 Fallback" include turns completed by another model when Anthropic's safeguards intervene. Anthropic says this "likely reduces" its own scores on the benchmarks where it triggers.

Refusals. AA footnotes that for some rows "the model or provider declined some tasks on safety grounds" and that "refused tasks are estimated from other models' token use." Derived costs for those rows are estimates on top of estimates.

Scope. The occupations are the five highest-earning predominantly digital ones per sector. Each task is a one-shot deliverable with a 250-turn cap. Nothing here tests multi-session work. AA-Briefcase covers "multi-week knowledge work projects, each with many linked tasks and thousands of input source files", and GDPval-AA does not.

One harness. OpenAI's paper says scaffolding changes results. GDPval-AA fixes Stirrup for every model, so scores describe Stirrup runs only. A vendor's own agent harness can score differently, and those results are not on this board.

Contamination. OpenAI open-sourced the 220 gold tasks, so they are public. AA has published no private GDPval-AA split and no contamination audit. Compare AutomationBench-AA, which uses a "657-task held-out split", and Index v4.2, which added "more private test sets to prevent gaming."

The vendor's own caution. Anthropic, whose models top the board, writes that "at these levels of capability we've found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest." The company that benefits most from the board is telling you not to read its margins literally.

OpenAI GDPval. The original has 1,320 tasks. Human experts grade blind against human-made deliverables and report win and win-or-tie rates against human work. GDPval-AA runs only the 220 public tasks, swaps expert graders for LLM judges, ranks models against each other instead of against humans, and runs every model in one open harness. GDPval asks whether a model matches a professional. GDPval-AA asks which model beats which.

AA-Briefcase v1.1. Briefcase uses the same Stirrup harness on multi-week projects with many linked tasks and thousands of input files. It carries 15% of Index v4.3.2 against GDPval-AA's 10%. If your work spans long projects with large document sets, Briefcase is the closer match.

AutomationBench-AA. This one tests SaaS workflows across simulated business apps through REST APIs, with a 50-turn cap, a 657-task held-out split and 5% of the index. Anthropic notes that Zapier ran its own AutomationBench figures without fallback models, so safeguard interventions counted as failures there.

Vals Index. Vals AI's index measures accuracy on verifiable finance, coding, legal and tax tasks, weighted by GDP as (8.0 Finance + 5.6 Coding + 1.2 Legal + 0.5 Tax) / 15.3. Its components include Finance Agent v2, Excel Modeling Benchmark, Terminal-Bench 4.0, Vibe Code Bench, Code Migration, Legal Research Bench, HLAB and Tax Agent Bench. The Vals Index does not include GDPval.

ModelOrganizationScoreHarness and setupDateSource and notes
Gemini 4 ArgonGoogle68.90%Vals Index, accuracyOct 2, 2026Vals board; $15.68 per test; 21st on GDPval-AA (1627 at High)
Claude Sonnet 5.5Anthropic67.04%Vals Index, accuracyOct 2, 2026Vals board; $21.34 per test
Claude Opus 5.5Anthropic66.97%Vals Index, accuracyOct 2, 2026Vals board; $32.14 per test
Claude Fable 5.1Anthropic65.83%Vals Index, accuracyOct 2, 2026Vals board; $28.71 per test
Claude Opus 5Anthropic63.67%Vals Index, accuracyOct 2, 2026Vals board; $19.31 per test
Vals AIVals Index accuracyvals.ai

Accuracy on verifiable answers, not pairwise Elo; compare ranks across the two boards, never the numbers.

Gemini 4 Argon is first on the Vals Index and 21st on GDPval-AA. Vals checks whether an answer is correct. GDPval-AA asks a judge which of two deliverables is better.

Terminal-Bench 4.0 and SWE-bench. Both score with code-verified test suites, not judged preference. GDPval-AA has no ground truth to check against, which is why it needs judges and Elo at all. Related agentic evaluations on this site include GAIA, DRACO for deep research and the Harvey legal agent benchmark.

What it means for teams choosing a model

Opus 5.5 leads for report, deck and spreadsheet work. It scores 1866 at Max on the independent board. The next non-Anthropic row is Grok 4.7 Xhigh at 1715, and the best OpenAI row is GPT-5.6 Sol Max at 1611.

Run Opus 5.5 at Xhigh, not Max. Xhigh scores 1837 for a derived $3.93 per task against $8.34 at Max, and the gap sits inside the intervals. Sonnet 5.5 Xhigh, at 1731 for $2.27, is the cheapest Anthropic row above 1700. Sonnet 5.5 Max costs $8.48, about what Opus 5.5 Max costs, for 27 fewer Elo, so skip it.

Benchmark at more than one effort level. Effort is the biggest lever on Claude models. Opus 5.5 spans 631 Elo from Low to Max, Sonnet 5.5 spans 660, Fable 5.1 spans 289 and Astra 176. A team that tests a model at one effort has tested one point on a ladder.

Among OpenAI models, pick GPT-6.1 Sol Max. It scores 1575 for a derived $0.82 per task, one fifth of Astra Max's cost and 65 above GPT-6 Sol Max. No OpenAI row exceeds 1611, which is 255 below Opus 5.5 Max. If you are committed to OpenAI for deliverable-style work, test GPT-6.1 Sol before Astra.

For high-volume, cost-bound work, test MiMo-V2.6-Pro. It scores 1686 for a derived $0.15 per task, the cheapest point on the unlabeled cost chart, and beats Gemini 4 Argon High at 1627 and $1.28.

Check your own judge. GDPval-AA ranks models by what frontier-model judges prefer. If your reviewers weigh accuracy over polish, the way the Vals Index does, your ranking will differ. Build a golden dataset of your own deliverables and keep a human in the loop for grading. The frontier models and LLM evaluation pages cover the wider setup.

The Klu LLM leaderboard tracks these models across benchmarks. To run the same kind of pairwise deliverable eval on your own reports and spreadsheets, see how teams set it up for operations workflows in Klu.