What is GDPval-AA v2.1?
GDPval-AA is Artificial Analysis's agentic, model-versus-model version of OpenAI's GDPval. OpenAI introduced GDPval in a paper submitted October 5, 2025. Its tasks come from the representative work of industry professionals who average 14 years of experience, and they cover the majority of U.S. Bureau of Labor Statistics work activities for 44 occupations across the top 9 sectors by contribution to U.S. GDP. Earlier LLM benchmarks scored multiple choice, string matches or test suites. GDPval grades the finished work product, whether a document, spreadsheet, presentation or diagram, by blinded expert comparison.
OpenAI open-sourced a 220-task gold subset along with an automated grading service. Artificial Analysis (AA) takes those 220 tasks, runs any model on them inside its open-source Stirrup agent harness with shell access and web browsing, has frontier-model judges rank the anonymized deliverables against each other, and publishes Elo ratings with 95% confidence intervals. The current version is v2.1, per the AA changelog entry of September 19, 2026.
On the live board, retrieved October 6, 2026 with 282 models listed, Claude Opus 5.5 at Max effort leads at 1866 Elo. Eight of the top ten configurations are Anthropic's. GDPval-AA v2.1 carries a 10% weight in the AA Intelligence Index v4.3.2, down from 20% in v4.1.
Despite the name, nothing here is weighted by GDP or dollars. The score is a blind pairwise preference between deliverables, and every task counts the same.
How the tasks, harness and scoring work
The full OpenAI GDPval set has 1,320 tasks across 44 occupations and nine sectors, including finance and insurance, government, healthcare, manufacturing and information. Epoch AI describes the occupations as "the five highest-earning predominantly digital occupations" in each sector. In OpenAI's own grading, "a domain expert sees only the task and two unlabeled deliverables (the model's output and a human expert's), and ranks them without knowing which is which." Each task ends in a win, tie or loss against the human baseline, and OpenAI's leaderboard reports a win rate and a wins-plus-ties rate.
GDPval-AA keeps the 220 public tasks and changes the rest. A run has two stages.
- Task submission. The model works in an agentic loop through Stirrup, which AA calls "our open source reference agent harness", with a shell and a web browser. It produces files: documents, slides, diagrams and spreadsheets.
- Pairwise grading. An LLM judge sees two anonymized submissions for the same task and picks a winner. Since v2 the judges come from "a rotating panel of frontier-model judges."
AA fits Elo ratings from those pairwise results. Since v2.1 the fit uses a Crowd-BT model and pins DeepSeek V4.1 Flash at Max effort to 1600, so every other score is a distance from that one model. The board prints a 95% interval per row, shown as a half-width such as ±26.
AA states that "all evaluations are conducted independently by Artificial Analysis." For two Anthropic launches the numbers are AA-measured but vendor-published, because AA says it "supported Anthropic to evaluate Claude Opus 5 ahead of release" (AA Opus 5 article).
The index takes GDPval-AA scores "frozen at the time of a model's addition and normalized as clamp((Elo - 500) / 2000)". In Index v4.3.2 the Agents category is 30% of the total: AA-Briefcase v1.1 at 15%, GDPval-AA v2.1 at 10% and AutomationBench-AA at 5%. Coding and Scientific Reasoning are 20% each, and General is 30%.
AA has not published the wall-clock limit per task, a token cap, the sandbox network policy, the judge identities, the judge-panel size or any per-occupation scores.
Versions
| Version | Date | What changed | Scale |
|---|---|---|---|
| v1 | Not stated | 100-turn limit | Not stated |
| v2 | June 15, 2026, with Index v4.1 | "Re-baselines ELO to human performance at 1000, introduces a rotating panel of frontier-model judges, and raises the turn limit from 100 to 250"; 20% of Index v4.1, "the highest weighted evaluation" | Human expert = 1000 |
| v2.1 | September 19, 2026 | "GDPval-AA Elo is now anchored to DeepSeek V4.1 Flash (max) at 1600, with updates to fitting methodology for GDPval-AA and AA-Briefcase to improve Elo stability"; same 220 tasks; Crowd-BT fitting | DeepSeek V4.1 Flash (Max) = 1600 |
Sources: AA Index v4.1 article, AA changelog, AA methodology.
The v2 and v2.1 numbers sit on different scales. A v2 score of 1800 and a v2.1 score of 1800 mean different things, and v2.1 has no human-expert reference point at all.
How to read the run labels
Some rows on the board carry a label inside the model name. Opus 5.5, Sonnet 5.5 and Fable 5.1 rows read "(Max, Default Fallback)" or similar. The Fable 5 row reads "(Max, Opus 4.8 Fallback)". Every other row has no label. The page carries no legend or footnote that defines these strings.
AA's Opus 5 article says Opus 5 has "support for server-side fallback as with Fable 5. Intelligence Index evaluations were run with Opus 4.8 fallback enabled." Anthropic's Opus 5.5 launch page adds: "Claude Opus 5.5 was evaluated with its production safeguards enabled. When they intervened, cybersecurity tasks were completed by Claude Opus 4.8, and biology and frontier LLM development tasks were completed by Claude Opus 5. This likely reduces Claude Opus 5.5's performance on these benchmarks." On AutomationBench-AA, the same page data shows Opus 5.5 at Xhigh with 236 of 12,973 LLM turns and 21 of 657 tasks served by a fallback model.
Three things are missing:
- AA does not say which model "Default Fallback" names.
- AA publishes no count of GDPval-AA turns or tasks served by a fallback model.
- The Opus 5 rows carry no label, yet AA's Opus 5 article says its Index runs had Opus 4.8 fallback enabled. A missing label does not prove a no-fallback run.
The tables below split rows by label, because a fallback-labeled row includes turns completed by a different model. A gap inside one table compares runs with the same label. A gap across tables is a position difference on one board.
GDPval-AA leaderboard (October 2026)
Every row below comes from the AA live board, retrieved October 6, 2026. All are AA-run on GDPval-AA v2.1 with Stirrup, shell and web, the same 220 tasks and the same rating scale. Rank is position by Elo among all 282 models.
Default Fallback rows
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 5.5 (Max) | Anthropic | 1866 ±26 | Stirrup, Max, Default Fallback | Oct 6, 2026 | AA board, rank 1 |
| Claude Sonnet 5.5 (Max) | Anthropic | 1839 ±24 | Stirrup, Max, Default Fallback | Oct 6, 2026 | AA board, rank 2 |
| Claude Opus 5.5 (Xhigh) | Anthropic | 1837 ±25 | Stirrup, Xhigh, Default Fallback | Oct 6, 2026 | AA board, rank 3 |
| Claude Fable 5.1 (Max) | Anthropic | 1758 ±20 | Stirrup, Max, Default Fallback | Oct 6, 2026 | AA board, rank 4 |
| Claude Fable 5.1 (Xhigh) | Anthropic | 1735 ±21 | Stirrup, Xhigh, Default Fallback | Oct 6, 2026 | AA board, rank 5 |
| Claude Sonnet 5.5 (Xhigh) | Anthropic | 1731 ±26 | Stirrup, Xhigh, Default Fallback | Oct 6, 2026 | AA board, rank 6 |
| Claude Opus 5.5 (High) | Anthropic | 1707 ±23 | Stirrup, High, Default Fallback | Oct 6, 2026 | AA board, rank 10 |
| Claude Fable 5.1 (High) | Anthropic | 1635 ±20 | Stirrup, High, Default Fallback | Oct 6, 2026 | AA board, rank 18 |
| Claude Opus 5.5 (Medium) | Anthropic | 1586 ±22 | Stirrup, Medium, Default Fallback | Oct 6, 2026 | AA board, rank 34 |
Comparable to each other on one v2.1 scale; these rows include turns served by an unnamed fallback model, so gaps to unlabeled rows are board positions, not same-condition gaps.
Opus 5.5 Max leads Sonnet 5.5 Max by 27 and Opus 5.5 Xhigh by 29. Both gaps fall inside the ±24 to ±26 intervals, so the top three are not separable on this benchmark. Opus 5.5 Max leads Fable 5.1 Max by 108, which is well clear of the noise.
Unlabeled rows, ranks 8 to 30
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Grok 4.7 (Xhigh) | SpaceXAI | 1715 ±23 | Stirrup, Xhigh | Oct 6, 2026 | AA board, rank 8 |
| Grok 4.7 (High) | SpaceXAI | 1710 ±20 | Stirrup, High | Oct 6, 2026 | AA board, rank 9 |
| MiMo-V2.6-Pro | Xiaomi | 1686 ±23 | Stirrup, no effort listed | Oct 6, 2026 | AA board, rank 12 |
| Muse Spark 1.3 (Max) | Meta | 1684 ±20 | Stirrup, Max | Oct 6, 2026 | AA board, rank 13 |
| Qwen3.8 Max (0902) | Alibaba | 1671 ±21 | Stirrup, no effort listed | Oct 6, 2026 | AA board, rank 14 |
| GLM-5.3 (Max) | Z.ai | 1653 ±20 | Stirrup, Max | Oct 6, 2026 | AA board, rank 15 |
| GLM 5.3 Flash | Z.ai | 1647 ±18 | Stirrup, no effort listed | Oct 6, 2026 | AA board, rank 16 |
| Grok 4.6 (Xhigh) | SpaceXAI | 1647 ±20 | Stirrup, Xhigh | Oct 6, 2026 | AA board, rank 17 |
| Qwen3.8-Flash-Next | Alibaba | 1633 ±20 | Stirrup, no effort listed | Oct 6, 2026 | AA board, rank 19 |
| Muse Spark 1.3 (Xhigh) | Meta | 1630 ±20 | Stirrup, Xhigh | Oct 6, 2026 | AA board, rank 20 |
| Gemini 4 Argon (High) | 1627 ±22 | Stirrup, High | Oct 6, 2026 | AA board, rank 21 | |
| Ling 3.1 Flash | InclusionAI | 1622 ±18 | Stirrup, no effort listed | Oct 6, 2026 | AA board, rank 22 |
| Grok 4.6 (High) | SpaceXAI | 1622 ±18 | Stirrup, High | Oct 6, 2026 | AA board, rank 23 |
| Grok 4.6 (Medium) | SpaceXAI | 1617 ±20 | Stirrup, Medium | Oct 6, 2026 | AA board, rank 24 |
| Qwen3.8 2.4T A95B | Alibaba | 1613 ±20 | Stirrup, no effort listed | Oct 6, 2026 | AA board, rank 26 |
| GPT-5.6 Sol (Max) | OpenAI | 1611 ±17 | Stirrup, Max | Oct 6, 2026 | AA board, rank 27; best OpenAI row |
| MiMo-V2.6-Flash | Xiaomi | 1611 ±23 | Stirrup, no effort listed | Oct 6, 2026 | AA board, rank 28 |
| Qwen3.8 Max | Alibaba | 1608 ±19 | Stirrup, no effort listed | Oct 6, 2026 | AA board, rank 29 |
| DeepSeek V4.1 Flash (Max) | DeepSeek | 1600 (anchor) | Stirrup, Max | Oct 6, 2026 | AA board, rank 30; scale anchor, fixed at 1600 |
Comparable to each other on one v2.1 scale with no fallback label; not comparable to v2 scores, where human experts sat at 1000.
Grok 4.7 Xhigh at 1715 is the best non-Anthropic configuration. It leads Gemini 4 Argon High by 88 and GPT-5.6 Sol Max by 104. MiMo-V2.6-Pro is the surprise of this group, 29 points behind Grok 4.7 Xhigh at a small fraction of the cost.
Claude Opus 5
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 5 (Max) | Anthropic | 1724 ±20 | Stirrup, Max, no label | Oct 6, 2026 | AA board, rank 7 |
| Claude Opus 5 (Xhigh) | Anthropic | 1692 ±20 | Stirrup, Xhigh, no label | Oct 6, 2026 | AA board, rank 11 |
| Claude Opus 5 (High) | Anthropic | 1596 ±20 | Stirrup, High, no label | Oct 6, 2026 | AA board, rank 31 |
Same v2.1 scale as the tables above; AA's Opus 5 article says its Index runs had Opus 4.8 fallback enabled, so these are not no-fallback runs despite the missing label.
On the board, Opus 5.5 Max sits 142 points above Opus 5 Max. The two carry different labels, so read that as a board position, not a controlled same-condition gap.
Opus 4.8 Fallback row
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Fable 5 (Max) | Anthropic | 1613 ±17 | Stirrup, Max, Opus 4.8 Fallback | Oct 6, 2026 | AA board, rank 25; June 2026 release |
Same v2.1 scale; the only row that names its fallback model, Opus 4.8.
OpenAI flagships and other rows below rank 30
These are the OpenAI flagship rows below rank 30 plus Step 5 Preview and Kimi K3 Max, the two other rows in this range with turn data in the page payload.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Step 5 Preview | StepFun | 1587 ±26 | Stirrup, no effort listed | Oct 6, 2026 | AA board, rank 33; 63.7 turns, derived $0.60 per task |
| GPT-6.1 Sol (Max) | OpenAI | 1575 ±20 | Stirrup, Max | Oct 6, 2026 | AA board, rank 35; 24.9 turns, derived $0.82 per task |
| GPT-5.6 Sol (Xhigh) | OpenAI | 1572 ±18 | Stirrup, Xhigh | Oct 6, 2026 | AA board, rank 36; no turn or token data in payload |
| GPT-6 Astra (Max) | OpenAI | 1542 ±25 | Stirrup, Max | Oct 6, 2026 | AA board, rank 40; 24.2 turns, derived $4.21 per task |
| Kimi K3 (Max) | Kimi | 1537 ±20 | Stirrup, Max | Oct 6, 2026 | AA board, rank 41; 37.5 turns, derived $1.97 per task |
| GPT-6 Sol (Max) | OpenAI | 1510 ±24 | Stirrup, Max | Oct 6, 2026 | AA board, rank 44; 39.9 turns, derived $1.19 per task |
Same v2.1 scale and no fallback label, so directly comparable to the unlabeled ranks 8 to 30 table.
OpenAI does poorly here. Its best row, GPT-5.6 Sol Max at 1611, ranks 27th. GPT-6.1 Sol Max leads GPT-6 Astra Max by 33, inside the intervals, and leads GPT-6 Sol Max by 65. Astra Max, OpenAI's newest flagship, ranks 40th of 282.
Effort ladders
Values are Elo with the 95% half-width, from the same board. Anthropic 5.5 and Fable 5.1 rows carry "Default Fallback"; Opus 5, OpenAI and Grok rows carry no label.
| Model | Low | Medium | High | Xhigh | Max | Max minus Low |
|---|---|---|---|---|---|---|
| Claude Opus 5.5 | 1235 ±22 | 1586 ±22 | 1707 ±23 | 1837 ±25 | 1866 ±26 | 631 |
| Claude Sonnet 5.5 | 1179 ±23 | 1324 ±23 | 1551 ±24 | 1731 ±26 | 1839 ±24 | 660 |
| Claude Fable 5.1 | 1469 ±20 | 1549 ±20 | 1635 ±20 | 1735 ±21 | 1758 ±20 | 289 |
| Claude Opus 5 | 1304 ±22 | 1492 ±20 | 1596 ±20 | 1692 ±20 | 1724 ±20 | 420 |
| GPT-6 Astra | 1366 ±26 | 1468 ±24 | 1485 ±26 | 1517 ±26 | 1542 ±25 | 176 |
| GPT-6.1 Sol | 1297 ±19 | 1433 ±19 | 1486 ±19 | 1510 ±18 | 1575 ±20 | 278 |
| GPT-6 Sol | 1204 ±24 | 1350 ±25 | 1396 ±23 | 1457 ±24 | 1510 ±24 | 306 |
| GPT-5.6 Sol | 1305 ±18 | 1422 ±18 | 1505 ±18 | 1572 ±18 | 1611 ±17 | 306 |
Grok 4.7 has Low 1589 ±24, High 1710 ±20 and Xhigh 1715 ±23, with no Medium or Max row. Grok 4.6 has Low 1409, Medium 1617, High 1622 and Xhigh 1647.
Within each row the label is constant, so the steps are same-condition comparisons; across rows the fallback split above still applies.
The Claude ladders are steep. Sonnet 5.5 at Low scores 1179, below every OpenAI Max row, and at Max it is second on the board. Astra's ladder is the flattest at 176 points, which says extra test-time compute buys Astra little on this task set. On the old v2 scale AA reported that Opus 5's "effort levels span 407 Elo points, with output token usage ranging around 8x from low to max effort"; the same ladder spans 420 on v2.1.
Anthropic's launch page claims: "At default effort (medium), Opus 5.5 beats GPT-6 Astra at max effort for about a fifth of the cost per task." The board agrees on the order, with Opus 5.5 Medium at 1586 against Astra Max at 1542. That 44-point lead is inside the ±22 and ±25 intervals. The page data has no token counts for Opus 5.5 Medium, so the "fifth of the cost" figure is Anthropic's alone.
Cost per task and turns
AA's page has a cost chart but does not print cost per task as a number. Every cost on this page is derived from the board's embedded page data. Turn counts sit at gdpvalBreakdown.avgTurns, and token totals over the 220 tasks sit at canonicalEvalTokenCounts.gdpval, which holds input, answer, reasoning and cacheableInput, the cache-hit part of input. Prices are USD per million tokens. Only 30 rows carry this breakdown.
The formula, per task in USD, is [(input - cacheableInput) x input price + cacheableInput x cache-hit price + (answer + reasoning) x output price] / 1,000,000 / 220. For Opus 5.5 Max that is (3,793,360,879 - 3,726,668,094) x $4 + 3,726,668,094 x $0.20 + (14,463,522 + 26,714,602) x $20, divided by 1,000,000 and by 220, which gives $8.34.
Anthropic costs are lower bounds. The payload has no GDPval cache-write token count, so cache-write premiums are missing, and fallback-model pricing is not included. AA also footnotes that refused tasks are estimated from other models' token use.
Default Fallback rows (all Anthropic)
| Label | Elo | Derived cost / task | Avg turns | Output tokens / task | Frontier |
|---|---|---|---|---|---|
| Opus 5.5 (Max) | 1866 | $8.34 | 85.7 | 187k | Frontier |
| Sonnet 5.5 (Max) | 1839 | $8.48 | 92.3 | 285k | |
| Opus 5.5 (Xhigh) | 1837 | $3.93 | 58.9 | 92k | Frontier |
| Fable 5.1 (Max) | 1758 | $8.85 | 60.4 | 106k | |
| Fable 5.1 (Xhigh) | 1735 | $6.54 | 50.7 | 79k | |
| Sonnet 5.5 (Xhigh) | 1731 | $2.27 | 43.6 | 93k | Frontier |
The cost frontier here runs Sonnet 5.5 Xhigh ($2.27, 1731), Opus 5.5 Xhigh ($3.93, 1837), Opus 5.5 Max ($8.34, 1866). Sonnet 5.5 Max and Fable 5.1 Max both cost more than Opus 5.5 Xhigh and score no higher. Opus 5.5 Xhigh costs 47% of Opus 5.5 Max.
Unlabeled rows
| Label | Elo | Derived cost / task | Avg turns | Output tokens / task | Frontier |
|---|---|---|---|---|---|
| Grok 4.7 (Xhigh) | 1715 | $4.50 | 54.4 | 136k | Frontier |
| Grok 4.7 (High) | 1710 | $3.68 | 48.8 | 118k | Frontier |
| MiMo-V2.6-Pro | 1686 | $0.15 | 46.0 | 88k | Frontier |
| Muse Spark 1.3 (Max) | 1684 | $1.64 | 74.5 | 85k | |
| Qwen3.8 Max (0902) | 1671 | $3.54 | 79.4 | 122k | |
| GLM-5.3 (Max) | 1653 | $2.22 | 70.0 | 94k | |
| GLM 5.3 Flash | 1647 | $0.30 | 80.9 | 126k | |
| Gemini 4 Argon (High) | 1627 | $1.28 | 35.4 | 75k | |
| DeepSeek V4.1 Flash (Max) | 1600 | $0.25 | 72.1 | 116k | |
| GPT-6.1 Sol (Max) | 1575 | $0.82 | 24.9 | 46k |
Only three of these ten sit on the cost frontier: MiMo-V2.6-Pro ($0.15, 1686), Grok 4.7 High ($3.68, 1710) and Grok 4.7 Xhigh ($4.50, 1715). Paying 30 times more for Grok 4.7 Xhigh than MiMo-V2.6-Pro buys 29 Elo.
Two reference points sit outside both groups. Claude Opus 5 Max scores 1724 at $6.44 with 60.5 turns and 97k output tokens per task. GPT-6 Astra Max scores 1542 at $4.21 with 24.2 turns and 40k output tokens.
A few readings follow from these numbers. Opus 5.5 Max costs more than Opus 5 Max, $8.34 against $6.44, because it takes 85.7 turns against 60.5 and emits 187k output tokens against 97k. GPT-6.1 Sol Max costs one fifth of Astra Max and scores 33 higher, inside the intervals.
Low turn counts go with low scores. Among the 30 rows with turn data, five average 40 turns or fewer per task. They are Astra at 24.2, GPT-6.1 Sol at 24.9, Gemini 4 Argon at 35.4, Kimi K3 at 37.5 and GPT-6 Sol at 39.9, and all five score 1627 or lower. The top Claude configurations take 60 to 86. AA reports the turn gap and gives no cause for it.
On the full Intelligence Index, not GDPval alone, AA's GPT-6 Astra article puts Astra Max at 53, level with Fable 5.1 at Max with fallback, for $3.26 against $7.63 per Index task.
Historical results
Anthropic launch snapshot (September 22, 2026)
Anthropic's Opus 5.5 launch page published GDPval-AA v2.1 scores that AA measured. Anthropic states that all its results use "adaptive thinking at max effort" unless noted. It does not state the effort for the GPT rows.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 5.5 | Anthropic | 1846 | Stirrup, adaptive thinking at max effort | Sep 22, 2026 | Anthropic launch page; live board 1866 (Max), +20 |
| Claude Fable 5.1 | Anthropic | 1735 | Stirrup, adaptive thinking at max effort | Sep 22, 2026 | Anthropic launch page; live board 1758 (Max); 1735 is the board's Xhigh value |
| Claude Opus 5 | Anthropic | 1708 | Stirrup, adaptive thinking at max effort | Sep 22, 2026 | Anthropic launch page; live board 1724 (Max), +16 |
| GPT-5.6 Sol | OpenAI | 1588 | Stirrup, effort not stated | Sep 22, 2026 | Anthropic launch page; live board 1611 (Max), +23 |
| GPT-6 Astra | OpenAI | 1542 | Stirrup, effort not stated | Sep 22, 2026 | Anthropic launch page; live board 1542 (Max), unchanged |
Same v2.1 scale as the live board, but an earlier fit; AA refit the ratings after launch, so rows moved by 0 to 23 points.
The launch page says: "On GDPval-AA v2.1, a test of real-world work across 44 occupations, Opus 5.5 scores 1846 Elo, ahead of Fable 5.1 and Opus 5." The page also has an Elo-versus-cost chart, with Elo from 1200 to 1800 against "Estimated cost per task (USD, log scale)" from $0.20 to $10, but the point values are not in the page text.
GDPval-AA v2 results (June to July 2026)
| Date | Model | Organization | Elo (v2) | Source and notes |
|---|---|---|---|---|
| Jun 15, 2026 | Claude Fable 5 (with fallback) | Anthropic | 1818 | AA Index v4.1 article |
| Jun 15, 2026 | Claude Opus 4.8 | Anthropic | 1638 | Same |
| Jun 15, 2026 | GPT-5.5 (xhigh) | OpenAI | 1531 | Same |
| Jul 24, 2026 | Claude Opus 5 (max) | Anthropic | 1861 | AA Opus 5 article; +114 over Fable 5, ">100 points ahead of Claude Fable 5 and GPT-5.6 Sol (max)" |
v2 scale anchored at human experts = 1000; not comparable to any v2.1 table on this page.
AA's September 9, 2026 Astra article gives no absolute v2 score for GPT-6 Astra. It says: "Compared to GPT-5.6 Sol, we observe a drop of ~45 Elo points in GDPval-AA v2 - a benchmark we adapted from OpenAI's dataset measuring economically valuable tasks across 44 occupations. We observed GPT-6 Astra using significantly fewer turns than other models in GDPval tasks - 24 per task at max effort, compared to 45 for GPT-5.6 Sol and 60 turns per task for Claude Fable 5.1 and Claude Opus 5." On v2.1 the board points the same way, with Astra Max 69 below GPT-5.6 Sol Max, outside the intervals.
Timeline
- October 5, 2025. OpenAI's GDPval paper reports that frontier performance "is improving roughly linearly over time", that the best models are "approaching industry experts in deliverable quality", and that "increased reasoning effort, increased task context, and increased scaffolding improves model performance."
- June 15, 2026. GDPval-AA v2 ships with Index v4.1. Fable 5 with fallback leads at 1818.
- July 24, 2026. Claude Opus 5 at max reaches 1861 on v2. The same article puts Opus 5 at 1720 on AA-Briefcase, 146 over Fable 5. Across the whole Intelligence Index, Opus 5 at max costs $2.03 per task against $2.75 for Fable 5 with fallback, $1.80 for Opus 4.8 at max and $1.53 for Sonnet 5 at max.
- September 7, 2026. Index v4.3 moves Terminal-Bench to v4.0 and adds AutomationBench-AA.
- September 14, 2026. AA-Briefcase is announced in Capability Indices v1.1.
- September 19, 2026. GDPval-AA v2.1 changes the anchor and fitting method.
- September 22, 2026. Anthropic's launch page shows Opus 5.5 at 1846 against Opus 5 at 1708 on the v2.1 snapshot, a 138-point gap.
- October 6, 2026. The live board in the tables above, now including GPT-6.1 Sol, released September 29, at 1575.
Failure modes and limitations
A relative scale. Elo ranks models against each other around one anchor. On v2.1 that anchor is DeepSeek V4.1 Flash at 1600; on v2 it was human experts at 1000. A GDPval-AA score means nothing without a version and a date.
Version changes move the whole scale. Opus 5 Max is 1861 on v2 and 1724 on v2.1. The model did not get worse. The scale moved.
Judge dependence. Scores come from LLM judges picking winners. v2 switched to a rotating panel of frontier-model judges, and AA has not published who the judges are or why it made the change. See LLM-as-a-judge for the general problem.
Precision. The 95% intervals run ±17 to ±26 on the top 30 rows. The Xhigh-to-Max step falls inside the intervals for Opus 5.5 (+29) and Fable 5.1 (+23), and outside them for Sonnet 5.5 (+108). For Opus 5.5 and Fable 5.1, the board cannot tell Max from Xhigh. For Sonnet 5.5, it can.
Fallback runs. Rows labeled "Default Fallback" or "Opus 4.8 Fallback" include turns completed by another model when Anthropic's safeguards intervene. Anthropic says this "likely reduces" its own scores on the benchmarks where it triggers.
Refusals. AA footnotes that for some rows "the model or provider declined some tasks on safety grounds" and that "refused tasks are estimated from other models' token use." Derived costs for those rows are estimates on top of estimates.
Scope. The occupations are the five highest-earning predominantly digital ones per sector. Each task is a one-shot deliverable with a 250-turn cap. Nothing here tests multi-session work. AA-Briefcase covers "multi-week knowledge work projects, each with many linked tasks and thousands of input source files", and GDPval-AA does not.
One harness. OpenAI's paper says scaffolding changes results. GDPval-AA fixes Stirrup for every model, so scores describe Stirrup runs only. A vendor's own agent harness can score differently, and those results are not on this board.
Contamination. OpenAI open-sourced the 220 gold tasks, so they are public. AA has published no private GDPval-AA split and no contamination audit. Compare AutomationBench-AA, which uses a "657-task held-out split", and Index v4.2, which added "more private test sets to prevent gaming."
The vendor's own caution. Anthropic, whose models top the board, writes that "at these levels of capability we've found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest." The company that benefits most from the board is telling you not to read its margins literally.
How GDPval-AA compares to related benchmarks
OpenAI GDPval. The original has 1,320 tasks. Human experts grade blind against human-made deliverables and report win and win-or-tie rates against human work. GDPval-AA runs only the 220 public tasks, swaps expert graders for LLM judges, ranks models against each other instead of against humans, and runs every model in one open harness. GDPval asks whether a model matches a professional. GDPval-AA asks which model beats which.
AA-Briefcase v1.1. Briefcase uses the same Stirrup harness on multi-week projects with many linked tasks and thousands of input files. It carries 15% of Index v4.3.2 against GDPval-AA's 10%. If your work spans long projects with large document sets, Briefcase is the closer match.
AutomationBench-AA. This one tests SaaS workflows across simulated business apps through REST APIs, with a 50-turn cap, a 657-task held-out split and 5% of the index. Anthropic notes that Zapier ran its own AutomationBench figures without fallback models, so safeguard interventions counted as failures there.
Vals Index. Vals AI's index measures accuracy on verifiable finance, coding, legal and tax tasks, weighted by GDP as (8.0 Finance + 5.6 Coding + 1.2 Legal + 0.5 Tax) / 15.3. Its components include Finance Agent v2, Excel Modeling Benchmark, Terminal-Bench 4.0, Vibe Code Bench, Code Migration, Legal Research Bench, HLAB and Tax Agent Bench. The Vals Index does not include GDPval.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Gemini 4 Argon | 68.90% | Vals Index, accuracy | Oct 2, 2026 | Vals board; $15.68 per test; 21st on GDPval-AA (1627 at High) | |
| Claude Sonnet 5.5 | Anthropic | 67.04% | Vals Index, accuracy | Oct 2, 2026 | Vals board; $21.34 per test |
| Claude Opus 5.5 | Anthropic | 66.97% | Vals Index, accuracy | Oct 2, 2026 | Vals board; $32.14 per test |
| Claude Fable 5.1 | Anthropic | 65.83% | Vals Index, accuracy | Oct 2, 2026 | Vals board; $28.71 per test |
| Claude Opus 5 | Anthropic | 63.67% | Vals Index, accuracy | Oct 2, 2026 | Vals board; $19.31 per test |
Accuracy on verifiable answers, not pairwise Elo; compare ranks across the two boards, never the numbers.
Gemini 4 Argon is first on the Vals Index and 21st on GDPval-AA. Vals checks whether an answer is correct. GDPval-AA asks a judge which of two deliverables is better.
Terminal-Bench 4.0 and SWE-bench. Both score with code-verified test suites, not judged preference. GDPval-AA has no ground truth to check against, which is why it needs judges and Elo at all. Related agentic evaluations on this site include GAIA, DRACO for deep research and the Harvey legal agent benchmark.
What it means for teams choosing a model
Opus 5.5 leads for report, deck and spreadsheet work. It scores 1866 at Max on the independent board. The next non-Anthropic row is Grok 4.7 Xhigh at 1715, and the best OpenAI row is GPT-5.6 Sol Max at 1611.
Run Opus 5.5 at Xhigh, not Max. Xhigh scores 1837 for a derived $3.93 per task against $8.34 at Max, and the gap sits inside the intervals. Sonnet 5.5 Xhigh, at 1731 for $2.27, is the cheapest Anthropic row above 1700. Sonnet 5.5 Max costs $8.48, about what Opus 5.5 Max costs, for 27 fewer Elo, so skip it.
Benchmark at more than one effort level. Effort is the biggest lever on Claude models. Opus 5.5 spans 631 Elo from Low to Max, Sonnet 5.5 spans 660, Fable 5.1 spans 289 and Astra 176. A team that tests a model at one effort has tested one point on a ladder.
Among OpenAI models, pick GPT-6.1 Sol Max. It scores 1575 for a derived $0.82 per task, one fifth of Astra Max's cost and 65 above GPT-6 Sol Max. No OpenAI row exceeds 1611, which is 255 below Opus 5.5 Max. If you are committed to OpenAI for deliverable-style work, test GPT-6.1 Sol before Astra.
For high-volume, cost-bound work, test MiMo-V2.6-Pro. It scores 1686 for a derived $0.15 per task, the cheapest point on the unlabeled cost chart, and beats Gemini 4 Argon High at 1627 and $1.28.
Check your own judge. GDPval-AA ranks models by what frontier-model judges prefer. If your reviewers weigh accuracy over polish, the way the Vals Index does, your ranking will differ. Build a golden dataset of your own deliverables and keep a human in the loop for grading. The frontier models and LLM evaluation pages cover the wider setup.
The Klu LLM leaderboard tracks these models across benchmarks. To run the same kind of pairwise deliverable eval on your own reports and spreadsheets, see how teams set it up for operations workflows in Klu.