ARC-AGI-2 and ARC-AGI-3

The ARC Prize Foundation's skill-acquisition benchmarks, which measure how efficiently a model learns a skill it has never seen, through static grid puzzles and interactive games

ReasoningBinary pass rate

Stephen M. Walker II · Co-Founder / CEO · October 6, 2026 · 14 min read

What are ARC-AGI-3 and ARC-AGI-2?

ARC-AGI is the ARC Prize Foundation's benchmark series for measuring how efficiently an AI system picks up a skill it has never seen. François Chollet created ARC-AGI and co-founded the Foundation with Mike Knoop (about ARC Prize). Every task draws only on Core Knowledge priors, meaning objects, geometry, basic physics and agents. There is no language to parse and no domain knowledge to look up, and untrained humans solved every task before it went in.

The design is the whole argument for the benchmark. A knowledge test like GPQA cannot tell a model that memorized a skill from one that worked it out on the spot. ARC tasks are new by construction, so retrieval does not help.

Two versions are live, and they test different things.

ARC-AGI-2, launched March 24, 2025, is a static set of grid puzzles. The model sees a few input and output grids, infers the transformation rule, and applies it to a new input. It is the abstraction test most people picture when they hear "ARC."

ARC-AGI-3, launched March 25, 2026, is interactive. The agent plays turn-based games on a 64x64 grid with no instructions. Nobody tells it the goal. It has to explore, work out what winning means, and then win in as few actions as it can, because the score compares its action count with the number of actions humans needed.

As of October 2026, GPT-6 Astra leads ARC-AGI-2 at 95.0% for $1.12 per task, and GPT-6.1 Sol sits 0.83 points behind at $0.254, so Astra costs 4.4 times as much for that gap. On ARC-AGI-3, Astra leads the Standard harness at 62.71% and the Provider Adapter harness at 99.95%. At launch in March 2026, under an earlier no-harness protocol, no frontier model reached 1% on ARC-AGI-3. The best was Claude Opus 4.6 at 0.50%.

Every current score on this page is ARC Prize Verified. The Foundation ran each model itself on its Semi-Private set, so nothing here is vendor-reported. The sources are the ARC-AGI-3 Technical Report, the ARC-AGI-2 paper, the Verified Testing Policy, ARC Prize's GPT-6 Astra post and the per-model results pages at arcprize.org/results, all accessed October 6, 2026.

Versions

VersionFormatVerified setMetricTop Verified score, Oct 2026
ARC-AGI-1Static grid puzzlesSemi-Private, 100 tasksBinary pass rate98.5%, saturated
ARC-AGI-2Static grid puzzles that need composed rulesSemi-Private, 120 tasksBinary pass rate95.0%
ARC-AGI-3Interactive turn-based games, goal withheldSemi-Private, 55 environmentsRelative Human Action Efficiency (RHAE)62.71% Standard, 99.95% Adapter

ARC-AGI-1 scores still appear on the results pages, but the top scores now span 97.5% to 98.5% across five organizations. It no longer ranks anything. ARC-AGI-2 is the static test to read, and ARC-AGI-3 is the interactive one.

How ARC-AGI-2 works

Each task is a few demonstration pairs plus a test input, and the model outputs a grid. Scoring is binary per task, with no partial credit. ARC Prize runs models without tools by default. It does not enable code execution or web search behind the model, and it rules out web search because search could leak Semi-Private tasks. Any tool use has to be declared in the model configuration.

The dataset is three separate pools of 120 tasks each, called Public Eval, Semi-Private Eval and Private Eval. Verified leaderboard scores come from Semi-Private. Each model gets one run, and ARC Prize does not average across runs. A score is accepted only if the Public and Semi-Private results agree within 3 points.

Humans set the ceiling. The ARC-AGI-2 paper reports a mean human solve time of 2.7 minutes per task, and a task stayed in only if at least two independent people solved it. A perfect score is within human reach.

ARC-AGI-2 replaced ARC-AGI-1 as the headline test because the old one stopped separating models. The new tasks require composing several rules, chaining steps, applying a rule only in the right context, and assigning symbolic meaning to shapes on the fly. A model that pattern-matches one transformation gets nowhere.

How ARC-AGI-3 works

An ARC-AGI-3 environment is a game. The agent sees a 64x64 grid, and each cell takes one of 16 colors. The full action set is five keys, Undo and a coordinate click, and each environment exposes a subset. Play is turn-based, so response latency does not count against the agent. Tool calls, reasoning steps and retries are not actions.

The agent is never told the objective. Each environment has at least six levels, the first being a tutorial, and humans finish an environment in about 20 minutes or less. The Technical Report defines three sets.

SetEnvironmentsUsed for
Public Demo25Public play, and the per-game pass/fail grids on each results page
Semi-Private55Verified runs of frontier models through external APIs
Fully Private55The Kaggle competition, deliberately out of distribution from the public set

The public set is easier. GPT-5.6 Sol at max effort averages 13.33% on Public and 7.78% on Semi-Private, and ARC Prize accepts a gap of up to 15 points between the two on ARC-AGI-3. The leaderboard numbers below are all Semi-Private.

Scoring with RHAE

ARC-AGI-3 scores each level by comparing the agent's action count with the human baseline:

level score = min(1.15, (human actions / AI actions)^2)

The environment score is a weighted average of its levels, where level l has weight l, capped at the weighted share of levels the agent completed. The benchmark score is the mean across environments.

The square makes the metric harsh. An agent that needs 10 times the human action count earns 1% for that level. At 5 times it earns 4%, and that is also where the leaderboard stops it: an agent is terminated at 5x the human median actions per level. The Technical Report says this cap has a negligible effect on scores because the square has already pushed a level that slow down to 4%. The 1.15 ceiling caps the reward for beating humans on efficiency.

The human baseline comes from members of the public. The Technical Report says ARC Prize tested 10 people per environment, set each level's baseline at the upper-median best human action count, and admitted an environment only if at least two people fully solved it. The Astra post describes the baseline as the median action count among players who completed each level, drawn from about 500 members of the general public.

Two harnesses, two questions

The Verified board reports ARC-AGI-3 under two harnesses and never mixes them.

The Standard harness is, in ARC Prize's words, "a minimal, provider-neutral interface" where the model "carries forward notes it chooses to keep." Every model gets the same system prompt, no tools and no ARC-AGI-3-specific scaffold.

The Provider Adapter harness uses "provider-designed context-management features, such as preserving opaque reasoning state between requests and using compaction for longer conversations." ARC Prize's own summary is that "these harnesses answer different evaluation questions."

Standard asks what a model does with a plain interface and its own notes. Adapter asks what it does when the vendor's context management keeps its reasoning state alive across a long game. On this benchmark, the gap between the two is larger than most gaps between models.

A community leaderboard accepts handcrafted harnesses, but those entries are self-reported and unverified. ARC Prize has a concrete reason for keeping them off the Verified board. On a variant of the TR87 environment, Opus 4.6 scored 0.0% with no harness and 97.1% with a harness built by Duke researchers. On BP35 it scored 0.0% both ways. ARC Prize concluded that harness gains on environments the builder has already seen do not transfer. A harness that replays human moves scores 100% on the public set.

What the cost figures cover

The two benchmarks report cost on different bases. ARC-AGI-1 and ARC-AGI-2 costs are USD per task. ARC-AGI-3 costs are a run total, as on the Grok 4.6 page: "2.11% on ARC-AGI-3 Semi-Private at a total cost of $5,612."

ARC Prize does not say whether an ARC-AGI-3 total covers Semi-Private alone or Public plus Semi-Private, or what one run includes. Its policy also states "We cap our evaluations at $10,000 USD per run," yet Astra's Standard runs are published at $26,098 to $49,791 and Opus 5's at $20,657. ARC Prize has not reconciled the two. Every ARC-AGI-3 cost below is the published figure, basis not stated.

Time is missing too. ARC Prize publishes no per-task time limit and no model wall-clock per task, only an aggregate speed ratio between Astra's two harnesses.

Current leaderboard

ARC-AGI-2, Semi-Private

Table note: one dataset, the 120-task Semi-Private set, one metric, binary pass rate, and no tools, with every row run once by ARC Prize. Rows compare with each other and not with ARC-AGI-1 or ARC-AGI-3. Each row is the model's best-scoring effort.

ModelOrganizationScoreHarness and setupDateSource and notes
GPT-6 AstraOpenAI95.0%No tools, max effort, single run2026-09-02ARC Prize results. $1.119 per task
GPT-6.1 SolOpenAI94.17%No tools, max effort, single run2026-09-29ARC Prize results. $0.254 per task. Page headline rounds to 94.2%
Claude Opus 5.5Anthropic93.33%No tools, high effort, single run2026-09-22ARC Prize results. $0.408 per task. Scores 91.67% at max effort
GPT-5.6 SolOpenAI92.5%No tools, max effort, single run2026-07-09ARC Prize results. $1.44 per task
Claude Opus 5Anthropic90.42%No tools, max effort, single run2026-07-24ARC Prize results. $2.06 per task
Claude Fable 5.1Anthropic90.0%No tools, max and xhigh tie, single run2026-09-01ARC Prize results. $4.49 per task at max, $3.12 at xhigh
GPT-6 SolOpenAI89.58%No tools, max effort, single run2026-09-22ARC Prize results. $0.439 per task
Gemini 3.8 FlashGoogle89.17%No tools, high effort, single run2026-09-02ARC Prize results. $0.40 per task
Claude Fable 5Anthropic89.17%No tools, max effort, single run2026-08-03ARC Prize results. $5.45 per task. Scores as of Aug 3, 2026; model released Jun 9, 2026
GPT-5.6 TerraOpenAI83.9%No tools, max effort, single run2026-07-09ARC Prize results. $1.09 per task. ARC Prize notes "xhigh performed worse than max"
Grok 4.6xAI67.08%No tools, xhigh effort, single run2026-08-11ARC Prize results. $0.757 per task
ARC PrizeARC-AGI-2 Semi-Private, 120 tasksBinary pass rate
ARC-AGI-2: score vs cost per taskSemi-Private, 120 tasks, no tools, ARC Prize Verified
90%92%94%96%$0.1$0.2$0.5$1$2$5
FrontierOpenAIBehind the frontier
Cost per task, log scale, lower to the right
Source: ARC Prize results pages, accessed October 6, 2026. Cost is USD per task from the score data on each model's results page; no derivation was needed. ARC-AGI-3 is not charted because its costs are run totals with no stated basis.

The top of ARC-AGI-2 is compressed and the prices are not. The four best scores, Astra at 95.0%, GPT-6.1 Sol at 94.17%, Opus 5.5 at 93.33% and GPT-5.6 Sol at 92.5%, span 2.5 points. Their per-task costs run from $0.254 to $1.44, a 5.7x spread.

Only three points sit on the cost frontier. GPT-6.1 Sol at high effort scores 91.67% for $0.135, GPT-6.1 Sol at max scores 94.17% for $0.254, and Astra at max scores 95.0% for $1.119. Everything else is beaten on both axes by something on that line. Opus 5.5 at high is close, but GPT-6.1 Sol at max is cheaper and 0.84 points more accurate. Gemini 3.8 Flash at high effort, 89.17% for $0.40, loses to GPT-6.1 Sol at high on both price and score.

ARC-AGI-2 by reasoning effort

Table note: same dataset and metric as the table above, one row per model, each cell a score and its cost per task at one effort level.

Modelmaxxhighhighmediumlow
GPT-6 Astra95.0%, $1.11993.33%, $0.82992.08%, $0.66892.08%, $0.48085.42%, $0.416
GPT-6.1 Sol94.17%, $0.25491.67%, $0.17891.67%, $0.13586.67%, $0.10076.67%, $0.086
Claude Opus 5.591.67%, $1.85292.5%, $0.67193.33%, $0.40887.5%, $0.33770.14%, $0.241
Claude Fable 5.190.0%, $4.49290.0%, $3.1288.75%, $1.66786.25%, $1.22478.33%, $0.952
GPT-5.6 Sol92.5%, $1.4490.0%, $1.0485.42%, $0.7467.08%, $0.4742.5%, $0.32
ARC PrizeARC-AGI-2 Semi-Private, 120 tasksBinary pass rate, max effort

Astra also has a run with reasoning off, at 59.58% for $0.370 per task. The ladders show how unevenly test-time compute pays off. Opus 5.5 peaks at high effort, 93.33% for $0.408, and drops to 91.67% at max for $1.852. GPT-6.1 Sol gains nothing from high to xhigh. GPT-5.6 Sol, by contrast, falls from 92.5% at max to 67.08% at medium.

ARC-AGI-3, Standard harness

Table note: one set, Semi-Private, one metric, RHAE, and the Standard harness across all rows. Not comparable to the Provider Adapter table or to the March 2026 launch scores. Run cost is the published figure, basis not stated.

ModelOrganizationScoreHarness and setupDateSource and notes
GPT-6 AstraOpenAI62.71%Standard, max effort2026-09-02ARC Prize results, Astra post. $26,098 per run
GPT-6 AstraOpenAI59.34%Standard, xhigh effort2026-09-02Same pages. $37,317 per run
GPT-6 AstraOpenAI54.82%Standard, high effort2026-09-02Same pages. $40,705 per run
GPT-6 AstraOpenAI38.59%Standard, medium effort2026-09-02Same pages. $48,090 per run
GPT-6 AstraOpenAI35.18%Standard, reasoning none2026-09-02Same pages. $49,791 per run
GPT-6 AstraOpenAI17.45%Standard, low effort2026-09-02Same pages. $38,166 per run
GPT-6.1 SolOpenAI52.73%Standard, max effort2026-09-29ARC Prize results. $7,632 per run
GPT-6.1 SolOpenAI39.93%Standard, xhigh effort2026-09-29Same page. $9,104 per run
GPT-6.1 SolOpenAI26.72%Standard, high effort2026-09-29Same page. $9,758 per run
GPT-6.1 SolOpenAI10.57%Standard, medium effort2026-09-29Same page. $7,370 per run
GPT-6.1 SolOpenAI3.92%Standard, low effort2026-09-29Same page. $7,111 per run
Claude Opus 5Anthropic30.16%Standard, high effort, only effort run2026-07-24ARC Prize results. $20,657 per run
Gemini 3.8 FlashGoogle10.37%Standard, high effort2026-09-02ARC Prize results. $4,398 per run
GPT-5.6 SolOpenAI7.78%Standard, max effort2026-07-09ARC Prize results. $25,064 per run. First model to win a public game, ft09 at 87%
GPT-6 SolOpenAI4.62%Standard, max effort2026-09-22ARC Prize results. $5,554 per run
Grok 4.6xAI2.11%Standard, xhigh effort2026-08-11ARC Prize results. $5,612, described on the page as "total cost"
GPT-5.6 TerraOpenAI0.80%Standard, max effort2026-07-09ARC Prize results. $7,918 per run
ARC PrizeStandardARC-AGI-3 Semi-PrivateRHAE

Astra leads by 9.98 points over GPT-6.1 Sol, and its published run cost is 3.4 times higher. Opus 5 is the only Anthropic model on this table, at 30.16%. Claude Opus 5.5 has an ARC-AGI-2 score but an empty ARC-AGI-3 column, Claude Fable 5.1's page lists ARC-AGI-1 and ARC-AGI-2 only, and Claude Fable 5's page says "ARC-AGI-3 testing is ongoing."

ARC-AGI-3, Provider Adapter harness

Table note: same set and metric as the Standard table, different harness. Compare rows only within this table. Only OpenAI and Google have Adapter runs. Run cost is the published figure, basis not stated.

ModelOrganizationScoreHarness and setupDateSource and notes
GPT-6 AstraOpenAI99.95%Provider Adapter, high effort2026-09-02ARC Prize results, Astra post. $18,817 per run
GPT-6 AstraOpenAI98.55%Provider Adapter, max effort2026-09-02Same pages. $17,332 per run
GPT-6 AstraOpenAI98.44%Provider Adapter, xhigh effort2026-09-02Same pages. $18,147 per run
GPT-6 AstraOpenAI98.44%Provider Adapter, medium effort2026-09-02Same pages. $19,285 per run
GPT-6 AstraOpenAI98.03%Provider Adapter, low effort2026-09-02Same pages. $21,298 per run
GPT-6 AstraOpenAI96.72%Provider Adapter, reasoning none2026-09-02Same pages. $23,457 per run
GPT-6.1 SolOpenAI96.37%Provider Adapter, xhigh effort2026-09-29ARC Prize results. $4,360 per run
GPT-6.1 SolOpenAI96.18%Provider Adapter, max effort2026-09-29Same page. $3,817 per run
GPT-6.1 SolOpenAI94.99%Provider Adapter, high effort2026-09-29Same page. $5,404 per run
GPT-6.1 SolOpenAI91.02%Provider Adapter, medium effort2026-09-29Same page. $5,835 per run
GPT-6.1 SolOpenAI82.79%Provider Adapter, low effort2026-09-29Same page. $6,622 per run
Gemini 3.8 FlashGoogle35.00%Provider Adapter, high effort2026-09-02ARC Prize results. $4,522 per run
Gemini 3.8 FlashGoogle24.20%Provider Adapter, medium effort2026-09-02Same page. $3,826 per run
Gemini 3.8 FlashGoogle15.07%Provider Adapter, low effort2026-09-02Same page. $3,158 per run
GPT-6 SolOpenAI23.03%Provider Adapter, max effort2026-09-22ARC Prize results. $8,722 per run
ARC PrizeProvider AdapterARC-AGI-3 Semi-PrivateRHAE

Astra at high effort is effectively at the ceiling, 3.6 points above GPT-6.1 Sol's best Adapter run. In the Astra post, ARC Prize compared the two harnesses on the 167 game and reasoning-level pairs that Astra solved under both, across Public and Semi-Private. Adapter runs were about 3.66 times faster by aggregate recorded elapsed time and used 49% fewer total tokens. Astra at max effort on the Adapter used fewer actions than the human baseline on 96.0% of levels, and 51.7% fewer actions per level on average.

ARC-AGI-1, Semi-Private (saturated)

Table note: a different dataset from ARC-AGI-2, with 100 tasks, kept here to show saturation, not to rank models. Rows that tie a model's best score are listed separately.

ModelOrganizationScoreHarness and setupDateSource and notes
GPT-6 AstraOpenAI98.5%No tools, xhigh effort2026-09-02ARC Prize results. $0.347 per task
GPT-6 AstraOpenAI98.5%No tools, high effort2026-09-02Same page. $0.284 per task. Max effort scores 97.5% at $0.433
GPT-6.1 SolOpenAI98.5%No tools, xhigh effort2026-09-29ARC Prize results. $0.074 per task
GPT-6.1 SolOpenAI98.5%No tools, high effort2026-09-29Same page. $0.059 per task
Claude Opus 5.5Anthropic98.5%No tools, high effort2026-09-22ARC Prize results. $0.158 per task
Claude Fable 5Anthropic98.5%No tools, xhigh effort2026-06-09ARC Prize results. $1.022 per task. Date is the model release date on the page
Claude Fable 5Anthropic98.5%No tools, max effort2026-06-09Same page. $2.11 per task
Gemini 3.8 FlashGoogle98.5%No tools, high effort2026-09-02ARC Prize results. $0.206 per task
Claude Fable 5.1Anthropic97.5%No tools, max effort2026-09-01ARC Prize results. $1.40 per task
Claude Opus 5Anthropic97.5%No tools, max effort2026-07-24ARC Prize results. $0.70 per task
Claude Opus 5Anthropic97.5%No tools, high effort2026-07-24Same page. $0.47 per task
GPT-6 SolOpenAI95.5%No tools, max effort2026-09-22ARC Prize results. $0.141 per task
Grok 4.6xAI87.5%No tools, medium effort2026-08-11ARC Prize results. $0.301 per task
ARC PrizeARC-AGI-1 Semi-Private, 100 tasksBinary pass rate

Historical progression

ARC-AGI-1 and o3-preview, December 2024

Table note: o3-preview was a preview model that ARC Prize tested with OpenAI directing the compute levels, so these are not Verified rows of a released model. Public, with 400 tasks, and Semi-Private, with 100, are different sets, and each compute setting is a separate configuration.

ModelOrganizationScoreHarness and setupDateSource and notes
o3-previewOpenAI75.7%ARC-AGI-1 Semi-Private, high efficiency, 6 samples2024-12-20ARC Prize. $26 per task, $2,680 retail total. Within ARC-AGI-Pub budget rules
o3-previewOpenAI87.5%ARC-AGI-1 Semi-Private, low efficiency, 1,024 samples2024-12-20Same post. 172x the compute of the 6-sample run. $4,560 per task, $456,000 total
o3-previewOpenAI82.8%ARC-AGI-1 Public, high efficiency, 6 samples2024-12-20Same post. $167 per task
o3-previewOpenAI91.5%ARC-AGI-1 Public, low efficiency, 1,024 samples2024-12-20Same post. $1,900 per task
ARC Prizearcprize.org

In December 2024, a top ARC-AGI-1 result cost thousands of dollars per task. In September 2026, GPT-6.1 Sol scores 98.5% on the same Semi-Private set for $0.059 per task, under the Verified no-tools protocol rather than OpenAI-directed sampling.

ARC-AGI-3 at launch, March 2026

Table note: Semi-Private, no harness, from Table 2 of the Technical Report. This is a different protocol from the Standard harness, so these rows do not compare with the current Standard or Adapter tables.

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 4.6 (Max)Anthropic0.50%No harness, Semi-Private2026-03Technical Report, Table 2
Gemini 3.1 Pro PreviewGoogle0.40%No harness, Semi-Private2026-03Same report
GPT-5.4 (High)OpenAI0.20%No harness, Semi-Private2026-03Same report
Grok-4.20xAI0.10%No harness, Semi-Private2026-03Same report
ARC PrizeNo harnessARC-AGI-3 Semi-Privatearcprize.org

Within a single protocol, the 2026 Verified results do show a trend. On the ARC-AGI-3 Standard harness, the best score went from 7.78% for GPT-5.6 Sol on July 9, to 30.16% for Opus 5 on July 24, to 62.71% for Astra on September 2. On ARC-AGI-2 the best score moved from 92.5% in July to 95.0% in September, a much smaller step because there is little room left.

Failure modes and limitations

Training exposure. ARC Prize states in the Technical Report that ARC-AGI-1 and ARC-AGI-2 are likely over-represented in training data. Its evidence is a Gemini 3 Deep Think reasoning chain that used the correct ARC integer-to-color mapping, 3 for green and 6 for magenta, though the verification prompt never mentioned ARC-AGI or the mapping. The cause it gives is that ARC tasks are verifiable, so a lab can generate millions of synthetic tasks, solve them and train on the results. That is a direct reason to read ARC-AGI-2 scores in the mid-90s with less confidence than ARC-AGI-3 scores.

Semi-Private leakage. Tasks go to vendor APIs under zero-data-retention agreements, and ARC Prize says plainly what that implies: "because tasks are sent to external APIs, we acknowledge the possibility of limited leakage over time. This is why we call it the 'Semi-Private' set." Its defenses are zero data retention and "the release of successive ARC-AGI benchmark versions on a roughly annual basis." Only the Fully Private set stays sealed.

Harness sensitivity. The Duke harness result above shows a handcrafted harness can take a single environment from 0.0% to 97.1%. The Verified board shows the same effect at full scale with vendor harnesses. Moving from Standard to Provider Adapter takes GPT-6 Sol from 4.62% to 23.03%, Gemini 3.8 Flash from 10.37% to 35.00%, GPT-6.1 Sol from 52.73% to 96.37% at xhigh effort, and Astra from 62.71% to 99.95%. Any ARC-AGI-3 number quoted without its harness is close to meaningless.

Effort is not monotonic. Astra on the Standard harness scores 35.18% with reasoning off, beating its own low-effort run at 17.45%. On the Adapter, Astra at high effort beats Astra at max. On ARC-AGI-2, Opus 5.5 scores 93.33% at high, 92.5% at xhigh and 91.67% at max. ARC Prize reports that GPT-5.6 Terra's "xhigh performed worse than max."

Cost runs backwards on Astra Standard. Astra's max-effort run cost $26,098 and its medium run $48,090, 1.84 times as much for a lower score. ARC Prize explains: "Higher reasoning levels generally cost less because Astra solves games in fewer actions, reducing the total number of model calls and tokens." On an interactive benchmark, a model that thinks longer per step takes fewer steps, and the bill follows the step count.

The run cost cap. The policy caps runs at $10,000, while Astra's Standard runs are published at $26,098 to $49,791 and Opus 5's at $20,657. ARC Prize does not say what a run includes.

Tools change what is measured. In PRO-LONG, a code-sandbox harness from an ARC-AGI-3 red-teaming partner, Astra wrote per-game board parsers, search algorithms and solvers. ARC Prize's verdict: "PRO-LONG's results should be understood as the combined performance of the model and its tools," under conditions that differ from the human testing.

What a saturated score would not prove. ARC Prize writes that "saturating the benchmark would not represent 'proof of achieving AGI'," says of Astra "we are not claiming that it is AGI," and notes that ARC-AGI-3 "does not represent the complexity and open-endedness of the real world." A 99.95% Adapter score is a strong result on 55 games, not a verdict on artificial general intelligence.

What is missing. ARC Prize has not published model wall-clock time per task, or what share of ARC-AGI-3 failures come from poor exploration versus inferring the wrong goal. That second number would tell agent builders where to spend effort, and nobody has it.

Most LLM benchmarks give the model a question and check the answer. ARC-AGI-3 gives it neither.

GPQA asks expert-written multiple-choice science questions. It is a knowledge test, and a model with more biology in its weights does better. ARC-AGI uses no language and no domain knowledge, so that advantage disappears.

Humanity's Last Exam collects hard closed-form academic questions with a single correct answer. ARC-AGI-3 has no question at all and withholds the goal, so the model has to discover what counts as success before it can pursue it.

GAIA tests tool-using assistants on questions that need the web and files. Official ARC-AGI-3 runs allow no tools and score action efficiency against human play, which is a measure of how fast an agent learns an environment, not how well it uses a browser.

OSWorld 2 and Terminal-Bench 4 are agentic computer-use benchmarks, closer to ARC-AGI-3 in that the model takes actions over many steps. Both state the instruction and the target end state. ARC-AGI-3 states neither.

Put together, ARC-AGI-2 is the cleanest test of static rule induction, and ARC-AGI-3 tests whether an agent can work out the task itself. Neither tells you how a model handles your documents or your codebase.

What it means for teams choosing a model

On static rule induction, the frontier models are close and the cost decides. GPT-6.1 Sol at max effort scores 94.17% for $0.254 per task. Astra at max adds 0.83 points for 4.4 times the cost. For most teams, Sol is the better buy, and GPT-6.1 Sol at high effort, 91.67% at $0.135, is the cheapest strong result on the board, 8.3 times cheaper than Astra for 3.33 fewer points.

Claude Opus 5.5 at high effort reaches 98% of Astra's score at 36% of its cost, 93.33% for $0.408. GPT-6.1 Sol at max beats it on both axes, so Opus 5.5 is off the frontier. Within Anthropic's lineup it is the clear ARC-AGI-2 value. Fable 5.1 at xhigh costs 7.6 times as much for 3.3 fewer points, and Fable 5.1 at max costs 11 times as much.

On goal-unknown interactive work with a neutral interface, no model is close to solved. Astra leads at 62.71%, GPT-6.1 Sol follows at 52.73%, Opus 5 sits at 30.16% and Gemini 3.8 Flash at 10.37%. No Claude model after Opus 5 has an ARC-AGI-3 score, so a team choosing between Opus 5.5 and an OpenAI model for exploratory agent work has no ARC-AGI-3 data on the Anthropic side.

The harness moves the score as much as the model does. Provider-managed context, with reasoning state carried between requests and compaction for long games, lifts GPT-6.1 Sol from 52.73% to 96.37% and Astra from 62.71% to 99.95%. GPT-6.1 Sol at xhigh on the Adapter reaches 96.37% for $4,360 per run, 3.6 points below Astra's best Adapter run at 23% of its cost. If you are building long-horizon agents, turn on the provider's context management before you rank models. Otherwise you are measuring your harness.

Do not default to max effort. Opus 5.5 scores higher at high than at max, at 22% of the max-effort cost, and Astra on the Adapter does better at high than at max. I would run an effort sweep on your own tasks before setting a production default, because the best setting differs by model and by workload. And ignore ARC-AGI-1 when comparing current models, since the top five organizations all sit between 97.5% and 98.5%.

Public scores give you a shortlist. To see which model and effort level hold up on your own tasks, compare current models on the LLM leaderboard and run the same kind of evaluation on your own workflow in Klu, as described on the research page.