What is Internal Research Debugging?
Internal Research Debugging is an OpenAI-internal evaluation that asks whether a model can find and resolve real bugs from OpenAI's own research experiments. The set has 41 bugs whose original fixes, in OpenAI's words, "took hours to days to debug by experienced OpenAI researchers." It also includes 6 alignment-auditing tasks, where the model has to rediscover misaligned behavior or broken environments that OpenAI found in real experiments, without being told what to look for.
OpenAI built the eval for the AI Self-Improvement category of its Preparedness Framework, the part of its safety process that tracks whether models are getting good at accelerating AI research itself. It first appeared in the GPT-5.6 System Card on July 9, 2026, and OpenAI reused it in the GPT-6 Astra System Card on September 3 and the GPT-6.1 Sol addendum on September 29.
The rationale in the Astra card is plain and, I think, correct about why debugging is a good early signal:
"Bugs in a research experiment can waste compute and significantly increase the amount of time required to test research hypotheses. Many debugging tasks also require searching through large quantities of information – but do not require novel infrastructure – which leads us to expect they may be an early bellwether for increases in research capability."
That last point is the interesting one. Building a new training system needs infrastructure a model can't conjure inside a sandbox. Debugging a broken run mostly needs patience, reading, and the judgment to tell a real cause from a red herring. If models get better at research, this is where it shows first.
OpenAI also says why it swapped this suite in for older measures. Monorepo-Bench was saturating, and OPQA turned out on review to contain problems "that were not solvable under test conditions and thus made overall results harder to interpret." The new suite, per the Astra card, captures "realistic, end-to-end tasks that newer models can attempt, rather than older tasks."
Only OpenAI runs it. The tasks, rubric, and harness are private, and no outside party can reproduce a score.
How the tasks and scoring work
Each task drops the model into a failed research experiment and asks it to work out what went wrong and fix it. Table 18 of the Astra card describes the capability as "Can models find and resolve real bugs in internal OpenAI research experiments that took researchers hours to days to fix?" That makes it an agent task over a large codebase and a pile of experiment artifacts, closer to an on-call investigation than a coding puzzle. The volume of logs and code involved is why OpenAI reads high scores as "better ability to search large codebases, inspect experiments, and identify likely causes of failures," and why context handling matters as much as code-writing skill.
The alignment-auditing tasks add a twist. The model isn't told there is misaligned behavior to find. It has to notice that a training environment is bad or that a model in the experiment learned something it shouldn't have, which ties the eval to alignment work as well as engineering. OpenAI states the 41 and the 6 separately and says only that the eval "also includes" the audit tasks. It does not say how the two groups combine into one score.
Scoring is a mean rubric score from 0 to 100 percent, higher is better. The Sol addendum labels its axis "Rubric score (average reward)" and the Astra card uses "Mean rubric reward (%) (higher is better)." Rubric scoring means partial credit. A model that finds the right subsystem but proposes a wrong fix can earn some points, which is different from the pass-or-fail unit tests of SWE-bench.
OpenAI has not published the tool set, scaffold, container contents, time or step limits, reasoning effort behind each reported score, attempts per task, grader identity (human or LLM judge), rubric contents, per-task weights, or what the cost axis in its charts is measured per. Every number below should be read with that list in mind.
Versions
There is no version numbering. The GPT-5.6 and GPT-6 Astra cards describe the tasks in identical text, and the Astra card says it uses "the suite of evaluations that we released in our GPT-5.6 Sol launch." Scores across the three documents are from the same eval. OpenAI does note in the GPT-5.6 card that comparison values for earlier models come "from recent snapshots," so an older model's number can shift slightly between cards.
Thresholds
OpenAI compares each model against an "indicative threshold for High capability" in AI Self-Improvement. The numeric threshold is not published. GPT-6 Astra scores 78.05% "while still being below our indicative threshold for High capability," and the Sol addendum says GPT-6.1 Sol "remains below the High threshold." The Astra appendix (A.8.1.3.1) uses a different name, saying Astra "remains below the Critical threshold."
Current leaderboard
GPT-6 Astra leads at 78.05%, 2.53 points ahead of GPT-6.1 Sol at 75.52%. GPT-5.6 Sol sits at 68.32%, and GPT-6 Sol scored 64.20%, below the older GPT-5.6 Sol.
All four values come from one sentence in section 9.1.3.1 of the Sol addendum, and Figure 37 of that addendum plots the same four bars. The date column is the publication date of the card where each model first appears, not the date of the run.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | 78.05% | OpenAI internal harness; settings not published | 2026-09-03 | Astra card 10.1.3.1; restated in Sol addendum 9.1.3.1 |
| GPT-6.1 Sol | OpenAI | 75.52% | OpenAI internal harness; settings not published | 2026-09-29 | Sol addendum 9.1.3.1, "improves meaningfully on GPT-6 Sol" |
| GPT-5.6 Sol | OpenAI | 68.32% | OpenAI internal harness; settings not published | 2026-07-09 | Sol addendum 9.1.3.1; first reported in the GPT-5.6 card |
| GPT-6 Sol | OpenAI | 64.20% | OpenAI internal harness; settings not published | 2026-09-03 | Sol addendum 9.1.3.1; 4.12 points below GPT-5.6 Sol |
Comparable within this table only as vendor-reported OpenAI runs on one private task set; there are no error bars, and no non-OpenAI model has a score.
The Astra card's Figure 84 adds OpenAI's smaller Luna models. Its labeled bars for Astra, GPT-5.6 Sol, and GPT-6 Sol match the table above, so the Luna rows sit on the same scale.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| GPT-5.6 Luna | OpenAI | 50.8% | OpenAI internal harness; settings not published | 2026-07-09 | Astra card Figure 84 bar label |
| GPT-6 Luna | OpenAI | 46.62% | OpenAI internal harness; settings not published | 2026-09-03 | Astra card Figure 84 bar label; appendix A.8.1.3.1 says GPT-6 Sol and Luna "score slightly lower than their predecessors" |
Comparable to the first table because Figure 84 carries the same Astra and Sol values; it omits GPT-6.1 Sol, Terra, GPT-5.5, and GPT-5.4.
The GPT-6 generation split in two. Astra jumped 9.73 points over GPT-5.6 Sol, while GPT-6 Sol fell 4.12 points below GPT-5.6 Sol and GPT-6 Luna fell 4.18 below GPT-5.6 Luna. GPT-6.1 Sol fixed the Sol regression, landing 11.32 points above GPT-6 Sol and 7.20 above GPT-5.6 Sol. A new model number didn't guarantee a better debugger in this generation, and a team that moved from GPT-5.6 Sol to GPT-6 Sol took a step backward on this eval.
Score against cost
The only dollar chart OpenAI published for this eval is Figure 37 of the GPT-5.6 card. It plots rubric score against "API Cost (USD)" for five models, each as a curve of markers. The card doesn't label the markers, though each curve's spread matches what you'd expect from different reasoning-effort settings. It also doesn't say whether cost is per task or per full run. The values below were read from the plotted markers, calibrated to the axis ticks; Sol's top marker reads 68.2% against the labeled 68.32%, so expect about half a point of error.
Terra is the efficient choice in this generation. Its top marker hits 67.7% at $4.86, within half a point of Sol's 68.2% at $6.64, for 27% less plotted cost. Below about $1.30 the Sol, Terra, and Luna curves overlap, so at low budgets the model choice barely moves the score. GPT-5.5 and GPT-5.4 are dominated. GPT-5.6 Luna reaches 50.7% at $1.70, matching GPT-5.5's best of 50.0% at $2.90.
Score against output tokens
OpenAI has not published dollar costs for any GPT-6 or GPT-6.1 model on this eval, and list prices don't tell you how many tokens each run consumed, so there is no cost chart for Astra. What exists is Figure 52 of the Astra card, which plots rubric score against output tokens for Astra and GPT-5.6 Sol. The card doesn't state the unit basis, such as mean per task. Values are read from the figure and calibrated to the labeled 78.05% and 68.32% end points.
| Model | Marker | Output tokens | Score |
|---|---|---|---|
| GPT-6 Astra | 1 | 2.3K | 54.9% |
| GPT-6 Astra | 2 | 5.6K | 71.1% |
| GPT-6 Astra | 3 | 6.9K | 75.6% |
| GPT-6 Astra | 4 | 10.1K | 78.1% |
| GPT-5.6 Sol | 1 | 2.0K | 30.2% |
| GPT-5.6 Sol | 2 | 5.0K | 48.1% |
| GPT-5.6 Sol | 3 | 8.9K | 55.2% |
| GPT-5.6 Sol | 4 | 12.9K | 60.7% |
| GPT-5.6 Sol | 5 | 20.1K | 68.3% |
Comparable only between these two models on Figure 52; output tokens are not cost, and the unit basis is unpublished.
This is the clearest efficiency gap OpenAI has published for this eval. Astra at about 5.6K output tokens scores 71.1%, already above GPT-5.6 Sol's best at about 20K. To reach roughly 55%, Astra needs about 2.3K tokens and Sol needs about 8.9K, close to four times as many. Astra gets far more debugging done per output token.
Historical progression
The eval is three months old, so the history is short. All rows are OpenAI, vendor-reported, on the same eval.
| Date | Model | Score | Source and notes |
|---|---|---|---|
| 2026-07-09 | GPT-5.4 | ~47.7% | GPT-5.6 card Figure 37, top marker read from figure |
| 2026-07-09 | GPT-5.5 | ~50.0% | GPT-5.6 card Figure 37, top marker read from figure |
| 2026-07-09 | GPT-5.6 Luna | 50.8% | Labeled in Astra card Figure 84; reads 50.7% on GPT-5.6 Figure 37 |
| 2026-07-09 | GPT-5.6 Terra | ~67.7% | GPT-5.6 card Figure 37, top marker read from figure |
| 2026-07-09 | GPT-5.6 Sol | 68.32% | Labeled in Sol addendum; reads 68.2% on GPT-5.6 Figure 37 |
| 2026-09-03 | GPT-6 Luna | 46.62% | Astra card Figure 84 |
| 2026-09-03 | GPT-6 Sol | 64.20% | Astra card Figure 84; Sol addendum 9.1.3.1 |
| 2026-09-03 | GPT-6 Astra | 78.05% | Astra card 10.1.3.1, "improves meaningfully over GPT-5.6 Sol" |
| 2026-09-29 | GPT-6.1 Sol | 75.52% | Sol addendum 9.1.3.1 |
Rows marked ~ were read from plotted markers and carry about half a point of error; GPT-5.4 and GPT-5.5 values are "from recent snapshots" per the GPT-5.6 card.
The top score climbed about 18 points from GPT-5.5 to GPT-5.6 Sol, then another 9.73 points to GPT-6 Astra about eight weeks later. The GPT-5.6 card summed up the first jump as Sol and Terra improving "meaningfully over GPT-5.5 and GPT-5.4 on real internal research debugging tasks." The Sol line itself went 68.32, then 64.20, then 75.52.
Documented failure modes and limits
No model solves the set. The Sol addendum says "All models still solve only a subset of difficult debugging tasks that can take experienced researchers hours or days to resolve," and the Astra and GPT-5.6 cards both say "research debugging remains not fully solved." Even Astra leaves about a fifth of the available rubric credit on the table.
No error bars on a small set. With 41 bugs and 6 audit tasks, a few tasks swing the mean. OpenAI publishes no per-task scores, weights, or confidence intervals for this eval, even though the Astra card shows 95% bootstrap intervals on some other figures. The 2.53-point gap between Astra and GPT-6.1 Sol has no published uncertainty attached.
Nobody else can check it. Tasks, rubric, and harness are private, and every score comes from OpenAI. OpenAI has also not published a contamination analysis.
The cost axis is underspecified. GPT-5.6 Figure 37 states neither the cost unit basis nor the effort level per marker, and no dollar figure exists for GPT-6 models.
Single vendor. Only OpenAI models are scored. Anthropic's Claude Fable 5.1 and Mythos 5.1 materials and Google DeepMind's Gemini 4 methodology page don't report this eval, and they couldn't, since it isn't public.
The threshold is a moving reference. OpenAI calls its High threshold "indicative," gives no number, and names a different threshold (Critical) for Astra in the appendix than in the main text. Treat the threshold language as OpenAI's safety judgment, not a measurement you can compare against.
How it compares to related benchmarks
SWE-bench. SWE-bench uses public GitHub issues in open-source repositories and scores pass or fail against unit tests. Internal Research Debugging uses private ML-experiment bugs, partial-credit rubrics, and problems that took OpenAI researchers hours to days. SWE-bench gives you a multi-lab leaderboard; this eval gives you harder, more realistic tasks from one lab.
KernelGen 1P. KernelGen sits in the same OpenAI suite. The agent gets "a kernel-development environment, benchmark harness, reference materials, and performance/correctness tests" and has to write a correct, faster kernel while "avoiding invalid shortcuts such as host-side compute or benchmark spoofing." The two evals don't track each other. On KernelGen, the Sol addendum's Figure 38 gives GPT-5.6 Sol 61.08%, GPT-6 Sol 38.12%, GPT-6 Astra 66.72%, and GPT-6.1 Sol 60.44%. GPT-6.1 Sol "performs comparably to GPT-5.6 Sol" on kernels but scores 7.20 points higher on debugging. A model that writes good GPU code isn't automatically the one that finds the bug in your training run.
NanoGPT. Also in the suite, NanoGPT gives the model one H100 and rewards training-loop speedups with clamp((t_baseline - t_trial) / t_baseline, 0, 1). "The current best human solution achieves a score of 72.38%." It measures optimization speed, not diagnosis. OpenAI's own caveat there, that "the task is constrained to a small training setup and does not demonstrate the ability to design, derisk, and operate frontier scale pretraining runs," is about NanoGPT, not about debugging.
PostTrainBench Lite and MLE-Bench Revised. The other two suite members use public tasks. PostTrainBench Lite gives an open-source base model, one H100, internet access, a Hugging Face dataset cache, and five hours across 12 model and benchmark combinations, with an LLM judge that zeroes the reward for cheating. MLE-Bench Revised has 72 public ML-competition problems scored by percentile rank against a harness-generated reference distribution. Both test building; Internal Research Debugging tests fixing.
Terminal-Bench and GAIA. Terminal-Bench 4 and Terminal-Bench Science are public terminal tasks with verifiable outputs and results from many labs. GAIA asks general-assistant questions with short answers and no codebase or experiment artifacts at all. For a cross-vendor read on agentic engineering, those are where to look.
What it means for teams choosing a model
If your work looks like this eval, meaning ML experiments that fail in non-obvious ways and long investigations through logs and code, GPT-6 Astra is the strongest documented choice at 78.05%. GPT-6.1 Sol is 2.53 points behind. OpenAI publishes no cost for either model on this eval, so it can't tell you which one wins on price per fixed bug. Run both on a sample of your own past bugs before committing.
Within the Sol line, GPT-6.1 Sol beats GPT-6 Sol by 11.32 points and GPT-5.6 Sol by 7.20. If you're still on GPT-6 Sol for debugging work, you're on a model that scored 4.12 points below GPT-5.6 Sol.
Effort settings move the score a lot. GPT-5.6 Sol's curve runs from about 30% to 68% across its markers, and Astra's runs from about 55% to 78%. Astra at about 5.6K output tokens already beats GPT-5.6 Sol at about 20K. Budget for test-time compute and measure the effort level you plan to ship, not just the top score.
In the GPT-5.6 generation, Terra gets within half a point of Sol at 27% less plotted cost, so Terra was the better buy for debugging there.
Every model still misses bugs that took experts hours to days. Keep a human reviewer on root-cause calls for failed runs, especially when the model's explanation sounds confident.
This eval can't settle a choice between OpenAI, Anthropic, and Google, because only OpenAI models have scores. Use public agentic benchmarks with multi-lab results for that, then test on your own bugs. The LLM leaderboard tracks the public numbers across frontier models. To build the same kind of eval from your own failed experiments and score candidates against a rubric, the research workflow page shows how that runs in Klu.