GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Stephen M. Walker II · Co-Founder / CEO

GPQA: A Graduate-Level Google-Proof Q&A Benchmark

GPQA, or Graduate-Level Google-Proof Q&A Benchmark, is a challenging dataset designed to evaluate the capabilities of Large Language Models (LLMs) and scalable oversight mechanisms. Introduced by researchers, GPQA comprises 448 multiple-choice questions across the domains of biology, physics, and chemistry, crafted by domain experts to ensure high quality and difficulty.

There are three variations of the test dataset with variable question lengths: extended (546), main (448), and diamond (198). The original research benchmark compared zero-shot, few-shot, CoT, and search variations.

GPQA Diamond Vendor-Reported Comparison (July 9, 2026)

The table below reproduces the no-tools GPQA Diamond comparison published with GPT-5.6 on July 9, 2026 and retrieved on July 15, 2026. It is a vendor-reported comparison, not a neutral leaderboard. Google's model page separately reports the same 94.3% result for Gemini 3.1 Pro Preview.

ProviderModelGPQA Diamond accuracyTools
OpenAIGPT-5.6 Sol94.6%None
AnthropicClaude Mythos Preview94.6%None
GoogleGemini 3.1 Pro Preview94.3%None
AnthropicClaude Mythos 594.1%None
OpenAIGPT-5.593.6%None

GPQA results remain sensitive to the evaluation harness, prompt, answer extraction, sampling, and reasoning settings. Vendor comparison tables may use settings chosen for each model, so small score differences should not be treated as definitive rank gaps without a controlled reproduction.

GPQA Eval Leaderboard

The GPQA Eval Leaderboard, updated as of June 26, 2024, showcases AI model performance on the challenging 198-question Diamond Set of the GPQA benchmark. Anthropic's Claude 3.5 Sonnet leads with 59.4% zero-shot Chain-of-Thought accuracy, followed by OpenAI's GPT-4 Opus (0513) at 53.6% and Anthropic's Claude 3 Opus at 50.4%. These results demonstrate significant progress in Anthropic and OpenAI models' ability to handle complex, graduate-level scientific questions across biology, physics, and chemistry.

Historical GPQA Diamond Zero-Shot-CoT Snapshot (June 26, 2024)

The page archived the following 198-question Diamond Set comparison as a June 26, 2024 snapshot; the GPT-4o model label was corrected in the repository on July 10, 2024. Here, zero-shot chain-of-thought (CoT) means the model was prompted to reason through the problem without worked examples before selecting an answer. The original page did not cite an immutable source or record the exact prompts, harness, answer extraction, sampling, reasoning settings, tool access, or complete model snapshots. A dash means the archived table did not report a score.

OrganizationModelDiamond Set (198 questions)
AnthropicClaude 3.5 Sonnet59.4% Zero-shot CoT
OpenAIGPT-4o (0513)53.6% Zero-shot CoT
AnthropicClaude 3 Opus50.4% Zero-shot CoT
OpenAIGPT-4 Turbo (0409)48.0% Zero-shot CoT
GoogleGemini 1.5 Pro (052024)46.2% Zero-shot CoT
GoogleGemini 1.5 Pro (022024)41.4% Zero-shot CoT
AnthropicClaude 3 Sonnet40.4% Zero-shot CoT
GoogleGemini 1.5 Pro39.5% Zero-shot CoT
OpenAIGPT-4 (0314)35.7% Zero-shot CoT
AnthropicClaude 3 Haiku33.3% Zero-shot CoT
MetaLlama-2-70B-chat31.1% Zero-shot CoT
OpenAIGPT-3.528.1% Zero-shot CoT
GoogleGemini 1.0 Ultra
GoogleGemini 1.0 Pro

This table is retained as a broad historical comparison, not a current ranking. Its zero-shot-CoT protocol and 2024 model versions differ from the explicitly no-tools 2026 vendor snapshot above, and the historical table does not even document whether tools were disabled. Absolute scores and rank changes across the two sections are therefore not directly comparable.

Key Features and Performance Insights

  • Expert-Level Difficulty — The questions are designed to be extremely challenging, with domain experts (those with or pursuing PhDs in the relevant fields) achieving an accuracy of 65% (74% when discounting clear mistakes identified in retrospect). This level of difficulty is intended to reflect graduate-level understanding in the respective sciences.
  • Google-Proof Nature — Highly skilled non-expert validators, despite having unrestricted web access and spending over 30 minutes per question on average, only reached a 34% accuracy rate. This "Google-proof" characteristic underscores the benchmark's resistance to simple lookup or shallow web searches, aiming at deeper understanding and reasoning.
  • Performance of AI Systems — The strongest GPT-4 based baseline model achieved a 39% accuracy, highlighting the significant challenge GPQA poses even to state-of-the-art AI systems. This gap between expert human performance and AI capabilities underscores the need for advanced scalable oversight methods to ensure AI systems can provide reliable and truthful information, especially in complex scientific domains.

Current context: The 39% GPT-4-based result is a historical baseline from GPQA's introduction. By the July 2026 vendor-reported no-tools comparison above, reported leading results reached 94.6%. This increase makes evaluation protocol disclosure and harder, less-contaminated scientific tests increasingly important.

How does GPQA compare to other benchmarks like GAIA and BASIS

GAIA: Real-World AI Assistant Assessment GAIA (General AI Assistant Benchmark) evaluates AI systems on practical, real-world tasks that encompass reasoning, multi-modal processing, web browsing, and tool utilization. Despite being conceptually simple for humans, who achieve 92% accuracy, GAIA poses significant challenges for AI, with GPT-4 (with plugins) scoring only 15%. This stark performance gap underscores GAIA's effectiveness in benchmarking AI systems' robustness and adaptability across diverse, everyday scenarios, emphasizing the need for AI to match or exceed average human performance on practical tasks.

BASIS: Frontier of Scientific AI Capabilities BASIS (Benchmark for Advanced Scientific Inquiry Systems) pushes the boundaries of AI evaluation in scientific domains, surpassing even GPQA in complexity. Tailored for assessing AI systems expected to perform at or beyond human expert level, BASIS focuses on tasks demanding advanced scientific inquiry and reasoning. This benchmark is crucial for developing and evaluating AI systems capable of contributing meaningfully to cutting-edge scientific research and problem-solving, potentially accelerating breakthroughs across various scientific disciplines.

Current comparison context: Where GAIA tests broad, practical assistant tasks, GPQA is narrower and deeper: no-tools GPQA Diamond runs isolate graduate-level scientific reasoning in biology, physics, and chemistry. GPQA also has search-enabled evaluation variants, so tool access must be reported with every score. The benchmarks are complementary and expose different aspects of practical versatility and specialized reasoning.

Objectives and Implications

GPQA, the Graduate-Level Google-Proof Q&A Benchmark, rigorously evaluates Large Language Models (LLMs) through 448 meticulously crafted multiple-choice questions spanning biology, physics, and chemistry. This benchmark probes LLMs' capacity for deep comprehension and sophisticated reasoning within these scientific domains, serving as a critical metric for scalable oversight mechanisms. GPQA's design specifically targets the development of robust methodologies enabling human experts to effectively supervise and validate AI outputs, particularly in domains where AI capabilities may surpass human expertise.

The advent of GPQA represents a significant milestone in AI assessment, directly addressing the critical need for models capable of processing and generating precise information in specialized scientific fields. As AI technology advances, GPQA and similar benchmarks become indispensable tools for quantifying progress towards AI systems capable of meaningful contributions to scientific research. These benchmarks drive the evolution of increasingly sophisticated AI architectures, aiming to minimize the disparity between AI-generated content and human expert knowledge. Ultimately, GPQA's rigorous standards promote the development of AI systems that can reliably produce truthful and accurate scientific information, potentially accelerating the pace of scientific discovery and innovation.

More terms

Continue exploring the glossary.

Learn how teams define, measure, and improve LLM systems.

Glossary term

What is Human Intelligence?

Human Intelligence refers to the mental quality that consists of the abilities to learn from experience, adapt to new situations, understand and handle abstract concepts, and use knowledge to manipulate one's environment. It is a complex ability influenced by various factors, including genetics, environment, culture, and education.
Read term

Glossary term

What is particle swarm optimization?

Particle swarm optimization (PSO) is a computational method that optimizes a problem by iteratively trying to improve a candidate solution with regard to a given measure of quality. It is a population-based stochastic optimization technique developed by Dr. Eberhart and Dr. Kennedy in 1995, inspired by social behavior of bird flocking or fish schooling.
Read term

It's time to build

Collaborate with your team on reliable Generative AI features.
Want expert guidance? Book a 1:1 onboarding session from your dashboard.

Talk to sales