Glossary term
𝕏 Grok-2 Beta Release
Top tip
This page documents the original Grok-2 beta release from August 2024. x.ai has since shipped multiple newer generations (including Grok 3 and Grok 4); see x.ai/news and the x.ai model docs for current models.
What was Grok-2?
Grok-2 was a language model released by x.ai in August 2024, building upon the capabilities of its predecessor, Grok-1.5. Released on the 𝕏 platform, Grok-2 and its smaller counterpart, Grok-2 mini, were designed to provide intelligent chat assistants with improved reasoning, chat, and coding functionality compared to earlier Grok models.
The table below reflects benchmark scores x.ai reported at the August 2024 release.
| Benchmark | Grok-1.5 | Grok-2 mini‡ | Grok-2‡ |
|---|---|---|---|
| GPQA | 35.9% | 51.0% | 56.0% |
| MMLU | 81.3% | 86.2% | 87.5% |
| MMLU-Pro | 51.0% | 72.0% | 75.5% |
| MATH§ | 50.6% | 73.0% | 76.1% |
| HumanEval¶ | 74.1% | 85.7% | 88.4% |
| MMMU | 53.6% | 63.2% | 66.1% |
| MathVista | 52.8% | 68.1% | 69.0% |
| DocVQA | 85.6% | 93.2% | 93.6% |

At launch, Grok-2 mini was made available on x.com, with both models later released to x.ai's Enterprise API.
Performance Benchmarks (August 2024, at release)
The scores below were reported by x.ai at Grok-2's August 2024 release and reflect a snapshot in time against the frontier models available then. They do not reflect the current state of any of these model families, several of which have since been superseded by newer generations.

The Grok-2 models were evaluated across a series of academic benchmarks, including reasoning, reading comprehension, math, science, and coding. Both Grok-2 and Grok-2 mini showed improvements over the prior Grok-1.5 model, and were competitive at the time with other leading models in areas such as graduate-level science knowledge (GPQA), general knowledge (MMLU, MMLU-Pro), and math competition problems (MATH). Grok-2 also performed strongly on vision-based tasks, including visual math reasoning (MathVista) and document-based question answering (DocVQA).
| Benchmark | Grok-2 | Gemini Pro 1.5 | Llama 3 405B | GPT-4o | Claude 3.5 Sonnet |
|---|---|---|---|---|---|
| GPQA | 56.0% | 46.2% | 51.1% | 53.6% | 59.6% |
| MMLU | 87.5% | 85.9% | 88.6% | 88.7% | 88.3% |
| MMLU-Pro | 75.5% | 69.0% | 73.3% | 72.6% | 76.1% |
| MATH§ | 76.1% | 67.7% | 73.8% | 76.6% | 71.1% |
| HumanEval¶ | 88.4% | 71.9% | 89.0% | 90.2% | 92.0% |
| MMMU | 66.1% | 62.2% | 64.5% | 69.1% | 68.3% |
| MathVista | 69.0% | 63.9% | — | 63.8% | 67.7% |
| DocVQA | 93.6% | 93.1% | 92.2% | 92.8% | 95.2% |
Real-Time Information Integration

At release, Grok-2 integrated real-time information from the 𝕏 platform, letting it draw on up-to-date posts when responding to queries.
Enterprise API Access
Following the initial 𝕏 platform launch, Grok-2 and Grok-2 mini became accessible through x.ai's enterprise API, giving developers a way to integrate Grok-2's capabilities into their own applications.
Where Grok-2 fits today
Grok-2 was x.ai's flagship model as of August 2024. It has since been superseded by later Grok generations (including Grok 3 and Grok 4); check x.ai/news and the x.ai model docs for current model offerings and benchmarks.
†GPT-4-Turbo and GPT-4o scores are from the May 2024 release.
††Claude 3 Opus and Claude 3.5 Sonnet scores are from the June 2024 release.
‡ Grok-2 MMLU, MMLU-Pro, MMMU and MathVista were evaluated using 0-shot CoT.
§ For MATH, we present maj@1 results.
¶ For HumanEval, we report pass@1 benchmark scores.
More terms
Continue exploring the glossary.
Glossary term
What is binary classification?
It's time to build
Collaborate with your team on reliable Generative AI features.
Want expert guidance? Book a 1:1 onboarding session from your dashboard.