August 14, 2024

𝕏 Grok-2 Beta Release

Stephen M. Walker II · Co-Founder / CEO

Top tip

This page documents the original Grok-2 beta release from August 2024. x.ai has since shipped multiple newer generations (including Grok 3 and Grok 4); see x.ai/news and the x.ai model docs for current models.

What was Grok-2?

Grok-2 was a language model released by x.ai in August 2024, building upon the capabilities of its predecessor, Grok-1.5. Released on the 𝕏 platform, Grok-2 and its smaller counterpart, Grok-2 mini, were designed to provide intelligent chat assistants with improved reasoning, chat, and coding functionality compared to earlier Grok models.

The table below reflects benchmark scores x.ai reported at the August 2024 release.

BenchmarkGrok-1.5Grok-2 mini‡Grok-2‡
GPQA35.9%51.0%56.0%
MMLU81.3%86.2%87.5%
MMLU-Pro51.0%72.0%75.5%
MATH§50.6%73.0%76.1%
HumanEval¶74.1%85.7%88.4%
MMMU53.6%63.2%66.1%
MathVista52.8%68.1%69.0%
DocVQA85.6%93.2%93.6%
Grok-2 Factuality Preference

At launch, Grok-2 mini was made available on x.com, with both models later released to x.ai's Enterprise API.

Performance Benchmarks (August 2024, at release)

The scores below were reported by x.ai at Grok-2's August 2024 release and reflect a snapshot in time against the frontier models available then. They do not reflect the current state of any of these model families, several of which have since been superseded by newer generations.

Grok-2 Win Rate

The Grok-2 models were evaluated across a series of academic benchmarks, including reasoning, reading comprehension, math, science, and coding. Both Grok-2 and Grok-2 mini showed improvements over the prior Grok-1.5 model, and were competitive at the time with other leading models in areas such as graduate-level science knowledge (GPQA), general knowledge (MMLU, MMLU-Pro), and math competition problems (MATH). Grok-2 also performed strongly on vision-based tasks, including visual math reasoning (MathVista) and document-based question answering (DocVQA).

BenchmarkGrok-2Gemini Pro 1.5Llama 3 405BGPT-4oClaude 3.5 Sonnet
GPQA56.0%46.2%51.1%53.6%59.6%
MMLU87.5%85.9%88.6%88.7%88.3%
MMLU-Pro75.5%69.0%73.3%72.6%76.1%
MATH§76.1%67.7%73.8%76.6%71.1%
HumanEval¶88.4%71.9%89.0%90.2%92.0%
MMMU66.1%62.2%64.5%69.1%68.3%
MathVista69.0%63.9%—63.8%67.7%
DocVQA93.6%93.1%92.2%92.8%95.2%

Real-Time Information Integration

Grok-2 Real-Time Information Integration

At release, Grok-2 integrated real-time information from the 𝕏 platform, letting it draw on up-to-date posts when responding to queries.

Enterprise API Access

Following the initial 𝕏 platform launch, Grok-2 and Grok-2 mini became accessible through x.ai's enterprise API, giving developers a way to integrate Grok-2's capabilities into their own applications.

Where Grok-2 fits today

Grok-2 was x.ai's flagship model as of August 2024. It has since been superseded by later Grok generations (including Grok 3 and Grok 4); check x.ai/news and the x.ai model docs for current model offerings and benchmarks.


† GPT-4-Turbo and GPT-4o scores are from the May 2024 release.

†† Claude 3 Opus and Claude 3.5 Sonnet scores are from the June 2024 release.

‡ Grok-2 MMLU, MMLU-Pro, MMMU and MathVista were evaluated using 0-shot CoT.

§ For MATH, we present maj@1 results.

¶ For HumanEval, we report pass@1 benchmark scores.

More terms

Continue exploring the glossary.

Learn how teams define, measure, and improve LLM systems.

Glossary term

What is lazy learning?

Lazy learning is a method in machine learning where the generalization of the training data is delayed until a query is made to the system. This is in contrast to eager learning, where the system tries to generalize the training data before receiving queries.
Read term

Glossary term

What is binary classification?

Binary classification is a type of supervised learning algorithm in machine learning that categorizes new observations into one of two classes. It's a fundamental task in machine learning where the goal is to predict which of two possible classes an instance of data belongs to. The output of binary classification is a binary outcome, where the result can either be positive or negative, often represented as 1 or 0, true or false, yes or no, etc.
Read term

It's time to build

Collaborate with your team on reliable Generative AI features.
Want expert guidance? Book a 1:1 onboarding session from your dashboard.

Talk to sales   →