Anthropic Claude 3

Stephen M. Walker II · Co-Founder / CEO

Top tip

Claude 3, released in March 2024, was a significant upgrade over Claude 2, though it did not offer substantial real-world advantages over GPT-4 Turbo. Anthropic has since released newer generations, including Claude 3.5 and Claude 4.

Anthropic released the Claude 3 model family in March 2024, establishing new benchmarks in a variety of cognitive tasks at the time. This lineup includes Claude 3 Haiku, Claude 3 Sonnet, and Claude 3 Opus, each model offering enhanced performance levels for different user needs in terms of intelligence, speed, and cost.

  • Claude 3 Haiku — The fastest and most compact model, designed for near-instant responsiveness, particularly suited for simple queries and requests.
  • Claude 3 Sonnet — Offers a balance between speed and intelligence, excelling in tasks that require rapid responses like knowledge retrieval or sales automation. It is twice as fast as previous Claude models.
  • Claude 3 Opus — The most intelligent model in the family, outperforming peers on benchmarks such as undergraduate and graduate-level knowledge and reasoning, basic mathematics, and more. It exhibits near-human comprehension and fluency in complex tasks.
Klu Anthropic Claude 3 Model Family Sizes

Availability

Anthropic's Claude 3 series introduces a significant leap in Claude-series capabilities, featuring multimodal understanding that processes both text and visual inputs, including photos, charts, graphs, and technical diagrams. This advancement allows for a broader application in various fields, enhancing the AI's ability to interact with and understand complex information.

Klu Claude AI Logic

At launch, Claude 3 was accessible in 159 countries via claude.ai and its API, with Opus and Sonnet available immediately and Haiku following shortly after. Anthropic positioned the models as strong performers for live customer chats, auto-completions, and data extraction tasks, citing benchmark results against competitors like OpenAI's GPT-4. Early independent tests left some room for skepticism.

Claude 3 Benchmark Performance

According to Anthropic's published launch data, the Claude 3 model family outperformed predecessors like GPT-4 and Gemini 1.0 Ultra across various cognitive tasks at the time. The Claude 3 Opus model, the most sophisticated in the lineup, showcased strong performance in key AI benchmarks such as MMLU (undergraduate-level knowledge), GPQA (graduate-level reasoning), GSM8K (basic mathematics), and HumanEval (coding), indicating its prowess in understanding complex tasks, mathematical reasoning, and coding, with near-human comprehension and fluency. Additionally, its advanced computer vision capabilities allow for effective information extraction from visual materials.

Klu Claude 3 Vision

At the time, Claude 3 Opus surpassed the then-leading models from OpenAI and Google on these benchmarks. It performed strongly on undergraduate and graduate knowledge, and grade school math, while the Claude 3 Sonnet model showed a notable ability to understand scientific diagrams, useful for enterprise operations and data analysis.

TestClaude 3 OpusGPT-4Gemini 1.0 Ultra
Undergraduate level knowledge MMLU86.8% 5-shot86.4% 5-shot83.7% 5-shot
Graduate level reasoning GPQA, Diamond50.4% 0-shot CoT35.7% 0-shot CoT
Grade school math GSM8K95.0% 0-shot CoT92.0% 5-shot CoT94.4% maj@32
Math problem-solving MATH60.1% 0-shot CoT52.9% 4-shot53.2% 4-shot
Multilingual math MGSM90.7% 0-shot74.5% 8-shot79.0% 8-shot
Code HumanEval84.9% 0-shot67.0% 0-shot74.4% 0-shot
Reasoning over text DROP, F1 score83.1% 3-shot80.9% 3-shot82.4% Variable shots
Mixed evaluations BIG-Bench-Hard86.8% 3-shot CoT83.1% 3-shot CoT83.6% 3-shot CoT
Knowledge Q&A ARC-Challenge96.4% 25-shot96.3% 25-shot
Common Knowledge HellaSwag95.4% 10-shot95.3% 10-shot87.8% 10-shot

Claude 3's multimodal nature, capable of processing both text and images, significantly expands its application range. Anthropic emphasizes AI safety, actively working to mitigate bias and ensure neutrality, making Claude 3 models a top choice for both enterprise and consumer applications due to their performance, safety features, and broad applicability.

Klu Claude 3 Vision
TestClaude 3 OpusGPT-4 VisionGemini 1.0 Ultra
Math & reasoning MMMU (val)59.4%56.8%59.4%
Document visual Q&A ANLS score, test89.3%88.4%90.9%
Math MathVista (testmini)50.5% CoT49.9%53.0%
Science diagrams AI2D, test88.1%78.2%79.5%
Chart Q&A Relaxed accuracy (test)80.8% 0-shot CoT78.5% 4-shot CoT80.8%

Claude 3 Model Series

The Claude 3 models, particularly Opus, excel in common AI evaluation benchmarks like MMLU, GPQA, and GSM8K, demonstrating near-human comprehension and fluency in complex tasks. These models also improve in analysis, forecasting, content creation, code generation, and multilingual communication.

At launch, Anthropic highlighted the economic potential of Claude 3, particularly the Opus model, in roles such as an economic analyst. In our own testing against GPT-4 Turbo at the time, we found little evidence to back this claim. Amazon reported over 10,000 organizations were already using Bedrock for generative AI applications, and the introduction of Claude 3 models was expected to accelerate adoption further.

Accuracy

Klu Claude 3 Hard Questions

Claude models have historically had issues with accuracy due to their ability to be very creative, most noticeable when iterating on complex storylines. Claude 3 improved on this, with Opus roughly doubling the rate of correct answers on complex, factual questions compared to Claude 2. Anthropic indicated citation support for verifying answers was planned for a future update.

Context length up to 1 million tokens

Klu Claude 3 Retrieval Recall

The Claude 3 models launched with a 200K context window, with Anthropic noting the potential for inputs over 1 million tokens for select customers. The models demonstrated robust recall capabilities, with Opus achieving near-perfect recall in the 'Needle In A Haystack' evaluation.

Speed

At launch, Claude 3 models delivered near-instantaneous results for live chats, auto-completions, and data extraction, with Haiku being the fastest and most cost-effective for its intelligence level. Sonnet offered roughly double the speed of prior Claude models with higher intelligence, while Opus matched their speed but with significantly enhanced intelligence.

Vision

These models also featured vision capabilities, processing various visual formats and offering this capability to enterprise customers with visually encoded knowledge bases.

Instruction Following, Reduced Refusals & AI Safety

The Claude 3 models were designed to be user-friendly, adept at following complex instructions, and capable of producing structured outputs like JSON for various applications.

Claude 2 had a tendency to refuse non-harmful prompts, a side effect of Anthropic's AI Safety measures. Claude 3 improved on this, showing a better understanding of prompts and less frequent unnecessary refusals.

Klu Claude 3 Refusals vs Claude 2

Anthropic's stated commitment to responsible AI development was reflected in the Claude 3 models, which prioritized reducing biases and enhancing safety across various risk scenarios. At launch, these models operated under Anthropic's AI Safety Level 2, with continuous monitoring to evaluate and manage risk levels. The design of Claude 3 models emphasized neutrality and bias mitigation, with Anthropic reporting a lower bias rate compared to previous versions, along with safety features intended to handle sensitive prompts more effectively.

Planned updates (as of March 2024 launch)

The initial release of the Claude 3 models featured a 200K context window, with Anthropic noting plans to extend input capabilities up to 1 million tokens for certain customers. At the time, Google's Gemini 1.5 was also expanding its context window, and Anthropic was expected to continue pushing its generally available context lengths to remain competitive.

Anthropic announced plans to introduce Tool Use (function calling), interactive coding (REPL), and advanced agentic capabilities in subsequent updates.

Anthropic indicated at launch that the intelligence of the Claude 3 models was far from reaching its peak, with frequent updates planned to expand the models' functionalities, especially for enterprise applications and large-scale deployments.

Claude 3 models launched with availability on major platforms including Amazon Bedrock and Google Cloud's Vertex AI, alongside backing from Amazon and Google that supported the models' distribution across industries.

Pricing at launch

At launch, each model had distinct features and costs, catering to different needs:

  • Opus was the most intelligent, positioned for complex tasks at $15 per million tokens for input and $75 for output.
  • Sonnet offered a balance of intelligence and speed, positioned for enterprise workloads at $3 per million tokens for input and $15 for output.
  • Haiku offered the fastest response for simple queries at $0.25 per million tokens for input and $1.25 for output.

At release, Opus and Sonnet were available first, with Haiku following shortly after. Sonnet powered the free experience on claude.ai, with Opus reserved for Claude Pro subscribers. Availability extended to Amazon Bedrock and Google Cloud's Vertex AI Model Garden.

Anthropic has since released newer model generations, including Claude 3.5 and Claude 4; for current models and pricing, see Anthropic's official documentation.

More terms

Continue exploring the glossary.

Learn how teams define, measure, and improve LLM systems.

Glossary term

What is an abstract data type?

An Abstract Data Type (ADT) is a mathematical model for data types, defined by its behavior from the point of view of a user of the data. It is characterized by a set of values and a set of operations that can be performed on these values. The term "abstract" is used because the data type provides an implementation-independent view. This means that the user of the data type doesn't need to know how that data type is implemented, they only need to know what operations can be performed on it.
Read term

Glossary term

What is Algorithmic Probability?

Algorithmic probability, also known as Solomonoff probability, is a mathematical method of assigning a prior probability to a given observation. It was invented by Ray Solomonoff in the 1960s and is used in inductive inference theory and analyses of algorithms.
Read term

It's time to build

Collaborate with your team on reliable Generative AI features.
Want expert guidance? Book a 1:1 onboarding session from your dashboard.

Talk to sales