LMSYS Chatbot Arena Leaderboard

A ranking of large language models built from blind human pairwise votes, started by LMSYS at UC Berkeley and now run by Arena

Chat and preferenceHistoricalArena Elo rating

Stephen M. Walker II · Co-Founder / CEO · 6 min read

This page covers results from July 2023, and newer models aren't on these boards

What is the LMSYS Chatbot Arena Leaderboard?

The LMSYS Chatbot Arena Leaderboard is a comprehensive ranking platform that assesses the performance of large language models (LLMs) in conversational tasks. It uses blind pairwise human votes to evaluate models like GPT-4, Claude, and others, providing a clear view of their strengths and weaknesses in real-world applications.

The leaderboard is updated regularly, reflecting the latest advancements in AI technology. It includes models from leading organizations such as OpenAI, Anthropic, Google, and Meta, showcasing their capabilities in engaging and informative conversations.

Key features of the LMSYS Chatbot Arena Leaderboard include:

  • Diverse Evaluation Metrics — The leaderboard uses multiple metrics for evaluation, including Chatbot Arena Elo, which is based on user feedback, and other performance indicators.
  • Regular Updates — The leaderboard is frequently updated to reflect the latest developments in LLM technology, ensuring that it remains a relevant and valuable resource for AI researchers and developers.
  • Community Engagement — Users can participate in the evaluation process by providing feedback on chatbot interactions, contributing to the dynamic nature of the leaderboard.

LMSYS Chatbot Arena Leaderboard (September 2024)

Historical table retained with its original breadth. Its Arena, MT-Bench, and MMLU columns have incompatible units and imperfect provenance; they are preserved as published and must not be combined.

ModelArena Elo ratingMT-bench (score)MMLULicense
o1-preview1355+12/-112991Proprietary
ChatGPT-4o-latest (2024-09-03)1335+5/-610213Proprietary
o1-mini1324+12/-93009Proprietary
Gemini-1.5-Pro-Exp-08271299+5/-428229Proprietary
Grok-2-08-131294+4/-423999Proprietary
GPT-4o-2024-05-131285+3/-390695Proprietary
GPT-4o-mini-2024-07-181273+3/-330434Proprietary
Claude 3.5 Sonnet1269+3/-362977Proprietary
Gemini-1.5-Flash-Exp-08271269+4/-422264Proprietary
Grok-2-Mini-08-131267+4/-522041Proprietary
Gemini Advanced App (2024-05-14)1267+3/-352218Proprietary
Meta-Llama-3.1-405b-Instruct-fp81266+4/-431280Llama 3.1 Community
Meta-Llama-3.1-405b-Instruct-bf161264+6/-85865Llama 3.1 Community
GPT-4o-2024-08-061263+4/-322562Proprietary
Gemini-1.5-Pro-0011259+3/-380656Proprietary
GPT-4-Turbo-2024-04-091257+3/-292973Proprietary
GPT-4-1106-preview12519.402023/4Proprietary
Mistral-Large-24071250—2024/7Mistral Research
Athene-70b1250—2024/7CC-BY-NC-4.0
Meta-Llama-3.1-70b-Instruct1249—2023/12Llama 3.1 Community
Claude 3 Opus12489.4587.1Proprietary
GPT-4-0125-preview12459.38—Proprietary
Yi-Large-preview1240——Proprietary
Gemini-1.5-Flash-0011227—78.9Proprietary
Deepseek-v2-API-06281219——DeepSeek
Gemma-2-27b-it1218——Gemma license
Yi-Large1212——Proprietary
Gemini App (2024-01-24)1209——Proprietary
Nemotron-4-340B-Instruct1209——NVIDIA Open Model
GLM-4-05201207——Proprietary
Llama-3-70b-Instruct1206—82.0Llama 3 Community
Claude 3 Sonnet12019.2287.0Proprietary
LMSYS Chatbot ArenaArena Elo rating

Vision Leaderboard

The Vision Leaderboard ranks top LLMs based on their performance in vision-based conversations. It evaluates models using metrics like Arena Elo, MT-bench score, and MMLU, providing insights into their strengths and weaknesses.

ModelArena ScoreOrganizationKnowledge Cutoff
Gemini-1.5-Pro-Exp-08271231 (+9/-6)Google2023/11
GPT-4o-2024-05-131209 (+6/-6)OpenAI2023/10
Gemini-1.5-Flash-Exp-08271208 (+11/-12)Google2023/11
Claude 3.5 Sonnet1191 (+6/-4)Anthropic2024/4
Gemini-1.5-Pro-0011151 (+8/-6)Google2023/11
GPT-4-Turbo-2024-04-091151 (+7/-4)OpenAI2023/12
GPT-4o-mini-2024-07-181120 (+6/-5)OpenAI2023/10
Gemini-1.5-Flash-8b-Exp-08271110 (+9/-10)Google2023/11
Qwen2-VL-72B1085 (+26/-19)AlibabaUnknown
Claude 3 Opus1075 (+5/-6)Anthropic2023/8
Gemini-1.5-Flash-0011072 (+7/-6)Google2023/11
InternVL2-26b1068 (+8/-7)OpenGVLab2024/7
Claude 3 Sonnet1048 (+6/-6)Anthropic2023/8
Yi-Vision1039 (+15/-15)01 AI2024/7
qwen2-vl-7b-instruct1037 (+23/-21)AlibabaUnknown
Reka-Flash-Preview-202406111024 (+8/-6)Reka AIUnknown
Reka-Core-202405011015 (+5/-6)Reka AIUnknown
InternVL2-4b1010 (+9/-8)OpenGVLab2024/7
LLaVA-v1.6-34B1000 (+9/-7)LLaVA2024/1
Claude 3 Haiku1000 (+7/-6)Anthropic2023/8
LLaVA-OneVision-qwen2-72b-ov-sft992 (+16/-13)LLaVA2024/8
CogVLM2-llama3-chat-19b990 (+13/-12)Zhipu AI2024/7
MiniCPM-v 2_6976 (+15/-13)OpenBMB2024/7
Phi-3.5-vision-instruct916 (+11/-10)Microsoft2024/8
Phi-3-Vision-128k-Instruct874 (+15/-12)Microsoft2024/3
LMSYS Chatbot ArenaVisionArena Score

The updated leaderboard continues to evaluate a wide spectrum of models from renowned AI research organizations such as OpenAI, Anthropic, Google, Meta, and Reka AI. It provides a comprehensive overview of model performance, considering metrics like Arena Elo rating, MT-bench score, and MMLU score. This latest update underscores the ongoing competition and rapid innovation in AI, with new models consistently pushing the boundaries of performance benchmarks. As of September 2024, the LMSYS Chatbot Arena Leaderboard remains an essential resource for tracking the state-of-the-art in LLM capabilities, offering valuable insights into the evolving landscape of artificial intelligence.

Status in 2026

As of October 2026, the leaderboard is active and runs at arena.ai/leaderboard/text under the name Arena, with its most recent update dated October 2, 2026. Arena ranks models from blind pairwise human votes fitted with a statistical rating model. It also offers a style-control view that adjusts for answer length and formatting. Arena now operates as an independent company that grew out of the LMSYS research group. FastChat, the original LMSYS code base, is no longer where Arena is developed. A 2025 paper, "The Leaderboard Illusion", criticized Arena for allowing private testing of many model variants with selective disclosure, and for sampling that favors proprietary models.

How does the LMSYS Chatbot Arena Leaderboard work?

The LMSYS Chatbot Arena Leaderboard evaluates LLMs through blind pairwise votes from users. Participants can engage with chatbots and provide feedback, which is then used to calculate the Arena Elo rating. This process ensures that the leaderboard reflects both human preferences and objective performance metrics.

The leaderboard is an essential resource for developers and researchers, offering insights into the strengths and weaknesses of various models. It helps identify areas for improvement and guides the development of more advanced conversational AI systems.

What is the purpose of the LMSYS Chatbot Arena Leaderboard?

The purpose of the LMSYS Chatbot Arena Leaderboard is to provide a transparent and dynamic evaluation of LLMs in conversational settings. By ranking models from blind user votes, it offers a comprehensive view of model performance, helping to drive innovation and improvement in AI technology.

The leaderboard is designed to foster collaboration and knowledge sharing among AI researchers and developers, promoting the development of more effective and engaging conversational models.

Future Directions for the LMSYS Chatbot Arena Leaderboard

Future directions for the LMSYS Chatbot Arena Leaderboard include expanding the range of evaluation metrics, incorporating more diverse conversational scenarios, and enhancing user engagement. By continuously evolving, the leaderboard aims to remain at the forefront of AI evaluation, providing valuable insights into the capabilities of the latest LLMs.

Current Arena successor context

The following material is additive and documents the successor project's current snapshots and revised methodology.

The LMSYS Chatbot Arena Leaderboard is a ranking platform that assesses the performance of large language models (LLMs) in conversational tasks through blind human pairwise comparisons and statistical rating models.

The project began as an academic effort from LMSYS (Large Model Systems Organization) at UC Berkeley. It was later rebranded LMArena in 2024 and then simply Arena, and now operates as an independent company (arena.ai) rather than a university research project. The snapshots below use the successor project's July 2026 data.

Key features of the leaderboard included:

  • Diverse Evaluation Metrics — The leaderboard used multiple metrics for evaluation, including Chatbot Arena Elo, which was based on user feedback, alongside other performance indicators such as MT-bench.
  • Frequent Updates — Rankings were updated as new models were added to the arena, so any static snapshot ages quickly.
  • Community Engagement — Users could participate in the evaluation process by voting on head-to-head chatbot responses, contributing to the dynamic nature of the rankings.

Arena Text Overall Snapshot (July 14, 2026)

This text ranking is the style-controlled Overall view in the official text data pinned to its July 14, 2026 revision. Style control reduces the influence of response-style features on preference ratings, while the confidence interval communicates statistical uncertainty.

RankModelRating95% confidence intervalVotes
1claude-fable-51507.51500.1–1515.07,959
2claude-opus-4-6-thinking1503.61499.8–1507.459,871
3claude-opus-4-7-thinking1502.91498.6–1507.247,141
4claude-opus-4-61497.71494.0–1501.463,636
5claude-opus-4-71494.11489.8–1498.448,248
ArenaText, style-controlled OverallRatinghuggingface.co

Arena Vision Overall Snapshot (July 12, 2026)

The vision ranking is a separate style-controlled Overall snapshot from the official vision data pinned to its July 12, 2026 revision. It reflects human preferences on image-and-text conversations, so its ratings should not be compared numerically with the text table.

RankModelRating95% confidence intervalVotes
1claude-fable-51318.11307.4–1328.84,383
2claude-opus-4-7-thinking1304.01296.9–1311.218,567
3claude-opus-4-6-thinking1299.31292.2–1306.318,430
4claude-opus-4-71298.61291.5–1305.619,049
5claude-opus-4-61298.41291.6–1305.122,778
ArenaVision, style-controlled OverallRatinghuggingface.co

Both tables are dated snapshots of a dynamic preference leaderboard. Rankings can move as votes accumulate, models enter or leave the arena, and the statistical methodology changes. The Arena leaderboard publishes the continuously updated view.

Historical LMSYS Text Snapshot (repository archive: September 25, 2024)

The repository page archived on September 25, 2024 contained the following broader LMSYS Chatbot Arena text ranking. The old table mixed Arena ratings with mislabeled MT-Bench, MMLU, uncertainty, vote-count, and release-date fields, so only the model names and Arena ratings that can be coherently recovered are preserved here. The archive did not record a versioned leaderboard export or the precise ranking-method revision.

RankModelArena rating
1o1-preview1355
2ChatGPT-4o-latest (2024-09-03)1335
3o1-mini1324
4Gemini-1.5-Pro-Exp-08271299
5Grok-2-08-131294
6GPT-4o-2024-05-131285
7GPT-4o-mini-2024-07-181273
8Claude 3.5 Sonnet1269
9Gemini-1.5-Flash-Exp-08271269
10Grok-2-Mini-08-131267
11Gemini Advanced App (2024-05-14)1267
12Meta-Llama-3.1-405b-Instruct-fp81266
13Meta-Llama-3.1-405b-Instruct-bf161264
14GPT-4o-2024-08-061263
15Gemini-1.5-Pro-0011259
16GPT-4-Turbo-2024-04-091257
17GPT-4-1106-preview1251
18Mistral-Large-24071250
19Athene-70b1250
20Meta-Llama-3.1-70b-Instruct1249
21Claude 3 Opus1248
22GPT-4-0125-preview1245
23Yi-Large-preview1240
24Gemini-1.5-Flash-0011227
25Deepseek-v2-API-06281219
26Gemma-2-27b-it1218
27Yi-Large1212
28Gemini App (2024-01-24)1209
29Nemotron-4-340B-Instruct1209
30GLM-4-05201207
31Llama-3-70b-Instruct1206
32Claude 3 Sonnet1201
LMSYS Chatbot ArenaTextArena rating

Historical LMSYS Vision Snapshot (repository archive: September 25, 2024)

The same repository archive included this separate vision-conversation ranking. Its displayed ratings and asymmetric uncertainty values are retained, while unrelated knowledge-cutoff metadata has been omitted because it was not part of the Arena evaluation. The archive did not specify whether its uncertainty values use the same confidence construction as the current dataset.

RankModelArena ratingReported uncertainty
1Gemini-1.5-Pro-Exp-08271231+9/-6
2GPT-4o-2024-05-131209+6/-6
3Gemini-1.5-Flash-Exp-08271208+11/-12
4Claude 3.5 Sonnet1191+6/-4
5Gemini-1.5-Pro-0011151+8/-6
6GPT-4-Turbo-2024-04-091151+7/-4
7GPT-4o-mini-2024-07-181120+6/-5
8Gemini-1.5-Flash-8b-Exp-08271110+9/-10
9Qwen2-VL-72B1085+26/-19
10Claude 3 Opus1075+5/-6
11Gemini-1.5-Flash-0011072+7/-6
12InternVL2-26b1068+8/-7
13Claude 3 Sonnet1048+6/-6
14Yi-Vision1039+15/-15
15qwen2-vl-7b-instruct1037+23/-21
16Reka-Flash-Preview-202406111024+8/-6
17Reka-Core-202405011015+5/-6
18InternVL2-4b1010+9/-8
19LLaVA-v1.6-34B1000+9/-7
20Claude 3 Haiku1000+7/-6
21LLaVA-OneVision-qwen2-72b-ov-sft992+16/-13
22CogVLM2-llama3-chat-19b990+13/-12
23MiniCPM-v 2_6976+15/-13
24Phi-3.5-vision-instruct916+11/-10
25Phi-3-Vision-128k-Instruct874+15/-12
LMSYS Chatbot ArenaVisionArena rating

These historical ratings predate the current style-controlled text and vision datasets. Changes in vote pools, model aliases, category filters, statistical estimation, style control, and uncertainty calculation mean the 2024 and 2026 numbers are not directly comparable. Manual maintenance should locate the exact September 2024 LMSYS exports, recover vote counts and intervals only from coherent source fields, document the historical methodology revision, and refresh the broader tables from pinned current datasets when full-row coverage is needed.

How did the LMSYS Chatbot Arena Leaderboard work?

The leaderboard evaluated LLMs through blind user voting. Participants compared responses from two anonymized models side by side and voted for the better one, and those preferences were aggregated into statistical ratings with uncertainty intervals.

This approach carried over to Arena, the leaderboard's successor: it remains a resource for developers and researchers to compare model strengths and weaknesses and track progress in conversational AI.

What was the purpose of the LMSYS Chatbot Arena Leaderboard?

The purpose of the LMSYS Chatbot Arena Leaderboard was to provide a transparent, crowd-sourced evaluation of LLMs in conversational settings. By aggregating user preferences into statistical ratings, it offered a comparative view of model performance that helped drive innovation and improvement across the industry.

The leaderboard also fostered collaboration and knowledge sharing among AI researchers and developers, encouraging the development of more effective and engaging conversational models.

From LMSYS to Arena

The Chatbot Arena project has since evolved beyond its original academic scope: it was rebranded LMArena in 2024, later shortened to Arena, and now operates as an independent company (arena.ai) rather than a UC Berkeley research project. The July 2026 tables above preserve exact published values from the successor project's official dataset.