Overview

57 学科多选题知识评测,覆盖人文、社科、STEM、医学等领域,共 14042 道题,评估模型的广泛世界知识与问题解决能力。

Metrics

MetricUnitDirection
accuracy%↑ Higher is better

Sources

Model Score Ranking

#ModelVendorScore
1GPT-4o (2024-05-13)openai89.8
2Claude 3.5 Sonnetanthropic88.7
3GPT-4oopenai88.7
4Llama 3.1 405Bmeta88.6
5DeepSeek V3deepseek88.5
6Sonar Reasoningother87.7
7GPT-4o (2024-08-06)openai87.6
8Grok-2xai87.5
9Claude 3.5 Sonnet (2024-10-22)anthropic87.2
10Gemini 1.5 Pro 002google87
11Gemini 2.0 Flash Thinkinggoogle86.5
12Gemini 2.0 Flashgoogle86.3
13o1openai86.3
14Qwen1.5 110Balibaba86.1
15Qwen2.5 72Balibaba86.1
16Gemini 1.5 Progoogle85.9
17GPT-4openai85.8
18Claude 3 Opus (2024-02-29)anthropic85.7
19Yi Visionother85.5
20o1 Previewopenai85.3
21Command Nightlycohere85.2
22GPT-4 Vision Previewopenai85
23Gemini 1.0 Progoogle84.5
24GPT-4 Turboopenai84.4
25Yi Large Turboother84.4
26GLM-4 Plusother84.3
27Jamba 1.5 Largeother84.3
28Claude 3 Sonnet (2024-02-29)anthropic84
29Mistral Large 2mistral84
30GPT-4 Visionopenai83.9
31Hermes 3 Llama 3.1 405Bother83.9
32Grok-2 Visionxai83.4
33Llama 3.3 70Bmeta83.4
34Mistral Mediummistral82.9
35Yi Largeother82.8
36GPT-4 0125 Previewopenai82.6
37Sonar Largeother82.5
38Command R+ (08-2024)cohere82.3
39Mistral Largemistral82
40Qwen1.5 32Balibaba81.9
41GPT-4 1106 Previewopenai81.8
42Claude 3 Sonnetanthropic81.5
43Llama 3.1 Nemotron 70Bmeta81.5
44Nous Hermes 2 Mixtral 8x7Bother81.4
45DBRX Instructother81.3
46Qwen1.5 72Balibaba81.3
47Claude 3 Opusanthropic81.2
48Mistral Smallmistral81.2
49Gemini 1.5 Flash 002google81.1
50Sonar Hugeother81
51GPT-4 32Kopenai80.8
52Grok-2 Minixai80.6
53NVIDIA Llama 3.1 Nemotron 70Bother80.4
54Gemini 1.0 Ultragoogle80.3
55Qwen2 72Balibaba80.2
56GLM-4 Flashother80
57Sonar Smallother80
58DBRX Baseother79.9
59GPT-4o miniopenai79.7
60Llama 3 70Bmeta79.5
61Zephyr ORPO 141B Alphaother79.5
62Command Rcohere79.3
63WizardLM Team WizardLM 2 8x22Bother79.2
64GLM-4 Airother78.9
65Microsoft WizardLM 2 8x22Bother78.9
66Gemini 1.0 Flashgoogle78.7
67Jamba 1.5other78.5
68Qwen2.5 32Balibaba78.5
69Orca 2 13Bother78.4
70Qwen1.5 14Balibaba78.4
71DeepSeek V2deepseek78.2
72o1 miniopenai78.2
73Claude 3 Haikuanthropic77.9
74Jamba 1.5 Miniother77.9
75Command R (08-2024)cohere77.8
76Mixtral 8x22Bmistral77.8
77Qwen2 57Balibaba77.7
78Claude 3.5 Haikuanthropic77.6
79Gemini 1.5 Flash-8B 002google77.5
80Gemma 2 27Bgoogle77.4
81StableCode 3Bother77.4
82Mixtral 8x7Bmistral77.2
83Claude 3 Haiku (2024-03-07)anthropic77.1
84Mistral Small 3mistral76.9
85Nous Hermes 2 Yi 34Bother76.9
86Falcon 180Bother76.7
87Hermes 3 Llama 3.1 70Bother76.7
88DeepSeek V2 Chatdeepseek76.6
89Gemini 1.5 Flash-8Bgoogle76.6
90Llama 3.2 90B Visionmeta76.5
91Jamba Instructother76.2
92Yi 1.5 34Bother76.1
93Gemini 1.5 Flashgoogle76
94Phi-3.5 MoEother75.7
95StableLM 2 12Bother75.7
96Llama 3.1 70Bmeta75.6
97Qwen2.5 14Balibaba75.4
98DeepSeek LLM 67Bdeepseek75.2
99Codestralmistral75.1
100Command R+cohere75
101StarChat2 15B v0.1other74.9
102Code Llama 34Bmeta74.5
103Qwen2.5 7Balibaba73.6
104GLM-4V 9Bother73.4
105OLMo 7B SFTother73.4
106Hermes 3 Llama 3.1 8Bother73.1
107Llama 3.1 8Bmeta73
108Mistral 7B v0.3mistral72.6
109DeepSeek Coder V2deepseek72.3
110Phi-3 Smallother71.8
111TigerBot 70B Chatother71.7
112Argilla Notus 7B v1other71.4
113Llama 2 7Bmeta70.8
114GLM-4 9B Chatother70.7
115OpenChat 3.5 1210other70.5
116Gemma 7Bgoogle69.6
117DeepSeek Math 7Bdeepseek69.3
118Mistral Nemomistral69.2
119OpenChat 3.6 8Bother69.2
120Phi-4other69.2
121Mathstral 7Bmistral69.1
122Microsoft WizardMath 7B v1other69.1
123Open-Platypusother68.7
124Mistral 7B v0.1mistral68.4
125PaLM 2google68.4
126OLMo 2 1124 7Bother68.3
127Phi-1other68.1
128Yi 6Bother68
129Llama 3 8Bmeta67.8
130Yi 1.5 9Bother67.8
131StarCoder2 7Bother67.4
132Claude 2anthropic66.3
133Llemma 7Bother66
134Code Llama 13Bmeta65.9
135DeepSeek Coder 33Bdeepseek65.9
136Microsoft WizardLM 2 7Bother65.7
137OLMo 7B Instructother65.6
138OLMo 1.7 7Bother65.5
139Zephyr 7B Betaother65.1
140Code Bisongoogle65
141Command R7Bcohere64.9
142Ministral 8Bmistral64.9
143Mistral 7B v0.2mistral64.4
144Nous Capybara 7Bother64.4
145Gemma 2 9Bgoogle64.3
146GPT-3.5 Turbo 16Kopenai64.2
147DeepSeek Coder 7Bdeepseek63.6
148Nous Hermes 2 Solar 10.7Bother63.4
149Qwen2 7Balibaba63.4
150Grok Vision Betaxai63.3
151Chat Bisongoogle63.1
152OLMo 7Bother62.6
153Jurassic-2 Ultraother62.5
154Code Llama 70Bmeta62.1
155Phi-3 Mediumother62.1
156Llama 3.2 11B Visionmeta61.9
157GPT-3.5 Turboopenai61.1
158Yi 1.5 6Bother61.1
159Claude 2.1anthropic60.9
160StarCoder2 15Bother60.6
161StarCoder2 3Bother60.6
162Flan-T5 XLother60.4
163Microsoft WizardCoder Python 34Bother60
164Codestral Mambamistral59.3
165Flan-UL2other58.7
166Baichuan2 13B Chatother58.6
167Code Llama 7Bmeta58.5
168Phi-1.5other58.4
169ChatGLM3 6Bother58
170Baichuan2 7B Chatother57.9
171Phi-2other57.5
172Claude Instant 1anthropic57.1
173Qwen2.5 3Balibaba56.8
174Falcon 40Bother56.7
175Llama 2 13Bmeta56.4
176Phi-3 Visionother56.4
177Jurassic-2 Midother55.9
178Phi-3.5 Miniother55.7
179Yi 34Bother55.5
180MPT 30Bother55.2
181Qwen2.5 1.5Balibaba54.5
182Mistral Tinymistral54.2
183Text Bisongoogle53.9
184Nous Capybara 34Bother53.2
185Llama 2 70Bmeta53
186Command Lightcohere52.8
187StableLM 2 1.6Bother52.7
188Llama 3.2 3Bmeta52.5
189Zephyr 7B Alphaother52.2
190Flan-T5 XXLother52
191Capybara 1.5Bother51.7
192MPT 7Bother51.7
193Grok Betaxai51
194Llama Guard 2 8Bmeta51
195StableLM 3 4Bother50.6
196GPT-3.5openai50.3
197Llama 3.2 1Bmeta48.6
198Gemma 2Bgoogle48
199StableLM Zephyr 3Bother44.3
200Phi-3 Miniother43.6
201Ministral 3Bmistral42.4
202Qwen2.5 0.5Balibaba41.5
203Qwen2 1.5Balibaba40.5
204Llama Guard 3 8Bmeta40.1
205Embed English v3cohere0

MMLU (Massive Multitask Language Understanding)

説明

57 学科多选题知识评测,覆盖人文、社科、STEM、医学等领域,共 14042 道题,评估模型的广泛世界知识与问题解决能力。

コア仕様

カテゴリライセンス最終更新
knowledgeMIT2024-01-01

ベンチマーク

単位
%

ソース

公式URL