Overview

评估大型语言模型严格遵循指令的能力,包含500个可验证的指令

Metrics

MetricUnitDirection
prompt_strict_acc%↑ Higher is better
prompt_loose_acc%↑ Higher is better

Sources

Model Score Ranking

#ModelVendorScore
1Gemini 1.5 Pro 002google87.9
2GPT-4o (2024-05-13)openai86.7
3Gemini 2.0 Flashgoogle85.9
4GPT-4o (2024-08-06)openai85.5
5o1 Previewopenai84.4
6Qwen1.5 110Balibaba83.4
7Sonar Reasoningother83.1
8Gemini 2.0 Flash Thinkinggoogle82.8
9o1openai82.2
10Claude 3.5 Sonnet (2024-10-22)anthropic81.6
11Gemini 1.0 Ultragoogle81.6
12Yi Visionother81.4
13Command Nightlycohere80.8
14StableLM 2 12Bother80
15Qwen1.5 14Balibaba79.9
16DBRX Baseother79.3
17Llama 3.3 70Bmeta79.3
18Yi 1.5 34Bother79.3
19Claude 3 Opusanthropic79.2
20GLM-4 Plusother79.2
21Jamba Instructother79.2
22Sonar Hugeother78.7
23DeepSeek V2deepseek78.6
24Hermes 3 Llama 3.1 405Bother78.2
25GPT-4 Turboopenai77.9
26Mistral Mediummistral77.4
27Gemini 1.5 Flash 002google77.2
28Grok-2 Visionxai76.9
29Mistral Small 3mistral76.9
30Qwen2 72Balibaba76.9
31DeepSeek LLM 67Bdeepseek76.8
32Qwen1.5 72Balibaba76.4
33Mistral Largemistral75.8
34GLM-4 Airother75.7
35Claude 3 Opus (2024-02-29)anthropic75.5
36Claude 3 Sonnet (2024-02-29)anthropic75.5
37Claude 3 Haiku (2024-03-07)anthropic75.3
38Microsoft WizardLM 2 8x22Bother75.3
39Nous Hermes 2 Mixtral 8x7Bother75.3
40Gemini 1.0 Progoogle75.1
41Sonar Largeother74.9
42GPT-4 32Kopenai74.7
43GPT-4openai74.6
44Llama 3.1 Nemotron 70Bmeta74.5
45Llama 3.2 90B Visionmeta74.5
46Claude 3 Sonnetanthropic74.3
47Gemini 1.5 Flash-8B 002google74.3
48Phi-3.5 MoEother74.3
49Orca 2 13Bother74.2
50Command Rcohere74.1
51Qwen2.5 14Balibaba74.1
52Mistral Smallmistral74
53GPT-4 1106 Previewopenai73.9
54Hermes 3 Llama 3.1 70Bother73.8
55Llama 3.1 70Bmeta73.7
56Claude 3.5 Haikuanthropic73.6
57GPT-4 Visionopenai73.5
58Grok-2 Minixai73.5
59Qwen2.5 32Balibaba73.5
60Yi Largeother73.4
61DBRX Instructother73.1
62o1 miniopenai73.1
63GPT-4 Vision Previewopenai72.7
64Gemma 2 27Bgoogle72.5
65Llama 3 70Bmeta72.2
66Command R+ (08-2024)cohere72.1
67Mixtral 8x7Bmistral72.1
68Sonar Smallother72.1
69Jamba 1.5 Miniother71.9
70Gemini 1.0 Flashgoogle71.8
71DeepSeek V2 Chatdeepseek71.5
72GPT-4 0125 Previewopenai71.4
73Qwen2 57Balibaba71.2
74Claude 3 Haikuanthropic70.8
75Jamba 1.5other70.8
76Jamba 1.5 Largeother70.8
77WizardLM Team WizardLM 2 8x22Bother70.7
78Yi Large Turboother70.7
79Falcon 180Bother70.5
80Command R (08-2024)cohere70.3
81Zephyr ORPO 141B Alphaother69.7
82NVIDIA Llama 3.1 Nemotron 70Bother69.6
83OLMo 7Bother68.4
84Nous Hermes 2 Yi 34Bother68.3
85Llama 3 8Bmeta68.2
86Qwen1.5 32Balibaba68.2
87Gemini 1.5 Flash-8Bgoogle68.1
88Gemini 1.5 Flashgoogle68
89GPT-4o miniopenai67.9
90GLM-4V 9Bother67.4
91Phi-3 Mediumother67.4
92Command R7Bcohere67.2
93Gemma 7Bgoogle66.9
94GLM-4 Flashother65.9
95Codestral Mambamistral65.5
96Yi 1.5 9Bother65.4
97Gemma 2 9Bgoogle64.3
98Code Llama 70Bmeta63.8
99Mathstral 7Bmistral63.6
100OLMo 7B Instructother63.6
101StarChat2 15B v0.1other63.3
102Llama 3.1 8Bmeta62.1
103Qwen2 7Balibaba61.8
104Code Llama 7Bmeta61.2
105OLMo 1.7 7Bother61.2
106DeepSeek Math 7Bdeepseek61.1
107Phi-3 Smallother61
108Mistral 7B v0.3mistral60.9
109Hermes 3 Llama 3.1 8Bother60.7
110GLM-4 9B Chatother59.9
111MPT 7Bother59.9
112Phi-4other59.6
113StarCoder2 15Bother59.6
114Text Bisongoogle59.5
115Codestralmistral59.4
116Code Bisongoogle59.2
117Mistral 7B v0.2mistral59.2
118Mistral 7B v0.1mistral58.9
119DeepSeek Coder 33Bdeepseek58.8
120Yi 1.5 6Bother58.8
121Code Llama 34Bmeta58.7
122Llama 2 7Bmeta58.5
123Llama 3.2 11B Visionmeta58.4
124Microsoft WizardMath 7B v1other57.7
125GPT-3.5 Turboopenai57.4
126Qwen2 1.5Balibaba57.4
127Falcon 40Bother57.2
128Llama 2 70Bmeta57.1
129Microsoft WizardLM 2 7Bother56.9
130DeepSeek Coder V2deepseek56.7
131OpenChat 3.6 8Bother56.7
132Phi-3 Visionother56.4
133Grok Betaxai56.2
134OLMo 7B SFTother56.2
135Ministral 8Bmistral56.1
136Mistral Nemomistral56
137OLMo 2 1124 7Bother56
138Qwen2.5 3Balibaba56
139Flan-T5 XLother55.7
140TigerBot 70B Chatother55.7
141Flan-UL2other55.6
142Nous Hermes 2 Solar 10.7Bother55.6
143DeepSeek Coder 7Bdeepseek55.4
144Qwen2.5 7Balibaba55.4
145StableLM Zephyr 3Bother55.3
146StarCoder2 3Bother55.2
147Claude 2anthropic55
148Claude 2.1anthropic55
149Phi-3 Miniother54.9
150Yi 6Bother54.9
151StarCoder2 7Bother54
152Command Lightcohere53.9
153Microsoft WizardCoder Python 34Bother53.8
154Phi-1other53.8
155Claude Instant 1anthropic53.2
156Llemma 7Bother52.9
157Nous Capybara 34Bother52.9
158GPT-3.5openai52.8
159Capybara 1.5Bother52.4
160StableCode 3Bother51.8
161Mistral Tinymistral51.7
162Llama Guard 2 8Bmeta51.6
163StableLM 3 4Bother51.6
164Flan-T5 XXLother51.3
165MPT 30Bother51.2
166Zephyr 7B Betaother51.2
167Zephyr 7B Alphaother50.9
168OpenChat 3.5 1210other50.4
169Baichuan2 7B Chatother50.1
170Code Llama 13Bmeta50.1
171Llama Guard 3 8Bmeta50
172Llama 2 13Bmeta49.9
173Jurassic-2 Midother49.8
174Phi-2other48.8
175Phi-3.5 Miniother48.7
176Llama 3.2 3Bmeta48.6
177Baichuan2 13B Chatother48.5
178GPT-3.5 Turbo 16Kopenai48.5
179Qwen2.5 0.5Balibaba48.5
180Chat Bisongoogle48.2
181Yi 34Bother48.2
182Phi-1.5other48
183Jurassic-2 Ultraother47.9
184Ministral 3Bmistral47.8
185Grok Vision Betaxai46.8
186Gemma 2Bgoogle46.4
187Open-Platypusother46
188PaLM 2google44.7
189Nous Capybara 7Bother44.2
190Llama 3.2 1Bmeta44.1
191StableLM 2 1.6Bother44.1
192Argilla Notus 7B v1other42.8
193ChatGLM3 6Bother42.5
194Qwen2.5 1.5Balibaba41.9
195Embed English v3cohere0

IFEval

説明

评估大型语言模型严格遵循指令的能力,包含500个可验证的指令

コア仕様

カテゴリライセンス最終更新
reasoningApache-2.02024-02-01

ベンチマーク

単位
%
%

ソース

公式URL