Key Specifications
| Vendor | google |
|---|
| Version | 2.0-flash-thinking |
|---|
| Release Date | 2024-12-19 |
|---|
| Context Window | 1.048576e+06 tokens |
|---|
| Input Modalities | text, image |
|---|
| Output Modalities | text |
|---|
| License | Proprietary |
|---|
| Documentation | https://ai.google.dev/gemini-api/docs |
|---|
Benchmark Performance
| Benchmark | Score | Unit | Evaluated At | Notes | Source |
|---|
| MMLU | 86.5 | % | 2024-12-19 | 5-shot | view |
| HUMANEVAL | 87.2 | pass@1 | 2024-12-19 | — | view |
| GSM8K | 91.3 | % | 2024-12-19 | 0-shot CoT | view |
| MATH | 56.6 | % | 2024-12-19 | 0-shot CoT | view |
| BBH | 87.9 | % | 2024-12-19 | 3-shot CoT | view |
| GPQA | 61 | % | 2024-12-19 | 0-shot | view |
| IFEVAL | 82.8 | % | 2024-12-19 | prompt_strict | view |
| ARC | 95 | % | 2024-12-19 | challenge | view |
| MUSR | 70.8 | % | 2024-12-19 | 0-shot | view |
| WINOGRANDE | 88.1 | % | 2024-12-19 | 0-shot | view |
Pricing
| Tier | Price | Currency |
|---|
| Input | $0.1 / Mtok | USD |
| Output | $0.4 / Mtok | USD |
| Cache Read | $0 / Mtok | USD |
| Cache Write | $0 / Mtok | USD |
Source:
https://ai.google.dev/pricing
· as of 2024-12-19
Compliance
- Data Residency: US
- SOC2: ✓
- HIPAA: ✗
- GDPR: ✓
- ISO 27001: ✓
Gemini 2.0 Flash Thinking
モデル概要
Google Gemini 2.0 Flash Thinking 实验版, 1M 上下文, 链式思维推理, 在数学与编码上接近 o1。
コア仕様
| ベンダー | バージョン | リリース日 | コンテキストウィンドウ | 入力モダリティ | 出力モダリティ | ライセンス |
|---|
| Google | 2.0-flash-thinking | 2024-12-19 | 1048K | text, image | text | Proprietary |
ベンチマークパフォーマンス
| ベンチマーク | スコア | 単位 | 備考 |
|---|
| MMLU (Massive Multitask Language Understanding) | 86.5 | % | 5-shot |
| HumanEval | 87.2 | pass@1 | — |
| GSM8K (Grade School Math 8K) | 91.3 | % | 0-shot CoT |
| MATH | 56.6 | % | 0-shot CoT |
| BBH (BIG-Bench Hard) | 87.9 | % | 3-shot CoT |
| GPQA | 61.0 | % | 0-shot |
| IFEval | 82.8 | % | prompt_strict |
| ARC | 95.0 | % | challenge |
| MUSR | 70.8 | % | 0-shot |
| WinoGrande | 88.1 | % | 0-shot |
料金
| 入力 | 出力 | キャッシュ読み取り | キャッシュ書き込み |
|---|
| — | — | — | — |
100万トークンあたり
強み
- MMLU score 86.5, strong knowledge reasoning.
- HumanEval 87.2, excellent code generation.
- GSM8K 91.3, robust math reasoning.
- 支持文本、图像、音频多模态输入。
弱み
ユースケース
- 代码生成与调试
- 视觉与图像理解
- Agent 工作流与工具调用
参考文献