Key Specifications

Vendordeepseek
Versionv3
Release Date2024-12-26
Context Window64000 tokens
Input Modalitiestext
Output Modalitiestext
LicenseDeepSeek License
Documentationhttps://api-docs.deepseek.com/

Benchmark Performance

BenchmarkScoreUnitEvaluated AtNotesSource
MMLU88.5%2024-12-265-shotview
HUMANEVAL82.6pass@12024-12-26view
GSM8K89.3%2024-12-260-shot CoTview
MATH61.6%2024-12-260-shot CoTview
BBH84.9%2024-12-263-shot CoTview

Pricing

TierPriceCurrency
Input$0.27 / MtokUSD
Output$1.1 / MtokUSD
Cache Read$0.07 / MtokUSD
Cache Write$0.27 / MtokUSD

Source: https://api-docs.deepseek.com/quick_start/pricing · as of 2024-12-26

Compliance

  • Data Residency: CN
  • SOC2: ✗
  • HIPAA: ✗
  • GDPR: ✗
  • ISO 27001: ✗

DeepSeek V3

모델 개요

DeepSeek V3 是开源 MoE 架构模型,总参数 671B、活跃参数 37B,64K 上下文窗口,在 MMLU、HumanEval、MATH 等基准上达到闭源旗舰水平,价格仅为同级模型的 1/10。

핵심 사양

공급업체버전출시일컨텍스트 창입력 모달리티출력 모달리티라이선스
Deepseekv32024-12-2664KtexttextDeepSeek License

벤치마크 성능

벤치마크점수단위비고
MMLU (Massive Multitask Language Understanding)88.5%5-shot
HumanEval82.6pass@1
GSM8K (Grade School Math 8K)89.3%0-shot CoT
MATH61.6%0-shot CoT
BBH (BIG-Bench Hard)84.9%3-shot CoT

가격

입력출력캐시 읽기캐시 쓰기

백만 토큰당

강점

  • MMLU score 88.5, strong knowledge reasoning.
  • HumanEval 82.6, excellent code generation.
  • GSM8K 89.3, robust math reasoning.
  • 采用 MoE 混合专家架构。

약점

  • 闭源专有模型,不支持自托管。

사용 사례

  • 代码生成与调试

참고문헌