Key Specifications

Vendordeepseek
Versionv3
Release Date2024-12-26
Context Window64000 tokens
Input Modalitiestext
Output Modalitiestext
LicenseDeepSeek License
Documentationhttps://api-docs.deepseek.com/

Benchmark Performance

BenchmarkScoreUnitEvaluated AtNotesSource
MMLU88.5%2024-12-265-shotview
HUMANEVAL82.6pass@12024-12-26view
GSM8K89.3%2024-12-260-shot CoTview
MATH61.6%2024-12-260-shot CoTview
BBH84.9%2024-12-263-shot CoTview

Pricing

TierPriceCurrency
Input$0.27 / MtokUSD
Output$1.1 / MtokUSD
Cache Read$0.07 / MtokUSD
Cache Write$0.27 / MtokUSD

Source: https://api-docs.deepseek.com/quick_start/pricing · as of 2024-12-26

Compliance

  • Data Residency: CN
  • SOC2: ✗
  • HIPAA: ✗
  • GDPR: ✗
  • ISO 27001: ✗

DeepSeek V3

Panoramica del modello

DeepSeek V3 是开源 MoE 架构模型,总参数 671B、活跃参数 37B,64K 上下文窗口,在 MMLU、HumanEval、MATH 等基准上达到闭源旗舰水平,价格仅为同级模型的 1/10。

Specifiche principali

FornitoreVersioneData di rilascioFinestra di contestoModalità di inputModalità di outputLicenza
Deepseekv32024-12-2664KtexttextDeepSeek License

Prestazioni benchmark

BenchmarkPunteggioUnitàNote
MMLU (Massive Multitask Language Understanding)88.5%5-shot
HumanEval82.6pass@1
GSM8K (Grade School Math 8K)89.3%0-shot CoT
MATH61.6%0-shot CoT
BBH (BIG-Bench Hard)84.9%3-shot CoT

Prezzi

InputOutputLettura cacheScrittura cache

per milione di token

Punti di forza

  • MMLU score 88.5, strong knowledge reasoning.
  • HumanEval 82.6, excellent code generation.
  • GSM8K 89.3, robust math reasoning.
  • 采用 MoE 混合专家架构。

Punti deboli

  • 闭源专有模型,不支持自托管。

Casi d’uso

  • 代码生成与调试

Riferimenti