Models & Pricing

MODELdeepseek-flash(1)deepseek-v4-pro
BASE URL (OpenAI Format)https://api.deepseek.com
BASE URL (Anthropic Format)https://api.deepseek.com/anthropic
MODEL VERSIONDeepSeek-V4.1-FlashDeepSeek-V4-Pro-0813
THINKING MODESupports both non-thinking and thinking (default) modes
See Thinking Mode for how to switch
CONTEXT LENGTH1M
MAX OUTPUTMAXIMUM: 384K
FEATURESJson Output✓✓
Tool Calls✓✓
Responses API✓✓
Anthropic API✓✓
Chat Prefix Completion(Beta)✓✓
FIM Completion(Beta)Non-thinking mode onlyNon-thinking mode only
Vision✓Not supported
PRICING(2)1M INPUT TOKENS
(CACHE HIT)
OFF-PEAK$0.003$0.022
PEAK$0.006$0.044
1M INPUT TOKENS
(CACHE MISS)
OFF-PEAK$0.15$0.66
PEAK$0.3$1.32
1M OUTPUT TOKENSOFF-PEAK$0.6$1.98
PEAK$1.2$3.96
Concurrency Limit(3)2500500

(2) Off-peak rates are half of the peak rates.