Models & Pricing
| MODEL | deepseek-flash(1) | deepseek-v4-pro | ||
| BASE URL (OpenAI Format) | https://api.deepseek.com | |||
| BASE URL (Anthropic Format) | https://api.deepseek.com/anthropic | |||
| MODEL VERSION | DeepSeek-V4.1-Flash | DeepSeek-V4-Pro-0813 | ||
| THINKING MODE | Supports both non-thinking and thinking (default) modes See Thinking Mode for how to switch | |||
| CONTEXT LENGTH | 1M | |||
| MAX OUTPUT | MAXIMUM: 384K | |||
| FEATURES | Json Output | ✓ | ✓ | |
| Tool Calls | ✓ | ✓ | ||
| Responses API | ✓ | ✓ | ||
| Anthropic API | ✓ | ✓ | ||
| Chat Prefix Completion(Beta) | ✓ | ✓ | ||
| FIM Completion(Beta) | Non-thinking mode only | Non-thinking mode only | ||
| Vision | ✓ | Not supported | ||
| PRICING(2) | 1M INPUT TOKENS (CACHE HIT) | OFF-PEAK | $0.003 | $0.022 |
| PEAK | $0.006 | $0.044 | ||
| 1M INPUT TOKENS (CACHE MISS) | OFF-PEAK | $0.15 | $0.66 | |
| PEAK | $0.3 | $1.32 | ||
| 1M OUTPUT TOKENS | OFF-PEAK | $0.6 | $1.98 | |
| PEAK | $1.2 | $3.96 | ||
| Concurrency Limit(3) | 2500 | 500 | ||
(2) Off-peak rates are half of the peak rates.