Flagship 1M context multimodal architecture with ultra-cheap prefix caching ($0.003/M)
| Scenario | Input Tokens | Output Tokens | Uncached Cost | With Prompt Caching |
|---|---|---|---|---|
| Short Chat Query | 1,000 | 500 | $0.00060 | $0.00041 |
| Document Summarization | 10,000 | 2,000 | $0.00360 | $0.00165 |
| Codebase & Context Analysis | 100,000 | 20,000 | $0.0360 | $0.0165 |
| Batch Corpus Processing | 1,000,000 | 100,000 | $0.2800 | $0.0850 |
Ideal for production workloads demanding balanced capabilities, deep context depth (1.05M tokens), and reliability from DeepSeek. Excellent when predictable tokenomics and prompt caching support are paramount.
If your use-case requires sub-second streaming latency or ultra-high frequency classification at micro-cent pricing, consider lighter budget options such as Gemini Flash-Lite or Claude Haiku. For deep formal logic, consider dedicated reasoning models like o3.
For DeepSeek-V4.1-Flash, 1 million input tokens costs $0.20, while 1 million output tokens costs $0.80. If using prompt caching, repetitive input prefixes are discounted to $0.01 per million.
DeepSeek-V4.1-Flash features a maximum context window of 1,048,576 tokens (~786,432 words), with a maximum output limit of 384,000 tokens per completion.
DeepSeek-V4.1-Flash utilizes the DeepSeek Byte-Level BPE (100k vocabulary). Token counting on TokenMath runs client-side to ensure maximum privacy.
Yes. DeepSeek-V4.1-Flash supports prompt caching with a cached input rate of $0.01/1M (saving up to 98% on repeated input context).