DeepSeek V4 Flash
DeepSeek V4 Flash is worth evaluating for coding, rag, and agents when its provider route and context window match the workload.
Use it for
- Teams evaluating coding, rag, and agents
- Workloads that can use a 1m context window
- Buyers comparing 4 tracked provider routes
Do not use it for
- Vision or document-understanding workloads
- Family
- DeepSeek V4
- Released
- 2026-04-24
- Context
- 1m
- Max output
- 384,000
- Parameters
- 284B
- Architecture
- Mixture of Experts
- Specialization
- general
- Openness
- Open source
- License
- MITOSI-approvedCommercial use: permitted
- Weights
- Available
- Code
- Unknown
- Training
- Pretrained
Cheapest of 5 routes · OpenRouter · cache read $0.018
About
DeepSeek V4 Flash is a 284B parameter (13B activated) Mixture-of-Experts language model with 1M-token context. Features a hybrid attention architecture combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) for efficient long-context inference. Supports thinking and non-thinking modes. Legacy API aliases deepseek-chat and deepseek-reasoner map to this model's non-thinking and thinking modes respectively. Pricing: $0.14/1M input, $0.28/1M output (cache hit: $0.0028/1M input). MIT licensed.
DeepSeek V4 Flash is an open-source model in the DeepSeek V4 family. The structured metadata tracks a 1m-token context window, reasoning, function calling, tool use, and structured outputs. This page tracks provider routes through DeepSeek Platform, OpenRouter, Microsoft Foundry, and 2 more, with the cheapest tracked route listed at $0.09 input and $0.18 output per 1M tokens. Headline tracked benchmarks include Google-Proof Q&A 88.1, MMLU PRO 86.4, and SWE-bench Verified 79.0.
Top use-case fit: coding, agents, and build tasks
Coding
Q/$ B4 relevant benchmarks in the decision map.
RAG
Included by capability and metadata signals in the decision map.
Agents
Q/$ A1 relevant benchmark in the decision map.
Provider price ladder
Compare all 5Compare API pricing across 4 providers for input and output tokens, batch, and cached reads when available.
| Provider | Input / 1M | Output / 1M | Cache | Route |
|---|---|---|---|---|
| OpenRouter | $0.090 | $0.180 | read $0.018 | Serverless |
| DeepSeek Platform | $0.140 | $0.280 | read $0.0028 | Serverless |
| Novita AI | $0.140 | $0.280 | read $0.028 | Serverless |
| Vercel AI Gateway | $0.140 | $0.280 | read $0.0028 | Serverless |
Available via routers & gateways(8)
LiteLLM
GatewayOpen-source Python SDK and proxy server that unifies 100+ LLM APIs behind a single OpenAI-compatible interface, with load balancing, cost tracking, and configurable failover.
OpenRouter
HybridUnified hybrid gateway to 400+ models from 60+ providers via a single OpenAI-compatible API, with optional auto-routing that selects the best model per prompt.
Portkey
GatewayProduction AI gateway routing to 1,600+ LLMs with failover, load balancing, semantic caching, and guardrails; Apache 2.0 core is fully self-hostable with the complete feature set.
Azure AI Foundry Model Router
RouterMicrosoft Azure AI Foundry's native model router that uses a trained ML model to route each prompt in real time to the optimal Azure-hosted model, with Balanced/Cost/Quality mode selection and automatic failover.
Helicone
GatewayObservability-first AI gateway with routing, caching, rate limiting, and request tracing; Apache 2.0 open-source core with a managed hosted tier for logging and analytics.
Kong AI Gateway
GatewayMulti-LLM AI gateway built on Kong Gateway 3.x, adding semantic routing, load balancing, guardrails, and MCP traffic analytics as plugins over Kong's existing API management platform.
Capabilities
Benchmark peer barsfor Coding
Benchmark scores(16)
| Benchmark | Score | Version | Evaluation | Source |
|---|---|---|---|---|
| Google-Proof Q&A | 88.1 | GPQA Diamond pass@1, V4-Flash Think MaxObserved 2026-07-03 | — | Source |
| MMLU PRO | 86.4 | MMLU-Pro EM, V4-Flash Think HighObserved 2026-07-03 | — | Source |
| SWE-bench Verified | 79.0 | SWE-bench Verified resolved, V4-Flash Think MaxObserved 2026-07-03 | — | Source |
| SWE-bench Pro | 52.6 | SWE-bench Pro resolved, V4-Flash Think MaxObserved 2026-07-03 | — | Source |
| LiveCodeBench | 91.6 | LiveCodeBench pass@1, V4-Flash Think MaxObserved 2026-07-03 | — | Source |
| HumanEval | 69.5 | HumanEval pass@1, DeepSeek-V4-Flash-Base, 0-shotObserved 2026-07-03 | — | Source |
| Massive Multitask Language Understanding | 88.7 | MMLU EM, DeepSeek-V4-Flash-Base, 5-shotObserved 2026-07-03 | — | Source |
| Terminal-Bench 2.0 | 56.9 | Terminal Bench 2.0 accuracy, V4-Flash Think MaxObserved 2026-07-03 | — | Source |
| GeneBench-Pro | 2.4 | xhighObserved 2026-06-30 | — | Source |
| Humanity's Last Exam | 34.8 | Humanity's Last Exam pass@1, no tools, V4-Flash Think MaxObserved 2026-07-03 | — | Source |
| Humanity's Last Exam — With Tools | 45.1 | Humanity's Last Exam with tools pass@1, V4-Flash Think MaxObserved 2026-07-03 | — | Source |
| Chatbot Arena | 1437.0 | LMArena Chatbot Arena text leaderboard Elo, deepseek-v4-flash-thinkingObserved 2026-07-03 | — | Source |
| SWE-bench Multilingual | 73.3 | SWE-bench Multilingual resolved, V4-Flash Think MaxObserved 2026-07-03 | — | Source |
| MCP-Atlas | 69.0 | MCPAtlas pass@1, V4-Flash Think MaxObserved 2026-07-03 | — | Source |
| BrowseComp | 73.2 | BrowseComp pass@1, V4-Flash Think MaxObserved 2026-07-03 | — | Source |
| Mathematics Aptitude Test of Heuristics | 57.4 | MATH EM, DeepSeek-V4-Flash-Base, 4-shotObserved 2026-07-03 | — | Source |
Migration checks
No linked migration route is available for this model yet.
Rankings & picks(8)
Compare DeepSeek V4 Flash with other models
Comparison and alternatives
Browse all comparisons →Show all 69 popular comparisonssorted by 7-day search impressions
Frequently asked questions
What is the context window of DeepSeek V4 Flash?
DeepSeek V4 Flash has a context window of 1m tokens.
What is the max output of DeepSeek V4 Flash?
DeepSeek V4 Flash can generate up to 384,000 output tokens.
How much does DeepSeek V4 Flash cost?
DeepSeek V4 Flash pricing ranges from $0.09/1M to $0.19/1M input tokens depending on the provider.
When was DeepSeek V4 Flash released?
DeepSeek V4 Flash was released on 2026-04-24.
Which providers offer DeepSeek V4 Flash?
DeepSeek V4 Flash is available from 5 providers: DeepSeek Platform, OpenRouter, Microsoft Foundry, Vercel AI Gateway, Novita AI.
What benchmarks has DeepSeek V4 Flash been tested on?
DeepSeek V4 Flash has been evaluated on 16 benchmarks, including Google-Proof Q&A, MMLU PRO, SWE-bench Verified, SWE-bench Pro, LiveCodeBench.
Cheapest of 5 routes · OpenRouter · cache read $0.018