Deepseek V4 Flash 0731
Deepseek V4 Flash 0731 is a dated checkpoint of DeepSeek's V4 Flash weights, with notably stronger agentic capabilities than the April preview. Deepseek V4 Flash 0731 scores 82.7 on Terminal-Bench, pairs a hybrid attention architecture with a context window of 1.0M tokens, and routes across Fireworks, DeepSeek, Baseten, DeepInfra, Novita AI, Alibaba Cloud, Wafer, GMICloud on AI Gateway. Your use is subject to DeepSeek's Terms & Privacy Policies.
import { streamText } from 'ai'
const result = streamText({ model: 'deepseek/deepseek-v4-flash-0731', prompt: 'Why is the sky blue?'})Playground
Try out Deepseek V4 Flash 0731 by DeepSeek. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.
Deepseek V4 Flash 0731
Providers
Route requests across multiple providers. Copy a provider slug to set your preference. Visit the docs for more info. Using a provider means you agree to their terms, listed under Legal.
| Provider |
|---|
P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the docs for more info.
P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the docs for more info.
Direct request success rate on AI Gateway and per-provider. Visit the docs for more info.
More models by DeepSeek
| Model |
|---|
About Deepseek V4 Flash 0731
Deepseek V4 Flash 0731 became available on AI Gateway on July 31, 2026 as a dated checkpoint of DeepSeek V4 Flash, the efficiency tier of DeepSeek's V4 generation. The V4 series arrived in April 2026 with two variants: V4 Pro for agentic coding, formal mathematical reasoning, and long-horizon workflows, and V4 Flash for high-volume, latency-sensitive work. V4 Flash performs close to V4 Pro on reasoning and holds up on simpler agent tasks at a smaller parameter size.
This checkpoint carries the updated weights that followed the April preview, and the agentic gains are the reason to pin this checkpoint. On Terminal-Bench, a benchmark that measures how well a model completes real tasks in a terminal, Deepseek V4 Flash 0731 scores 82.7, up 25.8 points from 56.9 in the April preview. That lands directly on coding-agent work such as fixing failing tests in a repository and opening a pull request.
The V4 architecture combines Compressed Sparse Attention (CSA) with Heavily Compressed Attention (HCA), and uses ManifoldConstrained Hyper-Connections (mHC) in place of standard residual connections. The combination targets efficient inference at the context window of 1.0M tokens. Maximum output is 1M tokens per response, which gives long reasoning chains and tool-call sequences room to finish in a single call.
AI Gateway routes Deepseek V4 Flash 0731 across Fireworks, DeepSeek, Baseten, DeepInfra, Novita AI, Alibaba Cloud, Wafer, GMICloud, with retries and automatic failover. Provider coverage grew after the updated weights first shipped, so check the list on this page for who serves the checkpoint today. Prefer a specific provider by passing an order array under providerOptions.gateway, which is also how you route to a provider running a promotional rate while one is active.
Set the model to deepseek/deepseek-v4-flash-0731 in the AI SDK, Chat Completions API, Responses API, Messages API, or other API formats, from TypeScript or Python. Authenticate with an AI Gateway API key or OIDC token, so you don't need a separate DeepSeek platform account. AI Gateway mirrors provider pricing with no markup and adds no platform fee on inference, including for Bring Your Own Key requests.
What To Consider When Choosing a Provider
- Configuration: The
0731suffix names a dated checkpoint.deepseek/deepseek-v4-flash-0731stays on that snapshot of the weights, whiledeepseek/deepseek-v4-flashpicks up whatever DeepSeek ships as current. Pin the checkpoint when you need reproducible behavior across deploys and evaluation runs, and use the undated model ID when you want new weights automatically. - Configuration: Provider coverage differs between the two IDs, including which backing providers offer Zero Data Retention. Check the provider list on this page, and use the
orderoption underproviderOptions.gatewaywhen you need traffic to prefer one provider. Individual providers sometimes run time-limited promotional rates that end on a set date, so read the pricing panel on this page rather than assuming a past rate still applies. - Zero Data Retention: AI Gateway supports Zero Data Retention for this model via direct gateway requests (BYOK is not included). To configure this, check the documentation.
- Authentication: AI Gateway authenticates requests using an API key or OIDC token. You do not need to manage provider credentials directly.
When to Use Deepseek V4 Flash 0731
Best for
- Pinned Production Behavior: A dated checkpoint holds model behavior steady across deploys and evaluation runs
- Coding Agent Tasks: Terminal-Bench gains land on repository work such as fixing tests and opening pull requests
- High-Volume Short-Form Traffic: Classification, routing, and extraction pipelines at efficiency-tier pricing
- Long-Context Workloads: A context window of 1.0M tokens and up to 1M tokens output tokens cover large inputs
- Provider-Controlled Routing: The
orderoption pins traffic to a preferred provider for cost or compliance reasons
Consider alternatives when
- Automatic Weight Updates: DeepSeek V4 Flash tracks the current weights with no model ID change on your side
- Deeper Reasoning Work: DeepSeek V4 Pro targets complex reasoning, multi-step problem solving, and agentic planning
- Dedicated Reasoning Specialist: DeepSeek-R1 remains the open-weights option for extended chain-of-thought workloads
Conclusion
Deepseek V4 Flash 0731 pins the V4 Flash checkpoint that carries the stronger agentic scores, at efficiency-tier pricing and a context window of 1.0M tokens. Pin Deepseek V4 Flash 0731 when reproducible behavior matters, and switch to deepseek/deepseek-v4-flash when you'd rather pick up new weights automatically.