Skip to content
Dashboard

Deepseek V4 Flash 0731

Deepseek V4 Flash 0731 is a dated checkpoint of DeepSeek's V4 Flash weights, with notably stronger agentic capabilities than the April preview. Deepseek V4 Flash 0731 scores 82.7 on Terminal-Bench, pairs a hybrid attention architecture with a context window of 1.0M tokens, and routes across Fireworks, DeepSeek, Baseten, DeepInfra, Novita AI, Alibaba Cloud, Wafer, GMICloud on AI Gateway. Your use is subject to DeepSeek's Terms & Privacy Policies.

ReasoningTool UseImplicit Caching
import { streamText } from 'ai'
const result = streamText({
model: 'deepseek/deepseek-v4-flash-0731',
prompt: 'Why is the sky blue?'
})
Read docs

About Deepseek V4 Flash 0731

Deepseek V4 Flash 0731 became available on AI Gateway on July 31, 2026 as a dated checkpoint of DeepSeek V4 Flash, the efficiency tier of DeepSeek's V4 generation. The V4 series arrived in April 2026 with two variants: V4 Pro for agentic coding, formal mathematical reasoning, and long-horizon workflows, and V4 Flash for high-volume, latency-sensitive work. V4 Flash performs close to V4 Pro on reasoning and holds up on simpler agent tasks at a smaller parameter size.

This checkpoint carries the updated weights that followed the April preview, and the agentic gains are the reason to pin this checkpoint. On Terminal-Bench, a benchmark that measures how well a model completes real tasks in a terminal, Deepseek V4 Flash 0731 scores 82.7, up 25.8 points from 56.9 in the April preview. That lands directly on coding-agent work such as fixing failing tests in a repository and opening a pull request.

The V4 architecture combines Compressed Sparse Attention (CSA) with Heavily Compressed Attention (HCA), and uses ManifoldConstrained Hyper-Connections (mHC) in place of standard residual connections. The combination targets efficient inference at the context window of 1.0M tokens. Maximum output is 1M tokens per response, which gives long reasoning chains and tool-call sequences room to finish in a single call.

AI Gateway routes Deepseek V4 Flash 0731 across Fireworks, DeepSeek, Baseten, DeepInfra, Novita AI, Alibaba Cloud, Wafer, GMICloud, with retries and automatic failover. Provider coverage grew after the updated weights first shipped, so check the list on this page for who serves the checkpoint today. Prefer a specific provider by passing an order array under providerOptions.gateway, which is also how you route to a provider running a promotional rate while one is active.

Set the model to deepseek/deepseek-v4-flash-0731 in the AI SDK, Chat Completions API, Responses API, Messages API, or other API formats, from TypeScript or Python. Authenticate with an AI Gateway API key or OIDC token, so you don't need a separate DeepSeek platform account. AI Gateway mirrors provider pricing with no markup and adds no platform fee on inference, including for Bring Your Own Key requests.