DeepSeek-V4-Flash-0731(deepseek-v4-flash-0731) is an open-source MoE large language model developed by the Chinese AI company DeepSeek, with support for a million-token context window. It is designed for coding, complex reasoning, tool use, agentic workflows, and long-document processing. Its advantages include strong performance with fewer active parameters and improved efficiency through DSpark speculative decoding. Compared with DeepSeek V4-Flash Preview, it offers significantly stronger coding and agent capabilities, while outperforming DeepSeek V4-Pro Preview on several benchmarks with fewer active parameters.
Pricing
Input Modalities
- Text
Output Modalities
- Text
Context length
- 1M tokens
Capabilities
- Thinking
- Streaming
- Tool calling
- Web search
- URL context
- Code interpreter
- Computer use
- File search
- Memory tool
- Structured outputs
- Citations
- Prompt caching
- Background mode
- Server-side sessions
Providers
Baidu baidu-deepseek-v4-flash-0731
Off-peak14:00–00:00 UTC
Pricing$0.2112$0.6336
Cache Read$0.0211/M tokens
Peak00:00–14:00 UTC
Pricing$0.4226$1.2678
Cache Read$0.0423/M tokens
Context1M
Max output384K
Latency2.5S
Throughput94.0TPS
Uptime
99.99% uptime 2 days ago
100.00% uptime yesterday
100.00% uptime today
Bytedance doubao-deepseek-v4-flash-0731
Pricing$0.4226$1.2678
Cache$0.1409
Context1M
Max output1M
Latency3.8S
Throughput47.1TPS
Uptime
100.00% uptime 2 days ago
100.00% uptime yesterday
100.00% uptime today
Deepinfra deepinfra-deepseek-v4-flash-0731
Pricing$0.198$0.396
Cache$0.0396
Context1M
Max output384K
Latency1.8S
Throughput33.5TPS
Uptime
99.99% uptime 2 days ago
99.99% uptime yesterday
100.00% uptime today
DeepSeek deep-deepseek-v4-flash-0731
Off-peak04:00–06:00, 10:00–01:00 UTC
Pricing$0.2324$0.6972
Cache Read$0.0077/M tokens
Peak01:00–04:00, 06:00–10:00 UTC
Pricing$0.4648$1.3944
Cache Read$0.0155/M tokens
Context1M
Max output384K
Latency1.9S
Throughput81.4TPS
Uptime
100.00% uptime 2 days ago
100.00% uptime yesterday
100.00% uptime today
Azure azure-deepseek-v4-flash-0731
Off-peak14:00–00:00 UTC
Pricing$0.2112$0.6336
Cache Read$0.0211/M tokens
Peak00:00–14:00 UTC
Pricing$0.4226$1.2678
Cache Read$0.0423/M tokens
Context1M
Max output384K
Latency13.0S
Throughput42.3TPS
Uptime
99.98% uptime 2 days ago
100.00% uptime yesterday
100.00% uptime today
Alibaba Cloud alicloud-deepseek-v4-flash-0731
Off-peak14:00–00:00 UTC
Pricing$0.2112$0.6336
Cache Read$0.0211/M tokens
Peak00:00–14:00 UTC
Pricing$0.4226$1.2678
Cache Read$0.0423/M tokens
Context1M
Max output384K
Latency2.5S
Throughput8.1TPS
Uptime
100.00% uptime 2 days ago
100.00% uptime yesterday
100.00% uptime today
Wafer wafer-deepseek-v4-flash-0731-fast
Pricing$0.28$1.4
Cache$0.07
Context1M
Max output384K
Latency2.3S
Throughput53.3TPS
Uptime
100.00% uptime 2 days ago
99.23% uptime yesterday
100.00% uptime today
Performance for deepseek-v4-flash-0731
Uptime is the percentage of requests that succeeded over the past 72 hours. AIHubMix continuously monitors every provider and automatically retries with the next-best provider when one returns an error or responds too slowly; Latency is total round-trip time (lower is better); Throughput is how fast the model writes (tokens per second, higher is better).
Uptime
Loading...
Latency
Loading...
Throughput
Loading...
Try this model
Python
