V3 Ultra-Fast Version,The current price is a limited-time 50% discount and will return to the original price on July 31st. The original price is: input: $0.55/M, output: $2.2/M. The model provider is the Sophnet platform. DeepSeek V3 Fast is a high-TPS, ultra-fast version of DeepSeek V3 0324, featuring full-precision (non-quantized) performance, enhanced code and math capabilities, and faster responses!
DeepSeek V3 0324 is a powerful Mixture-of-Experts (MoE) model with a total parameter count of 671B, activating 37B parameters per token.
It adopts Multi-Head Latent Attention (MLA) and the DeepSeekMoE architecture to achieve efficient inference and economical training costs.
It innovatively implements a load balancing strategy without auxiliary loss and sets multi-token prediction training targets to enhance performance.
The model is pre-trained on 14.8 trillion diverse, high-quality tokens and further optimized through supervised fine-tuning and reinforcement learning stages to fully realize its capabilities.
Comprehensive evaluations show that DeepSeek V3 outperforms other open-source models and rivals leading closed-source models in performance.
The entire training process only requires 2.788M H800 GPU hours and remains highly stable, with no irrecoverable loss spikes or rollbacks.
This model was retired on 2026-09-02.
Requests return 404 model_retired. All model retirements
Pricing
- Input Tokens: $0.56 /M tokens
- Output Tokens: $2.24 /M tokens
Input Modalities
- Text
Output Modalities
- Text
Capabilities
- Tools
- Tool calling
- Structured outputs
Providers
Sophnet DeepSeek-V3-Fast
Pricing$0.56$2.24
Context32K
Max output32K
Latency1.5S
Throughput150.0TPS
Uptime
0.00% uptime 3 days ago
0.00% uptime 2 days ago
0.00% uptime yesterday
Performance for DeepSeek-V3-Fast
Uptime is the percentage of requests that succeeded over the past 72 hours. AIHubMix continuously monitors every provider and automatically retries with the next-best provider when one returns an error or responds too slowly; Latency is total round-trip time (lower is better); Throughput is how fast the model writes (tokens per second, higher is better).
Uptime
Loading...
Latency
Loading...
Throughput
Loading...
Try this model
Python
