by DeepSeek
V3 Ultra-Fast Version,The current price is a limited-time 50% discount and will return to the original price on July 31st. The original price is: input: $0.55/M, output: $2.2/M. The model provider is the Sophnet platform. DeepSeek V3 Fast is a high-TPS, ultra-fast version of DeepSeek V3 0324, featuring full-precision (non-quantized) performance, enhanced code and math capabilities, and faster responses! DeepSeek V3 0324 is a powerful Mixture-of-Experts (MoE) model with a total parameter count of 671B, activating 37B parameters per token. It adopts Multi-Head Latent Attention (MLA) and the DeepSeekMoE architecture to achieve efficient inference and economical training costs. It innovatively implements a load balancing strategy without auxiliary loss and sets multi-token prediction training targets to enhance performance. The model is pre-trained on 14.8 trillion diverse, high-quality tokens and further optimized through supervised fine-tuning and reinforcement learning stages to fully realize its capabilities. Comprehensive evaluations show that DeepSeek V3 outperforms other open-source models and rivals leading closed-source models in performance. The entire training process only requires 2.788M H800 GPU hours and remains highly stable, with no irrecoverable loss spikes or rollbacks.
V3 Ultra-Fast Version,The current price is a limited-time 50% discount and will return to the original price on July 31st. The original price is: input: $0.55/M, output: $2.2/M. The model provider is the Sophnet platform. DeepSeek V3 Fast is a high-TPS, ultra-fast version of DeepSeek V3 0324, featuring full-precision (non-quantized) performance, enhanced code and math capabilities, and faster responses! DeepSeek V3 0324 is a powerful Mixture-of-Experts (MoE) model with a total parameter count of 671B, activating 37B parameters per token. It adopts Multi-Head Latent Attention (MLA) and the DeepSeekMoE architecture to achieve efficient inference and economical training costs. It innovatively implements a load balancing strategy without auxiliary loss and sets multi-token prediction training targets to enhance performance. The model is pre-trained on 14.8 trillion diverse, high-quality tokens and further optimized through supervised fine-tuning and reinforcement learning stages to fully realize its capabilities. Comprehensive evaluations show that DeepSeek V3 outperforms other open-source models and rivals leading closed-source models in performance. The entire training process only requires 2.788M H800 GPU hours and remains highly stable, with no irrecoverable loss spikes or rollbacks.
DeepSeek-V3-Fast has a 32,000 token context window.
On AIHubMix, DeepSeek-V3-Fast costs $0.56 per million input tokens and $2.24 per million output tokens.
DeepSeek-V3-Fast accepts text input.
DeepSeek-V3-Fast supports tool calling, function calling and structured outputs. Per-protocol parameter support is listed in the capability table on this page.
DeepSeek-V3-Fast is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to DeepSeek-V3-Fast — no other code changes needed.
DeepSeek-V3-Fast is developed by DeepSeek. AIHubMix aggregates it alongside models from other providers behind one API and one bill.
DeepSeek-V4-Flash official API release. Agent capabilities have been greatly enhanced…
DeepSeek-V4 features an ultra-long context of one million characters and achieves leading…
DeepSeek-V4 features an ultra-long context of one million characters and achieves leading…
DeepSeek-V3.2 is an efficient large language model equipped with DeepSeek Sparse…
DeepSeek-V3.2 is an efficient large language model equipped with DeepSeek Sparse…
DeepSeek-V3.1 non-thinking mode has now been updated to the DeepSeek-V3.1-Terminus…
Use DeepSeek-V3-Fast via the AIHubMix unified API — one interface for every major LLM.