DeepSeek-V3-Fast

by DeepSeek

V3 Ultra-Fast Version,The current price is a limited-time 50% discount and will return to the original price on July 31st. The original price is: input: $0.55/M, output: $2.2/M. The model provider is the Sophnet platform. DeepSeek V3 Fast is a high-TPS, ultra-fast version of DeepSeek V3 0324, featuring full-precision (non-quantized) performance, enhanced code and math capabilities, and faster responses! DeepSeek V3 0324 is a powerful Mixture-of-Experts (MoE) model with a total parameter count of 671B, activating 37B parameters per token. It adopts Multi-Head Latent Attention (MLA) and the DeepSeekMoE architecture to achieve efficient inference and economical training costs. It innovatively implements a load balancing strategy without auxiliary loss and sets multi-token prediction training targets to enhance performance. The model is pre-trained on 14.8 trillion diverse, high-quality tokens and further optimized through supervised fine-tuning and reinforcement learning stages to fully realize its capabilities. Comprehensive evaluations show that DeepSeek V3 outperforms other open-source models and rivals leading closed-source models in performance. The entire training process only requires 2.788M H800 GPU hours and remains highly stable, with no irrecoverable loss spikes or rollbacks.

API Pricing

Input$0.56 / 1M tokens
Output$2.24 / 1M tokens

Specifications

Context window32,000 tokens
Modalitiestext
Featurestool calling, function calling, structured outputs

FAQ

What is DeepSeek-V3-Fast?

V3 Ultra-Fast Version,The current price is a limited-time 50% discount and will return to the original price on July 31st. The original price is: input: $0.55/M, output: $2.2/M. The model provider is the Sophnet platform. DeepSeek V3 Fast is a high-TPS, ultra-fast version of DeepSeek V3 0324, featuring full-precision (non-quantized) performance, enhanced code and math capabilities, and faster responses! DeepSeek V3 0324 is a powerful Mixture-of-Experts (MoE) model with a total parameter count of 671B, activating 37B parameters per token. It adopts Multi-Head Latent Attention (MLA) and the DeepSeekMoE architecture to achieve efficient inference and economical training costs. It innovatively implements a load balancing strategy without auxiliary loss and sets multi-token prediction training targets to enhance performance. The model is pre-trained on 14.8 trillion diverse, high-quality tokens and further optimized through supervised fine-tuning and reinforcement learning stages to fully realize its capabilities. Comprehensive evaluations show that DeepSeek V3 outperforms other open-source models and rivals leading closed-source models in performance. The entire training process only requires 2.788M H800 GPU hours and remains highly stable, with no irrecoverable loss spikes or rollbacks.

What is the context length of DeepSeek-V3-Fast?

DeepSeek-V3-Fast has a 32,000 token context window.

How much does DeepSeek-V3-Fast cost?

On AIHubMix, DeepSeek-V3-Fast costs $0.56 per million input tokens and $2.24 per million output tokens.

What modalities does DeepSeek-V3-Fast support?

DeepSeek-V3-Fast accepts text input.

What features does DeepSeek-V3-Fast support?

DeepSeek-V3-Fast supports tool calling, function calling and structured outputs. Per-protocol parameter support is listed in the capability table on this page.

How do I call DeepSeek-V3-Fast via API?

DeepSeek-V3-Fast is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to DeepSeek-V3-Fast — no other code changes needed.

Who develops DeepSeek-V3-Fast?

DeepSeek-V3-Fast is developed by DeepSeek. AIHubMix aggregates it alongside models from other providers behind one API and one bill.

More from DeepSeek

deepseek-v4-flash-0731

by DeepSeek

DeepSeek-V4-Flash official API release. Agent capabilities have been greatly enhanced…

$0.1/1M in · $0.2/1M out
1,000,000 tokens context

deepseek-v4-flash

by DeepSeek

DeepSeek-V4 features an ultra-long context of one million characters and achieves leading…

$0.15/1M in · $0.31/1M out
1,000,000 tokens context

deepseek-v4-pro

by DeepSeek

DeepSeek-V4 features an ultra-long context of one million characters and achieves leading…

$0.46/1M in · $0.93/1M out
1,000,000 tokens context

deepseek-v3.2

by DeepSeek

DeepSeek-V3.2 is an efficient large language model equipped with DeepSeek Sparse…

$0.3/1M in · $0.45/1M out
128,000 tokens context

deepseek-v3.2-think

by DeepSeek

DeepSeek-V3.2 is an efficient large language model equipped with DeepSeek Sparse…

$0.3/1M in · $0.45/1M out
128,000 tokens context

DeepSeek-V3.1-Terminus

by DeepSeek

DeepSeek-V3.1 non-thinking mode has now been updated to the DeepSeek-V3.1-Terminus…

$0.56/1M in · $1.68/1M out
160,000 tokens context

Use DeepSeek-V3-Fast via the AIHubMix unified API — one interface for every major LLM.