Models

DeepSeek Launches V4 Pro 0813 with 5x Coding Leap

DeepSeek has quietly launched its flagship DeepSeek V4 Pro 0813 model, delivering a massive fivefold boost in agentic coding performance alongside a significant price increase.

AlphaSignal4 days agoModels
Image: AlphaSignal

DeepSeek has quietly transitioned its flagship reasoning model, DeepSeek V4 Pro 0813, to general availability by updating its API endpoint without an official announcement. The model features 1.6 trillion total parameters with roughly 49 billion active parameters per token, and it has been pre-trained on more than 32 trillion tokens using the Muon optimizer. It retains the preview version's architecture, including a DSpark speculative decoding module and a hybrid attention system combining Compressed Sparse Attention and Heavily Compressed Attention. This setup supports a 1,048,576-token context window and up to 384,000 output tokens, requiring only 27 percent of V3.2's per-token inference FLOPs and 10 percent of its KV cache at maximum context.

The model's weights have been released on Hugging Face under the MIT license. Despite having no architectural changes, V4 Pro 0813 achieves massive performance gains solely through post-training improvements, which utilize a two-stage pipeline of expert cultivation and on-policy distillation. On the Artificial Analysis Intelligence Index, the model scored 53, ranking third among all models. Its agentic coding capabilities saw the most dramatic growth, with its DeepSWE score exploding from 12.8 to 62.7, and its Terminal Bench 2.1 score rising from 72.1 to 87.9.

For practitioners, these capabilities come with a steep cost adjustment. DeepSeek has raised prices by 264 percent, setting rates at $1.32 per million input tokens and $3.96 per million output tokens, while cache-hit pricing has surged 12-fold. Furthermore, because the model lacks vision support and bills thinking tokens at output rates, actual agentic workflows may prove costlier than expected. Developers must also weigh whether the model justifies its premium over the cheaper V4 Flash 0731, which scores only one point lower on the Intelligence Index despite having 3.8 times fewer active parameters.

This is our own summary of reporting by AlphaSignal

More in Models