Alibaba Scales Qwen AI Models to 2.4 Trillion Parameters
Alibaba Cloud has scaled its Qwen AI family from a modest 7-billion-parameter model to a massive 2.4-trillion-parameter release, reshaping the open-source landscape for global developers.

Alibaba Cloud has expanded its Qwen artificial intelligence family to a massive scale, culminating in the release of Qwen3.8-2.4T-A95B. This open-source base model boasts 2.4 trillion parameters with 95 billion active parameters, though it requires a commercial license for providers earning over 50 million dollars. The rapid evolution of the Qwen ecosystem, which began in April 2023 with the enterprise beta of Tongyi Qianwen, has now delivered over 400 open models, surpassing 1 billion downloads and inspiring more than 200,000 derivative models on Hugging Face.
The journey to this milestone involved a relentless release cadence. In 2024, Alibaba shipped Qwen1.5 and Qwen2, introducing mixture-of-experts architectures like the Qwen2-57B-A14B and expanding context windows to 128K tokens. By September 2024, Qwen2.5 arrived with pretraining data scaled to 18 trillion tokens. The subsequent Qwen3 family, launched in April 2025, introduced hybrid thinking capabilities trained on 36 trillion tokens across 119 languages. Alibaba also experimented with novel architectures, such as the Qwen3-Next-80B-A3B, which combined Gated DeltaNet linear attention with gated attention to slash training costs to 10 percent of Qwen3-32B while boosting inference throughput tenfold.
Alibaba has consistently targeted frontier performance. The QwQ-32B reasoning model achieved comparable performance to the 671-billion-parameter DeepSeek-R1 while requiring only 24 gigabytes of VRAM compared to DeepSeek's 1,500 gigabytes. In early 2026, the closed Qwen3-Max-Thinking model proved comparable to top-tier models like GPT-5.2-Thinking, Claude Opus 4.5, and Gemini 3 Pro across 19 benchmarks. While early models were entirely open, Alibaba later adopted a two-tier strategy. High-end models like Qwen3.7-Max and Qwen3.7-Plus, priced at 0.4 dollars and 1.6 dollars per million tokens, remain proprietary, while mid-sized variants like Qwen3.8-27B remain under permissive Apache 2.0 licenses.
For AI practitioners, the Qwen trajectory provides a highly flexible, cost-effective alternative to Western frontier models. Developers can deploy highly efficient local models like the Qwen3.5-35B-A3B, which outperforms older 235-billion-parameter flagships, or leverage multimodal tools like the 7-billion-parameter Qwen-Image-2.1 and the speech-enabled Qwen2.5-Omni-7B. With Qwen 4 currently in training and future iterations projected to reach 5 to 10 trillion parameters, developers have a reliable open foundation that is already being integrated into major platforms like Apple Intelligence in China.
This is our own summary of reporting by MarkTechPost



