Hardware

Z.ai Runs GLM-5.3 Flash Model on 100,000 Chinese Chips

Chinese startup Z.ai used 100,000 domestic chips to power its GLM-5.3 Flash model, proving that Chinese hardware can successfully handle high-volume AI inference workloads.

AI Business2 days agoHardware
Image: AI Business

Z.ai officially launched its GLM-5.3 Flash model under the MIT open source license on August 26, following an August 20 preview under the code name Ox Alpha. The open-weight, multimodal model features 320 billion parameters and is optimized for code synthesis, vision-driven agentic tasks, and long-context processing. To handle online queries for the model, Z.ai deployed 100,000 Chinese-made chips, demonstrating that domestic hardware can support large-scale inference and reduce reliance on Nvidia.

The model offers a highly competitive pricing structure. Until September 9, GLM-5.3 Flash costs $0.075 per million input tokens and $0.25 per million output tokens, rising to $0.15 and $0.50 respectively after that date. This is significantly cheaper than OpenAI's GPT-5.6 Luna, which costs $0.20 per million input tokens and $2 per million output tokens, and Anthropic's Claude Opus 5, priced at $5 per million input tokens and $5 per million output tokens. The model's strong performance in blind developer tests on OpenRouter, where it became the most downloaded model, underscores its practical viability.

This hardware shift comes amid tightening regulatory environments. Beijing banned foreign AI chips from state-funded data centers in November 2025 and ordered major tech firms to stop purchasing Nvidia hardware in September 2025. While Nvidia maintains a strong lead in training frontier models, analysts note that the gap is narrowing for inference. By optimizing software and model architectures, Chinese firms can achieve high performance without needing exact hardware parity with Western chipmakers.

For enterprise practitioners, the success of GLM-5.3 Flash highlights the importance of designing flexible AI systems. Rather than relying on a single model provider, developers should build architectures that route workloads dynamically based on cost, quality, and capability. However, Western enterprises must still weigh these cost efficiencies against ongoing data security, procurement, and regulatory concerns associated with Chinese-hosted AI services.

This is our own summary of reporting by AI Business

More in Hardware