Zhipu AI releases GLM-5.3-Flash model, powered by 100,000 domestically made chips

Booth of Zhipu at the exhibition hall of the 2025 World Artificial Intelligence Conference on July 28, 2025 Photo: VCGChinese AI startup Zhipu AI said that its newly launched GLM-5.3-Flash, the first native multimodal model in its GLM-5 series, is powered entirely by 100,000 domestically made chips, underscoring China's push to build large-scale AI infrastructure using homegrown semiconductors.The company said that GLM-5.3-Flash has been officially released as an open-source model under the MIT license. It supports a 1-million-token context window and is built with a 300-billion-parameter architecture. Zhipu AI said the model can be deployed on all domestically-produced AI chips, according to information shared with the Global Times on Thursday.In benchmark testing, the model scored 57 points on the Artificial Analysis Intelligence Index, on par with Anthropic's Claude Opus 4.8. Zhipu AI also said GLM-5.3-Flash is priced at $0.15 per million input tokens and $0.50 per million output tokens, around one-tenth of GLM-5.3 and roughly one-fortieth of Claude Opus 4.8.The Zhipu team said this was mainly due to multiple architectural upgrades implemented simultaneously in GLM-5.3-Flash. By adopting a hybrid architecture of sparse attention and linear attention, the team was able to significantly reduce the cost of long-context services while maintaining precise long-context capabilities.Before its official debut, the GLM-5.3-Flash model was tested anonymously as Ox Alpha — known as "Niu Lai" in Chinese developer circles — on overseas platforms OpenRouter and OpenCode, where it recorded over 60 trillion tokens in usage and swept to the top of online usage charts over the week, with all request traffic backed by domestic-chip computing power, financial media outlet Caixin reported.This shows that Chinese models can attract developers through genuine cost reduction and efficiency gains at the architectural level, even completely stripped of identity labels. Developers are voting for this model with real paid usage volume, because when using it, no one knows who the vendor is, Tian Feng, former dean of SenseTime's Intelligence Industry Research Institute, told the Global Times on Thursday.Zhipu AI said the company has, for the first time, tried using a large chip cluster to provide services in GLM-5.3-Flash, with these chips connected via its self-developed high-bandwidth interconnect network.In its technical documentation shared with the Global Times, Zhipu said its inference service ran on a cluster of more than 100,000 domestic chips and that "hardware efficiency and per-token cost have reached a level comparable to that of mainstream Nvidia GPUs.""This proves that domestic chips can fully and efficiently support frontier-model inference in large-scale scenarios," the company said.Semiconductor research firm SemiAnalysis commented on X that with every request carried by domestic silicon at Nvidia-comparable efficiency, "the CUDA moat is once again being tested," following OpenAI's announcement of its in-house Jalapeño inference chip. Chinese media outlet LatePost reported the chips may come from Huawei, Moore Threads and Hygon, though Zhipu AI would not confirm the suppliers or models.The latest development highlights growing competition in the global AI model market, as Chinese developers increasingly seek to prove both technical capability and large-scale deployment capacity using homegrown semiconductor infrastructure, according to industry analysts.As of press time on Thursday, Zhipu AI's Hong Kong-listed shares rose more than 6.3 percent in morning trading.Alibaba announced to release its new multimodal model Qwen3.8-Flash that delivers stronger coding and office-task performance while cutting training costs. Qwen3.8-Flash supports a default context window of 262,144 tokens, expandable to 1 million tokens, enabling it to handle large files, lengthy conversations and extensive research materials, according to the company.According to Tian, the latest releases suggest that China's large language models are moving beyond headline performance and into a new phase of cost-efficient, large-scale deployment. The combination of native multimodal capability, a one-million-token context window, and inference on domestically produced chips indicates that Chinese AI developers are able to optimize both model architecture and infrastructure at the same time, Tian said, adding that if the efficiency gains continue, they could accelerate broader adoption among enterprises and developers, while strengthening China's long-term AI ecosystem resilience.Domestic large AI models reached a weekly usage of 36.84 trillion tokens between August 10 and 16, surpassing the US for 16 consecutive weeks to rank first globally, according to Shanghai-based Securities Times.