Z.ai releases GLM-5.3-Flash, an open-source model with 320 billion parameters that lands just three points behind the larger GLM-5.3 on Artificial Analysis's Intelligence Index, at a seventh of the cost. What's notable is that all of the inference traffic ran on Chinese AI chips instead of Nvidia hardware.

Z.ai's GLM-5.3-Flash runs natively on Chinese AI chips, hits 63.4 on DeepSWE, and offers API access at $0.15 per million input tokens.

For the past week, developers have been puzzling over a model called Ox Alpha. It appeared on...