The company says it handled all online traffic for the model on 100,000 domestically made chips, a claim CNBC could not independently verify

Z.ai's GLM-5.3-Flash runs natively on Chinese AI chips, hits 63.4 on DeepSWE, and offers API access at $0.15 per million input tokens.

GLM-5.3-Flash’s lower API service rates add fresh pressure on the country’s model developers as they race to lower inference costs to gain market share

Z.ai open-sources ‘Ox Alpha’ model as GLM-5.3-Flash - SiliconANGLE