The Chinese AI company said it had already launched the model globally in stealth mode for a week.

Z.ai's GLM-5.3-Flash runs natively on Chinese AI chips, hits 63.4 on DeepSWE, and offers API access at $0.15 per million input tokens.

GLM-5.3-Flash’s lower API service rates add fresh pressure on the country’s model developers as they race to lower inference costs to gain market share