Z.ai has released GLM-5.3-Flash, an MIT-licensed multimodal model with 320 billion total parameters and 18 billion active for each request. The model weights are available on Hugging Face, and the company advertises a context window of one million tokens.
Artificial Analysis scored the model at 57 on its Intelligence Index at maximum reasoning effort, three points behind the larger GLM-5.3. Its measured cost was $0.09 per task, compared with $0.68 for the larger model. Those figures describe one benchmark and pricing setup, so they should not be treated as a guarantee for every workload.
The infrastructure is also notable: Z.ai says all inference traffic ran on Chinese AI chips rather than Nvidia hardware, using its own serving software. Independent buyers still need to test quality, latency and hardware availability for their applications, and the model’s benchmark position includes reported weaknesses. The release nevertheless combines open weights, sparse activation and a non-Nvidia deployment path at a comparatively low measured cost.