Z.ai released GLM-5.3-Flash, a natively multimodal mixture-of-experts model with 320 billion total parameters and 18 billion active parameters, published under the MIT license. The model supports a one-million-token context window and, according to Z.ai, performs comparably to Claude Opus 4.8 on coding benchmarks at roughly one-tenth the price while outperforming its predecessor GLM-5.2. GLM-5.3-Flash is now available to GLM Coding Plan users with triple the quota of GLM-5.3, with model weights published on Hugging Face for deployment via SGLang, vLLM, and TokenSpeed.
