For six days in August, one of the busiest coding models on the internet had no name. Ox Alpha appeared free on OpenRouter and OpenCode on August 20, 2026, credited only to “a third-party provider who has chosen to remain anonymous.” It accepted text, images, and video, carried a 1M-token context window, and let developers tune its reasoning effort. On August 26, Z.ai ended the guessing: Ox Alpha was GLM-5.3-Flash, a 320B-parameter model with just 18B active, and what Z.ai calls the first natively multimodal model in its GLM-5 series. Z.ai says it served every request on 100,000 Chinese-made chips.
The draw was price. Z.ai says GLM-5.3-Flash scores 57 on Artificial Analysis’s Intelligence Index at about $0.045 per task, a level it claims once cost roughly ten times more, and it lands close to Claude Opus 4.8 on agentic coding. On OpenRouter it lists at $0.075 and $0.25 per million input and output tokens, and the weights now sit on Hugging Face under an MIT licence. For a week, developers pointed real repositories at an endpoint with no named owner and no model card.
A model nobody could name topped OpenRouter’s usage chart that week, and Z.ai collected real agent traffic plus a name-making buzz before it published anything. Developers got a cheap, fast coding model. What they did not get was a name they could check.
Sources:
- GLM-5.3-Flash: Frontier Intelligence, Flash Cost (Z.ai)
- GLM 5.3 Flash model page and pricing (OpenRouter)
- GLM-5.3-Flash weights on Hugging Face
- Z.ai shares surge 8% after releasing new AI model running only on Chinese chips (CNBC)
Disclaimer: For information only. Accuracy or completeness not guaranteed. Illegal use prohibited. Not professional advice or solicitation. Read more: /terms-of-service
Reuse
Citation
@misc{kabui2026,
author = {{Kabui, Charles}},
title = {The {Anonymous} {Coding} {Model} {Developers} {Used} for a
{Week} {Was} {Z.ai’s} {GLM-5.3} {Flash}},
date = {2026-09-14},
url = {https://toknow.ai/posts/ox-alpha-glm-5-3-flash-stealth-model-reveal/},
langid = {en-GB}
}
