
GLM-5.3-Flash: Let Your Agent See the UI It Built
Summary
Build a render-screenshot-critique loop with GLM-5.3-Flash native vision, for pennies a run.
For six days an unbranded model called ox-alpha sat on OpenRouter and OpenCode charging nothing, accepting text, images and video, and swallowing a million tokens of context. People burned through it. Z.ai says it became the most-used model of the week on both services. On August 26 the mask came off: ox-alpha is GLM-5.3-Flash, a 320B-parameter mixture-of-experts model with 18B active per token, released under the MIT license, and priced at $0.075 per million input tokens during the launch discount.
The benchmark chatter will fade in a week. The part worth learning is quieter and more useful: Z.ai trained this model to look at rendered output and revise its own work from what it sees. Not "describe this photo" vision. Screenshot-of-your-own-code vision. That turns a pattern most people talk about but rarely wire up into something you can actually afford to run in a loop.
Keep reading — it's free
Enter your email to keep reading — plus the best of AI & tech, daily. Free, forever.
Already a member? Sign in
Comments
Be the first to comment