If you looked the other way, you may have missed the stealth preview model that was free on OpenRouter.AI and OpenCode.AI, which was branded "ox-alpha", that was making a bit of a splash. Everyone had figured out that it was a new Z.ai multi-modal, but no one could figure out how they were giving away 5,835,184,092,873 prompt tokens and 94,402,126,602 completion tokens in a single day!
The answer when it launched as GLM-5.3-Flash is that its an very efficient Flash open-weights model that you can get hosted by a US hardware by San Francisco firm OpenCode.AI for an input cost of $0.07 / 1M and output cost of $0.25 / 1M.
Yet they have hidden one dark truth that I can now reveal. The ultimate proof that this model is different. I have hooked it into my personal fork of the rust based apache2 codex, which I prefer, and I asked it to tell me a joke, and it bombed:
In the era of all the models memorising a snake game and "tell me a joke" it made one up on the spot and bombed 💖💚💛❤️
I have promoted all the models "tell me a joke" on all the harnesses, and you get the same set of responses. They have memorised that. They do not make up stuff. They stick with some very common jokes. I have given up asking for "tell me another" as they only have a limited set.












