GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...
Top apps by tokens routed through Reroute in the last 30 days. Apps appear here when they send an X-Title header.
Once apps call this model with an X-Title header, they're ranked here by tokens.