reroute

Z.ai: GLM 5.3 FlashX

z-ai/glm-5.3-flashx

GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...

Modalities
In / Out Price
$0.37 / $1.25 per 1M
Context
1M
Released
Sep 18, 2026
Knowledge Cutoff
—

Pricing

What Reroute charges per unit. No markup on inference — you pay the carrier's price, metered from the carrier's own usage numbers.

Input tokens
$0.37 / M
Output tokens
$1.25 / M
Cache read
$0.09 / M

By provider

Each carrier sets its own price. The router weighs these against uptime when it picks one.

ProviderInput /MOutput /MCache read /MCache write /MDiscount
Z.AI$0.37$1.25$0.09——