reroute

Z.ai: GLM 5.3 Flash

z-ai/glm-5.3-flash

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

Modalities
In / Out Price
$0.15 / $0.50 per 1M
Context
1M
Released
Aug 26, 2026
Knowledge Cutoff
—

Pricing

What Reroute charges per unit. No markup on inference — you pay the carrier's price, metered from the carrier's own usage numbers.

Input tokens
$0.15 / M
Output tokens
$0.50 / M
Cache read
$0.03 / M

By provider

Each carrier sets its own price. The router weighs these against uptime when it picks one.

ProviderInput /MOutput /MCache read /MCache write /MDiscount
OpenInference$0.032$1.84$0.02——
Relace$0.04$0.50$0.025——
Sail Research$0.045$0.60$0.0285——
Sail Research$0.045$0.60$0.0285——
InferenceNet$0.06$0.50$0.04——
Reka$0.06$2.40$0.04——
Wafer$0.07$0.30$0.03——
DeepInfra$0.075$0.25$0.015—50%
Novita$0.084$0.28$0.0168—44%
StreamLake$0.087$0.29$0.0174—42%
GMICloud$0.09$0.30$0.018—40%
Decart$0.0915$0.305$0.0183—39%
DekaLLM$0.10$1$0.04——
Near AI$0.105$0.35$0.0245—30%
Morph$0.11$0.70$0.034——
Phala$0.112$0.375$0.0225—25%
Modal$0.15$0.50$0.03——
BaseTen$0.15$0.50$0.03——
Crusoe$0.15$0.50$0.03——
CoreWeave$0.15$0.50$0.05——
AtlasCloud$0.15$0.50$0.03——
Fireworks$0.15$0.50$0.03——
Friendli$0.15$0.50$0.03——
SiliconFlow$0.15$0.50$0.03——
DigitalOcean$0.15$0.50$0.03——
Together$0.15$0.50$0.03——
Parasail$0.15$0.50$0.03——
BaseTen$0.15$0.50$0.03——
Venice$0.15$0.50$0.03——
Z.AI$0.15$0.50$0.03——
Parasail$0.188$0.625$0.0375——
Inceptron$0.225$0.60$0.09——
Fireworks$0.225$0.75$0.045——