reroute

Z.ai: GLM 5.3 Flash

z-ai/glm-5.3-flash

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

Modalities
In / Out Price
$0.15 / $0.50 per 1M
Context
1M
Released
Aug 26, 2026
Knowledge Cutoff
—

Providers

Different carriers host the same model. Reroute routes each request to the best one on price and uptime, and fails over to the next if a carrier goes down.

ProviderContextInput /MOutput /MCache read /MLatencyThroughputUptime
OpenInferencefp41M$0.032$2.8$0.02——100.00%
Relace1M$0.04$0.50$0.025——100.00%
Sail Researchfp41M$0.045$0.60$0.0285——100.00%
Sail Researchfp41M$0.045$0.60$0.0285——100.00%
Reka262K$0.06$1$0.04——100.00%
InferenceNetfp41M$0.065$0.50$0.042——100.00%
Wafer1M$0.07$0.30$0.03——100.00%
DeepInfrafp41M$0.075$0.25$0.015——100.00%
Novitafp81M$0.084$0.28$0.0168——100.00%
StreamLakefp81M$0.087$0.29$0.0174——100.00%
GMICloudfp81M$0.09$0.30$0.018——100.00%
Decartfp41M$0.0945$0.315$0.0189——100.00%
DekaLLM1M$0.10$1$0.04——100.00%
Near AIfp81M$0.105$0.35$0.0245——100.00%
Phalanvfp41M$0.112$0.375$0.0225——100.00%
Modalnvfp41M$0.15$0.50$0.03——100.00%
BaseTenfp81M$0.15$0.50$0.03——90.00%
Crusoefp41M$0.15$0.50$0.03——100.00%
CoreWeavenvfp41M$0.15$0.50$0.05——100.00%
AtlasCloudfp81M$0.15$0.50$0.03——100.00%
Fireworks1M$0.15$0.50$0.03——100.00%
Friendli1M$0.15$0.50$0.03——100.00%
SiliconFlowfp81M$0.15$0.50$0.03——100.00%
DigitalOcean1M$0.15$0.50$0.03——100.00%
Together1M$0.15$0.50$0.03——100.00%
Parasailfp41M$0.15$0.50$0.03——100.00%
BaseTenfp81M$0.15$0.50$0.03——86.67%
Venice1M$0.15$0.50$0.03——100.00%
Z.AIfp81M$0.15$0.50$0.03——100.00%
Morphfp81M$0.168$0.70$0.034——76.67%
Parasailfp41M$0.188$0.625$0.0375——100.00%
Inceptronfp81M$0.225$0.60$0.09——100.00%
Fireworks1M$0.225$0.75$0.045——100.00%