Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...
Different carriers host the same model. Reroute routes each request to the best one on price and uptime, and fails over to the next if a carrier goes down.
| Provider | Context | Input /M | Output /M | Cache read /M | Latency | Throughput | Uptime |
|---|---|---|---|---|---|---|---|
| 131K | $0.042 | $0.22 | $0.021 | — | — | 100.00% | |
| 262K | $0.0517 | $0.225 | $0.0217 | — | — | 100.00% | |
| 262K | $0.06 | $0.33 | $0.04 | — | — | 100.00% | |
| 262K | $0.07 | $0.34 | — | — | — | 100.00% | |
| 262K | $0.0765 | $0.255 | $0.0425 | — | — | 100.00% | |
| 256K | $0.08 | $0.32 | $0.032 | — | — | 100.00% | |
| 262K | $0.10 | $0.30 | $0.05 | — | — | 100.00% | |
| 256K | $0.10 | $0.30 | $0.05 | — | — | 100.00% | |
| 256K | $0.13 | $0.40 | $0.05 | — | — | 100.00% | |
| 262K | $0.13 | $0.40 | $0.05 | — | — | 100.00% | |
| 262K | $0.13 | $0.40 | — | — | — | 100.00% | |
| 262K | $0.14 | $0.40 | $0.05 | — | — | 100.00% | |
| 262K | $0.15 | $0.60 | — | — | — | 100.00% |