Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
Different carriers host the same model. Reroute routes each request to the best one on price and uptime, and fails over to the next if a carrier goes down.
| Provider | Context | Input /M | Output /M | Cache read /M | Latency | Throughput | Uptime |
|---|---|---|---|---|---|---|---|
| 262K | $0.09 | $0.34 | $0.05 | — | — | 100.00% | |
| 262K | $0.10 | $0.34 | $0.10 | — | — | 100.00% | |
| 256K | $0.12 | $0.36 | $0.09 | — | — | 100.00% | |
| 131K | $0.12 | $0.37 | $0.012 | — | — | 80.00% | |
| 262K | $0.14 | $0.40 | $0.14 | — | — | 96.67% | |
| 262K | $0.14 | $0.40 | — | — | — | 100.00% | |
| 262K | $0.14 | $0.40 | — | — | — | 63.33% | |
| 262K | $0.15 | $0.40 | $0.06 | — | — | 100.00% | |
| 262K | $0.20 | $0.40 | — | — | — | 100.00% | |
| 262K | $0.361 | $1.09 | $0.181 | — | — | 100.00% | |
| 262K | $0.38 | $1.15 | — | — | — | 96.55% | |
| 262K | $0.75 | $1 | $0.20 | — | — | 100.00% | |
| 262K | $0.75 | $1 | $0.25 | — | — | 96.67% |