For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...
Different carriers host the same model. Reroute routes each request to the best one on price and uptime, and fails over to the next if a carrier goes down.
| Provider | Context | Input /M | Output /M | Cache read /M | Latency | Throughput | Uptime |
|---|---|---|---|---|---|---|---|
| 1M | $0.10 | $0.40 | $0.03 | — | — | 100.00% | |
| 1M | $0.10 | $0.40 | $0.025 | — | — | 100.00% | |
| 1M | $0.11 | $0.44 | $0.033 | — | — | — |