GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and...
Different carriers host the same model. Reroute routes each request to the best one on price and uptime, and fails over to the next if a carrier goes down.
| Provider | Context | Input /M | Output /M | Cache read /M | Latency | Throughput | Uptime |
|---|---|---|---|---|---|---|---|
| 1M | $2 | $8 | $0.50 | — | — | 100.00% | |
| 1M | $2 | $8 | $0.50 | — | — | 100.00% | |
| 1M | $2.20 | $8.80 | $0.55 | — | — | — |