DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context...
Different carriers host the same model. Reroute routes each request to the best one on price and uptime, and fails over to the next if a carrier goes down.
| Provider | Context | Input /M | Output /M | Cache read /M | Latency | Throughput | Uptime |
|---|---|---|---|---|---|---|---|
| 164K | $0.25 | $0.95 | $0.13 | — | — | 100.00% | |
| 164K | $0.27 | $1 | — | — | — | 100.00% | |
| 161K | $0.55 | $1.65 | $0.55 | — | — | 100.00% | |
| 131K | $0.60 | $1.70 | — | — | — | 100.00% | |
| 131K | $0.65 | $1.50 | — | — | — | 100.00% |