A 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA. The model is multilingual, supporting English, French, German, Spanish, Italian, Portuguese, Chinese, Japanese,...
Different carriers host the same model. Reroute routes each request to the best one on price and uptime, and fails over to the next if a carrier goes down.
| Provider | Context | Input /M | Output /M | Cache read /M | Latency | Throughput | Uptime |
|---|---|---|---|---|---|---|---|
| 131K | $0.018 | $0.03 | — | — | — | 100.00% | |
| 128K | $0.023 | $0.03 | $0.015 | — | — | 100.00% | |
| 131K | $0.029 | $0.03 | — | — | — | 100.00% | |
| 131K | $0.03 | $0.03 | — | — | — | 100.00% | |
| 131K | $0.15 | $0.15 | $0.015 | — | — | — |