The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...
Different carriers host the same model. Reroute routes each request to the best one on price and uptime, and fails over to the next if a carrier goes down.
| Provider | Context | Input /M | Output /M | Cache read /M | Latency | Throughput | Uptime |
|---|---|---|---|---|---|---|---|
| 262K | $0.39 | $2.34 | — | — | — | 100.00% | |
| 262K | $0.45 | $3 | $0.22 | — | — | 100.00% | |
| 262K | $0.50 | $3.60 | $0.30 | — | — | 100.00% | |
| 131K | $0.55 | $3.50 | $0.11 | — | — | 62.50% | |
| 262K | $0.55 | $3.50 | $0.225 | — | — | 100.00% | |
| 262K | $0.55 | $3.50 | $0.55 | — | — | 100.00% | |
| 256K | $0.60 | $3.60 | $0.12 | — | — | 100.00% | |
| 262K | $0.60 | $3.60 | — | — | — | 100.00% | |
| 262K | $0.60 | $3.60 | — | — | — | 100.00% | |
| 128K | $0.75 | $4.50 | — | — | — | 100.00% |