The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall...
Different carriers host the same model. Reroute routes each request to the best one on price and uptime, and fails over to the next if a carrier goes down.
| Provider | Context | Input /M | Output /M | Cache read /M | Latency | Throughput | Uptime |
|---|---|---|---|---|---|---|---|
| 262K | $0.08 | $0.75 | $0.04 | — | — | 100.00% | |
| 262K | $0.14 | $1 | $0.05 | — | — | 100.00% | |
| 262K | $0.15 | $1 | $0.05 | — | — | 100.00% | |
| 262K | $0.163 | $1.30 | — | — | — | 100.00% | |
| 262K | $0.225 | $1.80 | $0.225 | — | — | — | |
| 262K | $0.24 | $1.80 | $0.15 | — | — | — | |
| 256K | $0.313 | $1.25 | $0.156 | — | — | — |