The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of...
Different carriers host the same model. Reroute routes each request to the best one on price and uptime, and fails over to the next if a carrier goes down.
| Provider | Context | Input /M | Output /M | Cache read /M | Latency | Throughput | Uptime |
|---|---|---|---|---|---|---|---|
| 262K | $0.195 | $1.56 | — | — | — | 100.00% | |
| 262K | $0.25 | $2 | — | — | — | 100.00% | |
| 262K | $0.26 | $2.60 | — | — | — | 100.00% | |
| 262K | $0.27 | $2.16 | $0.27 | — | — | 100.00% | |
| 262K | $0.30 | $2.40 | $0.03 | — | — | 100.00% | |
| 262K | $0.30 | $2.40 | — | — | — | 100.00% |