Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, and video. With 32 billion parameters, it combines deep visual perception with advanced text...
Tokens processed for this model through Reroute, per day, over the last 30 days.
Daily token volume shows up here as soon as the first request for this model completes.