by Qwen
The Qwen3.5 native vision-language Flash series models are designed with a hybrid architecture that integrates linear attention mechanisms and sparse mixture-of-experts models, achieving higher inference efficiency. Compared with the 3 series, the models deliver leapfrog improvements in both pure-text and multimodal performance; they respond quickly and combine inference speed with high performance.
The Qwen3.5 native vision-language Flash series models are designed with a hybrid architecture that integrates linear attention mechanisms and sparse mixture-of-experts models, achieving higher inference efficiency. Compared with the 3 series, the models deliver leapfrog improvements in both pure-text and multimodal performance; they respond quickly and combine inference speed with high performance.
qwen3.5-flash has a 991,000 token context window.
On AIHubMix, qwen3.5-flash costs $0.028 per million input tokens and $0.282 per million output tokens. Cached input reads are billed at $0.0028 per million tokens.
qwen3.5-flash accepts text, image and video input.
qwen3.5-flash supports tool calling, function calling, structured outputs, web, long context and thinking. Per-protocol parameter support is listed in the capability table on this page.
qwen3.5-flash is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to qwen3.5-flash — no other code changes needed.
qwen3.5-flash is developed by Qwen. AIHubMix aggregates it alongside models from other providers behind one API and one bill.
Qwen Image 3.0(qwen-image-3.0) is an image generation and editing model developed by…
Qwen Image 3.0 Pro (qwen-image-3.0-pro) is Alibaba Cloud Qwen’s flagship image generation…
Qwen 3.8 Max(qwen3.8-max) is Alibaba Cloud’s flagship native vision-language model, built…
Qwen 3.8 Max Preview(Qwen3.8-Max-Preview) is the latest-generation foundation model in…
qwen-audio-3.0-tts-flash is a high-performance speech synthesis large model optimized for…
qwen-audio-3.0-tts-plus is a high-performance speech synthesis large model designed for…
Use qwen3.5-flash via the AIHubMix unified API — one interface for every major LLM.