Last verified: August 4, 2026. “Qwen Flash” is now a model tier rather than one fixed downloadable checkpoint. The name can refer to a current Qwen3.7 or Qwen3.6 Flash service, or to the older qwen-flash API family. Always record the provider, region, and exact model ID.
For a new QwenCloud integration, the current recommendation is qwen3.7-flash. Alibaba Cloud Model Studio documents qwen3.6-flash as its current broadly regional Flash model. The older qwen-flash ID remains listed in some Model Studio regions under Legacy Qwen.
Current Qwen Flash model IDs
| Provider | Model ID | Input to output | Documented service limits | Status |
|---|---|---|---|---|
| QwenCloud | qwen3.7-flash | Text, image, video to text | 1M context. Marketplace card: 131K max output; developer matrices: 64K, checked August 4, 2026. | Current mainline |
| QwenCloud | qwen3.7-flash-2026-07-15 | Text, image, video to text | 1M context and 64K max output in the developer matrices | Dated snapshot |
| Alibaba Cloud Model Studio | qwen3.6-flash | Text, image, video to text | 1M hosted context; provider-layer output label depends on the cited route | Current regional model; Qwen’s launch connects this API ID to the open Qwen3.6-35B-A3B release |
| Alibaba Cloud Model Studio, US scope | qwen3.6-flash-us | Use the supported protocol for the US route | Use the matching US model page and account limits | Documented regional ID; do not construct or move it across scopes |
| Alibaba Cloud Model Studio | qwen-flash | Text to text | 1M in the cited legacy matrix | Legacy mainline; separate from current Qwen3.7 and Qwen3.6 Flash IDs |
Rolling and dated IDs are separate strings. Do not assume that a rolling ID always stays on one snapshot, or copy an ID, context value, output ceiling, price, or lifecycle notice across providers.
Context, output, and modalities
The current Qwen3.7 and Qwen3.6 Flash services accept text, image, and video and return text. Their cited service documentation places them in a 1M-context category, but the usable input must leave room for system instructions, media tokens, reasoning, tools, and the answer.
Official Qwen3.7 output-limit discrepancy, checked August 4, 2026: QwenCloud’s model-specific Marketplace cards display a 131K maximum output for the rolling qwen3.7-max, qwen3.7-plus, and qwen3.7-flash routes. Its text and vision developer matrices display 64K for the corresponding Qwen3.7 services. Neither page showed a visible revision date.
Treat the number as documentation-path and endpoint specific. Use a conservative configured ceiling until the exact model, protocol, deployment scope, and account accept a tested value. This page does not claim an independently tested maximum.
The legacy qwen-flash family is documented under text generation. Do not send images or video to that ID merely because newer Flash models are multimodal. Also avoid confusing it with Qwen-VL, Coder, Omni, TTS, or audio Flash models.
Thinking mode, tools, and structured output
Tool support is provider-specific. Alibaba Cloud Model Studio lists legacy qwen-flash with hybrid thinking, Function Calling, built-in tools, and structured output. QwenCloud international lists legacy qwen-flash with a 1M context window, 32K maximum output, Function Calling, and structured output, but does not list built-in tools. Current Qwen3.7 and Qwen3.6 Flash services have their own route-specific feature matrices. Record the exact provider, model ID, endpoint, and thinking mode, and validate tool arguments and structured output in application code.
Do not publish hidden reasoning. For an evaluation, retain the user prompt, final answer, request settings, usage fields, and error/status data that the provider permits you to store.
Is Qwen Flash open source or self-hosted?
Hosted service IDs and repository IDs are not interchangeable. Qwen’s official April 15, 2026 launch says the open-weight Qwen3.6-35B-A3B release is available through Alibaba Cloud Model Studio as qwen3.6-flash. Use Qwen/Qwen3.6-35B-A3B for the downloadable repository and qwen3.6-flash for the hosted API route.
That official relationship does not make Marketplace context, tools, caching, rate limits, pricing, or data terms part of the open repository. No equivalent public repository was verified for the exact qwen3.7-flash or legacy qwen-flash service IDs. For local deployment, follow the Qwen download guide and the exact model card and license.
Which Flash model should you choose?
- New QwenCloud project: start with
qwen3.7-flashwhen the current selector and your account expose it. - Model Studio project: evaluate
qwen3.6-flashon the matching global route, orqwen3.6-flash-usonly for the documented US scope. - Reproducible production behavior: pin a dated snapshot when offered and revalidate before changing it.
- Existing
qwen-flashworkload: treat the old ID as a legacy compatibility target and run a migration set before switching.
Availability, data residency, base URL, key type, model ID, price, and rate limit are one route-specific configuration. Verify them in the workspace where the application will run.
A safe migration checklist
- Write down the old provider, region, base URL, exact model ID, and thinking setting.
- Keep provider and region constant during the first comparison where possible.
- Test final-answer quality, refusals, JSON validity, tool calls, long prompts, visual inputs, and streaming separately.
- Use a warm-up plus at least ten repeated requests before reporting median or P95 latency; one request is not a benchmark.
- Compare token usage and current provider pricing rather than relying on “cheapest” claims.
- Pin a dated snapshot if an unannounced mainline change would be unacceptable.
Verification status and test limits
We checked the model names, lifecycle labels, modalities, developer matrices, Marketplace cards, Qwen3.6 launch, and open repository on August 4, 2026. No authenticated inference or maximum-length request was completed for this update, so we publish no independent output ceiling, speed ranking, cost ranking, or quality result.
A publishable live test must record the provider, deployment scope, endpoint, key type, exact model ID, protocol, media, thinking setting, configured output cap, date, status, request ID, usage, and limitations without exposing credentials.
Official sources
- QwenCloud text-model matrix
- QwenCloud vision-model matrix
- Qwen3.7 Flash Marketplace card
- QwenCloud model changelog
- Qwen3.6-35B-A3B launch and hosted-ID mapping
- Official Qwen3.6-35B-A3B repository
- Alibaba Cloud Model Studio text models
- Alibaba Cloud Model Studio regional model IDs and rate limits
What is the current Qwen Flash model?
For QwenCloud, the current mainline is qwen3.7-flash. Alibaba Cloud Model Studio documents qwen3.6-flash on supported routes. The older qwen-flash ID is a legacy family.
Does Qwen Flash have a 1M context window?
The cited service matrices document a 1M context for the current Flash families and legacy qwen-flash. Maximum output is route-specific: checked August 4, 2026, QwenCloud’s Qwen3.7 Marketplace card and developer matrices disagreed at 131K versus 64K.
Can I download Qwen Flash and run it locally?
Use repository IDs, not hosted aliases, for local deployment. Qwen officially connects the hosted qwen3.6-flash access path with the open Qwen/Qwen3.6-35B-A3B release, but their identifiers and route-level specifications remain separate. No matching public repository was verified for the exact qwen3.7-flash or legacy qwen-flash service IDs.

