Last verified: August 23, 2026 for both Qwen and Google sources. The current Qwen vs Google Gemini comparison begins with a lifecycle correction. QwenCloud launched qwen3.8-max on August 3 as its flagship. Google made gemini-3.7-flash generally available on August 13 and now presents it as the latest, more capable Flash model. gemini-3.6-flash and gemini-3.5-flash-lite remain generally available, while gemini-3.1-pro-preview remains a Preview model.
There is no defensible single winner without a controlled test. Gemini 3.7 Flash and 3.5 Flash-Lite have explicit multimodal and tool matrices, while Gemini 3.6 Flash remains a supported GA predecessor. Qwen offers a current hosted flagship, lower-priced published Qwen3.7 tiers and a broader ecosystem that also includes separately released open-weight checkpoints. The right route depends on input types, required tools, region, data terms, cost and whether local deployment matters.
For the shared source-first method and links to every current matchup, use the Qwen AI comparison hub.
Qwen vs Gemini at a glance
| Dimension | Qwen | Google Gemini |
|---|---|---|
| Current headline model | qwen3.8-max, launched August 3, 2026 |
gemini-3.7-flash is the latest GA Flash model; gemini-3.6-flash remains GA and gemini-3.1-pro-preview remains Preview |
| Lower-cost current tier | qwen3.7-flash |
gemini-3.5-flash-lite |
| Documented context | 1M for Qwen3.8 Max and the current Qwen3.7 Max, Plus and Flash services | 1,048,576 input tokens for Gemini 3.7 Flash and 3.5 Flash-Lite |
| Documented max output | Qwen3.8 Max: 131K. Qwen3.7 official-source discrepancy: the QwenCloud Marketplace cards for Max, Plus, and Flash show 131K, while the developer text and vision matrices show 64K. Both paths were checked August 23, 2026; neither exposed a visible update date. Treat the ceiling as route-specific and configure conservatively until the exact endpoint and account are tested. | 65,536 output tokens for Gemini 3.7 Flash and 3.5 Flash-Lite |
| Multimodal API inputs | Qwen’s current vision-language services are model-specific; verify the exact ID | Text, image, video, audio and PDF for both current GA models compared here |
| Local deployment | Possible with separately released open-weight Qwen checkpoints | The Gemini API models compared here are managed Google services |
| Lifecycle choice | Use a rolling alias for updates or a dated snapshot where one is offered and reproducibility matters | Prefer GA models for production; treat Preview lifecycle and shutdown risk separately |
Current Qwen lineup
QwenCloud’s August 3 changelog describes qwen3.8-max as its most capable flagship so far: a native vision-language, 2.4T mixture-of-experts model with hybrid thinking enabled by default and a 1M context. Its current text-generation guide lists thinking, function calling, built-in tools and structured output.
For workloads where published price and tiering matter, qwen3.7-plus remains the documented balanced route and qwen3.7-flash the lower-cost route. Both list 1M context plus thinking, function calling, built-in tools and structured output. Their public maximum-output documentation is inconsistent; use the dated 131K-versus-64K source disclosure in the table above and configure conservatively until the exact endpoint and account are tested. See the Qwen model directory for exact IDs and status.
The model-specific QwenCloud Marketplace page checked August 23 lists a 131K maximum output and PAYG list prices of $2 input and $6 output per 1M tokens for qwen3.8-max. These figures belong to QwenCloud Marketplace, not Alibaba Cloud Model Studio. Do not transfer them to a regional Model Studio estimate; use the Qwen pricing guide for route-separated costs.
Current Gemini lineup
Gemini 3.7 Flash
gemini-3.7-flash became generally available on August 13, 2026. Google presents it as the latest and more capable Flash model. Its model page documents text, image, video, audio and PDF input with text output, a 1,048,576-token input limit and a 65,536-token output limit.
The same page lists caching, code execution, file search, function calling, Google Maps grounding, search grounding, structured outputs, thinking and URL context. Capability labels describe the API surface, not independently measured accuracy or reliability; test the exact files, tools and schemas used by the application.
Gemini 3.6 Flash and Gemini 3.5 Flash-Lite
gemini-3.6-flash remains generally available, but it is no longer the latest Flash model after Gemini 3.7 Flash reached GA. Existing integrations should not be described as obsolete solely because a newer GA model exists; compare behavior, price and lifecycle before migrating.
gemini-3.5-flash-lite remains Google’s GA lower-cost model for high-throughput work. Its official matrix documents text, image, video, audio and PDF input, a 1,048,576-token input limit and a 65,536-token output limit, along with thinking and listed tool capabilities.
Lifecycle note for older names
gemini-3-flash-preview remains an older Preview-era name in Google’s lifecycle documentation. gemini-3.1-pro-preview is still a Preview model. “No shutdown date announced” does not convert a Preview into a GA release; production comparisons should label lifecycle status explicitly and prefer current GA IDs unless a specific Preview capability is required.
API pricing
Prices below are public standard rates per 1 million tokens, rechecked on August 23, 2026. They exclude taxes, negotiated agreements, tool calls, cache storage unless shown, priority processing and consumer subscriptions. Gemini output pricing includes thinking tokens. The Qwen3.8 row is a QwenCloud Marketplace rate, not a Model Studio regional quote.
| Model | Rate period or input scope | Input | Cached input | Output |
|---|---|---|---|---|
qwen3.8-max; QwenCloud Marketplace |
1M context; 131K maximum output | $2.00 | Separate Marketplace cache rates; not compared here | $6.00 |
qwen3.7-max |
0–991K | $2.50 | Route-specific | $7.50 |
qwen3.7-plus |
Up to 256K | $0.40 | Route-specific | $1.60 |
qwen3.7-plus |
Over 256K–1M | $1.20 | Route-specific | $4.80 |
qwen3.7-flash |
Up to 32K | $0.03 | Route-specific | $0.13 |
qwen3.7-flash |
Over 32K–256K | $0.10 | Route-specific | $0.40 |
qwen3.7-flash |
Over 256K–1M | $0.20 | Route-specific | $0.80 |
gemini-3.7-flash |
Introductory standard rate through December 31, 2026 | $0.75 | $0.075 plus $0.50 per 1M cached tokens per hour | $3.75 |
gemini-3.7-flash |
Standard rate from January 1, 2027 | $1.50 | $0.15 plus $1.00 per 1M cached tokens per hour | $7.50 |
gemini-3.5-flash-lite |
Standard paid tier | $0.30 | $0.03 plus cache storage | $2.50 |
Google’s Batch and Flex rates for Gemini 3.7 Flash are currently $0.375 input, $0.0375 cached input and $1.875 output per million tokens through December 31, 2026; those rates double on January 1, 2027. These asynchronous or flexible service tiers do not promise the same latency as standard requests.
Search and Maps grounding can add per-query charges after the included allowance. QwenCloud also lists separate built-in-tool fees. A token-only table is therefore not a total-cost calculator. Review the exact provider route and the current pricing reference.
Multimodal work: what the specifications establish
Google explicitly lists text, image, video, audio and PDF inputs for Gemini 3.7 Flash and Gemini 3.5 Flash-Lite. QwenCloud describes Qwen3.8 Max as native vision-language and lists Qwen3.7 multimodal services, but support must still be verified against the exact Qwen ID and route.
This does not prove one model reads a particular chart, video, scanned document or audio clip more accurately. A useful evaluation should include representative files, expected answers, OCR edge cases, temporal questions for video, long-PDF retrieval and structured extraction. Record tokenization and output quality instead of inferring performance from an “accepts PDF” checkmark.
Tools, grounding and agents
Gemini 3.7 Flash and 3.5 Flash-Lite document function calling, code execution, search grounding, Maps grounding, file search, URL context and structured output. QwenCloud documents function calling and built-in tools for its current flagship, Plus and Flash tiers. The implementation difference matters more than the checklist: permissions, schemas, retries, citations and tool-call pricing determine production reliability.
For Qwen implementation examples and provider boundaries, use the Qwen API guide. Never assume a parameter or tool name transfers unchanged between Gemini’s API and a Qwen OpenAI-compatible route.
Data-use and deployment boundaries
Google’s Gemini API pricing page marks content on the free tier as used to improve Google products and content on the paid tier as not used to improve products. Enterprise and cloud terms can differ, so organizations should verify the contract attached to the actual project rather than relying on a comparison article.
Qwen’s ecosystem additionally includes separately published open-weight checkpoints. A hosted alias such as qwen3.8-max is not automatically a downloadable checkpoint. If local or controlled deployment is required, consult the Qwen download guide and verify the exact model card, license and hardware plan.
Which should you choose?
- Start with Gemini 3.7 Flash when you need Google’s latest GA Flash model and its documented multimodal, grounding and tool integrations; account for the introductory rate ending December 31, 2026.
- Keep or evaluate Gemini 3.6 Flash when an existing GA integration is already validated, but do not present it as Google’s latest Flash model.
- Evaluate Gemini 3.5 Flash-Lite for high-volume extraction or agent tasks where its lower standard rate fits the workload.
- Evaluate Qwen3.8 Max when the current Qwen flagship is the capability target; confirm route availability and use the QwenCloud Marketplace $2 input / $6 output rate only for that route.
- Evaluate Qwen3.7 Plus or Flash when you need a currently priced Qwen tier with a documented 1M context and built-in tools.
- Use a separately released Qwen checkpoint when deployment control is a hard requirement.
Methodology and limitations
This page was updated from official QwenCloud and Google developer documentation available on August 23, 2026. The review covered current model IDs, lifecycle status, published inputs, context, output limits, capabilities and list pricing. It did not include login-dependent consumer-app testing, latency measurement or scored output comparisons. Provider marketing statements and benchmarks are not independent evidence.
For a live comparison, use the same files and prompts, disable or enable tools consistently, pin model versions where possible, record thinking settings and output schemas, and publish the scoring rubric plus failures. Consumer Gemini, Google AI Studio, Vertex AI and the Gemini Developer API can have different features, quotas and terms; QwenCloud, Alibaba Cloud and third-party Qwen deployments can also differ.
Frequently asked questions
What is the current Google Gemini Flash model for production?
gemini-3.7-flash is Google’s latest generally available Flash model as of August 23, 2026. gemini-3.6-flash and gemini-3.5-flash-lite also remain GA, while gemini-3.1-pro-preview remains Preview.
Does Gemini 3.7 Flash have a 1M context window?
Its official model page lists a 1,048,576-token input limit and a 65,536-token output limit. Effective product limits and usable input can still depend on the route, files, tools and account.
What does Gemini 3.7 Flash cost?
The introductory standard rate through December 31, 2026 is $0.75 input, $0.075 cached input plus $0.50 per million cached tokens per hour, and $3.75 output per million tokens. From January 1, 2027, those rates become $1.50, $0.15 plus $1.00 per hour, and $7.50 respectively. Check Google’s current pricing page before budgeting.
Which is cheaper, Qwen or Gemini?
Some published Qwen rates are lower than some Gemini rates, but total cost depends on task quality, request bands, output length, tools, cache storage, retries and service tier. The temporary Gemini 3.7 Flash introductory rate also changes on January 1, 2027. Do not use a QwenCloud Marketplace row as a Model Studio regional price.
Can Gemini 3.7 Flash read PDFs, video and audio?
Google’s API model page lists PDF, video and audio as supported input types, along with text and images; output is text. That capability listing does not guarantee accuracy on every file, so test representative inputs.
Can I run Qwen or Gemini locally?
Some separately released Qwen checkpoints can be self-hosted under their specific licenses. The hosted Qwen aliases and Gemini API models compared here should not be treated as downloadable weights.

