Last verified: August 4, 2026 for Qwen3.8 Max Marketplace pricing and output limits; competitor-side sources were last checked August 3, 2026. The current Qwen vs Google Gemini comparison begins with a lifecycle correction. QwenCloud launched qwen3.8-max on August 3 as its new flagship. Google’s current production Flash choices are the generally available gemini-3.6-flash and gemini-3.5-flash-lite. The older gemini-3-flash-preview still appears in Google’s lifecycle table, but Google names Gemini 3.6 Flash as its recommended replacement.
There is no defensible single winner without a controlled test. Gemini 3.6 Flash and 3.5 Flash-Lite have unusually explicit multimodal and tool matrices. Qwen offers a current hosted flagship, lower-priced published Qwen3.7 tiers and a broader ecosystem that also includes separately released open-weight checkpoints. The right route depends on input types, required tools, region, data terms, cost and whether local deployment matters.
For the shared source-first method and links to every current matchup, use the Qwen AI comparison hub.
Qwen vs Gemini at a glance
| Dimension | Qwen | Google Gemini |
|---|---|---|
| Current headline model | qwen3.8-max, launched August 3, 2026 |
gemini-3.6-flash is the current GA Flash model; gemini-3.1-pro-preview remains a preview |
| Lower-cost current tier | qwen3.7-flash |
gemini-3.5-flash-lite |
| Documented context | 1M for Qwen3.8 Max and the current Qwen3.7 Max, Plus and Flash services | 1,048,576 input tokens for Gemini 3.6 Flash and 3.5 Flash-Lite |
| Documented max output | Qwen3.8 Max: 131K. Qwen3.7 official-source discrepancy: the QwenCloud Marketplace cards for Max, Plus, and Flash show 131K, while the developer text and vision matrices show 64K. Both paths were checked August 4, 2026; neither exposed a visible update date. Treat the ceiling as route-specific and configure conservatively until the exact endpoint and account are tested. | 65,536 output tokens for Gemini 3.6 Flash and 3.5 Flash-Lite |
| Multimodal API inputs | Qwen’s current vision-language services are model-specific; verify the exact ID | Text, image, video, audio and PDF for both current GA models compared here |
| Local deployment | Possible with separately released open-weight Qwen checkpoints | The Gemini API models compared here are managed Google services |
| Lifecycle choice | Use a rolling alias for updates or a dated snapshot where one is offered and reproducibility matters | Prefer GA models for production; treat Preview lifecycle and shutdown risk separately |
Current Qwen lineup
QwenCloud’s August 3 changelog describes qwen3.8-max as its most capable flagship so far: a native vision-language, 2.4T mixture-of-experts model with hybrid thinking enabled by default and a 1M context. Its current text-generation guide lists thinking, function calling, built-in tools and structured output.
For workloads where published price and tiering matter, qwen3.7-plus remains the documented balanced route and qwen3.7-flash the lower-cost route. Both list 1M context plus thinking, function calling, built-in tools and structured output. Their public maximum-output documentation is inconsistent; use the dated 131K-versus-64K source disclosure in the table above and configure conservatively until the exact endpoint and account are tested. See the Qwen model directory for exact IDs and status.
The model-specific QwenCloud Marketplace page checked August 4 lists a 131K maximum output and PAYG list prices of $2 input and $6 output per 1M tokens for qwen3.8-max. These figures belong to QwenCloud Marketplace, not Alibaba Cloud Model Studio. Do not transfer them to a regional Model Studio estimate; use the Qwen pricing guide for route-separated costs.
Current Gemini lineup
Gemini 3.6 Flash
gemini-3.6-flash became generally available on July 21, 2026. Google’s model page documents text, image, video, audio and PDF input with text output, a 1,048,576-token input limit and a 65,536-token output limit.
The same page lists caching, code execution, file search, function calling, Google Maps grounding, search grounding, structured outputs, thinking and URL context. Computer Use is marked Preview, while image generation and the Live API are not supported by this model. Those boundaries matter: multimodal understanding does not automatically include media generation or real-time voice.
Gemini 3.5 Flash-Lite
gemini-3.5-flash-lite is Google’s GA lower-cost model for high-throughput work. Its official matrix documents the same input types and token limits as Gemini 3.6 Flash, along with thinking and the listed tool capabilities. Google describes it as optimized for subagent tasks, document parsing, extraction and cost-sensitive execution.
What happened to older Gemini names?
Google’s deprecation table recommends gemini-3.6-flash as the replacement for gemini-3-flash-preview. It also lists gemini-3.1-pro-preview as a Preview model without an announced shutdown date. “No shutdown date announced” does not convert a Preview into a GA release. A production comparison should label its lifecycle explicitly.
API pricing
Prices below are public standard rates per 1 million tokens. The Qwen3.8 Marketplace row was checked August 4, 2026; the remaining rows retain their August 3 source check. They exclude taxes, negotiated agreements, tool calls, storage, priority processing and consumer subscriptions. Gemini output pricing includes thinking tokens. The Qwen3.8 row is a QwenCloud Marketplace rate, not a Model Studio regional quote.
| Model | Input scope | Input | Cached input | Output |
|---|---|---|---|---|
qwen3.8-max; QwenCloud Marketplace |
1M context; 131K maximum output | $2.00 | Separate Marketplace cache rates; not compared here | $6.00 |
qwen3.7-max |
0–991K | $2.50 | Route-specific | $7.50 |
qwen3.7-plus |
Up to 256K | $0.40 | Route-specific | $1.60 |
qwen3.7-plus |
Over 256K–1M | $1.20 | Route-specific | $4.80 |
qwen3.7-flash |
Up to 32K | $0.03 | Route-specific | $0.13 |
qwen3.7-flash |
Over 32K–256K | $0.10 | Route-specific | $0.40 |
qwen3.7-flash |
Over 256K–1M | $0.20 | Route-specific | $0.80 |
gemini-3.6-flash |
Standard paid tier | $1.50 | $0.15 plus cache storage | $7.50 |
gemini-3.5-flash-lite |
Standard paid tier | $0.30 | $0.03 plus cache storage | $2.50 |
Google also publishes Batch and Flex rates for the current models. For Gemini 3.6 Flash, both are $0.75 input and $3.75 output per million tokens. For Gemini 3.5 Flash-Lite, both are $0.15 input and $1.25 output. These asynchronous or flexible service tiers are not the same latency promise as standard requests.
Search and Maps grounding can add per-query charges after the included allowance. QwenCloud also lists separate built-in-tool fees. A token-only table is therefore not a total-cost calculator. Review the exact provider route and the current pricing reference.
Multimodal work: what the specifications establish
Google explicitly lists text, image, video, audio and PDF inputs on the two current GA Gemini model pages. QwenCloud describes Qwen3.8 Max as native vision-language and lists Qwen3.7 multimodal services, but support must still be verified against the exact Qwen ID and route.
This does not prove one model reads a particular chart, video, scanned document or audio clip more accurately. A useful evaluation should include representative files, expected answers, OCR edge cases, temporal questions for video, long-PDF retrieval and structured extraction. Record tokenization and output quality instead of inferring performance from an “accepts PDF” checkmark.
Tools, grounding and agents
Gemini 3.6 Flash and 3.5 Flash-Lite document function calling, code execution, search grounding, Maps grounding, file search, URL context and structured output. QwenCloud documents function calling and built-in tools for its current flagship, Plus and Flash tiers. The implementation difference matters more than the checklist: permissions, schemas, retries, citations and tool-call pricing determine production reliability.
For Qwen implementation examples and provider boundaries, use the Qwen API guide. Never assume a parameter or tool name transfers unchanged between Gemini’s API and a Qwen OpenAI-compatible route.
Data-use and deployment boundaries
Google’s Gemini API pricing page marks content on the free tier as used to improve Google products and content on the paid tier as not used to improve products. Enterprise and cloud terms can differ, so organizations should verify the contract attached to the actual project rather than relying on a comparison article.
Qwen’s ecosystem additionally includes separately published open-weight checkpoints. A hosted alias such as qwen3.8-max is not automatically a downloadable checkpoint. If local or controlled deployment is required, consult the Qwen download guide and verify the exact model card, license and hardware plan.
Which should you choose?
- Start with Gemini 3.6 Flash when you need Google’s current GA multimodal model and its documented grounding and tool integrations.
- Evaluate Gemini 3.5 Flash-Lite for high-volume extraction or agent tasks where its lower standard rate fits the workload.
- Evaluate Qwen3.8 Max when the current Qwen flagship is the capability target; confirm route availability and use the QwenCloud Marketplace $2 input / $6 output rate only for that route.
- Evaluate Qwen3.7 Plus or Flash when you need a currently priced Qwen tier with a documented 1M context and built-in tools.
- Use a separately released Qwen checkpoint when deployment control is a hard requirement.
Methodology and limitations
This page was rebuilt from official QwenCloud and Google developer documentation available on August 3, 2026. The review covered current model IDs, lifecycle status, published inputs, context, output limits, capabilities and list pricing. It did not include login-dependent consumer-app testing, latency measurement or scored output comparisons. Provider marketing statements and benchmarks are not independent evidence.
For a live comparison, use the same files and prompts, disable or enable tools consistently, pin model versions where possible, record thinking settings and output schemas, and publish the scoring rubric plus failures. Consumer Gemini, Google AI Studio, Vertex AI and the Gemini Developer API can have different features, quotas and terms; QwenCloud, Alibaba Cloud and third-party Qwen deployments can also differ.
Frequently asked questions
What is the current Google Gemini model for production?
Google lists Gemini 3.6 Flash and Gemini 3.5 Flash-Lite as generally available. Gemini 3.1 Pro Preview remains a Preview model, so its lifecycle should be treated separately.
Does Gemini 3.6 Flash have a 1M context window?
Its official model page lists a 1,048,576-token input limit and a 65,536-token output limit. Effective product limits and usable input can still depend on the route, files, tools and account.
Which is cheaper, Qwen or Gemini?
Some published Qwen rates are lower than the standard Gemini rates shown here, including QwenCloud Marketplace’s $2 input / $6 output Qwen3.8 Max row, but total cost depends on task quality, request bands, output length, tools, cache, retries and service tier. Do not use the Marketplace row as a Model Studio regional price.
Can Gemini 3.6 Flash read PDFs, video and audio?
Google’s API model page lists PDF, video and audio as supported input types, along with text and images. Output is text. That capability listing does not guarantee accuracy on every file, so test representative inputs.
Can I run Qwen or Gemini locally?
Some separately released Qwen checkpoints can be self-hosted under their specific licenses. The hosted Qwen aliases and Gemini API models compared here should not be treated as downloadable weights.

