Qwen Max is a hosted model tier, not one permanent model or a downloadable checkpoint. The current production flagship is qwen3.8-max, launched by QwenCloud on August 3, 2026. qwen3.7-max remains active and retains a separate operational role: it is still the formal replacement named for the affected older Qwen3 Max IDs scheduled to retire on October 10, 2026.
Exact identifiers must remain separate. qwen3.8-max, qwen3.8-max-preview, qwen3.7-max, qwen3-max, qwen-max, and dated snapshots can differ in lifecycle, access route, context, modalities, features, prices, and reproducibility. QwenCloud now documents one explicit exception: qwen3.8-max-preview is retired, and calls using that old Token Plan ID are automatically routed to qwen3.8-max. This routing must not be generalized to any other older Max alias or retirement notice.
Last verified: August 23, 2026. Model aliases, snapshots, regions, limits, modes, prices, and retirement dates can change. Confirm the exact ID in the official model directory and in the console for the region and product used by your application.
Independence notice: Qwen-AI.chat is an independent informational website and demo. It is not operated, authorized, or endorsed by Alibaba, Alibaba Cloud, or the Qwen team. This page documents hosted identifiers and first-party checkpoints; it does not imply that this site’s chat interface runs any model listed here.
Qwen Max names and IDs at a glance
| Name or ID | What it identifies | Access and status | Do not confuse it with |
|---|---|---|---|
| Qwen Max | A hosted Qwen product tier or category | Use an exact service model ID in requests | A checkpoint name, fixed parameter count, or family-wide license |
qwen3.8-max |
Production hosted Qwen3.8 Max ID | Current hosted flagship; 1M context and managed multimodal/tool features on documented QwenCloud routes | Qwen/Qwen3.8-2.4T-A95B, which is related but not interchangeable |
qwen3.8-max-preview |
Retired Token Plan model ID | Old calls are automatically routed to qwen3.8-max; Credits and usage statistics are calculated as production |
A separate live preview or a reproducible pinned model |
Qwen/Qwen3.8-2.4T-A95B |
Official open-weight Qwen3.8 checkpoint | 2.4T total, 95B active; text-only; thinking-only; 262,144 native context; custom Qwen3.8-Max License | The hosted qwen3.8-max service ID or its full managed feature set |
qwen3.7-max |
Hosted Qwen3.7 Max mainline ID | Active, up to 1M context, and the formal replacement for the retiring older Max IDs listed below | A retired model, qwen3-max, or an open-weight repository |
qwen3.6-max-preview, qwen3-max-preview, qwen3-max |
Older hosted Max mainline IDs | Scheduled to retire October 10, 2026; formal replacement: qwen3.7-max |
qwen3.8-max as an undocumented automatic replacement |
qwen3-max-2026-01-23, qwen3-max-2025-09-23 |
Older hosted Qwen3 Max snapshots | Scheduled to retire October 10, 2026; formal replacement: qwen3.7-max |
Indefinitely supported pinned versions |
qwen-max |
Legacy hosted rolling Model Studio ID | The exact model card lists 32,768 context, 30,720 maximum input, 8,192 maximum output, and qwen-max-2025-01-25 as the rolling equivalent; Alibaba’s general legacy matrix conflicts by listing 128K |
Qwen2.5-Max as a release name, qwen3-max, or qwen3.7-max |
qwen-max-2025-01-25 |
Qwen2.5-Max launch snapshot and the current exact card’s stated equivalent for the rolling qwen-max alias |
Alibaba also cites the dated endpoint as a deprecated historical snapshot; do not call the snapshot directly without current route verification | The rolling qwen-max string or Qwen/Qwen2.5-72B-Instruct |
Qwen/Qwen2.5-72B-Instruct |
Downloadable Qwen2.5 text checkpoint | 72.7B parameters with its repository-specific custom Qwen License | Qwen2.5-Max or any hosted Max ID |
Qwen/Qwen3-235B-A22B |
Downloadable Qwen3 MoE checkpoint | 235B total and 22B activated parameters under Apache 2.0 | The hosted qwen3-max service ID |
Which Qwen Max ID should a developer use?
- New highest-capability evaluation: start with
qwen3.8-maxwhere that exact production ID, endpoint, and billing route are available. Treat QwenCloud’s performance description as provider positioning until the model is tested on representative application data. - Existing
qwen3.7-maxintegration: there is no retirement-driven requirement to leave it based on the August 3 policy. Keep it when its snapshots, regional availability, cost, latency, or validated behavior fit the application; testqwen3.8-maxas a separate migration candidate. - Existing
qwen3-max,qwen3-max-preview, orqwen3.6-max-previewintegration: the formal retirement replacement remainsqwen3.7-max. Test and migrate before October 10, 2026. Do not substituteqwen3.8-maxinto the compliance plan unless the provider updates the notice or your own validation supports a separate change. - Existing affected Qwen3 Max snapshot:
qwen3-max-2026-01-23andqwen3-max-2025-09-23also formally point toqwen3.7-maxfor the October 10 retirement. - Existing
qwen-maxintegration: keep the literal rolling alias distinct from the Qwen2.5-Max release name and from newer Max generations. Its exact card currently states 32,768 context and namesqwen-max-2025-01-25as the equivalent snapshot, while the general legacy matrix lists 128K and other documentation treats the dated endpoint as deprecated. Configure against the exact card, test the live route, and plan an explicit migration rather than assuming either limit or mapping. - Self-hosting: choose an exact first-party weight repository. Hosted Max service identifiers are not Hugging Face checkpoint IDs.
Use the Qwen API guide for route-specific invocation, the Qwen pricing guide for current costs, and the Qwen download directory for separately published weights.
qwen3.8-max: the current production flagship
QwenCloud released the exact ID qwen3.8-max on August 3, 2026. Its release notes describe a native vision-language Max model with a 2.4-trillion-parameter Mixture-of-Experts architecture, hybrid thinking enabled by default, and a 1M-token context window. QwenCloud positions it as its most capable flagship to date and as an improvement over the Qwen3.7 series. These are official provider specifications and claims; this page does not convert them into an independently reproduced ranking.
The current QwenCloud model matrix marks Thinking, Function Calling, built-in tools, and structured output for qwen3.8-max. Confirm the exact API reference and account access before implementation because newly launched model pages and API references can update at different times.
qwen3.8-max is the production hosted identifier. qwen3.8-max-preview has ended its preview period and is officially retired. QwenCloud still accepts the old Token Plan string but automatically routes calls to qwen3.8-max, with Credits and usage statistics calculated as production. Update configurations rather than treating the retired ID as a separate model. See the dedicated Qwen3.8 Max guide, API guide, and pricing guide for implementation details.
Qwen also publishes Qwen/Qwen3.8-2.4T-A95B and an official FP8 variant. The model card says the hosted Qwen3.8-Max service is based on this open checkpoint, but adds vision input, non-thinking support, a 1M context length by default, official built-in tools, and other managed features. The open checkpoint is text-only and thinking-only and uses the custom Qwen3.8-Max License. A related open release does not make the API ID and repository ID interchangeable.
qwen3.7-max: active predecessor and formal legacy replacement
qwen3.7-max remains a listed hosted model even though it is no longer the newest flagship. It continues to provide documented 1M-context deployments and its May 20 and June 8 snapshots. More importantly for lifecycle accuracy, the current deprecation policy still names qwen3.7-max as the formal replacement for the affected Qwen3 Max mainline IDs and snapshots scheduled to retire on October 10, 2026. The dedicated Qwen3.7 Max guide owns its detailed capabilities, pricing, and API caveats.
| Mainline ID | qwen3.7-max |
|---|---|
| Documented snapshots | qwen3.7-max-2026-05-20 and qwen3.7-max-2026-06-08 |
| Real-time context | Up to 1M tokens in the Model Studio pricing tables reviewed |
| Modes | Thinking and non-thinking modes are documented for listed deployments; confirm behavior in the selected region and interface |
| Distribution | Hosted proprietary Model Studio service; no downloadable weights are implied by the API ID |
qwen3.7-max-2026-05-20 and also listed the June 8 snapshot. A provider-managed mainline alias can be remapped, while a dated snapshot identifies a specific served version.[3]Context caveat: the 1M figure belongs to documented real-time qwen3.7-max deployments. Alibaba Cloud limits qwen3.7-max batch requests to 256K context tokens. Do not advertise one universal context limit without naming the inference route.[9]
qwen3-max: a separate 256K hosted model
qwen3-max is not a shorter spelling of qwen3.7-max. Alibaba Cloud’s pricing tables give the former up to 256K context and the latter up to 1M in the documented real-time services. The same tables mapped qwen3-max to qwen3-max-2026-01-23 on the verification date and also listed qwen3-max-2025-09-23.[3]
QwenCloud schedules qwen3.6-max-preview, qwen3-max-preview, and qwen3-max for retirement on October 10, 2026, with qwen3.7-max as the documented replacement. The same date and replacement apply to qwen3-max-2026-01-23 and qwen3-max-2025-09-23. The later launch of qwen3.8-max does not alter that formal chain unless the provider updates the policy. After retirement, inference calls to an affected ID are documented to fail, so applications must migrate explicitly rather than expect automatic substitution. See the dedicated Qwen3-Max migration guide.
The hosted ID also does not identify Qwen/Qwen3-235B-A22B. The latter is a downloadable Qwen3 checkpoint with a public model card, explicit architecture details, and an Apache 2.0 license. Similar family branding does not make the hosted service and open-weight checkpoint interchangeable.
qwen-max: a legacy rolling alias with conflicting documentation
Alibaba Cloud’s exact qwen-max model card lists 32,768 tokens of context, 30,720 maximum input tokens, 8,192 maximum output tokens, and identifies qwen-max-2025-01-25 as the rolling alias’s equivalent snapshot.[11] Alibaba Cloud’s general legacy-model matrix conflicts by listing qwen-max at 128K context.[12] Neither page turns the rolling service ID into a downloadable checkpoint.
The dated ID needs separate lifecycle handling. It was the API name in the January 2025 Qwen2.5-Max release, and the exact current card uses it as an equivalence label; however, Alibaba’s error documentation also presents that dated endpoint as a deprecated historical snapshot.[10] Do not copy qwen-max-2025-01-25 into a new request without confirming it in the target region and workspace.
Safe rule: preserve the exact model string in documentation and configuration, use 32,768 as the conservative limit for the rolling alias unless the exact route explicitly documents and accepts more, and do not attach the 1M limit of qwen3.7-max, the 256K limit of qwen3-max, or any open-weight parameter count or license to qwen-max.
Qwen2.5-Max and its actual launch ID
Qwen announced Qwen2.5-Max on January 28, 2025 as a large-scale mixture-of-experts model trained on more than 20 trillion tokens and post-trained with supervised fine-tuning and reinforcement learning from human feedback. The release provided hosted access through Alibaba Cloud and stated the API model name precisely: qwen-max-2025-01-25.[5]
“Qwen2.5-Max” was the release name; qwen-max-2025-01-25 was the launch snapshot used in the example request. The official announcement did not publish a parameter count, so this page does not assign one. In particular, it must not be described as 72B or as a renamed Qwen/Qwen2.5-72B-Instruct checkpoint.
The snapshot format is also a lifecycle boundary. Alibaba Cloud’s retirement policy uses qwen-max-2025-01-25 as an example of a date-stamped Qwen snapshot.[4] Its error documentation then names the same ID as an example of a deprecated historical snapshot and says the endpoint is no longer available after deprecation.[10] Do not copy this launch ID into a new integration.
Why Qwen2.5-72B is not Qwen Max
Qwen/Qwen2.5-72B-Instruct is a first-party downloadable text checkpoint. Its model card states 72.7B total parameters, 70.0B excluding embeddings, a 131,072-token full sequence with documented YaRN setup, and the repository’s custom Qwen License.[7] Those facts belong to that repository, not to Qwen2.5-Max or any hosted qwen-*-max ID.
| Question | Hosted Qwen Max ID | Open-weight Qwen checkpoint |
|---|---|---|
| How is it accessed? | Alibaba Cloud Model Studio endpoint and regional credentials | First-party weight repository and a compatible inference runtime |
| What name goes in the request? | An Alibaba model ID such as qwen3.7-max |
The served checkpoint ID or an explicit local alias |
| Are weights included? | No; an API catalog entry does not provide checkpoint files | Yes, when the first-party repository publishes them |
| Which license applies? | Hosted-service terms and applicable provider terms | The license in that exact repository |
| Can specifications be copied across? | No. Context, modes, parameter count, architecture, license, and retirement status must stay attached to the exact identifier. | |
Alibaba Cloud describes Model Studio’s ready-to-use Qwen services as proprietary.[6] For downloadable model families, use the separate Qwen2.5 checkpoint guide or Qwen3 checkpoint guide and verify the exact repository license.
Migration checklist for a Max integration
- Inventory the literal ID: record every production, staging, batch, evaluation, and fallback configuration containing
qwen-max,qwen3-max,qwen3.7-max,qwen3.8-max-preview, orqwen3.8-max. - Match the region: confirm the workspace, API key, base URL, deployment scope, and selected model are available in the same Model Studio region.
- Choose alias or snapshot deliberately: a mainline alias can move to another provider-managed version; a snapshot is clearer for reproducibility but can still be retired.
- Test mode behavior: compare thinking and non-thinking settings, reasoning-token handling, streaming, stop conditions, and output length.
- Retest tools and JSON: validate function names, argument schemas, parallel calls, malformed outputs, retries, and application-side authorization.
- Measure your workload: evaluate answer quality, latency, token use, context truncation, caching, rate limits, and cost with representative data.
- Separate real-time and batch limits: do not send a batch request sized for the 1M real-time limit; the documented
qwen3.7-maxbatch cap is 256K. - Remove unsupported claims: do not publish a parameter count, architecture, self-hosting path, or open-source license for a hosted Max ID unless a first-party source states it for that exact ID.
Frequently asked questions
Is Qwen Max one model?
No. Qwen Max is a hosted product tier. qwen3.8-max, qwen3.8-max-preview, qwen3.7-max, qwen3-max, qwen-max, and dated snapshots are separate identifiers whose specifications and lifecycle status must not be merged.
Which Qwen Max ID should a new integration evaluate?
Evaluate the production qwen3.8-max when the current flagship tier is required and the exact ID is available through the intended QwenCloud route. Also evaluate qwen3.7-max when regional availability, documented snapshots, established behavior, or the formal migration path from retiring Qwen3 Max IDs matters. Do not silently replace one ID with the other.
Did qwen3.8-max replace qwen3.7-max?
Qwen3.8 Max replaced Qwen3.7 Max in flagship positioning, not through a published retirement of qwen3.7-max. The current policy still lists Qwen3.7 Max and names it as the replacement for older Qwen3 Max IDs. Treat a future lifecycle change only as official after it appears in the provider’s deprecation documentation.
Is qwen-max the same as Qwen2.5-Max?
No. Qwen2.5-Max is a release name, qwen-max is a rolling service ID, and qwen-max-2025-01-25 is a dated snapshot. The exact current card names that snapshot as the rolling alias’s equivalent, while Alibaba also describes the dated endpoint as deprecated elsewhere. Preserve the literal ID and verify the live regional route instead of treating the three labels as interchangeable.
Is Qwen2.5-Max the 72B Qwen2.5 checkpoint?
No. Qwen2.5-Max was a hosted large-scale mixture-of-experts release. Qwen/Qwen2.5-72B-Instruct is a separate downloadable 72.7B checkpoint with its own model card and custom Qwen License.
Are qwen3-max and qwen3.7-max interchangeable?
No. qwen3-max has a documented 256K real-time context limit and is scheduled for retirement on October 10, 2026. qwen3.7-max is its documented replacement and has up to 1M real-time context in the reviewed Model Studio tables.
Is Qwen3-Max the same as Qwen3-235B-A22B?
No. qwen3-max is a hosted Alibaba Model Studio ID. Qwen/Qwen3-235B-A22B is a separate downloadable checkpoint with 235B total parameters, 22B activated parameters, and an Apache 2.0 license.
Can I download Qwen Max weights?
The hosted Max IDs are service identifiers and cannot be downloaded as API IDs. However, Qwen now publishes the related Qwen/Qwen3.8-2.4T-A95B checkpoint and an official FP8 variant. Qwen states that hosted Qwen3.8-Max is based on this model but adds features that the downloadable checkpoint does not provide. Download only from the official repository, read the custom license, and do not assume that self-hosting reproduces the hosted service’s modalities, tools, context defaults, pricing, or behavior.
Does qwen3.7-max support one million tokens?
Alibaba Cloud documents up to 1M context for listed real-time qwen3.7-max deployments. Its batch-inference documentation caps qwen3.7-max requests at 256K, so the route must be stated with the limit.
What is the difference between a mainline ID and a snapshot?
A mainline ID such as qwen3.7-max can be mapped by the provider to a served version. A dated snapshot identifies a specific version more clearly, but it can still be retired and is not guaranteed to remain available indefinitely.
Does the same Qwen Max setup work in every region?
No. Model availability, workspace, API key, base URL, deployment scope, modes, limits, and pricing can differ by region. Copy the exact ID and endpoint from the regional Model Studio console used by your application.
Official sources reviewed
- QwenCloud model releases — August 3 production launch of
qwen3.8-max. - QwenCloud text-generation model directory — current portfolio position and capability matrix.
- QwenCloud Token Plan — separate preview status for
qwen3.8-max-preview. - QwenCloud model deprecation policy — formal October 10 replacement chains.
- Alibaba Cloud Model Studio: exact Qwen Max model card — rolling alias limits and equivalent snapshot.
- Alibaba Cloud Model Studio: text-generation model matrix — conflicting general legacy-model context entry.
- Alibaba Cloud Model Studio: Model inference pricing — Max IDs, snapshots, deployment scopes, modes, and context tiers. Consult the source for prices because they can change.
- Alibaba Cloud Model Studio: Model decommissioning policy — October 10, 2026 retirements, replacement ID, and snapshot lifecycle.
- Qwen Team: Qwen2.5-Max release — release date, hosted access, training description, and launch API ID.
- Alibaba Cloud: What is Model Studio — hosted proprietary Qwen services and Qwen Max positioning.
- Qwen/Qwen2.5-72B-Instruct model card — checkpoint ID, parameter count, context configuration, and repository license.
- Qwen/Qwen3-235B-A22B model card — checkpoint ID, total and activated parameters, context, and Apache 2.0 license.
- Alibaba Cloud Model Studio: Batch inference — supported Max IDs and the 256K batch cap for
qwen3.7-max. - Alibaba Cloud Model Studio: Error codes — the deprecated
qwen-max-2025-01-25example and unavailable endpoint behavior.
Editorial review: The Qwen-AI.chat Editorial Team cross-checked every model name, API ID, snapshot, context statement, distribution method, license label, and retirement date against the first-party sources listed above. No third-party parameter estimates or benchmark rankings are used. Corrections can be submitted through the contact page.

