Qwen3-Coder: Open Models, API Retirement, and Setup

Qwen3-Coder is a family of text-based coding language models developed by the Qwen team. It includes large mixture-of-experts Instruct checkpoints, the separate Qwen3-Coder-Next architecture, downloadable model variants, and hosted Alibaba Cloud Model Studio IDs.

Last verified: August 3, 2026.

The open Qwen3-Coder Instruct checkpoints are non-thinking models and do not generate <think> blocks. Hosted Model Studio services must be treated as separate interfaces.

Hosted retirement notice: Alibaba Cloud and QwenCloud schedule qwen3-coder-plus and five listed Qwen3-Coder hosted snapshots or service IDs for retirement on October 10, 2026, with qwen3.7-plus as the replacement. This provider lifecycle does not retire downloaded Hugging Face weights.

Official open Qwen3-Coder checkpoints

Exact Hugging Face IDTypeParametersContextThinking behaviorLicense
Qwen/Qwen3-Coder-NextInstruct80B total / 3B active262,144 nativeNon-thinking onlyApache 2.0
Qwen/Qwen3-Coder-Next-BaseBase80B total / 3B active262,144 nativeBase checkpointApache 2.0
Qwen/Qwen3-Coder-480B-A35B-InstructInstruct480B total / 35B active262,144 native; documented 1M extension with YaRNNon-thinking onlyApache 2.0
Qwen/Qwen3-Coder-30B-A3B-InstructInstruct30.5B total / 3.3B active262,144 native; documented 1M extension with YaRNNon-thinking onlyApache 2.0

FP8 and GGUF variants are also published for selected checkpoints. Their precision, runtime support, memory use, and performance should be evaluated separately from the original BF16 repositories.

Active-parameter labels describe mixture-of-experts computation. An 80B-total model does not become a 3B download merely because about 3B parameters are activated during token processing.

What is Qwen3-Coder-Next?

Qwen3-Coder-Next is not a renamed 30B or 480B checkpoint. It is built on Qwen3-Next-80B-A3B-Base and uses a hybrid architecture combining Gated DeltaNet layers, attention layers, and a mixture-of-experts design.

  • 80B total parameters.
  • Approximately 3B active parameters.
  • 262,144-token native context.
  • Text input and text output.
  • Non-thinking Instruct behavior.
  • Separate Base and Instruct repositories.

The checkpoint-specific Qwen3-Coder-Next model card documents 262,144 tokens natively. Its official vLLM and SGLang launch examples focus on that 256K range and advise reducing the context to 32,768 if a server cannot start.

The family-level Qwen3-Coder repository says the native 256K context can be extended to 1M with YaRN. Treat 1M as an optional context-extension claim, not a native guarantee: validate the exact checkpoint, runtime, YaRN configuration, memory use, retrieval accuracy, and output budget before presenting or deploying it.

Does open Qwen3-Coder use Thinking mode?

The open Instruct checkpoints listed above support non-thinking mode only. They do not generate <think></think> blocks, and adding enable_thinking=False is unnecessary.

This rule applies to the downloadable open checkpoints. A hosted service with a similar lowercase model ID can expose a different service-level interface and must not be used as evidence that the local checkpoint has hybrid Thinking behavior.

The documented 358-language claim

The official Qwen3-Coder repository lists support for 358 coding-language and syntax labels. The published list contains programming languages as well as markup, data, configuration, and file formats such as JSON, YAML, Markdown, CSV, and Dockerfile.

Therefore, “358 coding languages and formats in the official list” is more accurate than claiming equal proficiency across 358 programming languages. The number describes published coverage, not guaranteed performance on every syntax or toolchain.

Fill-in-the-middle support

The official repository states that fill-in-the-middle is supported across Qwen3-Coder versions. The documented prompt structure is:

<|fim_prefix|>existing code before the gap
<|fim_suffix|>existing code after the gap
<|fim_middle|>

Applications should use the tokenizer associated with the exact checkpoint because special-token IDs and chat templates are part of the model interface.

Tool calling and the qwen3_coder parser

Qwen3-Coder can propose structured tool calls, but the surrounding application must validate and execute them. The model does not independently edit files, run commands, browse websites, or deploy software.

Official vLLM and SGLang deployments use the qwen3_coder tool-call parser.

pip install "vllm>=0.15.0"

vllm serve Qwen/Qwen3-Coder-Next \
  --port 8000 \
  --max-model-len 32768 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder

The model card documents a 262,144-token default maximum for Coder-Next but also advises reducing the configured context if the server cannot start. The example above uses 32,768 to avoid presenting the largest context as a low-memory default.

Call the local endpoint with Node.js

npm install openai
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: "EMPTY",
  baseURL: "http://localhost:8000/v1",
});

const completion = await client.chat.completions.create({
  model: "Qwen/Qwen3-Coder-Next",
  messages: [
    {
      role: "user",
      content: (
        "Write a TypeScript function that validates an email-like "
        + "identifier. Explain its limitations after the code."
      )
    }
  ],
  max_tokens: 1024,
});

console.log(completion.choices[0].message.content);

Open checkpoints versus hosted Model Studio IDs

Hugging Face repository IDs and lowercase Alibaba Cloud Model Studio IDs are different interfaces. A provider can retire a hosted ID without deleting or disabling an already downloaded open checkpoint with a similar name.

Alibaba’s current service matrix and lifecycle notice establish the following status as checked on August 3, 2026:

Hosted Model Studio IDDocumented service interfaceLifecycle status
qwen3-coder-plus1M context; Thinking and Function Calling supportedScheduled for retirement October 10, 2026; replacement qwen3.7-plus
qwen3-coder-plus-2025-09-23Dated Plus snapshot; service matrix groups it with the 1M Plus lineScheduled for retirement October 10, 2026; replacement qwen3.7-plus
qwen3-coder-plus-2025-07-22Dated Plus snapshotScheduled for retirement October 10, 2026; replacement qwen3.7-plus
qwen3-coder-next256K context; service-level Thinking and Function Calling supportedScheduled for retirement October 10, 2026; replacement qwen3.7-plus
qwen3-coder-30b-a3b-instruct256K context; Thinking unsupported; Function Calling supportedScheduled for retirement October 10, 2026; replacement qwen3.7-plus
qwen3-coder-480b-a35b-instruct256K context; Thinking unsupported; Function Calling supportedScheduled for retirement October 10, 2026; replacement qwen3.7-plus
qwen3-coder-flash1M context; Thinking and Function Calling supportedNot listed in the October 10 Qwen3-Coder retirement table checked on August 3, 2026
qwen3-coder-flash-2025-07-28Dated Flash snapshotNot listed in that retirement table; verify current region and alias mapping

The absence of a Flash ID from this retirement notice is not a promise of permanent availability. Provider aliases, regions, quotas, prices, and lifecycle notices can change. Record the provider, endpoint, region, exact model ID, and verification date together.

The service-level Thinking entry for hosted qwen3-coder-next does not change the behavior of the open Qwen/Qwen3-Coder-Next checkpoint, whose official model card says it is non-thinking and does not generate <think> blocks.

Migrate a hosted call to qwen3.7-plus

Alibaba lists qwen3.7-plus as the replacement for the retiring hosted Qwen3-Coder IDs. Test it against representative repository, tool-calling, latency, cost, and safety cases before switching production traffic. Set the base URL for the exact Model Studio region and workspace documented for your account.

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.DASHSCOPE_API_KEY,
  baseURL: process.env.DASHSCOPE_BASE_URL,
});

const completion = await client.chat.completions.create({
  model: "qwen3.7-plus",
  messages: [
    {
      role: "system",
      content: "You are a careful coding assistant.",
    },
    {
      role: "user",
      content: "Review this patch for correctness and security risks.",
    },
  ],
});

console.log(completion.choices[0].message.content);

This example is documentation-derived and was not executed for this update. Do not hard-code API credentials. Verify the current model ID, endpoint, region, availability, and request fields in the Qwen API guide, then check current costs in the Qwen pricing guide. Use the download guide for open weights and the Models hub for the current family map.

Is Qwen3-Coder multimodal?

The Qwen3-Coder models covered here are text models. They do not natively accept screenshots or other image inputs. A coding agent can use an external vision model or screenshot-analysis tool, but that does not make Qwen3-Coder itself multimodal.

Deployment limitations

  • Generated code can contain functional or security defects.
  • Tool calls must be authorized, validated, and executed by the application.
  • Native 256K context can require substantial KV-cache memory.
  • A MoE active-parameter figure is not the full checkpoint size.
  • The official max_new_tokens=65536 examples do not establish a universal output guarantee for every runtime.
  • Quantization can change quality, memory use, throughput, and supported context.
  • Repository-wide benchmark claims should be read with their evaluation date and configuration.

Frequently asked questions

Does Qwen3-Coder-Next generate Thinking blocks?

The open Qwen/Qwen3-Coder-Next checkpoint is explicitly non-thinking and does not generate <think> blocks.

Does Qwen3-Coder-Next support 1M context locally?

Its checkpoint-specific card documents 262,144 tokens natively. The family repository says 256K can be extended to 1M with YaRN, but the checkpoint’s deployment examples focus on 256K. Treat 1M as optional, non-native context extension that requires checkpoint- and runtime-specific configuration and validation.

Which hosted Qwen3-Coder IDs retire on October 10, 2026?

Alibaba lists qwen3-coder-plus, qwen3-coder-next, qwen3-coder-30b-a3b-instruct, qwen3-coder-plus-2025-09-23, qwen3-coder-plus-2025-07-22, and qwen3-coder-480b-a35b-instruct. The documented replacement is qwen3.7-plus.

Does hosted retirement disable downloaded Qwen3-Coder weights?

No. The notice applies to provider-hosted service IDs. A downloaded open checkpoint remains a separate artifact that you operate under its own license and runtime requirements.

Are qwen3-coder-plus and qwen3-coder-flash downloadable weights?

They are hosted Model Studio IDs. They should not be described as Hugging Face checkpoint names or treated as identical to the 30B, 80B, or 480B open repositories.

Does Qwen3-Coder support exactly 358 programming languages equally?

The official repository publishes a list of 358 coding-language and syntax labels, including file and configuration formats. The list does not guarantee equal proficiency across every entry.

Official sources

Verification scope: documentation, repository, model-card, and lifecycle review only. No local checkpoint, Model Studio API, QwenCloud API, latency, tool-call, or code-quality test was executed for this patch, so no original output or performance result is claimed.

For the preceding generation, see Qwen2.5-Coder. For general Qwen models, review the Qwen3 guide.

This independent technical reference is not operated by Alibaba Cloud or the Qwen team. Test generated code and tool calls in a controlled environment.

Leave a Reply

Your email address will not be published. Required fields are marked *