Alibaba now ships two Qwen image lines that don’t share a version sequence. Qwen-Image-2.1, released September 20, 2026, is the open-weights model: 7B parameters in the generator, native transparent output, editing with up to 10 references, downloadable from Hugging Face. Qwen Image 3.0 and 3.0 Pro, announced July 22 and on the API since August 4, 2026, are hosted models with per-image pricing and no published weights. The number says “3.0 is newer,” but the honest framing is that they answer different questions: do you want to run the model or rent it?
This page puts the two side by side on license, cost, capabilities, and operations, and ends with a decision table. For the 2.1 feature walkthrough see What is Qwen-Image-2.1; for code, the how-to guide. Whichever you choose, both end up as an HTTP endpoint your app calls, and Apidog can hold both behind one collection so switching is a config change.
Side by side
| Qwen-Image-2.1 | Qwen Image 3.0 / 3.0 Pro | |
|---|---|---|
| Released | September 20, 2026 | Announced July 22, 2026; API August 4, 2026 |
| Distribution | Open weights (Hugging Face, ModelScope) | Hosted API only; no published weights [VERIFY] |
| License / terms | Qwen Research License; commercial use needs a separate license | Standard API terms of service |
| Price | Your GPU time | qwen-image-3.0: $0.03 per image (1K or 2K). qwen-image-3.0-pro: $0.04 at 1K, $0.075 at 2K. Input images $0.003 each [VERIFY] |
| Generator size | 7B (32 Single-Stream DiT layers) + Qwen3-VL 8B encoder | Not disclosed |
| Max resolution | 2048 x 2048 default; up to 2752 x 1536 | 2048 x 2048 |
| Transparent output | Native RGBA generation and editing | Not advertised |
| Reference images | Up to 10 | Input images supported (priced per image); limit not published [VERIFY] |
| Local editing | Circles, painted annotations, separate mask | Not advertised in the launch material |
| Text rendering | Improved typography, style, and layout | Headline feature: fine text around 10 px, 12 languages (self-reported) |
| Outputs per call | One (you loop) | Up to 6 |
| Prompt length | Not stated | About 4.5K tokens on Pro (official claim) |
| Prompt rewriting | Separate PE-T2I / PE-I2I models you run yourself | Handled server-side |
| Where to try | HF Space, Qwen Chat, ComfyUI | Model Studio console, Qwen Chat |
Sources for the 3.0 column are Alibaba’s July announcement and the QwenCloud docs as summarized by LLM Stats; re-check the price table on Model Studio before you budget, because image pricing is region-specific.
License is the first filter
For most teams this row ends the comparison. The 2.1 license says you “shall not use the Materials for any commercial purpose without obtaining a separate commercial license.” That’s different from the Apache 2.0 terms the original Qwen-Image shipped under, and it means a customer-facing feature built on 2.1 needs an email to Qwen’s business team and a signed agreement before launch.
Qwen Image 3.0 is a paid API with standard commercial terms. You pay per image and you can ship.
So: research, evaluation, internal tooling, or a project where you’re willing to negotiate a license, 2.1 is on the table. A product feature shipping this quarter with no legal budget, 3.0 is the default.
Cost: a GPU bill vs three cents an image
The 3.0 standard tier is $0.03 per image at either resolution, so 10,000 images a month is $300, plus $0.003 per input image on edits. Pro doubles at 2K.
Self-hosting 2.1 has no per-image price, but the arithmetic is only favorable at volume and with a GPU you’d otherwise leave idle. Qwen publishes no throughput numbers, so measure your own: at 40 steps and 2048 x 2048 on a single data-center GPU, count images per hour, divide the hourly rate, and compare to $0.03. Editing with many references is where 2.1’s KV cache reuse helps, since reference images are encoded once per request rather than per step. If your GPU is shared with other workloads, remember that image generation will hog it during bursts.
A rough rule: under a few thousand images a month, the API is cheaper once you count engineering time. Above tens of thousands a month with steady load and a licensing agreement in place, self-hosting wins. Between those, it depends on whether you already run GPUs.
Capabilities: transparency vs dense text
The two models were built for different jobs.
Pick 2.1 when the output needs an alpha channel. Stickers, product cutouts, UI assets, layered compositions, extracting a subject from a photo as RGBA: 2.1 does this natively in the same model, including editing a transparent layer without losing its transparency. 3.0’s launch material doesn’t mention alpha output. Doing it on 3.0 means a separate background-removal step with the halo artifacts that come with it.
Pick 2.1 when editing is the workload. Ten reference images, region selection by circle, paint, or mask, and fidelity work on faces and products are the center of the 2.1 release. The launch post shows group photos from six portraits, outfits from five product shots, and a furnished room from ten furniture photos.
Pick 3.0 when the image is mostly text. Dashboards, wireframes, posters, and slide-like layouts with fine text in many languages are what 3.0 was tuned for. 2.1 improved typography too, but 3.0’s “10 px text in 12 languages” is the stronger published claim, and its 4.5K-token prompt budget on Pro suits long layout specs.
Pick 3.0 when you need several candidates per call. Up to six outputs per request is a workflow feature 2.1 doesn’t have; with 2.1 you loop over seeds.
Neither vendor page prints a benchmark table with numbers. The 2.1 post shows a Qwen-Image-Bench chart; the 3.0 claims are self-reported. Treat both as prompts for your own evaluation, which the Apidog section makes repeatable.
Operations: what you own
Self-hosting 2.1 means owning: model downloads and updates, a GPU with enough memory (Qwen publishes no VRAM table; the README offers CPU offload and points at vLLM-Omni with FP8, SGLang, and LightX2V), a serving wrapper, queueing under burst load, and the two optional prompt-rewriting models if you want one-line prompts to work. In return you get data locality, no rate limits but your own, and a fixed model version that nobody changes under you.
Using 3.0 means owning: an Alibaba Cloud account and region choice, API keys, rate-limit handling, and the risk that the hosted model changes behavior between your test run and production. In return you get zero infrastructure and a per-image invoice.
The Nano Banana 2 API and gpt-image-2.5 API guides go through the same hosted trade-offs for Google and OpenAI, if you’re comparing across vendors rather than within Qwen.
Decision table
| Your situation | Choose |
|---|---|
| Shipping a commercial feature this quarter, no legal budget | 3.0 |
| Need transparent PNGs or RGBA editing | 2.1 (with a license if commercial) |
| Editing with many reference images | 2.1 |
| Text-heavy layouts in several languages | 3.0 |
| Research, evaluation, internal tools | 2.1 |
| Under a few thousand images a month | 3.0 |
| Tens of thousands a month, GPUs already running, license signed | 2.1 |
| Data can’t leave your network | 2.1 |
| Want six candidates per request | 3.0 |
Run both behind one collection
The practical answer for many teams is both: 2.1 for the transparent-asset pipeline and internal tooling, 3.0 for the customer-facing text layouts. That works cleanly if the two live behind one contract.
In Apidog, define a POST /v1/images endpoint once, with prompt, aspect, transparent, seed, and a references file field, and set up two environments: base_url pointing at your 2.1 wrapper (the how-to guide has a 40-line FastAPI version) and a second pointing at an adapter for Model Studio’s 3.0 endpoint. Then:
- Send the same prompt set to both with fixed seeds and save the responses as test cases.
- Assert on what differs:
Content-Type, response size floors, and for 2.1 theX-Image-Mode: RGBAheader when transparency is requested. - Record response time per request; Apidog shows it on every run, which is the throughput number Qwen doesn’t publish.
- Run the scenario weekly. When 3.0’s hosted behavior shifts or you swap 2.1 for a quantized build, the diff shows up in the report, not in support tickets.
- Mock the endpoint so the frontend builds against the contract while the model choice is still open.
Download Apidog and import the endpoint to start comparing.
FAQ
Is Qwen Image 3.0 open source? No public weights have been released for 3.0 or 3.0 Pro; they’re hosted API models [VERIFY]. Qwen-Image-2.1 is the open-weights release.
Can I use Qwen-Image-2.1 commercially? Only with a separate commercial license from Qwen. The default Qwen Research License excludes commercial use.
Which is cheaper? At low volume, 3.0 at $0.03 per image. At high, steady volume on hardware you already run, 2.1, after you’ve secured a license and measured your own throughput.
Does 3.0 generate transparent images? It isn’t advertised. Native RGBA output is a 2.1 feature; see What is Qwen-Image-2.1.
Can I try either for free? 2.1 has a Hugging Face Space and Qwen Chat access; the free-tier guide lists the options. 3.0 is available in Qwen Chat and via Model Studio’s trial quota, which varies by region.
Where to go next
2.1 and 3.0 are a fork, not an upgrade path. The license decides the first question, the workload decides the second, and volume decides the third. Whichever you land on, put it behind a contract you test: the how-to guide builds the 2.1 side, and the same Apidog collection can front the 3.0 side the day you need it.



