How to use Qwen-Image-2.1: diffusers code, transparent output, and an API you can test

Run Qwen-Image-2.1 with diffusers: text-to-image, RGBA transparent output, editing with up to 10 references, a FastAPI wrapper, and Apidog tests for transparency and seed reproducibility.

INEZA Felin-Michel

INEZA Felin-Michel

28 September 2026

How to use Qwen-Image-2.1: diffusers code, transparent output, and an API you can test

Apidog for Enterprise

On-Premises Deploy

SSO & RBAC

SOC 2 Compliant

Explore Apidog Enterprise

Qwen-Image-2.1 is the open-weights image model Alibaba released on September 20, 2026: one 7B-parameter generator that handles text-to-image, editing with up to 10 reference images, and native transparent (RGBA) output. This guide gets you from pip install to a working HTTP endpoint. It covers the four reference-code paths from the GitHub README, the settings that matter, a small FastAPI wrapper so the model is callable like any other image API, and how to test that endpoint in Apidog so prompt changes and model updates don’t break your app.

If you want the background first, What is Qwen-Image-2.1 covers the architecture and the license. The short version of the license: research and non-commercial use only unless you get a separate commercial agreement from Qwen. Everything below is fine for evaluation.

button

Before you start

Requirement Detail
Python packages torch>=2.4.0, transformers>=5.17, diffusers from GitHub main, accelerate, pillow
Pipeline class QwenImage21Pipeline (one class for generation and editing)
Weights Qwen/Qwen-Image-2.1, bf16 safetensors
GPU Not specified by Qwen; the reference code targets one CUDA device in bfloat16, with enable_model_cpu_offload() as the fallback
Default output 2048 x 2048; 40 inference steps
Optional Qwen-Image-2.1-PE-T2I / PE-I2I prompt-rewriting models

Install:

pip install "torch>=2.4.0" "transformers>=5.17" accelerate pillow
pip install git+https://github.com/huggingface/diffusers

The diffusers integration landed in a dedicated PR on launch day, so a PyPI release older than September 20 won’t have the pipeline class.

Step 1: text-to-image

import torch
from diffusers import QwenImage21Pipeline

pipe = QwenImage21Pipeline.from_pretrained(
    "Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16
).to("cuda")

image = pipe(
    prompt='A neon shop sign that reads "QWEN IMAGE 2.1", rainy night, reflections on wet pavement',
    num_inference_steps=40,
    generator=torch.Generator("cuda").manual_seed(42),
).images[0]
image.save("t2i_example.png")

Two things to notice. The prompt puts the sign text in quotes; Qwen’s text rendering is the reason people pick this line of models, and quoting the literal string is the convention from earlier releases. And the seed is explicit. Keep it that way in every request you intend to test, because a fixed seed is what makes an image endpoint reproducible enough to assert on.

Pass width and height from the supported table when you need a non-square output:

Ratio Size
1:1 2048 x 2048
4:3 / 3:4 2400 x 1792 / 1792 x 2400
3:2 / 2:3 2528 x 1696 / 1696 x 2528
16:9 / 9:16 2752 x 1536 / 1536 x 2752

Step 2: transparent output

Transparency is prompt-driven. The README’s recommended phrasing is literal, so use it:

image = pipe(
    prompt=(
        "This is an RGBA image with transparency. A cute cartoon dragon sticker. "
        "The image has alpha channel and the background is transparent."
    ),
    num_inference_steps=40,
    generator=torch.Generator("cuda").manual_seed(42),
).images[0]
image.save("transparent_example.png")

Check the result rather than trusting it:

assert image.mode == "RGBA", image.mode
alpha = image.getchannel("A")
print("transparent pixels:", sum(1 for p in alpha.getdata() if p == 0))

That assertion is the first test you’ll port into Apidog later. A model that silently returns RGB when you asked for RGBA is a bug your users will find before you do.

Step 3: editing, with one image or up to ten

The same pipeline edits when you pass image:

from PIL import Image

input_image = Image.open("input.png")
edited = pipe(
    prompt="Change the background to a sunset beach",
    image=input_image,
    num_inference_steps=40,
    generator=torch.Generator("cuda").manual_seed(42),
).images[0]
edited.save("edit_example.png")

For multiple references, pass a list (the launch post’s limit is 10):

refs = [Image.open(f"ref_{i}.png") for i in range(3)]
result = pipe(
    prompt="These three characters are sitting around a campfire in a forest",
    image=refs,
    num_inference_steps=40,
    generator=torch.Generator("cuda").manual_seed(42),
).images[0]
result.save("multi_ref_example.png")

Local editing works three ways in the launch post: colored circles you reference by color in the prompt, painted annotations, or the untouched original plus a separate mask image passed as two inputs. The README doesn’t ship a dedicated mask example, so the two-input form is the one to try first: image=[original, mask] with a prompt that describes what goes in the masked region [VERIFY against the README once a mask example lands].

Editing is also where 2.1’s speed work shows. Reference images and the instruction are static across denoising steps, so the model computes their key-value cache once and reuses it. Ten references cost noticeably less than ten times one.

Step 4: fit it on your GPU

Qwen hasn’t published VRAM numbers. If the bf16 pipeline doesn’t fit, the README offers:

pipe = QwenImage21Pipeline.from_pretrained("Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16)
pipe.enable_model_cpu_offload()

For serving, the README points at vLLM-Omni (with FP8), SGLang, and LightX2V. ComfyUI has native support with a template workflow if you’d rather not write Python at all. And if you have no GPU, the free options cover the hosted demo and Qwen Chat.

Step 5: wrap it as an HTTP API

Application code shouldn’t import diffusers. Put the pipeline behind a small service so it has a contract you can version, mock, and test. This FastAPI wrapper is about 40 lines and returns PNG bytes:

# server.py
import io, torch
from fastapi import FastAPI, UploadFile, File, Form
from fastapi.responses import Response
from PIL import Image
from diffusers import QwenImage21Pipeline

app = FastAPI()
pipe = QwenImage21Pipeline.from_pretrained(
    "Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16
).to("cuda")

SIZES = {"1:1": (2048, 2048), "16:9": (2752, 1536), "9:16": (1536, 2752)}

@app.post("/v1/images")
async def generate(
    prompt: str = Form(...),
    aspect: str = Form("1:1"),
    transparent: bool = Form(False),
    seed: int = Form(42),
    steps: int = Form(40),
    references: list[UploadFile] = File(default=[]),
):
    if transparent and not prompt.startswith("This is an RGBA image"):
        prompt = ("This is an RGBA image with transparency. " + prompt +
                  " The image has alpha channel and the background is transparent.")
    refs = [Image.open(io.BytesIO(await f.read())) for f in references[:10]]
    w, h = SIZES.get(aspect, SIZES["1:1"])
    kwargs = dict(prompt=prompt, num_inference_steps=steps,
                  generator=torch.Generator("cuda").manual_seed(seed))
    if refs:
        kwargs["image"] = refs if len(refs) > 1 else refs[0]
    else:
        kwargs.update(width=w, height=h)
    image = pipe(**kwargs).images[0]
    buf = io.BytesIO()
    image.save(buf, format="PNG")
    return Response(buf.getvalue(), media_type="image/png",
                    headers={"X-Image-Mode": image.mode, "X-Seed": str(seed)})

Run it with uvicorn server:app --port 8000. The two response headers, X-Image-Mode and X-Seed, exist so a test can check transparency and reproducibility without decoding the PNG. That’s the only product-specific choice in the wrapper; the rest is a plain multipart endpoint.

Step 6: test the endpoint in Apidog

Now it’s an API, and the same discipline you’d apply to the gpt-image-2.5 API or the Nano Banana 2 API applies here. In Apidog:

  1. Create the endpoint as POST {{base_url}}/v1/images with a multipart body: prompt, aspect, transparent, seed, steps, and a repeatable references file field. Put base_url in an environment so the same collection points at your laptop, the GPU box, or a mock.
  2. Send a text-to-image request with seed=42 and the neon-sign prompt. Confirm 200, Content-Type: image/png, and X-Image-Mode: RGB.
  3. Add assertions in the post-processor: status is 200, X-Image-Mode equals RGBA when transparent=true, response body size is above a floor (a 2K PNG that comes back at 2 KB is a blank image), and X-Seed echoes what you sent.
  4. Send the transparency case and a three-reference edit the same way, attaching the images in the file field. Save each as a test case.
  5. Run them as a test scenario on a schedule or in CI. When you swap in a quantized build or a future 2.2, the suite tells you in minutes whether transparency still works and whether the seed still reproduces.
  6. Mock it while the GPU is busy. Apidog’s smart mock returns a canned PNG for the same contract, so the frontend keeps building.

Because Apidog also generates the OpenAPI spec and docs from the endpoint you defined, the wrapper’s contract becomes shareable the moment it works. Download Apidog and import the endpoint above to start.

Optional: prompt rewriting with PE-T2I

The demo Space turns one-line requests into long structured prompts using Qwen-Image-2.1-PE-T2I, a fine-tuned Qwen3.5-VL 9B that returns JSON with an expanded English prompt and a recommended aspect ratio. Run it as a second service in front of /v1/images, or skip it and write full prompts yourself. If you add it, test it separately: it’s a text API with a JSON contract, and a broken rewriter produces bad images that look like a generator bug.

FAQ

Does one pipeline do both generation and editing? Yes. QwenImage21Pipeline generates when called with a prompt alone and edits when you pass image (a single PIL image or a list of up to 10).

How do I get a transparent PNG? Start the prompt with “This is an RGBA image with transparency” and say the background is transparent. Check image.mode == "RGBA" on the result.

What are the recommended settings? 40 inference steps and bfloat16, per the README. Guidance values aren’t listed for 2.1; earlier Qwen-Image releases used true_cfg_scale=4.0, so try that if outputs look under-guided [VERIFY].

Can I use this in a commercial product? Not under the default license. Qwen-Image-2.1 ships under the Qwen Research License; commercial use requires a separate license from Qwen. Details in What is Qwen-Image-2.1.

Is there a hosted API instead? Qwen Image 3.0 and 3.0 Pro are Alibaba’s hosted image models with per-image pricing. The 2.1 vs 3.0 comparison covers when to self-host and when to rent.

Where to go next

You now have four working calls, a wrapper with a stable contract, and a test suite that checks the two properties that matter most for this model, transparency and reproducibility. Next, decide whether the research license fits your use, or whether the hosted 3.0 API is the better fit, and keep both behind the same Apidog collection so switching is a base URL change, not a rewrite.

Explore more

GPT-6.1 Sol vs Claude Sonnet 5.5: same $2/$10 price, launched a day apart, and no shared benchmark

GPT-6.1 Sol vs Claude Sonnet 5.5: same $2/$10 price, launched a day apart, and no shared benchmark

GPT-6.1 Sol vs Claude Sonnet 5.5: same $2/$10 price, launched a day apart, no shared benchmark. Specs, vendor claims, and how to test both yourself.

30 September 2026

Is GPT-6.1 Sol Free?

Is GPT-6.1 Sol Free?

Is GPT-6.1 Sol free? No: it's for Plus and up in ChatGPT Work and Codex, and the API has no free tier. Here are the cheapest routes, with real cost math.

30 September 2026

How to Use Claude Sonnet 5.5 for Free: Every Route That Works (and the Ones That Don't)

How to Use Claude Sonnet 5.5 for Free: Every Route That Works (and the Ones That Don't)

Is Claude Sonnet 5.5 free? Yes on Claude.ai (web, iOS, Android). Every free route checked, plus what isn't: Claude Code, the API, and Copilot Free.

29 September 2026

Practice API Design-first in Apidog

Discover an easier way to build and use APIs

How to use Qwen-Image-2.1: diffusers code, transparent output, and an API you can test