Qwen-Image-2.1 is the open-weights image model Alibaba released on September 20, 2026: one 7B-parameter generator that handles text-to-image, editing with up to 10 reference images, and native transparent (RGBA) output. This guide gets you from pip install to a working HTTP endpoint. It covers the four reference-code paths from the GitHub README, the settings that matter, a small FastAPI wrapper so the model is callable like any other image API, and how to test that endpoint in Apidog so prompt changes and model updates don’t break your app.
If you want the background first, What is Qwen-Image-2.1 covers the architecture and the license. The short version of the license: research and non-commercial use only unless you get a separate commercial agreement from Qwen. Everything below is fine for evaluation.
Before you start
| Requirement | Detail |
|---|---|
| Python packages | torch>=2.4.0, transformers>=5.17, diffusers from GitHub main, accelerate, pillow |
| Pipeline class | QwenImage21Pipeline (one class for generation and editing) |
| Weights | Qwen/Qwen-Image-2.1, bf16 safetensors |
| GPU | Not specified by Qwen; the reference code targets one CUDA device in bfloat16, with enable_model_cpu_offload() as the fallback |
| Default output | 2048 x 2048; 40 inference steps |
| Optional | Qwen-Image-2.1-PE-T2I / PE-I2I prompt-rewriting models |
Install:
pip install "torch>=2.4.0" "transformers>=5.17" accelerate pillow
pip install git+https://github.com/huggingface/diffusers
The diffusers integration landed in a dedicated PR on launch day, so a PyPI release older than September 20 won’t have the pipeline class.
Step 1: text-to-image
import torch
from diffusers import QwenImage21Pipeline
pipe = QwenImage21Pipeline.from_pretrained(
"Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16
).to("cuda")
image = pipe(
prompt='A neon shop sign that reads "QWEN IMAGE 2.1", rainy night, reflections on wet pavement',
num_inference_steps=40,
generator=torch.Generator("cuda").manual_seed(42),
).images[0]
image.save("t2i_example.png")
Two things to notice. The prompt puts the sign text in quotes; Qwen’s text rendering is the reason people pick this line of models, and quoting the literal string is the convention from earlier releases. And the seed is explicit. Keep it that way in every request you intend to test, because a fixed seed is what makes an image endpoint reproducible enough to assert on.
Pass width and height from the supported table when you need a non-square output:
| Ratio | Size |
|---|---|
| 1:1 | 2048 x 2048 |
| 4:3 / 3:4 | 2400 x 1792 / 1792 x 2400 |
| 3:2 / 2:3 | 2528 x 1696 / 1696 x 2528 |
| 16:9 / 9:16 | 2752 x 1536 / 1536 x 2752 |
Step 2: transparent output
Transparency is prompt-driven. The README’s recommended phrasing is literal, so use it:
image = pipe(
prompt=(
"This is an RGBA image with transparency. A cute cartoon dragon sticker. "
"The image has alpha channel and the background is transparent."
),
num_inference_steps=40,
generator=torch.Generator("cuda").manual_seed(42),
).images[0]
image.save("transparent_example.png")
Check the result rather than trusting it:
assert image.mode == "RGBA", image.mode
alpha = image.getchannel("A")
print("transparent pixels:", sum(1 for p in alpha.getdata() if p == 0))
That assertion is the first test you’ll port into Apidog later. A model that silently returns RGB when you asked for RGBA is a bug your users will find before you do.
Step 3: editing, with one image or up to ten
The same pipeline edits when you pass image:
from PIL import Image
input_image = Image.open("input.png")
edited = pipe(
prompt="Change the background to a sunset beach",
image=input_image,
num_inference_steps=40,
generator=torch.Generator("cuda").manual_seed(42),
).images[0]
edited.save("edit_example.png")
For multiple references, pass a list (the launch post’s limit is 10):
refs = [Image.open(f"ref_{i}.png") for i in range(3)]
result = pipe(
prompt="These three characters are sitting around a campfire in a forest",
image=refs,
num_inference_steps=40,
generator=torch.Generator("cuda").manual_seed(42),
).images[0]
result.save("multi_ref_example.png")
Local editing works three ways in the launch post: colored circles you reference by color in the prompt, painted annotations, or the untouched original plus a separate mask image passed as two inputs. The README doesn’t ship a dedicated mask example, so the two-input form is the one to try first: image=[original, mask] with a prompt that describes what goes in the masked region [VERIFY against the README once a mask example lands].
Editing is also where 2.1’s speed work shows. Reference images and the instruction are static across denoising steps, so the model computes their key-value cache once and reuses it. Ten references cost noticeably less than ten times one.
Step 4: fit it on your GPU
Qwen hasn’t published VRAM numbers. If the bf16 pipeline doesn’t fit, the README offers:
pipe = QwenImage21Pipeline.from_pretrained("Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16)
pipe.enable_model_cpu_offload()
For serving, the README points at vLLM-Omni (with FP8), SGLang, and LightX2V. ComfyUI has native support with a template workflow if you’d rather not write Python at all. And if you have no GPU, the free options cover the hosted demo and Qwen Chat.
Step 5: wrap it as an HTTP API
Application code shouldn’t import diffusers. Put the pipeline behind a small service so it has a contract you can version, mock, and test. This FastAPI wrapper is about 40 lines and returns PNG bytes:
# server.py
import io, torch
from fastapi import FastAPI, UploadFile, File, Form
from fastapi.responses import Response
from PIL import Image
from diffusers import QwenImage21Pipeline
app = FastAPI()
pipe = QwenImage21Pipeline.from_pretrained(
"Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16
).to("cuda")
SIZES = {"1:1": (2048, 2048), "16:9": (2752, 1536), "9:16": (1536, 2752)}
@app.post("/v1/images")
async def generate(
prompt: str = Form(...),
aspect: str = Form("1:1"),
transparent: bool = Form(False),
seed: int = Form(42),
steps: int = Form(40),
references: list[UploadFile] = File(default=[]),
):
if transparent and not prompt.startswith("This is an RGBA image"):
prompt = ("This is an RGBA image with transparency. " + prompt +
" The image has alpha channel and the background is transparent.")
refs = [Image.open(io.BytesIO(await f.read())) for f in references[:10]]
w, h = SIZES.get(aspect, SIZES["1:1"])
kwargs = dict(prompt=prompt, num_inference_steps=steps,
generator=torch.Generator("cuda").manual_seed(seed))
if refs:
kwargs["image"] = refs if len(refs) > 1 else refs[0]
else:
kwargs.update(width=w, height=h)
image = pipe(**kwargs).images[0]
buf = io.BytesIO()
image.save(buf, format="PNG")
return Response(buf.getvalue(), media_type="image/png",
headers={"X-Image-Mode": image.mode, "X-Seed": str(seed)})
Run it with uvicorn server:app --port 8000. The two response headers, X-Image-Mode and X-Seed, exist so a test can check transparency and reproducibility without decoding the PNG. That’s the only product-specific choice in the wrapper; the rest is a plain multipart endpoint.
Step 6: test the endpoint in Apidog
Now it’s an API, and the same discipline you’d apply to the gpt-image-2.5 API or the Nano Banana 2 API applies here. In Apidog:
- Create the endpoint as
POST {{base_url}}/v1/imageswith a multipart body:prompt,aspect,transparent,seed,steps, and a repeatablereferencesfile field. Putbase_urlin an environment so the same collection points at your laptop, the GPU box, or a mock. - Send a text-to-image request with
seed=42and the neon-sign prompt. Confirm200,Content-Type: image/png, andX-Image-Mode: RGB. - Add assertions in the post-processor: status is 200,
X-Image-ModeequalsRGBAwhentransparent=true, response body size is above a floor (a 2K PNG that comes back at 2 KB is a blank image), andX-Seedechoes what you sent. - Send the transparency case and a three-reference edit the same way, attaching the images in the file field. Save each as a test case.
- Run them as a test scenario on a schedule or in CI. When you swap in a quantized build or a future 2.2, the suite tells you in minutes whether transparency still works and whether the seed still reproduces.
- Mock it while the GPU is busy. Apidog’s smart mock returns a canned PNG for the same contract, so the frontend keeps building.
Because Apidog also generates the OpenAPI spec and docs from the endpoint you defined, the wrapper’s contract becomes shareable the moment it works. Download Apidog and import the endpoint above to start.
Optional: prompt rewriting with PE-T2I
The demo Space turns one-line requests into long structured prompts using Qwen-Image-2.1-PE-T2I, a fine-tuned Qwen3.5-VL 9B that returns JSON with an expanded English prompt and a recommended aspect ratio. Run it as a second service in front of /v1/images, or skip it and write full prompts yourself. If you add it, test it separately: it’s a text API with a JSON contract, and a broken rewriter produces bad images that look like a generator bug.
FAQ
Does one pipeline do both generation and editing? Yes. QwenImage21Pipeline generates when called with a prompt alone and edits when you pass image (a single PIL image or a list of up to 10).
How do I get a transparent PNG? Start the prompt with “This is an RGBA image with transparency” and say the background is transparent. Check image.mode == "RGBA" on the result.
What are the recommended settings? 40 inference steps and bfloat16, per the README. Guidance values aren’t listed for 2.1; earlier Qwen-Image releases used true_cfg_scale=4.0, so try that if outputs look under-guided [VERIFY].
Can I use this in a commercial product? Not under the default license. Qwen-Image-2.1 ships under the Qwen Research License; commercial use requires a separate license from Qwen. Details in What is Qwen-Image-2.1.
Is there a hosted API instead? Qwen Image 3.0 and 3.0 Pro are Alibaba’s hosted image models with per-image pricing. The 2.1 vs 3.0 comparison covers when to self-host and when to rent.
Where to go next
You now have four working calls, a wrapper with a stable contract, and a test suite that checks the two properties that matter most for this model, transparency and reproducibility. Next, decide whether the research license fits your use, or whether the hosted 3.0 API is the better fit, and keep both behind the same Apidog collection so switching is a base URL change, not a rewrite.



