Skip to content

Qwen Image

Qwen Image is a diffusion-transformer image family from the Qwen team that encodes prompts with the Qwen2.5-VL 7B vision-language model and decodes with a 16-channel VAE. InvokeAI supports two variants:

  • Qwen Image (e.g. Qwen Image 2512) — text-to-image generation.
  • Qwen Image Edit (e.g. Qwen Image Edit 2511) — text-guided editing driven by one or more reference images. Because its text encoder is a vision-language model, it reads the reference images alongside your instruction (“replace the background with a beach”, “put this jacket on the person in image 2”).

The variant is detected automatically on install. Both variants can also run text-to-image, image-to-image, inpainting and outpainting on the Canvas; only the Edit variant accepts reference images.

Qwen Image 2512, Qwen Image Edit 2511, the Lightning LoRAs and the Qwen2.5-VL text encoder are released under Apache 2.0, which allows commercial use. None of the downloads is gated.

This is a summary, not legal advice. Read the full license on the model’s HuggingFace page before using a model commercially.

Qwen Image is a large model: the full Diffusers pipeline is a ~40 GB download. To run it on a consumer GPU, use a GGUF transformer together with the standalone VAE and fp8 text encoder:

Transformer buildDownload
Q2_K~7.5 GB
Q4_K_M~13 GB
Q6_K~17 GB
Q8_0~22 GB

The Qwen2.5-VL encoder is loaded in addition to the transformer (~7 GB for the fp8 build). If your GPU cannot hold both, enable Low-VRAM mode. Single-file checkpoints that ship scaled fp8 weights can be kept in fp8 with FP8 Storage.

The easiest path is the Qwen Image bundle in the Model Manager’s Starter Models. It installs the standalone VAE and fp8 encoder, the Diffusers and the Q4_K_M / Q8_0 GGUF builds of both Qwen Image 2512 and Qwen Image Edit 2511, and the Lightning LoRAs for each. The other GGUF builds (Q2_K, Q6_K) are available as individual starter models.

Qwen Image needs three components:

ComponentDiffusers installGGUF / single-file install
Transformerbundled in the pipelinethe .gguf / .safetensors file
VAE (Qwen Image VAE)bundledinstalled separately
Text encoder (Qwen2.5-VL)bundledinstalled separately
  • Diffusers (Qwen/Qwen-Image-2512, Qwen/Qwen-Image-Edit-2511): a single ~40 GB install that includes everything.
  • GGUF (from unsloth/Qwen-Image-2512-GGUF and unsloth/Qwen-Image-Edit-2511-GGUF): the file contains only the transformer. The GGUF starter models install the Qwen Image VAE (~250 MB) and the Qwen2.5-VL Encoder (fp8 scaled) (~7 GB) as dependencies, so you never need the full pipeline.
  • Single-file .safetensors checkpoints (bf16/fp16, and ComfyUI-style fp8-scaled or nvfp4 builds) can be installed by path or URL; like GGUF, they need a standalone VAE and encoder.

Three Qwen2.5-VL encoder builds are available as starter models: fp8 scaled (~7 GB, the one the bundle installs), NVFP4 (~5.7 GB download, about 6.7 GB once loaded), and Diffusers (~16 GB, full precision). The single-file encoders ship with their tokenizer and config inside InvokeAI.

When a GGUF or single-file Qwen Image model is selected, pick the Qwen VL Encoder and VAE in the Components section next to the model. Alternatively, set Component source to an installed Diffusers Qwen Image model and InvokeAI takes the missing VAE and encoder from it. A Diffusers main model needs neither.

Selecting a Qwen Image model applies these defaults:

  • Steps: 40
  • CFG Scale: 4
  • Size: 1024×1024; width and height snap to multiples of 16.

There is no scheduler choice — Qwen Image always uses its own flow-matching schedule, so the Scheduler control is hidden.

The negative prompt is used only when CFG Scale is above 1; at CFG 1 it is ignored.

The Qwen Image Lightning and Qwen Image Edit Lightning starter LoRAs are distillation LoRAs that cut generation to 4 or 8 steps. Use the one that matches your model variant, and set:

  • Steps: 4 or 8, matching the LoRA
  • CFG Scale: 1

The LoRAs are trained for a fixed schedule shift of 3. In the workflow editor, set Shift on the Denoise - Qwen Image node to 3.0 when using them; leave it empty for the base model’s default schedule. The Generate tab does not expose this setting.

With a Qwen Image Edit model selected, add up to five reference images, then describe the change you want in the prompt. The text-to-image variant does not accept reference images.

  • All reference images are given to the Qwen2.5-VL encoder together with your prompt, so you can refer to them in the instruction.
  • The first reference image is also encoded by the VAE and supplied to the transformer as a pixel-level reference.

In the workflow editor, reference images go into the Reference Images input of Prompt - Qwen Image; connect an Image to Latents - Qwen Image node to the Reference Latents input of Denoise - Qwen Image to add the pixel-level reference.

Qwen Image LoRAs (diffusers PEFT, Kohya and LoKR formats) are supported and apply to the transformer. In the workflow editor use Apply LoRA - Qwen Image or Apply LoRA Collection - Qwen Image.

Qwen Image works with PiD Decode, which replaces the VAE decode with a 4× super-resolving pixel diffusion. The PiD Decoder Qwen-Image (2K to 4K) and PiD 1.5 Decoder Qwen-Image (2K to 4K) decoders (including an int8 build of the latter) are starter models.

NodePurpose
Main Model - Qwen ImageLoads the transformer, VAE and Qwen VL encoder (or a component source)
Prompt - Qwen ImageEncodes the prompt and any reference images
Denoise - Qwen ImageRuns sampling; accepts reference latents, img2img latents and masks
Image to Latents - Qwen ImageVAE encode
Latents to Image - Qwen ImageVAE decode
Latents to Image - Qwen-Image + PiD (4x SR)PiD super-resolution decode

The Prompt - Qwen Image node also has a Quantization option (none, int8, nf4) that quantizes the Qwen2.5-VL encoder on load to reduce its VRAM use.

This site was designed and developed by Aether Fox Studio.