Skip to content

Z-Image

Z-Image is a 6B-parameter text-to-image family from Tongyi-MAI that understands both English and Chinese prompts. It encodes prompts with a Qwen3 4B text encoder and decodes with the same 16-channel VAE as FLUX.1. InvokeAI supports two variants:

  • Z-Image Turbo is distilled for fast, low-step generation with CFG disabled (CFG Scale 1.0). This is the checkpoint the starter models install.
  • Z-Image Base is the undistilled foundation model. It needs more steps, and supports CFG and negative prompts.

A Diffusers pipeline’s variant is detected on install. Single-file, GGUF and SDNQ files are registered as Turbo. If you install a Base checkpoint in one of those formats, change its Variant to Z-Image Base in the Model Manager so the right defaults and schedulers apply.

Z-Image works in the Generate tab, on the Canvas (text-to-image, image-to-image, inpainting and outpainting) and in the workflow editor.

Z-Image Turbo, Z-Image (Base) and their Qwen3 text encoder are released under Apache 2.0, which allows commercial use. None of the downloads is gated.

This is a summary, not legal advice. Read the full license on the model’s HuggingFace page before using a model commercially.

See the System Requirements table. In short: a Q4_K GGUF or NVFP4 transformer runs in about 8 GB of VRAM, while the Q8 and full-precision (BF16) builds want 16 GB or more. The full Diffusers pipeline is roughly a 33 GB download.

For more headroom you can enable FP8 Storage on a full-precision Z-Image model, use an SDNQ build, or turn on Low-VRAM mode.

The easiest path is the Z-Image Turbo bundle in the Model Manager. It installs the Q4_K GGUF transformer, the quantized Qwen3 encoder, the FLUX VAE and both Z-Image ControlNets.

Z-Image needs three components:

ComponentDiffusers / SDNQ pipelineGGUF / single-file install
Transformerbundledthe .gguf / .safetensors file
Text encoder (Qwen3 4B)bundledinstalled separately
VAE (FLUX.1)bundledinstalled separately (any FLUX VAE)

The starter models:

StarterFormatSizeAlso installs
Z-Image TurboDiffusers pipeline~33 GBnothing, everything is bundled
Z-Image Turbo (quantized)GGUF Q4_K~4 GBQwen3 encoder (GGUF Q6_K, ~3.3 GB) + FLUX VAE
Z-Image Turbo (Q8)GGUF Q8_0~6.6 GBQwen3 encoder (GGUF Q6_K) + FLUX VAE
Z-Image Turbo (NVFP4)Comfy-Org nvfp4 single file~4.5 GBQwen3 encoder (FP4 mixed, ~3.5 GB) + FLUX VAE
Z-Image Turbo (SDNQ uint4 + SVD)SDNQ Diffusers pipeline~5 GBnothing, everything is bundled

A full-precision Z-Image Qwen3 Text Encoder (~8 GB) is also available as a separate starter.

The NVFP4 build keeps its quantized layers packed in memory and decodes them on the fly, so it needs about a third of the full-precision transformer’s memory.

When the selected Z-Image model is not a self-contained pipeline, the Components section next to the model shows three slots:

  • Component source: a Diffusers (or SDNQ) Z-Image pipeline to borrow the encoder and VAE from.
  • Qwen3 Encoder: a standalone Qwen3 4B encoder.
  • VAE: Z-Image decodes with the FLUX VAE, so FLUX-base VAEs are listed here.

You need either a component source or both a Qwen3 encoder and a VAE before you can generate.

Selecting a Z-Image model applies these defaults:

VariantStepsCFG ScaleSchedulerSize
Z-Image Turbo91.0Euler1024×1024
Z-Image Base504.0Euler1024×1024
  • CFG Scale: 1.0 turns CFG off, which is what Turbo is trained for. Values below 1.0 are not accepted.
  • Negative prompt: only used when CFG Scale is above 1.0. With Turbo at 1.0 it has no effect.
  • Schedulers: Euler (the default and recommended), Heun (2nd order) (better quality, about twice as slow) and LCM. LCM works with Turbo only and is not offered for Z-Image Base.
  • Resolution: width and height must be multiples of 16.
  • Prompt weighting: Compel-style prompt syntax does not apply to Z-Image. Write prompts in plain language.

In the workflow editor, Denoise - Z-Image also has a Shift input that overrides the timestep shift. Leave it blank to have it calculated from the image size, which is the recommended setting.

Two Z-Image ControlNets are available as starters (both are part of the bundle):

  • Z-Image ControlNet Union handles Canny, HED, Depth, Pose and MLSD control images.
  • Z-Image ControlNet Tile is useful for upscaling and adding detail.

On the Canvas, add a Control Layer while a Z-Image model is selected. The layer uses a Z-Image ControlNet with a default weight of 0.75. The weight ranges from 0 to 2, and 0.65 to 0.80 is the recommended range. Only one Z-Image control layer can be used per generation.

In the workflow editor, the same thing is the Z-Image ControlNet node feeding the Control input of Denoise - Z-Image. That node also needs the VAE input connected, which it uses to encode the control image.

Z-Image supports Regional Guidance layers on the Canvas with positive prompts. Regional negative prompts, auto-negative and regional reference images are not supported for Z-Image.

In the workflow editor, Prompt - Z-Image takes an optional mask, and Denoise - Z-Image accepts either one conditioning or a collection of them.

Z-Image LoRAs are supported in Kohya, PEFT/Diffusers and LyCORIS formats. They are applied to both the transformer and the Qwen3 encoder. In the workflow editor, use Apply LoRA - Z-Image or Apply LoRA Collection - Z-Image.

Z-Image Turbo tends to produce similar images across different seeds. The Seed Variance Enhancer - Z-Image node, available in the workflow editor only, adds reproducible, seed-based noise to the prompt conditioning to increase variety:

  • Strength: 0 is off, 0.1 (the default) is subtle, 0.5 is strong.
  • Randomize Percent: the share of embedding values that receive noise (default 50).

Connect it between Prompt - Z-Image and the denoise node’s positive conditioning.

Z-Image has no PiD decoder of its own and uses the FLUX PiD decoders, since it shares FLUX.1’s VAE. See PiD Decode.

This site was designed and developed by Aether Fox Studio.