Skip to content

Model Formats and Families

There are a bewildering number of model families, types and formats. This page provides a brief overview of the models that InvokeAI supports: the formats they come in, the components that make up a model, and the families of models, each with its own page under Local Models.

Model weights are distributed in two main layouts, and InvokeAI installs both:

  • Diffusers — A folder, usually a HuggingFace repository, laid out the way the HuggingFace diffusers library expects. Each component of the model (the denoiser, the VAE, the text encoders, the tokenizers and the scheduler settings) lives in its own subfolder with its own configuration file. A Diffusers install is self-contained: everything the model needs comes with it. Install one by pasting its repository ID, such as stabilityai/stable-diffusion-xl-base-1.0, into the Model Manager.
  • Checkpoint (single file) — One file, usually .safetensors (older models may use .ckpt or .pt), holding the weights of one or more components. This is the usual format on Civitai. Single-file checkpoints for Stable Diffusion 1.x and SDXL generally include the VAE and text encoders. Checkpoints for most newer architectures contain only the transformer, so you also have to install the VAE and text encoder that go with them, either separately or as part of the family’s starter bundle.

Newer models are often distributed in quantized forms, which store the weights at reduced precision to save disk space and VRAM. InvokeAI recognizes several of these, including GGUF files (named for their quantization level, such as Q4_K_M or Q8_0), fp8, int8 and nf4 builds. Like other single-file checkpoints, a quantized file usually carries only the transformer and needs separately installed companions. Lower precision trades a little image quality for a large saving in memory; each model family’s page lists the builds it supports.

InvokeAI recognizes the format from the file’s contents, so you install them like any other model:

  • GGUF (.gguf, llama.cpp quantization such as Q4_K or Q8_0): weights stay packed in memory and are unpacked layer by layer during generation. Supported for the FLUX.1, FLUX.2, Z-Image, Krea-2, Qwen-Image and Wan transformers, and for the T5, Qwen3, Qwen3-VL, Mistral and Gemma 2 text encoders.
  • NF4 (bitsandbytes): 4-bit, CUDA only.
  • SDNQ: see SDNQ Quantization.
  • int8 (int8_convrot, from ComfyUI repackages): the weights stay int8 in memory on every device, including Apple Silicon, and need no setting to turn on. Supported for FLUX.1, FLUX.2 Klein, Z-Image, Krea-2, Ideogram 4, MiniMax H3, LTX-2 and the PiD decoders, and for the Qwen3 8B, Qwen3-VL and LTX-2 Gemma encoders.
  • nvfp4 (4-bit, in ComfyUI’s checkpoint format, such as Comfy-Org’s builds; other nvfp4 exports such as AWQ are refused): the weights stay packed in memory and are unpacked per layer, so no special GPU is needed. Supported for the Z-Image, Krea-2, Qwen-Image and LTX-2 transformers and for the Qwen3, Qwen2.5-VL and FLUX.2 Mistral encoders; the Ideogram 4 nvfp4 files are refused at install.
  • fp8 (plain or fp8_scaled): InvokeAI keeps the weights in fp8 when FP8 Storage is on, which installation does for you, or when fp8_compute is enabled.
  • MXFP8 (fp8 with one scale per block of 32 weights): loaded by unpacking to BF16. Turn FP8 Storage off for these; see the FP8 Storage caution.

Text encoders for Qwen-Image, Z-Image, Krea-2 and Ideogram 4 bring their tokenizer and configuration with InvokeAI, so single-file and GGUF encoders load without network access.

A single-file Stable Diffusion 1.x, 2.x or SDXL model can be converted to Diffusers format with the Convert to diffusers action in the model’s detail view. Conversion replaces the single file with an equivalent Diffusers folder. Converting is optional: InvokeAI runs the single-file checkpoint directly.

What we call “a model” is really a pipeline of several cooperating networks, called submodels or components. The prompt is split into tokens and turned into numbers by the text encoder. The denoiser then starts from random noise and, over a series of steps, gradually shapes it into an image guided by those numbers. It works in a compressed “latent” space rather than on pixels, so at the end the VAE decodes the result into the image you see.

A Diffusers install carries all of these components together. With single-file and quantized installs you often pick some of them yourself. They appear as separate models in the Model Manager, and generation panels show slots for them (for example VAE, Text encoder or Model components) when the selected model needs them. Several families share components: FLUX.2 and Ideogram 4 use the same VAE, for example, so one download can serve both.

ComponentWhat it does
Denoiser (UNet or transformer)The core of the model and by far the largest part. At each step it predicts how to remove noise from the latent image, steered by the encoded prompt. Older models (SD 1.x, SDXL) use a convolutional UNet; newer ones use a diffusion transformer (DiT). Some models use two: the Wan 2.2 A14B models hand off from a high-noise expert to a low-noise expert partway through, and Ideogram 4 runs a conditional and an unconditional transformer side by side.
VAE (variational autoencoder)Translates between pixels and latents. The encoder half compresses an input image (for image-to-image, inpainting or a video’s first frame) into latents; the decoder half turns the finished latents back into pixels. A different VAE can change color and fine detail.
TokenizerSplits the prompt into tokens, the word pieces the text encoder understands. It is a small vocabulary file rather than a neural network, and it always comes with its text encoder.
Text encoderTurns the tokens into embeddings that tell the denoiser what to draw. Older models use CLIP; newer ones use large language models such as T5, Qwen3, Mistral or Gemma, which understand long, detailed prompts much better. These can be as large as the denoiser, and several families let you run them from a separately installed, quantized file. SDXL and FLUX.1 combine two encoders.
Image encoderTurns a reference image into embeddings the denoiser can use, the way a text encoder does for text. Used by IP-Adapters and FLUX Redux (CLIP Vision or SigLIP) to carry the style or content of a reference image. Vision-language text encoders such as Qwen-VL read reference images directly, which is how image-editing models like Qwen Image Edit see the image they edit.
SchedulerThe algorithm that decides how much noise to remove at each step. It is configuration rather than a network; you can pick a different scheduler for many models in the generation settings.
Audio VAEVideo-with-audio models only (LTX-2, MiniMax H3). The audio counterpart of the VAE: it encodes and decodes the latent form of the soundtrack that is denoised alongside the video.
VocoderLTX-2 only. Converts the decoded audio spectrogram into a playable waveform. MiniMax H3’s audio VAE produces a waveform directly and needs no vocoder.
ConnectorsLTX-2 only. A small network that adapts the Gemma text encoder’s output, taken from every layer, into the form the LTX-2 transformer expects.
Latent upsamplerLTX-2 only. Enlarges the latent video between the two passes of two-stage generation, so the second pass can refine at higher resolution.

Every main model belongs to a family, also called its base or architecture. The family determines which components, LoRAs, ControlNets and other add-ons work with the model: a LoRA trained for SDXL cannot be used with FLUX.1, for example. InvokeAI detects the family when you install a model, shows it as a colored badge in the Model Manager, and only offers add-ons that match the selected main model.

Families differ widely in size, speed and specialty. Older families such as SD 1.5 and SDXL run on modest hardware and have huge libraries of community fine-tunes, LoRAs and ControlNets. Newer families follow long prompts more faithfully and render text better, but need more VRAM. Many newer families also come in distilled “Turbo” or “Lightning” variants that trade some flexibility for much faster generation.

FamilyGeneratesHighlightsLicenseDetails
Stable Diffusion 1.x / 2.xImagesThe original open model family. Small and fast, with the largest library of fine-tunes, LoRAs, ControlNets and IP-Adapters. Best at 512×512.OpenRAIL-M; varies by fine-tuneSD 1.5 / 2.x
SDXLImagesHigher-resolution successor to SD 1.5 (1024×1024) with a similarly rich ecosystem of add-ons. An optional refiner model, used in the workflow editor, polishes fine detail.OpenRAIL++-M; varies by fine-tuneSDXL
Stable Diffusion 3.5ImagesMedium and Large transformer models from Stability AI, with three text encoders.Community, free under $1M revenueSD 3.5
FLUX.1ImagesDev, Schnell, Kontext (image editing) and Krea variants, plus Fill (inpainting, workflow editor only) and Redux, with ControlNets and IP-Adapters.schnell: Apache 2.0; others non-commercialFLUX.1
FLUX.2ImagesKlein 4B and 9B, and the much larger [dev]. Generates and edits images from reference images.Klein 4B: Apache 2.0; Klein 9B and [dev] non-commercialFLUX.2
CogView4ImagesTHUDM’s 6B text-to-image model with a GLM text encoder. Basic support, without LoRAs, ControlNets or reference images.Apache 2.0CogView4
Z-ImageImages6B model in a fast, few-step Turbo variant and a Base variant. Understands English and Chinese prompts, with ControlNets.Apache 2.0Z-Image
ERNIE-ImageImagesBaidu’s 8B text-to-image model, in standard and distilled Turbo variants.Apache 2.0ERNIE-Image
Qwen ImageImagesText-to-image (2512) and instruction-based image editing (Edit 2511), with fast Lightning variants.Apache 2.0Qwen Image
Krea-2Images12B model in fast Turbo and undistilled Raw variants, with training-free style reference.Community, free under $1M revenueKrea-2
Ideogram 4ImagesStructured prompts that place described elements in regions of the image.Non-commercialIdeogram 4
AnimaImagesCompact 2B anime-focused model, with ControlNet-LLLite adapters for inpainting and sketch guidance.Non-commercialAnima
Wan 2.2Video, imagesText-to-video and image-to-video, including first/last-frame interpolation. Can also generate single images.Apache 2.0Wan 2.2
LTX-2Video with audioVideo with a synchronized soundtrack, audio-to-video and video-to-audio, keyframes and extension.Community, free under $10M revenueLTX-2
MiniMax H3Video with audioVideo with a stereo soundtrack, from first/last frames or from up to 12 reference images and videos.Community, with conditionsMiniMax H3

The video families are described in the Video Generation section, and are also listed under Local Models in the sidebar.

This site was designed and developed by Aether Fox Studio.