Skip to content

Stable Diffusion 1.5 / 2.x

Stable Diffusion 1.x and 2.x are the original latent-diffusion UNet models. They are small and fast by current standards, run on modest hardware, and have the largest library of community fine-tunes, LoRAs, embeddings and ControlNets.

  • SD 1.x (in practice, SD 1.5 and its fine-tunes) is trained at 512×512 and uses a single CLIP text encoder. It is fully supported: text-to-image, image-to-image, inpainting and outpainting, control layers, reference images, regional guidance and upscaling.
  • SD 2.x is treated as legacy. Models you install yourself load and generate, but InvokeAI ships no SD 2.x starter models and some Canvas features are not available for it (see SD 2.x).

InvokeAI identifies the family, the variant (normal or inpainting; SD 2.x also has a depth variant) and the prediction type (epsilon or v-prediction) automatically on install.

ModelLicenseCommercial use
SD 1.5 base, Dreamshaper 8, CyberRealistic v4.1, ReV AnimatedCreativeML OpenRAIL-MAllowed, subject to the license’s use restrictions
Deliberate v5CC BY-NC-ND 4.0Companies and services must contact the author; the author places no restrictions on private individuals

None of the starter downloads is gated.

This is a summary, not legal advice. Read the full license on the model’s HuggingFace page before using a model commercially.

SD 1.5 is the lightest local image model InvokeAI supports: the System Requirements table lists 4 GB VRAM and 8 GB RAM as the minimum at 512×512. On smaller cards, see Low-VRAM mode.

The easiest path is the Stable Diffusion 1.5 bundle in the Model Manager’s starter models. It installs:

  • Dreamshaper 8 — a popular, versatile main model
  • EasyNegative — a textual-inversion embedding for the negative prompt
  • Three IP-Adapters for reference images — Standard, Precise and Face — plus their image encoder
  • A set of ControlNets: Hard Edge Detection (canny), Inpainting, Line Drawing (mlsd), Depth Map, Lighting Detection (Normals), Segmentation Map, Lineart, Lineart Anime, Pose Detection (openpose), Contour Detection (scribble), Soft Edge Detection (softedge), Remix (shuffle) and Tile
  • SwinIR, a 4× upscaler used by the Upscale tool

Other SD 1.5 starter models can be installed individually:

Starter modelNotes
CyberRealistic v4.1Photorealistic; installs its companion CyberRealistic Negative v3 embedding
ReV AnimatedFantasy and anime styles
Dreamshaper 8 (inpainting)Inpainting version of Dreamshaper 8
Deliberate v5Versatile general-purpose model
Deliberate v5 (inpainting)Inpainting version of Deliberate v5
QRCode Monster v2 (SD1.5)ControlNet for scannable, stylized QR codes
T2I-Adapters: Hard Edge Detection (canny), Sketch, Depth MapLighter-weight alternatives to ControlNet

SD 1.x and 2.x main models are self-contained: the UNet, text encoder and VAE all come in one install. Both formats are supported:

  • Diffusers folders (for example a HuggingFace repo ID such as Lykon/dreamshaper-8).
  • Single-file checkpoints (.safetensors / .ckpt), the format most community fine-tunes are published in.

No HuggingFace token is needed for the starter models.

If a model was identified incorrectly, open it in the Model Manager and correct its Variant or Prediction Type.

Selecting an SD 1.5 model applies these defaults, which you can override globally or per model in the Model Manager’s Default Settings:

SettingSD 1.5 default
Width × Height512 × 512
Steps30
CFG Scale7
SchedulerEuler Ancestral
  • Resolution — dimensions must be multiples of 8. An image size equivalent to 512×512 in pixel count is recommended. For larger output, use the Upscale tool, or try HiDiffusion.
  • Negative prompt — fully supported.
  • Schedulers — the full standard list (DDIM, DPM++ 2M Karras, Euler, Euler Ancestral, UniPC, and so on) is available.
  • Prompt syntax — SD prompts support Compel weighting and prompt functions.

The Advanced section of the Generate panel adds these SD-specific controls:

  • CLIP Skip — skips the last layers of the text encoder (up to 12 for SD 1.5). Some fine-tunes, especially anime models, are designed to be used with it.
  • CFG Rescale — intended for models trained with zero-terminal SNR; a value around 0.7 is suggested for those. Leave it at 0 for ordinary models.
  • Seamless Tiling — generates images that tile without visible seams, on the X axis, Y axis or both.
  • VAE and VAE Precision — override the model’s bundled VAE, and choose fp32 or fp16 for encoding and decoding. VAE Precision is fp32 unless the model’s default settings in the Model Manager say fp16.
  • HiDiffusion — improves structure at higher resolutions. See HiDiffusion.

Models with an inpainting variant (such as Dreamshaper 8 (inpainting) and Deliberate v5 (inpainting)) are detected automatically. They are made for filling masked areas and extending images on the Canvas, and usually blend new content into the surrounding image better than a normal model does. The Inpainting ControlNet is an alternative that works with ordinary SD 1.5 models.

  • LoRAs for SD 1.x (and SD 2.x), including LyCORIS variants, are supported. Add them in the LoRA section of the Generate panel. LoRAs made for a different model family are not applied.
  • Textual-inversion embeddings are supported in the positive and negative prompts. Type < in a prompt box to pick from the compatible embeddings you have installed — for example, EasyNegative in the negative prompt.

SD 1.5 supports ControlNet and T2I-Adapter models as control layers on the Canvas. A control layer guides composition from an image — edges, depth, pose, line art and so on. See Adding a gallery image to a layer for creating a control layer from an image.

SD 1.5 supports up to five reference images, using IP-Adapter models:

  • Standard Reference — references an image with looser precision.
  • Precise Reference — references an image more precisely.
  • Face Reference — a precise reference adapted for faces.

Each reference image has a Mode — Style, Composition or Both — and a Weight. Under Advanced, you can set the range of steps the reference is active for and, in Style mode, choose a Style variant (Standard, Strong or Precise).

The starter IP-Adapters carry their own image encoder. The CLIP Vision selector next to the model matters only for IP-Adapters installed from a single checkpoint file; if the chosen encoder is not installed, it is downloaded on first use.

SD 1.5 supports regional guidance on the Canvas: each region can have its own positive prompt, its own negative prompt, and its own IP-Adapter reference images.

The Upscale tool supports SD 1.5 main models. It requires a Tile ControlNet for SD 1.5 (the starter Tile ControlNet is part of the bundle), and the bundle’s SwinIR model can serve as the upscaler.

The workflow library includes Text to Image - SD1.5 and MultiDiffusion SD1.5. In the workflow editor, SD 1.x and 2.x models are loaded with the Main Model - SD1.5, SD2 node and prompted with Prompt - SD1.5. Other nodes that work with them include Denoise - SD1.5, SDXL, Apply LoRA - SD1.5, ControlNet - SD1.5, SD2, SDXL, T2I-Adapter - SD1.5, SDXL, IP-Adapter - SD1.5, SDXL, Apply CLIP Skip - SD1.5, SDXL and Apply Seamless - SD1.5, SDXL.

SD 2.x models (SD 2.0 and 2.1 and their fine-tunes) are supported for models you install yourself, from a Diffusers folder or a single-file checkpoint. There are no SD 2.x starter models.

  • Defaults — selecting an SD 2.x model applies 768×768, 30 steps, CFG 7 and Euler Ancestral. That suits the 768px v-prediction checkpoints. For the 512px -base checkpoints, change the size to 512×512, or set it as the model’s default in the Model Manager.
  • Prediction type — SD 2.x checkpoints may be epsilon or v-prediction. InvokeAI detects this on install from the checkpoint or its scheduler config. If a model’s output looks wrong, check its Prediction Type in the Model Manager.
  • CLIP Skip goes up to 24, and CFG Rescale, Seamless Tiling, the negative prompt and regional prompts work as for SD 1.5.
  • Not available for SD 2.x on the Canvas: control layers, reference images (IP-Adapter), HiDiffusion and the Upscale tool. In the workflow editor, the ControlNet - SD1.5, SD2, SDXL node does accept SD 2.x ControlNets.
  • LoRAs and embeddings made for SD 2.x are supported.
This site was designed and developed by Aether Fox Studio.