Skip to content

CogView4

CogView4 is THUDM’s 6B-parameter diffusion-transformer text-to-image model. It uses a GLM language model as its text encoder and a 16-channel VAE. InvokeAI supports the base CogView4-6B checkpoint. CogView4 support is basic: plain generation and Canvas image editing work, but the adapters and controls that other families offer are not available.

CogView4-6B is released under Apache 2.0, which allows commercial use. The download is not gated.

This is a summary, not legal advice. Read the full license on the model’s HuggingFace page before using a model commercially.

The Diffusers download is about 31 GB, including the GLM text encoder and VAE. See System Requirements for general guidance, and the Low VRAM guide if you run out of memory.

Install CogView4 from the Model Manager’s starter models. Its source is THUDM/CogView4-6B. This family has no starter bundle.

CogView4 installs in Diffusers format only, as a single folder that contains the transformer, GLM text encoder, tokenizer and VAE. No other models are needed, and the components cannot be replaced individually. Single-file and GGUF CogView4 models are not recognized.

When you select CogView4, InvokeAI applies the defaults from the model’s reference example:

SettingDefault
Steps50
CFG Scale3.5
Resolution1024 × 1024
  • CFG Scale applies true classifier-free guidance.
  • Width and height must be multiples of 32.
  • Negative prompts are supported and always used.
  • No scheduler choice. CogView4 uses its own flow-matching sampler, so the Scheduler control is hidden.
  • Prompt weighting (Compel syntax) does not apply to CogView4. See Prompt Syntax. The GLM encoder reads up to 1024 tokens, so long, descriptive prompts are fine.

On the Canvas, CogView4 supports text-to-image, image-to-image, inpainting and outpainting. LoRAs, ControlNet, reference images, regional guidance and PiD decoding are not available for CogView4.

The CogView4 nodes are Main Model - CogView4, Prompt - CogView4, Denoise - CogView4, Image to Latents - CogView4 and Latents to Image - CogView4. Each prompt needs its own Prompt - CogView4 node. Connect them to the denoise node’s positive and negative conditioning inputs. Both are required.

This site was designed and developed by Aether Fox Studio.