Skip to content

Ideogram 4

Ideogram 4 is an open-weight text-to-image model with a distinctive structured JSON prompt: instead of a single sentence, the model is trained to read an overall scene description plus a list of regions, each with a bounding box and its own text. Invoke assembles this JSON for you.

Ideogram 4 is released under the Ideogram 4 Non-Commercial License. The weights may not be used commercially, and outputs may not be used to build competing models. Ideogram claims no rights in images you generate, but the license does not explicitly grant commercial use of them. The ideogram-ai repositories are gated; the Comfy-Org copy is not, but the same license applies.

This is a summary, not legal advice. Read the full license on the model’s HuggingFace page before using a model commercially.

The weights are gated on HuggingFace under a non-commercial license. Open the model page, accept the terms, and make sure your HuggingFace token is set up in the config file before installing. Two bundled builds are available:

Paste either repo ID into the Model Manager’s HuggingFace / URL field to install. Each folder carries everything Ideogram 4 needs.

Comfy-Org’s Comfy-Org/Ideogram-4 repackage is ungated (the same non-commercial license still applies) and ships the parts separately. Ideogram 4 runs two transformers — a conditional branch and an unconditional one, both active at every step — so a single-file install is four models, and each starter installs all four together:

  • diffusion_models/ideogram4_<build>.safetensors — the conditional branch, selected as the model
  • diffusion_models/ideogram4_unconditional_<build>.safetensors — the unconditional branch
  • text_encoders/qwen3vl_8b_fp8_scaled.safetensors — the Qwen3-VL 8B encoder (not the 4B one Krea-2 uses). A llama.cpp GGUF of the 8B encoder works as well; see Qwen3-VL GGUF encoders.
  • vae/flux2-vae.safetensors — the 32-channel VAE, shared with FLUX.2, which is why it installs under the FLUX.2 base

Pick the unconditional branch, the encoder and the VAE in the Components section next to the model — and pick the unconditional branch of the same build as the model. The nvfp4 repackage in that repo is not supported yet and is refused at install.

Which branch a file holds is read from the file itself, and from the filename for a build that does not record it, so keep these names as published.

Two builds of the transformers are supported, and they behave differently in memory:

download, per branchresident, per branch
Ideogram 4 (single file, int8) — int8_convrot8.9 GB8.9 GB, on every device
Ideogram 4 (single file, fp8) — fp8_scaled8.6 GB8.7 GB with fp8_compute or FP8 Storage; 17.3 GB with neither

Both branches are resident at once, so double those figures for the pair.

Ideogram 4 keeps both transformer branches loaded for the whole generation. With the single-file builds on a CUDA or ROCm GPU (with partial loading on, the default), InvokeAI checks whether the pair fits next to the working memory a generation needs. If it does not, it keeps only part of the second branch on the GPU and streams the rest in each step. On a 24 GB card with the fp8 pair this keeps about 2.4 GB free instead of 1.2 GB. The bundled nf4 and fp8 folders are not affected. Three things follow:

  • A streamed layer computes on a slightly different path, so images at the same seed can differ from those generated with both branches fully on the GPU.
  • If generations still run out of memory, raise device_working_mem_gb in invokeai.yaml (for example to 8).
  • If you set max_cache_vram_gb, InvokeAI does not limit the second branch this way, because the cache already plans from that limit.

When an Ideogram 4 model is selected, Invoke builds the structured JSON prompt automatically:

  • The positive prompt becomes the overall scene description.
  • Each enabled Regional Guidance layer on the Canvas contributes one element: its drawn box becomes the region’s bounding box and its prompt becomes that region’s description. Draw a box where you want something and describe it there.
  • To drive the model directly, paste a raw JSON object into the prompt box — anything starting with { is passed through unchanged.

The exact JSON that was encoded is stored in the image metadata as Structured Caption, and can be recalled straight back into the prompt box from the metadata viewer.

  • Sampler Preset — the primary quality/speed control. Quality (48 steps), Default (20 steps), and Turbo (12 steps) each bundle a step count, a guidance schedule, and the schedule shift.
  • Advanced overrides (all optional, leave on Auto to use the preset’s values): Steps, Guidance Scale, Schedule Shift (mu), and a Color Palette that biases the generated colors.
This site was designed and developed by Aether Fox Studio.