Skip to content

Anima

Anima is a 2B-parameter, anime-focused text-to-image model built on NVIDIA’s Cosmos Predict2 diffusion transformer. Instead of a large text encoder, it pairs a small Qwen3 0.6B encoder with a built-in LLM adapter that translates the encoder’s output for the transformer. It decodes with the 16-channel Wan 2.1 / Qwen Image VAE.

InvokeAI supports Anima for text-to-image, and on the Canvas for image-to-image, inpainting, outpainting and regional prompts.

Anima is released under the CircleStone Labs Non-Commercial License, and, because it is built on Cosmos-Predict2, the NVIDIA Open Model License also applies. The weights may not be used commercially; the model card says images you generate may be. The download is not gated.

This is a summary, not legal advice. Read the full license on the model’s HuggingFace page before using a model commercially.

Anima is small by current standards: the transformer is a ~4.5 GB download, the Qwen3 encoder ~1.2 GB and the VAE ~200 MB. On low-VRAM GPUs, Low-VRAM mode and FP8 Storage both apply to Anima.

Install the Anima bundle from the Model Manager’s Starter Models. It contains:

  • Anima Base 1.0 — the main model (single-file checkpoint, ~4.5 GB)
  • Anima Qwen3 0.6B Text Encoder (~1.2 GB)
  • Anima QwenImage VAE (~200 MB)
  • Anima LLLite Inpainting and Anima LLLite Sketch — ControlNet-LLLite adapters (see below)

Installing Anima Base 1.0 on its own also installs the encoder and VAE as dependencies. Anima is distributed as a single file only, so the three components are always separate models:

ComponentModel
Transformerthe Anima checkpoint (Cosmos Predict2 DiT + LLM adapter)
Text encodera Qwen3 0.6B encoder
VAEthe 16-channel Wan 2.1 / Qwen Image VAE

Select the Qwen3 Encoder and VAE in the Components section next to the model; both are required. Only the 0.6B Qwen3 encoder is offered — the larger Qwen3 encoders used by Z-Image and FLUX.2 are not compatible. The same VAE file may be installed under the Anima, Qwen Image or Wan base, and all three are listed. FLUX VAEs are not compatible.

Anima also needs a T5-XXL tokenizer, which ships with InvokeAI; no T5 model has to be installed.

Selecting an Anima model applies these defaults:

  • Steps: 35
  • CFG Scale: 4.5 (the minimum is 1, which turns CFG off)
  • Scheduler: Euler
  • Size: 1024×1024; width and height snap to multiples of 8.

The Scheduler menu offers Anima’s own set: Euler, Heun (2nd order), DPM++ 2M, DPM++ 2M SDE, ER-SDE and LCM.

The negative prompt is used only when CFG Scale is above 1; at CFG 1 it is ignored.

On the Canvas, Regional Guidance layers with a positive prompt work with Anima: each region’s prompt is steered to its masked area, alongside the global prompt. The restriction is applied on alternating transformer blocks, leaving the others unrestricted to keep the image coherent. Regional negative prompts and regional reference images are not supported.

In the workflow editor, the Prompt - Anima node has an optional mask input, and Denoise - Anima accepts a collection of conditionings for its positive and negative inputs.

Anima LoRAs (Kohya / LyCORIS format) are supported. They apply to the transformer and, where the LoRA includes text-encoder layers, to the Qwen3 encoder. In the workflow editor use Apply LoRA - Anima or Apply LoRA Collection - Anima.

Anima supports kohya-ss’s ControlNet-LLLite adapters, small models (8–66 MB) that condition the transformer on an image. These are available in the workflow editor only: add an Anima ControlNet-LLLite node, give it the conditioning image and a control model, and connect it to the Control LLLite input of Denoise - Anima. Several adapters can be combined by collecting them, but each adapter model may be used only once per generation. Each node has Weight, Begin Step Percent and End Step Percent settings.

Starter modelPurpose
Anima LLLite InpaintingConditions on the masked image during inpainting/outpainting. Requires the node’s mask input (white = area to inpaint).
Anima LLLite SketchMixed scribble / HED / lineart / grayscale conditioning
Anima LLLite Depth (Preview3)Depth
Anima LLLite Scribble (Preview3)Scribble
Anima LLLite Lineart (Preview3)Lineart
Anima LLLite Pose (Preview3)Pose
NodePurpose
Main Model - AnimaLoads the transformer, Qwen3 encoder and VAE
Prompt - AnimaEncodes a prompt, with an optional regional mask
Denoise - AnimaRuns sampling; accepts img2img latents, masks and LLLite adapters
Image to Latents - AnimaVAE encode
Latents to Image - AnimaVAE decode
Anima ControlNet-LLLiteConfigures one ControlNet-LLLite adapter
This site was designed and developed by Aether Fox Studio.