Skip to content

New Model Type Integration Checklist

This guide covers every step needed to add a new model architecture to InvokeAI:

  • the model manager, which identifies and loads the files
  • the architecture declaration
  • the invocations that form the generation graph
  • the webv2 frontend
  • the tests and CI gates that enforce completeness

Most architectures load their transformer, VAE and text encoder through classes from diffusers and transformers. If the pinned releases lack the new model’s classes, handle the dependency first:

  • Wait for a stable release that contains the classes. Do not pin a git commit; pyproject.toml pins diffusers to an exact version.
  • Bump in a separate change. A new release can change the numerical behavior of every existing architecture. The change has four parts:
    • Update pyproject.toml and uv.lock, including raised transitive minimums such as huggingface-hub.
    • Update tests/backend/model_manager/load/test_diffusers_0XX_compatibility.py. It asserts the exact pinned version and the classes InvokeAI imports.
    • Run the full test suite.
    • Run smoke generations for the existing architectures.
  • Check the transformers requirement in the model card against the pin. Model cards often name the version the authors tested, not a hard minimum. Verify that the encoder and processor behave correctly on the pinned version before bumping it.

File: invokeai/backend/model_manager/taxonomy.py

  1. Add the BaseModelType member.

    invokeai/backend/model_manager/taxonomy.py
    class BaseModelType(str, Enum):
    ...
    Krea2 = "krea-2"
    """Indicates the model is associated with the Krea 2 model architecture, including Krea-2-Turbo."""
    NewModel = "new-model"
    """Indicates the model is associated with the NewModel architecture."""

    The value is stored in every user’s model database, so it cannot be renamed later. It also names the declaration module, with - replaced by _, so it must yield a valid Python module name. Use lowercase letters, digits and - only, and no dots: qwen-image-2-1, not qwen-image-2.1.

  2. Add a variant enum (if needed).

    Add one only when models of the architecture differ in ways that matter at load or generation time, for example distilled versus base. Variant strings are resolved without the base (configs/factory.py), so every value must be unique across all variant enums. That is why Krea2VariantType.Turbo is "krea2_turbo" and not "turbo".

    invokeai/backend/model_manager/taxonomy.py
    class NewModelVariantType(str, Enum):
    """NewModel variants."""
    Turbo = "new_model_turbo"
    """Distilled: few steps, CFG off."""
    Base = "new_model_base"
    """Undistilled base model: more steps, CFG on."""

    List the enum in all three places:

    • AnyVariant in taxonomy.py
    • variant_type_adapter in taxonomy.py
    • ModelRecordChanges.variant in invokeai/app/services/model_records/model_records_base.py

    tests/backend/architectures/test_variants.py checks that the three agree and that the values are unique.

  3. Add encoder types (only for a new text encoder).

    There is no generic text-encoder type. Each encoder family has its own ModelType member, and a ModelFormat member for its folder layout. Existing examples are Qwen3Encoder, Qwen3VLEncoder, QwenVLEncoder, MistralEncoder and T5Encoder.

    Reuse an existing type whenever the encoder is a model InvokeAI already supports. Krea-2, Ideogram 4 and MiniMax H3 all use ModelType.Qwen3VLEncoder, told apart by Qwen3VLVariantType. Check whether the encoder weights are a stock checkpoint before adding a type.


File: invokeai/backend/architectures/defs/new_model.py

Each architecture declares what it is in exactly one register(...) call. Modules under defs/ are discovered automatically, so there is no list to edit.

At boot, validate() runs from ApiDependencies.initialize. It refuses to start the app in two cases:

  • a BaseModelType has no module
  • a module omits a required facet

It checks only that the facets are present, not that their values are right.

FacetRequiredWhat it declaresRead by
LatentSpaceFacetyesLatent channels, spatial compression, latent→RGB projectionStep previews (app/util/step_callback.py)
ConditioningFacetyesThe *ConditioningInfo class the text encoder writesThe safe-globals list for conditioning deserialization
DefaultSettingsFacetyesSteps, CFG or guidance, scheduler and size per variant (None is the fallback); optional by_name_hintStored on the model config at identification; prefills the UI
ModalityFacetyesModes (txt2img, img2img, inpaint, outpaint, video modes) and metadata_slugMust equal GENERATION_MODES; which canvas tools the UI offers
FeaturesFacetyesNegative prompt, dimension_grid, guidance label and range, scheduler set, control kinds, reference images, regional guidance, and moreServed to webv2 at GET /api/v2/models/capabilities
VaeFacetnoVAE bases (and channel counts) the architecture decodes withModel loader VAE field, VAE picker, loader validation
VariantFacetnoVariant enum per ModelTypeVariant consistency tests
UNetDownscaleFacetnoUNet downscale factorSD-family UNets only

The facet classes live in invokeai/backend/architectures/facets/. Modeled on defs/krea_2.py:

invokeai/backend/architectures/defs/new_model.py
"""What the new-model architecture declares."""
from invokeai.backend.architectures.facets.conditioning import ConditioningFacet
from invokeai.backend.architectures.facets.default_settings import DefaultSettingsFacet
from invokeai.backend.architectures.facets.features import FeaturesFacet, NegativePrompt
from invokeai.backend.architectures.facets.latent_space import FLUX2_32, LatentSpaceFacet
from invokeai.backend.architectures.facets.modality import ModalityFacet
from invokeai.backend.architectures.facets.vae import VaeCompatibility, VaeFacet
from invokeai.backend.architectures.facets.variant import VariantFacet
from invokeai.backend.architectures.registry import register
from invokeai.backend.model_manager.configs.default_settings import MainModelDefaultSettings
from invokeai.backend.model_manager.taxonomy import BaseModelType, ModelType, NewModelVariantType
from invokeai.backend.stable_diffusion.diffusion.conditioning_data import NewModelConditioningInfo
register(
BaseModelType.NewModel,
# NewModel decodes with the FLUX.2 VAE, so it shares that latent space.
LatentSpaceFacet(FLUX2_32),
ConditioningFacet(NewModelConditioningInfo),
DefaultSettingsFacet(
{
# Values from the model card.
NewModelVariantType.Base: MainModelDefaultSettings(
scheduler="euler", steps=28, cfg_scale=4.0, width=1024, height=1024
),
# Turbo, and any model whose variant is unknown.
None: MainModelDefaultSettings(scheduler="euler", steps=8, cfg_scale=1.0, width=1024, height=1024),
}
),
ModalityFacet(frozenset({"txt2img", "img2img", "inpaint", "outpaint"}), metadata_slug="new_model"),
FeaturesFacet(
negative_prompt=NegativePrompt(visible=True, usage="cfg-gated"),
dimension_grid=16, # equals `multiple_of` on new_model_denoise.width/height
guidance_label="CFG",
scheduler_set="flow",
),
VaeFacet(frozenset({VaeCompatibility(BaseModelType.Flux2)})),
# Variant values must be globally unique; see taxonomy.py.
VariantFacet({ModelType.Main: NewModelVariantType}),
)

Things to get right:

  • Latent space. If the architecture uses a VAE InvokeAI already supports, reuse its shared space from facets/latent_space.py:

    • FLUX_16: FLUX.1, Z-Image
    • FLUX2_32: FLUX.2, ERNIE-Image, Ideogram 4
    • WAN21_16: Wan 2.1, Qwen-Image, Krea-2, Anima
    • WAN22_48

    Otherwise, declare a new LatentSpace in that file. Compute its projection with scripts/generate_vae_linear_approximation.py rather than guessing it.

  • Grid and guidance are pinned to the denoise node by tests:

    • dimension_grid must equal the multiple_of on the node’s width and height.
    • guidance_min and guidance_max must match the ge and le of the guidance field.
    • A constraint that the node only enforces inside invoke(), per variant, goes in dimension_grid_by_variant.
  • VaeFacet replaces the default rule, under which an architecture accepts only its own base. List every accepted base explicitly, including the architecture’s own if it applies. Z-Image, for example, accepts only FLUX VAEs.

  • Mode strings are persisted in image metadata and cannot be renamed. metadata_slug is usually the base value with _ (z_image), but not always (krea2).

  • Default settings are copied onto the model config when a model is identified (configs/factory.py), so model configs need no code for them. Use by_name_hint only for sub-models that nothing on disk can tell apart.


Folder: invokeai/backend/model_manager/configs/

Configs identify a model on disk. Every non-abstract subclass of Config_Base registers itself through Config_Base.__init_subclass__; there is no decorator.

  1. Main model configs (configs/main.py)

    One class per format. Each combines a format mixin, Main_Config_Base and Config_Base:

    invokeai/backend/model_manager/configs/main.py
    class Main_Diffusers_NewModel_Config(Diffusers_Config_Base, Main_Config_Base, Config_Base):
    """Model config for NewModel diffusers models."""
    base: Literal[BaseModelType.NewModel] = Field(BaseModelType.NewModel)
    variant: NewModelVariantType = Field()
    @classmethod
    def from_model_on_disk(cls, mod: ModelOnDisk, override_fields: dict[str, Any]) -> Self:
    raise_if_not_dir(mod)
    raise_for_override_fields(cls, override_fields)
    # The pipeline class name implies the base.
    raise_for_class_name(common_config_paths(mod.path), {"NewModelPipeline"})
    # `_get_variant` reads whatever distinguishes the variants, e.g. a flag in model_index.json.
    variant = override_fields.pop("variant", None) or cls._get_variant(mod)
    repo_variant = override_fields.pop("repo_variant", None) or cls._get_repo_variant_or_raise(mod)
    return cls(**override_fields, variant=variant, repo_variant=repo_variant)
    class Main_Checkpoint_NewModel_Config(Checkpoint_Config_Base, Main_Config_Base, Config_Base):
    """Model config for NewModel single-file checkpoints."""
    base: Literal[BaseModelType.NewModel] = Field(default=BaseModelType.NewModel)
    format: Literal[ModelFormat.Checkpoint] = Field(default=ModelFormat.Checkpoint)
    variant: NewModelVariantType = Field()
    @classmethod
    def from_model_on_disk(cls, mod: ModelOnDisk, override_fields: dict[str, Any]) -> Self:
    raise_if_not_file(mod)
    raise_for_override_fields(cls, override_fields)
    state_dict = mod.load_state_dict()
    if not _has_new_model_keys(state_dict):
    raise NotAMatchError("state dict does not look like a NewModel model")
    if _has_ggml_tensors(state_dict):
    raise NotAMatchError("state dict looks like GGUF quantized")
    variant = override_fields.pop("variant", None) or _get_new_model_variant_from_name(mod.path.name)
    return cls(**override_fields, variant=variant)

    If GGUF files exist, add a Main_GGUF_NewModel_Config with format: Literal[ModelFormat.GGUFQuantized] that requires _has_ggml_tensors.

    The identification helpers are in configs/identification_utils.py:

    • NotAMatchError: “not this config”
    • InvalidMatchError: “this config, but the file is broken”
    • raise_if_not_file / raise_if_not_dir
    • raise_for_override_fields
    • raise_for_class_name
    • common_config_paths
    • state_dict_has_any_keys_*
  2. Detection helpers

    Detect the architecture from keys that only it has, and from shapes where keys are shared. Strip ComfyUI prefixes such as model.diffusion_model. from each key first, with _strip_comfyui_key_prefix in configs/main.py. It reads the shared COMFYUI_KEY_PREFIXES, so detection and the loaders agree on which prefixes exist. Exclude LoRA suffixes, so that a LoRA for the architecture is not mistaken for a main model.

    invokeai/backend/model_manager/configs/main.py
    def _has_new_model_keys(state_dict: dict[str | int, Any]) -> bool:
    """True for NewModel transformer weights, False for its LoRAs and for every other architecture."""
    # `txt_in.text_norm` is unique to NewModel; `img_in` alone is shared with Qwen-Image and Krea-2.
    ...

    Configs must exclude each other. Identification iterates Config_Base.CONFIG_CLASSES, which is a set, so the order in the AnyModelConfig union decides nothing. When two configs match the same file, matches_sort_key breaks the tie, which amounts to chance. For every existing config whose heuristic the new files could satisfy, add a negative check on one side or both. The same applies to LoRA and VAE configs, whose heuristics are often loose, for example “any key starting with transformer_blocks.”. Cover each exclusion with a detection test.

  3. VAE config (only for a new VAE) (configs/vae.py)

    VAE_Checkpoint_Config_Base and VAE_Diffusers_Config_Base detect the plain AutoencoderKL VAEs: SD 1, SD 2 (single files only) and SDXL at 4 latent channels, FLUX.1 and SD3 at 16. Within one latent width the weights cannot tell these bases apart, so _VAE_FAMILIES groups them, and an explicit base override chooses within a family but never across one. For 16 channels, a folder’s config.json decides by scaling_factor/shift_factor, and a config that matches neither is not filed at all. A 16-channel single file goes to the backbone named in its file name, folder or install source (configs/backbone_names.py), else to FLUX.1. At 4 channels the older heuristics still apply: the file name for a single file, the SDXL config values or name for a folder. A new architecture that reuses one of these networks with its own constants (CogView 4 would) belongs in that table and in the diffusers constants check, not in a new detector.

    Newer architectures with their own network subclass Checkpoint_Config_Base, Config_Base directly, as VAE_Checkpoint_Flux2_Config does:

    invokeai/backend/model_manager/configs/vae.py
    class VAE_Checkpoint_NewModel_Config(Checkpoint_Config_Base, Config_Base):
    """Model config for NewModel VAE checkpoints."""
    type: Literal[ModelType.VAE] = Field(default=ModelType.VAE)
    format: Literal[ModelFormat.Checkpoint] = Field(default=ModelFormat.Checkpoint)
    base: Literal[BaseModelType.NewModel] = Field(default=BaseModelType.NewModel)
    cpu_only: bool | None = Field(default=None, description="Whether this model should run on CPU only")
    @classmethod
    def from_model_on_disk(cls, mod: ModelOnDisk, override_fields: dict[str, Any]) -> Self:
    raise_if_not_file(mod)
    raise_for_override_fields(cls, override_fields)
    if not _is_new_model_vae(mod.load_state_dict()):
    raise NotAMatchError("state dict does not look like a NewModel VAE")
    return cls(**override_fields)

    Then exclude the new VAE in VAE_Checkpoint_Config_Base._validate_looks_like_vae, and in every family detector it could also satisfy. Wan, Qwen-Image and Anima VAEs share one layout family.

  4. Text encoder config (only for a new encoder type)

    Create configs/<encoder>.py with base=BaseModelType.Any, the new ModelType, and one config per format: folder (its own ModelFormat), single-file checkpoint, and GGUF if it exists. qwen3_encoder.py, qwen3_vl_encoder.py and mistral_encoder.py are complete examples.

  5. AnyModelConfig union (configs/factory.py)

    Add each new config to the union, in the shape the other entries use:

    invokeai/backend/model_manager/configs/factory.py
    Annotated[Main_Diffusers_NewModel_Config, Main_Diffusers_NewModel_Config.get_tag()],

Folder: invokeai/backend/model_manager/load/model_loaders/

Loaders turn a config into an in-memory model. Every module in this folder is imported automatically. A loader registers for a (base, type, format) triple. Lookup falls back to base=BaseModelType.Any, which is how encoder loaders serve every architecture.

  1. Diffusers-format main model

    Subclass GenericDiffusersLoader. It resolves each submodel’s class from model_index.json. Override _load_model only for what the pipeline gets wrong: tokenizer quirks, config patches, dtype, or dropping unused modules.

    invokeai/backend/model_manager/load/model_loaders/new_model.py
    @ModelLoaderRegistry.register(base=BaseModelType.NewModel, type=ModelType.Main, format=ModelFormat.Diffusers)
    class NewModelDiffusersModel(GenericDiffusersLoader):
    """Loads NewModel submodels from a diffusers pipeline folder."""
    def _load_model(self, config: AnyModelConfig, submodel_type: Optional[SubModelType] = None) -> AnyModel:
    if submodel_type is None:
    raise Exception("A submodel type must be provided when loading main pipelines.")
    model_path = Path(config.path)
    load_class = self.get_hf_load_class(model_path, submodel_type)
    dtype = TorchDevice.choose_bfloat16_safe_dtype(TorchDevice.choose_torch_device())
    result: AnyModel = load_class.from_pretrained(model_path / submodel_type.value, torch_dtype=dtype)
    return self._apply_fp8_layerwise_casting(result, config, submodel_type)
  2. Single-file and GGUF main models

    Subclass ModelLoader, build the transformer under accelerate.init_empty_weights(), convert the keys to the diffusers layout, and load with assign=True. Krea2CheckpointModel and Krea2GGUFCheckpointModel in krea2.py show the full pattern:

    • stripping the ComfyUI prefix with CheckpointPrefix.detect(sd).strip(sd) (invokeai/backend/model_manager/checkpoint_prefix.py) rather than a hand-written loop
    • reserving RAM before materializing, exactly once: build the empty model first, take the fp8 and nvfp4 side channels out of the state dict, then reserve, and only then fold or widen anything. A fold before the reservation lands on a cache that was never asked for the room. For fp8, nvfp4 and dense files the reservation is reserve_for_load (invokeai/backend/quantization/load_plan.py).
    • quantized formats: ComfyUI fp8_scaled, int8_convrot and nvfp4 (helpers in invokeai/backend/quantization/), and GGUF through gguf_sd_loader. For int8, call install_int8_convrot_layers instead of its individual steps. It makes the reservation itself through its reserve argument, so do not also call reserve_for_load; the steps’ order matters, and skipping the first one loads a mixed checkpoint’s fp8 weights without their scale.
    • a loader that handles no quantization side channel (a VAE, an adapter, a non-strict load_state_dict) calls reject_quantized_side_channel so a quantized file is refused rather than loaded wrongly
    • opting into FP8 storage with _apply_fp8_layerwise_casting

    Declare FP8 Storage support at registration. The Model Manager only shows the FP8 Storage toggle where fp8_storage_verdict(base, type, format) (invokeai/backend/model_manager/load/fp8_capability.py) says it works. A loader that does not apply the cast says so where it registers, with a reason: fp8_storage=Unimplemented("...") for work not done yet, NotApplicable("...") for a deliberate decision. tests/backend/model_manager/load/test_fp8_capability.py walks the registry and fails for a loader that neither casts nor declares.

    Community single-file checkpoints often use key names that differ from diffusers, for example fused projections. Verify the conversion numerically against the diffusers module on a small config.

  3. VAE loader (only for a new VAE)

    Build the diffusers VAE class and load the converted state dict. vae.py and flux.py handle existing VAEs from native and ComfyUI layouts.

  4. Text encoder loader (only for a new encoder type)

    Register with base=BaseModelType.Any and the encoder’s ModelType, and serve both SubModelType.TextEncoder and SubModelType.Tokenizer. Vendor tokenizers and configs the loader needs, as invokeai/backend/qwen2_5_vl/ does, so a single-file install loads without network access. Add the vendored files to [tool.setuptools.package-data] in pyproject.toml, or a wheel install fails with FileNotFoundError.


A text encoder saves a ConditioningFieldData to disk, and the denoise node reads it back. Each architecture has its own conditioning types, and they must be threaded through five places:

  1. Info class and union (invokeai/backend/stable_diffusion/diffusion/conditioning_data.py)

    Add a dataclass with a to(device, dtype) method, and add it to the ConditioningFieldData.conditionings union. The ConditioningFacet then puts it on the safe-globals list that deserialization needs. A class missing from that list fails with an UnpicklingError inside the denoise node, after the encoder has already run.

    invokeai/backend/stable_diffusion/diffusion/conditioning_data.py
    @dataclass
    class NewModelConditioningInfo:
    """NewModel text conditioning."""
    prompt_embeds: torch.Tensor
    """Shape: (batch_size, seq_len, hidden_size)."""
    prompt_embeds_mask: torch.Tensor | None = None
    """Shape: (batch_size, seq_len). True for valid tokens."""
    def to(self, device: torch.device | None = None, dtype: torch.dtype | None = None):
    self.prompt_embeds = self.prompt_embeds.to(device=device, dtype=dtype)
    if self.prompt_embeds_mask is not None:
    self.prompt_embeds_mask = self.prompt_embeds_mask.to(device=device)
    return self
  2. Field (invokeai/app/invocations/fields.py): NewModelConditioningField with conditioning_name and an optional mask for regional prompting. Add node field descriptions to FieldDescriptions in the same file.

  3. Output (invokeai/app/invocations/primitives.py): NewModelConditioningOutput, registered with @invocation_output("new_model_conditioning_output").

  4. Encoder field (invokeai/app/invocations/model.py, only for a new encoder type): for example Qwen3VLEncoderField, with tokenizer, text_encoder and loras.

  5. Public API (invokeai/invocation_api/__init__.py): export the info class, and list it in __all__.


Invocations expose the architecture as nodes in the generation graph.

  1. Model loader (new_model/new_model_model_loader.py)

    Outputs the submodel identifiers. The transformer always comes from the main model; the VAE and encoder come from standalone models if selected, otherwise from the Diffusers main model. The field hints ui_model_base and ui_model_type drive the model pickers. The VAE field must use accepted_vae_bases(), which tests/backend/architectures/test_vae.py enforces.

    invokeai/app/invocations/new_model/new_model_model_loader.py
    @invocation_output("new_model_model_loader_output")
    class NewModelModelLoaderOutput(BaseInvocationOutput):
    transformer: TransformerField = OutputField(description=FieldDescriptions.transformer, title="Transformer")
    qwen3_vl_encoder: Qwen3VLEncoderField = OutputField(
    description=FieldDescriptions.qwen3_vl_encoder, title="Qwen3-VL Encoder"
    )
    vae: VAEField = OutputField(description=FieldDescriptions.vae, title="VAE")
    @invocation("new_model_model_loader", title="Main Model - NewModel", tags=["model", "new_model"],
    category="model", version="1.0.0", classification=Classification.Prototype)
    class NewModelModelLoaderInvocation(BaseInvocation):
    model: ModelIdentifierField = InputField(
    description=FieldDescriptions.main_model, input=Input.Direct,
    ui_model_base=BaseModelType.NewModel, ui_model_type=ModelType.Main, title="Transformer",
    )
    vae_model: Optional[ModelIdentifierField] = InputField(
    default=None, description="Standalone VAE model.", input=Input.Direct,
    ui_model_base=accepted_vae_bases(BaseModelType.NewModel), ui_model_type=ModelType.VAE, title="VAE",
    )
    qwen3_vl_encoder_model: Optional[ModelIdentifierField] = InputField(
    default=None, description="Standalone Qwen3-VL Encoder model.", input=Input.Direct,
    ui_model_type=ModelType.Qwen3VLEncoder, title="Qwen3-VL Encoder",
    )
    def invoke(self, context: InvocationContext) -> NewModelModelLoaderOutput:
    transformer = self.model.model_copy(update={"submodel_type": SubModelType.Transformer})
    # Validate each standalone component (base, type, variant, accepts_vae(...)) with an error
    # that names what the user picked; fall back to the Diffusers main model's submodels.
    ...
  2. Text encoder (text_encoder/new_model_text_encoder.py)

    Loads tokenizer and encoder through context.models.load(...), applies LoRAs to the encoder if the architecture has any, extracts the hidden states the transformer was trained on, and saves the conditioning:

    invokeai/app/invocations/text_encoder/new_model_text_encoder.py
    @invocation("new_model_text_encoder", title="Prompt - NewModel", tags=["prompt", "conditioning", "new_model"],
    category="conditioning", version="1.0.0", classification=Classification.Prototype,
    idle_gpu_offloadable=True)
    class NewModelTextEncoderInvocation(BaseInvocation):
    prompt: str = InputField(description="Text prompt.", ui_component=UIComponent.Textarea)
    mask: TensorField | None = InputField(default=None, description="Regional mask.", input=Input.Connection)
    qwen3_vl_encoder: Qwen3VLEncoderField = InputField(
    title="Qwen3-VL Encoder", description=FieldDescriptions.qwen3_vl_encoder, input=Input.Connection
    )
    @torch.no_grad()
    def invoke(self, context: InvocationContext) -> NewModelConditioningOutput:
    prompt_embeds, prompt_mask = self._encode(context) # template, layer choice, padding
    conditioning_data = ConditioningFieldData(
    conditionings=[NewModelConditioningInfo(prompt_embeds=prompt_embeds.cpu(), prompt_embeds_mask=prompt_mask)]
    )
    conditioning_name = context.conditioning.save(conditioning_data)
    return NewModelConditioningOutput.build(conditioning_name, mask=self.mask)

    Copy the reference pipeline’s encoding exactly and cover it with a test: the prompt template, the prefix tokens it drops, the hidden-state layer (and whether it is before or after the final norm), the padding side and the maximum length. Small deviations here degrade images without failing.

  3. Denoise (new_model/new_model_denoise.py)

    invokeai/app/invocations/new_model/new_model_denoise.py
    @invocation("new_model_denoise", title="Denoise - NewModel", tags=["image", "new_model"],
    category="image", version="1.0.0", classification=Classification.Prototype)
    class NewModelDenoiseInvocation(BaseInvocation):
    latents: Optional[LatentsField] = InputField(default=None, description=FieldDescriptions.latents,
    input=Input.Connection)
    denoise_mask: Optional[DenoiseMaskField] = InputField(default=None, description=FieldDescriptions.denoise_mask,
    input=Input.Connection)
    denoising_start: float = InputField(default=0.0, ge=0, le=1, description=FieldDescriptions.denoising_start)
    denoising_end: float = InputField(default=1.0, ge=0, le=1, description=FieldDescriptions.denoising_end)
    transformer: TransformerField = InputField(description=FieldDescriptions.transformer, input=Input.Connection)
    positive_conditioning: NewModelConditioningField | list[NewModelConditioningField] = InputField(
    description=FieldDescriptions.positive_cond, input=Input.Connection
    )
    negative_conditioning: NewModelConditioningField | list[NewModelConditioningField] | None = InputField(
    default=None, description=FieldDescriptions.negative_cond, input=Input.Connection
    )
    cfg_scale: float | list[float] = InputField(default=1.0, description=FieldDescriptions.cfg_scale)
    # `multiple_of` must equal FeaturesFacet.dimension_grid.
    width: int = InputField(default=1024, gt=0, multiple_of=16, description="Width of the generated image.")
    height: int = InputField(default=1024, gt=0, multiple_of=16, description="Height of the generated image.")
    steps: int = InputField(default=8, gt=0, description=FieldDescriptions.steps)
    seed: int = InputField(default=0, description="Randomness seed for reproducibility.")
    def invoke(self, context: InvocationContext) -> LatentsOutput:
    # 1. Noise from the seed; init latents for img2img; clip the schedule by denoising_start/end.
    # 2. Load the transformer with a working-memory estimate; apply LoRA patches.
    # 3. Loop: model prediction, optional CFG, scheduler step,
    # RectifiedFlowInpaintExtension merge when a denoise_mask is set, step callback.
    # 4. Save and return the latents.
    ...

    Report progress with context.util.sd_step_callback(state, BaseModelType.NewModel); it renders previews through the architecture’s LatentSpaceFacet. The denoise node is also where these are checked: interrupts (context.util.is_canceled()), inpainting, img2img through denoising_start, and a working-memory estimate passed to model_on_device(working_mem_bytes=...).

  4. VAE encode and decode (vae/new_model_image_to_latents.py, vae/new_model_latents_to_image.py)

    Apply the VAE’s latent normalization, scaling and shift or per-channel mean and std, exactly as the reference pipeline does. Handle tiling and give a working-memory estimate. The estimate helpers live in invokeai/backend/util/vae_working_memory.py, with calibration scripts in scripts/calibrate_*_working_memory.py. Node types conventionally end in _i2l and _l2i.

  5. LoRA loaders (if LoRA is supported, see Optional features)

    new_model_lora_loader and new_model_lora_collection_loader. webv2 uses the collection loader.


  • Backend package (optional). Put reusable math in invokeai/backend/new_model/: noise, packing, position IDs, schedules and text encoding. FLUX, FLUX.2 and Krea-2 have one; Qwen-Image and Z-Image keep the loop in the denoise node. Use a package once code is shared between nodes or needs its own tests.
  • Schedulers. Flow-matching architectures use diffusers’ FlowMatchEulerDiscreteScheduler, or the shared maps in invokeai/backend/flux/schedulers.py: FLUX_SCHEDULER_MAP, ZIMAGE_SCHEDULER_MAP and ERNIE_IMAGE_SCHEDULER_MAP, each with name values and labels. Declare which set the UI offers in FeaturesFacet.scheduler_set. Set scheduler_applies_to_graph=True only if the denoise node really takes a scheduler field. Reproduce the reference sigma schedule and shift (mu, time-shift type, terminal shift) exactly, and test it against the reference scheduler.
  • External noise (optional). Only if the denoise node accepts a noise tensor: extend LatentNoiseType and its shape and grid logic in invokeai/app/invocations/latent_noise.py, and extend noise_type in invocations/noise.py. Validate with validate_noise_tensor_shape.
  • Inpainting. Rectified-flow models use RectifiedFlowInpaintExtension (invokeai/backend/rectified_flow/rectified_flow_inpaint_extension.py).
  • Previews. LatentSpaceFacet provides them; there is nothing to add in the step callback.

File: invokeai/app/invocations/metadata.py

  1. Add the mode strings to GENERATION_MODES.

    invokeai/app/invocations/metadata.py
    GENERATION_MODES = Literal[
    ...
    "new_model_txt2img",
    "new_model_img2img",
    "new_model_inpaint",
    "new_model_outpaint",
    ]

    They must equal the ModalityFacet declaration: <metadata_slug>_<mode> for each mode. tests/backend/architectures/test_modality.py compares the two. webv2 writes <slug>_txt2img, and the canvas derives the other modes by replacing txt2img, so declare all four image modes if the canvas supports them.

  2. Metadata fields (if needed).

    CoreMetadataInvocation (core_metadata) accepts extra fields, so architecture-specific keys such as the chosen encoder or VAE need no declaration. Declare a field only if it needs validation. Adding one requires a node version bump and follows the CORE_METADATA_VERSION rules documented in Media Metadata.


Folder: invokeai/backend/model_manager/starter_models/

The catalogue is a package with one module per architecture:

  • types.py holds StarterModel.
  • common.py holds encoders and upscalers that several architectures share. Import a shared dependency from there instead of declaring a second copy; if an encoder you need is declared in another architecture’s module, move it to common.py.
invokeai/backend/model_manager/starter_models/new_model.py
"""NewModel starter models."""
from invokeai.backend.model_manager.starter_models.common import qwen3_vl_encoder_4b
from invokeai.backend.model_manager.starter_models.types import StarterModel
from invokeai.backend.model_manager.taxonomy import BaseModelType, ModelFormat, ModelType, NewModelVariantType
new_model_vae = StarterModel(
name="NewModel VAE",
base=BaseModelType.NewModel,
source="org/NewModel::vae/diffusion_pytorch_model.safetensors", # a file inside a repo
description="NewModel VAE. ~300MB",
type=ModelType.VAE,
)
new_model_turbo = StarterModel(
name="NewModel Turbo",
base=BaseModelType.NewModel,
source="org/NewModel-Turbo", # a whole diffusers repo
description="NewModel Turbo, full diffusers pipeline. ~20GB",
type=ModelType.Main,
variant=NewModelVariantType.Turbo,
)
new_model_turbo_gguf_q4_k_m = StarterModel(
name="NewModel Turbo (Q4_K_M GGUF)",
base=BaseModelType.NewModel,
source="https://huggingface.co/org/NewModel-Turbo-GGUF/resolve/main/new_model_turbo-Q4_K_M.gguf",
description="GGUF ships only the transformer; the VAE and encoder are installed as dependencies. ~7GB",
type=ModelType.Main,
format=ModelFormat.GGUFQuantized,
variant=NewModelVariantType.Turbo,
dependencies=[new_model_vae, qwen3_vl_encoder_4b],
)

Then import the models in the package’s __init__.py:

  • Add them to STARTER_MODELS. It is a curated order, the order the install dialog shows, so insert them where they belong instead of appending.
  • Optionally add a STARTER_BUNDLES entry keyed by the base.
  • Keep source values unique. The module asserts that.

Put license restrictions in the description; the Ideogram 4 and MiniMax H3 descriptions are examples.


webv2 does not keep per-architecture tables of its own. Grid, default settings, negative-prompt behavior, guidance label and range, scheduler set, VAE acceptance, control kinds and reference-image limits all come from the backend’s FeaturesFacet, DefaultSettingsFacet and VaeFacet, served at GET /api/v2/models/capabilities. What webv2 owns is the graph topology and the per-family UI.

Paths below are relative to invokeai/frontend/webv2/src/features/generation/core/ unless stated otherwise.

  1. Regenerate the capabilities fixture

    Terminal window
    REGEN_CAPABILITIES_FIXTURE=1 uv run --no-sync pytest tests/backend/architectures/test_capabilities_fixture.py

    This rewrites __fixtures__/architectureCapabilities.json. The mock backend and the unit tests read it.

  2. Mark the base as buildable

    • Add it to KnownGenerationModelBase in contracts.ts.
    • Add it to SUPPORTED_GENERATE_BASES in supportedBases.ts.
    • Update the pinned list in supportedBases.test.ts.
  3. Write the graph builder (graph.ts)

    Add buildNewModelGraph and register it in GRAPH_BUILDERS, which satisfies Record<SupportedGenerateBase, …>, so the compiler reports a missing entry. buildZImageGraph and buildKrea2Graph are good templates:

    invokeai/frontend/webv2/src/features/generation/core/graph.ts
    const buildNewModelGraph = (
    settings: GenerateSettings,
    model: MainModelConfig,
    outputIsIntermediate: boolean,
    projectSettings: GenerationProjectSettings
    ): BackendGraphContract => {
    const vaeModel = getCompatibleVae(settings, model);
    const graph: BackendGraphContract = { edges: [], id: createId('new_model_graph'), nodes: {} };
    const { negativePrompt, positivePrompt, seed } = addPromptAndSeedNodes(graph);
    const useCfg = settings.cfgScale > 1;
    const modelLoader = addNode(graph, {
    id: 'model_loader',
    model,
    type: 'new_model_model_loader',
    vae_model: vaeModel ?? undefined,
    });
    const posCond = addNode(graph, { id: 'pos_cond', type: 'new_model_text_encoder' });
    const posCondCollect = addNode(graph, { id: 'pos_cond_collect', type: 'collect' });
    const size = getDenoiseSize(settings, model);
    const denoise = addNode(graph, {
    cfg_scale: settings.cfgScale,
    denoising_end: 1,
    denoising_start: 0,
    height: size.height,
    id: 'denoise_latents',
    steps: settings.steps,
    type: 'new_model_denoise',
    width: size.width,
    });
    // Creates the `canvas_output` node.
    const output = addImageOutputNode(graph, 'new_model_l2i', outputIsIntermediate);
    addEdge(graph, modelLoader, 'transformer', denoise, 'transformer');
    addEdge(graph, modelLoader, 'qwen3_vl_encoder', posCond, 'qwen3_vl_encoder');
    addEdge(graph, positivePrompt, 'value', posCond, 'prompt');
    addEdge(graph, posCond, 'conditioning', posCondCollect, 'item');
    addEdge(graph, posCondCollect, 'collection', denoise, 'positive_conditioning');
    // …the same chain for `negative_prompt` → `neg_cond` → `neg_cond_collect` when useCfg…
    addEdge(graph, seed, 'value', denoise, 'seed');
    addEdge(graph, denoise, 'latents', output, 'latents');
    addEdge(graph, modelLoader, 'vae', output, 'vae');
    addMetadata(graph, output, settings, model, 'new_model_txt2img', projectSettings, {
    vae: vaeModel ?? undefined,
    });
    return graph;
    };

    The builder produces the txt2img graph only; the canvas grafts img2img, inpaint and outpaint onto it. That depends on these conventions:

    • Node ids positive_prompt, negative_prompt, seed, model_loader, pos_cond, neg_cond, pos_cond_collect, neg_cond_collect and denoise_latents.
    • An output node canvas_output whose vae input has an edge (requireVaeSource).
    • addDecodeOutput if the architecture supports PiD decoding, otherwise addImageOutputNode.

    Validation that only one architecture needs belongs in getModelFamilyValidationReasons (baseGenerationPolicies.ts).

  4. Component pickers (encoder, VAE, Diffusers component source)

    • Add a case to getBaseComponentSectionPolicy in baseGenerationPolicies.ts. Its default returns an empty policy silently.
    • Add encoder filters in componentCompatibility.ts. If a main model can bring its own components, update isBundledMainForBase. VAE acceptance comes from the backend.
    • A new settings key must be added in all of these places:
      • GenerateSettings (types.ts)
      • GenerateComponentValueKey, COMPONENT_SETTING_LABELS, getComponentPolicyContext and getDefaultGenerateSettings (baseGenerationPolicies.ts)
      • normalizeGenerateSettings and cloneGenerateWidgetValues (settings.ts)
  5. Canvas

    • Map the encode node in CANVAS_I2L_NODE_TYPES (canvas/compileCanvasGraph.ts). The map is Partial, so a missing entry only fails at runtime.
    • If the architecture uses the strength^0.2 denoising-start curve, add it to canvasDenoisingStart.
    • Add a case to BASE_CASES in canvas/compileCanvasGraph.test.ts.
  6. Recall (src/workbench/image-actions/)

    • In imageRecall.ts, add the component keys the builder writes into metadata to hasComponentModels and to the component patch.
    • Add the base to GUIDANCE_BASES in recallParameters.ts if it takes the shared guidance slider.

    Nothing checks recall against what the builder writes, so test it.

  7. Model registry and labels

    • src/features/models/core/baseIdentity.ts: MODEL_BASES, plus its pinned test
    • src/features/models/core/types.ts: ModelBase
    • src/features/models/core/taxonomy.ts: variant values
    • src/features/models/core/relationships.ts: which encoder and VAE types may link to the base
    • src/features/workflow/core/modelRequirements.ts: BASE_LABELS
    • New user-facing strings go in invokeai/frontend/webv2/public/locales/en.json.
  8. Per-family UI (only for controls unique to the architecture)

    src/features/generation/ui/GenerateRenderSection.tsx branches on the family for controls such as Krea-2 seed variance and Wan settings. Controls every architecture has come from capabilities; they need no code.

  9. Regenerate the graph contract

    Terminal window
    pnpm -C invokeai/frontend/webv2 test src/features/generation/core/graphCoverage.test.ts -u

    This rewrites __snapshots__/generateGraphNodeTypes.json. Then raise SUPPORTED_BASE_COUNTS["generate"] in tests/app/invocations/test_frontend_graph_node_types.py, and LITERAL_FLOORS if needed. That Python test checks every node type, edge field and literal value the builders emit against the backend’s invocation registry.


  • Backend config: add LoRA_LyCORIS_NewModel_Config(LoRA_LyCORIS_Config_Base, Config_Base) in configs/lora.py. The key heuristics of LoRA configs are loose, so make the new config and the existing ones exclude each other, by module names and by shapes (inner dimension). Cover this with stripped LoRA fixtures in tests/model_identification/stripped_models/.
  • Conversion: add a lora_model_from_new_model_state_dict in invokeai/backend/patches/lora_conversions/ and a branch for the base in load/model_loaders/lora.py. Handle every naming convention in circulation: diffusers/PEFT, Kohya, and ComfyUI, including fused projections.
  • Nodes: new_model_lora_loader and new_model_lora_collection_loader. The denoise node patches the transformer; on quantized weights it uses sidecar patching (requires_sidecar_patching).
  • webv2: addTransformerLoraCollectionLoader(...) in the builder. Add variant rules in isLoraCompatibleWithModel (settings.ts).
  • Backend: set FeaturesFacet.max_reference_images (and reference_images_require_variant if only one variant supports them). Wire the images into the text encoder and/or the denoise node, as the reference pipeline does.
  • webv2: extend the reference-image config union (types.ts), getDefaultReferenceImageConfig and getReferenceImageConfigSupported (baseGenerationPolicies.ts), and wire the images in the builder.
  • Control: set FeaturesFacet.control_kinds. Add the config (configs/controlnet.py, base class ControlNet_Checkpoint_Config_Base) and the node. In webv2, update canvas/controlValidation.ts, canvas/addControlLayers.ts and src/workbench/widgets/layers/controlModelOptions.ts.
  • Regional guidance: set FeaturesFacet.supports_regional_guidance and honor the conditioning mask in the denoise node. In webv2, update canvas/addRegionalGuidance.ts.

Declare PiDDecoderVariantType in the VariantFacet, and extend configs/pid_decoder.py, load/model_loaders/pid_decoder.py and backend/pid/decode.py. In webv2, update pid.ts and pidGraph.ts.


Several tests pin hand-maintained tables, so a new architecture has to extend them. They fail with a message that names the missing entry.

TestWhat to add
tests/backend/architectures/test_features.pyDENOISE_NODE: the node that owns width and height
tests/backend/architectures/test_guidance_range.pyGUIDANCE_FIELD (node and field the slider feeds), or NO_GUIDANCE_SLIDER
tests/backend/architectures/test_latent_space.pyDECLARED_LATENT_SPACES
tests/backend/architectures/test_default_settings.pyRows in DEFAULT_SETTINGS_MATRIX
tests/backend/architectures/test_conditioning.pyThe pinned numbers of conditioning types and architectures
tests/backend/architectures/test_variants.pyNothing if the variant lists and unique values are right; it checks them
tests/backend/architectures/test_modality.pyNothing if GENERATION_MODES matches ModalityFacet
tests/backend/architectures/test_capabilities_fixture.pyRegenerate the webv2 fixture (see section 11)
tests/app/invocations/test_frontend_graph_node_types.pySUPPORTED_BASE_COUNTS, possibly LITERAL_FLOORS
tests/backend/model_manager/load/test_diffusers_0XX_compatibility.pyNew classes, when the dependency is bumped

Beyond those tables, test what can silently go wrong:

  • detection of every format, including negative cases against similar architectures
  • key conversion against the diffusers module
  • the text-encoder template
  • the sampling schedule against the reference scheduler

Use real lightweight modules with small configs rather than mocks.

CI gates for generated artifacts. New or changed invocations change the OpenAPI schema. CI (openapi-checks.yml, typegen-checks.yml) requires the shared package’s invokeai/frontend/api/openapi.json and invokeai/frontend/api/schema.ts to be current. Regenerate both from invokeai/frontend/api in the repository’s Python environment after installing that package’s locked dependencies:

Terminal window
pnpm install --frozen-lockfile
set -o pipefail
python ../../../scripts/generate_openapi_schema.py > openapi.json && pnpm format:openapi
python ../../../scripts/generate_openapi_schema.py | pnpm typegen

Before review:

  • Ruff (uv tool run ruff@0.11.2 check and format --check)
  • the focused tests, then the full Python suite
  • webv2 lint:oxc, lint:tsc, architecture:check and Vitest
  • a real generation in every mode on a GPU, compared against the reference pipeline with the same seed and settings
  • the browser flows in webv2: install, generate, canvas modes, recall

A minimal txt2img integration with a Diffusers main model, reusing an existing VAE and encoder:

  • Directoryinvokeai
    • Directoryapp/invocations
      • fields.py NewModelConditioningField
      • primitives.py NewModelConditioningOutput
      • metadata.py GENERATION_MODES
      • Directorynew_model
        • __init__.py required, or the package’s nodes do not load
        • new_model_model_loader.py
        • new_model_denoise.py
      • Directorytext_encoder
        • new_model_text_encoder.py
      • Directoryvae
        • new_model_latents_to_image.py
    • app/services/model_records/model_records_base.py only with a variant enum
    • Directorybackend
      • Directoryarchitectures/defs
        • new_model.py the architecture declaration
      • Directorymodel_manager
        • taxonomy.py
        • Directoryconfigs
          • main.py
          • factory.py
        • Directoryload/model_loaders
          • new_model.py
        • Directorystarter_models
          • new_model.py
          • __init__.py
      • stable_diffusion/diffusion/conditioning_data.py
    • invocation_api/__init__.py
    • Directoryfrontend/webv2/src/features
      • Directorygeneration/core
        • contracts.ts
        • supportedBases.ts
        • graph.ts
        • baseGenerationPolicies.ts
        • __fixtures__/architectureCapabilities.json regenerated
        • __snapshots__/generateGraphNodeTypes.json regenerated
      • Directorymodels/core
        • baseIdentity.ts
        • types.ts
      • workflow/core/modelRequirements.ts
  • Directorytests
    • backend/architectures the tables in section 13
    • app/invocations/test_frontend_graph_node_types.py

For img2img, inpaint and outpaint, add:

  • Directoryinvokeai
    • Directoryapp/invocations/vae
      • new_model_image_to_latents.py
    • Directoryfrontend/webv2/src/features/generation/core/canvas
      • compileCanvasGraph.ts CANVAS_I2L_NODE_TYPES
      • compileCanvasGraph.test.ts BASE_CASES

Values are taken from the declarations under invokeai/backend/architectures/defs/.

FLUX.1FLUX.2Z-ImageQwen-ImageKrea-2
Latent spaceFLUX_16 (16 ch, 8×)FLUX2_32 (32 ch, 8×)FLUX_16 (16 ch, 8×)WAN21_16 (16 ch, 8×)WAN21_16 (16 ch, 8×)
Text encoderCLIP + T5Qwen3 (Klein), Mistral (Dev)Qwen3Qwen2.5-VL 7BQwen3-VL 4B
VAE acceptedFLUX.1FLUX.2FLUX.1Qwen-Image, AnimaQwen-Image, Anima
Guidance sliderGuidance (distilled)Guidance (max 20)CFG (Turbo: 1.0)CFG 4.0CFG (Turbo: 1.0)
Negative promptneverneverCFG-gatedCFG-gatedCFG-gated
dimension_grid1616161616
Scheduler setflowflowflow (Base: flow-no-lcm)standardflow
Reference images55—5 (edit variant)—
Codebackend/flux/backend/flux2/in the nodesin the nodesbackend/krea2/
This site was designed and developed by Aether Fox Studio.