New Model Type Integration Checklist
This guide covers every step needed to add a new model architecture to InvokeAI:
- the model manager, which identifies and loads the files
- the architecture declaration
- the invocations that form the generation graph
- the webv2 frontend
- the tests and CI gates that enforce completeness
1. Upstream dependencies
Section titled “1. Upstream dependencies”Most architectures load their transformer, VAE and text encoder through classes from diffusers and transformers. If the pinned releases lack the new model’s classes, handle the dependency first:
- Wait for a stable release that contains the classes. Do not pin a git commit;
pyproject.tomlpinsdiffusersto an exact version. - Bump in a separate change. A new release can change the numerical behavior of every existing architecture. The change has four parts:
- Update
pyproject.tomlanduv.lock, including raised transitive minimums such ashuggingface-hub. - Update
tests/backend/model_manager/load/test_diffusers_0XX_compatibility.py. It asserts the exact pinned version and the classes InvokeAI imports. - Run the full test suite.
- Run smoke generations for the existing architectures.
- Update
- Check the
transformersrequirement in the model card against the pin. Model cards often name the version the authors tested, not a hard minimum. Verify that the encoder and processor behave correctly on the pinned version before bumping it.
2. Taxonomy
Section titled “2. Taxonomy”File: invokeai/backend/model_manager/taxonomy.py
-
Add the
BaseModelTypemember.invokeai/backend/model_manager/taxonomy.py class BaseModelType(str, Enum):...Krea2 = "krea-2""""Indicates the model is associated with the Krea 2 model architecture, including Krea-2-Turbo."""NewModel = "new-model""""Indicates the model is associated with the NewModel architecture."""The value is stored in every user’s model database, so it cannot be renamed later. It also names the declaration module, with
-replaced by_, so it must yield a valid Python module name. Use lowercase letters, digits and-only, and no dots:qwen-image-2-1, notqwen-image-2.1. -
Add a variant enum (if needed).
Add one only when models of the architecture differ in ways that matter at load or generation time, for example distilled versus base. Variant strings are resolved without the base (
configs/factory.py), so every value must be unique across all variant enums. That is whyKrea2VariantType.Turbois"krea2_turbo"and not"turbo".invokeai/backend/model_manager/taxonomy.py class NewModelVariantType(str, Enum):"""NewModel variants."""Turbo = "new_model_turbo""""Distilled: few steps, CFG off."""Base = "new_model_base""""Undistilled base model: more steps, CFG on."""List the enum in all three places:
AnyVariantintaxonomy.pyvariant_type_adapterintaxonomy.pyModelRecordChanges.variantininvokeai/app/services/model_records/model_records_base.py
tests/backend/architectures/test_variants.pychecks that the three agree and that the values are unique. -
Add encoder types (only for a new text encoder).
There is no generic text-encoder type. Each encoder family has its own
ModelTypemember, and aModelFormatmember for its folder layout. Existing examples areQwen3Encoder,Qwen3VLEncoder,QwenVLEncoder,MistralEncoderandT5Encoder.Reuse an existing type whenever the encoder is a model InvokeAI already supports. Krea-2, Ideogram 4 and MiniMax H3 all use
ModelType.Qwen3VLEncoder, told apart byQwen3VLVariantType. Check whether the encoder weights are a stock checkpoint before adding a type.
3. Architecture declaration
Section titled “3. Architecture declaration”File: invokeai/backend/architectures/defs/new_model.py
Each architecture declares what it is in exactly one register(...) call. Modules under defs/ are discovered automatically, so there is no list to edit.
At boot, validate() runs from ApiDependencies.initialize. It refuses to start the app in two cases:
- a
BaseModelTypehas no module - a module omits a required facet
It checks only that the facets are present, not that their values are right.
| Facet | Required | What it declares | Read by |
|---|---|---|---|
LatentSpaceFacet | yes | Latent channels, spatial compression, latent→RGB projection | Step previews (app/util/step_callback.py) |
ConditioningFacet | yes | The *ConditioningInfo class the text encoder writes | The safe-globals list for conditioning deserialization |
DefaultSettingsFacet | yes | Steps, CFG or guidance, scheduler and size per variant (None is the fallback); optional by_name_hint | Stored on the model config at identification; prefills the UI |
ModalityFacet | yes | Modes (txt2img, img2img, inpaint, outpaint, video modes) and metadata_slug | Must equal GENERATION_MODES; which canvas tools the UI offers |
FeaturesFacet | yes | Negative prompt, dimension_grid, guidance label and range, scheduler set, control kinds, reference images, regional guidance, and more | Served to webv2 at GET /api/v2/models/capabilities |
VaeFacet | no | VAE bases (and channel counts) the architecture decodes with | Model loader VAE field, VAE picker, loader validation |
VariantFacet | no | Variant enum per ModelType | Variant consistency tests |
UNetDownscaleFacet | no | UNet downscale factor | SD-family UNets only |
The facet classes live in invokeai/backend/architectures/facets/. Modeled on defs/krea_2.py:
"""What the new-model architecture declares."""
from invokeai.backend.architectures.facets.conditioning import ConditioningFacetfrom invokeai.backend.architectures.facets.default_settings import DefaultSettingsFacetfrom invokeai.backend.architectures.facets.features import FeaturesFacet, NegativePromptfrom invokeai.backend.architectures.facets.latent_space import FLUX2_32, LatentSpaceFacetfrom invokeai.backend.architectures.facets.modality import ModalityFacetfrom invokeai.backend.architectures.facets.vae import VaeCompatibility, VaeFacetfrom invokeai.backend.architectures.facets.variant import VariantFacetfrom invokeai.backend.architectures.registry import registerfrom invokeai.backend.model_manager.configs.default_settings import MainModelDefaultSettingsfrom invokeai.backend.model_manager.taxonomy import BaseModelType, ModelType, NewModelVariantTypefrom invokeai.backend.stable_diffusion.diffusion.conditioning_data import NewModelConditioningInfo
register( BaseModelType.NewModel, # NewModel decodes with the FLUX.2 VAE, so it shares that latent space. LatentSpaceFacet(FLUX2_32), ConditioningFacet(NewModelConditioningInfo), DefaultSettingsFacet( { # Values from the model card. NewModelVariantType.Base: MainModelDefaultSettings( scheduler="euler", steps=28, cfg_scale=4.0, width=1024, height=1024 ), # Turbo, and any model whose variant is unknown. None: MainModelDefaultSettings(scheduler="euler", steps=8, cfg_scale=1.0, width=1024, height=1024), } ), ModalityFacet(frozenset({"txt2img", "img2img", "inpaint", "outpaint"}), metadata_slug="new_model"), FeaturesFacet( negative_prompt=NegativePrompt(visible=True, usage="cfg-gated"), dimension_grid=16, # equals `multiple_of` on new_model_denoise.width/height guidance_label="CFG", scheduler_set="flow", ), VaeFacet(frozenset({VaeCompatibility(BaseModelType.Flux2)})), # Variant values must be globally unique; see taxonomy.py. VariantFacet({ModelType.Main: NewModelVariantType}),)Things to get right:
-
Latent space. If the architecture uses a VAE InvokeAI already supports, reuse its shared space from
facets/latent_space.py:FLUX_16: FLUX.1, Z-ImageFLUX2_32: FLUX.2, ERNIE-Image, Ideogram 4WAN21_16: Wan 2.1, Qwen-Image, Krea-2, AnimaWAN22_48
Otherwise, declare a new
LatentSpacein that file. Compute its projection withscripts/generate_vae_linear_approximation.pyrather than guessing it. -
Grid and guidance are pinned to the denoise node by tests:
dimension_gridmust equal themultiple_ofon the node’s width and height.guidance_minandguidance_maxmust match thegeandleof the guidance field.- A constraint that the node only enforces inside
invoke(), per variant, goes indimension_grid_by_variant.
-
VaeFacetreplaces the default rule, under which an architecture accepts only its own base. List every accepted base explicitly, including the architecture’s own if it applies. Z-Image, for example, accepts only FLUX VAEs. -
Mode strings are persisted in image metadata and cannot be renamed.
metadata_slugis usually the base value with_(z_image), but not always (krea2). -
Default settings are copied onto the model config when a model is identified (
configs/factory.py), so model configs need no code for them. Useby_name_hintonly for sub-models that nothing on disk can tell apart.
4. Model configs
Section titled “4. Model configs”Folder: invokeai/backend/model_manager/configs/
Configs identify a model on disk. Every non-abstract subclass of Config_Base registers itself through Config_Base.__init_subclass__; there is no decorator.
-
Main model configs (
configs/main.py)One class per format. Each combines a format mixin,
Main_Config_BaseandConfig_Base:invokeai/backend/model_manager/configs/main.py class Main_Diffusers_NewModel_Config(Diffusers_Config_Base, Main_Config_Base, Config_Base):"""Model config for NewModel diffusers models."""base: Literal[BaseModelType.NewModel] = Field(BaseModelType.NewModel)variant: NewModelVariantType = Field()@classmethoddef from_model_on_disk(cls, mod: ModelOnDisk, override_fields: dict[str, Any]) -> Self:raise_if_not_dir(mod)raise_for_override_fields(cls, override_fields)# The pipeline class name implies the base.raise_for_class_name(common_config_paths(mod.path), {"NewModelPipeline"})# `_get_variant` reads whatever distinguishes the variants, e.g. a flag in model_index.json.variant = override_fields.pop("variant", None) or cls._get_variant(mod)repo_variant = override_fields.pop("repo_variant", None) or cls._get_repo_variant_or_raise(mod)return cls(**override_fields, variant=variant, repo_variant=repo_variant)class Main_Checkpoint_NewModel_Config(Checkpoint_Config_Base, Main_Config_Base, Config_Base):"""Model config for NewModel single-file checkpoints."""base: Literal[BaseModelType.NewModel] = Field(default=BaseModelType.NewModel)format: Literal[ModelFormat.Checkpoint] = Field(default=ModelFormat.Checkpoint)variant: NewModelVariantType = Field()@classmethoddef from_model_on_disk(cls, mod: ModelOnDisk, override_fields: dict[str, Any]) -> Self:raise_if_not_file(mod)raise_for_override_fields(cls, override_fields)state_dict = mod.load_state_dict()if not _has_new_model_keys(state_dict):raise NotAMatchError("state dict does not look like a NewModel model")if _has_ggml_tensors(state_dict):raise NotAMatchError("state dict looks like GGUF quantized")variant = override_fields.pop("variant", None) or _get_new_model_variant_from_name(mod.path.name)return cls(**override_fields, variant=variant)If GGUF files exist, add a
Main_GGUF_NewModel_Configwithformat: Literal[ModelFormat.GGUFQuantized]that requires_has_ggml_tensors.The identification helpers are in
configs/identification_utils.py:NotAMatchError: “not this config”InvalidMatchError: “this config, but the file is broken”raise_if_not_file/raise_if_not_dirraise_for_override_fieldsraise_for_class_namecommon_config_pathsstate_dict_has_any_keys_*
-
Detection helpers
Detect the architecture from keys that only it has, and from shapes where keys are shared. Strip ComfyUI prefixes such as
model.diffusion_model.from each key first, with_strip_comfyui_key_prefixinconfigs/main.py. It reads the sharedCOMFYUI_KEY_PREFIXES, so detection and the loaders agree on which prefixes exist. Exclude LoRA suffixes, so that a LoRA for the architecture is not mistaken for a main model.invokeai/backend/model_manager/configs/main.py def _has_new_model_keys(state_dict: dict[str | int, Any]) -> bool:"""True for NewModel transformer weights, False for its LoRAs and for every other architecture."""# `txt_in.text_norm` is unique to NewModel; `img_in` alone is shared with Qwen-Image and Krea-2....Configs must exclude each other. Identification iterates
Config_Base.CONFIG_CLASSES, which is a set, so the order in theAnyModelConfigunion decides nothing. When two configs match the same file,matches_sort_keybreaks the tie, which amounts to chance. For every existing config whose heuristic the new files could satisfy, add a negative check on one side or both. The same applies to LoRA and VAE configs, whose heuristics are often loose, for example “any key starting withtransformer_blocks.”. Cover each exclusion with a detection test. -
VAE config (only for a new VAE) (
configs/vae.py)VAE_Checkpoint_Config_BaseandVAE_Diffusers_Config_Basedetect the plainAutoencoderKLVAEs: SD 1, SD 2 (single files only) and SDXL at 4 latent channels, FLUX.1 and SD3 at 16. Within one latent width the weights cannot tell these bases apart, so_VAE_FAMILIESgroups them, and an explicitbaseoverride chooses within a family but never across one. For 16 channels, a folder’sconfig.jsondecides byscaling_factor/shift_factor, and a config that matches neither is not filed at all. A 16-channel single file goes to the backbone named in its file name, folder or install source (configs/backbone_names.py), else to FLUX.1. At 4 channels the older heuristics still apply: the file name for a single file, the SDXL config values or name for a folder. A new architecture that reuses one of these networks with its own constants (CogView 4 would) belongs in that table and in the diffusers constants check, not in a new detector.Newer architectures with their own network subclass
Checkpoint_Config_Base, Config_Basedirectly, asVAE_Checkpoint_Flux2_Configdoes:invokeai/backend/model_manager/configs/vae.py class VAE_Checkpoint_NewModel_Config(Checkpoint_Config_Base, Config_Base):"""Model config for NewModel VAE checkpoints."""type: Literal[ModelType.VAE] = Field(default=ModelType.VAE)format: Literal[ModelFormat.Checkpoint] = Field(default=ModelFormat.Checkpoint)base: Literal[BaseModelType.NewModel] = Field(default=BaseModelType.NewModel)cpu_only: bool | None = Field(default=None, description="Whether this model should run on CPU only")@classmethoddef from_model_on_disk(cls, mod: ModelOnDisk, override_fields: dict[str, Any]) -> Self:raise_if_not_file(mod)raise_for_override_fields(cls, override_fields)if not _is_new_model_vae(mod.load_state_dict()):raise NotAMatchError("state dict does not look like a NewModel VAE")return cls(**override_fields)Then exclude the new VAE in
VAE_Checkpoint_Config_Base._validate_looks_like_vae, and in every family detector it could also satisfy. Wan, Qwen-Image and Anima VAEs share one layout family. -
Text encoder config (only for a new encoder type)
Create
configs/<encoder>.pywithbase=BaseModelType.Any, the newModelType, and one config per format: folder (its ownModelFormat), single-file checkpoint, and GGUF if it exists.qwen3_encoder.py,qwen3_vl_encoder.pyandmistral_encoder.pyare complete examples. -
AnyModelConfigunion (configs/factory.py)Add each new config to the union, in the shape the other entries use:
invokeai/backend/model_manager/configs/factory.py Annotated[Main_Diffusers_NewModel_Config, Main_Diffusers_NewModel_Config.get_tag()],
5. Model loaders
Section titled “5. Model loaders”Folder: invokeai/backend/model_manager/load/model_loaders/
Loaders turn a config into an in-memory model. Every module in this folder is imported automatically. A loader registers for a (base, type, format) triple. Lookup falls back to base=BaseModelType.Any, which is how encoder loaders serve every architecture.
-
Diffusers-format main model
Subclass
GenericDiffusersLoader. It resolves each submodel’s class frommodel_index.json. Override_load_modelonly for what the pipeline gets wrong: tokenizer quirks, config patches, dtype, or dropping unused modules.invokeai/backend/model_manager/load/model_loaders/new_model.py @ModelLoaderRegistry.register(base=BaseModelType.NewModel, type=ModelType.Main, format=ModelFormat.Diffusers)class NewModelDiffusersModel(GenericDiffusersLoader):"""Loads NewModel submodels from a diffusers pipeline folder."""def _load_model(self, config: AnyModelConfig, submodel_type: Optional[SubModelType] = None) -> AnyModel:if submodel_type is None:raise Exception("A submodel type must be provided when loading main pipelines.")model_path = Path(config.path)load_class = self.get_hf_load_class(model_path, submodel_type)dtype = TorchDevice.choose_bfloat16_safe_dtype(TorchDevice.choose_torch_device())result: AnyModel = load_class.from_pretrained(model_path / submodel_type.value, torch_dtype=dtype)return self._apply_fp8_layerwise_casting(result, config, submodel_type) -
Single-file and GGUF main models
Subclass
ModelLoader, build the transformer underaccelerate.init_empty_weights(), convert the keys to the diffusers layout, and load withassign=True.Krea2CheckpointModelandKrea2GGUFCheckpointModelinkrea2.pyshow the full pattern:- stripping the ComfyUI prefix with
CheckpointPrefix.detect(sd).strip(sd)(invokeai/backend/model_manager/checkpoint_prefix.py) rather than a hand-written loop - reserving RAM before materializing, exactly once: build the empty model first, take the fp8 and nvfp4 side channels out of the state dict, then reserve, and only then fold or widen anything. A fold before the reservation lands on a cache that was never asked for the room. For fp8, nvfp4 and dense files the reservation is
reserve_for_load(invokeai/backend/quantization/load_plan.py). - quantized formats: ComfyUI
fp8_scaled,int8_convrotandnvfp4(helpers ininvokeai/backend/quantization/), and GGUF throughgguf_sd_loader. For int8, callinstall_int8_convrot_layersinstead of its individual steps. It makes the reservation itself through itsreserveargument, so do not also callreserve_for_load; the steps’ order matters, and skipping the first one loads a mixed checkpoint’s fp8 weights without their scale. - a loader that handles no quantization side channel (a VAE, an adapter, a non-strict
load_state_dict) callsreject_quantized_side_channelso a quantized file is refused rather than loaded wrongly - opting into FP8 storage with
_apply_fp8_layerwise_casting
Declare FP8 Storage support at registration. The Model Manager only shows the FP8 Storage toggle where
fp8_storage_verdict(base, type, format)(invokeai/backend/model_manager/load/fp8_capability.py) says it works. A loader that does not apply the cast says so where it registers, with a reason:fp8_storage=Unimplemented("...")for work not done yet,NotApplicable("...")for a deliberate decision.tests/backend/model_manager/load/test_fp8_capability.pywalks the registry and fails for a loader that neither casts nor declares.Community single-file checkpoints often use key names that differ from diffusers, for example fused projections. Verify the conversion numerically against the diffusers module on a small config.
- stripping the ComfyUI prefix with
-
VAE loader (only for a new VAE)
Build the diffusers VAE class and load the converted state dict.
vae.pyandflux.pyhandle existing VAEs from native and ComfyUI layouts. -
Text encoder loader (only for a new encoder type)
Register with
base=BaseModelType.Anyand the encoder’sModelType, and serve bothSubModelType.TextEncoderandSubModelType.Tokenizer. Vendor tokenizers and configs the loader needs, asinvokeai/backend/qwen2_5_vl/does, so a single-file install loads without network access. Add the vendored files to[tool.setuptools.package-data]inpyproject.toml, or a wheel install fails withFileNotFoundError.
6. Conditioning plumbing
Section titled “6. Conditioning plumbing”A text encoder saves a ConditioningFieldData to disk, and the denoise node reads it back. Each architecture has its own conditioning types, and they must be threaded through five places:
-
Info class and union (
invokeai/backend/stable_diffusion/diffusion/conditioning_data.py)Add a dataclass with a
to(device, dtype)method, and add it to theConditioningFieldData.conditioningsunion. TheConditioningFacetthen puts it on the safe-globals list that deserialization needs. A class missing from that list fails with anUnpicklingErrorinside the denoise node, after the encoder has already run.invokeai/backend/stable_diffusion/diffusion/conditioning_data.py @dataclassclass NewModelConditioningInfo:"""NewModel text conditioning."""prompt_embeds: torch.Tensor"""Shape: (batch_size, seq_len, hidden_size)."""prompt_embeds_mask: torch.Tensor | None = None"""Shape: (batch_size, seq_len). True for valid tokens."""def to(self, device: torch.device | None = None, dtype: torch.dtype | None = None):self.prompt_embeds = self.prompt_embeds.to(device=device, dtype=dtype)if self.prompt_embeds_mask is not None:self.prompt_embeds_mask = self.prompt_embeds_mask.to(device=device)return self -
Field (
invokeai/app/invocations/fields.py):NewModelConditioningFieldwithconditioning_nameand an optionalmaskfor regional prompting. Add node field descriptions toFieldDescriptionsin the same file. -
Output (
invokeai/app/invocations/primitives.py):NewModelConditioningOutput, registered with@invocation_output("new_model_conditioning_output"). -
Encoder field (
invokeai/app/invocations/model.py, only for a new encoder type): for exampleQwen3VLEncoderField, withtokenizer,text_encoderandloras. -
Public API (
invokeai/invocation_api/__init__.py): export the info class, and list it in__all__.
7. Invocations
Section titled “7. Invocations”Invocations expose the architecture as nodes in the generation graph.
-
Model loader (
new_model/new_model_model_loader.py)Outputs the submodel identifiers. The transformer always comes from the main model; the VAE and encoder come from standalone models if selected, otherwise from the Diffusers main model. The field hints
ui_model_baseandui_model_typedrive the model pickers. The VAE field must useaccepted_vae_bases(), whichtests/backend/architectures/test_vae.pyenforces.invokeai/app/invocations/new_model/new_model_model_loader.py @invocation_output("new_model_model_loader_output")class NewModelModelLoaderOutput(BaseInvocationOutput):transformer: TransformerField = OutputField(description=FieldDescriptions.transformer, title="Transformer")qwen3_vl_encoder: Qwen3VLEncoderField = OutputField(description=FieldDescriptions.qwen3_vl_encoder, title="Qwen3-VL Encoder")vae: VAEField = OutputField(description=FieldDescriptions.vae, title="VAE")@invocation("new_model_model_loader", title="Main Model - NewModel", tags=["model", "new_model"],category="model", version="1.0.0", classification=Classification.Prototype)class NewModelModelLoaderInvocation(BaseInvocation):model: ModelIdentifierField = InputField(description=FieldDescriptions.main_model, input=Input.Direct,ui_model_base=BaseModelType.NewModel, ui_model_type=ModelType.Main, title="Transformer",)vae_model: Optional[ModelIdentifierField] = InputField(default=None, description="Standalone VAE model.", input=Input.Direct,ui_model_base=accepted_vae_bases(BaseModelType.NewModel), ui_model_type=ModelType.VAE, title="VAE",)qwen3_vl_encoder_model: Optional[ModelIdentifierField] = InputField(default=None, description="Standalone Qwen3-VL Encoder model.", input=Input.Direct,ui_model_type=ModelType.Qwen3VLEncoder, title="Qwen3-VL Encoder",)def invoke(self, context: InvocationContext) -> NewModelModelLoaderOutput:transformer = self.model.model_copy(update={"submodel_type": SubModelType.Transformer})# Validate each standalone component (base, type, variant, accepts_vae(...)) with an error# that names what the user picked; fall back to the Diffusers main model's submodels.... -
Text encoder (
text_encoder/new_model_text_encoder.py)Loads tokenizer and encoder through
context.models.load(...), applies LoRAs to the encoder if the architecture has any, extracts the hidden states the transformer was trained on, and saves the conditioning:invokeai/app/invocations/text_encoder/new_model_text_encoder.py @invocation("new_model_text_encoder", title="Prompt - NewModel", tags=["prompt", "conditioning", "new_model"],category="conditioning", version="1.0.0", classification=Classification.Prototype,idle_gpu_offloadable=True)class NewModelTextEncoderInvocation(BaseInvocation):prompt: str = InputField(description="Text prompt.", ui_component=UIComponent.Textarea)mask: TensorField | None = InputField(default=None, description="Regional mask.", input=Input.Connection)qwen3_vl_encoder: Qwen3VLEncoderField = InputField(title="Qwen3-VL Encoder", description=FieldDescriptions.qwen3_vl_encoder, input=Input.Connection)@torch.no_grad()def invoke(self, context: InvocationContext) -> NewModelConditioningOutput:prompt_embeds, prompt_mask = self._encode(context) # template, layer choice, paddingconditioning_data = ConditioningFieldData(conditionings=[NewModelConditioningInfo(prompt_embeds=prompt_embeds.cpu(), prompt_embeds_mask=prompt_mask)])conditioning_name = context.conditioning.save(conditioning_data)return NewModelConditioningOutput.build(conditioning_name, mask=self.mask)Copy the reference pipeline’s encoding exactly and cover it with a test: the prompt template, the prefix tokens it drops, the hidden-state layer (and whether it is before or after the final norm), the padding side and the maximum length. Small deviations here degrade images without failing.
-
Denoise (
new_model/new_model_denoise.py)invokeai/app/invocations/new_model/new_model_denoise.py @invocation("new_model_denoise", title="Denoise - NewModel", tags=["image", "new_model"],category="image", version="1.0.0", classification=Classification.Prototype)class NewModelDenoiseInvocation(BaseInvocation):latents: Optional[LatentsField] = InputField(default=None, description=FieldDescriptions.latents,input=Input.Connection)denoise_mask: Optional[DenoiseMaskField] = InputField(default=None, description=FieldDescriptions.denoise_mask,input=Input.Connection)denoising_start: float = InputField(default=0.0, ge=0, le=1, description=FieldDescriptions.denoising_start)denoising_end: float = InputField(default=1.0, ge=0, le=1, description=FieldDescriptions.denoising_end)transformer: TransformerField = InputField(description=FieldDescriptions.transformer, input=Input.Connection)positive_conditioning: NewModelConditioningField | list[NewModelConditioningField] = InputField(description=FieldDescriptions.positive_cond, input=Input.Connection)negative_conditioning: NewModelConditioningField | list[NewModelConditioningField] | None = InputField(default=None, description=FieldDescriptions.negative_cond, input=Input.Connection)cfg_scale: float | list[float] = InputField(default=1.0, description=FieldDescriptions.cfg_scale)# `multiple_of` must equal FeaturesFacet.dimension_grid.width: int = InputField(default=1024, gt=0, multiple_of=16, description="Width of the generated image.")height: int = InputField(default=1024, gt=0, multiple_of=16, description="Height of the generated image.")steps: int = InputField(default=8, gt=0, description=FieldDescriptions.steps)seed: int = InputField(default=0, description="Randomness seed for reproducibility.")def invoke(self, context: InvocationContext) -> LatentsOutput:# 1. Noise from the seed; init latents for img2img; clip the schedule by denoising_start/end.# 2. Load the transformer with a working-memory estimate; apply LoRA patches.# 3. Loop: model prediction, optional CFG, scheduler step,# RectifiedFlowInpaintExtension merge when a denoise_mask is set, step callback.# 4. Save and return the latents....Report progress with
context.util.sd_step_callback(state, BaseModelType.NewModel); it renders previews through the architecture’sLatentSpaceFacet. The denoise node is also where these are checked: interrupts (context.util.is_canceled()), inpainting, img2img throughdenoising_start, and a working-memory estimate passed tomodel_on_device(working_mem_bytes=...). -
VAE encode and decode (
vae/new_model_image_to_latents.py,vae/new_model_latents_to_image.py)Apply the VAE’s latent normalization, scaling and shift or per-channel mean and std, exactly as the reference pipeline does. Handle tiling and give a working-memory estimate. The estimate helpers live in
invokeai/backend/util/vae_working_memory.py, with calibration scripts inscripts/calibrate_*_working_memory.py. Node types conventionally end in_i2land_l2i. -
LoRA loaders (if LoRA is supported, see Optional features)
new_model_lora_loaderandnew_model_lora_collection_loader. webv2 uses the collection loader.
8. Sampling
Section titled “8. Sampling”- Backend package (optional). Put reusable math in
invokeai/backend/new_model/: noise, packing, position IDs, schedules and text encoding. FLUX, FLUX.2 and Krea-2 have one; Qwen-Image and Z-Image keep the loop in the denoise node. Use a package once code is shared between nodes or needs its own tests. - Schedulers. Flow-matching architectures use diffusers’
FlowMatchEulerDiscreteScheduler, or the shared maps ininvokeai/backend/flux/schedulers.py:FLUX_SCHEDULER_MAP,ZIMAGE_SCHEDULER_MAPandERNIE_IMAGE_SCHEDULER_MAP, each with name values and labels. Declare which set the UI offers inFeaturesFacet.scheduler_set. Setscheduler_applies_to_graph=Trueonly if the denoise node really takes a scheduler field. Reproduce the reference sigma schedule and shift (mu, time-shift type, terminal shift) exactly, and test it against the reference scheduler. - External noise (optional). Only if the denoise node accepts a noise tensor: extend
LatentNoiseTypeand its shape and grid logic ininvokeai/app/invocations/latent_noise.py, and extendnoise_typeininvocations/noise.py. Validate withvalidate_noise_tensor_shape. - Inpainting. Rectified-flow models use
RectifiedFlowInpaintExtension(invokeai/backend/rectified_flow/rectified_flow_inpaint_extension.py). - Previews.
LatentSpaceFacetprovides them; there is nothing to add in the step callback.
9. Metadata and generation modes
Section titled “9. Metadata and generation modes”File: invokeai/app/invocations/metadata.py
-
Add the mode strings to
GENERATION_MODES.invokeai/app/invocations/metadata.py GENERATION_MODES = Literal[..."new_model_txt2img","new_model_img2img","new_model_inpaint","new_model_outpaint",]They must equal the
ModalityFacetdeclaration:<metadata_slug>_<mode>for each mode.tests/backend/architectures/test_modality.pycompares the two. webv2 writes<slug>_txt2img, and the canvas derives the other modes by replacingtxt2img, so declare all four image modes if the canvas supports them. -
Metadata fields (if needed).
CoreMetadataInvocation(core_metadata) accepts extra fields, so architecture-specific keys such as the chosen encoder or VAE need no declaration. Declare a field only if it needs validation. Adding one requires a node version bump and follows theCORE_METADATA_VERSIONrules documented in Media Metadata.
10. Starter models
Section titled “10. Starter models”Folder: invokeai/backend/model_manager/starter_models/
The catalogue is a package with one module per architecture:
types.pyholdsStarterModel.common.pyholds encoders and upscalers that several architectures share. Import a shared dependency from there instead of declaring a second copy; if an encoder you need is declared in another architecture’s module, move it tocommon.py.
"""NewModel starter models."""
from invokeai.backend.model_manager.starter_models.common import qwen3_vl_encoder_4bfrom invokeai.backend.model_manager.starter_models.types import StarterModelfrom invokeai.backend.model_manager.taxonomy import BaseModelType, ModelFormat, ModelType, NewModelVariantType
new_model_vae = StarterModel( name="NewModel VAE", base=BaseModelType.NewModel, source="org/NewModel::vae/diffusion_pytorch_model.safetensors", # a file inside a repo description="NewModel VAE. ~300MB", type=ModelType.VAE,)
new_model_turbo = StarterModel( name="NewModel Turbo", base=BaseModelType.NewModel, source="org/NewModel-Turbo", # a whole diffusers repo description="NewModel Turbo, full diffusers pipeline. ~20GB", type=ModelType.Main, variant=NewModelVariantType.Turbo,)
new_model_turbo_gguf_q4_k_m = StarterModel( name="NewModel Turbo (Q4_K_M GGUF)", base=BaseModelType.NewModel, source="https://huggingface.co/org/NewModel-Turbo-GGUF/resolve/main/new_model_turbo-Q4_K_M.gguf", description="GGUF ships only the transformer; the VAE and encoder are installed as dependencies. ~7GB", type=ModelType.Main, format=ModelFormat.GGUFQuantized, variant=NewModelVariantType.Turbo, dependencies=[new_model_vae, qwen3_vl_encoder_4b],)Then import the models in the package’s __init__.py:
- Add them to
STARTER_MODELS. It is a curated order, the order the install dialog shows, so insert them where they belong instead of appending. - Optionally add a
STARTER_BUNDLESentry keyed by the base. - Keep
sourcevalues unique. The module asserts that.
Put license restrictions in the description; the Ideogram 4 and MiniMax H3 descriptions are examples.
11. Frontend (webv2)
Section titled “11. Frontend (webv2)”webv2 does not keep per-architecture tables of its own. Grid, default settings, negative-prompt behavior, guidance label and range, scheduler set, VAE acceptance, control kinds and reference-image limits all come from the backend’s FeaturesFacet, DefaultSettingsFacet and VaeFacet, served at GET /api/v2/models/capabilities. What webv2 owns is the graph topology and the per-family UI.
Paths below are relative to invokeai/frontend/webv2/src/features/generation/core/ unless stated otherwise.
-
Regenerate the capabilities fixture
Terminal window REGEN_CAPABILITIES_FIXTURE=1 uv run --no-sync pytest tests/backend/architectures/test_capabilities_fixture.pyThis rewrites
__fixtures__/architectureCapabilities.json. The mock backend and the unit tests read it. -
Mark the base as buildable
- Add it to
KnownGenerationModelBaseincontracts.ts. - Add it to
SUPPORTED_GENERATE_BASESinsupportedBases.ts. - Update the pinned list in
supportedBases.test.ts.
- Add it to
-
Write the graph builder (
graph.ts)Add
buildNewModelGraphand register it inGRAPH_BUILDERS, whichsatisfies Record<SupportedGenerateBase, …>, so the compiler reports a missing entry.buildZImageGraphandbuildKrea2Graphare good templates:invokeai/frontend/webv2/src/features/generation/core/graph.ts const buildNewModelGraph = (settings: GenerateSettings,model: MainModelConfig,outputIsIntermediate: boolean,projectSettings: GenerationProjectSettings): BackendGraphContract => {const vaeModel = getCompatibleVae(settings, model);const graph: BackendGraphContract = { edges: [], id: createId('new_model_graph'), nodes: {} };const { negativePrompt, positivePrompt, seed } = addPromptAndSeedNodes(graph);const useCfg = settings.cfgScale > 1;const modelLoader = addNode(graph, {id: 'model_loader',model,type: 'new_model_model_loader',vae_model: vaeModel ?? undefined,});const posCond = addNode(graph, { id: 'pos_cond', type: 'new_model_text_encoder' });const posCondCollect = addNode(graph, { id: 'pos_cond_collect', type: 'collect' });const size = getDenoiseSize(settings, model);const denoise = addNode(graph, {cfg_scale: settings.cfgScale,denoising_end: 1,denoising_start: 0,height: size.height,id: 'denoise_latents',steps: settings.steps,type: 'new_model_denoise',width: size.width,});// Creates the `canvas_output` node.const output = addImageOutputNode(graph, 'new_model_l2i', outputIsIntermediate);addEdge(graph, modelLoader, 'transformer', denoise, 'transformer');addEdge(graph, modelLoader, 'qwen3_vl_encoder', posCond, 'qwen3_vl_encoder');addEdge(graph, positivePrompt, 'value', posCond, 'prompt');addEdge(graph, posCond, 'conditioning', posCondCollect, 'item');addEdge(graph, posCondCollect, 'collection', denoise, 'positive_conditioning');// …the same chain for `negative_prompt` → `neg_cond` → `neg_cond_collect` when useCfg…addEdge(graph, seed, 'value', denoise, 'seed');addEdge(graph, denoise, 'latents', output, 'latents');addEdge(graph, modelLoader, 'vae', output, 'vae');addMetadata(graph, output, settings, model, 'new_model_txt2img', projectSettings, {vae: vaeModel ?? undefined,});return graph;};The builder produces the txt2img graph only; the canvas grafts img2img, inpaint and outpaint onto it. That depends on these conventions:
- Node ids
positive_prompt,negative_prompt,seed,model_loader,pos_cond,neg_cond,pos_cond_collect,neg_cond_collectanddenoise_latents. - An output node
canvas_outputwhosevaeinput has an edge (requireVaeSource). addDecodeOutputif the architecture supports PiD decoding, otherwiseaddImageOutputNode.
Validation that only one architecture needs belongs in
getModelFamilyValidationReasons(baseGenerationPolicies.ts). - Node ids
-
Component pickers (encoder, VAE, Diffusers component source)
- Add a
casetogetBaseComponentSectionPolicyinbaseGenerationPolicies.ts. Itsdefaultreturns an empty policy silently. - Add encoder filters in
componentCompatibility.ts. If a main model can bring its own components, updateisBundledMainForBase. VAE acceptance comes from the backend. - A new settings key must be added in all of these places:
GenerateSettings(types.ts)GenerateComponentValueKey,COMPONENT_SETTING_LABELS,getComponentPolicyContextandgetDefaultGenerateSettings(baseGenerationPolicies.ts)normalizeGenerateSettingsandcloneGenerateWidgetValues(settings.ts)
- Add a
-
Canvas
- Map the encode node in
CANVAS_I2L_NODE_TYPES(canvas/compileCanvasGraph.ts). The map isPartial, so a missing entry only fails at runtime. - If the architecture uses the
strength^0.2denoising-start curve, add it tocanvasDenoisingStart. - Add a case to
BASE_CASESincanvas/compileCanvasGraph.test.ts.
- Map the encode node in
-
Recall (
src/workbench/image-actions/)- In
imageRecall.ts, add the component keys the builder writes into metadata tohasComponentModelsand to the component patch. - Add the base to
GUIDANCE_BASESinrecallParameters.tsif it takes the shared guidance slider.
Nothing checks recall against what the builder writes, so test it.
- In
-
Model registry and labels
src/features/models/core/baseIdentity.ts:MODEL_BASES, plus its pinned testsrc/features/models/core/types.ts:ModelBasesrc/features/models/core/taxonomy.ts: variant valuessrc/features/models/core/relationships.ts: which encoder and VAE types may link to the basesrc/features/workflow/core/modelRequirements.ts:BASE_LABELS- New user-facing strings go in
invokeai/frontend/webv2/public/locales/en.json.
-
Per-family UI (only for controls unique to the architecture)
src/features/generation/ui/GenerateRenderSection.tsxbranches on the family for controls such as Krea-2 seed variance and Wan settings. Controls every architecture has come from capabilities; they need no code. -
Regenerate the graph contract
Terminal window pnpm -C invokeai/frontend/webv2 test src/features/generation/core/graphCoverage.test.ts -uThis rewrites
__snapshots__/generateGraphNodeTypes.json. Then raiseSUPPORTED_BASE_COUNTS["generate"]intests/app/invocations/test_frontend_graph_node_types.py, andLITERAL_FLOORSif needed. That Python test checks every node type, edge field and literal value the builders emit against the backend’s invocation registry.
12. Optional features
Section titled “12. Optional features”- Backend config: add
LoRA_LyCORIS_NewModel_Config(LoRA_LyCORIS_Config_Base, Config_Base)inconfigs/lora.py. The key heuristics of LoRA configs are loose, so make the new config and the existing ones exclude each other, by module names and by shapes (inner dimension). Cover this with stripped LoRA fixtures intests/model_identification/stripped_models/. - Conversion: add a
lora_model_from_new_model_state_dictininvokeai/backend/patches/lora_conversions/and a branch for the base inload/model_loaders/lora.py. Handle every naming convention in circulation: diffusers/PEFT, Kohya, and ComfyUI, including fused projections. - Nodes:
new_model_lora_loaderandnew_model_lora_collection_loader. The denoise node patches the transformer; on quantized weights it uses sidecar patching (requires_sidecar_patching). - webv2:
addTransformerLoraCollectionLoader(...)in the builder. Add variant rules inisLoraCompatibleWithModel(settings.ts).
Reference images
Section titled “Reference images”- Backend: set
FeaturesFacet.max_reference_images(andreference_images_require_variantif only one variant supports them). Wire the images into the text encoder and/or the denoise node, as the reference pipeline does. - webv2: extend the reference-image config union (
types.ts),getDefaultReferenceImageConfigandgetReferenceImageConfigSupported(baseGenerationPolicies.ts), and wire the images in the builder.
Control and regional guidance
Section titled “Control and regional guidance”- Control: set
FeaturesFacet.control_kinds. Add the config (configs/controlnet.py, base classControlNet_Checkpoint_Config_Base) and the node. In webv2, updatecanvas/controlValidation.ts,canvas/addControlLayers.tsandsrc/workbench/widgets/layers/controlModelOptions.ts. - Regional guidance: set
FeaturesFacet.supports_regional_guidanceand honor the conditioningmaskin the denoise node. In webv2, updatecanvas/addRegionalGuidance.ts.
PiD decoding
Section titled “PiD decoding”Declare PiDDecoderVariantType in the VariantFacet, and extend configs/pid_decoder.py, load/model_loaders/pid_decoder.py and backend/pid/decode.py. In webv2, update pid.ts and pidGraph.ts.
13. Tests and CI gates
Section titled “13. Tests and CI gates”Several tests pin hand-maintained tables, so a new architecture has to extend them. They fail with a message that names the missing entry.
| Test | What to add |
|---|---|
tests/backend/architectures/test_features.py | DENOISE_NODE: the node that owns width and height |
tests/backend/architectures/test_guidance_range.py | GUIDANCE_FIELD (node and field the slider feeds), or NO_GUIDANCE_SLIDER |
tests/backend/architectures/test_latent_space.py | DECLARED_LATENT_SPACES |
tests/backend/architectures/test_default_settings.py | Rows in DEFAULT_SETTINGS_MATRIX |
tests/backend/architectures/test_conditioning.py | The pinned numbers of conditioning types and architectures |
tests/backend/architectures/test_variants.py | Nothing if the variant lists and unique values are right; it checks them |
tests/backend/architectures/test_modality.py | Nothing if GENERATION_MODES matches ModalityFacet |
tests/backend/architectures/test_capabilities_fixture.py | Regenerate the webv2 fixture (see section 11) |
tests/app/invocations/test_frontend_graph_node_types.py | SUPPORTED_BASE_COUNTS, possibly LITERAL_FLOORS |
tests/backend/model_manager/load/test_diffusers_0XX_compatibility.py | New classes, when the dependency is bumped |
Beyond those tables, test what can silently go wrong:
- detection of every format, including negative cases against similar architectures
- key conversion against the diffusers module
- the text-encoder template
- the sampling schedule against the reference scheduler
Use real lightweight modules with small configs rather than mocks.
CI gates for generated artifacts. New or changed invocations change the OpenAPI schema. CI (openapi-checks.yml, typegen-checks.yml) requires the shared package’s invokeai/frontend/api/openapi.json and invokeai/frontend/api/schema.ts to be current. Regenerate both from invokeai/frontend/api in the repository’s Python environment after installing that package’s locked dependencies:
pnpm install --frozen-lockfileset -o pipefailpython ../../../scripts/generate_openapi_schema.py > openapi.json && pnpm format:openapipython ../../../scripts/generate_openapi_schema.py | pnpm typegenBefore review:
- Ruff (
uv tool run ruff@0.11.2 checkandformat --check) - the focused tests, then the full Python suite
- webv2
lint:oxc,lint:tsc,architecture:checkand Vitest - a real generation in every mode on a GPU, compared against the reference pipeline with the same seed and settings
- the browser flows in webv2: install, generate, canvas modes, recall
Summary: Minimal integration
Section titled “Summary: Minimal integration”A minimal txt2img integration with a Diffusers main model, reusing an existing VAE and encoder:
Directoryinvokeai
Directoryapp/invocations
- fields.py
NewModelConditioningField - primitives.py
NewModelConditioningOutput - metadata.py
GENERATION_MODES Directorynew_model
- __init__.py required, or the package’s nodes do not load
- new_model_model_loader.py
- new_model_denoise.py
Directorytext_encoder
- new_model_text_encoder.py
Directoryvae
- new_model_latents_to_image.py
- fields.py
- app/services/model_records/model_records_base.py only with a variant enum
Directorybackend
Directoryarchitectures/defs
- new_model.py the architecture declaration
Directorymodel_manager
- taxonomy.py
Directoryconfigs
- main.py
- factory.py
Directoryload/model_loaders
- new_model.py
Directorystarter_models
- new_model.py
- __init__.py
- stable_diffusion/diffusion/conditioning_data.py
- invocation_api/__init__.py
Directoryfrontend/webv2/src/features
Directorygeneration/core
- contracts.ts
- supportedBases.ts
- graph.ts
- baseGenerationPolicies.ts
- __fixtures__/architectureCapabilities.json regenerated
- __snapshots__/generateGraphNodeTypes.json regenerated
Directorymodels/core
- baseIdentity.ts
- types.ts
- workflow/core/modelRequirements.ts
Directorytests
- backend/architectures the tables in section 13
- app/invocations/test_frontend_graph_node_types.py
For img2img, inpaint and outpaint, add:
Directoryinvokeai
Directoryapp/invocations/vae
- new_model_image_to_latents.py
Directoryfrontend/webv2/src/features/generation/core/canvas
- compileCanvasGraph.ts
CANVAS_I2L_NODE_TYPES - compileCanvasGraph.test.ts
BASE_CASES
- compileCanvasGraph.ts
Reference: existing implementations
Section titled “Reference: existing implementations”Values are taken from the declarations under invokeai/backend/architectures/defs/.
| FLUX.1 | FLUX.2 | Z-Image | Qwen-Image | Krea-2 | |
|---|---|---|---|---|---|
| Latent space | FLUX_16 (16 ch, 8×) | FLUX2_32 (32 ch, 8×) | FLUX_16 (16 ch, 8×) | WAN21_16 (16 ch, 8×) | WAN21_16 (16 ch, 8×) |
| Text encoder | CLIP + T5 | Qwen3 (Klein), Mistral (Dev) | Qwen3 | Qwen2.5-VL 7B | Qwen3-VL 4B |
| VAE accepted | FLUX.1 | FLUX.2 | FLUX.1 | Qwen-Image, Anima | Qwen-Image, Anima |
| Guidance slider | Guidance (distilled) | Guidance (max 20) | CFG (Turbo: 1.0) | CFG 4.0 | CFG (Turbo: 1.0) |
| Negative prompt | never | never | CFG-gated | CFG-gated | CFG-gated |
dimension_grid | 16 | 16 | 16 | 16 | 16 |
| Scheduler set | flow | flow | flow (Base: flow-no-lcm) | standard | flow |
| Reference images | 5 | 5 | — | 5 (edit variant) | — |
| Code | backend/flux/ | backend/flux2/ | in the nodes | in the nodes | backend/krea2/ |