Skip to content

Media Metadata

Every image and video InvokeAI produces carries its generation record inside the file, so a download dropped into any Invoke instance is recallable and remixable without access to the database that produced it. This page is the reference for that contract: where the record lives in each container, the record’s fields, and how the format evolves.

The same three strings are written to every media file. Each value is compact UTF-8 JSON.

KeyContentRequired
invokeai_metadataThe generation record (schema below)when the graph produced one
invokeai_workflowThe WorkflowWithoutID the session ran, verbatimonly when a workflow was run
invokeai_graphThe executed Graph, verbatimwhen available
ContainerWhere the keys liveWriterReader
PNGtEXt chunks, one per keyDiskImageFileStorage.save via Pillow’s PngInfoextract_metadata_from_image reads Image.info
MP4QuickTime keyed metadata: moov/udta/meta with an mdta keys table and ilst UTF-8 data atomsDiskVideoFileStorage.save remuxes through the bundled ffmpeg (-movflags use_metadata_tags, stream copy, +faststart)invokeai.app.util.mp4_metadata.read_mp4_tags, a pure-Python box walk

Both readers feed extract_metadata in invokeai/app/api/extract_metadata.py, which owns the validation rules: the metadata must parse to a JSON object, the workflow must validate as a WorkflowWithoutID, and the graph must be an object with a nodes object and an edges list. Anything else degrades to “absent”. On upload, a client-supplied metadata form field wins over the embedded copy; workflow and graph are never client-overridable.

The MP4 layout is the one ffprobe, exiftool and mediainfo display. Like PNG chunks it survives a container-aware copy — for ffmpeg that is a remux with -movflags use_metadata_tags; a plain -c copy or a re-encode drops it.

A flat JSON object, the output of the core_metadata node. Video is a profile of the same record as images, discriminated by generation_mode, not a second format. Every key is optional; readers ignore keys they do not know and never reject a record for having them.

Common shapes:

ModelIdentifier = { key, hash, name, base, type, submodel_type? }
ImageRef = { image_name }
VideoRef = { video_name }
LoRA = { model: ModelIdentifier, weight }
H3Reference = { kind: "image" | "video", image_name?, video_name?, detail?, conditioning?, start_frame?, end_frame? }
KeyTypeMeaning
metadata_versionsemver stringThe record schema version, currently 1.0.0. Absent on records written before versioning; they are read with the same rules plus the aliases listed below.
app_versionstringThe InvokeAI release that wrote the record. Informational.

The version is over the record, not the node that writes it. A minor bump adds keys or widens a value set: any 1.x reader parses any other 1.x. A major bump renames, removes or retypes a key. Readers are tolerant by construction: they read the keys they know, ignore the rest, and never reject a record for its version, so a newer record still recalls what it can. The current value is CORE_METADATA_VERSION in invokeai/app/invocations/metadata.py.

KeyType
generation_modestring. Video values: wan_t2v, wan_i2v, wan_interpolate, wan_extend_video, minimax_h3_t2v, minimax_h3_i2v, minimax_h3_lf2v, minimax_h3_flf2v, minimax_h3_extend_video, minimax_h3_ref2v. A record is video metadata iff this is one of them.
positive_prompt, negative_promptstring; negative_prompt is absent for MiniMax H3, which has no CFG
seed, stepsinteger
cfg_scalenumber (Wan only)
width, heightinteger
modelModelIdentifier: the main (transformer) model
vaeModelIdentifier: a standalone VAE override
lorasLoRA[]
KeyTypeNotes
num_framesinteger
fpsintegerThe delivered frame rate. MiniMax H3 records its fixed 24.
first_frame_image, last_frame_imageImageRefIn extend mode the first frame is the one extracted from the source clip, so recall ignores it when source_video is set.
source_videoVideoRefThe clip an extension continues
source_video_start_frame, source_video_end_frameintegerInclusive trim bounds of the source clip
wan_guidance_scale_low_noisenumberThe low-noise expert’s CFG when it differed from cfg_scale
wan_t5_encoder_modelModelIdentifierStandalone UMT5-XXL used with a single-file Wan main
wan_transformer_low_noiseModelIdentifierStandalone low-noise expert used with a single-file Wan main
wan_component_sourceModelIdentifierWan Diffusers install that supplied VAE and encoder to a single-file main
minimax_h3_transformer_modelModelIdentifierLegacy: single-file transformer recorded as an override of a Diffusers model
minimax_h3_text_encoder_modelModelIdentifierStandalone Qwen3-VL text encoder
minimax_h3_component_sourceModelIdentifierH3 Diffusers install that supplied VAEs and tokenizer to a single-file transformer
minimax_h3_hybrid_base_modelModelIdentifierThe FL2VA transformer overlaid on a Ref2VA main
minimax_h3_hybrid_start_blockintegerFirst block taken from the hybrid base
minimax_h3_referencesH3Reference[]Ref2VA references in conditioning order
media_originstringSet by the upload route, not the graph: audio_upload marks an audio file wrapped into a video

Earlier writers spelled three Wan keys differently. Readers accept both spellings; the current graph builders write only the canonical ones.

CanonicalAlso read as
wan_guidance_scale_low_noiseguidance_scale_low_noise
wan_t5_encoder_modelwan_t5_encoder
wan_transformer_low_noisetransformer_low_noise
  • Models in a video record resolve against the local catalog by key, then by hash, then by name + base + type. Keys are minted per install, so a record that travelled with a file names keys the receiving catalog has never seen; the BLAKE3 hash is the content identity and resolves the same weights wherever they are installed. The name step is the last resort and the only one that can pick a different file of the same model. Image recall still resolves by key only.
  • Media references (image_name, video_name) are names in the originating gallery, exactly as an image record’s init_image is. Recall restores everything else and clears media it cannot find rather than substituting.
  • Workflow and graph load into the editor on any install, subject to the same model and media resolution.

The Video Recall API takes its parameters under these same key names from an external process, with models given by name.

  1. Add a typed Optional field to CoreMetadataInvocation under its canonical name, stamp it from the graph builder, and read it in the recall mappers.
  2. Bump CORE_METADATA_VERSION’s minor component. Bump the major only for a rename, removal or type change, and add the old spelling to the readers’ alias table.
  3. Update the tables on this page.
This site was designed and developed by Aether Fox Studio.