Skip to content

Video Workflows

InvokeAI ships twenty-six ready-made video workflows — ten for Wan 2.2, six for MiniMax H3 and ten for LTX-2 — so you can start generating video without wiring up nodes yourself. Each one is a complete recipe — open it from the Workflows library, pick your models, type a prompt (and/or drop in an image), and press Invoke.

This page helps you choose the right workflow and run it. For how the models work under the hood, VRAM details, and troubleshooting, see the Video Generation guide.


1. Install the video models. Open the Model Manager and install a starter bundle: Wan 2.2 Text-to-Video, Wan 2.2 Image-to-Video, MiniMax H3, and/or LTX-2.5. The Video Generation guide explains the bundles and what fits your GPU.

2. You’ll select models when you open a workflow. The shipped workflows come with their model slots left empty on purpose — every InvokeAI install stores models a little differently, so a workflow can’t point at yours automatically. When you open one, look at its Notes panel (the note icon in the workflow editor): it lists exactly which model to pick in each slot. Drop those in once and you’re set.


The Wan workflows come in three families — Text to Video, Image to Video, and Extend Video — each with a high-quality option, a “concept LoRA” variant, and a lighter low-VRAM option — plus an Interpolate workflow that bridges two images and an Extend Video to Image variant that steers a continuation toward a target frame. The six MiniMax H3 workflows cover the same conditioning modes — plus reference-conditioned generation — and add sound. The ten LTX-2 workflows cover every one of those modes with sound, add audio-to-video and video-to-audio, and include a two-stage high-resolution text-to-video and a LoRA-collection variant.

I want to…UseNotes
Make a clip from a text descriptionText to Video - Wan 2.2 LightningBest quality, fast. Start here.
Make a clip with a soundtrackText to Video - LTX-2 or Text to Video - MiniMax H3Both generate synchronized stereo audio with every video.
Make a high-resolution clipText to Video - LTX-2 Two-StageGenerates at half size, upscales, then refines at 1024p or 1536p. Slow.
Animate an existing imageImage to Video - Wan 2.2 Lightning, First Frame to Video - MiniMax H3 or First Frame to Video - LTX-2Your image becomes the first frame.
Make a video that ends on an imageLast Frame to Video - MiniMax H3 or Last Frame to Video - LTX-2The video builds toward your image.
Make a video between two imagesInterpolate 2 Images to Video - Wan 2.2 Lightning, First and Last Frame to Video - MiniMax H3 or First and Last Frame to Video - LTX-2Provide a start and end image; the model interpolates between them.
Make a video longer than one generationExtend Video - Wan 2.2 Lightning, Extend Video - MiniMax H3 or Extend Video - LTX-2Continues an existing video and stitches the pieces together. H3 and LTX-2 continue the soundtrack too.
Make a clip in the style/likeness of reference mediaReference to Video - MiniMax H3Conditions on up to 3 reference videos and 9 images (Ref2VA transformer required); reference order changes the result.
Extend a video toward a specific end frameExtend Video to Image - Wan 2.2 Lightning or Extend Video to Image - LTX-2Continues a video and interpolates the new part to a target image.
Make a picture for a soundtrack — a song, a voice recordingAudio to Video - LTX-2The audio is kept as-is; the video is generated to match it, including lip movement.
Add a soundtrack to a silent or existing clipVideo to Audio - LTX-2Your footage is kept as-is; a new stereo soundtrack is generated for it.
Add my own style/subject LoRAsthe w/ Concept LoRAs variant of a Wan workflow, or Text to Video - LTX-2 w/ LoRA CollectionSame as the base workflow, plus your LoRAs.
Run on a smaller GPU (≈12–16 GB)a TI2V-5B (Low Quality) workflowSmaller model, lower memory, lower quality.

If you’re not sure, start with Text to Video - Wan 2.2 Lightning to get a feel for it, then move to Image-to-Video and Extend — or straight to Text to Video - MiniMax H3 if you want audio and have the hardware for it.


Text to Video - Wan 2.2 Lightning

Type a prompt, get a ~5-second clip. Uses the high-quality A14B model with the Lightning speed-up (just 4 steps). The recommended starting point.

Text to Video - …w/ Concept LoRAs

The same workflow, plus extra slots for your own concept LoRAs (a style or character you’ve trained or downloaded). Every LoRA slot must be filled — if you don’t have concept LoRAs to add, use the plain version instead.

Text to Video - Wan 2.2 TI2V-5B (Low Quality)

Uses the smaller TI2V-5B model. Lower quality and slower per result, but fits comfortably on 12–16 GB GPUs. Good for quick drafts or smaller cards.

Image to Video - Wan 2.2 Lightning

Drop in a starting image and a prompt; the model animates outward from your image as the first frame. Uses the A14B model + Lightning. The recommended image-to-video starting point.

Image to Video - …w/ Concept LoRAs

Same as above, with slots for your own concept LoRAs.

Image to Video - Wan 2.2 TI2V-5B (Low Quality)

Image-to-video on the smaller TI2V-5B model. Lower quality, lower VRAM — the low-memory option for animating an image.

Interpolate 2 Images to Video - Wan 2.2 Lightning

Provide a start image and an end image; the model generates a clip that begins on the first and animates smoothly to the second (first-last-frame interpolation). Great for morphing between two stills or bridging two shots. Uses I2V-A14B + Lightning (4-step).

The Wan models are trained on short (~5 second) clips, so the way to make something longer is to continue a video and stitch the pieces together. These workflows do that for you.

Extend Video - Wan 2.2 Lightning

Provide an existing video and use the sliders to choose where to continue from. The workflow generates a new clip starting from that point and joins it onto the original with a smooth transition. Run it repeatedly to keep growing the video.

Extend Video - …w/ Concept LoRAs

Same as above, with slots for your own concept LoRAs so the continuation keeps your style or subject.

Extend Video to Image - Wan 2.2 Lightning

Continue a video toward a target image. Provide a starting video and a destination image; the new segment interpolates from the video’s last frame to your image (FLF2V), then joins onto the original with a smooth transition. Combines the Extend and Interpolate ideas — useful for steering where a continuation ends up.


One model family, six workflows, every conditioning mode — and every output is an MP4 with a jointly generated stereo soundtrack. All default to a Turbo step-distillation LoRA (6 steps for FL2VA with the original Turbo LoRA, 8 with either LightX2V release); H3 has no CFG or negative prompt to tune, and always renders at 24 fps. Reference to Video needs the Ref2VA transformer selected in the model loader; the other five use FL2VA.

Text to Video - MiniMax H3

Type a prompt, get a clip with sound. H3’s vision-language text encoder follows detailed scene descriptions well — describe the audio you expect, too.

First Frame to Video - MiniMax H3

Drop in a starting image; the video animates outward from it, soundtrack included.

Last Frame to Video - MiniMax H3

The reverse — provide an ending image and the video builds toward it. Unique to H3; no Wan workflow can do this.

First and Last Frame to Video - MiniMax H3

Provide both; the model interpolates a clip (with audio) that starts on the first image and lands on the second.

Extend Video - MiniMax H3

Continue an existing video with a new H3 segment, joined on with a smooth transition. The continuation renders at H3’s fixed 24 fps.


Every LTX-2 workflow outputs an MP4 with a jointly generated stereo soundtrack (or, for audio-to-video, your own). Each works with either transformer: the Dev build runs the guided 30-step recipe, and the Distilled build a fixed 8 steps that ignore the Steps and CFG fields. The canvas comes from your image, source video or clip where there is one — except in Audio to Video, which generates the picture — and otherwise the Width and Height fields set only the aspect ratio. Target Resolution sets the size.

Text to Video - LTX-2

Type a prompt, get a ~5-second clip with sound. Describe the audio you expect, not just the picture.

Text to Video - LTX-2 Two-Stage

The high-resolution version: generates at half the canvas, doubles it with LTX-2’s latent upscaler, then refines it at full size. Defaults to 1024p (1792×1024 at 16:9); 1536p also works, but 512p–768p fail here — use the plain Text to Video workflow for those. Much slower — minutes on Distilled, far longer on Dev.

Text to Video - LTX-2 w/ LoRA Collection

Text-to-video plus a list of your own LTX-2 LoRAs, each with its own weight. Leave the list empty to run the base model.

First Frame to Video - LTX-2

Your image becomes the opening frame; the video and its soundtrack continue from it.

Last Frame to Video - LTX-2

Provide an ending image and the video builds toward it.

First and Last Frame to Video - LTX-2

Provide both; the clip starts on the first image and lands on the second. Give the two images the same aspect ratio.

Extend Video - LTX-2

Continue an existing clip. The continuation starts from the source’s last few frames and its closing sound, and joins onto it with a cross-fade over the same span, so the picture and the soundtrack both carry across the join. Context Frames sets how much of the source is carried in.

Extend Video to Image - LTX-2

The same, with a Last Frame image: the continuation ends on your image. Give it the source clip’s aspect ratio.

Audio to Video - LTX-2

Drop in a clip — or an audio file uploaded to the gallery — and LTX-2 generates a picture for its soundtrack, down to lip movement. The soundtrack decides the length, and the output carries the original recording. Say in the prompt who or what is making the sound.

Video to Audio - LTX-2

Drop in a clip and LTX-2 generates a soundtrack for it. The clip decides the length and frame rate, and the output is your original footage with the new audio.


  1. Open the Workflows tab and load one of the workflows from the library.

  2. Read the Notes (note icon) and select the models it lists — the main model, VAE, and text encoder, plus the Lightning LoRA pair on the A14B workflows, the H3 loader’s three selections plus the Turbo LoRA on the MiniMax workflows, or the LTX-2 loader’s three selections on the LTX-2 workflows.

  3. Set your inputs:

    • Text to Video — type a positive prompt describing the scene and motion.
    • Image to Video / First (or Last) Frame to Video — drop a starting (or ending) image into the image field, and optionally add a prompt.
    • Extend Video — provide a starting video and trim it with the Start Frame and End Frame sliders (they appear once the video is loaded); the continuation picks up from the end of the trimmed range.
  4. (Optional) adjust resolution. The defaults are chosen to balance speed and memory. If you hit an out-of-memory error, lower the resolution preset — see the VRAM tips below.

  5. Press Invoke. The finished MP4 appears in the gallery and plays inline in the viewer — with its audio track, on H3.


A rough guide — exact numbers depend on resolution and your system. See Video Generation → OOM errors for the full picture.

Your GPUBest choice
24 GB and upAny Wan workflow at 720p; the MiniMax H3 workflows (with enable_partial_loading and generous system RAM).
16–24 GBA14B Lightning workflows at 480p; TI2V-5B at up to 720p.
12–16 GBThe TI2V-5B (Low Quality) workflows.

If you get an out-of-memory error, lower the resolution first (it helps far more than reducing the frame count), and switch to a TI2V-5B workflow if you’re still over budget. On H3, switch the resolution to 768 lowres for drafts.


These workflows are starting points. Once you’re comfortable, the Video Generation guide covers the Video panel (the no-graph way to run all of this), the model variants, how image conditioning works, extending clips into longer videos, recommended parameters, and troubleshooting.

Read the Video Generation guide
This site was designed and developed by Aether Fox Studio.