
This workflow turns a pair of images into a short, coherent video by leveraging Wan 2.1’s inpainting-based video diffusion. You provide a start frame and an end frame; the WanFunInpaintToVideo node then guides the model to synthesize the in‑between motion while preserving subject and scene details anchored in your images. Text prompts (via CLIPTextEncode) steer style and content, while negative prompts help suppress artifacts.
Under the hood, Step1 loads the core components: UNETLoader (Wan2.1 / Wan), VAELoader/VAEDecode for latent <-> pixel space conversion, and CLIPLoader for text conditioning. Your start/end frames are brought in with LoadImage and encoded using CLIPVisionLoader + CLIPVisionEncode to produce image conditioning for the UNet. KSampler drives the diffusion process using ModelSamplingSD3 for the correct sigma schedule, with CFGZeroStar and SkipLayerGuidanceDiT stabilizing guidance. The Attention Booster group applies UNetTemporalAttentionMultiply to strengthen temporal consistency across frames. Finally, CreateVideo assembles frames and SaveVideo writes the finished clip.
FAQ
Frequently Asked Questions
Seedance 2.0: Reference to Video


Nano Banana 2 Lite: Text to Image

Seedream 5.0 Pro: Image Edit

Krea-2: Text to Image
4K Seedance 2.0 - Reference to Video


Z-Image-Turbo Text to Image
Grok: Image Edit
Nano Banana 2 Lite: Image Edit
Grok: Video generation

Seedance2.0 4K: Reference to Video

Grok Imagine Image Quality: Generation
LTX 2.3 - Lipdub LoRA + Voice Clone
Seedance 2.0 Mini: Reference to Video
SCAIL-2: Character Replacement

Ideogram v4: Text to Image
Googly Eyes

Seedance 2.0 - Viral Videos Character Swap
Gemini Omni Flash: Image to Video

Nano Banana 2: Image Edit


cinematic_annotate_video
HappyHorse 1.1: Text to Video
HappyHorse 1.1: Image to Video
Beeble SwitchX: Video Edit

3x3 Contact Sheet
Restore Archival Footage - LTX 2.3 Dearchive LoRA













