
This workflow turns a pair of images into a short, coherent video by leveraging Wan 2.1’s inpainting-based video diffusion. You provide a start frame and an end frame; the WanFunInpaintToVideo node then guides the model to synthesize the in‑between motion while preserving subject and scene details anchored in your images. Text prompts (via CLIPTextEncode) steer style and content, while negative prompts help suppress artifacts.
Under the hood, Step1 loads the core components: UNETLoader (Wan2.1 / Wan), VAELoader/VAEDecode for latent <-> pixel space conversion, and CLIPLoader for text conditioning. Your start/end frames are brought in with LoadImage and encoded using CLIPVisionLoader + CLIPVisionEncode to produce image conditioning for the UNet. KSampler drives the diffusion process using ModelSamplingSD3 for the correct sigma schedule, with CFGZeroStar and SkipLayerGuidanceDiT stabilizing guidance. The Attention Booster group applies UNetTemporalAttentionMultiply to strengthen temporal consistency across frames. Finally, CreateVideo assembles frames and SaveVideo writes the finished clip.
API
このワークフローをコードから利用
すべてのComfyワークフローはJSONグラフです。以下のペイロードは、このワークフローそのものであり、ComfyUIが実行する内容と同じです — このURLから取得し、バージョン管理に保存したり、ComfyUIで読み込んでノードごとに実行できます。
Comfy SDKを使ってTypeScriptまたはPythonから実行できます。同じコードでComfy Cloudや自身でホストしたComfyUIのどちらにも対応 — ベースURLだけが異なります。
// Install (beta)
npm i @comfyorg/sdk
// Run "Wan 2.1 インペイント" (TypeScript)
import { Comfy } from "@comfyorg/sdk";
const client = new Comfy({ apiKey: "comfyui-..." });
// This workflow, exported in API format (see note below)
const wf = await client.workflows.fromFile("wan2.1_fun_inp_api.json");
const asset = client.assets.fromFile("input.png");
wf.setInput("72", "image", asset); // LoadImage
wf.setInput("6", "text", "your prompt here"); // CLIPTextEncode
const job = await client.run(wf);
await job.getOutputs("80")[0].toFile("output.png"); // SaveVideoSDKはAPI形式のワークフローを受け取ります:このワークフローをComfyUIで開き、「ファイル → ワークフローをエクスポート(API)」を使用してください。
FAQ













