
InfiniteTalk: Audio-Driven Full-Body Video Dubbing turns a single image and one or more audio tracks into a lip‑synced, full‑body video while preserving the original identity, background, and camera motion. The core of the pipeline is WanInfiniteTalkToVideo, which fuses features from the loaded Wan2.1 I2V model with an audio embedding to drive body and facial motion. Step1 - Load Models prepares the UNet and VAE (UNETLoader, VAELoader) for image‑to‑video synthesis, and loads text and audio backbones (CLIPLoader with CLIPTextEncode for camera/stage prompts, and AudioEncoderLoader with AudioEncoderEncode for voice features). You import the starting frame via LoadImage, then draw masks per character in the MaskEditor to localize motion to each speaker.
API
Use this workflow from code
Every Comfy workflow is a JSON graph. The payload below is this workflow, exactly as ComfyUI runs it — fetch it from the URL, keep it in version control, or load it in ComfyUI and run it node by node.
Run it from TypeScript or Python with the Comfy SDK. The same code targets Comfy Cloud or a ComfyUI you host yourself — only the base URL changes.
// Install (beta)
npm i @comfyorg/sdk
// Run "InfiniteTalk: Audio-Driven Full-Body Video Dubbing" (TypeScript)
import { Comfy } from "@comfyorg/sdk";
const client = new Comfy({ apiKey: "comfyui-..." });
// This workflow, exported in API format (see note below)
const wf = await client.workflows.fromFile("video_wan2_1_infinitetalk_api.json");
const asset = client.assets.fromFile("input.png");
wf.setInput("32", "image", asset); // LoadImage
const job = await client.run(wf);
await job.getOutputs("162")[0].toFile("output.png"); // SaveVideoThe SDK takes a workflow in API format: open this workflow in ComfyUI and use File → Export Workflow (API).
Comfy Cloud API access requires a plan with an API key.
FAQ













