This ComfyUI workflow turns a single image and a voice recording into a lip-synced talking video using the LTX-2.3 model. You load a portrait with LoadImage and provide speech via LoadAudio or capture it live with RecordAudio. Both streams feed the LTX-2.3 generator node (98ee9e5b-467b-40aa-a534-36033f27d0b4), which synthesizes a sequence of frames where the subject speaks in time with the audio. The resulting frames are encoded to an MP4 using SaveVideo.
Under the hood, LTX-2.3 conditions on the visual identity from your reference image and the temporal features of the provided audio to drive mouth shapes and subtle facial motions over time. The node typically exposes settings like output resolution and FPS, and many builds also include a seed for reproducibility and an optional text prompt to guide motion or style. The MarkdownNote in the graph documents quick tips and links to the official Lightricks model repositories so you can download the required weights.
FAQ
Frequently Asked Questions
Seedance 2.0: Reference to Video


Nano Banana 2 Lite: Text to Image

Seedream 5.0 Pro: Image Edit

Krea-2: Text to Image
4K Seedance 2.0 - Reference to Video


Z-Image-Turbo Text to Image
Grok: Image Edit
Nano Banana 2 Lite: Image Edit
Grok: Video generation

Seedance2.0 4K: Reference to Video

Grok Imagine Image Quality: Generation
LTX 2.3 - Lipdub LoRA + Voice Clone
Seedance 2.0 Mini: Reference to Video
SCAIL-2: Character Replacement

Ideogram v4: Text to Image
Googly Eyes

Seedance 2.0 - Viral Videos Character Swap
Gemini Omni Flash: Image to Video

Nano Banana 2: Image Edit


cinematic_annotate_video
HappyHorse 1.1: Text to Video
HappyHorse 1.1: Image to Video
Beeble SwitchX: Video Edit

3x3 Contact Sheet
Restore Archival Footage - LTX 2.3 Dearchive LoRA













