LTX 2.3_ID_lora_CelebVHQ

This ComfyUI tutorial workflow turns a single image into a short, voice-driven video with strong visual identity and perceived lip alignment. It combines Lightricks LTX 2.3 for video generation with an ID-LoRA (CelebVHQ) to keep the subject’s face and styling consistent across frames, while Qwen 3 TTS produces a matching voice track. The pipeline is organized into clear groups: Load Image, Create Prompts, Prompt Extraction, Create Voice, and Create Video.

Under the hood, GeminiNode analyzes your uploaded image (via LoadImage) and drafts two complementary texts: a speech line with expressive cues and a scene/appearance description. RegexExtract then parses these into an ID-LoRA-ready format with three tagged sections: [VISUAL], [SPEECH], and [SOUNDS]. FB_Qwen3TTSCustomVoice uses the [SPEECH] and style hints from [SOUNDS] to synthesize a voice track, saved with SaveAudioMP3. The LTX 2.3 video node (the custom node in this graph) consumes the structured prompt plus your reference image to render a talking video, with the CelebVHQ ID-LoRA improving identity retention in dynamic, in-the-wild scenarios. PreviewAny lets you iterate quickly, and SaveVideo writes the final clip. Local users can optionally load the LTX spatial upscaler (ltx-2.3-spatial-upscaler-x2-1.1.safetensors) for cleaner details at higher resolution.

FAQ

Frequently Asked Questions

View all workflows
Character
Cinematic
Image to Video
Lip Sync
Multiple Angles
Portrait
Style Reference
Style Transfer
Text to Video
Video Generation
Video
Showing 30 of 579 templates