- Ana Sayfa
- İş Akışları
- LTX 2.3_ID_lora_CelebVHQ

This ComfyUI tutorial workflow turns a single image into a short, voice-driven video with strong visual identity and perceived lip alignment. It combines Lightricks LTX 2.3 for video generation with an ID-LoRA (CelebVHQ) to keep the subject’s face and styling consistent across frames, while Qwen 3 TTS produces a matching voice track. The pipeline is organized into clear groups: Load Image, Create Prompts, Prompt Extraction, Create Voice, and Create Video.
Under the hood, GeminiNode analyzes your uploaded image (via LoadImage) and drafts two complementary texts: a speech line with expressive cues and a scene/appearance description. RegexExtract then parses these into an ID-LoRA-ready format with three tagged sections: [VISUAL], [SPEECH], and [SOUNDS]. FB_Qwen3TTSCustomVoice uses the [SPEECH] and style hints from [SOUNDS] to synthesize a voice track, saved with SaveAudioMP3. The LTX 2.3 video node (the custom node in this graph) consumes the structured prompt plus your reference image to render a talking video, with the CelebVHQ ID-LoRA improving identity retention in dynamic, in-the-wild scenarios. PreviewAny lets you iterate quickly, and SaveVideo writes the final clip. Local users can optionally load the LTX spatial upscaler (ltx-2.3-spatial-upscaler-x2-1.1.safetensors) for cleaner details at higher resolution.
FAQ
























