
This workflow turns text prompts into music using the ACE-Step 1.5 XL Base (4B) model inside ComfyUI. It pairs UNETLoader with the acestep_v1.5_xl_base_bf16.safetensors diffusion model and VAELoader with the ace_1.5_vae.safetensors decoder. Your prompt is encoded by TextEncodeAceStepAudio1.5, an EmptyAceStep1.5LatentAudio node creates a blank latent audio clip at your chosen duration, and KSampler performs the denoising pass to synthesize music. VAEDecodeAudio reconstructs the waveform from latents, and SaveAudioMP3 writes the final track to disk.
Under the hood, ConditioningZeroOut provides a clean fallback when the negative prompt is empty, while ModelSamplingAuraFlow configures the model’s sampler to the correct flow/schedule so KSampler can produce stable, on-style results. PrimitiveNode and PrimitiveInt nodes supply simple controls for duration, steps, guidance (CFG), and seed. The workflow is organized into clear groups (Model, Duration, Prompt) so you can quickly load the right weights, set the clip length, write a prompt, and iterate rapidly by adjusting steps, CFG, and seed.
FAQ
Frequently Asked Questions
Seedance 2.0: Reference to Video


Nano Banana 2 Lite: Text to Image

Seedream 5.0 Pro: Image Edit

Krea-2: Text to Image
4K Seedance 2.0 - Reference to Video


Z-Image-Turbo Text to Image
Grok: Image Edit
Nano Banana 2 Lite: Image Edit
Grok: Video generation

Seedance2.0 4K: Reference to Video

Grok Imagine Image Quality: Generation
LTX 2.3 - Lipdub LoRA + Voice Clone
Seedance 2.0 Mini: Reference to Video
SCAIL-2: Character Replacement

Ideogram v4: Text to Image
Googly Eyes

Seedance 2.0 - Viral Videos Character Swap
Gemini Omni Flash: Image to Video

Nano Banana 2: Image Edit


cinematic_annotate_video
HappyHorse 1.1: Text to Video
HappyHorse 1.1: Image to Video
Beeble SwitchX: Video Edit

3x3 Contact Sheet
Restore Archival Footage - LTX 2.3 Dearchive LoRA













