MiniMax Music 3: Text to Music

This ComfyUI tutorial workflow turns structured text into full-length music using MiniMax Music 3. At its core is the MiniMax Music 3 generator node (UUID: ac99f841-a3de-4329-9564-953b81cf9e16), which accepts two key inputs: a Caption describing style, mood, tempo, instrumentation, and arrangement, and optional Lyrics formatted with section tags like [intro], [verse], [chorus], [bridge], and [outro]. The node renders a complete track—up to five minutes—with stable song structure and high-fidelity audio. A MarkdownNote block in the graph documents the prompt schema so you can quickly copy, tweak, and version your captions and lyrics inside the UI.

Under the hood, you can choose between two model weights: minimax_music3_dit_fp16.safetensors (recommended when VRAM allows) and minimax_music3_dit_int8_convrot.safetensors (optimized for lower VRAM). The generator node produces an audio tensor that is written to disk via SaveAudioAdvanced, where you control filename, format, and sample rate. The result is a finished, arranged song suitable for prototypes, background scores, and content production without leaving ComfyUI.

Tüm iş akışlarını gör
Character
Cinematic
Image to Video
Lip Sync
Multiple Angles
Portrait
Style Reference
Style Transfer
Text to Video
Video Generation
Video
Showing 30 of 30 templates