This ComfyUI tutorial workflow turns a single text prompt into a finished video with native stereo audio using MiniMax H3, an omni‑modal model that synthesizes visuals, voice, sound effects, and music in one pass. There are no file inputs to manage—just describe what you want, and the graph returns a roughly 15‑second clip at 24 fps, up to 2K resolution, with stereo sound baked in.

Under the hood, a MiniMax H3 subgraph (node 79dd8a95-ce9d-4c14-b264-2162e8bec5ce) handles generation and decoding. It uses the MiniMax H3 video and audio VAEs (minimax_h3_video_vae_fp16.safetensors and minimax_h3_audio_vae_fp32.safetensors) to reconstruct frames and stereo audio from the model’s latents. ResolutionSelector maps convenient presets to valid dimensions in multiples of 32 (e.g., 960×544, 1152×640, 1216×672), and SaveVideo multiplexes the frames and stereo audio stream into a single MP4. A MarkdownNote in the canvas provides quick tips and prompt examples. The result is a compact, prompt‑only pipeline ideal for rapid concept visualization, social posts, and audio‑synced storytelling.

API

Use this workflow from code

Every Comfy workflow is a JSON graph. The payload below is this workflow, exactly as ComfyUI runs it — fetch it from the URL, keep it in version control, or load it in ComfyUI and run it node by node.

e8099b642c9f.json
Fetching workflow JSON…

Run it from TypeScript or Python with the Comfy SDK. The same code targets Comfy Cloud or a ComfyUI you host yourself — only the base URL changes.

// Install (beta)
npm i @comfyorg/sdk

// Run this workflow (TypeScript)
import { Comfy } from "@comfyorg/sdk";

const client = new Comfy({ apiKey: "comfyui-..." });
const wf = await client.workflows.fromFile("workflow_api.json");
const job = await client.run(wf);
await job.getOutputs("<output-node-id>")[0].toFile("output.png");

The SDK takes a workflow in API format: open this workflow in ComfyUI and use File → Export Workflow (API).

Comfy Cloud API access requires a plan with an API key.

FAQ

Frequently Asked Questions

View all workflows
Character
Cinematic
Image to Video
Lip Sync
Multiple Angles
Portrait
Style Reference
Style Transfer
Text to Video
Video Generation
Video
Showing 30 of 30 templates