This ComfyUI tutorial workflow shows how to generate a complete, synchronized video-with-audio clip directly from text using the MiniMaxH3ImageToVideo node (node id 4c314f31-ecda-4b08-ae98-faaba1bf613f) and the FastVideo FastH3 8-Step V2 checkpoint. The model is DMD2‑distilled to run in just eight denoising steps, which dramatically reduces generation time while maintaining coherent motion and audio alignment. No image or audio files are required—only a text prompt. A ResolutionSelector node feeds clean width/height (and commonly FPS/duration presets) into the generator so you can switch aspect ratios without manually reconfiguring multiple inputs. The SaveVideo node then muxes the decoded video frames and the generated audio into a single MP4 output.
Under the hood, the generator produces temporally consistent video latents and synchronized audio latents from your text prompt. These are decoded by the MiniMax H3 VAEs—minimax_h3_video_vae_fp16.safetensors for frames and minimax_h3_audio_vae_fp32.safetensors for sound—before being passed to SaveVideo. Because the checkpoint is distilled for eight steps, you get fast turnarounds that are ideal for rapid prototyping, ideation, and storyboarding. A MarkdownNote in the canvas explains the task setup and links to the subgraph guide in case you want to inspect or customize the encapsulated nodes.
API
從程式碼使用此工作流
每個 Comfy 工作流都是一個 JSON 圖。下方的 payload 就是此工作流,與 ComfyUI 執行時完全相同 — 你可以從此 URL 取得、放入版本控制,或在 ComfyUI 載入並逐節點執行。
使用 Comfy SDK 以 TypeScript 或 Python 執行。相同程式碼可用於 Comfy Cloud 或你自架的 ComfyUI — 只需更改 base URL。
// Install (beta)
npm i @comfyorg/sdk
// Run this workflow (TypeScript)
import { Comfy } from "@comfyorg/sdk";
const client = new Comfy({ apiKey: "comfyui-..." });
const wf = await client.workflows.fromFile("workflow_api.json");
const job = await client.run(wf);
await job.getOutputs("<output-node-id>")[0].toFile("output.png");SDK 需使用 API 格式的工作流:請在 ComfyUI 開啟此工作流,並使用 檔案 → 匯出工作流(API)。
















