
This ComfyUI tutorial workflow demonstrates end-to-end speech-to-text using the FishAudioSpeechToText node. Audio is loaded with LoadAudio (a sample file, fish_audio_example.mp3, is pre-connected so you can run it immediately), then sent to the Fish Audio ASR service for transcription. The node automatically detects the spoken language across 80+ languages and returns a clean transcript string along with an optional JSON payload of timestamped segments. The transcript is saved to disk via SaveText, while a MarkdownNote provides on-canvas tips.
Technically, LoadAudio decodes your file into an audio tensor and forwards it to FishAudioSpeechToText. When precise_timestamps is enabled on that node, it also returns segments_json containing per-segment text with start/end times (and, when available, finer word-level timing). SaveText writes the main transcript string to your ComfyUI output folder for easy reuse in subtitles, captions, or indexing pipelines. This setup is lightweight, API-driven, and practical for rapid transcription, subtitle prep, or generating text prompts from recorded audio.
API
코드에서 이 워크플로우 사용하기
모든 Comfy 워크플로우는 JSON 그래프입니다. 아래 페이로드는 이 워크플로우 그 자체로, ComfyUI에서 실행되는 것과 동일합니다 — URL에서 가져오거나, 버전 관리에 보관하거나, ComfyUI에서 불러와 노드별로 실행할 수 있습니다.
Comfy SDK로 TypeScript 또는 Python에서 실행하세요. 동일한 코드는 Comfy Cloud 또는 직접 호스팅하는 ComfyUI 모두에 적용됩니다 — 기본 URL만 다릅니다.
// Install (beta)
npm i @comfyorg/sdk
// Run this workflow (TypeScript)
import { Comfy } from "@comfyorg/sdk";
const client = new Comfy({ apiKey: "comfyui-..." });
const wf = await client.workflows.fromFile("workflow_api.json");
const job = await client.run(wf);
await job.getOutputs("<output-node-id>")[0].toFile("output.png");SDK는 API 형식의 워크플로우를 사용합니다: 이 워크플로우를 ComfyUI에서 열고 파일 → 워크플로우 내보내기(API)를 사용하세요.













