Fish Audio: Speech to Text

This ComfyUI tutorial workflow demonstrates end-to-end speech-to-text using the FishAudioSpeechToText node. Audio is loaded with LoadAudio (a sample file, fish_audio_example.mp3, is pre-connected so you can run it immediately), then sent to the Fish Audio ASR service for transcription. The node automatically detects the spoken language across 80+ languages and returns a clean transcript string along with an optional JSON payload of timestamped segments. The transcript is saved to disk via SaveText, while a MarkdownNote provides on-canvas tips.

Technically, LoadAudio decodes your file into an audio tensor and forwards it to FishAudioSpeechToText. When precise_timestamps is enabled on that node, it also returns segments_json containing per-segment text with start/end times (and, when available, finer word-level timing). SaveText writes the main transcript string to your ComfyUI output folder for easy reuse in subtitles, captions, or indexing pipelines. This setup is lightweight, API-driven, and practical for rapid transcription, subtitle prep, or generating text prompts from recorded audio.

API

通过代码使用此工作流

每个 Comfy 工作流都是一个 JSON 图。下面的内容就是此工作流的完整数据,ComfyUI 运行时也是如此——你可以从该 URL 获取、放入版本控制,或在 ComfyUI 中加载并逐节点运行。

97f877b082ae.json
正在获取工作流 JSON…

使用 Comfy SDK 从 TypeScript 或 Python 运行。相同代码可用于 Comfy Cloud 或你自托管的 ComfyUI——只需更改基础 URL。

// Install (beta)
npm i @comfyorg/sdk

// Run this workflow (TypeScript)
import { Comfy } from "@comfyorg/sdk";

const client = new Comfy({ apiKey: "comfyui-..." });
const wf = await client.workflows.fromFile("workflow_api.json");
const job = await client.run(wf);
await job.getOutputs("<output-node-id>")[0].toFile("output.png");

SDK 需要 API 格式的工作流:在 ComfyUI 中打开此工作流,使用 文件 → 导出工作流(API)。

Comfy Cloud API 访问需要包含 API 密钥的套餐。

查看所有工作流
Showing 30 of 30 templates