ElevenLabs: Text to Speech

This ComfyUI workflow leverages the power of ElevenLabs to transform written text into ultra-realistic speech. Utilizing nodes like ElevenLabsTextToSpeech and ElevenLabsVoiceSelector, users can either select from a range of preset voices or upload a sample to clone a specific voice for synthesis. The workflow is versatile, allowing for both text-to-speech conversion and voice cloning, making it a powerful tool for content creators, educators, and developers who need high-quality audio outputs. By integrating nodes such as LoadAudio and SaveAudioMP3, users can manage audio inputs and outputs efficiently, ensuring a seamless experience from text input to audio file generation.

API

Use this workflow from code

Every Comfy workflow is a JSON graph. The payload below is this workflow, exactly as ComfyUI runs it — fetch it from the URL, keep it in version control, or load it in ComfyUI and run it node by node.

api_elevenlabs_text_to_speech.json
Fetching workflow JSON…

Run it from TypeScript or Python with the Comfy SDK. The same code targets Comfy Cloud or a ComfyUI you host yourself — only the base URL changes.

// Install (beta)
npm i @comfyorg/sdk

// Run "ElevenLabs: Text to Speech" (TypeScript)
import { Comfy } from "@comfyorg/sdk";

const client = new Comfy({ apiKey: "comfyui-..." });

// This workflow, exported in API format (see note below)
const wf = await client.workflows.fromFile("api_elevenlabs_text_to_speech_api.json");

const job = await client.run(wf);
await job.getOutputs("215")[0].toFile("output.png"); // SaveAudioMP3

The SDK takes a workflow in API format: open this workflow in ComfyUI and use File → Export Workflow (API).

Comfy Cloud API access requires a plan with an API key.

FAQ

Frequently Asked Questions

View all workflows
Character
Cinematic
Image to Video
Lip Sync
Multiple Angles
Portrait
Style Reference
Style Transfer
Text to Video
Video Generation
Video
Showing 30 of 30 templates