@elevenlabs/text-to-speech

Collections of Skills to assist building with ElevenLabs

View in AI SkillSafe app
Scanned · no findings
322 downloads
0 stars
0 demos
SKILL.md
nametext-to-speech
descriptionConvert text to speech using ElevenLabs voice AI. Use when generating audio from text, creating voiceovers, building voice apps, or synthesizing speech in 90+ languages.
licenseMIT
compatibilityRequires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).
metadata{"openclaw": {"requires": {"env": ["ELEVENLABS_API_KEY"]}, "primaryEnv": "ELEVENLABS_API_KEY"}}

ElevenLabs Text-to-Speech

Generate natural speech from text - supports 90+ languages, multiple models for quality vs latency tradeoffs. Examples default to eleven_v4, the latest and highest-quality model.

Setup: See Installation Guide. For JavaScript, use @elevenlabs/* packages only.

Quick Start

Python

from elevenlabs import ElevenLabs

client = ElevenLabs()

audio = client.text_to_speech.convert(
    text="Hello, welcome to ElevenLabs!",
    voice_id="JBFqnCBsd6RMkjVDRZzb",  # George
    model_id="eleven_v4"
)

with open("output.mp3", "wb") as f:
    for chunk in audio:
        f.write(chunk)

JavaScript

import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import { createWriteStream } from "fs";
import { Readable } from "stream";

const client = new ElevenLabsClient();
const audio = await client.textToSpeech.convert("JBFqnCBsd6RMkjVDRZzb", {
  text: "Hello, welcome to ElevenLabs!",
  modelId: "eleven_v4",
});
// convert() returns a web ReadableStream — bridge it to a Node stream to write to disk
Readable.fromWeb(audio).pipe(createWriteStream("output.mp3"));

CLI

Use say to play text immediately with the default voice and eleven_v3 model:

elevenlabs say "Hello!"

Pipe text into say when another command produces the input:

echo "The build finished successfully." | elevenlabs say

Use the API command when you need to set request parameters directly:

elevenlabs text-to-speech convert --voice-id JBFqnCBsd6RMkjVDRZzb \
  --text "Hello!" --model-id eleven_v4 --output output.mp3

The CLI reads ELEVENLABS_API_KEY from the environment automatically.

Models

Model ID Languages Latency Best For
eleven_v4 90+ Standard Default. Highest quality, emotional range, audio tags, voice cloning accuracy
eleven_v4_turbo 90+ ~100ms Real-time v4 quality for agents and interactive apps (via Text to Dialogue WebSocket)
eleven_flash_v2_5 32 ~75ms Lowest latency, lowest cost, stream-input WebSocket
eleven_flash_v2 English ~75ms English-only, lowest latency
eleven_multilingual_v2 29 Standard Previous generation; most stable on very long-form generations
eleven_v3 70+ Standard Previous generation expressive model

eleven_turbo_v2_5 and eleven_turbo_v2 are superseded by the Flash models (same output, lower latency) — use eleven_flash_v2_5 / eleven_flash_v2 instead.

Eleven v4 notes: supports only stability and similarity_boost voice settings (no style or speed), does not support SSML, and is not available on the stream-input TTS WebSocket. Direct delivery with audio tags such as [whispers], [laughs], [sarcastic].

Voice IDs

Use pre-made voices or create custom voices in the dashboard.

Popular voices:

  • JBFqnCBsd6RMkjVDRZzb - George (male, narrative)
  • EXAVITQu4vr4xnSDxMaL - Sarah (female, soft)
  • onwK4e9ZLuTAKqWW03F9 - Daniel (male, authoritative)
  • XB0fDUnXU5powFXDhCwa - Charlotte (female, conversational)
voices = client.voices.get_all()
for voice in voices.voices:
    print(f"{voice.voice_id}: {voice.name}")

Voice Settings

Fine-tune how the voice sounds:

  • Stability: How consistent the voice stays. Lower values = more emotional range and variation, but can sound unstable. Higher = steady, predictable delivery.
  • Similarity boost: How closely to match the original voice sample. Higher values sound more like the original but may amplify audio artifacts.
  • Style: Exaggerates the voice's unique style characteristics (v2/v3 models; not supported on Eleven v4).
  • Speed: Speech speed multiplier (v2/v3 models; not supported on Eleven v4).
  • Speaker boost: Post-processing that enhances clarity and voice similarity.
from elevenlabs import VoiceSettings

audio = client.text_to_speech.convert(
    text="Customize my voice settings.",
    voice_id="JBFqnCBsd6RMkjVDRZzb",
    model_id="eleven_v4",
    voice_settings=VoiceSettings(
        stability=0.5,          # Lower = more expressive, higher = more consistent
        similarity_boost=0.75,  # Higher = closer to the reference voice
    )
)

Language Selection

Use language_code with models that support language enforcement to guide pronunciation and text normalization. Unsupported language codes are ignored, and language_code is not supported on eleven_multilingual_v2.

audio = client.text_to_speech.convert(
    text="Bonjour, comment allez-vous?",
    voice_id="JBFqnCBsd6RMkjVDRZzb",
    model_id="eleven_v4",
    language_code="fr"  # ISO 639-1 code
)

Text Normalization

Controls how numbers, dates, and abbreviations are converted to spoken words. For example, "01/15/2026" becomes "January fifteenth, twenty twenty-six":

  • "auto" (default): Model decides based on context
  • "on": Always normalize (use when you want natural speech)
  • "off": Speak literally (use when you want "zero one slash one five...")
audio = client.text_to_speech.convert(
    text="Call 1-800-555-0123 on 01/15/2026",
    voice_id="JBFqnCBsd6RMkjVDRZzb",
    model_id="eleven_v4",
    apply_text_normalization="on"
)

Request Stitching

When generating long audio in multiple requests, the audio can have pops, unnatural pauses, or tone shifts at the boundaries. Request stitching solves this by letting each request know what comes before/after it:

# First request
audio1 = client.text_to_speech.convert(
    text="This is the first part.",
    voice_id="JBFqnCBsd6RMkjVDRZzb",
    model_id="eleven_v4",
    next_text="And this continues the story."
)

# Second request using previous context
audio2 = client.text_to_speech.convert(
    text="And this continues the story.",
    voice_id="JBFqnCBsd6RMkjVDRZzb",
    model_id="eleven_v4",
    previous_text="This is the first part."
)

Output Formats

Format Description
mp3_44100_128 MP3 44.1kHz 128kbps (default) - compressed, good for web/apps
mp3_44100_192 MP3 44.1kHz 192kbps (Creator+) - higher quality compressed
mp3_44100_64 MP3 44.1kHz 64kbps - lower quality, smaller files
mp3_22050_32 MP3 22.05kHz 32kbps - smallest MP3 files
pcm_16000 Raw PCM 16kHz - use for real-time processing
pcm_22050 Raw PCM 22.05kHz
pcm_24000 Raw PCM 24kHz - good balance for streaming
pcm_44100 Raw PCM 44.1kHz (Pro+) - CD quality
pcm_48000 Raw PCM 48kHz (Pro+) - highest quality
ulaw_8000 μ-law 8kHz - standard for phone systems (Twilio, telephony)
alaw_8000 A-law 8kHz - telephony (alternative to μ-law)
opus_48000_64 Opus 48kHz 64kbps - efficient streaming codec
wav_44100 WAV 44.1kHz - uncompressed with headers

Streaming

Use the stream method to receive audio chunks as they're generated:

audio_stream = client.text_to_speech.stream(
    text="This text will be streamed as audio.",
    voice_id="JBFqnCBsd6RMkjVDRZzb",
    model_id="eleven_v4"
)

for chunk in audio_stream:
    play_audio(chunk)

For real-time apps where latency matters most, use eleven_v4_turbo (~100ms) over the Text to Dialogue WebSocket, or eleven_flash_v2_5 (~75ms) over HTTP streaming or the stream-input WebSocket. See references/streaming.md.

Error Handling

try:
    audio = client.text_to_speech.convert(
        text="Generate speech",
        voice_id="invalid-voice-id",
        model_id="eleven_v4"
    )
except Exception as e:
    print(f"API error: {e}")

Common errors:

  • 401: Invalid API key
  • 422: Invalid parameters (check voice_id, model_id)
  • 429: Rate limit exceeded

Tracking Costs

Monitor character usage via response headers (x-character-count, request-id):

response = client.text_to_speech.convert.with_raw_response(
    text="Hello!", voice_id="JBFqnCBsd6RMkjVDRZzb", model_id="eleven_v4"
)
audio = response.parse()
print(f"Characters used: {response.headers.get('x-character-count')}")

References

Embed badges

Add these to your README to show the skill's verification status.

SkillSafe verified badge
Verified badge
[![SkillSafe verified badge](https://api.skillsafe.ai/v1/badge/@elevenlabs/text-to-speech/verified)](https://skillsafe.ai/skill/@elevenlabs/text-to-speech/)
Installs badge
Installs badge
[![Installs badge](https://api.skillsafe.ai/v1/badge/@elevenlabs/text-to-speech/installs)](https://skillsafe.ai/skill/@elevenlabs/text-to-speech/)
Scan badge
Scan badge
[![Scan badge](https://api.skillsafe.ai/v1/badge/@elevenlabs/text-to-speech/scan)](https://skillsafe.ai/skill/@elevenlabs/text-to-speech/)
Eval pass rate badge
Eval pass rate
[![Eval pass rate badge](https://api.skillsafe.ai/v1/badge/@elevenlabs/text-to-speech/eval)](https://skillsafe.ai/skill/@elevenlabs/text-to-speech/)