SV Premium

Migration guides

Migrate text-to-speech from ElevenLabs to Inworld TTS

Rebuild voices and streaming calls, check pronunciation and audio formats, and run a blind listening test before a staged rollout.

From ElevenLabs to Inworld TTS · Published 2026-10-08

Before you commit

Hands-on effort
1–3 days of hands-on integration
Elapsed time
Allow 1–2 weeks for blind listening tests, latency measurement and staged rollout.
Examples cover
Inworld TTS-2 API · ElevenLabs Multilingual v2 example
Content review
2026-10-09
Integration testing
Not recorded. Rehearse the commands and rollback on staging.

Prerequisites

  • Original recordings and voice rights, both API accounts, production scripts and a streaming test harness.
  • A data-protection decision and listeners who can compare quality without knowing the provider.

Don’t migrate yet if…

  • You require a self-serve DPA or EU residency that the proposed Inworld plan does not offer.
  • Voice quality, cloning rights, pronunciation or measured latency fail your acceptance criteria.

ElevenLabs lists its standard models (v4, v3, Multilingual v2) at $80 per million characters on every API plan. The plan fee buys an allowance at that same rate. Inworld charges $25 per million for Realtime TTS-2 on demand, or $20 down to $12.50 on monthly plans.

At 5 million characters a month, the estimates are about $400 on ElevenLabs and $100 on Inworld’s Builder plan. At 20 million, they’re about $1,600 and $300 on Inworld’s Developer plan. ElevenLabs ran a v4 launch discount until 12 October 2026; these figures use the regular list prices. Compare your character usage in the calculator.

Inworld supports HTTP and WebSocket streaming, instant voice cloning and word timestamps. It costs less per character, but your voices will not sound identical after the move. If you don’t need cloning, Kokoro-82M on DeepInfra costs $0.62 per million characters and is another option to test. For self-hosted cloning, Chatterbox is MIT-licensed and runs on your own GPUs.

Before changing providers, compare the voices with listeners who use your product, measure streaming latency from your own servers, and check pronunciation and audio formats.

0. Before you start

1. Pick and rebuild voices

ElevenLabs voice IDs don’t carry over. For each voice in production:

2. Swap the code

Before (ElevenLabs HTTP streaming):

import requests
r = requests.post(
    f"https://api.elevenlabs.io/v1/text-to-speech/{voice_id}/stream",
    params={"output_format": "pcm_24000"},
    headers={"xi-api-key": ELEVENLABS_API_KEY},
    json={"text": text, "model_id": "eleven_multilingual_v2",
          "voice_settings": {"stability": 0.5, "similarity_boost": 0.75}},
    stream=True,
)
r.raise_for_status()
for chunk in r.iter_content(chunk_size=4096):
    player.write(chunk)  # raw 16-bit PCM, 24 kHz

After (Inworld HTTP streaming):

import base64, json, requests
r = requests.post(
    "https://api.inworld.ai/tts/v1/voice:stream",
    headers={"Authorization": f"Basic {INWORLD_API_KEY}"},
    json={"text": text, "voiceId": "Ashley", "modelId": "inworld-tts-2",
          "audioConfig": {"audioEncoding": "PCM", "sampleRateHertz": 24000},
          "deliveryMode": "BALANCED"},
    stream=True,
)
r.raise_for_status()
for line in r.iter_lines():
    if not line:
        continue
    msg = json.loads(line)
    if "error" in msg:
        raise RuntimeError(msg["error"].get("message"))
    audio = msg.get("result", {}).get("audioContent")
    if audio:
        player.write(base64.b64decode(audio))  # raw 16-bit PCM, 24 kHz

Three differences catch people out:

For voice agents that feed LLM tokens through ElevenLabs’ wss://api.elevenlabs.io/v1/text-to-speech/{voice_id}/stream-input, use Inworld’s WebSocket API at wss://api.inworld.ai/tts/v1/voice:streamBidirectional. Its model is different: you open a context with create, send sendText messages, then flushContext or closeContext, and wait for contextClosed before you stop reading. autoMode: true with autoModeStrategy: "SENTENCE_BOUNDARY" (preview) does the sentence buffering for you.

3. Feature mapping

ElevenLabs Inworld
voice_id in the URL path voiceId in the body
eleven_v4, eleven_v3, eleven_multilingual_v2 inworld-tts-2
eleven_flash_v2_5, eleven_v4_turbo inworld-tts-2-flash ($15/M on demand; no steering)
stability deliveryMode: STABLE, BALANCED, CREATIVE (TTS-2 only)
similarity_boost, style, use_speaker_boost No equivalent. Similarity depends on the clone sample
speed audioConfig.speakingRate (0.5–1.5)
v3 audio tags instruction field or inline [bracket] steering tags (TTS-2 only)
output_format=mp3_44100_128 audioEncoding: "MP3", sampleRateHertz: 44100, bitRate: 128000
pcm_24000 audioEncoding: "PCM", sampleRateHertz: 24000
ulaw_8000 audioEncoding: "MULAW", sampleRateHertz: 8000
opus_* OGG_OPUS
/stream (HTTP) POST /tts/v1/voice:stream (NDJSON)
/with-timestamps, character alignment timestampType: "CHARACTER" or "WORD", returned in timestampInfo
Pronunciation dictionaries (pronunciation_dictionary_locators) Inline IPA only, e.g. /kriːt/, one word per pair of slashes
<break time="1s" /> Same tag, up to 20 per request, 10 s max each
apply_text_normalization applyTextNormalization: ON, OFF
previous_text synthesisContext.previousRequests
Instant / professional voice cloning Instant cloning; professional cloning in beta

Pronunciation. Inworld has no stored dictionary. Export your ElevenLabs dictionaries, convert alias rules to plain text replacements and phoneme rules to English IPA (Inworld rejects ARPAbet), and apply them in your own code before each request.

Timestamps. ElevenLabs returns alignment.characters with character_start_times_seconds. Inworld’s CHARACTER mode returns characterAlignment with characterStartTimeSeconds. WORD mode adds phoneme detail. Set timestampTransportStrategy explicitly: SYNC keeps timestamps with their audio, and ASYNC is faster but can send them after it.

4. Blind listening test

  1. Pick 30–50 real scripts from production logs, including numbers, dates, product names, long paragraphs and short agent replies.
  2. Render each with the current ElevenLabs voice and your two or three Inworld candidates, in the same format and sample rate. Normalize loudness so the louder clip doesn’t win.
  3. Play randomized, unlabeled A/B pairs to at least five listeners who use the product, not just engineers. Ask which they prefer and flag mispronunciations.
  4. Measure time-to-first-byte from your own servers: start a timer at the request, stop it at the first decoded audio chunk, over at least 100 requests per provider at your usual concurrency. Inworld’s published latencies are server-side P90 and exclude the network.

Pick the voice that wins or ties. Fix mispronunciations with the IPA list from step 3 and re-test those clips.

5. Rollback plan

Put the provider behind a flag (TTS_PROVIDER=elevenlabs|inworld), with per-voice mapping in config, and ramp by percentage of sessions, not of requests, so one conversation never switches voices midway. Cache generated audio keyed by provider, voice, model, settings and a text hash. Cached ElevenLabs audio keeps serving repeat content during the ramp, and rolling back means flipping the flag. Keep ElevenLabs on the cheapest plan that covers your rollback traffic until the ramp finishes.

6. Clean up

Found a command or version mismatch? Report a correction with the version and steps to reproduce it. Remove secrets and customer data first.