Migrate text-to-speech from ElevenLabs to Inworld TTS
Rebuild voices and streaming calls, check pronunciation and audio formats, and run a blind listening test before a staged rollout.
From ElevenLabs to Inworld TTS · Published 2026-10-08
Before you commit
- Hands-on effort
- 1–3 days of hands-on integration
- Elapsed time
- Allow 1–2 weeks for blind listening tests, latency measurement and staged rollout.
- Examples cover
- Inworld TTS-2 API · ElevenLabs Multilingual v2 example
- Content review
- 2026-10-09
- Integration testing
- Not recorded. Rehearse the commands and rollback on staging.
Prerequisites
- Original recordings and voice rights, both API accounts, production scripts and a streaming test harness.
- A data-protection decision and listeners who can compare quality without knowing the provider.
Don’t migrate yet if…
- You require a self-serve DPA or EU residency that the proposed Inworld plan does not offer.
- Voice quality, cloning rights, pronunciation or measured latency fail your acceptance criteria.
ElevenLabs lists its standard models (v4, v3, Multilingual v2) at $80 per million characters on every API plan. The plan fee buys an allowance at that same rate. Inworld charges $25 per million for Realtime TTS-2 on demand, or $20 down to $12.50 on monthly plans.
At 5 million characters a month, the estimates are about $400 on ElevenLabs and $100 on Inworld’s Builder plan. At 20 million, they’re about $1,600 and $300 on Inworld’s Developer plan. ElevenLabs ran a v4 launch discount until 12 October 2026; these figures use the regular list prices. Compare your character usage in the calculator.
Inworld supports HTTP and WebSocket streaming, instant voice cloning and word timestamps. It costs less per character, but your voices will not sound identical after the move. If you don’t need cloning, Kokoro-82M on DeepInfra costs $0.62 per million characters and is another option to test. For self-hosted cloning, Chatterbox is MIT-licensed and runs on your own GPUs.
Before changing providers, compare the voices with listeners who use your product, measure streaming latency from your own servers, and check pronunciation and audio formats.
0. Before you start
- Check your ElevenLabs renewal date. Characters reset each billing cycle, so plan the cutover to land just before a renewal, then cancel for the end of that cycle.
- Pull a month of character usage per model from the ElevenLabs usage page. That’s your calculator input. Inworld bills per input character, counted in UTF-16 code units.
- List what you use: voice IDs (library, cloned, professional clones), model IDs,
voice_settings, output formats, pronunciation dictionaries, timestamps, WebSocket streaming. - Data protection. On Inworld, a DPA, EU data residency and a BAA are Enterprise-only. If you need a self-serve DPA, stop here and pick another target.
1. Pick and rebuild voices
ElevenLabs voice IDs don’t carry over. For each voice in production:
- Library voices: shortlist two or three Inworld built-in voices with a similar age, accent and pace (
GET https://api.inworld.ai/tts/v1/voices, or the TTS Playground). - Cloned voices: re-clone from your original recordings, not from ElevenLabs output. Instant voice cloning takes 3 to 30 seconds of audio (wav, mp3 or webm, up to 12 MB). Professional voice cloning is in beta and only on the larger plans.
- Consent. Inworld’s Service Specific Terms require you to have the rights to any voice you upload, and its Acceptable Use Policy forbids impersonation without consent. The consent you collected for ElevenLabs may name ElevenLabs as the processor, so check its wording and get fresh written consent if needed.
2. Swap the code
Before (ElevenLabs HTTP streaming):
import requests
r = requests.post(
f"https://api.elevenlabs.io/v1/text-to-speech/{voice_id}/stream",
params={"output_format": "pcm_24000"},
headers={"xi-api-key": ELEVENLABS_API_KEY},
json={"text": text, "model_id": "eleven_multilingual_v2",
"voice_settings": {"stability": 0.5, "similarity_boost": 0.75}},
stream=True,
)
r.raise_for_status()
for chunk in r.iter_content(chunk_size=4096):
player.write(chunk) # raw 16-bit PCM, 24 kHz
After (Inworld HTTP streaming):
import base64, json, requests
r = requests.post(
"https://api.inworld.ai/tts/v1/voice:stream",
headers={"Authorization": f"Basic {INWORLD_API_KEY}"},
json={"text": text, "voiceId": "Ashley", "modelId": "inworld-tts-2",
"audioConfig": {"audioEncoding": "PCM", "sampleRateHertz": 24000},
"deliveryMode": "BALANCED"},
stream=True,
)
r.raise_for_status()
for line in r.iter_lines():
if not line:
continue
msg = json.loads(line)
if "error" in msg:
raise RuntimeError(msg["error"].get("message"))
audio = msg.get("result", {}).get("audioContent")
if audio:
player.write(base64.b64decode(audio)) # raw 16-bit PCM, 24 kHz
Three differences catch people out:
- Auth.
INWORLD_API_KEYis the Base64 credential copied from the Portal. Send it as is, without encoding it again. - Framing. ElevenLabs streams raw audio bytes. Inworld streams newline-delimited JSON with Base64 audio in
result.audioContent. A stream can return HTTP 200 and still end with anerrorline. - Text limits. 4,000 characters per streaming request and 2,000 on the non-streaming
POST /tts/v1/voice. ElevenLabs allows 10,000 on Multilingual v2, so long-form jobs need splitting. Pass the preceding text insynthesisContext.previousRequeststo keep the prosody consistent across chunks.
For voice agents that feed LLM tokens through ElevenLabs’ wss://api.elevenlabs.io/v1/text-to-speech/{voice_id}/stream-input, use Inworld’s WebSocket API at wss://api.inworld.ai/tts/v1/voice:streamBidirectional. Its model is different: you open a context with create, send sendText messages, then flushContext or closeContext, and wait for contextClosed before you stop reading. autoMode: true with autoModeStrategy: "SENTENCE_BOUNDARY" (preview) does the sentence buffering for you.
3. Feature mapping
| ElevenLabs | Inworld |
|---|---|
voice_id in the URL path |
voiceId in the body |
eleven_v4, eleven_v3, eleven_multilingual_v2 |
inworld-tts-2 |
eleven_flash_v2_5, eleven_v4_turbo |
inworld-tts-2-flash ($15/M on demand; no steering) |
stability |
deliveryMode: STABLE, BALANCED, CREATIVE (TTS-2 only) |
similarity_boost, style, use_speaker_boost |
No equivalent. Similarity depends on the clone sample |
speed |
audioConfig.speakingRate (0.5–1.5) |
| v3 audio tags | instruction field or inline [bracket] steering tags (TTS-2 only) |
output_format=mp3_44100_128 |
audioEncoding: "MP3", sampleRateHertz: 44100, bitRate: 128000 |
pcm_24000 |
audioEncoding: "PCM", sampleRateHertz: 24000 |
ulaw_8000 |
audioEncoding: "MULAW", sampleRateHertz: 8000 |
opus_* |
OGG_OPUS |
/stream (HTTP) |
POST /tts/v1/voice:stream (NDJSON) |
/with-timestamps, character alignment |
timestampType: "CHARACTER" or "WORD", returned in timestampInfo |
Pronunciation dictionaries (pronunciation_dictionary_locators) |
Inline IPA only, e.g. /kriːt/, one word per pair of slashes |
<break time="1s" /> |
Same tag, up to 20 per request, 10 s max each |
apply_text_normalization |
applyTextNormalization: ON, OFF |
previous_text |
synthesisContext.previousRequests |
| Instant / professional voice cloning | Instant cloning; professional cloning in beta |
Pronunciation. Inworld has no stored dictionary. Export your ElevenLabs dictionaries, convert alias rules to plain text replacements and phoneme rules to English IPA (Inworld rejects ARPAbet), and apply them in your own code before each request.
Timestamps. ElevenLabs returns alignment.characters with character_start_times_seconds. Inworld’s CHARACTER mode returns characterAlignment with characterStartTimeSeconds. WORD mode adds phoneme detail. Set timestampTransportStrategy explicitly: SYNC keeps timestamps with their audio, and ASYNC is faster but can send them after it.
4. Blind listening test
- Pick 30–50 real scripts from production logs, including numbers, dates, product names, long paragraphs and short agent replies.
- Render each with the current ElevenLabs voice and your two or three Inworld candidates, in the same format and sample rate. Normalize loudness so the louder clip doesn’t win.
- Play randomized, unlabeled A/B pairs to at least five listeners who use the product, not just engineers. Ask which they prefer and flag mispronunciations.
- Measure time-to-first-byte from your own servers: start a timer at the request, stop it at the first decoded audio chunk, over at least 100 requests per provider at your usual concurrency. Inworld’s published latencies are server-side P90 and exclude the network.
Pick the voice that wins or ties. Fix mispronunciations with the IPA list from step 3 and re-test those clips.
5. Rollback plan
Put the provider behind a flag (TTS_PROVIDER=elevenlabs|inworld), with per-voice mapping in config, and ramp by percentage of sessions, not of requests, so one conversation never switches voices midway. Cache generated audio keyed by provider, voice, model, settings and a text hash. Cached ElevenLabs audio keeps serving repeat content during the ramp, and rolling back means flipping the flag. Keep ElevenLabs on the cheapest plan that covers your rollback traffic until the ramp finishes.
6. Clean up
- Cancel the ElevenLabs subscription for the end of the current cycle. Unused characters don’t carry over once the plan ends.
- ElevenLabs’ help center says you keep a commercial license, with no time limit, for audio generated while you were subscribed. Audio generated outside a paid subscription isn’t covered. Download anything you still need first, because ElevenLabs doesn’t guarantee access to stored history on the free tier.
- Delete cloned voices you no longer need on ElevenLabs, and revoke the API keys.
- On Inworld, the pricing page lists a commercial license on every plan, including On-Demand, and the terms assign Outputs to you. If you buy credits, consider auto-reload on the Billing page, which tops up the balance when it falls below a threshold you set.
Found a command or version mismatch? Report a correction with the version and steps to reproduce it. Remove secrets and customer data first.