Platform

Model Playground

Under Models in the sidebar, listen to a voice or watch a transcript come back before you ever put it in front of a caller. Nothing here touches an agent; it is the fastest way to answer "does this voice sound right" or "does it understand this accent".

Generate speech

Open Models → Text to Speech. Type or paste text in the box, pick a voice from the picker, and select Generate speech. The audio plays back in the panel on the right as soon as it is ready, and stays there so you can compare it against the next voice you try.

Screenshot: The Text to Speech playground: text box on the left, voice picker and Settings above it, generated audio on the right.

Settings exposes the same knobs an agent's Voice tab does for that provider (language, model, speed where the provider supports it), so a setting that sounds right here is the one to carry into the agent.

Choosing a voice

Use Search voices to filter by name or language. Every voice shows a Copy voice ID action — that ID is what you paste into an agent's Voice tab, so trying a voice here and then setting it on an agent is a two-click round trip, not a search all over again.

Cloned voices (see the Library guide for how to clone one) show up in the same picker once they finish processing, mixed in with the platform voices rather than in a separate list.

Test live transcription

Open Models → Speech to Text and select Start listening. Your browser will ask for microphone access; once granted, speak, and the words appear as you say them. This is a real streaming connection to the same transcription engine a phone call uses, not a canned demo.

Screenshot: The Speech to Text playground mid-session: 'Listening.' status, live partial transcript, and the session timer.

Select Stop listening when you are done, or let the session time out on its own.

Session limits and languages

An authenticated session is capped at 5 minutes; the public playground on the marketing site is capped at 45 seconds and limited to a small number of concurrent listeners, to keep it usable for everyone trying it at once. Neither cap is something you configure — they exist so a forgotten open tab does not run, or bill, indefinitely.

The playground listens in whatever language the audio is spoken in; there is no language picker to set first, unlike an agent's Transcription tab where you can pin a language hint for a caller you already know speaks one language most of the time.

Cloning a voice

Voice cloning lives under Configurations → Voices, next to the playground rather than inside it: upload a clean sample of the voice you want, give it a name, and it appears in the Text to Speech voice picker once processing finishes, ready to try the same way as any built-in voice.

How playground usage is billed

Both playgrounds draw from the same wallet a live call does — text to speech by the character generated, speech to text by the second the microphone session is open, whether or not anything was said in that second (see How each service is billed). Trying five voices back to back is five separate charges, not one.

Tip: a service credit (a coupon issued for TTS or STT specifically) is drawn from before your wallet balance is touched, in the playground exactly as on a call — there is nothing separate to redeem here.

Next