Changelog
API updates, new endpoints, and breaking changes.
v1.2.0speech-to-textSeptember 2026
- +Speech-to-Text API: transcribe audio with Valmiki, our transcription engine. 100+ languages with automatic detection and punctuation.
- +POST /v1/transcribe takes a finished recording, as a raw body or multipart upload, and returns the transcript with its duration. WAV, MP3, FLAC, OGG, Opus, WebM, AAC and M4A, up to 10 MB and 10 minutes.
- +WS /v1/transcribe/stream takes 16 kHz mono PCM audio and returns partial and final transcripts as they are recognised. Browsers that cannot set a WebSocket header may pass the key as api_key in the query string.
- +Transcription is billed per second of audio, because a transcriber charges for the time it is listening whether or not anyone is speaking. A recording is billed for its own duration; a stream for as long as it is open. Speech-to-text credits are drawn before your wallet.
- +Audio we reject (unsupported format, over 10 MB, over 10 minutes, unreadable) is never billed. An Idempotency-Key makes a retry safe: the transcript comes back again with seconds_billed of 0.
v1.1.0callsSeptember 2026
- +Call history API: GET /v1/calls returns your calls newest first, with filters for agent, status, direction, phone number and date range. Paginated with a cursor rather than page numbers, so a page never repeats or skips a call while new ones arrive.
- +Call detail: GET /v1/calls/{call_id} returns the models used, the itemised cost, the transcript and every tool the agent invoked with its arguments, result and duration. Use the include parameter to choose which sections are returned.
- +Transcripts: GET /v1/calls/{call_id}/transcript returns the transcript on its own, as turns and as a single block of text.
- +Recordings: GET /v1/calls/{call_id}/recording redirects to the audio, or returns a time limited link as JSON with redirect=false. Links are valid for 15 minutes.
- +Voice previews now have a stable address. preview_url on each voice points at GET /v1/voices/{voice_id}/preview, which needs no API key and can be used directly as the src of an audio element.
- +Every call endpoint is scoped to the tenant that owns the API key.
v1.0.0launchJune 2026
- +Public API launched: Text-to-Speech, Voices. Agents, Agent Tools, and Call Recordings coming soon.
- +Pāṇini TTS: 250+ languages, 250-300ms latency, streaming synthesis over HTTP.
- +Voice cloning from as little as 10 seconds of reference audio.
- +Agent builder with tool library (coming soon): attach webhook or function tools to any agent.
- +Bearer token authentication with per-key scoping.