Skip to main content
The Rime API authenticates every request with a bearer token in the Authorization header: Authorization: Bearer YOUR_API_KEY. See API authentication for how to create a key. When you set Accept: text/event-stream, the /v1/rime-tts route streams the response as server-sent events. SSE is specific to Mist v2, so use streaming HTTP or WebSockets for other models.

Fixed headers

text/event-stream
required

Variable parameters

string
required
Must be a voice from the Rime voice catalog.
string
required
The text you’d like spoken. Character limit per request is 1,000 via the API and in the dashboard UI.
string
Set to mistv2.
string
One of mp3, mulaw, or pcm
string
default:"eng"
If provided, the language must match the language spoken by the selected speaker. Verify the pairing in the Rime voice catalog.
bool
default:"false"
When set to true, adds pauses between words enclosed in angle brackets. The number inside the brackets specifies the pause duration in milliseconds.
Example: “Hi. <200> I’d love to have a conversation with you.” adds a 200ms pause between the first and second sentences.
bool
default:"false"
When set to true, you can specify the phonemes for a word enclosed in curly brackets.
Example: “{h’El.o} World” pronounces “Hello” with the phonemes you supplied. Learn more about custom pronunciation.
string
Comma-separated list of speed values applied to words in square brackets. Values < 1.0 speed up speech, > 1.0 slow it down. Example: “This is [slow] and [fast]”, use “3, 0.5” to make “slow” slower and “fast” faster.
int
The value, if provided, must be between 4000 and 44100. Default: 22050
float
default:"1.0"
Adjusts the speed of speech. Lower than 1.0 is faster and higher than 1.0 is slower.Note: this is the legacy Mist v2 convention. Coda and Mist v3 invert it, so for those models higher than 1.0 is faster.
bool
default:"false"
mist/mistv2 only. Skips text normalization before synthesis. This reduces latency, but digits and abbreviations reach the model unexpanded and may be mispronounced.

Example output

chunk events carry base64-encoded audio in data. Decode each one and append the bytes in order to rebuild the audio. The timestamps event gives word-level timings, and done marks the end of the stream, so stop reading when it arrives.