Skip to main content
The Rime API authenticates every request with a bearer token in the Authorization header: Authorization: Bearer YOUR_API_KEY. See API authentication for how to create a key. The API returns a single JSON object, {"audioContent": "..."}, whose audioContent field holds base64-encoded MP3 audio. Decode it before you save or play it. Only Mist v1 and Mist v2 return this envelope, and only when Accept doesn’t name a streamable audio type, which is why the examples send application/json. The response arrives after synthesis finishes, so use streaming HTTP if you want playback to start sooner.

Fixed parameters

mp3
required

Variable parameters

string
required
Must be a voice from the Rime voice catalog.
string
required
The text you’d like spoken. Character limit per request is 1,000 via the API and in the dashboard UI.
string
Set to mistv2.
string
default:"eng"
If provided, the language must match the language spoken by the selected speaker. Verify the pairing in the Rime voice catalog.
bool
default:"false"
When set to true, adds pauses between words enclosed in angle brackets. The number inside the brackets specifies the pause duration in milliseconds.
Example: “Hi. <200> I’d love to have a conversation with you.” adds a 200ms pause between the first and second sentences.
bool
default:"false"
When set to true, you can specify the phonemes for a word enclosed in curly brackets.
Example: “{h’El.o} World” pronounces “Hello” with the phonemes you supplied. Learn more about custom pronunciation.
string
Comma-separated list of speed values applied to words in square brackets. Values < 1.0 speed up speech, > 1.0 slow it down. Example: For text “This is [slow] and [fast]”, use “3, 0.5” to make “slow” slower and “fast” faster.
int
The value, if provided, must be between 4000 and 44100. Default: 22050
float
default:"1.0"
Adjusts the speed of speech. Lower than 1.0 is faster and higher than 1.0 is slower.Note: this is the legacy Mist v2 convention. Coda and Mist v3 invert it, so for those models higher than 1.0 is faster.
bool
default:"false"
mist/mistv2 only. Skips text normalization before synthesis. This reduces latency, but digits and abbreviations reach the model unexpanded and may be mispronounced.