Skip to main content
Rime is voice AI made for human conversation: natural, low-latency speech for production voice agents, IVR, and telephony. Its flagship model, Coda, is trained on real full-duplex conversations between real people, not voice actors or audiobook narrators, so it holds up in production. Coda earns top-rated voice quality in human evaluations and streams at sub-100ms model latency. Deploy it in the cloud, your own VPC, or fully on-premises, across 600+ voices and 50+ languages.

Rime at a glance

Rime is a text-to-speech platform for real-time voice: conversational agents, IVR, telephony, and any product that needs to speak. The same speech API runs in Rime’s cloud, in your virtual private cloud, or fully on-premises, and scales from a free trial to enterprise volume.

Talk to the team about enterprise, on-prem, or volume pricing

Custom deployments, compliance (SOC 2, HIPAA), SLAs, and dedicated support for production scale.

Start with Coda

Coda is the default model for new projects and the recommended destination for Arcana traffic. Existing Arcana integrations can switch by setting modelId: coda; the rest of the API contract stays the same. Arcana is being sunset on August 15, 2026; see Models for migration details. Coda pairs an LLM backbone with a dedicated speech inference engine trained on full-duplex conversations between real people. It delivers top-rated voice quality in human evaluations, sub-100ms model latency (sub-200ms end to end over the cloud API), and one voice lineup across English, Spanish, French, Portuguese, German, and Japanese.

Generate your first audio with Coda

Send one request with modelId: coda and get natural speech back in under five minutes.

Make voice part of your product

Generate speech

Turn text into playable or downloadable audio in formats suited to web, mobile, telephony, and media workflows.

Stream speech in real time

Use HTTP for simple streaming responses or persistent WebSockets for the tightest conversational loop, with timestamps and interruption handling.

Choose voices, models, and languages

Browse distinct speakers and select the speech model, language, and sound that fit your application.

Shape how speech sounds

Guide delivery, pacing, spelling, pauses, pronunciation, and the way numbers, dates, and other text are spoken.

Build voice agents

Begin with a complete LiveKit tutorial or connect Rime directly from Next.js, Vite, Express, Node.js, or FastAPI.

Connect voice platforms

Integrate with LiveKit, Pipecat, Vapi, Daily, SignalWire, VideoSDK, and other voice application frameworks.

Create a custom voice

Make a branded enterprise voice available to the same speech API as Rime’s voice catalog.

Run Rime in your environment

Use Rime’s regional cloud endpoints or deploy supported models on your own infrastructure.

Build with the tools you use

Dashboard

Explore, preview, and save voices; generate speech and inspect normalized text. Create API tokens, follow setup progress, manage teams, and review usage and billing.

HTTP and WebSocket APIs

Generate complete audio responses, stream audio, receive word timestamps, list voices, and use Rime’s language tools.

Application code

Use standard HTTP and WebSocket clients from Python, JavaScript, Go, or any server-side language; no Rime-specific SDK is required.

CLI

Generate, play, and save speech, inspect usage, and test response speed from the terminal.

MCP server

Let Claude, Codex, and compatible tools browse voices, generate samples, inspect pronunciation and text normalization, and scaffold integrations.

Application integrations

Use Rime with supported voice-agent frameworks, telephony services, deployment platforms, and application builders.
One Rime API token authenticates the API, CLI, MCP tools, and supported integrations. Tokens, team access, and billing are managed in the Rime dashboard.