Skip to main content
Run Rime on your own NVIDIA host as a paired API service and TTS model service. Each model deployment needs one of each service; multiple pairs can share a machine when capacity allows.
On-prem deployment requires registry access, a Rime license, and API credentials. Contact help@rime.ai before starting. The running services also need outbound HTTPS access to Rime’s usage and license endpoints.

Architecture

On-prem components The API service handles HTTP and WebSocket requests and verifies the license. The TTS service performs model inference. On-prem keeps all audio and text within your own network; nothing is sent to Rime’s cloud. This supports data-residency requirements and compliance with data privacy and protection regulations.

Published latency figures

  • Coda: Available on-premises. Setup details and performance numbers coming in a follow-up release.
  • Mist v2: Our tests have shown median latency of 175ms with randomly generated sentences between 40 and 50 characters on A10Gs and similar GPUs.
  • Arcana: See performance tuning.

Prerequisites

Hardware requirements

  • GPU
    • For Mist
      • NVIDIA T4, L4, A10, or higher
    • For Arcana
      • NVIDIA A100, H100 MIG 3g.40gb, or higher
  • Storage
    • 50 GB storage
  • CPU
    • 8 vCPUs
  • Memory requirements
    • 32 GiB

Software requirements

  • Supported Linux Distributions
    • Debian 12 (bookworm), x86_64
    • Ubuntu Server 24.04 (noble), x86_64
  • NVIDIA drivers
    • Minimum: 525.60.13
    • Recommended: 570.133.20 or higher
  • Docker
  • NVIDIA Container Toolkit

Installations

NVIDIA drivers

Follow https://www.nvidia.com/en-us/drivers to install the latest NVIDIA drivers, or use the following instructions on Debian-based systems:
NVIDIA Driver Installation (Debian-based)

Docker

Follow https://docs.docker.com/engine/install to install Docker on your system. Optionally, add the current user to the docker group for convenience: https://docs.docker.com/engine/install/linux-postinstall. The code snippets below assume that you can run docker as the current login.

NVIDIA Container Toolkit

Follow https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html to install the NVIDIA Container Toolkit. Note that you should follow both the Installation and the Configuration sections.

Verification

To verify that you have all the prerequisites installed, run the following command:
Verify Prerequisites
You should see your GPU listed in the output, alongside the driver version and CUDA version.

Firewall requirements

The Rime API service listens on port 8000 for HTTP, port 8002 for binary WebSockets, and port 8003 for JSON WebSockets. You will also need to allow the following outbound traffic in your firewall rules:
  • https://optimize.rime.ai/usage: registers on-prem usage with our servers.
  • https://optimize.rime.ai/license: verifies that your on-prem license is active.
  • us-docker.pkg.dev on port 443: container image registry.

API and registry credentials

Generate an API key

Create a Rime API key in the dashboard. Rime supplies the Artifact Registry key and license access separately.

Deployment

The following sections select the model and API images, configure the service pair, and start it with Docker Compose.

Artifact Registry login

Use the Artifact Registry key provided by Rime:
Log in to Artifact Registry

Container images

TTS service

Arcana

The Arcana images can be found at us-docker.pkg.dev/rime-labs/arcana/v2/<language>:<tag>.
  • The support languages are: en, es, fr, de, si.
  • The latest version is 20260420.

Arcana v3 (multilingual)

The Arcana v3 images can be found at us-docker.pkg.dev/rime-labs/arcana/v3/ennea:<tag>.
  • The support languages are: en, es, fr, pt, de, ja, ta, si, he.
  • The latest version is 20260420.
For Arcana only, you can also load the engine and data packages from different containers:
  • us-docker.pkg.dev/rime-labs/engine/arcana:<tag>
  • us-docker.pkg.dev/rime-labs/package/arcana/<language>:<tag>

Coda (multilingual)

The Coda v1 images can be found at us-docker.pkg.dev/rime-labs/coda/v1/coda:<tag>.
  • The support languages are: en, es, fr, pt, de, ja.
  • The latest version is 20260517.

Mist v3 (multilingual)

The Mist v3 images can be found at us-docker.pkg.dev/rime-labs/mist/v3/omni:<tag>
  • The support languages are: de, en, es, fr.
  • The latest version is 20260420.

API service

The latest image version is:
  • us-docker.pkg.dev/rime-labs/api/service:20260424

Docker Compose configuration

A simple way of deploying on a machine is to use Docker Compose. Create a compose.yml file with your editor of choice to define the services and their configurations:
compose.yml
When running on Kubernetes, ensure that MODEL_URL points to http://0.0.0.0:8080/invocations instead of the Docker Compose service name.

Multi-model backend

If you want to serve multiple Arcana languages via a single API instance, you can create a compose.yml like the following:
compose.yml
Note that the ARCANA_{LANG}_MODEL_URL environment variable must point to the container running the Arcana image for that language, but you should still point MODEL_URL to a default model container. The model environment variables currently supported are:
The API will route to these model backends based on the request parameter lang.

Authentication configuration

By default, callers must pass their Rime API key in every request via the Authorization: Bearer <key> header. Two additional environment variables let you configure authentication at the deployment level instead.

Pre-configuring the API key (RIME_API_KEY)

If you set RIME_API_KEY, the API service will use it to authenticate with the Rime license server automatically, and callers do not need to include an API key in their requests. You can supply it as an environment variable:
compose.yml
Or mount it as a secret file at /secrets/rime_api_key inside the container:
compose.yml
When neither is provided, the per-request Authorization header pathway remains active as normal.

Alternate API key header (API_KEY_HEADER)

On platforms that intercept the Authorization header, set API_KEY_HEADER to the name of an alternate header that callers will use to pass their Rime API key:
compose.yml
Callers then authenticate with:

Platform API key (PLATFORM_API_KEY)

On platforms that require authenticated inter-container requests, set PLATFORM_API_KEY so the API service can reach the model backend. It can also be mounted as a secret at /secrets/platform_api_key:
compose.yml

Start Docker Compose

Start Docker Compose
Allow approximately five minutes for model warm-up before sending the first synthesis request.

Verify health

Health check
A ready deployment returns apiStatus: "ok", licenseStatus: "valid", and modelReachable: true:

AWS video walkthrough

Requests and response formats

HTTP requests

Request:
Request example
Response:
Response format
Sample response file: result.txt

Receiving a response in MP3 format

Request:
Request example
Response: Sample response file: result.mp3

Receiving a response in PCM (raw) format

Request:
Request example
Response: Sample response file: result.pcm

WebSocket endpoints

JSON websockets

The JSON WebSocket endpoint compatible with coda, arcana, and mist models will be served at port 8003. For example, ws://localhost:8003, which is equivalent to our cloud JSON WebSocket API. See the Coda JSON WebSocket docs, Arcana JSON WebSocket docs, and Mist JSON WebSocket docs depending on which model backend you have configured.

Non-JSON websockets

The non-JSON WebSocket endpoint will be served at port 8002. For example, ws://localhost:8002, which is equivalent to our cloud WebSocket API.

Deprecated

A deprecated WebSocket endpoint is served on port 8001. It is only compatible with the Mist model family; use the current WebSocket endpoints for new integrations.