> ## Documentation Index
> Fetch the complete documentation index at: https://docs.rime.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Quickstart

> Deploy a licensed Rime API and TTS service pair on an NVIDIA host with Docker Compose.

Run Rime on your own NVIDIA host as a paired API service and TTS model service. Each model deployment needs one of each service; multiple pairs can share a machine when capacity allows.

<Warning>
  On-prem deployment requires registry access, a Rime license, and API credentials. Contact [help@rime.ai](mailto:help@rime.ai) before starting. The running services also need outbound HTTPS access to Rime's usage and license endpoints.
</Warning>

## Architecture

<img src="https://mintcdn.com/rimelabs/DVHs1HOnPvW2NRCW/images/on-prem-components_update.png?fit=max&auto=format&n=DVHs1HOnPvW2NRCW&q=85&s=ff1fd5753b0596d22bb7e1a279f406e5" alt="On-prem components" width="1582" height="806" data-path="images/on-prem-components_update.png" />

The API service handles HTTP and WebSocket requests and verifies the license. The TTS service performs model inference.

On-prem keeps all audio and text inside your own network: no audio or text reaches Rime's cloud. That supports data-residency requirements and compliance with data privacy and protection regulations.

## Published latency figures

* **Coda:** This page documents the current image and service-pair setup. Rime has not published Coda on-prem performance numbers.
* **Mist v2:** Rime measured median latency of **175ms** with randomly generated sentences between 40 and 50 characters on A10Gs and similar GPUs.
* **Arcana:** See [performance tuning](/docs/on-prem/performance).

# Prerequisites

## Hardware requirements

* GPU
  * For Mist
    * NVIDIA T4, L4, A10, or higher
  * For Coda
    * Confirm GPU requirements with Rime before provisioning
  * For Arcana
    * NVIDIA A100, H100 MIG `3g.40gb`, or higher
* Storage
  * 50 GB storage
* CPU
  * 8 vCPUs
* Memory requirements
  * 32 GiB

## Software requirements

* Supported Linux Distributions
  * Debian 12 (`bookworm`), x86\_64
  * Ubuntu Server 24.04 (`noble`), x86\_64
* NVIDIA drivers
  * Minimum: `525.60.13`
  * Recommended: `570.133.20` or higher
* Docker
* NVIDIA Container Toolkit

### Installations

#### NVIDIA drivers

Follow [https://www.nvidia.com/en-us/drivers](https://www.nvidia.com/en-us/drivers) to install the latest NVIDIA drivers, or use the following instructions on Debian-based systems:

```bash NVIDIA Driver Installation (Debian-based) theme={null}
# Update packages
sudo apt-get update

# Install basic toolchain and kernel headers
sudo apt-get install -y gcc make wget linux-headers-$(uname -r)

# Download and install the NVIDIA driver.
NVIDIA_DRIVER_VERSION=580.95.05
NVIDIA_DRIVER_PATH=/opt/NVIDIA-Linux-x86_64-${NVIDIA_DRIVER_VERSION}.run
sudo rm -f "${NVIDIA_DRIVER_PATH}"
sudo wget "https://us.download.nvidia.com/tesla/${NVIDIA_DRIVER_VERSION}/NVIDIA-Linux-x86_64-${NVIDIA_DRIVER_VERSION}.run" -O "${NVIDIA_DRIVER_PATH}"
sudo chmod +x "${NVIDIA_DRIVER_PATH}"
sudo "${NVIDIA_DRIVER_PATH}" --silent --no-questions
```

#### Docker

Follow [https://docs.docker.com/engine/install](https://docs.docker.com/engine/install) to install Docker on your system.

Optionally, add the current user to the `docker` group for convenience: [https://docs.docker.com/engine/install/linux-postinstall](https://docs.docker.com/engine/install/linux-postinstall).
The code snippets below assume that you can run `docker` as the current login.

#### NVIDIA Container Toolkit

Follow [https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html) to install the NVIDIA Container Toolkit.

Follow both the **Installation** and the **Configuration** sections.

#### Verification

To verify that you have all the prerequisites installed, run the following command:

```bash Verify Prerequisites theme={null}
docker run --rm --gpus all nvidia/cuda:12.8.1-base-ubi9 nvidia-smi
```

You should see your GPU listed in the output, alongside the driver version and CUDA version.

## Firewall requirements

The Rime API service listens on port 8000 for HTTP, port 8002 for binary WebSockets, and port 8003 for JSON WebSockets.

Allow the following outbound traffic in your firewall rules:

* `https://optimize.rime.ai/usage`: the API service calls this endpoint to register on-prem usage with Rime's usage service.
* `https://optimize.rime.ai/license`: the API service calls this endpoint to verify that your on-prem license is active.
* `us-docker.pkg.dev` on port 443: container image registry.

# API and registry credentials

## Generate an API key

In the Rime dashboard, create an API key on the [**API Tokens**](https://app.rime.ai/tokens/) page. Rime supplies the Artifact Registry key and license access separately.

# Deployment

Select the model and API images, configure the service pair, then start it with Docker Compose.

## Artifact Registry login

Use the Artifact Registry key provided by Rime:

```bash Log in to Artifact Registry theme={null}
cat KEY-FILE | docker login -u _json_key --password-stdin https://us-docker.pkg.dev
```

## Container images

### Image tags

<Warning>
  All model images and the API image are published together under a single `YYYYMMDD` tag. Deploy the **same tag** across every container; mixing tags pairs components that were never released or tested together.

  The current release tag is `20260801`.
</Warning>

Substitute that tag for `<tag>` in the image references below, and update all containers together when you upgrade.

### TTS service

#### Arcana

<Note>The August 15 Arcana cutoff applies to the cloud API. The Arcana images below remain available for on-prem deployments.</Note>

The Arcana images can be found at `us-docker.pkg.dev/rime-labs/arcana/v2/<language>:<tag>`.

* The supported languages are: `en`, `es`, `fr`, `de`, `ar`, `hi`, `si`.

#### Arcana v3 (multilingual)

The Arcana v3 images can be found at `us-docker.pkg.dev/rime-labs/arcana/v3/ennea:<tag>`.

* The supported languages are: `en`, `es`, `fr`, `pt`, `de`, `ja`, `si`, `he`.

For Arcana only, you can also load the engine and data packages from different containers:

* `us-docker.pkg.dev/rime-labs/engine/arcana:<tag>`
* `us-docker.pkg.dev/rime-labs/package/arcana/<language>:<tag>`

#### Coda (multilingual)

The Coda v1 images can be found at `us-docker.pkg.dev/rime-labs/coda/v1/coda:<tag>`.

* The supported languages are: `en`, `es`, `fr`, `pt`, `de`, `ja`, `ar`.

<Note>Hindi is available through the cloud Coda API but is not included in the current `20260801` on-prem image.</Note>

#### Mist v3 (multilingual)

The Mist v3 images can be found at `us-docker.pkg.dev/rime-labs/mist/v3/omni:<tag>`

* The supported languages are: `de`, `en`, `es`, `fr`.

### API service

* `us-docker.pkg.dev/rime-labs/api/service:<tag>`

Use the same `YYYYMMDD` tag as the model images.

## Docker Compose configuration

Create a `compose.yml` file that defines the services and their configurations:

```yaml compose.yml theme={null}
version: '3.8'
services:
  api:
    image: us-docker.pkg.dev/rime-labs/api/service:<tag>
    depends_on:
      - model
    ports:
      - "8000:8000"
      - "8002:8002" # binary websockets api
      - "8003:8003" # json websockets api
    restart: unless-stopped
    environment:
      - MODEL_URL=http://model:8080/invocations

  model:
    image: us-docker.pkg.dev/rime-labs/<model>/<version>/<language>:<tag>
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              capabilities: [gpu]
              count: all
    restart: unless-stopped
```

> When running on Kubernetes, ensure that `MODEL_URL` points to `http://0.0.0.0:8080/invocations` instead of the Docker Compose service name.

### Multi-model backend

If you want to serve multiple Arcana languages via a single API instance, you can create a `compose.yml` like the following:

```yaml compose.yml theme={null}
services:
  en-api:
    image: us-docker.pkg.dev/rime-labs/api/service:<tag>
    depends_on:
      - en-model
      - es-model
    ports:
      - "8000:8000"
      - "8001:8001"
      - "8002:8002"
      - "8003:8003"
    restart: unless-stopped
    environment:
      - MODEL_URL=http://en-model:8080/invocations
      - ARCANA_ENG_MODEL_URL=http://en-model:8080/invocations
      - ARCANA_SPA_MODEL_URL=http://es-model:8080/invocations
  en-model:
    image: us-docker.pkg.dev/rime-labs/arcana/v2/en:<tag>
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              capabilities: [gpu]
              count: all
    restart: unless-stopped

  es-model:
    image: us-docker.pkg.dev/rime-labs/arcana/v2/es:<tag>
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              capabilities: [gpu]
              count: all
    restart: unless-stopped
```

Note that the `ARCANA_{LANG}_MODEL_URL` environment variable must point to the container running the Arcana image for that language,
but you should still point `MODEL_URL` to a default model container. The model environment variables currently supported are:

```
ARCANA_ENG_MODEL_URL
ARCANA_SPA_MODEL_URL
ARCANA_FRA_MODEL_URL
ARCANA_GER_MODEL_URL
```

The API will route to these model backends based on the request parameter <a href="https://docs.rime.ai/api-reference/arcana/streaming-mp3#param-lang">lang</a>.

### Authentication configuration

By default, callers must pass their Rime API key in every request via the `Authorization: Bearer <key>` header. Two additional environment variables let you configure authentication at the deployment level instead.

#### Pre-configuring the API key (`RIME_API_KEY`)

If you set `RIME_API_KEY`, the API service will use it to authenticate with the Rime license server automatically, and callers do not need to include an API key in their requests.

You can supply it as an environment variable:

```yaml compose.yml theme={null}
environment:
  - MODEL_URL=http://model:8080/invocations
  - RIME_API_KEY=<your-rime-api-key>
```

Or mount it as a secret file at `/secrets/rime_api_key` inside the container:

```yaml compose.yml theme={null}
services:
  api:
    image: us-docker.pkg.dev/rime-labs/api/service:<tag>
    environment:
      - MODEL_URL=http://model:8080/invocations
    volumes:
      - /run/secrets/rime_api_key:/secrets/rime_api_key:ro
```

If you set neither, callers keep passing their API key in the per-request `Authorization` header.

#### Alternate API key header (`API_KEY_HEADER`)

On platforms that intercept the `Authorization` header, set `API_KEY_HEADER` to the name of an alternate header that callers will use to pass their Rime API key:

```yaml compose.yml theme={null}
environment:
  - MODEL_URL=http://model:8080/invocations
  - API_KEY_HEADER=x-my-platform-rime-api-key
```

Callers then authenticate with:

```bash theme={null}
curl -H "x-my-platform-rime-api-key: <your-rime-api-key>" ...
```

#### Platform API key (`PLATFORM_API_KEY`)

On platforms that require authenticated inter-container requests, set `PLATFORM_API_KEY` so the API service can reach the model backend. You can also mount it as a secret at `/secrets/platform_api_key`:

```yaml compose.yml theme={null}
environment:
  - MODEL_URL=http://model:8080/invocations
  - PLATFORM_API_KEY=<your-platform-api-key>
```

### Start Docker Compose

```bash Start Docker Compose theme={null}
docker compose up -d
```

<Note>Allow approximately five minutes for model warm-up before sending the first synthesis request.</Note>

### Verify health

```bash Health check theme={null}
curl http://localhost:8000/health
```

A ready deployment returns `apiStatus: "ok"`, `licenseStatus: "valid"`, and `modelReachable: true`:

```json theme={null}
{
  "apiStatus": "ok",
  "timestamp": "2026-07-24T18:00:00.000Z",
  "licenseStatus": "valid",
  "modelReachable": true
}
```

## AWS video walkthrough

<iframe src="https://drive.google.com/file/d/1zzrPCVIDsiTMNY_pyb4Z2TezgCTl4Qc1/preview" width="560" height="315" allow="autoplay" allowfullscreen />

# Requests and response formats

## HTTP requests

**Request:**

```bash Request example theme={null}
curl -X POST "http://localhost:8000" \
  -H "Authorization: Bearer <API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
  "text": "I would love to have a conversation with you. The new model is out.",
  "speaker": "joy",
  "modelId": "mist"
}' -o result_mist.txt
```

**Response:**

```json Response format theme={null}
{"audioContent":{"model_output":"<base64>"}}
```

Sample response file: [`result.txt`](https://drive.google.com/file/d/1GW2D8pm5witYMQdKQrPvp_OW5PM0TxNj/view?usp=drive_link)

## Receiving a response in MP3 format

**Request:**

```bash Request example theme={null}
curl -X POST "http://localhost:8000" -H "Authorization: Bearer <API KEY>" -H "Content-Type: application/json" -H "Accept: audio/mpeg" -d '{
  "text": "I would love to have a conversation with you.",
  "speaker": "joy",
  "modelId": "mist"
}' -o result.mp3
```

**Response:**

Sample response file: [`result.mp3`](https://drive.google.com/file/d/1iwmWB1byBXknmNvNmvB_SgBpvnU4bChJ/view?usp=sharing)

### Receiving a response in PCM (raw) format

**Request:**

```bash Request example theme={null}
curl -X POST "http://localhost:8000" -H "Authorization: Bearer <API KEY>" -H "Content-Type: application/json" -H "Accept: audio/L16" -d '{
  "text": "I would love to have a conversation with you.",
  "speaker": "joy",
  "modelId": "mist"
}' -o result.pcm
```

**Response:**

Sample response file: [`result.pcm`](https://drive.google.com/file/d/1pwkGW9jCe1TN9GF5yQa6j619vc8WfjGu/view?usp=drive_link)

### WebSocket endpoints

#### JSON websockets

The API service serves the JSON WebSocket endpoint for `coda`, `arcana`, and `mist` models on port `8003`. For example, `ws://localhost:8003` is equivalent to [Rime's cloud JSON WebSocket API](/docs/websockets).

<Warning>**Spanish is not accepted on the on-prem JSON WebSocket ports.** Sending `lang=spa` to port 8003 (or 8001) returns HTTP 400, for every model, including Coda and Arcana deployments that serve Spanish voices. This restriction does not exist on the cloud endpoint. For Spanish on-prem, use the binary WebSocket endpoint or the HTTP endpoint, which lose word timestamps and context IDs.</Warning>

See the [Coda JSON WebSocket docs](/api-reference/coda/websockets-json), [Arcana JSON WebSocket docs](/api-reference/arcana/websockets-json), and [Mist JSON WebSocket docs](/api-reference/mistv2/websockets-json) depending
on which model backend you have configured.

#### Non-JSON websockets

The API service serves the non-JSON WebSocket endpoint on port `8002`. For example, `ws://localhost:8002` is equivalent to [Rime's cloud WebSocket API](/docs/websockets).

#### Deprecated Mist endpoint (port 8001)

The API service also serves a deprecated WebSocket endpoint on port `8001`. It is compatible only with the Mist model family. Use the [current WebSocket endpoints](/docs/websockets) for new integrations.
