Skip to main content
Unmute requires GPU acceleration to run the speech-to-text, text-to-speech, and LLM models with acceptable latency. This guide covers GPU requirements and configuration.

Hardware Requirements

Minimum Requirements

  • GPU: NVIDIA GPU with CUDA support
  • VRAM: At least 16 GB
  • Architecture: x86_64 (aarch64 is not supported)

Memory Usage by Service

When running all services, approximate VRAM usage:
The default docker-compose.yml uses Llama 3.2 1B which fits in 16GB VRAM. If using larger models like Mistral Small 3.2 24B, you’ll need more VRAM.

Operating System Support

Windows: Native Windows is not supported (#84). Use WSL (Windows Subsystem for Linux) instead.macOS: Not supported (#74). macOS does not have NVIDIA GPU support.
Supported platforms:
  • Linux (native)
  • Windows with WSL 2

Docker Setup

Install NVIDIA Container Toolkit

The NVIDIA Container Toolkit allows Docker containers to access your GPU.
1

Install the Container Toolkit

Follow the official NVIDIA Container Toolkit installation guide.For Ubuntu/Debian:
2

Verify the Installation

Test that Docker can access your GPU:
You should see output showing your GPU(s), similar to:

Configure GPU Access in Docker Compose

The docker-compose.yml file configures GPU access for the AI services:

Multi-GPU Configuration

Running services on separate GPUs significantly improves latency. On unmute.sh, TTS latency decreases from ~750ms (single L40S GPU) to ~450ms (multi-GPU setup).

Single GPU Setup (Default)

By default, all services share GPU(s) using count: all:

Dedicated GPU per Service

If you have 3+ GPUs, assign one GPU to each service for optimal performance:
1

Check Available GPUs

List your GPUs:
Example output:
2

Update docker-compose.yml

Modify the stt, tts, and llm services to use dedicated GPUs:
3

Restart Services

Apply the changes:
Docker will automatically distribute services across available GPUs when using count: 1. You don’t need to manually specify device IDs.

Memory Optimization

If you’re running out of GPU memory, adjust these settings in docker-compose.yml:

LLM Memory Settings

integer
default:"1536"
Maximum context length for the LLM. Lower values use less memory but support shorter conversations.
float
default:"0.4"
Percentage of GPU memory to allocate (0.0-1.0). Lower values leave more memory for other services.

Switch to a Smaller Model

Use a smaller LLM model:
Available models (by size):
  • meta-llama/Llama-3.2-1B-Instruct - ~6 GB VRAM
  • google/gemma-3-1b-it - ~6 GB VRAM (note: slower on vLLM)
  • google/gemma-3-12b-it - ~12 GB VRAM
  • mistralai/Mistral-Small-3.2-24B-Instruct-2506 - ~24 GB VRAM

Dockerless Setup

For dockerless deployment, ensure CUDA 12.1+ is installed:
1

Install CUDA

Install CUDA 12.1 or later:
  • Via conda: conda install cuda -c nvidia/label/cuda-12.1.0
  • Or download from NVIDIA’s website
2

Verify Installation

3

Run Services

The dockerless scripts automatically detect and use available GPUs:

Troubleshooting

If nvidia-smi works but Docker can’t access the GPU:
  1. Verify NVIDIA Container Toolkit is installed
  2. Restart Docker: sudo systemctl restart docker
  3. Check Docker runtime: docker info | grep -i runtime
  4. Try the verification command again:
If services crash with OOM errors:
  1. Check VRAM usage: nvidia-smi
  2. Reduce --gpu-memory-utilization for the LLM
  3. Lower --max-model-len for shorter conversations
  4. Use a smaller LLM model
  5. Stop other GPU-intensive applications
For Windows WSL users:
  1. Ensure you’re using WSL 2 (not WSL 1)
  2. Update to the latest NVIDIA driver for Windows
  3. Install NVIDIA CUDA on WSL following Microsoft’s guide
  4. Don’t install NVIDIA drivers inside WSL - use the Windows driver