Skip to main content

Hardware Requirements

GPU Requirements

Unmute requires a CUDA-capable NVIDIA GPU. CPU-only deployment is not supported.
VRAM: 16GB minimumExample GPUs:
  • NVIDIA RTX 4090 (24GB)
  • NVIDIA RTX 3090 (24GB)
  • NVIDIA L40S (48GB)
  • NVIDIA A100 (40GB/80GB)
  • NVIDIA RTX A6000 (48GB)
Memory Breakdown:
  • STT: 2.5GB VRAM
  • TTS: 5.3GB VRAM
  • LLM: 6.1GB+ VRAM (model dependent)
  • Overhead: ~2GB for CUDA and buffers
With 16GB VRAM, use Llama 3.2 1B with --gpu-memory-utilization=0.4 and --max-model-len=1536

Architecture Requirements

x86_64 Only

Unmute is built for x86_64 (AMD64) architecture.Not Supported:
  • ARM64 (aarch64) - No support planned
  • Apple Silicon (M1/M2/M3) - No support planned
This is due to dependencies on CUDA and compiled Rust binaries.

System Memory

Recommended RAM: 16GB+ system memoryWhile models run on GPU, the host needs memory for:
  • Docker containers and Python processes
  • Model loading and initialization
  • Audio buffering and WebSocket connections

Software Requirements

Operating System

Docker (Docker Compose)

1

Install Docker

Follow the official Docker installation guide for your platform.Linux (Ubuntu/Debian):
Windows: Install Docker Desktop with WSL 2 backend
2

Install Docker Compose

Docker Compose is included with Docker Desktop. On Linux:
Verify installation:
3

Install NVIDIA Container Toolkit

Required for GPU access from Docker containers:
See NVIDIA’s official guide for other distributions.
4

Verify GPU Access

Test that Docker can access your GPU:

Dockerless Setup (Alternative)

If you prefer to run services without Docker:
This is more complex and requires manual dependency management. Docker Compose is recommended.
Required Tools:
  • uv: Python package manager
  • cargo: Rust toolchain (for STT/TTS servers)
  • pnpm: Node package manager (for frontend)
  • CUDA 12.1: For Rust processes
Start Services:
Access at http://localhost:3000

Configuration Requirements

Hugging Face Access

Model Access Token

Unmute downloads models from Hugging Face Hub:Required:
  1. Hugging Face account
  2. Accept licenses for models you’ll use:
  3. Generate access token with read access
  4. Set environment variable:
Security: Never use tokens with write access in production deployments

Network Requirements

Ports

Docker Compose:
  • Port 80: Traefik (HTTP traffic)
Dockerless:
  • Port 3000: Frontend
  • Port 8000: Backend WebSocket
Optional:
  • Port 9090: Prometheus metrics
  • Port 3001: Grafana dashboards

Bandwidth

Per User:
  • Audio upstream: ~16 KB/s
  • Audio downstream: ~16 KB/s
  • Total: ~32 KB/s bidirectional
For 10 concurrent users:
  • ~320 KB/s (~2.5 Mbps)

Browser Requirements

WebRTC & WebSocket Support

Recommended Browsers:
  • Chrome 90+
  • Firefox 88+
  • Edge 90+
  • Safari 14+ (requires HTTPS)
Required Features:
  • WebSocket support
  • WebRTC (for optional WebRTC mode)
  • Microphone access (requires HTTPS or localhost)
  • Web Audio API
Modern browsers require HTTPS or localhost for microphone access. Use SSH port forwarding for remote access over HTTP.

Model Requirements

Default Models

Model: Kyutai STT 1B (English/French)Specifications:
  • Size: ~2GB download
  • VRAM: 2.5GB
  • Languages: English, French
  • Architecture: Transformer (16 layers, 2048 d_model)
  • Latency: 6-token delay (~200ms)
Configuration (stt.toml):

Performance Targets

Latency

Single GPU (L40S):
  • STT: ~200ms
  • LLM: ~500ms (model dependent)
  • TTS: ~750ms
  • Total: ~1450ms
Multi-GPU:
  • STT: ~200ms
  • LLM: ~500ms
  • TTS: ~450ms
  • Total: ~1150ms

Throughput

Per Backend Instance:
  • Max concurrent users: 4
  • Limited by Python GIL
Scaling Strategy:
  • Run multiple backend replicas
  • Each replica handles 4 users
  • Load balance with Traefik
  • Example: 10 replicas = 40 users

Optional Components

The “Dev (news)” character requires a NewsAPI key:
  1. Sign up at newsapi.org
  2. Get your free API key
  3. Add to environment:
Without this, the news character won’t have current topics.
Optional for Docker Swarm deployments:
  • Used for service registration and health checks
  • Required for multi-node setups
  • Not needed for single-machine Docker Compose
Optional monitoring stack:
  • Prometheus: Metrics collection
  • Grafana: Visualization dashboards
  • Pre-configured for Unmute metrics
  • Included in Docker Swarm setup
See services/prometheus/ and services/grafana/

Deployment Comparison

Choose the deployment method that matches your resources and use case:

Start with Docker Compose

We strongly recommend starting with Docker Compose:
  • Fastest setup (5-10 minutes)
  • Fully supported by Kyutai team
  • Easy to troubleshoot
  • Perfect for learning and development
Switch to other methods only when you need:
  • Fine-grained control (Dockerless)
  • Multi-machine scaling (Docker Swarm)

Ready to Start?

Quick Start Guide

Follow step-by-step instructions to get Unmute running

Join the Community

Star the repo, report issues, and contribute