Skip to main content

Backend Architecture

FastAPI Application

File: unmute/main_websocket.py The main backend server is a FastAPI application that handles:
Concurrency Control (main_websocket.py:71):
Each WebSocket connection acquires the semaphore, limiting concurrent sessions to prevent resource exhaustion.

UnmuteHandler

File: unmute/unmute_handler.py The core orchestration class that manages a single conversation session. Key Responsibilities:
  • Receive audio from frontend
  • Route audio to STT
  • Manage conversation state
  • Generate LLM responses
  • Stream TTS audio back to frontend
  • Handle interruptions
Architecture:
Service Management: Services (STT, TTS, LLM) are managed as “Quests” for clean lifecycle:

WebSocket Protocol Handler

Files:
  • unmute/main_websocket.py:380-404 - Main route handler
  • unmute/main_websocket.py:406-492 - Receive loop
  • unmute/main_websocket.py:512-582 - Emit loop
Two-Loop Architecture:
  1. Receive Loop - Handles incoming messages:
  2. Emit Loop - Sends messages to frontend:

Quest Manager

File: unmute/quest_manager.py Manages lifecycle of background services with clean cancellation.
Features:
  • Async context manager for automatic cleanup
  • Service initialization with retries
  • Graceful shutdown on errors
  • Quest removal (for interruptions)
Example Usage:

Chatbot State Manager

File: unmute/llm/chatbot.py Manages conversation history and state transitions.
Message Format:

Service Discovery

File: unmute/service_discovery.py Finds available service instances with capacity.
Capacity Handling:
  • Services can reject with Error message when at capacity
  • Backend tries next instance
  • Raises MissingServiceAtCapacity if all exhausted

Metrics Collection

File: unmute/metrics.py Prometheus metrics using prometheus_client. Counter Metrics:
Gauge Metrics:
Histogram Metrics:
Integration:

Audio Processing

Opus Codec:
Audio Format Conversion:

Voice Management

File: unmute/tts/voices.py Loads and validates voices from voices.yaml.
Voice Sources:
  1. File: Pre-recorded audio on server
  2. Freesound: Creative Commons audio from Freesound.org
  3. Custom: User-uploaded voice cloning
API Endpoint (main_websocket.py:200):

Voice Cloning

File: unmute/tts/voice_cloning.py Generates voice embeddings from uploaded audio.
Upload Endpoint (main_websocket.py:240):

Recording System

File: unmute/recorder.py Optional conversation recording for debugging/analysis.
Privacy:
  • Audio data anonymized (only sample counts, not PCM)
  • User can opt-out via allow_recording: false
  • Recordings stored in RECORDINGS_DIR (configurable)

Timer Utilities

File: unmute/timer.py Stopwatch for accurate timing measurements.
Usage:

Exception Handling

File: unmute/exceptions.py Custom exceptions for service failures.
Error Reporting (main_websocket.py:334):

CORS Configuration

File: unmute/main_websocket.py:84 CORS middleware for local development.

Middleware

Upload Size Limiting

File: unmute/main_websocket.py:213 Limits voice upload file size.

Prometheus Instrumentation

File: unmute/main_websocket.py:74 Automatic HTTP request metrics.

Configuration

File: unmute/kyutai_constants.py Centralized configuration from environment variables.

Next Steps