Skip to main content

Overview

While Unmute uses WebSockets for message transport, it leverages Web Audio API and Opus encoding for efficient real-time audio processing. The system handles:
  • Microphone input capture and encoding
  • Real-time audio streaming with low latency
  • Audio decoding and playback
  • Voice activity detection (VAD)
  • Audio buffering and synchronization
Unmute does not use WebRTC peer-to-peer connections. Instead, it uses WebSocket for transport with Opus audio encoding, which provides similar efficiency for the client-server architecture.

Audio Pipeline Architecture

Frontend Audio Processing

Audio Processor Setup

The useAudioProcessor hook (frontend/src/app/useAudioProcessor.ts) manages the complete audio pipeline:

Microphone Input Processing

Configuration (useAudioProcessor.ts:83-104):
Key Settings:
boolean
Enabled to prevent feedback from speakers
boolean
Disabled - handled by backend processing
boolean
Enabled for consistent audio levels
number
24kHz - balances quality and bandwidth
number
20ms frames for low latency

Opus Encoding

The frontend uses the opus-recorder library to encode microphone input:
The encoded Opus data is then base64-encoded and sent via WebSocket:

Audio Output Processing

Decoder Setup (useAudioProcessor.ts:56-77):
Audio Worklet Processing: The audio-output-processor worklet handles the actual audio playback, buffering incoming frames and outputting them at the correct rate.

Audio Analysis for Visualization

Both input and output audio streams are analyzed for visualization:
These analyzers provide frequency domain data used by the circular audio visualizers in the UI.

Backend Audio Processing

Opus Decoding

The backend uses the sphn library for Opus stream processing (unmute/main_websocket.py:415-477):
First Packet Detection: The backend waits for the first Opus packet by checking bit 2 of byte 5 in the Opus stream. This ensures proper stream synchronization and prevents processing stale data from previous connections.

Opus Encoding

For outgoing audio, the backend encodes PCM audio to Opus (unmute/main_websocket.py:520-558):

Audio Buffering and Synchronization

Backend Buffering

The TTS system manages audio buffering to prevent stuttering (unmute/tts/text_to_speech.py:88-94):
number
Buffer size in seconds to prevent stuttering while maintaining low latency

Frontend Buffering

The audio output processor worklet handles buffering on the client side, ensuring smooth playback even with network jitter.

Audio Format Specifications

Input Audio (Microphone)

string
Opus
string
24kHz
string
Mono (1 channel)
string
20ms
string
Adaptive (Opus encoder automatic)
string
Base64-encoded over WebSocket

Output Audio (TTS)

string
Opus
string
24kHz
string
Mono (1 channel)
string
Variable (based on TTS output)
string
Base64-encoded over WebSocket

Cloudflare TURN Configuration

Although Unmute doesn’t use WebRTC peer connections, it includes utilities for obtaining TURN server credentials from Cloudflare (unmute/webrtc_utils.py):
This utility is available for future use if Unmute moves to a peer-to-peer WebRTC architecture.

Performance Considerations

Latency Optimization

  1. Small Frame Sizes: 20ms frames minimize encoding latency
  2. Streaming Mode: streamPages: true sends data immediately without waiting for complete pages
  3. Low Complexity: encoderComplexity: 0 trades some quality for lower CPU usage and latency
  4. Minimal Buffering: AUDIO_BUFFER_SEC = FRAME_TIME_SEC * 4 keeps buffer small

Bandwidth Optimization

  1. 24kHz Sample Rate: Lower than 48kHz but sufficient for voice
  2. Mono Audio: Single channel reduces bandwidth by 50%
  3. Opus Codec: Highly efficient compression for speech
  4. Adaptive Bitrate: Opus automatically adjusts based on audio characteristics

CPU Optimization

  1. Web Workers: Encoding/decoding runs in separate threads
  2. Audio Worklets: Audio processing runs on high-priority audio thread
  3. Async Processing: Backend uses asyncio.to_thread for CPU-intensive operations

Debugging Audio Issues

Enable Developer Mode

Press D in the frontend to enable developer mode, which shows:
  • Debug dictionary with internal state
  • Additional logging in the console

Check Audio Levels

The circular visualizers show audio activity:
  • User circle (right): Should pulse when speaking
  • Assistant circle (left): Should pulse during TTS output

Common Issues

  • Check microphone permissions
  • Verify microphone is not muted in system settings
  • Check browser DevTools console for errors
  • Ensure echoCancellation is properly configured
  • Network issues may be causing packet loss
  • Backend TTS may be slower than real-time
  • Increase AUDIO_BUFFER_SEC value
  • Check CPU usage on backend
  • This can occur when TTS is slower than real-time
  • Buffering in the audio pipeline causes delayed playback
  • Adjust AUDIO_BUFFER_SEC to balance latency vs stability
  • Ensure echoCancellation: true is set
  • Use headphones to prevent speaker feedback
  • Lower speaker volume