Skip to main content

Overview

The Speech-to-Text (STT) component uses Kyutai’s streaming ASR (Automatic Speech Recognition) models to transcribe user speech in real-time with low latency and integrated Voice Activity Detection (VAD). Key Features:
  • Real-time streaming transcription
  • ~2.5 second algorithmic delay (configurable)
  • Integrated pause prediction (VAD)
  • WebSocket-based binary protocol (MessagePack)
  • Word-level timestamps

Architecture

STT Service

Technology: Rust (moshi-server) Location: services/moshi-server/ Model: Kyutai Streaming ASR Deployment: Docker container with GPU access

Service Configuration

Docker Compose (docker-compose.yml:78):
Resource Usage:
  • VRAM: ~2.5 GB
  • Concurrent streams: Limited by capacity management

Python Client

File: unmute/stt/speech_to_text.py

SpeechToText Class

Connection Flow

Startup Sequence

File: speech_to_text.py:130

Sending Audio

File: speech_to_text.py:105
Frame Size: 480 samples (20ms @ 24kHz) Encoding: MessagePack with use_single_float=True for efficiency

Receiving Messages

File: speech_to_text.py:175 The STT client is an async iterator:

Message Types

Client → Server

Audio Message

Example:

Marker Message

Used for latency measurements - server echoes back.

Server → Client

Ready Message

Sent immediately after connection. Signals server is ready.

Word Message

Example:

Step Message

Sent every audio frame (20ms). Contains VAD scores. PRS Array:
  • prs[0]: Probability of speech continuing
  • prs[1]: Probability of speech ending soon
  • prs[2]: Pause prediction score (0-1, higher = more likely pause)
Example:

Error Message

Sent when server cannot accept connection (at capacity).

Voice Activity Detection

Pause Prediction

The backend uses prs[2] from Step messages to detect pauses. File: unmute/unmute_handler.py:372
Threshold: 0.6 (configurable) Smoothing: Exponential moving average

Exponential Moving Average

File: unmute/stt/exponential_moving_average.py
Purpose: Smooth noisy VAD scores to prevent false pause detections.

Interruption Detection

File: unmute/unmute_handler.py:352 Two methods for detecting user interruption:
  1. STT Word: Any word from STT during bot speaking
  2. VAD-based: Pause prediction drops below threshold
Cooldown Period: First 3 seconds (VAD-based only)
  • Prevents echo cancellation issues
  • STT word-based interruption always works

Flushing

File: unmute/unmute_handler.py:340 When a pause is detected, the STT needs to be “flushed” to process remaining audio:
Why Zeros?: The STT has an internal delay buffer (~2.5s). Sending zeros pushes remaining audio through the model. Flush Timing:

Timing & Latency

Time to First Token (TTFT)

File: speech_to_text.py:203
Typical: 50-100ms after first audio frame

Algorithmic Delay

Constant: STT_DELAY_SEC = 2.5 (configurable)
  • Purpose: Look-ahead for better accuracy
  • Trade-off: Higher delay = better accuracy, higher latency
  • Current Time: Tracked via step messages

Metrics

File: unmute/metrics.py

STT-Specific Metrics

Real-Time Factor

Calculated during flush:

Integration with UnmuteHandler

Startup

File: unmute/unmute_handler.py:422

Message Loop

File: unmute/unmute_handler.py:436

Error Handling

Connection Failures

WebSocket Disconnect

Graceful Shutdown

File: speech_to_text.py:157

Testing & Debugging

Dummy STT

File: unmute/stt/dummy_speech_to_text.py For testing without GPU:

Example Script

File: unmute/scripts/stt_from_file_example.py Transcribe audio file:

Performance Tuning

Delay Configuration

Adjust STT_DELAY_SEC in environment:

VAD Threshold

Adjust pause detection sensitivity:

EMA Parameters

Adjust smoothing for VAD scores:

Next Steps