Skip to main content
Unmute’s backend uses a WebSocket protocol based on the OpenAI Realtime API, making it possible to build custom frontends or integrate Unmute into your own applications.

Protocol Overview

The Unmute backend communicates over WebSocket using a JSON-based event protocol. The protocol handles:
  • Real-time bidirectional audio streaming
  • Speech transcription events
  • Session configuration
  • Response generation status
  • Error handling

WebSocket Connection

Endpoint Details

string
required
/v1/realtime
string
required
realtime (WebSocket subprotocol)
number
8000 (development), 80 (production via Traefik)

Establishing a Connection

Connect to the Unmute backend using the WebSocket API with the realtime subprotocol:

Message Structure

All messages follow a common event structure defined in unmute/openai_realtime_api_events.py:

Client to Server Events

Messages your frontend sends to the backend.

Session Configuration

Required: Send this before the backend will start processing audio.
object
required
Defines the character’s conversation behavior. Can be:
  • {"type": "smalltalk", "language": "en"} - General conversation
  • {"type": "constant", "text": "Custom instructions"} - Custom personality
  • {"type": "quiz_show"} - Quiz game mode
  • {"type": "news"} - Tech news discussion
  • {"type": "guess_animal"} - Guessing game
  • {"type": "unmute_explanation"} - Unmute Q&A
string
required
Path to the voice file on the server (e.g., from voices.yaml)
boolean
required
Whether to allow conversation recording

Audio Input Streaming

Send user microphone audio to the backend:
Audio Format Requirements:
  • Codec: Opus
  • Sample Rate: 24kHz
  • Channels: Mono
  • Encoding: Base64-encoded bytes

Example: Capturing and Sending Audio

Server to Client Events

Messages the backend sends to your frontend.

Session Updated

Confirms session configuration was applied:

Response Created

Indicates the assistant has started generating a response:

Audio Response Streaming

Receive generated speech audio:
Audio Format: Same as input (Opus, 24kHz, mono, base64-encoded)

Example: Playing Audio Response

Audio Response Complete

Text Response Streaming

Receive the text being generated (useful for subtitles/debugging):

Text Response Complete

Transcription Streaming

Real-time transcription of user speech:
The start_time field is an Unmute extension not present in the standard OpenAI Realtime API.

Speech Detection Events

VAD Interruption

Indicates the user interrupted the assistant’s response:
This is an Unmute-specific event not in the OpenAI Realtime API.

Error Events

Connection Lifecycle

1

Health Check (Optional)

Before establishing WebSocket, verify backend is running:
2

Establish WebSocket Connection

Connect with the realtime subprotocol:
3

Configure Session

Send session.update with character and voice settings. The backend will not process audio until this is sent.
4

Stream Audio

Begin sending microphone audio via input_audio_buffer.append events.
5

Handle Responses

Process incoming audio, text, and transcription events from the backend.
6

Graceful Shutdown

Close the WebSocket connection when done:

Reference Implementation

Unmute includes reference client implementations you can study:

Next.js Frontend

The official frontend implementation:
  • Location: frontend/src/app/Unmute.tsx
  • Framework: React with Next.js
  • Features: Full WebSocket handling, audio recording, playback, UI

Python Load Test Client

A simpler client for testing and benchmarking:
  • Location: unmute/loadtest/loadtest_client.py
  • Use case: Automated testing, latency measurement
  • Language: Python with asyncio

OpenAI Realtime API Compatibility

Unmute’s protocol is inspired by the OpenAI Realtime API but includes some differences:

Unmute Extensions

These event types are specific to Unmute:
  • unmute.interrupted_by_vad
  • unmute.response.text.delta.ready
  • unmute.response.audio.delta.ready
  • unmute.additional_outputs
  • unmute.input_audio_buffer.append_anonymized

Simplified Parameters

Some OpenAI parameters are simplified or omitted for Unmute’s specific use case. See unmute/openai_realtime_api_events.py for the complete event schema.

Future Compatibility

The goal is to make Unmute fully compatible with the OpenAI Realtime API so frontends can work with both backends interchangeably. Contributions to improve compatibility are welcome!

Example: Minimal Custom Client

Debugging Tips

Enable subtitles in the official frontend by pressing S to see real-time transcription and text responses.
Enable dev mode by setting ALLOW_DEV_MODE = true in frontend/src/hooks/useKeyboardShortcuts.ts, then press D to see detailed debug information.
Check backend logs for detailed WebSocket event information:

Further Reading

  • Protocol Documentation: docs/browser_backend_communication.md in the source repository
  • Event Definitions: unmute/openai_realtime_api_events.py
  • OpenAI Realtime API: platform.openai.com/docs/guides/realtime