Skip to main content

Overview

Server events are messages sent from the Unmute backend to the client. These events stream audio responses, transcriptions, and status updates.

session.updated

Confirms that the session configuration was successfully updated.

Response Fields

string
required
Always "session.updated"
string
required
Unique event identifier
object
required
The updated session configuration (mirrors the session.update request)

Example


response.created

Indicates that the assistant has started generating a response.

Response Fields

string
required
Always "response.created"
string
required
Unique event identifier
object
required
Response metadata
string
required
Always "realtime.response"
string
required
Response status. One of: "in_progress", "completed", "cancelled", "failed", "incomplete"
string
required
Voice identifier being used for this response
array
default:"[]"
Conversation history (array of message objects)

Example


response.audio.delta

Streams generated speech audio to the client.

Response Fields

string
required
Always "response.audio.delta"
string
required
Unique event identifier
string
required
Base64-encoded Opus audio chunkAudio Specifications:
  • Codec: Opus
  • Sample Rate: 24 kHz
  • Channels: Mono
  • Encoding: Base64 string

Example

Implementation Notes

  • Audio chunks are sent as they become available from the text-to-speech system
  • Due to Opus buffering, not every PCM chunk results in output
  • Chunks should be decoded and played in sequence

JavaScript Example


response.audio.done

Indicates that audio streaming for the current response has completed.

Response Fields

string
required
Always "response.audio.done"
string
required
Unique event identifier

Example


response.text.delta

Streams the text being generated (for display or debugging).

Response Fields

string
required
Always "response.text.delta"
string
required
Unique event identifier
string
required
Text chunk being generated

Example


response.text.done

Indicates that text generation is complete and provides the full text.

Response Fields

string
required
Always "response.text.done"
string
required
Unique event identifier
string
required
Complete generated text

Example


conversation.item.input_audio_transcription.delta

Streams real-time transcription of user speech.

Response Fields

string
required
Always "conversation.item.input_audio_transcription.delta"
string
required
Unique event identifier
string
required
Transcription text chunk
number
required
Timestamp when speech started (Unmute extension)

Example


input_audio_buffer.speech_started

Indicates that speech was detected in the user’s audio input. Note: Based on speech-to-text detection, not voice activity detection (VAD). This ensures the event is only sent when actual speech is transcribed.

Response Fields

string
required
Always "input_audio_buffer.speech_started"
string
required
Unique event identifier

Example


input_audio_buffer.speech_stopped

Indicates that a pause was detected in the user’s audio input. Note: Based on voice activity detection (VAD).

Response Fields

string
required
Always "input_audio_buffer.speech_stopped"
string
required
Unique event identifier

Example


unmute.interrupted_by_vad

Indicates that the voice activity detector interrupted the assistant’s response generation because the user started speaking. Unmute Extension: This event is specific to Unmute.

Response Fields

string
required
Always "unmute.interrupted_by_vad"
string
required
Unique event identifier

Example


unmute.response.text.delta.ready

Indicates that a text delta is ready for processing. Unmute Extension: This event is specific to Unmute.

Response Fields

string
required
Always "unmute.response.text.delta.ready"
string
required
Unique event identifier
string
required
Text chunk that is ready

Example


unmute.response.audio.delta.ready

Indicates that an audio delta is ready with sample count information. Unmute Extension: This event is specific to Unmute.

Response Fields

string
required
Always "unmute.response.audio.delta.ready"
string
required
Unique event identifier
integer
required
Number of audio samples in this chunk

Example


unmute.additional_outputs

Provides additional debug or metadata outputs from the system. Unmute Extension: This event is specific to Unmute and used for debugging.

Response Fields

string
required
Always "unmute.additional_outputs"
string
required
Unique event identifier
any
required
Additional output data (structure varies)

Example


error

Reports errors during the WebSocket session.

Response Fields

string
required
Always "error"
string
required
Unique event identifier
object
required
Error details
string
required
Error type (e.g., "invalid_request_error", "fatal")
string
Error code (optional)
string
required
Human-readable error message
string
Parameter that caused the error (optional)
any
Additional error details (Unmute extension, optional)

Example: Invalid JSON

Example: Fatal Error

Note: Fatal errors typically result in the WebSocket connection being closed by the server.

Next Steps

Client Events

Events sent from client to server

Session Management

Configure voice and conversation settings