Skip to main content

Overview

The llm_utils module provides utilities for interacting with Large Language Models via the OpenAI-compatible API, preprocessing conversation history, and streaming text responses.

Constants

INTERRUPTION_CHAR

Character appended to assistant messages when the bot is interrupted by the user.

USER_SILENCE_MARKER

Marker inserted into user messages when they remain silent for an extended period.

Classes

VLLMStream

Streaming LLM client for chat completions.
AsyncOpenAI
required
AsyncOpenAI client instance
float
default:"1.0"
Sampling temperature (0.0 to 2.0). Lower values are more deterministic.

Methods

chat_completion
Generates streaming chat completion.
list[dict[str, str]]
required
Conversation history in OpenAI format. Each dict should have “role” and “content” keys.
Returns: AsyncIterator[str] - Stream of text chunks Example:

Functions

get_openai_client

Creates an AsyncOpenAI client instance.
str
default:"LLM_SERVER"
Base URL of the LLM server
str | None
default:"KYUTAI_LLM_API_KEY"
API key for authentication. Defaults to “EMPTY” for vLLM servers that don’t require keys.
Returns: AsyncOpenAI client Example:

autoselect_model

Automatically selects an LLM model from the server. Returns: str - Model identifier Raises: ValueError if multiple models are available (requires explicit selection) Notes:
  • Uses KYUTAI_LLM_MODEL environment variable if set
  • Otherwise queries the server and selects the model if only one is available
  • Result is cached for performance

preprocess_messages_for_llm

Preprocesses conversation history before sending to the LLM.
list[dict[str, str]]
required
Raw conversation history with “role” and “content” keys
Returns: list[dict[str, str]] - Cleaned conversation history Processing steps:
  1. Removes messages containing only the INTERRUPTION_CHAR
  2. Strips INTERRUPTION_CHAR suffix from interrupted messages
  3. Merges consecutive messages from the same role
  4. Adds dummy “Hello.” user message if needed for model compatibility
  5. Removes USER_SILENCE_MARKER prefix when user continues talking
Example:

rechunk_to_words

Rechunks a text stream into whole words for better TTS pronunciation.
AsyncIterator[str]
required
Stream of text chunks (may break mid-word)
Returns: AsyncIterator[str] - Stream of complete words Behavior:
  • Spaces are included with the following word: "foo bar baz" → "foo", " bar", " baz"
  • Multiple whitespace characters are merged into a single space
  • Buffers partial words until whitespace is encountered
Example:

Protocol

LLMStream

Protocol for LLM streaming clients. Any class implementing chat_completion() can be used as an LLM stream.

Complete Example

Advanced Usage: Integration with TTS

Environment Variables

  • KYUTAI_LLM_MODEL: Model identifier to use (if not set, auto-selects)
  • KYUTAI_LLM_API_KEY: API key for LLM server
  • LLM_SERVER: Base URL of the LLM server

Notes

  • The VLLMStream class auto-selects the model if not explicitly configured
  • Message preprocessing handles common conversation artifacts (interruptions, silence markers)
  • Word rechunking is essential for natural TTS pronunciation
  • All async functions should be run within an event loop
  • The client supports any OpenAI-compatible API (vLLM, llama.cpp, etc.)