Overview
Thellm_utils module provides utilities for interacting with Large Language Models via the OpenAI-compatible API, preprocessing conversation history, and streaming text responses.
Constants
INTERRUPTION_CHAR
USER_SILENCE_MARKER
Classes
VLLMStream
AsyncOpenAI
required
AsyncOpenAI client instance
float
default:"1.0"
Sampling temperature (0.0 to 2.0). Lower values are more deterministic.
Methods
chat_completion
list[dict[str, str]]
required
Conversation history in OpenAI format. Each dict should have “role” and “content” keys.
AsyncIterator[str] - Stream of text chunks
Example:
Functions
get_openai_client
str
default:"LLM_SERVER"
Base URL of the LLM server
str | None
default:"KYUTAI_LLM_API_KEY"
API key for authentication. Defaults to “EMPTY” for vLLM servers that don’t require keys.
AsyncOpenAI client
Example:
autoselect_model
str - Model identifier
Raises: ValueError if multiple models are available (requires explicit selection)
Notes:
- Uses
KYUTAI_LLM_MODELenvironment variable if set - Otherwise queries the server and selects the model if only one is available
- Result is cached for performance
preprocess_messages_for_llm
list[dict[str, str]]
required
Raw conversation history with “role” and “content” keys
list[dict[str, str]] - Cleaned conversation history
Processing steps:
- Removes messages containing only the
INTERRUPTION_CHAR - Strips
INTERRUPTION_CHARsuffix from interrupted messages - Merges consecutive messages from the same role
- Adds dummy “Hello.” user message if needed for model compatibility
- Removes
USER_SILENCE_MARKERprefix when user continues talking
rechunk_to_words
AsyncIterator[str]
required
Stream of text chunks (may break mid-word)
AsyncIterator[str] - Stream of complete words
Behavior:
- Spaces are included with the following word:
"foo bar baz"→"foo"," bar"," baz" - Multiple whitespace characters are merged into a single space
- Buffers partial words until whitespace is encountered
Protocol
LLMStream
chat_completion() can be used as an LLM stream.
Complete Example
Advanced Usage: Integration with TTS
Environment Variables
KYUTAI_LLM_MODEL: Model identifier to use (if not set, auto-selects)KYUTAI_LLM_API_KEY: API key for LLM serverLLM_SERVER: Base URL of the LLM server
Notes
- The
VLLMStreamclass auto-selects the model if not explicitly configured - Message preprocessing handles common conversation artifacts (interruptions, silence markers)
- Word rechunking is essential for natural TTS pronunciation
- All async functions should be run within an event loop
- The client supports any OpenAI-compatible API (vLLM, llama.cpp, etc.)