Overview
Tool calling allows the LLM to invoke external functions, such as:- Weather API lookups
- Database queries
- Web searches
- Smart home controls
- Calendar operations
- Custom business logic
Architecture
The tool-calling proxy:- Receives text generation requests from Unmute backend
- Forwards to the LLM with tool definitions
- Detects when the LLM wants to call a tool
- Executes the tool and injects results
- Continues generation with tool results
- Returns final response to Unmute
Implementation Approach
The recommended approach is to create a FastAPI server that:- Exposes an OpenAI-compatible API endpoint
- Wraps your LLM server (VLLM, Ollama, etc.)
- Intercepts tool calls and executes them
- Streams results back to Unmute
Why This Works
Unmute expects streaming text responses from an OpenAI-compatible endpoint. As long as your proxy server provides this interface, Unmute doesn’t need to know about the tool calling happening behind the scenes.Step-by-Step Implementation
1
Create a Tool-Calling Proxy Server
Create a new FastAPI application that wraps your LLM:
2
Add Streaming Support
For real-time voice, streaming is essential:
3
Configure Unmute to Use Your Proxy
Update
docker-compose.yml to point to your proxy server:4
Test Tool Calling
Start a conversation and ask the assistant to use a tool:The tool call happens transparently behind the scenes.
Example Tools
Weather Lookup
Web Search
Database Query
LLM Server Compatibility
Tool calling support varies by LLM server:VLLM
Supports OpenAI-compatible function calling format
--enable-auto-tool-choice and --tool-call-parser:
Ollama
Supports tools parameter in recent versions
OpenAI API
Native support for function calling
Advanced Patterns
Multi-Step Tool Chains
Allow the LLM to call multiple tools in sequence:Conditional Tool Availability
Show different tools based on context:Error Handling
Considerations
Community Contributions
Tool calling support would make a great contribution to the Unmute project! If you build a robust tool-calling proxy server, consider:- Opening a pull request to add it to the main repository
- Documenting your approach for others
- Sharing example tool implementations
Reference Implementation
For a complete working example, check out these resources:- VLLM Tool Calling Docs: docs.vllm.ai/en/latest/features/tool_calling.html
- OpenAI Function Calling Guide: platform.openai.com/docs/guides/function-calling
- FastAPI Streaming: fastapi.tiangolo.com/advanced/custom-response/#streamingresponse
Next Steps
External LLM
Configure Unmute to use different LLM providers
Custom Frontend
Build your own client using the WebSocket protocol