Skip to main content

Overview

Unmute includes several debugging tools to help diagnose issues with audio quality, latency, service connectivity, and system behavior. This guide covers both development and production debugging techniques.

Development Mode

Enabling Dev Mode

Unmute includes a hidden debug view for development. Enable it by modifying frontend/src/app/useKeyboardShortcuts.ts:
Restart the frontend, then press D to toggle the debug view.

Debug Information Available

The debug view displays real-time information from self.debug_dict in unmute_handler.py. You can add custom debug data:

Subtitles Mode

Press S to enable subtitles showing:
  • User speech transcription (from STT)
  • Bot responses (text and audio)
Useful for verifying STT accuracy and response content.

Logging

Backend Logging

Unmute uses Python’s standard logging with structured output:
View logs:

Log Configuration

From loadtest_client.py example:
Change level=logging.DEBUG for verbose output.

Service-Specific Logs

TTS and STT services write logs to dedicated volumes:
Access logs on the host machine:

Recording Sessions

Enabling Recordings

Set the KYUTAI_RECORDINGS_DIR environment variable:
Unmute records:
  • User audio input
  • Bot audio output
  • Session metadata
Access recordings:

Recording Format

Recordings use the Recorder class (from unmute_handler.py):
Audio is saved in WAV format at 24kHz sample rate.

Load Testing

Running Load Tests

The loadtest_client.py script simulates realistic user conversations:
Parameters:
  • --n-workers: Parallel connections (simulates concurrent users)
  • --n-conversations: Total conversations to run
  • --audio-dir: Directory with test audio files (MP3 format)
  • --listen: Play received audio for manual verification

Interpreting Results

Load test output includes detailed timing:
Key metrics:
  • Mean/Median: Average performance
  • p90/p95: Tail latencies (worst-case scenarios)
  • Realtime factor: Less than 1.0 means TTS generates faster than playback (good)
  • OK fraction: Success rate (should be >0.95)

Debugging with Load Tests

This plays back audio and shows detailed logs for manual inspection.

Health Checks

Backend Health Endpoint

Healthy response:
Unhealthy response:

Service Discovery

Unmute uses service discovery to find available STT/TTS instances (from service_discovery.py):
Debug service discovery:

Common Issues

Issue: STT Not Transcribing

Symptoms: No subtitles appear, silence timeout triggers Debug steps:
  1. Check STT service logs: docker logs <stt_container>
  2. Verify microphone permissions in browser
  3. Enable subtitles (press S) to see if any text appears
  4. Check STT metrics: worker_stt_recv_words_total (should increase)
  5. Test with known audio file using load test
Common causes:
  • Microphone not connected/permitted
  • Echo cancellation consuming speech
  • STT service out of memory

Issue: High Latency

Symptoms: Delayed responses, choppy audio Debug steps:
  1. Check Grafana dashboards for latency spikes
  2. Monitor GPU utilization: nvidia-smi -l 1
  3. Check service metrics:
  4. Run load test to isolate bottleneck
  5. Review self.debug_dict in dev mode
Common causes:
  • GPU shared between services (use multi-GPU setup)
  • High context length (--max-model-len too large)
  • Network latency between services
  • Insufficient GPU memory causing swapping

Issue: TTS Audio Choppy

Symptoms: Audio stutters or drops frames Debug steps:
  1. Check output frame size in unmute_handler.py:
  2. Monitor TTS realtime factor:
    Should be less than 1.0 (faster than realtime)
  3. Check TTS service logs for errors
  4. Verify GPU not overloaded
Common causes:
  • Output frame size too large
  • TTS generation slower than realtime
  • Network congestion
  • CPU throttling

Issue: Service Connection Failures

Symptoms: worker_stt_misses or worker_tts_misses increasing Debug steps:
  1. Check service health:
  2. Test connectivity from backend:
  3. Check service discovery:
  4. Review service logs for crashes
Common causes:
  • Service crashed and restarting
  • GPU out of memory
  • Network misconfiguration
  • Too many concurrent requests

Issue: LLM Timeouts

Symptoms: worker_vllm_hard_errors increasing, responses cut off Debug steps:
  1. Check LLM service logs:
  2. Monitor GPU memory:
  3. Check request context length:
  4. Review LLM configuration in docker-compose.yml
Common causes:
  • Context window too large for GPU memory
  • --gpu-memory-utilization too high
  • Long conversation history exceeding --max-model-len
  • Model loading failure

Debugging Tools

1. Audio Debugging

From loadtest_client.py:
Use this to verify audio quality at each pipeline stage.

2. Timing Analysis

Unmute uses PhasesStopwatch for detailed timing:
Add custom stopwatches to measure specific operations.

3. WebSocket Message Inspection

Log all WebSocket messages:
Inspect message payloads to debug protocol issues.

4. Metrics Endpoint

Query Prometheus metrics directly:

Debug Environment Variables

From unmute_handler.py:
Uncomment and modify these in the source code for local debugging.

Debugging Production

Docker Swarm Debugging

Scaling for Debugging

Force Service Restart

Additional Resources

Getting Help

If you encounter issues:
  1. Check existing GitHub issues
  2. Review logs and metrics before reporting
  3. Include reproducible steps and error messages
  4. Share relevant configuration (docker-compose.yml, etc.)
Note: From the README: “If something isn’t working for you, don’t hesitate to open an issue. We’ll do our best to help you figure out what’s wrong.”