Skip to main content
Docker Swarm is used for production deployments like unmute.sh. While Docker Compose runs on a single machine, Docker Swarm scales across multiple nodes. Think of it as “multi-node Docker Compose.”
This deployment method is provided to show how Unmute.sh scales in production. Due to the complexity of multi-node debugging, the Kyutai team cannot provide support for Swarm deployments. Use at your own risk.

When to Use Docker Swarm

Choose Docker Swarm when you need:
  • High availability: Multiple replicas ensure uptime during updates or crashes
  • Horizontal scaling: Distribute load across multiple GPUs/machines
  • Production features: HTTPS, monitoring, authentication, load balancing
  • Traffic handling: Support for many concurrent users

Architecture

Main Application

Monitoring Stack

Setup Instructions

All commands should be executed from a client machine with access to your swarm nodes, not directly on the swarm nodes themselves.
1

Prepare GPU Nodes

Set up each GPU node in your swarm:
This script installs Docker, NVIDIA drivers, and configures the node for Swarm.
2

Initialize Swarm Manager

Designate one node as the manager (only needed once):
The manager node coordinates the swarm and must run certain services like Prometheus.
3

Add Worker Nodes

Get the join command from the manager:
This outputs a command like:
Run this command on each worker node to join the swarm.
4

Configure Environment Variables

Set up required environment variables on your client machine:
How to generate tokens:
5

Deploy the Stack

Run the deployment script:
These scripts build images and deploy the stack defined in swarm-deploy.yml.

Production URLs

Once deployed, services are available at: Monitoring services require Google authentication (configured via traefik-forward-auth).

Key Configuration Details

HTTPS with Let’s Encrypt

Traefik automatically obtains and renews SSL certificates:
swarm-deploy.yml

GPU Resource Allocation

Services reserve GPUs using generic resources:
swarm-deploy.yml
With 3 TTS replicas, this requires 3 GPUs total.

Service Replicas

Swarm configuration on unmute.sh:
  • Frontend: 5 replicas (no GPU needed)
  • Backend: 16 replicas (no GPU needed)
  • LLM: 2 replicas (2 GPUs total)
  • TTS: 3 replicas (3 GPUs total)
  • STT: 1 replica (1 GPU)
  • Voice cloning: 2 replicas (no GPU)

Load Balancing

The backend uses manual load balancing for TTS/STT via tasks.<service_name>:
swarm-deploy.yml

Scaling Operations

Adding More Resources

1

Add New Node to Swarm

2

Scale Service

This increases the LLM service to 10 replicas.
Swarm does not automatically rebalance containers across nodes. To redistribute containers after adding nodes, force a service restart.

Restarting a Service

Force update to restart all replicas:
Useful when:
  • New voices are added to voices.yaml
  • Configuration changes need to propagate
  • Rebalancing containers across nodes

Updating a Single Service

Update just the frontend image without touching other services:
Other useful updates:

Monitoring

The swarm deployment includes a complete monitoring stack:

Prometheus

Metrics collection configured via Docker socket:
swarm-deploy.yml
Services expose metrics via the prometheus-port label:
swarm-deploy.yml

Grafana

Dashboards are built into the Docker image:
swarm-deploy.yml
Changes to dashboards in the UI are lost on restart unless exported and added to the build context.

Cadvisor

Collects container metrics on every node:
swarm-deploy.yml

Advanced Configuration

Changing Docker Data Directory

If you need more disk space, change Docker’s data location:
  1. Edit /etc/docker/daemon.json on each node:
  2. Restart Docker:

Network Encryption

The overlay network is encrypted for security:
swarm-deploy.yml
This is important when nodes communicate over the public internet.

Redis for State Management

Unlike single-machine deployments, Swarm uses Redis for shared state:
swarm-deploy.yml

Troubleshooting

Service Won’t Start

Check service logs:
Check service status:

Not Enough Resources

If services can’t find GPUs or have insufficient CPU/memory:
Adjust resource reservations in swarm-deploy.yml.

Network Issues

Use the debugger service to test connectivity:

Differences from Docker Compose

Migration Path

To migrate from Docker Compose to Swarm:
  1. Start with one node: Initialize swarm on your existing machine
  2. Convert compose file: Most syntax is compatible, add deploy sections
  3. Test locally: Deploy to single-node swarm
  4. Add monitoring: Set up Prometheus, Grafana, Traefik
  5. Scale out: Add worker nodes and increase replicas

Next Steps