Skip to main content

Overview

Neurenix provides comprehensive support for distributed training across multiple GPUs and compute nodes. The framework integrates with industry-standard distributed training backends including MPI, Horovod, and DeepSpeed.

MPI Backend

The Message Passing Interface (MPI) backend provides low-level distributed computing primitives for parallel training.

MPIManager

The MPIManager class provides an interface to MPI functionality for distributed training.
Reference: neurenix/distributed/mpi.py:15

Collective Operations

Barrier Synchronization

Reference: neurenix/distributed/mpi.py:135

Broadcast

Reference: neurenix/distributed/mpi.py:148

All-Reduce

Reference: neurenix/distributed/mpi.py:168

All-Gather

Reference: neurenix/distributed/mpi.py:188

Scatter

Reference: neurenix/distributed/mpi.py:207

Context Manager

Reference: neurenix/distributed/mpi.py:227

Horovod Backend

Horovod provides a unified API for distributed training with automatic gradient aggregation.

HorovodManager

Reference: neurenix/distributed/horovod.py:15

Broadcasting Parameters

Reference: neurenix/distributed/horovod.py:207

Distributed Optimizer

Reference: neurenix/distributed/horovod.py:224

Complete Training Example

DeepSpeed Backend

DeepSpeed provides advanced optimization techniques including ZeRO optimization, pipeline parallelism, and mixed precision training.

DeepSpeedManager

Reference: neurenix/distributed/deepspeed.py:15

Initialize Model with DeepSpeed

Reference: neurenix/distributed/deepspeed.py:152

ZeRO Optimization Stages

ZeRO Stage 1: Optimizer state partitioning
  • Partitions optimizer states across GPUs
  • Reduces memory by ~4x for Adam
ZeRO Stage 2: Gradient partitioning
  • Partitions gradients in addition to optimizer states
  • Further reduces memory usage
ZeRO Stage 3: Parameter partitioning
  • Partitions model parameters across GPUs
  • Enables training very large models

Training with DeepSpeed

Data Parallel Training

For simple multi-GPU training on a single node:
Reference: neurenix/distributed/__init__.py:9

Best Practices

Device Placement

Learning Rate Scaling

Gradient Accumulation

Checkpointing

Environment Variables

Common environment variables for distributed training:

Launching Distributed Training

MPI Launch

Horovod Launch

DeepSpeed Launch

Performance Tips

  1. Use NCCL backend for GPU communication (fastest)
  2. Enable tensor fusion to reduce communication overhead
  3. Use gradient compression (FP16) to reduce bandwidth
  4. Batch data loading to maximize GPU utilization
  5. Profile communication to identify bottlenecks
  6. Use InfiniBand for multi-node training when available