Skip to main content

Overview

Data loaders provide efficient batching, shuffling, and parallel loading of datasets for training and evaluation.

DataLoader Class

Parameters

Dataset
required
Dataset to load data from.
int
default:"1"
Number of samples per batch.
bool
default:"False"
Whether to shuffle the data at the beginning of each epoch.
int
default:"0"
Number of worker processes for parallel data loading. 0 means data will be loaded in the main process.
bool
default:"False"
If True, the data loader will copy tensors into CUDA pinned memory before returning them. Useful for GPU training.
bool
default:"False"
Whether to drop the last incomplete batch if the dataset size is not divisible by the batch size.
Optional[Callable]
Function to merge a list of samples into a batch. If None, uses default collation.

Methods

iter

Return an iterator over the dataset.

len

Return the number of batches.
int
Number of batches in the data loader.

DistributedDataLoader

Data loader for distributed training across multiple devices.

Additional Parameters

int
default:"0"
Rank of the current process in distributed training.
int
default:"1"
Total number of processes in distributed training.

Utility Functions

default_collate

Default collation function that stacks samples into batches.
List[Any]
required
List of samples to collate.
Any
Collated batch.

worker_init_fn

Initialization function for data loader workers.

Example Usage

Performance Tips

num_workers: Use 2-8 workers for optimal performance. Too many workers can cause overhead.
pin_memory: Enable for GPU training to speed up host-to-device transfers.
prefetch: DataLoader automatically prefetches batches in the background for better throughput.
batch_size: Larger batch sizes improve GPU utilization but require more memory. Find the sweet spot for your hardware.

Common Patterns

Training Loop

Multi-GPU Training