Overview
Data loaders provide efficient batching, shuffling, and parallel loading of datasets for training and evaluation.DataLoader Class
Parameters
Dataset
required
Dataset to load data from.
int
default:"1"
Number of samples per batch.
bool
default:"False"
Whether to shuffle the data at the beginning of each epoch.
int
default:"0"
Number of worker processes for parallel data loading. 0 means data will be loaded in the main process.
bool
default:"False"
If True, the data loader will copy tensors into CUDA pinned memory before returning them. Useful for GPU training.
bool
default:"False"
Whether to drop the last incomplete batch if the dataset size is not divisible by the batch size.
Optional[Callable]
Function to merge a list of samples into a batch. If None, uses default collation.
Methods
iter
len
int
Number of batches in the data loader.
DistributedDataLoader
Additional Parameters
int
default:"0"
Rank of the current process in distributed training.
int
default:"1"
Total number of processes in distributed training.
Utility Functions
default_collate
List[Any]
required
List of samples to collate.
Any
Collated batch.