Skip to main content

Overview

The DatasetHub module provides utilities for loading datasets from URLs or file paths, supporting various formats and preprocessing options.

DatasetFormat Enum

Methods

from_extension

Determine format from file extension.
str
required
File extension (e.g., ‘csv’, ‘json’, ‘npy’).
DatasetFormat
The corresponding dataset format.

Dataset Class

Parameters

Any
required
The dataset content.
DatasetFormat
required
Format of the dataset.
str
Name of the dataset.
Dict
Additional information about the dataset.
Callable
Function to transform data samples.

Methods

len

Return the number of samples in the dataset.

getitem

Get a sample or batch from the dataset.
Union[int, slice]
required
Index or slice to retrieve.
Any
The sample or batch at the specified index.

to_tensor

Convert the dataset to a tensor.
str
default:"auto"
The framework to use (‘torch’, ‘tensorflow’, ‘numpy’, or ‘auto’).
Any
Tensor representation of the dataset.

DatasetHub Class

Main class for loading and managing datasets.

Parameters

Optional[str]
Directory to cache downloaded datasets. If None, uses default cache directory.

Methods

load_dataset

Load a dataset from a file path or URL.
Union[str, Path]
required
File path or URL to load the dataset from.
Optional[DatasetFormat]
Format of the dataset. If None, inferred from file extension.
bool
default:"True"
Whether to download the dataset if it’s a URL.
Dataset
The loaded dataset.

register_dataset

Register a dataset for easy loading by name.
str
required
Name to register the dataset under.
Union[str, Path]
required
File path or URL of the dataset.
DatasetFormat
required
Format of the dataset.
Optional[Dict]
Additional metadata about the dataset.

Convenience Functions

load_dataset

Convenience function to load a dataset using the default DatasetHub instance.

register_dataset

Convenience function to register a dataset using the default DatasetHub instance.

Example Usage

Supported Formats

CSV

Comma-separated values files

JSON

JSON and JSONL formats

NumPy

.npy and .npz arrays

Pickle

Python pickle files

Text

Plain text files

Images

JPG, PNG, BMP, GIF

Audio

WAV, MP3, OGG, FLAC

Video

MP4, AVI, MOV, MKV

SQL

SQLite databases

Best Practices

Caching: DatasetHub automatically caches downloaded datasets. Use a persistent cache directory for better performance.
Transformations: Apply transformations in the Dataset constructor for automatic preprocessing during data loading.
Large datasets: For datasets that don’t fit in memory, use lazy loading or streaming formats like Parquet.