Skip to main content
The dataset command provides comprehensive dataset management capabilities including listing, downloading, registering, splitting, and converting datasets.

Usage

Actions

Available Formats

  • csv - Comma-separated values
  • json - JSON format
  • npy - NumPy binary format
  • hdf5 - HDF5 format
  • tfrecord - TensorFlow record format
  • parquet - Apache Parquet format

Examples

List registered datasets

List in JSON format

Download a dataset

Download to specific location

Download from URL

Register a dataset

Register with format

Register with metadata

Get dataset info

Get info for local dataset

Split a dataset

Split with shuffling

Convert dataset format

Convert with explicit input format

Action Details

list

List all registered datasets in the local registry. Options:
  • --format: Output format (text, json)

download

Download a dataset from a URL or registered name. Arguments:
  • source: Dataset URL or registered name
Options:
  • --output: Output directory or file (default: data)
  • --format: Dataset format (auto-detected if not specified)

register

Register a dataset in the local registry for easy access. Arguments:
  • name: Dataset name
  • url: Dataset URL or file path
Options:
  • --format: Dataset format
  • --metadata: Metadata (JSON string or file path)

info

Get detailed information about a dataset. Arguments:
  • name: Dataset name or path
Options:
  • --format: Output format (text, json)

split

Split a dataset into train/validation/test sets. Arguments:
  • input: Input dataset file or directory
Options:
  • --output: Output directory (default: data)
  • --ratio: Split ratios (default: 0.8,0.2)
  • --shuffle: Shuffle data before splitting
  • --seed: Random seed for reproducibility

convert

Convert a dataset to a different format. Arguments:
  • input: Input dataset file or directory
  • output: Output file or directory
Options:
  • --input-format: Input format (auto-detected if not specified)
  • --output-format: Output format (required)

Error Handling

Dataset not found

Invalid split ratio

File not found

Use Cases

1. Download and prepare dataset

2. Register custom dataset

3. Convert dataset format

4. Create reproducible splits

5. Manage multiple datasets

Best Practices

1. Always shuffle when splitting

2. Use consistent split ratios

Standard splits:
  • 70/15/15: Balanced three-way split
  • 80/20: Simple train/val split
  • 80/10/10: More training data

3. Register datasets with metadata

4. Convert to efficient formats

For large datasets, use efficient formats:

5. Organize dataset directories

Integration Example

Complete dataset preparation pipeline:

See Also