About this project
Hugging Face Datasets is a lightweight Python library focused on two main capabilities: one-line dataloaders for public datasets, and efficient, reproducible data pre-processing.
Loading data
- `load_dataset()` fetches datasets from the Hugging Face Hub or from local files. The Hub hosts a large and growing collection of datasets covering image, audio, text (hundreds of languages and dialects), 3D medical imaging, video, and AI agent traces.
- Local formats supported include CSV, JSON, JSONL, Parquet, Arrow, HDF5, XML, plain text, PNG, JPEG, WAV, MP3, PDF and NIfTI.
- Datasets can also be built from Python objects: dictionaries, lists, Pandas DataFrames, or generators.
Processing data
- `dataset.map()` applies user functions to examples, with batching and multi-processing (`num_proc`) for parallel speedups.
- Results are cached, so repeated processing steps are not recomputed.
- The Apache Arrow backend provides zero-copy, memory-mapped storage, letting datasets exceed available RAM.
- Streaming mode (`streaming=True`) iterates over data on the fly without downloading it first, useful for very large or disk-constrained workloads.
Interoperability
- Native conversion to and from NumPy, Pandas, Polars, Arrow, PyTorch, TensorFlow, JAX and Spark.
- Built-in FAISS and Elasticsearch index support for similarity search.
- A `Json()` feature type handles flexible structured data.
Core classes
- `Dataset`: in-memory or memory-mapped, supports indexing, slicing, random access and caching.
- `IterableDataset`: lazy and streamable for out-of-core processing.
- Both are wrapped in `DatasetDict` / `IterableDatasetDict` for multi-split datasets such as train/test/validation.
Installation and extras
- Installable via pip (`pip install datasets`) or conda (`conda install -c huggingface -c conda-forge datasets`).
- Optional extras cover audio, vision, PDFs/NIfTI, and framework integrations for PyTorch, TensorFlow and JAX.
Community and reproducibility
- The project documents how to upload and share datasets on the Hub via the web, Python or Git.
- Users are encouraged to pin dataset revisions for reproducibility.
- Contributions are welcome, with a contributing guide covering issues, pull requests, code style (Ruff), testing and documentation.
- The library is described in an EMNLP 2021 system demonstration paper and has versioned Zenodo DOIs for citation.
Comments
0 Rating appears after 10 ratings
Sign in to join the discussion.