About this project

Hugging Face Datasets is a lightweight Python library focused on two main capabilities: one-line dataloaders for public datasets, and efficient, reproducible data pre-processing. Loading data - `load_dataset()` fetches datasets from the Hugging Face Hub or from local files. The Hub hosts a large and growing collection of datasets covering image, audio, text (hundreds of languages and dialects), 3D medical imaging, video, and AI agent traces. - Local formats supported include CSV, JSON, JSONL, Parquet, Arrow, HDF5, XML, plain text, PNG, JPEG, WAV, MP3, PDF and NIfTI. - Datasets can also be built from Python objects: dictionaries, lists, Pandas DataFrames, or generators. Processing data - `dataset.map()` applies user functions to examples, with batching and multi-processing (`num_proc`) for parallel speedups. - Results are cached, so repeated processing steps are not recomputed. - The Apache Arrow backend provides zero-copy, memory-mapped storage, letting datasets exceed available RAM. - Streaming mode (`streaming=True`) iterates over data on the fly without downloading it first, useful for very large or disk-constrained workloads. Interoperability - Native conversion to and from NumPy, Pandas, Polars, Arrow, PyTorch, TensorFlow, JAX and Spark. - Built-in FAISS and Elasticsearch index support for similarity search. - A `Json()` feature type handles flexible structured data. Core classes - `Dataset`: in-memory or memory-mapped, supports indexing, slicing, random access and caching. - `IterableDataset`: lazy and streamable for out-of-core processing. - Both are wrapped in `DatasetDict` / `IterableDatasetDict` for multi-split datasets such as train/test/validation. Installation and extras - Installable via pip (`pip install datasets`) or conda (`conda install -c huggingface -c conda-forge datasets`). - Optional extras cover audio, vision, PDFs/NIfTI, and framework integrations for PyTorch, TensorFlow and JAX. Community and reproducibility - The project documents how to upload and share datasets on the Hub via the web, Python or Git. - Users are encouraged to pin dataset revisions for reproducibility. - Contributions are welcome, with a contributing guide covering issues, pull requests, code style (Ruff), testing and documentation. - The library is described in an EMNLP 2021 system demonstration paper and has versioned Zenodo DOIs for citation.