About this project

Hugging Face Tokenizers is a high-performance tokenization library designed for both research and production. It provides a Rust core implementation with bindings for Python and Node.js, and aims to be the industry-standard tokenization engine. Key features include: - Support for multiple tokenization models: BPE, Unigram, WordPiece, and WordLevel. - Full pipeline components: Normalizer, PreTokenizer, Model, PostProcessor, and Decoder. - Optimized for speed and efficiency, with hardware-adapted kernels (NEON, AVX-512, SSE, SIMD128) and portable fallbacks. - Sub-crates for modular use: `tk-encode` (inference engine), `bitcannon` (bitstream pre-tokenization), `tk-serialize` (reader), `tk-convert` (legacy conversion), `tk-train` (training), and `bitmap_gen` (dev-only table generator). - Easy integration with Hugging Face models via `from_pretrained`. - Roadmap for v1.0.0 includes restoring training, batch encoding with zero-copy, dlpack API for PyTorch/JAX/NumPy interop, and more. Installation: `pip install --pre tokenizers` Usage example: ```python from tokenizers import Tokenizer tokenizer = Tokenizer.from_pretrained("meta-llama/Llama-3.1-8B") tokenizer.encode("Hello, y'all! How are you 😁 ?") ``` The library is actively developed by Hugging Face and the community, with a focus on performance and broad language support.