About this project
Hugging Face Tokenizers is a high-performance tokenization library designed for both research and production. It provides a Rust core implementation with bindings for Python and Node.js, and aims to be the industry-standard tokenization engine.
Key features include:
- Support for multiple tokenization models: BPE, Unigram, WordPiece, and WordLevel.
- Full pipeline components: Normalizer, PreTokenizer, Model, PostProcessor, and Decoder.
- Optimized for speed and efficiency, with hardware-adapted kernels (NEON, AVX-512, SSE, SIMD128) and portable fallbacks.
- Sub-crates for modular use: `tk-encode` (inference engine), `bitcannon` (bitstream pre-tokenization), `tk-serialize` (reader), `tk-convert` (legacy conversion), `tk-train` (training), and `bitmap_gen` (dev-only table generator).
- Easy integration with Hugging Face models via `from_pretrained`.
- Roadmap for v1.0.0 includes restoring training, batch encoding with zero-copy, dlpack API for PyTorch/JAX/NumPy interop, and more.
Installation: `pip install --pre tokenizers`
Usage example:
```python
from tokenizers import Tokenizer
tokenizer = Tokenizer.from_pretrained("meta-llama/Llama-3.1-8B")
tokenizer.encode("Hello, y'all! How are you 😁 ?")
```
The library is actively developed by Hugging Face and the community, with a focus on performance and broad language support.
Comments
0 Rating appears after 10 ratings
Sign in to join the discussion.