About this project
Dia is a 1.6B parameter text-to-speech model created by Nari Labs. It directly generates highly realistic dialogue from a transcript, supporting multiple speakers through [S1] and [S2] tags. The model can produce non-verbal communications such as laughter, coughing, sighing, and throat clearing. Users can condition output on audio prompts for emotion, tone control, and voice cloning. The project provides pretrained model checkpoints on Hugging Face, inference code, a Gradio UI, CLI, and Transformers integration. Currently, the model supports English generation only. Benchmark results on an RTX 4090 show realtime factors ranging from x0.9 to x2.2 depending on precision and compilation, with VRAM usage between 4.4GB and 7.9GB. The project is licensed under Apache 2.0 and is intended for research and educational use, with restrictions against identity misuse, deceptive content, and illegal activities.
Comments
0 Rating appears after 10 ratings
Sign in to join the discussion.