About this project

Real-Time Voice Cloning is an implementation of the paper "Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis" (SV2TTS), originally developed as a master's thesis. It clones a voice from a few seconds of audio and then generates arbitrary speech in real time. SV2TTS is a three-stage deep learning framework. The first stage creates a digital representation of a voice from a few seconds of audio (a GE2E-based speaker encoder). In the second and third stages, that representation is used as a reference to synthesize speech for arbitrary text, using a Tacotron-based synthesizer and a WaveRNN vocoder. The repository implements SV2TTS and the GE2E encoder itself, and builds on fatchord/WaveRNN for the vocoder and synthesizer components. The project provides a toolbox with both a GUI (demo_toolbox.py) and a command-line interface (demo_cli.py). Windows and Linux are supported. Setup requires ffmpeg for reading audio files and uv for Python package management; the toolbox can be run with either a CUDA extra for NVIDIA GPUs or a CPU extra. Pretrained models are downloaded automatically, with manual downloads available from Hugging Face. Optional datasets such as LibriSpeech train-clean-100 can be used for experimentation, though users may instead supply their own audio files or record directly in the toolbox. The README notes that the repository has aged, and that many paid SaaS services now offer better audio quality. It points readers to Papers with Code for recent speech synthesis research and to Chatterbox as a more current open-source alternative.