About this project

MockingBird is an open-source voice cloning and text-to-speech project built on PyTorch. It is a fork of Real-Time-Voice-Cloning, which supported only English, and extends it with Mandarin Chinese support tested on several datasets such as aidatatang_200zh, magicdata, aishell3 and data_aishell. Main capabilities described in the README: - Voice cloning and speech generation in real time, with a stated goal of cloning a voice from a short sample. - Chinese Mandarin support, tested with multiple public datasets. - Runs on Windows and Linux, and includes workaround instructions for M1 macOS. - A web server mode (web.py) for remote calls, a GUI toolbox (demo_toolbox.py), and a command-line generator (gen_voice.py). - Training pipelines for encoder, synthesizer and vocoder, plus options to reuse pretrained encoder/vocoder and community-shared synthesizer models. - Documentation of common issues such as VRAM limits, batch size tuning, dataset layout and training progress expectations. The README notes the repository is no longer actively updated by the author, who points to a separate cloud-hosted service. It is released under the MIT license. The project references several research papers and implementations, including SV2TTS, GlobalStyleToken, HiFi-GAN, Fre-GAN, WaveRNN, Tacotron and GE2E.