About this project
CrisperWhisper 2.0 is a speech recognition toolkit built around an explicit, per-call choice between two transcription styles. Verbatim mode writes down what was actually said in a consistent format, including fillers, repetitions, cut-offs, false starts and vocal events such as laughter. Intended mode produces the clean, readable version the speaker meant, formatting numbers, dates and emails as written text. The README presents this as the project's central design decision: most speech-to-text systems inherit a fixed style from training data, while this one exposes it as a parameter.
Beyond the two modes, the README describes word-level timings derived from supervised cross-attention alignment, reported at roughly 30 ms mean boundary error on read speech and 41 ms on conversational speech in the project's own benchmarks. A "verbatimize" capability takes audio plus an existing trusted clean transcript and inserts only the disfluencies and vocal events present in the audio, which the README frames as a way to convert clean corpora into verbatim datasets for TTS data, clinical speech analysis and dataset construction. Multilingual verbatim and intended modes are stated to work across most languages Whisper supports.
Longform audio is handled through what the README calls conditional continuation: each window continues from the words already transcribed, which the project says avoids duplicated or dropped words at chunk boundaries and avoids timestamp-token bookkeeping. A CTranslate2 runtime provides speculative decoding with a smaller draft model, and the README states hallucination mitigation for Whisper's looping-repetition failure mode is enabled by default.
Installation offers two extras: a CTranslate2 backend for NVIDIA GPUs on Linux, and a pure PyTorch backend for macOS, Windows and CPU. The first load downloads weights from HuggingFace and, on the ct2 backend, performs a one-time conversion that requires torch and transformers. Model sizes range from small to large, plus turbo, with a separate "Pro" line described as trained on additional proprietary data and supporting hotword boosting. The README also lists options such as dual-mode transcription in one pass, forced alignment for existing transcripts, and float16 or int8_float16 quantization.
Licensing is split: the inference code in the repository is MIT-licensed, while the model weights are released under a non-commercial research license, with commercial licensing available on request; the Pro models are commercial-license only. The README links to a paper, full documentation, model pages and several research posts explaining the benchmark, multilingual style control, the aligner, longform continuation, verbatimize and the inference stack. Benchmark tables in the README are the project's own reported results and should be read as such.
Comments
0 people shared their preference · Deer Point appears after 10 participants
Sign in to join the discussion.