About this project
EnviousWispr is a free, open-source AI dictation and speech-to-text app for macOS, built for Apple Silicon and designed to run entirely on-device. Audio is processed locally and never uploaded; no account or subscription is required, and the app works offline.
Core flow: press a global keybind from any app, speak, and the app transcribes locally, optionally polishes the text with an LLM, then copies it to the clipboard and can auto-paste into the active app. The README describes sub-second transcription, with the full keybind-to-paste flow typically around a second and a half when AI polish is enabled. Silero VAD detects when you stop speaking and ends recording automatically. Keybind modes include push-to-talk, toggle, and hands-free (double-press to lock for long-form dictation).
Two ASR engines are offered. Parakeet TDT v3 (NVIDIA NeMo, via FluidAudio) is the default, runs on the Apple Neural Engine, supports 25 European languages, and needs roughly 460 MB of disk. WhisperKit (OpenAI Whisper Large v3 Turbo, via argmax) runs on the GPU, supports 99+ languages with automatic language detection, and needs roughly 1.6 GB. Both run through CoreML; the first launch downloads and compiles the model.
AI polish is optional and by default stays on-device. Options include EG-1, the project's own fine-tuned cleanup model (about 2.9 GB, macOS 14+); S1-mini by Superwhisper (about 484 MB, English-oriented, with Tone, Structure and Context style settings); Apple Intelligence (macOS 26+, no extra download); and Ollama (local or hosted). Cloud polish via OpenAI, Gemini or Claude is bring-your-own-key and sends text only, never audio. The README states EG-1 is distributed under its own model license rather than the app's GPLv3, and is a fine-tuned derivative of Qwen3-4B-Instruct-2507.
Feature highlights from the README include: Live Preview of words in the recording pill while speaking; Escape Recovery, which keeps cancelled dictation text for 24 hours (text only, never audio); Snippets that paste saved text word-for-word with date, time or clipboard fill-ins; Transcribe a File for existing audio/video (m4a, mp3, wav, aiff, caf, mp4, mov, flac) with speaker turns when voices can be separated; custom vocabulary and vocabulary packs with Contacts import; Quick Add for saving a corrected spelling from anywhere on macOS; emoji dictation by name; deterministic formatting of numbers, dates and money in English; history with All/Dictations/Transcripts filters; menu bar presence; and Sparkle auto-updates.
The README also documents reliability work: the critical path (record, transcribe, paste) is kept separate from optional enhancements so failures in polish or history saving do not block delivery; a multi-step paste path falls back automatically for apps that resist standard pasting; onboarding requires Accessibility permission and re-checks if revoked; and the speech engine runs inside the app rather than a separate helper to avoid idle freezes. Releases are described as signed, notarized and Gatekeeper-checked.
Requirements are macOS 14 (Sonoma) or later and Apple Silicon (M1 or newer). Installation is via Homebrew cask (saurabhav88/tap/enviouswispr) or a DMG download; permissions for Microphone, Accessibility and (on first paste fallback) Automation are requested. Building from source uses Swift Package Manager for compilation, while the runnable .app is assembled through Tuist and the Xcode build engine; release DMGs are produced by a script requiring full Xcode 26+, mise and Tuist.
Architecture notes in the README describe a pipeline state machine (idle, recording, transcribing, polishing, complete), Swift 6 strict concurrency with actor isolation, deliberately separate Parakeet and WhisperKit backends, and a "Heart & Limbs" pattern where the critical path never fails and optional features degrade gracefully.
Privacy: audio is processed locally and not uploaded, though local recovery can temporarily retain audio. Anonymous product analytics (PostHog) and crash reporting (Sentry) are on by default and can each be disabled in settings; crash reports exclude dictated text and audio. The project is licensed GPLv3 for application code, with the EG-1 model weights under a separate community model license that restricts redistribution and use for training competing models. The name and logo are trademarks.
Comments
0 people shared their preference · Deer Point appears after 10 participants
Sign in to join the discussion.