WinSTT
A complete local voice toolkit for macOS, Linux, and Windows. Speech-to-text, text-to-speech, wake-word detection, and LLM-powered text processing — powered by Whisper, NeMo, and 70+ AI models, completely offline and entirely on your hardware.

The main window - 9-band audio visualizer with live hotkey, mic and model chips
See it in action
A focused look at the features you reach for every day - each one running locally on your machine.
Can you send over the meeting notes from this morning
yeah no i'm not gonna get the report done today its taking forever
Unfortunately, the report won't be ready today — it's taking longer than expected.
The quarterly review went well. Revenue grew twenty-four percent year over year, so let's keep the momentum going.
Your voice stays on your machine
Transcription runs entirely on your local hardware. Audio is processed in-memory by on-device AI models and never written to disk or sent anywhere - no usage analytics. Optional LLM cleanup is local (Ollama) unless you opt into OpenRouter. Anonymized crash reports are opt-out. You own your data.
Everything you need
A complete speech-to-text toolkit, running locally on your desktop
Real-Time Preview
See words appear as you speak. Dual-model architecture runs a fast model for live preview alongside a large model for final accuracy.
70+ AI Models
OpenAI Whisper, NVIDIA NeMo (Parakeet & Canary), Moonshine, Cohere, GigaAM, Vosk. Switch models from the UI - no restart.
Four Recording Modes
Push-to-talk, toggle, passive listen mode (loopback capture), and wake-word activation.
Platform-Native Packages
Download a macOS Apple Silicon DMG, Linux AppImage/deb/rpm packages, or Windows portable builds from the same release.
File Transcription
Drop audio files for batch transcription. Export as plain text or SRT subtitles with timestamps.
Dictionary & Snippets
Fuzzy-match correction nudges misheard names to the right spelling. Trigger words expand into full text.
Transcription History
A local dashboard of everything you've dictated - word stats, an activity heatmap, and a searchable log.
LLM Text Enhancement
Clean up dictation or run custom hotkey-triggered transforms - local Ollama or, opt-in, OpenRouter.
Text-to-Speech
Read selected text aloud with the bundled Kokoro-82M ONNX voice model - 54 voices across 9 languages.
Localized UI
Interface available in English, Spanish, French, Chinese, Hindi, and Arabic.