WinSTT
Local-first speech-to-text for macOS, Linux, and Windows — real-time dictation powered by on-device AI, with optional LLM cleanup and text-to-speech. Your audio never leaves your machine.
WinSTT turns your voice into text anywhere you can type. Press a hotkey, speak, and the transcription is pasted straight into the app you were using — a chat box, an editor, a terminal, a search bar. Speech recognition runs on-device in a single native app; nothing is sent to the cloud.
Your voice stays on your machine
Transcription runs entirely on local hardware. Audio is processed in memory by on-device ONNX models and never written to disk or uploaded — there are no usage analytics. Optional LLM cleanup runs locally via Ollama unless you explicitly opt into OpenRouter. Anonymized crash reports (Sentry) are opt-out. The whole codebase is MIT-licensed and auditable.
What you can do
Dictate anywhere
Push-to-talk, toggle, passive listen, or wake-word. A live preview appears as you speak; the polished final text lands at your cursor.
Real-time preview
A fast model streams words live while the main model nails the final accuracy.
LLM cleanup
Reshape dictation with tone presets and custom modifiers — local Ollama, OpenRouter, or Apple Intelligence.
Text-to-speech
Read any selection aloud with Kokoro-82M — 54 voices across 9 languages.
File transcription
Drop in audio files and export plain text or timestamped SRT subtitles.
Transcription history
A local dashboard of everything you've dictated — word stats, an activity heatmap, search, and karaoke playback of saved recordings.
Dictionary & snippets
Teach it names and jargon, fix recurring mis-hears, and expand short triggers into full text.
70+ models
Whisper, NeMo Parakeet/Canary, Moonshine, Cohere, GigaAM, Vosk, and more — all ONNX, swappable from the UI without a restart.
Local-first
Works fully offline. No accounts, no cloud, no analytics. CPU, DirectML, or OpenVINO.
See it in action
The recording overlay floats above whatever you're working in, showing the live transcription as you speak — either a bottom pill or a top dynamic island.
Start here
Quick start
Download, launch, and dictate your first sentence in under two minutes.
Install
Choose the macOS, Linux, or Windows package for your machine.
Recording modes
Push-to-talk, toggle, listen, and wake-word — and when to use each.
Choose a model
Browse 70+ STT models and pick the right accuracy/speed trade-off.
How it compares
WinSTT vs. cloud dictation apps and other local tools — privacy, cost, and features.