WinSTT logoWinSTT

WinSTT

Local-first speech-to-text for macOS, Linux, and Windows — real-time dictation powered by on-device AI, with optional LLM cleanup and text-to-speech. Your audio never leaves your machine.

WinSTT turns your voice into text anywhere you can type. Press a hotkey, speak, and the transcription is pasted straight into the app you were using — a chat box, an editor, a terminal, a search bar. Speech recognition runs on-device in a single native app; nothing is sent to the cloud.

The main window flow — press the hotkey and speak; the visualizer reacts and the transcription appears live.

Your voice stays on your machine

Transcription runs entirely on local hardware. Audio is processed in memory by on-device ONNX models and never written to disk or uploaded — there are no usage analytics. Optional LLM cleanup runs locally via Ollama unless you explicitly opt into OpenRouter. Anonymized crash reports (Sentry) are opt-out. The whole codebase is MIT-licensed and auditable.

70+
STT models, 11 families
85 ms
DirectML p50 (whisper-tiny-q4)
6
interface languages
0
bytes of telemetry

What you can do

See it in action

The recording overlay floats above whatever you're working in, showing the live transcription as you speak — either a bottom pill or a top dynamic island.

Floating bottom
Dynamic island

Start here

On this page