Quick Start
Download, launch, and dictate your first sentence in under two minutes — no Python, no accounts, no setup.
Go from a fresh download to text at your cursor in five steps — the onboarding wizard picks a model and tests your mic for you, then you hold a hotkey and speak.
- 1DownloadPortable build — no installer, no Python.
- 2OnboardPick local, choose a mic, run a quick test.
- 3Pick a modelThe tiny model is bundled and ready.
- 4Hold & speakPress the hotkey and start talking.
- 5Text at cursorRelease — it pastes where you were typing.
WinSTT is local-first. The installer bundles the speech engine and a starter model, so the only thing you need is a microphone. Audio is transcribed in memory by on-device ONNX models and never leaves your machine.
Nothing to install but the app
The app package ships the Rust speech engine and a vendored tiny model — there's
no Python, no separate service, no model download required to get your first transcription, and no
account to create. You can be fully offline.
Five steps to your first transcription
Download the app
Open the download menu and choose the package type for your desktop OS. It detects macOS, Linux, or Windows and only shows the matching release binaries. Not sure which one? See Install for platform notes and system requirements.
Launch — onboarding picks local vs cloud and tests your mic
On first launch the onboarding wizard walks you through three choices: run a local on-device model (the default — private and offline) or a cloud provider, select your microphone, and run a quick mic test so you can confirm WinSTT hears you before you dictate. Pick local and you're done in seconds.

Choose local for fully on-device transcription; the mic test confirms your input device works. Pick or download a model
The starter
tinymodel is bundled and ready, so you can skip straight to dictating. To choose a different one, open the model picker from the footer or the Model settings tab: it lists 70+ models grouped by maker, each row showing accuracy and speed bars, size, language coverage, and per-quantization download badges. Downloads stream in the background without interrupting the app.
Pick a model and quantization; downloads run in the background and most swaps are live — no restart. Hold the hotkey and speak
The default mode is Push-to-Talk Push-to-Talk: hold LCtrlLMeta and start talking. A recording overlay floats above whatever you're working in and shows a live preview as the words come in. The main window's footer shows your active hotkey, microphone, and model at a glance.
The overlay shows the live preview while you hold the hotkey and speak. Release — text appears at your cursor
Let go of the hotkey and WinSTT runs the final, accurate pass and pastes the polished text straight into wherever your cursor was — a chat box, an editor, a terminal, a search bar. That's it: hold, speak, release.

Between dictations the compact main window stays out of the way; the footer shows your hotkey, device, and model.
Don't like the default hotkey?
Push-to-Talk is bound to LCtrlLMeta out of the box. Change it — or switch to Toggle, Listen, or Wake-word triggering — on the Hotkey and recording modes pages.
If the hotkey does nothing
Some locked-down corporate machines block the low-level keyboard hooks Push-to-Talk relies on. Switch to Toggle mode (it uses a single key event) or run WinSTT as Administrator. See Troubleshooting for audio-device and download issues.
Related
Install
Pick the macOS, Linux, or Windows package that matches your machine.
Recording modes
Push-to-talk, toggle, listen, and wake-word — and when to use each.
Choose a model
Browse 70+ STT models and pick the right accuracy/speed trade-off.
Change the hotkey
Rebind Push-to-Talk and the re-paste shortcut to keys that suit you.
WinSTT
Local-first speech-to-text for macOS, Linux, and Windows — real-time dictation powered by on-device AI, with optional LLM cleanup and text-to-speech. Your audio never leaves your machine.
Install
Download the right WinSTT package for macOS, Linux, or Windows. No Python, no account, no separate speech server.