WinSTT application icon
100% Free / Open Source / MIT

WinSTT

A complete local voice toolkit for macOS, Linux, and Windows. Speech-to-text, text-to-speech, wake-word detection, and LLM-powered text processing — powered by Whisper, NeMo, and 70+ AI models, completely offline and entirely on your hardware.

100% local
Works offline
On-device AI
Free & open source
The WinSTT main window - the title bar shows the active hotkey, a 9-band audio visualizer fills the center, and the footer shows GPU, input device, and model.

The main window - 9-band audio visualizer with live hotkey, mic and model chips

See it in action

A focused look at the features you reach for every day - each one running locally on your machine.

message.txt

Can you send over the meeting notes from this morning

…this morninglarge-v3
Pastes as you speak
Words land as you speak - a fast model previews live while the accurate model finalizes.
Parakeet TDTNVIDIA · NeMo
600Mint8
Canary 1BNVIDIA · NeMo
1Bfp16
large-v3-turboOpenAI · Whisper
809Mfp16
MoonshineUseful Sensors
190Mint8
Browse 70+ speech models - maker, size, and quantization at a glance.
Ollama · Qwen 2.5ProfessionalPolite
RAW

yeah no i'm not gonna get the report done today its taking forever

CLEAN

Unfortunately, the report won't be ready today — it's taking longer than expected.

Strip filler and fix punctuation with a local LLM - and see exactly what changed.
Push-to-TalkCtrl + Space
ToggleTap on · tap off
ListenLoopback capture
Wake Word"Hey WinSTT"
Push-to-talk, toggle, passive listen, or wake-word activation.

The quarterly review went well. Revenue grew twenty-four percent year over year, so let's keep the momentum going.

USHeart1.0×
Read any text aloud with Kokoro - 54 voices across 9 languages.
Overall WPM
AI Impact248fixes made
Total Words24,31812 day streak
Activity
Words-per-minute, AI-fix impact, streaks, and a year-long activity graph.

Your voice stays on your machine

Transcription runs entirely on your local hardware. Audio is processed in-memory by on-device AI models and never written to disk or sent anywhere - no usage analytics. Optional LLM cleanup is local (Ollama) unless you opt into OpenRouter. Anonymized crash reports are opt-out. You own your data.

Everything you need

A complete speech-to-text toolkit, running locally on your desktop

Real-Time Preview

See words appear as you speak. Dual-model architecture runs a fast model for live preview alongside a large model for final accuracy.

Learn more about Real-Time Preview

70+ AI Models

OpenAI Whisper, NVIDIA NeMo (Parakeet & Canary), Moonshine, Cohere, GigaAM, Vosk. Switch models from the UI - no restart.

Learn more about 70+ AI Models

Four Recording Modes

Push-to-talk, toggle, passive listen mode (loopback capture), and wake-word activation.

Learn more about Four Recording Modes

Platform-Native Packages

Download a macOS Apple Silicon DMG, Linux AppImage/deb/rpm packages, or Windows portable builds from the same release.

Learn more about Platform-Native Packages

File Transcription

Drop audio files for batch transcription. Export as plain text or SRT subtitles with timestamps.

Learn more about File Transcription

Dictionary & Snippets

Fuzzy-match correction nudges misheard names to the right spelling. Trigger words expand into full text.

Learn more about Dictionary & Snippets

Transcription History

A local dashboard of everything you've dictated - word stats, an activity heatmap, and a searchable log.

Learn more about Transcription History

LLM Text Enhancement

Clean up dictation or run custom hotkey-triggered transforms - local Ollama or, opt-in, OpenRouter.

Learn more about LLM Text Enhancement

Text-to-Speech

Read selected text aloud with the bundled Kokoro-82M ONNX voice model - 54 voices across 9 languages.

Learn more about Text-to-Speech

Localized UI

Interface available in English, Spanish, French, Chinese, Hindi, and Arabic.

Learn more about Localized UI