General
Recording mode, audio feedback, display language, the visualizer and recording overlay, live-transcription placement, and startup behavior.
The Recording tab is where you choose how recording is triggered, how WinSTT looks while you dictate, and how it starts at sign-in. It is the busiest settings tab — and several controls appear or hide depending on the recording mode you pick.
Recording
This section sets the trigger strategy and the feedback you get while dictating. The first control swaps the rest of the section in and out.

pttgeneral.recordingModeA four-way switcher: Push-to-Talk Toggle Listen Wake Word. Push-to-talk holds to record; toggle starts/stops on a tap; listen transcribes system audio from a loopback device; wake word arms a keyword detector. Switching mode shows or hides the mode-specific controls below. See Recording modes for when to use each.
falsegeneral.manualToggleStopShown only in Toggle mode. When on, a toggle session runs continuously from the first press to the second — silence-based auto-stop and silence-timing punctuation tuning are both disabled. Use it for long-form dictation where soft pauses were cutting you off.
nullgeneral.loopbackDeviceIndexShown only in Listen mode. Picks which WASAPI loopback (system-audio) device is
transcribed. null uses the default render device. This lets WinSTT caption a call,
video, or anything playing through your speakers.

Wake-word controls (Wakeword mode)
These three controls appear only when the mode is Wakeword.

alexageneral.wakeWordA keyword that arms recording when spoken. The list unifies 18 free keywords across two
engines — a 2x badge means both Porcupine (PVP) and openWakeWord (OWW) agree
(highest accuracy); single-engine keywords carry just the PVP or OWW badge. The
detector backend is chosen automatically from the keyword you select.
0.6general.wakeWordSensitivityA 0.00–1.00 slider in 0.05 steps (21 positions). Lower is stricter — fewer false triggers, may miss soft pronunciations. Higher is more permissive. 0.60 is a sensible middle for most voices.
5general.wakeWordTimeoutSeconds the gate stays armed after a detection (1–30). If you say the wake word but never follow up, the engine returns to listening after this window — protection against stray noise triggering a long-tail recording.
falsegeneral.speakerDiarizationShown only in Listen mode. Colors each speaker in the transcript and tracks identities across the session. First use downloads ~32 MB of ONNX models. Toggles live — no restart.
Audio feedback
0general.systemAudioReductionWhileDictatingNow lives on the Output tab. A 6-step slider — Off, 20%, 40%, 60%, 80%, Mute — that ducks
your speaker volume while you dictate so playback doesn't bleed into the mic. Intermediate
values drop to (100 − value)% of the previous level; the original volume is restored when
recording stops.
truegeneral.recordingSoundHidden in Listen mode. Plays a short chime when recording starts and stops. Turning it on reveals the Sound Library below.
Custom sound library
With Recording Sound on, the Sound Library lists the built-in chime plus any clips you
add. Drag in or browse for a .wav/.mp3 (≤ 3 s), then rename, preview,
select, or delete each entry. The active clip is stored as
general.recordingSoundPath (empty = built-in); your uploads live in
general.recordingSoundLibrary. The chime plays through the output device set on the
Audio tab (general.outputDeviceId).
Display
This section controls the interface language and everything you see while recording.

enuseLocaleStoreThe interface language, independent of the language you dictate in. Six locales are fully translated — Arabic (العربية), English, Spanish (Español), French (Français), Hindi (हिन्दी), and Chinese (中文). Arabic renders right-to-left. (Additional seeded locales appear in the list but still display English until their translation passes land.)
Visualizer
The visualizer is the animated audio meter shown in the main window and the recording overlay. Pick the style that suits you — all five react live to your mic input.
bargeneral.visualizerTypeA five-way switcher: Bar, Grid, Radial, Wave, and Aura. The gallery below shows each reacting to speech.
The gallery below is rendered from the app's visualizer styles so each clip can stay high-resolution, consistent, and easy to compare.
9general.visualizerBarCountShown only when the type is Bar. The number of bars, 3–21 in odd steps (3, 5, 7 … 21). More bars read denser; fewer read calmer.
Recording overlay
on, xsgeneral.showRecordingOverlayA 6-step slider — Off · XS · S · M · L · XL — that both turns the floating overlay on
and sets its size. Position 0 (Off) hides the overlay and reverts overlay-only live-display
choices to in-app; XS–XL pick the visualizer height. Greyed out in Listen mode (the
overlay never shows there). Size is stored as general.visualizerSize.
floating-bottomgeneral.overlayModeGreyed out when the overlay is off. Two layouts (shown below): Floating bottom — a two-piece pill near the bottom of the primary display; and Dynamic island — a morphing capsule docked to the top-center that grows as the live preview fills in.
bothgeneral.liveTranscriptionDisplayWhere the live preview renders, as two checkboxes — In app (the main-window feed) and
In pill (inside the overlay). The combination maps to none, in-app, in-pill, or
both. The "In pill" option is disabled while the overlay is off.
Overlay position is platform-derived
general.overlayPosition (auto/none/top/bottom, default auto) gates which screen
edge the pill may appear on. On Windows and macOS auto resolves to the bottom edge; on
Linux it resolves to none because some compositors break the paste pipeline when an
always-on-top window appears mid-keystroke.
Startup
How WinSTT behaves at sign-in and when you close the window.

falsegeneral.autoStartLaunch WinSTT automatically when you sign in.
falsegeneral.startMinimizedStart hidden in the system tray instead of opening the main window.
truegeneral.minimizeToTrayClosing the window keeps WinSTT running in the tray rather than quitting. With this off, the close button exits the app entirely.
trueRestartgeneral.sendCrashReportsOpt-out anonymized crash/error reporting via Sentry. Audio and transcripts are never sent — only stack traces and error context.
Requires a restart
Toggling Send Crash Reports takes effect on the next launch — crash reporting is switched on or off once when WinSTT starts and can't be changed mid-session.
Related General settings on other tabs
A handful of controls live under the same general.* schema but are surfaced on the tab
that fits them best:
| Setting | Key | Default | Where it lives |
|---|---|---|---|
| Auto-submit after paste | general.autoSubmit | false | Quality |
| Auto-submit key | general.autoSubmitKey | enter | Quality |
| Context awareness | general.contextAwareness | false | Quality |
| Context deny-list | general.contextDenyList | 6 password managers | LLM cleanup |
| History max entries | general.historyMaxEntries | 1000 | History |
| Recording retention | general.recordingRetention | cap | History |
| Pre-release updates | general.receivePrereleaseUpdates | false | About tab |
| Output device | general.outputDeviceId | system default | Audio |
A few highlights, with their gotchas:
falsegeneral.autoSubmitWhen on, WinSTT presses a submit key right after the paste lands. Auto-Submit Key
(general.autoSubmitKey) chooses the combo — Enter for chat boxes or Ctrl+Enter for
IDE prompts. Off pastes and leaves the cursor where the target app puts it.
falsegeneral.contextAwarenessReads text from the focused window just before each dictation and feeds it to the LLM cleanup step so names and jargon are spelled correctly. This is platform-gated and opt-in: it shows a confirmation dialog before activating, and is only offered when it actually helps (a Whisper model or LLM cleanup is active).
6 entriesgeneral.contextDenyListAn app/host allow-out list for context capture. Each entry is an executable basename
(1password.exe) or a URL host suffix (bankofamerica.com, matches subdomains). Matching
windows have their captured text, HTML, and URL stripped before reaching the LLM. Seeded
with six common password managers.
1000general.historyMaxEntriesCap on persisted transcription-history rows (10–10,000). The main process trims oldest on each insert; larger histories slow the history dashboard.
capgeneral.recordingRetentionAuto-deletes saved recordings: Never keeps everything; Cap prunes recordings beyond History Max Entries; 3 days / 2 weeks / 3 months are absolute age cutoffs. Cleanup runs at startup and whenever the policy changes — "Never" still saves, it just never prunes.
falsegeneral.receivePrereleaseUpdatesOpt in to alpha/beta auto-updates. Alpha installs always keep updating to newer alphas regardless of this toggle; the knob only changes behavior on stable builds.
Reset
A single button at the bottom restores every setting to its factory default after a confirmation dialog.
Reset to Defaults has no undo
Confirming the reset clears all tabs — model, audio, quality, dictionary, snippets, LLM, TTS, and integrations — back to defaults. There is no undo; your API keys, custom modifiers, dictionary entries, and sound library are wiped.
Related
Recording modes
Push-to-talk, toggle, listen, and wake word explained in depth, with the controls each one unlocks.
Audio settings
Input device, VAD sensitivity, the mic-release policy, and the output device for chimes.
Quality settings
Smart Endpoint, auto-submit, context awareness, and the realtime preview timing.
Transcription history
The history dashboard, max-entry cap, and recording-retention policy.
Processing
Everything WinSTT does between speech and paste — live-preview timing, smart endpoint detection, optional LLM cleanup and hotkey transforms, context awareness, and paste behavior.
Hotkeys
The four global hotkeys — push-to-talk, re-paste, read-aloud, and transform — plus the while-held combos that cycle modes and cancel a pass.