File Transcription
Drag an audio or video file onto WinSTT — or pick one from the tray — and get a plain-text or timestamped SRT transcript next to it. Uses your main model, fully offline.
Drop an existing recording onto WinSTT and it transcribes the whole file in one
pass, writing the text to a .txt or .srt file beside the source — no dictation,
no hotkey, same on-device model you use for live dictation.
.srt or plain .txt is written right beside it.Two ways to start
Both routes feed the same engine; only how you hand it the file differs.
Drag and drop
Drag an audio or video file from Explorer onto the main window. While the file is over the window the visualizer area shows a "Drop file to transcribe" overlay; release to start. This is the fastest path when the window is visible.
Transcribe File… from the tray
Right-click the tray icon and choose Transcribe File…. A native picker titled "Select Audio or Video File to Transcribe" opens; select one file. Use this when the window is minimized to the tray.

A request runs to completion in the background and reports progress as it goes; the running dictation model stays connected the whole time.
Supported formats
WinSTT accepts common audio and video containers — the audio track is extracted and transcribed. Anything outside this list is rejected with an "Unsupported file format" message.
| Kind | Extensions |
|---|---|
| Audio | .mp3 · .wav · .flac · .m4a · .aac · .ogg · .wma |
| Video | .mp4 · .mkv · .avi · .mov · .wmv · .flv · .webm |
Output format
One setting, in the Output tab, decides what gets written.

txtgeneral.fileTranscriptionFormatsSelect any combination of plain text, subtitle, and structured-data outputs. At least one format always remains selected.
| Format | Contents | Best for |
|---|---|---|
| TXT | Continuous plain text, no timing | Notes, articles, pasting into a doc |
| SRT | Numbered cues with start/end timecodes | Captions and subtitles for video |
| VTT | WebVTT cues with start/end timecodes | Web players and browser captions |
| JSON | Schema-tagged segments and optional word timings | Automation and downstream processing |
| CSV | One row per segment | Spreadsheets and data analysis |
Where files land
The output filename is the source path with the format appended — transcribing
interview.mp3 produces outputs such as interview.mp3.txt, interview.mp3.srt,
or interview.mp3.json. Where each
file is written depends on one setting.
autogeneral.fileTranscriptionSaveLocationAuto writes the transcript next to the source file automatically, with no prompt. Ask opens a Save dialog each time so you can choose the folder and name; cancelling the dialog cancels the transcription. Auto is the default.
Same model as dictation
File transcription reuses your main model — whatever you picked in the Model tab. There's no separate file engine, so accuracy, language scope, and device all match your dictation setup. Larger files take proportionally longer; a heavier model trades speed for accuracy here just as it does live.
The server must be running
Transcription is dispatched to the local STT server over the WebSocket. If the server is still starting up or disconnected, the request is refused with a connection error — wait for the app to show connected, then retry.
Related
Choose a model
The main model that powers both dictation and file transcription.
Processing settings
Where the File Transcription format and other post-processing options live.
Dictation
Real-time voice typing with the hotkey — the everyday way to use WinSTT.
Transcription history
Browse, search, and replay everything you've dictated.
Dictation
Press your hotkey (or just speak) and WinSTT drops polished text at your cursor. The four modes — Push-to-Talk, Toggle, Listen, and Wake Word — decide what starts a recording; everything after is identical.
Text-to-Speech
Read any selected text aloud with the on-device Kokoro-82M voice — 54 voices across 9 languages, all synthesized locally.