Local-first dictation for Windows. Whisper and Qwen run on your own hardware, so the audio, the transcript and the text you paste never leave the machine you installed it on.
OmniVox turns speech into text entirely on your own machine. The models live on your disk, the audio never leaves it, and there is no account, no subscription and no telemetry to opt out of.
9Speech models
Whisper, from Tiny at 75 MB through Distil Large V3 — nine builds, English and multilingual. Pick the one your hardware can hold; every one of them runs locally.
99Languages
Transcribe in any of them, or translate straight into English on the way out.
0Calls to dictate
Capture, transcription, voice commands and structuring never dial out — not under any setting. The one network call OmniVox needs is the first-run model download.
Capture
01
Hold Ctrl + Alt anywhere in Windows and start talking.
The hotkey is registered with the operating system rather than with an app, so it fires the same in your editor, your terminal and the reply box you already had focus in. The pill sits above the taskbar as a 42-pixel slit and expands only while it is listening — nothing to switch to, nothing to switch back from.
Whisper runs natively on your own CPU or GPU. The audio never leaves the machine.
whisper.cpp, compiled in through whisper-rs — a native binary on your machine, not a wrapper around a hosted API. The catalog runs from Tiny at 75 MB to Distil Large V3, nine builds in all, and OmniVox measures the machine it lands on and matches one to it, instead of asking you to guess.
Fastest, lowest resource — usable on a machine with no GPU.
75 MBf16
Use
Base (English)Active
The everyday balance of speed and accuracy.
142 MBf16
Active
Medium (English)Included
Ships with OmniVox, so there is something to talk into at once.
1.5 GBf16
Use
Medium Q5 (English)Recommended
Best value for most machines — about a third of the memory.
539 MBq5_0
Download
Speech models · main window
Weights stream to a temporary file under a hard size ceiling and are only moved into place once the download completes — so a half-fetched model is never one OmniVox will try to load.
src-tauri/src/models/downloader.rs
Structure
03
Say “voxify” and the transcript comes back as a filled-in prompt.
A local Qwen3 1.7B Q8 reads what you actually said and reshapes it into slots — goal, files, constraints, urgency. Decoding is constrained by a GBNF grammar, so the model can only emit the shape you asked for. It cannot drift into prose, and it cannot invent a field the schema never had.
Fix the auth middleware so a stale JWT refreshes instead of 401ing.
Files
src/middleware/auth.ts
src/lib/tokens.ts
Constraints
Must not break the refresh-token flow
No new dependencies
Raw transcript
Paste⌘↵RawCopyEditDismiss
Structured panel · overlay, actual size
The installer is not code-signed. What the install script does instead is check the digest GitHub reports for the file it actually downloaded — a guarantee we can keep.
public/install.ps1
Land
04
The text arrives in the window you were already working in.
Type simulation rather than paste: OmniVox synthesises the keystrokes into the focused window, then puts back whatever was on your clipboard before it started. Around twenty-five voice commands ride the same path — say “new paragraph” and you get two Shift+Enters, not the words.
The default every fresh install starts in: a neutral writing style and your global vocabulary, carrying everything that isn’t worth a mode of its own.
The fallback when no binding matches
Vocabulary seeded with the identifiers you say out loud — verifyJWT, Tauri, Qwen3 — so the phonetic pass rewrites what Whisper heard into what you actually type.
Bind it to Code.exe · WindowsTerminal.exe
A formal writing style with its own snippets, so “sig” expands to your real sign-off and a dictated reply arrives finished rather than raw.
Bind it to OUTLOOK.EXE · slack.exe
A seeded clinical vocabulary, so drug names and anatomy are corrected against a list you control instead of guessed at phonetically.
Bind it to the app you chart in
Citations, party names and the formal register that goes with them, kept in one mode so they never leak into the next casual message you send.
Bind it to WINWORD.EXE · OUTLOOK.EXE
Vocabulary
Dictionary
Snippets
Writing style
App bindings
A mode is those five things saved together. OmniVox seeds a fresh install with the five above, then matches the foreground process against each mode’s bindings, so switching windows switches mode. Add as many of your own as you like — the list is a starting point, not a ceiling.
Context Mode
Right-click the pill · menu 196 px, actual size
Meeting notes
In development
Every line of the notes carries the moment it came from
OmniVox records the meeting, transcribes it on your own machine, and turns it into structured notes with timestamps that point back at the transcript.
It is built, but it has not shipped: meeting notes is not part of v0.6.3 and is not in the current download.
Your own rough notes, typed or dictated during the call, sit beside a searchable transcript you can jump through by timestamp and edit segment by segment.
Alt+M drops a marker the summariser treats as important. When the notes are written — overview, topics, decisions, action items, open questions, highlights — every item carries an [HH:MM:SS] reference back to the moment it came from, and a local validator strips any reference the model invented before it ever reaches the screen.
Six summary styles, plus your own instructions
General notes
Executive brief
One-on-one
Sales call
Interview
Standup
Transcript
38:02
You00:11:47
Can we lock the migration date before we scope anything else?
Meeting00:12:04
Two weeks after the audit closes works on our side.
You00:12:31Marker
Then the audit is the dependency, not the migration.
AI notes
General notes
Overview
The migration is gated on the audit closing, so the date moves with it.[00:00:42]
Topics
Migration window and its dependency on the audit[00:11:47]
Rollback if the cutover slips past Friday[00:19:08]
Action items
Confirm the audit close date[00:12:31]
YouDue Fri
Every reference matched the transcript
Meeting notes · main window · in development
Two channels, no bot
Your microphone and whatever Windows is playing, captured as two independent streams through WASAPI loopback. Meet, Zoom, Teams, Webex, Discord — it works with all of them because it listens to the machine rather than integrating with a service, and nothing joins your call.
Transcription stays local
The same Whisper engine and the same model dictation already uses. Audio is sealed to disk in recoverable 20-second chunks and each chunk is deleted the moment it has been transcribed, so no permanent recording is ever kept.
Transcribe after, or live
Run it after the meeting and Whisper never competes with your video call for CPU. Or run it live, transcribing each chunk as it seals, and have the transcript finished when the call is.
Attribution by source
Microphone segments are labelled “You”, system-audio segments “Meeting”, and renaming either applies to every segment from that source. This is attribution by where the audio came from — not neural diarization. It does not separate one voice in the room from another.
Follow-ups outlive the notes
Action items survive a regeneration of the summary rather than being rewritten with it, and they collect in a cross-meeting inbox carrying owners and due dates.
Five ways out
Polished Markdown notes, a full Markdown archive, a full JSON archive, WebVTT and SRT — so the meeting leaves in a format something else can read.
The cloud boundary
One part of this can leave your machine. The audio is not it.
The AI notes are the only step that can reach a hosted model, and only if you supply your own OpenRouter key — it is inactive until you do. Everything else, capture and transcription and search and export, runs with no network at all.
Sent, only with your key
Transcript text
The notes you typed
Never sent
Audio, in any form
Anything at all, with no key set
Key in Windows Credential Manager
Monthly cap · $2.00 default
Per-meeting cap · $0.05 default
Checked locally, before the request
Privacy architecture
Local by construction — every claim on this page is a line you can read in the source, not a policy you have to trust
OmniVox has no server, no account and nothing to sign into. Below is where to look if you would rather verify that than take it on faith.
On-device inference
whisper.cpp and llama.cpp are compiled into the binary through the whisper-rs and llama-cpp-2 crates. Every word you dictate is transcribed and structured by native code on your own silicon — there is no hosted service behind it and no key to paste.
An allow-list you can count
The Tauri webview's connect-src is 'self' plus three HuggingFace hosts, and they only ever serve model weights. Anything else the app reached for would be refused by the browser engine, not by a promise. Meeting notes adds exactly one more origin, openrouter.ai, and only when you have supplied your own key.
History, notes, dictionary entries and your context modes live in omnivox.db in your own app-data folder. Delete a dictation from the History page and it is gone from disk — there is no copy elsewhere.
No telemetry to opt out of
Cargo.toml and package.json list no analytics SDK, no crash reporter and no auth library. The dependency lists are short enough to read in a minute, which is the whole argument.
Downloads you can verify
The install one-liner checks the SHA-256 that GitHub publishes for the asset it just fetched. Model weights stream to a .part file under a hard size ceiling and are renamed into place only once complete.
Source-available
The repository is published for exactly one reason: so the five statements above can be checked instead of believed.
An always-on-top pad you can dictate into from any window. Cards or Note view, a voice-reactive mic orb, and contents that survive a restart.
v0.5.0
Dictated lists
Say “bullet point” and you get one. “Number item” counts itself upward, “end list” closes it, and the whole list lands as a single atomic paste.
v0.5.0
Structuring profiles
Agent Prompt, Email or Notes — one per context mode, so the same dictation becomes a prompt in your editor and a finished reply in your inbox.
v0.5.0
Keyboard-driven confirm pill
Enter confirms, Esc dismisses — from an overlay that never takes focus. A message about to be sent shows an editable confirm seeded with what was heard.
v0.4.0
Voice Command Mode
Speak a command and OmniVox performs it: launch or switch apps, run a search, fire keyboard and media shortcuts. Consequential actions wait behind a confirm.
v0.3.1
Phonetic correction
Vocabulary stopped being a hint and became a correction. Anything that sounds like one of your entries is rewritten to the entry exactly as you spelled it.
v0.2.11
Analytics
Lifetime words, day streaks, speaking pace and a contribution heatmap — all computed locally, from history that never leaves the machine.
v0.2.0
Structured Mode
Raw dictation into a slot-filled Markdown prompt through llama.cpp, held to shape by a GBNF grammar and degrading to plain text on timeout.
v0.5.0
Scratchpad
An always-on-top pad you can dictate into from any window. Cards or Note view, a voice-reactive mic orb, and contents that survive a restart.
v0.5.0
Dictated lists
Say “bullet point” and you get one. “Number item” counts itself upward, “end list” closes it, and the whole list lands as a single atomic paste.
v0.5.0
Structuring profiles
Agent Prompt, Email or Notes — one per context mode, so the same dictation becomes a prompt in your editor and a finished reply in your inbox.
v0.5.0
Keyboard-driven confirm pill
Enter confirms, Esc dismisses — from an overlay that never takes focus. A message about to be sent shows an editable confirm seeded with what was heard.
v0.4.0
Voice Command Mode
Speak a command and OmniVox performs it: launch or switch apps, run a search, fire keyboard and media shortcuts. Consequential actions wait behind a confirm.
v0.3.1
Phonetic correction
Vocabulary stopped being a hint and became a correction. Anything that sounds like one of your entries is rewritten to the entry exactly as you spelled it.
v0.2.11
Analytics
Lifetime words, day streaks, speaking pace and a contribution heatmap — all computed locally, from history that never leaves the machine.
v0.2.0
Structured Mode
Raw dictation into a slot-filled Markdown prompt through llama.cpp, held to shape by a GBNF grammar and degrading to plain text on timeout.
Dictation that stays local
No account, nothing to sign into. The installer is 18 MB because the models are fetched once, on first run, and never again.