OmniVox v0.6.3 is outSee what shipped

Dictation that never leaves your machine

Local-first dictation for Windows. Whisper and Qwen run on your own hardware, so the audio, the transcript and the text you paste never leave the machine you installed it on.

  • Windows 10/11
  • v0.6.3
  • 18 MB
  • Free during early access
The overlay — 42 px idle, shown larger here

OmniVox turns speech into text entirely on your own machine. The models live on your disk, the audio never leaves it, and there is no account, no subscription and no telemetry to opt out of.

  • 9Speech models

    Whisper, from Tiny at 75 MB through Distil Large V3 — nine builds, English and multilingual. Pick the one your hardware can hold; every one of them runs locally.

  • 99Languages

    Transcribe in any of them, or translate straight into English on the way out.

  • 0Calls to dictate

    Capture, transcription, voice commands and structuring never dial out — not under any setting. The one network call OmniVox needs is the first-run model download.

Capture

Hold Ctrl + Alt anywhere in Windows and start talking.

The hotkey is registered with the operating system rather than with an app, so it fires the same in your editor, your terminal and the reply box you already had focus in. The pill sits above the taskbar as a 42-pixel slit and expands only while it is listening — nothing to switch to, nothing to switch back from.

  • Ctrl
  • Alt
What shipped in v0.6.3

Idle 42 px · Listening 148 px

Transcribe

Whisper runs natively on your own CPU or GPU. The audio never leaves the machine.

whisper.cpp, compiled in through whisper-rs — a native binary on your machine, not a wrapper around a hosted API. The catalog runs from Tiny at 75 MB to Distil Large V3, nine builds in all, and OmniVox measures the machine it lands on and matches one to it, instead of asking you to guess.

whisper.cpp, the engine underneath

Speech models

9
Matched to your hardware on first run
Tiny (English)

Fastest, lowest resource — usable on a machine with no GPU.

75 MBf16
Use
Base (English)Active

The everyday balance of speed and accuracy.

142 MBf16
Active
Medium (English)Included

Ships with OmniVox, so there is something to talk into at once.

1.5 GBf16
Use
Medium Q5 (English)Recommended

Best value for most machines — about a third of the memory.

539 MBq5_0
Download

Speech models · main window

Weights stream to a temporary file under a hard size ceiling and are only moved into place once the download completes — so a half-fetched model is never one OmniVox will try to load.

src-tauri/src/models/downloader.rs

Structure

Say “voxify” and the transcript comes back as a filled-in prompt.

A local Qwen3 1.7B Q8 reads what you actually said and reshapes it into slots — goal, files, constraints, urgency. Decoding is constrained by a GBNF grammar, so the model can only emit the shape you asked for. It cannot drift into prose, and it cannot invent a field the schema never had.

Qwen3, the model doing the structuring
Structured· AI
Normal2 files
Goal

Fix the auth middleware so a stale JWT refreshes instead of 401ing.

Files
  • src/middleware/auth.ts
  • src/lib/tokens.ts
Constraints
  • Must not break the refresh-token flow
  • No new dependencies
Paste⌘↵RawCopyEditDismiss

Structured panel · overlay, actual size

The installer is not code-signed. What the install script does instead is check the digest GitHub reports for the file it actually downloaded — a guarantee we can keep.

public/install.ps1

Land

The text arrives in the window you were already working in.

Type simulation rather than paste: OmniVox synthesises the keystrokes into the focused window, then puts back whatever was on your clipboard before it started. Around twenty-five voice commands ride the same path — say “new paragraph” and you get two Shift+Enters, not the words.

Read the source on GitHub

Editing

19
new paragraphNew paragraphAnywhereAt end
delete last wordDelete last wordAnywhereAt end
bullet pointBullet list itemAnywhereAt end
select allSelect allAnywhereAt end
sendSend (Enter)AnywhereAt end

Voice commands · main window

Context modes

OmniVox for General

The default every fresh install starts in: a neutral writing style and your global vocabulary, carrying everything that isn’t worth a mode of its own.

The fallback when no binding matches

  • Vocabulary
  • Dictionary
  • Snippets
  • Writing style
  • App bindings

A mode is those five things saved together. OmniVox seeds a fresh install with the five above, then matches the foreground process against each mode’s bindings, so switching windows switches mode. Add as many of your own as you like — the list is a starting point, not a ceiling.

Context Mode

Right-click the pill · menu 196 px, actual size

Meeting notes

In development

Every line of the notes carries the moment it came from

OmniVox records the meeting, transcribes it on your own machine, and turns it into structured notes with timestamps that point back at the transcript.

It is built, but it has not shipped: meeting notes is not part of v0.6.3 and is not in the current download.

Notes beside the transcript

Your own rough notes, typed or dictated during the call, sit beside a searchable transcript you can jump through by timestamp and edit segment by segment.

Alt+M drops a marker the summariser treats as important. When the notes are written — overview, topics, decisions, action items, open questions, highlights — every item carries an [HH:MM:SS] reference back to the moment it came from, and a local validator strips any reference the model invented before it ever reaches the screen.

Six summary styles, plus your own instructions

  • General notes
  • Executive brief
  • One-on-one
  • Sales call
  • Interview
  • Standup

Transcript

38:02
You00:11:47

Can we lock the migration date before we scope anything else?

Meeting00:12:04

Two weeks after the audit closes works on our side.

You00:12:31Marker

Then the audit is the dependency, not the migration.

AI notes

General notes
Overview

The migration is gated on the audit closing, so the date moves with it.[00:00:42]

Topics
  • Migration window and its dependency on the audit[00:11:47]
  • Rollback if the cutover slips past Friday[00:19:08]
Action items

Confirm the audit close date[00:12:31]

YouDue Fri
Every reference matched the transcript

Meeting notes · main window · in development

Two channels, no bot

Your microphone and whatever Windows is playing, captured as two independent streams through WASAPI loopback. Meet, Zoom, Teams, Webex, Discord — it works with all of them because it listens to the machine rather than integrating with a service, and nothing joins your call.

Transcription stays local

The same Whisper engine and the same model dictation already uses. Audio is sealed to disk in recoverable 20-second chunks and each chunk is deleted the moment it has been transcribed, so no permanent recording is ever kept.

Transcribe after, or live

Run it after the meeting and Whisper never competes with your video call for CPU. Or run it live, transcribing each chunk as it seals, and have the transcript finished when the call is.

Attribution by source

Microphone segments are labelled “You”, system-audio segments “Meeting”, and renaming either applies to every segment from that source. This is attribution by where the audio came from — not neural diarization. It does not separate one voice in the room from another.

Follow-ups outlive the notes

Action items survive a regeneration of the summary rather than being rewritten with it, and they collect in a cross-meeting inbox carrying owners and due dates.

Five ways out

Polished Markdown notes, a full Markdown archive, a full JSON archive, WebVTT and SRT — so the meeting leaves in a format something else can read.

The cloud boundary

One part of this can leave your machine. The audio is not it.

The AI notes are the only step that can reach a hosted model, and only if you supply your own OpenRouter key — it is inactive until you do. Everything else, capture and transcription and search and export, runs with no network at all.

Sent, only with your key

  • Transcript text
  • The notes you typed

Never sent

  • Audio, in any form
  • Anything at all, with no key set
  • Key in Windows Credential Manager
  • Monthly cap · $2.00 default
  • Per-meeting cap · $0.05 default
  • Checked locally, before the request

Privacy architecture

Local by construction — every claim on this page is a line you can read in the source, not a policy you have to trust

OmniVox has no server, no account and nothing to sign into. Below is where to look if you would rather verify that than take it on faith.

On-device inference

whisper.cpp and llama.cpp are compiled into the binary through the whisper-rs and llama-cpp-2 crates. Every word you dictate is transcribed and structured by native code on your own silicon — there is no hosted service behind it and no key to paste.

An allow-list you can count

The Tauri webview's connect-src is 'self' plus three HuggingFace hosts, and they only ever serve model weights. Anything else the app reached for would be refused by the browser engine, not by a promise. Meeting notes adds exactly one more origin, openrouter.ai, and only when you have supplied your own key.

connect-src 'self' https://huggingface.co https://cdn-lfs.huggingface.co https://cdn-lfs-us-1.huggingface.co

One local SQLite file

History, notes, dictionary entries and your context modes live in omnivox.db in your own app-data folder. Delete a dictation from the History page and it is gone from disk — there is no copy elsewhere.

No telemetry to opt out of

Cargo.toml and package.json list no analytics SDK, no crash reporter and no auth library. The dependency lists are short enough to read in a minute, which is the whole argument.

Downloads you can verify

The install one-liner checks the SHA-256 that GitHub publishes for the asset it just fetched. Model weights stream to a .part file under a hard size ceiling and are renamed into place only once complete.

Source-available

The repository is published for exactly one reason: so the five statements above can be checked instead of believed.

Read the source

Release history

Everything here has already shipped

Each card names the release it landed in. Nothing on this row is planned work, and nothing is a demo that only runs on this page.

Read every release note
  • v0.5.0

    Scratchpad

    An always-on-top pad you can dictate into from any window. Cards or Note view, a voice-reactive mic orb, and contents that survive a restart.

  • v0.5.0

    Dictated lists

    Say “bullet point” and you get one. “Number item” counts itself upward, “end list” closes it, and the whole list lands as a single atomic paste.

  • v0.5.0

    Structuring profiles

    Agent Prompt, Email or Notes — one per context mode, so the same dictation becomes a prompt in your editor and a finished reply in your inbox.

  • v0.5.0

    Keyboard-driven confirm pill

    Enter confirms, Esc dismisses — from an overlay that never takes focus. A message about to be sent shows an editable confirm seeded with what was heard.

  • v0.4.0

    Voice Command Mode

    Speak a command and OmniVox performs it: launch or switch apps, run a search, fire keyboard and media shortcuts. Consequential actions wait behind a confirm.

  • v0.3.1

    Phonetic correction

    Vocabulary stopped being a hint and became a correction. Anything that sounds like one of your entries is rewritten to the entry exactly as you spelled it.

  • v0.2.11

    Analytics

    Lifetime words, day streaks, speaking pace and a contribution heatmap — all computed locally, from history that never leaves the machine.

  • v0.2.0

    Structured Mode

    Raw dictation into a slot-filled Markdown prompt through llama.cpp, held to shape by a GBNF grammar and degrading to plain text on timeout.

Dictation that stays local

No account, nothing to sign into. The installer is 18 MB because the models are fetched once, on first run, and never again.

irm https://tryomnivox.com/install.ps1 | iex

Windows 10 / 11 (x64)4 GB+ RAM · GPU optionalInternet once, to fetch models