← Back to site

Feature index

Everything OmniVox does, and what has actually shipped

This is the complete inventory, not a highlight reel. Every entry below states plainly whether it is in the release you can download today or whether it only exists in the development branch.

Entries
77
In v0.6.3
59
In development
18
Shipped · v0.6.3
In the current release, and in the installer this page links to. Where an entry names an earlier version instead, that is the tagged release it first landed in — it is still in v0.6.3.
In development
Built and running in the codebase, but not in any tag. It is not in v0.6.3 and not in the download — treat it as something coming, not something you are buying into today.

01

Dictation & capture

11 entries · all shipped

Push-to-talk dictation

Shipped · v0.6.3

Hold a global hotkey — Left Ctrl + Left Alt by default — speak, and release; the text lands in whatever window had focus. It is registered with the OS, so it fires in editors, terminals and reply boxes alike. Fully rebindable.

Hands-free toggle

Shipped · v0.6.3

Double-press the same combo within 400 ms to lock recording on; press again to stop.

Floating pill overlay

Shipped · v0.6.3

An always-on-top 42-pixel capsule that expands only while listening. Shows state, a live waveform, and a right-click menu with the mode selector and quick toggles.

Live preview

Shipped · v0.6.3

Words appear in the pill while you are still speaking. Opt-in.

Ghost mode

Shipped · v0.6.3

Hides the pill for screen recordings and screenshots while keeping it interactive. Recoverable from the tray.

Noise reduction

Shipped · v0.6.3

An RNNoise pre-pass strips fan noise, keyboard clatter and non-speech before transcription.

Audio ducking

Shipped · v0.6.3

Lowers system volume while you dictate so music does not bleed into the mic, then restores it. Default 70%.

Anti-aliased resampling

Shipped · v0.3.1

Mic audio is downsampled to 16 kHz through a windowed-sinc filter rather than linear interpolation.

Input device picker

Shipped · v0.6.3

Choose which microphone OmniVox listens to.

Launch at startup

Shipped · v0.6.3

Starts with Windows via the registry Run key.

Dark, light and system themes

Shipped · v0.6.3

Light is a warm-paper treatment.

02

How text lands

4 entries · all shipped

Three output modes

Shipped · v0.6.3

Clipboard (copy only), Type (a clipboard-verified paste that then restores whatever you had on your clipboard), or Both.

Ship Mode

Shipped · v0.6.3

Presses Enter after the text lands, so a dictated message sends itself. Type and Both only.

Writing style

Shipped · v0.6.3

Formal, Casual or Very Casual capitalisation and punctuation, overridable per context mode.

Filler removal

Shipped · v0.6.3

Drops “um”, “uh”, “you know” and stutter repeats. Turn it off for verbatim quoting.

03

Voice commands

6 entries · all shipped

Spoken phrases that become real keystrokes mid-dictation. Every command has its own on/off toggle, a rename field, and a trigger scope of Anywhere or At end.

Editing commands

Shipped · v0.6.3

“new line”, “new paragraph”, “delete last word”, “select all”, “copy that”, “cut that”, “undo that”, “redo that”, “press tab”, “press escape”, “press enter”.

Dictated lists

Shipped · v0.6.3

“bullet point” and “next bullet” produce dash-prefixed items; “number item” auto-counts 1. 2. 3.; “end list” closes it. The list lands as one atomic paste.

Spoken ordinals

Shipped · v0.6.3

Saying “First… Second…” renders a numbered list with the ordinal words stripped.

“Send”

Shipped · v0.6.3

As the final word, presses Enter. Independent of Ship Mode.

Opt-in pointer commands

Shipped · v0.6.3

Mouse click, right click, double click, scroll up, scroll down and switch window. Off by default.

Custom commands

Shipped · v0.6.3

Map your own phrase to a key combo, a mouse action or an app launch.

04

Command Mode

10 entries · all shippedShipped v0.4.0 · expanded in v0.5.0

A separate dedicated hotkey — Right Ctrl by default — where the whole utterance is an instruction rather than dictation. Off by default.

Launch and switch apps

Shipped · v0.4.0

“open Spotify”, “switch to Slack”, backed by a scanned app index.

Web search and navigation

Shipped · v0.4.0

“search the web for…”, “go to…”. A spoken URL must be grounded in what you actually said; anything ungrounded parks on the confirm pill.

Keyboard chords

Shipped · v0.4.0

Copy, paste, save, new tab, close tab, screenshot, refresh, find, next and previous tab, and more.

Media and volume

Shipped · v0.4.0

Play and pause, next and previous track, mute and unmute, volume up and down. Unmute is deterministic rather than a blind toggle.

Window control

Shipped · v0.4.0

Minimise, maximise, close, full screen, minimise everything.

Type and send

Shipped · v0.4.0

“type …” pastes for review; “tell Claude to …” pastes and presses Enter. Every send routes through the confirm pill.

Multi-step chaining

Shipped · v0.4.0

One utterance can chain steps: “open Spotify and play it”.

Confirm pill

Shipped · v0.4.0

Enter confirms, Esc dismisses, Ctrl+Enter sends an editable message, from an overlay that never takes focus.

“Stop” and “undo that”

Shipped · v0.4.0

Aborts a running command, or reverses the last action: re-closing a launched app, restoring a minimised window, undoing a non-submitted paste.

Identity-bound execution

Shipped · v0.6.3

Every side-effecting action is bound to a specific window handle and owning process, and re-verified as foreground before it fires.

05

Context modes

2 entries · all shipped

Context modes

Shipped · v0.6.3

Bundles of writing style, dictionary, snippets, vocabulary, app bindings and a Structured Mode profile, saved together. The built-ins are General, Programming, Business & Sales, Medical and Legal, each seeded with a domain vocabulary. Custom modes are unlimited.

Automatic switching

Shipped · v0.6.3

The active mode follows the foreground application when recording starts.

06

Vocabulary & dictionary

4 entries · all shipped

Phonetic auto-correction

Shipped · v0.3.1

Add a word like “OmniCue” and a deterministic post-transcription pass rewrites anything that sounds like it — same consonant skeleton, edit distance within 3 — to your exact spelling. Split compounds fuse back together.

Word replacements

Shipped · v0.6.3

“Heard as → Replace with” rules, seeded with common corrections.

Snippets

Shipped · v0.6.3

A trigger word expands to longer text. Global, or scoped to one mode.

Screen context

Shipped · v0.6.3

Opt-in. Reads visible text from the focused window through Windows UI Automation, extracts identifiers, file paths and CLI flags, and feeds them to Whisper so they transcribe verbatim. Never leaves the device.

07

Structured Mode

6 entries · all shippedShipped v0.2.0 · profiles in v0.5.0

A fully local LLM pass that reshapes a raw dictation into a finished document, held to shape by a GBNF grammar so the output is physically schema-valid.

Three profiles

Shipped · v0.6.3

Agent Prompt (goal, context, constraints, files, urgency), Email (recipient, subject, body, sign-off) and Notes (a hierarchical outline). One per context mode.

The Voxify gate

Shipped · v0.2.0

Optional: structure only the utterances you end with “voxify”. Fifteen-plus phonetic variants are accepted, so a Whisper misread still triggers.

Anti-fabrication

Shipped · v0.2.0

Output is grounded in the raw transcript: files must appear in what you said, slots are de-duplicated, and third-person is rewritten to first.

The preview panel

Shipped · v0.2.0

A 420-pixel panel with the Markdown preview, metadata chips, a raw-transcript drawer, and Edit, Paste, Copy, Raw and Dismiss actions. “Raw” pastes your exact pre-LLM words.

Graceful degradation

Shipped · v0.2.0

On timeout, a missing model or a parse failure it falls back to plain text rather than failing the dictation.

A warm KV cache

Shipped · v0.2.0

The prompt cache is held across dictations, released after five idle minutes and rebuilt in the background the moment you press record.

08

Models & hardware

6 entries · all shipped

Nine Whisper models

Shipped · v0.6.3

From Tiny (75 MB) through Medium, Medium Q5, Large V3 Turbo and Distil Large V3 (1.52 GB, roughly five times faster than large-v3). English and multilingual builds.

99 languages

Shipped · v0.6.3

The multilingual builds transcribe with automatic language detection, or translate into English on the way out.

Hardware-matched recommendation

Shipped · v0.6.3

OmniVox measures the machine and recommends a model tier rather than asking you to guess, and shows the active backend.

Two Qwen3 models

Shipped · v0.6.3

Qwen3 1.7B Instruct Q8_0 (1.83 GB, the default for Structured Mode) and Qwen3 0.6B Instruct Q8_0 (639 MB).

Vulkan GPU acceleration

Shipped · v0.6.3

One build flag covers both Whisper and the LLM, on AMD and NVIDIA alike. Roughly three to five times faster than CPU.

Honest fallback

Shipped · v0.6.3

If the GPU fails to load, the overlay says it is running on CPU instead of silently degrading, and every model load is logged with its backend and duration.

09

Meeting notes

17 entries · all in development

Recording a call end to end — both sides of the audio, a live transcript, and notes that cite it — on the same local Whisper engine dictation already uses.

Not in v0.6.3

Built and running in the development branch. None of it is part of v0.6.3, and none of it is in the download this page links to.

Dual-channel capture

In development

Two independent audio streams: your microphone, and Windows system output via WASAPI loopback. Works with Meet, Zoom, Teams, Webex, Discord or anything else playing audio.

Local transcription

In development

The same Whisper engine and model as dictation. Audio is sealed to disk in recoverable 20-second chunks, and each chunk is deleted the moment it has been transcribed.

Transcribe after or during

In development

Process everything on stop so Whisper never competes with your call, or transcribe live as chunks seal.

Speaker attribution by source

In development

Microphone segments are labelled “You” and system-audio segments “Meeting”. Rename any speaker and apply it to every segment from that source at once. This is source-based attribution, not neural diarization.

Call detection

In development

Polls for two independent signals — the Windows microphone consent store plus a window-title scan — before suggesting a recording. Gated behind a setting.

The meeting widget

In development

A draggable dock with a voice-reactive orb, dual channel meters, an elapsed timer, and pause, resume and stop.

Your notes alongside the transcript

In development

Type or dictate rough notes during the call, beside a searchable live transcript with timestamp jumping and per-segment editing.

Mark moment

In development

Alt+M drops a timestamped marker the summariser treats as an importance signal.

AI notes with evidence

In development

Overview, topics, decisions, action items, open questions and highlights, every item carrying an [HH:MM:SS] reference back to the transcript. Invented references are stripped by a local validator before anything renders.

Six summary styles

In development

General notes, Executive brief, One-on-one, Sales call, Interview and Standup, plus free-text custom instructions.

Ask about this meeting

In development

Grounded follow-up questions. Retrieval runs locally: only the most relevant segments are selected on your machine, and the panel tells you how many of how many were used.

Follow-ups

In development

Durable action items that survive note regeneration, with a cross-meeting inbox, owner filters, due dates and jump-to-source.

Version history

In development

The last fifty note snapshots, with hand-edits captured automatically and safe restore.

Recurring meetings

In development

Carry agenda, participants and open action items forward from the last meeting with the same title.

Templates, tags, favourites and trash

In development

Reusable meeting setups, full-text search across titles, notes and transcripts, and a restorable trash.

Recovery

In development

Sealed audio is resumed after a crash, and a potentially billed AI request is never silently repeated.

Five export formats

In development

Polished Markdown notes, a full Markdown archive, a full JSON archive, WebVTT and SRT.

10

Your notes, history and stats

4 entries · all shipped

Notes

Shipped · v0.6.3

An auto-saving editor in the main window; dictation inserts straight into the open note.

Scratchpad

Shipped · v0.6.3

A small always-on-top pad you can dictate into from anywhere, in Cards or Note view, with a voice-reactive orb and a capture toggle that routes your hotkey into the pad while you read elsewhere. Remembers size and position, and survives a restart.

History

Shipped · v0.6.3

A searchable local archive of every dictation, with delete and export.

Analytics

Shipped · v0.6.3

Lifetime words, characters, dictations and time recorded; day streaks, average pace and active days; a contribution heatmap, a peak-hours histogram and a 30-day trend. Computed locally — the full transcript text never crosses the IPC boundary.

11

Privacy & data

7 entries · 1 in development

Everything local by default

Shipped · v0.6.3

Dictation audio and transcription are entirely on-device under every configuration. Audio is never uploaded.

No account, no telemetry

Shipped · v0.6.3

There is no server to sign into, and no analytics SDK, crash reporter or auth library in the dependency lists.

One SQLite file

Shipped · v0.6.3

History, notes, dictionary entries and modes live in omnivox.db in your own app-data folder. Delete a dictation and it is gone from disk.

History can be turned off

Shipped · v0.6.3

Fail-closed on the save path, and existing rows are securely purged.

Retention policy

Shipped · v0.6.3

New installs keep 30 days by default, adjustable up to forever, with a bounded cleanup job every 24 hours.

Per-window command authorisation

Shipped · v0.6.3

Every internal command declares which of the app’s windows may call it.

Optional cloud summaries, bring your own key

In development

Meeting AI notes are the one feature that can leave the machine, and only if you supply your own OpenRouter key. It is inactive until you do. Audio is never sent — only transcript text and the notes you typed. The key is stored in Windows Credential Manager, and a monthly spend guardrail (default $2) and a per-meeting cap (default $0.05) are enforced locally before any request.