ZERO TELEMETRY · 100% ON-DEVICE · ZERO CLOUD

Speak to anything on your machine.

Real-time voice dictation and programmable Output Commands you trigger by name, out loud. whisper.cpp, Moonshine and Parakeet TDT run natively on your own CPU or GPU, hotkeys arrive through the XDG GlobalShortcuts portal, and a seven-step first-run wizard has the whole thing working before you ever open Settings.

Download v0.6.5
chmod +x VoxCtrl-*.AppImage && ./VoxCtrl-*.AppImage
Read Docs
Linux Native · Windows Beta
XDG Portal (Zero Keylogging)
Vulkan GPU Acceleration
12 Delivery Types + Command Router
CORE ARCHITECTURE

Engineered for absolute speed and privacy.

VoxCtrl is not another wrapper around an expensive cloud API. It is a native desktop service written in Rust and Tauri — Linux-first, now with a Windows build in early beta.

Three engines, all on-device

whisper.cpp for reference accuracy, Moonshine ONNX for noisy rooms, and NVIDIA Parakeet TDT 0.6B for ultra-fast non-autoregressive transcription. All run on your own CPU or GPU — nothing leaves the machine.

whisper-rs Moonshine ONNX Parakeet TDT Vulkan

XDG GlobalShortcuts Portal

Shortcuts register with your desktop environment. The OS tells VoxCtrl only when its own key fires. VoxCtrl cannot see your typing in browsers, terminals, or password managers.

No Keylogger No Root Required Wayland & X11

Output Commands, spoken by name

Say “VoxCtrl notes, remember to call the plumber” and the text lands in your notes instead of your cursor. Twelve delivery types — inject, file, exec, pipe, socket, dbus, http, webhook, mcp, speak, chat, clipboard — plus the Voice Command Router that picks between them.

12 Delivery Types Voice Command Router targets.toml Hot-Reload

Local Neural TTS Feedback

Hear synthesized replies aloud without cloud APIs. Choose from Breeze-TTS-2, Pocket-TTS voice cloning, Piper, Inflect Micro (38MB ONNX), or lightweight eSpeak-NG.

Pocket-TTS Clone Breeze-TTS-2 Piper

Built-in MCP Server

Exposes voice dictation and speech synthesis as high-level JSON-RPC tools to AI clients like Claude Desktop and Cursor via secure local Unix domain sockets.

Model Context Protocol Claude Desktop Cursor

Low-Latency Capture Loop

Built on CPAL to keep capture latency down. Voice Activity Detection stops the recording when you stop talking, and optional RNNoise suppression filters fans and keyboards off the capture path.

CPAL RNNoise 16 kHz mono VAD

Seven-step first-run wizard

New in v0.4.0. A fresh install picks an engine and model size (downloaded before you continue), records a hotkey and registers it with your desktop, chooses an overlay, runs a live dictation test, and optionally adds a voice — instead of dropping you into a settings window full of defaults nobody chose.

Saves as you go voxctrl --setup Re-runnable

Self-updating, checksum-verified

VoxCtrl asks GitHub once on launch whether a newer release exists, shows what changed, and — if you agree — downloads the build matching your install, verifies it against the published checksum, replaces itself and restarts. Nothing is replaced until a complete, verified file is on disk. One tick in Settings → General turns the check off.

Checksum verified No identifier sent Opt-out
OUTPUT COMMANDS · VOICE COMMAND ROUTER

Say the destination out loud.

Every place your speech can land is a named Output Command. Start dictation, say “VoxCtrl”, then the command's name, then what you want to send. Say nothing of the sort and dictation goes wherever your hotkey already points — so this costs you nothing until you want it.

trigger“VoxCtrl lead-in · optional, add this to my command name notes connector: payload remember to call the plumber”
Voice Command Router — live parse
whisper.cpp · on-device
“ ”
01 Transcription, tokenised by the router
trigger lead-in command name payload no match — plain dictation
02 Match & delivery
Settings → Visual · Command Trigger Overlay (3 s, configurable)
⚡ NOTES ▸ remember to call the plumber

Longest name wins

Candidates are sorted by length, so a command called Personal Notes is matched before one called Notes — you don't have to rename anything to disambiguate.

Natural phrasing is allowed

Up to ten filler words may sit between the trigger and the command name — add, put, send, this, to my, please… — and connectors like saying, that says or with are trimmed off the front of the payload.

The trigger survives bad transcription

voxctrl, vox ctrl, vox control all count, any “<word> control” phrase counts, and a leading token within a small edit distance of voxctrl is accepted too.

No trigger, no surprise

A dictation with no trigger phrase in it falls straight through to the target your hotkey is bound to. The Command router is a superset of plain injection, not a mode you have to remember to leave.

DELIVERY TYPES

Twelve ways to land a sentence.

Each Output Command declares one delivery type. Pick a tab to see what it does and the exact targets.toml block that configures it. A thirteenth type, command, is the router itself — it inspects the transcription and hands it to one of the others.

One file, hot-reloaded.

Commands live in targets.toml, or you can add them from Settings → Output Commands. Either way the router picks up the change immediately — there is no restart, and a hotkey binding can fan one utterance out to several commands at once via target_ids.

"VoxCtrl [command-name], [your text here...]"

The name the router matches is the command's id or its label — whichever you find easier to say.

~/.config/voxctrl/targets.toml Hot-Reloads Live
# Declare named output commands
[[target]]
id = "notes"
label = "notes"
delivery = "file"
file_path = "~/.notes.txt"
file_timestamp = true

[[target]]
id = "chat"
label = "chat"
delivery = "chat"
chat_url = "http://localhost:11434/v1/chat/completions"
chat_model = "llama3"

Keystroke Injection inject

Simulates typing directly into the active focused window. Uses native Wayland wtype, X11 xdotool, or clipboard Ctrl+V fallback.

Perfect for hands-free code dictation, emails, slack messages, and editor writing.
// TOML Target Configuration
[[target]]
id = "default"
label = "Focused Window"
delivery = "inject"
strip_newlines = false
DYNAMIC HUD OVERLAYS

Eight animated overlays. Zero visual clutter.

VoxCtrl provides instant visual feedback while the microphone is live. Choose your preferred aesthetic, or run silently with tray-only notifications. Every style is also a drop-in index.html + style.css pair in the overlays folder — edit it, and the change shows up on the next dictation with no restart.

GitHub Overlays

● Wayland Compositor Viewport Overlay: Position [Bottom Center]
FIRST-RUN SETUP WIZARD

Working before you open Settings.

New in v0.4.0: a first launch walks through seven screens and ends with a configuration that actually runs — not a settings window full of defaults nobody chose. Watch it below, or take the controls yourself.

THE SETTINGS PANEL

Nothing the wizard chose is locked in.

Eleven tabs hold every option the wizard offered plus the ones it skipped. A v0.4.0 settings audit went through this window control by control and either implemented or removed anything that did nothing — so what you see here is what the app actually does.

VoxCtrl Settings — General tab, showing the setup wizard, update check and MCP server controls VoxCtrl Settings — Output Commands tab, listing the Notes, Personal Notes, Say and Command targets

General

Re-run the first-launch wizard whenever you like, hold the single HuggingFace access token every gated voice model shares (Pocket-TTS, Breeze-TTS-2, VoxCPM2), decide whether VoxCtrl asks GitHub for a newer release on startup, and switch the local MCP JSON-RPC server on or off. The socket path is spelled out for both platforms: /tmp/voxctrl-mcp.sock on Linux, \\.\pipe\voxctrl-mcp on Windows.

EngineBackend (whisper.cpp, Moonshine, Parakeet TDT, Remote Speech Engine), model size and download, GPU offloading
Post-ProcessingS1-mini dictation cleanup, basic text cleanup (filler words, punctuation, list formatting), custom vocabulary, and snippets
HotkeysGestures, bindings, and a per-keybind S1-mini override
VisualOverlay style, position, and the command trigger overlay
AudioInput device, VAD, RNNoise suppression
TTSEngine, voice downloads, global stop key, and Model Memory
OpenAI APIAny OpenAI-compatible endpoint for LLM rewriting
Bug ReportAdded in v0.5.1, after these screenshots were taken
AboutVersion, build, and what this build can put on the GPU
EARLY BETA Windows 10 21H2+ / Windows 11

VoxCtrl runs on Windows now. It needs testers.

v0.5.0 shipped the first Windows build and v0.5.1 added the GPU one beside it. Both were built by one person on one machine — so the parts that touch Windows directly are the parts most likely to misbehave: the global shortcut, typing text into other windows, and the on-screen overlay. If you run Windows, that is exactly what would be most useful to hear about.

01

Install it

An NSIS installer, about 18 MB. It is not code-signed yet, so SmartScreen will say the publisher is unknown — More info → Run anyway. Nothing else to install first.

02

Check the microphone

Windows denies microphone access silently, with no prompt and no error. If dictation produces nothing at all, turn on Settings → Privacy & security → Microphone → Let desktop apps access your microphone. This catches most people once.

03

Try the awkward parts

Punctuation and symbols are the single most valuable test — anything with % ( ) + [ ] { } ^ ~ in it. Then a few different apps, all four gesture styles, and a paragraph or two long enough that the app switches from typing to pasting.

04

Tell me what broke — one button

Settings → Bug Report. Describe what happened and press a button; it gathers the log, your Windows version, CPU and GPU, which build you are on and your settings, and files it with no GitHub account needed.

⚑

The Bug Report page shows you the entire report first

What travels
  • The words you typed into the box
  • OS, CPU, GPU and which build you are running
  • The log file, scrubbed
  • Your settings, as classified values and counts
What never does
  • Anything you dictated, and any audio
  • API keys and access tokens — reported only as set / not set
  • Your username, hostname, IP or any file path
  • Custom vocabulary, snippets, prompts, and your commands' labels, URLs and shell commands

Nothing runs at startup, on a timer, or in the background — a report exists only because you pressed a button and read it first. Redaction works from an allowlist, and a test fails the build if a setting appears that nobody has classified, so a setting added later cannot leak by being forgotten. Reports are rate-limited (five a day, twenty a month); saving one to a file never is.

Get the Windows beta Read the 15-minute test plan One installer now: -windows-x86_64-webgpu.exe puts Moonshine on any Direct3D 12 GPU and falls back to the CPU when there isn't one
RELEASE TIMELINE

Everything that landed since v0.3.6.

Twenty releases since v0.3.6, from the Breeze-TTS-2 engine to a single Linux build, a Windows beta, a Remote Speech Engine and now on-device dictation cleanup.

  1. v0.6.1 – v0.6.517–19 Sep 2026Stability + one Windows build

    Fixed a WebKitGTK bug that left a blurry smear on screen as the overlay closed on released Linux AppImages, and two silent Windows failure modes — global hotkeys blocked by an elevated foreground window, and a microphone that opens successfully but delivers only silence. Mitigated a Hyprland WebKit crash and a stuck hold-to-talk gesture, and kept the D-Bus service reliably on the session bus. The separate Windows CPU-only installer is gone — the single -windows-x86_64-webgpu.exe build now covers both, accelerating Moonshine on any Direct3D 12 GPU and falling back to the CPU when there isn't one. VoxCtrl also added a proper MIT LICENSE file.

  2. v0.5.3 – v0.5.99–16 Sep 2026Remote Speech Engine + S1-mini

    Bring your own speech engine: any OpenAI-compatible /v1/audio/transcriptions server can now handle transcription instead. S1-mini arrives as an on-device dictation cleanup pass — a Qwen3-0.6B model, run through llama.cpp in an isolated sidecar process, that smooths hesitations and self-corrections into clean prose — alongside VoxCPM2 as a sixth TTS voice. Settings gained a dedicated Post-Processing tab for S1-mini and text cleanup, and overlay styles became drop-in customizable via an index.html + style.css pair.

  3. v0.5.28 Sep 2026

    One Linux AppImage instead of separate CPU and Vulkan builds. The Vulkan build accelerates any GPU through the host driver and runs on the CPU when there is none, so the same file is right either way — and people on the old CPU AppImage still get updates.

  4. v0.5.17 Sep 2026

    Settings → Bug Report: file a report without a GitHub account, after reading the whole thing. Plus the Windows GPU build, which puts Moonshine on any Direct3D 12 GPU.

  5. v0.5.06 Sep 2026first Windows build

    VoxCtrl's first Windows release. Settings → Engine now states plainly which engine this build can put on the GPU, rather than implying all of them can. And one unreadable value in the config no longer takes the rest of your settings down with it.

  6. v0.4.04 Sep 2026setup wizard

    The seven-step first-run wizard, self-updating from GitHub releases with checksum verification, Output Targets renamed to Output Commands with the spoken form explained in the app itself, and a settings audit that either implemented or removed every control that did nothing.

  7. v0.3.10 · v0.3.93 Sep 2026

    X11 fixes: the event selection had to be split or the server refused every key, and the backend now requires XInput 2.1.

  8. v0.3.82 Sep 2026

    The AppImage starts on stock Linux Mint 21 and Ubuntu 22.04 desktops — no libfuse2, no preparation.

  9. v0.3.71 Sep 2026

    The Breeze-TTS-2 engine, and a drastic speed-up for Piper.

SECURITY GUARANTEE

Why Linux power users trust VoxCtrl.

Most voice dictation software acts like malware—reading all raw input events or requiring root udev privileges. VoxCtrl changes that permanently.

The VoxCtrl Standard

XDG GlobalShortcuts Portal

  • Zero Keystroke Snooping: Global shortcuts are registered with your desktop environment. The OS grabs the key and tells VoxCtrl only that its own shortcut fired.
  • Zero Permissions to Grant: No root access, no /etc/udev/rules.d/ alterations, no adding your user to the dangerous input group.
  • Air-Gapped Operation: After downloading models, VoxCtrl operates 100% offline. Ambient audio never touches the wire.
Traditional Dictation Apps

Raw Input Sniffing & Cloud APIs

  • Global Keyloggers: Often read /dev/input/event* directly, granting read access to everything typed in browsers, terminals, and password managers.
  • Cloud Latency & Eavesdropping: Streaming audio to proprietary third-party cloud servers introduces lag and compromises conversational privacy.
  • Permanent System Degradation: Messing with udev rules and input permissions exposes your whole desktop to unprivileged rogue processes.
GET STARTED TODAY

One AppImage. Or the Windows beta.

As of v0.5.2 there is a single Linux build: the Vulkan AppImage accelerates any NVIDIA, AMD or Intel GPU through the host driver and falls back to the CPU when there is no Vulkan device — the same file is right whether or not you have a GPU. Run it once and VoxCtrl registers its own desktop entry and icon. No libfuse2, no udev rule, no permissions to grant.

Grab the latest build from https://github.com/JRufer/VoxCtrl/releases/latest

# 1. Grab the one Linux build from the latest release
#    https://github.com/JRufer/VoxCtrl/releases/latest
#    (Vulkan-accelerated, and falls back to the CPU when there is no GPU)

# 2. Make it executable
chmod +x VoxCtrl-linux-x86_64-vulkan.AppImage

# 3. Run it. The first-run wizard opens, and VoxCtrl registers its own
#    desktop entry and icon under ~/.local/share — no install step, no
#    udev rule, no permissions to grant.
./VoxCtrl-linux-x86_64-vulkan.AppImage
Linux: glibc 2.35+ (Ubuntu 22.04, Mint 21, Debian 12, Fedora 36, Arch)
A desktop with the XDG GlobalShortcuts portal (KDE Plasma, GNOME 48+, Hyprland)
PipeWire or PulseAudio · wtype (Wayland) or xdotool (X11) for typing
Windows 10 21H2+ or Windows 11 — installer not code-signed yet