Full voice mode for Claude Code, Codex, Antigravity, and Pi — running entirely on your Mac.
macOS · Apple Silicon · No cloud APIs · Nothing leaves your machine
Kokoro has no voice for Dutch, German, Polish, Russian or Ukrainian — it read them aloud in an English accent, which came out unintelligible rather than merely accented. A second on-device engine, Supertonic-3, now covers every one of the 24 languages Kokoro can't speak, Greek among them — 102 voices across 32 languages in all, with Kokoro still the default wherever it has a voice. The app routes on the voice you pick, and your agent is nudged to write its spoken reply in that voice's language. Both engines run on the Apple Neural Engine — Supertonic synthesizes at roughly 19× realtime.
Spoken replies reach beyond Claude Code and Codex to the Antigravity CLI and Pi. The Agents tab lists all four with their own status and Connect button — connect as many as you use, not one at a time. Connect wires up a speak tool, so speech starts mid-turn instead of waiting for the whole reply. Add a voice persona and each accent brings its own character; still 100% local, on the Apple Neural Engine.
Start with a word and your hands never leave the keyboard — or leave it entirely. Three seconds of silence submits; say "hold on" to cut in while it's speaking. Prefer keys? Press-to-Talk and Hold-to-Talk are one shortcut away.
Open Whisperer is a free, open-source, fully local voice mode for macOS — speech-to-text dictation and streaming text-to-speech for Claude Code, Codex, Antigravity, Pi, and any app on Apple Silicon, with nothing sent to the cloud.
Press-to-Talk, Hold-to-Talk, or fully Hands-Free. Say "initiate" to start, silence auto-submits, "hold on" interrupts.
WhisperKit for speech-to-text, Kokoro and Supertonic-3 for text-to-speech — all in-process on the Apple Neural Engine. No cloud APIs, no data leaves your Mac. Ever.
Responses begin speaking after the first sentence instead of the whole reply — with automatic fallback if anything's unavailable.
Open the Agents tab, hit Connect on each agent you use, and a speak tool is wired up. Auto-Focus brings your editor forward, Auto-Submit hits Enter for you. Just talk.
Whisper transcribes your speech across all 100 languages it supports, on-device. Pick from 102 voices across 32 languages — the full 54-voice Kokoro roster, plus every one of the 24 languages Kokoro can't speak — each accent a voice persona that colors the reply's tone.
A floating window with a real-time waveform, recording status, and your recent transcriptions. Always visible, never in the way.
Download the DMG, drag to Applications. Fully native — no Python. On first launch it downloads the Whisper and Kokoro models and runs them on the Apple Neural Engine.
Hold Ctrl and speak. Your words are transcribed locally and typed straight into Claude Code, Codex, Antigravity, or Pi.
Every reply is spoken aloud — now streaming, so it starts talking back on the very first sentence.
Yes. Open Whisperer is free and open source under the MIT license — no subscriptions, no accounts, no API keys.
No. Speech-to-text (WhisperKit) and text-to-speech (Kokoro and Supertonic-3) both run locally on your Mac. There are no cloud APIs and nothing leaves your machine.
Spoken replies work with Claude Code, Codex, Antigravity, and Pi. The Agents tab lists all four with their own Connect button, so you can connect as many as you use — Connect wires up a speak tool (a drop-in extension for Pi) so every response is read aloud. Dictation works with all your apps: your voice types into any macOS app.
An Apple Silicon Mac (M-series) running macOS. It's fully native — no Python. On first launch it downloads the on-device Whisper and Kokoro models and runs them in-process on the Apple Neural Engine.
Three ways: Press-to-Talk, Hold-to-Talk, or fully hands-free — say "initiate" to start, stay silent to submit, and say "hold on" to interrupt.
2.0.5 fixes the speaking indicator. The “Playing…” status and the speaker icon on the voice preview button only ever updated in Hands-Free — in Press-to-Talk and Hold-to-Talk they stayed dark. Playback state is now published inside the app rather than polled off disk, so both follow real playback in every mode, and react immediately instead of up to a third of a second late.
Leaving Hands-Free mid-listen no longer strands the last analyzer frame on the dictation overlay, redrawing at 30 fps instead of going quiet. The audio spectrum is also analyzed only when something draws it — every audio buffer used to be reduced to 96 frequency bands on both the microphone and playback paths, including for the default overlay style, which doesn't display them. Background timers are quieter, and paste_debug.log is now capped rather than growing for the life of an install.
2.0.4 fixes custom vocabulary. Setting any custom vocabulary used to make every dictation return an empty transcript — recording ran, nothing got typed. The app had been pinned to a fork of WhisperKit carrying a one-line fix for that; Argmax has now shipped the fix upstream in WhisperKit 1.1.0, with a more thorough implementation and ~200 new tests, so speech-to-text runs on upstream argmaxinc/WhisperKit and the fork is out of the dependency graph entirely.
Codex hooks move to ~/.codex/hooks.json — Codex warns when hooks are split between that file and an inline [hooks] table in config.toml, so reconnect Codex in Settings → Agents to move an existing install. Pi gains the persona and reply-language layers it was missing, so a spoken reply on Pi now matches Claude Code and Codex, and Claude Code no longer logs SDK auth failed on startup.
2.0.3 sets the app up for you. A four-pane sheet on first launch covers permissions, dictation, voice and connecting your coding agent, and carries model-download progress the whole way through — including the ~1.5 GB speech model, which used to download with nothing on screen explaining the wait. It's skippable at every step and reopenable any time from the menubar → Setup…, and the agent pane detects which agents you actually have and lists those first.
Personas are no longer a secret. Picking a voice has always attached a national character to every spoken reply, and nothing in the app ever said so. Now the voice picker names each language group's persona, the selected one is spelled out underneath, and Settings → Voice → Persona lets you pick a different one — it defaults to Automatic, the voice's own character, so nothing changes unless you change it. Multilingual (Supertonic) voices get no persona automatically, so this is currently the only way to give them any. Type is larger throughout the windows; the dictation overlay is deliberately untouched.
2.0.1 lets you pick how the app looks. A Themes card in Settings → General offers six palettes: Champagne (the new default — cream with champagne gold and a sky-blue tint), Cream (the original look), Light and Dark (neutral greys on white or near-black), Pastel (soft blue and rose), and Sky (soft sky blue with a champagne accent).
Your choice applies everywhere — Settings, the window chrome and the dictation overlay all follow it. The five fixed themes hold their look whatever macOS is set to; Cream is the one theme that keeps tracking your Mac's light/dark setting. The menu bar icon deliberately stays neutral and follows macOS, so it stays legible against a light or dark menu bar.
2.0.0 speaks languages Kokoro has no voice for. A second on-device engine, Supertonic-3, handles Dutch, German, Polish, Russian and Ukrainian, while Kokoro keeps its nine languages and stays the default — the app routes on the voice you pick, and your agent is nudged to write its spoken reply in that voice's language.
It also brings a real tabbed Settings window (Dictation · Voice · Agents · Advanced, plus General) and an Agents tab where all four agents connect at once. Dictation runs on WhisperKit large-v3 turbo — stronger noise robustness, ~99 languages, and clean punctuation and casing — with your custom vocabulary working at both layers, biasing the model as it listens and correcting the finished transcript. Still fully native, streaming, and 100% local.
Since we don't pay Apple $99/year for signing, you need to allow the app to run after installing:
xattr -cr /Applications/OpenWhisperer.app
Or clear quarantine on the DMG before opening it:
xattr -d com.apple.quarantine ~/Downloads/OpenWhisperer-2.0.5.dmg