● Local dictation for macOS
recording…
Double-tap fn, say what you're thinking, tap fn again. Your words land wherever your cursor is — typed out by Whisper large-v3-turbo running entirely on your Mac. No cloud, no account, no audio ever uploaded.
How it works
No window to find, no shortcut to memorize, no palette to summon. Whisper Freely lives in your menu bar and listens for one gesture, system-wide.
You hear a soft pop ◉ and the menu-bar icon turns red. The mic is live.
Dictate a sentence or a few paragraphs — into any app, any text field, mid-thought.
About two seconds later your words are pasted at your cursor, punctuated and cased.
Privacy
The only server involved is your Mac.
What you get
OpenAI's Whisper large-v3-turbo — the same family behind the best cloud dictation — running on your GPU via Apple's MLX framework.
Mail, Slack, your editor, a terminal — if you can type there, you can dictate there.
Dictate in English, Urdu, Japanese, or Portuguese and switch mid-day. Language is detected automatically.
Runs on the efficiency of Apple silicon — a few seconds of GPU per dictation, idle the rest of the time.
A single 🎙 in your menu bar. Red while recording, back to quiet when done. No dock icon, no windows.
Up to three minutes per take — dictate a full email or a memo section in one breath, not phrase by phrase.
Download
Version 1.0 · signed 520 MB app · requires an Apple Silicon Mac on macOS 14 or newer. The public download link goes live once Apple's notarization clears.
On first launch, macOS will ask for three permissions. Each maps to one thing the app visibly does:
To notice the fn double-tap — the app listens for that one gesture and ignores everything else you type.
To hear you while the red dot is on. The mic is opened when you start a dictation and closed the moment you stop.
To paste the finished text at your cursor. If you skip this, transcripts still land on your clipboard.
Questions
Yes. The model downloads once on first launch. After that you can dictate on a plane with Wi-Fi off — there is nothing to reach and no one to phone.
Whisper large-v3-turbo sits within a fraction of a word-error-point of the best cloud transcription on standard benchmarks, and it handles accents, jargon, and punctuation gracefully. Short answer: you stop checking its work.
They exist in RAM for the seconds between your two fn taps, get transcribed on the GPU, and are gone. Nothing is saved, uploaded, or retained — there is no code path that could.
A dictation costs a few seconds of GPU time via Apple's MLX framework, similar to exporting a photo. The rest of the time the app is idle, holding the model warm in memory so your next dictation starts instantly.
Not currently — Whisper Freely is built on MLX, which requires Apple silicon (M1 or later).