SMF·WISPR-FLOW
Can a prompt replace Wispr Flow?
Voice dictation — cross-platform dictation and meeting notes
Exhibit tracking slip
Verdict
Split the product in two and the verdict writes itself. The desktop dictation loop — hold a key, transcribe locally, paste into the focused app — is a clean, well-scoped build, and a local Whisper model is not the bottleneck: on Apple Silicon, `large-v3-turbo` transcribes 7.25 seconds of speech in 0.13 seconds, roughly 55× realtime. What you cannot rebuild in a weekend is the other half of what you're paying for: dictation on iPhone and Android, Windows support, and meeting notes with speaker identification. That's why this is a kinda and Superwhisper, which is desktop-only, is a yes.
Exhibit A — The prompt
Received on17.08.2026Build a push-to-talk dictation app for macOS on Apple Silicon. Holding Right Command records the microphone; releasing it transcribes locally and pastes the text into whatever field currently has focus. Ship it as an agent app (LSUIElement) started by a LaunchAgent, with a small indicator pinned to the right edge of the screen showing idle / listening / transcribing / error.
Run transcription with mlx-whisper (large-v3-turbo) in a long-lived Python child process that keeps the model in memory and speaks a line protocol over stdin/stdout: one WAV path in, one JSON line out. Do not reload the model per dictation.
Six things will break if you don't handle them explicitly:
1. Create the CGEventTap with `.listenOnly`. Any other mode consumes the events, and Right Command + C stops copying. Detect Right Command by the device-dependent flag bit 0x10, not by keyCode, and re-enable the tap on `tapDisabledByTimeout` or macOS will silently kill it after a few days.
2. Start recording on key-down, before the ~250 ms guard that filters accidental taps — otherwise the first word of every sentence is lost. Open the visual indicator only after the guard passes, so a normal Right Command shortcut doesn't make it flash.
3. `AVAudioEngine.prepare()` on an engine whose nodes were never touched raises an Objective-C exception that Swift cannot catch with `try`; the process dies. Access `inputNode` first to attach it to the graph.
4. mlx-whisper decodes files by shelling out to ffmpeg, which is absent from a LaunchAgent's minimal PATH. Read the WAV in Python yourself and pass a numpy float32 array to `transcribe()`.
5. macOS binds TCC permissions to `identifier + certificate leaf`, never the binary hash. Sign with a self-signed code-signing certificate, or microphone, accessibility and input-monitoring reset on every rebuild. Note that `security find-identity -v` hides self-signed certificates as untrusted — that is expected, and codesign accepts them.
6. Whisper hallucinates end-of-video boilerplate on silence. The only reliable fix is not calling it: measure speech above -45 dBFS in 20 ms windows and cancel the dictation below 400 ms of speech. Add a normalised blocklist for whatever still slips through.
If you make the indicator clickable, it must be a non-activating NSPanel sized to its own shape, otherwise it steals focus and the synthetic paste has nowhere to land.
Request microphone, accessibility and input-monitoring at startup, never mid-dictation, and keep the app alive while waiting rather than exiting — a LaunchAgent will otherwise restart it in a loop. Store transcripts as JSON Lines; never store audio.
Opening prefills the prompt — press enter to run it.
Exhibit B — What you lose
- B.1 dictation on iPhone and Android
- B.2 Windows support (a local build targets one OS at a time)
- B.3 meeting notes with speaker identification
- B.4 tone adaptation per application
- B.5 support when a macOS update breaks the event tap
Prior art
- DictéeLicense: MIT
- mlx-whisperLicense: MIT
- whisper.cppLicense: MIT
Exhibit C — Why people still pay: cross-platform reach
Because it follows you across every device you own, and because the free tier already covers 2,000 desktop dictations a week — enough that most people never hit the paywall. The subscription buys mobile and meetings, not the transcription itself.
Questions
Can I keep dictating on my phone?
No, and that's the honest reason this isn't a yes. Wispr Flow's iPhone and Android apps are half of what the subscription buys. A local build is a desktop utility bound to system-wide hotkeys and clipboard access — concepts that don't exist on mobile. If you dictate on the move, keep paying.
The free tier gives me 2,000 dictations a week. Why build anything?
For most people, don't — the free tier genuinely covers normal use. Build if your audio must not leave the machine (client calls, medical or legal notes, anything under an NDA), or if you want vocabulary tuned to your own jargon rather than a general model.
Wispr Flow feels instant. Will a local model?
Yes, and measurably so. On an M5 Max with the model resident in RAM, mlx-whisper large-v3-turbo returns 7.25 seconds of speech in 0.13 seconds — about 55× realtime, using roughly 250 MB of RAM. You wait for your own sentence, not for the transcription.
What actually costs the time in this build?
Not the AI. In a real end-to-end build, zero bugs came from the model and every one came from macOS: an event tap that swallows shortcuts, an audio API that throws an uncatchable exception, ffmpeg missing from a LaunchAgent's PATH, and permissions that reset on every rebuild unless you sign with a stable certificate. The prompt above encodes all of them.
Related tools
Receipt