Fastest Local Speech-to-Text on a Mac: 3 Compared
Want to go deeper than this article?
Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.
Voice working locally? Build the whole pipeline. Whisper, TTS, and voice cloning wired into real projects — hands-on courses. First chapter free, no card.
Short answer: on an M-series Mac the fastest local transcription is Parakeet on the Apple Neural Engine, not Whisper. In the third-party mac-whisper-speedtest run on a MacBook Pro M4 24GB, FluidAudio's Core ML Parakeet returned a short utterance in 0.1935s and parakeet-mlx in 0.4995s, versus 1.0230s for mlx-whisper (large-v3-turbo) and 1.2293s for whisper.cpp with Core ML. Whisper only wins when you need its 99-language coverage.
Two things you need before you touch a terminal. First, WhisperKit is no longer its own repository — it is a product inside argmaxinc/argmax-oss-swift, which means most install instructions you will find are wrong. Second, the benchmark above measures a single short clip on one machine and was last updated in August 2025. It ranks the architectures reliably. It does not tell you what your Air will do on a 60-minute file. We give you a script for that further down, because we did not have an Air and a Max on the bench to measure it for you.
The Short Answer, With The Caveat Attached
Pick by language first, then by speed. That ordering saves more time than any backend choice.
| If you need | Use | Why |
|---|---|---|
| English or one of 25 European languages, maximum speed | Parakeet TDT v3 via FluidAudio (Swift) or parakeet-mlx (Python) | Fastest published numbers on Apple Silicon by a wide margin |
| Anything outside those languages, or translation | WhisperKit large-v3-v20240930_626MB | 99-language Whisper coverage, Core ML/ANE optimised |
| A scriptable pipeline you already know | mlx_whisper CLI | One pip install, sensible CLI, no Xcode |
| Embedding in a C/C++ app or a non-Apple build too | whisper.cpp with WHISPER_COREML=1 | The portable option; Core ML moves the encoder to the ANE |
The uncomfortable part of that table is that "fastest speech-to-text on a Mac" and "Whisper on a Mac" are different questions with different answers. If you came here to make Whisper faster, the honest answer is often to stop using Whisper.
Reading articles is good. Building is better.
Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.
What Changed: WhisperKit Moved
As of 2026, WhisperKit is a library product inside argmaxinc/argmax-oss-swift — 6,326 stars, MIT licence, last pushed August 13, 2026. The repo describes itself as "On-device Speech AI for Apple Silicon" and now bundles three kits:
- WhisperKit — speech-to-text with OpenAI Whisper
- SpeakerKit — speaker diarization with Pyannote
- TTSKit — text-to-speech with Qwen-TTS
There is also an ArgmaxOSS umbrella product if you want all three. Practical consequences:
- The Swift Package URL is now
https://github.com/argmaxinc/argmax-oss-swift.git,from: "0.9.0". - The in-repo CLI target is
argmax-cli, notwhisperkit-cli. The Homebrew formula is stillbrew install whisperkit-cli. - Requirements: Xcode 16.0+ across the board, macOS 14.0+ for WhisperKit. SpeakerKit goes back to macOS 13.0; TTSKit needs macOS 15.0+.
If a tutorial tells you to add github.com/argmaxinc/WhisperKit as a package, it predates the consolidation. That single stale line is why so many people conclude "WhisperKit is broken".
The Four Stacks
Each stack is a different bet on which piece of silicon does the work. That is the whole taxonomy.
whisper.cpp — Metal by default, ANE optionally
ggml-org/whisper.cpp (52,975 stars, MIT, pushed August 14, 2026) runs inference fully on the GPU via Metal on Apple Silicon out of the box. Separately, you can build with -DWHISPER_COREML=1 to move the encoder onto the Apple Neural Engine — the README states this "can result in significant speed-up — more than x3 faster compared with CPU-only execution". Note the two limits: it is the encoder only, and the comparison baseline is CPU-only, not Metal.
One quirk worth knowing before you file a bug: the first run on a device is slow, because the ANE service compiles the Core ML model into a device-specific format. It is a one-time cost per model per machine.
WhisperKit — Core ML end to end, Swift-native
Apple-native, ships pre-converted Core ML model variants, and handles the awkward parts of long audio properly. The .incremental audio loading mode streams a large file from disk in bounded-memory chunks, splitting at VAD silence boundaries so the output matches a full-file run — only peak memory differs. From the CLI that is --incremental-loading. If you have ever watched a transcription job balloon past your RAM on a three-hour recording, that flag is the reason to pick WhisperKit.
Argmax recommends large-v3-v20240930_626MB for maximum multilingual accuracy across iOS and macOS, large-v3-v20240930_turbo for maximum speed and accuracy on macOS specifically, and tiny only for debugging.
MLX — Python, unified memory, two flavours
mlx-whisper (from ml-explore/mlx-examples) is the plain pip install mlx-whisper route with an mlx_whisper audio.mp3 CLI. Its default model is mlx-community/whisper-tiny — set --model or you are benchmarking the worst Whisper variant against everyone else's best.
Blaizzy/mlx-audio (7,750 stars, MIT, pushed August 17, 2026) is the broader library: TTS, STT and speech-to-speech on MLX, with backends for Whisper, Distil-Whisper, Parakeet, Canary, Moonshine and more under one API.
FluidAudio — ANE-only, Swift, Parakeet-first
FluidInference/FluidAudio (2,655 stars, Apache-2.0, pushed August 16, 2026) is explicitly built to keep inference on the Neural Engine and, in its own words, avoid "GPU/MPS entirely" — optimised for background and always-on workloads. Its default ASR model is Parakeet TDT v3, 0.6B parameters, 25 European languages, with a separate Japanese model and Mandarin options (SenseVoice, Paraformer). Install is Swift Package Manager, from: "0.12.4".
This is the engine behind a large share of the Mac dictation apps you have seen on Product Hunt. If your goal is "hold a hotkey, get text at the cursor", you are choosing between apps built on this and apps built on WhisperKit.
Published Numbers
Everything in this section is somebody else's measurement, labelled as such — we did not have M-series hardware on the bench for this page. Read them as evidence, not as our claim.
Third-party head-to-head (single short utterance, M4 24GB)
From anvanvan/mac-whisper-speedtest, the "large" tier on a MacBook Pro M4 24GB:
| Implementation | Time (s) | Model / config as reported |
|---|---|---|
| fluidaudio-coreml | 0.1935 | parakeet-tdt-0.6b-v2-coreml, Swift bridge |
| parakeet-mlx | 0.4995 | parakeet-tdt-0.6b-v2, MLX |
| mlx-whisper | 1.0230 | whisper-large-v3-turbo, no quantization |
| insanely-fast-whisper | 1.1324 | large-v3-turbo, mps, batch 12, fp16 compute, 4-bit |
| whisper.cpp | 1.2293 | large-v3-turbo-q5_0, coreml=True, 4 threads |
| lightning-whisper-mlx | 1.8160 | large, batch 12 |
| whisperkit | 2.2190 | large-v3, Swift bridge |
| whisper-mps | 5.3722 | large, mps |
| faster-whisper | 6.9613 | large-v3-turbo, CPU, int8 |
Four caveats you must hold onto:
- It is one short sentence. Fixed per-run overhead (model warm-up, audio load) dominates far more than it would on a 60-minute file. This is a dictation-latency benchmark, not a batch-throughput benchmark.
- WhisperKit was run on
large-v3, not the turbo variant Argmax recommends on macOS. That is not a like-for-like row, and it is the row people screenshot. - The repo was last pushed August 4, 2025. Every project in the table has shipped a year of changes since.
- faster-whisper ran on CPU because CTranslate2 has no Metal backend. That row is not a verdict on faster-whisper, it is a reminder that faster-whisper is a CUDA-first tool.
Maintainer-reported accuracy and real-time factor
From FluidAudio's own model documentation — again, their numbers, on their hardware:
| Model | Reported accuracy | Reported RTFx | Chip |
|---|---|---|---|
| Parakeet Unified 0.6B (batch) | 2.15% avg WER, LibriSpeech test-clean | 123x | M5 Pro |
| Parakeet TDT-CTC-110M | 3.01% WER, LibriSpeech test-clean | 96.5x | M2 |
| Parakeet Unified 0.6B (streaming) | 2.21% avg WER, LibriSpeech test-clean | — | — |
| Parakeet TDT Japanese | 6.85% CER, JSUT | 10.8x | M2 |
RTFx of 123x means one hour of audio in roughly 29 seconds, if it held across a full file — which is exactly the assumption a short-clip benchmark cannot test. LibriSpeech test-clean is also read audiobook speech: clean, single-speaker, close-mic. Your podcast with crosstalk will not produce a 2.15% WER on any of these. For how the two model families differ on messy audio, see Parakeet vs Whisper.
Run this on your own machine and stop paying every month
Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.
Install Each One
All four installs are short; the two that trip people up are the WhisperKit package URL and the mlx-whisper default model.
Parakeet on MLX (fastest Python route)
brew install ffmpeg
pip install parakeet-mlx -U # or: uv tool install parakeet-mlx -U
parakeet-mlx meeting.m4a --output-format srt --highlight-words
Defaults worth knowing, straight from the CLI reference: model mlx-community/parakeet-tdt-0.6b-v3, output format srt, chunk duration 120s with 15s overlap, bf16 precision, greedy decoding. On long files where memory is tight, --local-attention --local-attention-context-size 256 reduces intermediate memory so you can drop chunking entirely. --highlight-words gives you word-level timestamps in SRT/VTT.
WhisperKit (Swift / CLI)
brew install whisperkit-cli
Or as a package dependency:
dependencies: [
.package(url: "https://github.com/argmaxinc/argmax-oss-swift.git", from: "0.9.0"),
],
// then: .product(name: "WhisperKit", package: "argmax-oss-swift")
From a checkout, transcribing a long file without eating your RAM:
swift run argmax-cli transcribe \
--model large-v3-v20240930_626MB \
--audio-path "interview.wav" \
--incremental-loading
mlx-whisper
brew install ffmpeg
pip install mlx-whisper
# Set the model. The default is whisper-tiny.
mlx_whisper interview.wav --model mlx-community/whisper-large-v3-turbo -f srt
whisper.cpp with the Core ML encoder
pip install ane_transformers coremltools openai-whisper
xcode-select --install
./models/generate-coreml-model.sh base.en
cmake -B build -DWHISPER_COREML=1
cmake --build build -j --config Release
Confirm it took: the startup banner should print COREML = 1 and a line reading whisper_init_state: Core ML model loaded. If you see COREML = 0, you are on the Metal path and the build flag did not apply. Expect the very first run to stall while the ANE service compiles the model.
Benchmark Your Own Mac
Do this before you commit to a stack — it takes ten minutes and it is the only number that describes your machine. Use one real file, not a sentence.
# Make a 60-minute test file from any source recording.
ffmpeg -i source.m4a -t 3600 -ar 16000 -ac 1 -c:a pcm_s16le bench60.wav
# Time each stack on the same file.
time parakeet-mlx bench60.wav --output-format txt
time mlx_whisper bench60.wav --model mlx-community/whisper-large-v3-turbo -f txt
time ./build/bin/whisper-cli -m models/ggml-large-v3-turbo.bin -f bench60.wav -otxt
Real-time factor is just division, and it is the number to compare:
RTFx = audio duration ÷ wall-clock transcription time. A 3,600-second file finished in 60 seconds is 60x. Higher is faster.
To catch throttling rather than average it away, split the file and time the halves separately:
ffmpeg -i bench60.wav -t 1800 first30.wav
ffmpeg -i bench60.wav -ss 1800 second30.wav
time parakeet-mlx first30.wav --output-format txt
time parakeet-mlx second30.wav --output-format txt
If the second half is materially slower on identical-length audio, that gap is your thermal ceiling, not your model. Watch it live in another terminal with sudo powermetrics --samplers smc,cpu_power -i 5000.
For word error rate, transcribe a file you have a human transcript for and compare with jiwer — backend arguments are cheap, WER on your own audio is not.
Air vs Pro/Max: What Throttling Actually Costs
We did not measure this, and we are not going to guess at a percentage. Here is what to check instead, and why the answer is probably not the one you expect.
The intuition "fanless Air throttles, so buy a Pro" assumes the work lands on the GPU. On the ANE paths it largely does not. FluidAudio's stated design goal is minimising CPU usage and avoiding GPU/MPS entirely precisely so models can run as background, always-on workloads — the profile a fanless chassis handles best. whisper.cpp's Core ML build moves the encoder to the same silicon. The MLX and default Metal paths are the ones that will heat a passively cooled machine on a long file.
So the practical guidance:
- On an Air, prefer an ANE-resident stack (FluidAudio, or whisper.cpp built with
WHISPER_COREML=1) for anything longer than a few minutes. - On a Pro/Max with fans, the Metal and MLX paths become more attractive because sustained GPU load is survivable.
- Verify with the split-file test above rather than trusting either claim, including ours.
There is a second, sneakier constraint on the smallest Macs: an 8GB Air is not memory-starved by a 0.6B ASR model, but it will be if you also keep an LLM resident. If you are building a full local pipeline on Apple Silicon, our Mac local AI setup guide covers the memory budgeting, and the Apple M4 guide covers what the chip tiers actually buy you.
Memory Footprint
Rough arithmetic, so you can size before you download: parameter count times bytes per parameter gives you the weights, and activations sit on top.
| Model | Params | Weights at stated precision | Notes |
|---|---|---|---|
| Parakeet TDT v3 | 0.6B | ~1.2GB at bf16 (arithmetic) | parakeet-mlx default precision is bf16 |
| Parakeet TDT-CTC-110M | 110M | ~0.2GB at bf16 (arithmetic) | Fused preprocessor+encoder, iOS compatible |
| Parakeet EOU (streaming) | 120M | ~0.24GB at bf16 (arithmetic) | 160/320/1280ms chunk variants |
| WhisperKit large-v3 turbo (compressed) | — | 626MB on disk (Argmax naming) | The size is in the variant name |
The 626MB figure is the only one there we can quote directly rather than derive — Argmax encodes it in the model name. The rest is params × bytes and should be treated as a floor, not a peak. Peak unified-memory use during a long transcription depends heavily on chunking: WhisperKit's --incremental-loading and parakeet-mlx's --chunk-duration both exist to keep that peak bounded, and both are the first thing to reach for on a 16GB machine.
What We Could Not Verify
Being straight about this is more useful than a fabricated table. We did not have M-series hardware on the bench for this page, so the following are open:
- Our own RTF numbers per chip tier. Everything above is attributed to the projects' own docs or to the third-party speedtest repo.
- WER on identical audio across all five stacks. Nobody publishes that comparison on the same file, and the LibriSpeech figures come from a clean read-speech corpus that flatters every model.
- Sustained throughput over 60 minutes on a fanless Air. The split-file test above is the exact procedure we would have run.
- Peak unified-memory use per stack. Derived floors only; measure with Activity Monitor during a real run.
If you run the benchmark section on your own machine, that data is more authoritative than anything on this page, because it describes your Mac.
Verdict
- Try Parakeet before you try to make Whisper fast. On Apple Silicon it is the faster architecture in every published comparison we found, and for English dictation the accuracy is not the compromise people assume — 2.15% WER on read speech, per FluidAudio.
- Use Whisper for language coverage, not for speed. 99 languages and translation is a real advantage over Parakeet v3's 25. If you need it, WhisperKit's
large-v3-v20240930_626MBis the Mac-optimised way to get it. - Fix your package URL before you debug anything else. WhisperKit lives in
argmaxinc/argmax-oss-swiftnow, and the CLI target isargmax-cli. - Set
--modelon mlx-whisper. The tiny default has generated more bad takes about MLX than any actual limitation of MLX. - On a fanless Air, stay on the ANE. It is the path designed for sustained background load, and it is where the long-file win most likely is.
Once transcription is solved, the next two problems are usually who said what and how to get subtitles out of it — WhisperX handles word-level alignment and diarization, and local AI subtitles with Whisper covers the burn-in stage. For always-on dictation rather than file transcription, local voice typing compares the app layer built on top of these engines.
Sources
- argmaxinc/argmax-oss-swift — WhisperKit/TTSKit/SpeakerKit package, install paths, model variants (verified August 18, 2026)
- FluidInference/FluidAudio — Core ML/ANE ASR, model catalogue with WER and RTFx figures
- senstella/parakeet-mlx — CLI reference and defaults
- ggml-org/whisper.cpp — Core ML build instructions and the >3x encoder claim
- ml-explore/mlx-examples and Blaizzy/mlx-audio — MLX Whisper and multi-backend STT
- anvanvan/mac-whisper-speedtest — third-party M4 head-to-head (last pushed August 4, 2025)
FAQ
Voice working locally? Build the whole pipeline.
Whisper, TTS, and voice cloning wired into real projects — hands-on courses. First chapter free, no card.
Replace the speech-AI subscription
Local Speech Studio covers TTS, voice cloning and transcription end to end — including which licences actually let you sell what you make.
Liked this? 25 full AI courses are waiting.
From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.
Build Real AI on Your Machine
RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.
Want the structured version?
Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.
Keep going
- PILLARmodels/coqui-tts
- audio.cpp: Local TTS and Speech-to-Text, No Python
- Best Local Speech-to-Text Models: 4 Tested on One File
- Best Local TTS Models 2026: 8 Open-Source Voices Tested
- Best Local TTS Without a GPU: Real-Time on CPU
- Build a $10K/Month AI Podcast: Whisper + Bark + Coqui TTS
- Build a Local Voice Assistant: Whisper + Ollama + Piper
- Chatterbox TTS Setup: Free ElevenLabs Killer (MIT, 2026)
- Coqui TTS Python Guide: pip install + XTTS API Examples
- Dub Videos Into Any Language Locally: pyVideoTrans + Whisper
Comments (0)
No comments yet. Be the first to share your thoughts!