★ Reading this for free? Get 25 structured AI courses + per-chapter AI tutor — the first chapter of every course free, no card.Start free in 30 secondsOr own every course: $149 once
Speech

Fastest Local Speech-to-Text on a Mac: 3 Compared

October 4, 2026
13 min read
LocalAimaster Research Team

Want to go deeper than this article?

Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.

📚AI Learning Path

Voice working locally? Build the whole pipeline. Whisper, TTS, and voice cloning wired into real projects — hands-on courses. First chapter free, no card.

Start free
Or own it for life — Lifetime $149, pay once

Short answer: on an M-series Mac the fastest local transcription is Parakeet on the Apple Neural Engine, not Whisper. In the third-party mac-whisper-speedtest run on a MacBook Pro M4 24GB, FluidAudio's Core ML Parakeet returned a short utterance in 0.1935s and parakeet-mlx in 0.4995s, versus 1.0230s for mlx-whisper (large-v3-turbo) and 1.2293s for whisper.cpp with Core ML. Whisper only wins when you need its 99-language coverage.

Two things you need before you touch a terminal. First, WhisperKit is no longer its own repository — it is a product inside argmaxinc/argmax-oss-swift, which means most install instructions you will find are wrong. Second, the benchmark above measures a single short clip on one machine and was last updated in August 2025. It ranks the architectures reliably. It does not tell you what your Air will do on a 60-minute file. We give you a script for that further down, because we did not have an Air and a Max on the bench to measure it for you.


The Short Answer, With The Caveat Attached

Pick by language first, then by speed. That ordering saves more time than any backend choice.

If you needUseWhy
English or one of 25 European languages, maximum speedParakeet TDT v3 via FluidAudio (Swift) or parakeet-mlx (Python)Fastest published numbers on Apple Silicon by a wide margin
Anything outside those languages, or translationWhisperKit large-v3-v20240930_626MB99-language Whisper coverage, Core ML/ANE optimised
A scriptable pipeline you already knowmlx_whisper CLIOne pip install, sensible CLI, no Xcode
Embedding in a C/C++ app or a non-Apple build toowhisper.cpp with WHISPER_COREML=1The portable option; Core ML moves the encoder to the ANE

The uncomfortable part of that table is that "fastest speech-to-text on a Mac" and "Whisper on a Mac" are different questions with different answers. If you came here to make Whisper faster, the honest answer is often to stop using Whisper.


Reading articles is good. Building is better.

Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.

What Changed: WhisperKit Moved

As of 2026, WhisperKit is a library product inside argmaxinc/argmax-oss-swift — 6,326 stars, MIT licence, last pushed August 13, 2026. The repo describes itself as "On-device Speech AI for Apple Silicon" and now bundles three kits:

  • WhisperKit — speech-to-text with OpenAI Whisper
  • SpeakerKit — speaker diarization with Pyannote
  • TTSKit — text-to-speech with Qwen-TTS

There is also an ArgmaxOSS umbrella product if you want all three. Practical consequences:

  • The Swift Package URL is now https://github.com/argmaxinc/argmax-oss-swift.git, from: "0.9.0".
  • The in-repo CLI target is argmax-cli, not whisperkit-cli. The Homebrew formula is still brew install whisperkit-cli.
  • Requirements: Xcode 16.0+ across the board, macOS 14.0+ for WhisperKit. SpeakerKit goes back to macOS 13.0; TTSKit needs macOS 15.0+.

If a tutorial tells you to add github.com/argmaxinc/WhisperKit as a package, it predates the consolidation. That single stale line is why so many people conclude "WhisperKit is broken".


The Four Stacks

Each stack is a different bet on which piece of silicon does the work. That is the whole taxonomy.

whisper.cpp — Metal by default, ANE optionally

ggml-org/whisper.cpp (52,975 stars, MIT, pushed August 14, 2026) runs inference fully on the GPU via Metal on Apple Silicon out of the box. Separately, you can build with -DWHISPER_COREML=1 to move the encoder onto the Apple Neural Engine — the README states this "can result in significant speed-up — more than x3 faster compared with CPU-only execution". Note the two limits: it is the encoder only, and the comparison baseline is CPU-only, not Metal.

One quirk worth knowing before you file a bug: the first run on a device is slow, because the ANE service compiles the Core ML model into a device-specific format. It is a one-time cost per model per machine.

WhisperKit — Core ML end to end, Swift-native

Apple-native, ships pre-converted Core ML model variants, and handles the awkward parts of long audio properly. The .incremental audio loading mode streams a large file from disk in bounded-memory chunks, splitting at VAD silence boundaries so the output matches a full-file run — only peak memory differs. From the CLI that is --incremental-loading. If you have ever watched a transcription job balloon past your RAM on a three-hour recording, that flag is the reason to pick WhisperKit.

Argmax recommends large-v3-v20240930_626MB for maximum multilingual accuracy across iOS and macOS, large-v3-v20240930_turbo for maximum speed and accuracy on macOS specifically, and tiny only for debugging.

MLX — Python, unified memory, two flavours

mlx-whisper (from ml-explore/mlx-examples) is the plain pip install mlx-whisper route with an mlx_whisper audio.mp3 CLI. Its default model is mlx-community/whisper-tiny — set --model or you are benchmarking the worst Whisper variant against everyone else's best.

Blaizzy/mlx-audio (7,750 stars, MIT, pushed August 17, 2026) is the broader library: TTS, STT and speech-to-speech on MLX, with backends for Whisper, Distil-Whisper, Parakeet, Canary, Moonshine and more under one API.

FluidAudio — ANE-only, Swift, Parakeet-first

FluidInference/FluidAudio (2,655 stars, Apache-2.0, pushed August 16, 2026) is explicitly built to keep inference on the Neural Engine and, in its own words, avoid "GPU/MPS entirely" — optimised for background and always-on workloads. Its default ASR model is Parakeet TDT v3, 0.6B parameters, 25 European languages, with a separate Japanese model and Mandarin options (SenseVoice, Paraformer). Install is Swift Package Manager, from: "0.12.4".

This is the engine behind a large share of the Mac dictation apps you have seen on Product Hunt. If your goal is "hold a hotkey, get text at the cursor", you are choosing between apps built on this and apps built on WhisperKit.


Published Numbers

Everything in this section is somebody else's measurement, labelled as such — we did not have M-series hardware on the bench for this page. Read them as evidence, not as our claim.

Third-party head-to-head (single short utterance, M4 24GB)

From anvanvan/mac-whisper-speedtest, the "large" tier on a MacBook Pro M4 24GB:

ImplementationTime (s)Model / config as reported
fluidaudio-coreml0.1935parakeet-tdt-0.6b-v2-coreml, Swift bridge
parakeet-mlx0.4995parakeet-tdt-0.6b-v2, MLX
mlx-whisper1.0230whisper-large-v3-turbo, no quantization
insanely-fast-whisper1.1324large-v3-turbo, mps, batch 12, fp16 compute, 4-bit
whisper.cpp1.2293large-v3-turbo-q5_0, coreml=True, 4 threads
lightning-whisper-mlx1.8160large, batch 12
whisperkit2.2190large-v3, Swift bridge
whisper-mps5.3722large, mps
faster-whisper6.9613large-v3-turbo, CPU, int8

Four caveats you must hold onto:

  1. It is one short sentence. Fixed per-run overhead (model warm-up, audio load) dominates far more than it would on a 60-minute file. This is a dictation-latency benchmark, not a batch-throughput benchmark.
  2. WhisperKit was run on large-v3, not the turbo variant Argmax recommends on macOS. That is not a like-for-like row, and it is the row people screenshot.
  3. The repo was last pushed August 4, 2025. Every project in the table has shipped a year of changes since.
  4. faster-whisper ran on CPU because CTranslate2 has no Metal backend. That row is not a verdict on faster-whisper, it is a reminder that faster-whisper is a CUDA-first tool.

Maintainer-reported accuracy and real-time factor

From FluidAudio's own model documentation — again, their numbers, on their hardware:

ModelReported accuracyReported RTFxChip
Parakeet Unified 0.6B (batch)2.15% avg WER, LibriSpeech test-clean123xM5 Pro
Parakeet TDT-CTC-110M3.01% WER, LibriSpeech test-clean96.5xM2
Parakeet Unified 0.6B (streaming)2.21% avg WER, LibriSpeech test-clean——
Parakeet TDT Japanese6.85% CER, JSUT10.8xM2

RTFx of 123x means one hour of audio in roughly 29 seconds, if it held across a full file — which is exactly the assumption a short-clip benchmark cannot test. LibriSpeech test-clean is also read audiobook speech: clean, single-speaker, close-mic. Your podcast with crosstalk will not produce a 2.15% WER on any of these. For how the two model families differ on messy audio, see Parakeet vs Whisper.


Own it instead of renting it

Run this on your own machine and stop paying every month

Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.

Install Each One

All four installs are short; the two that trip people up are the WhisperKit package URL and the mlx-whisper default model.

Parakeet on MLX (fastest Python route)

brew install ffmpeg
pip install parakeet-mlx -U          # or: uv tool install parakeet-mlx -U

parakeet-mlx meeting.m4a --output-format srt --highlight-words

Defaults worth knowing, straight from the CLI reference: model mlx-community/parakeet-tdt-0.6b-v3, output format srt, chunk duration 120s with 15s overlap, bf16 precision, greedy decoding. On long files where memory is tight, --local-attention --local-attention-context-size 256 reduces intermediate memory so you can drop chunking entirely. --highlight-words gives you word-level timestamps in SRT/VTT.

WhisperKit (Swift / CLI)

brew install whisperkit-cli

Or as a package dependency:

dependencies: [
    .package(url: "https://github.com/argmaxinc/argmax-oss-swift.git", from: "0.9.0"),
],
// then: .product(name: "WhisperKit", package: "argmax-oss-swift")

From a checkout, transcribing a long file without eating your RAM:

swift run argmax-cli transcribe \
  --model large-v3-v20240930_626MB \
  --audio-path "interview.wav" \
  --incremental-loading

mlx-whisper

brew install ffmpeg
pip install mlx-whisper

# Set the model. The default is whisper-tiny.
mlx_whisper interview.wav --model mlx-community/whisper-large-v3-turbo -f srt

whisper.cpp with the Core ML encoder

pip install ane_transformers coremltools openai-whisper
xcode-select --install

./models/generate-coreml-model.sh base.en
cmake -B build -DWHISPER_COREML=1
cmake --build build -j --config Release

Confirm it took: the startup banner should print COREML = 1 and a line reading whisper_init_state: Core ML model loaded. If you see COREML = 0, you are on the Metal path and the build flag did not apply. Expect the very first run to stall while the ANE service compiles the model.


Benchmark Your Own Mac

Do this before you commit to a stack — it takes ten minutes and it is the only number that describes your machine. Use one real file, not a sentence.

# Make a 60-minute test file from any source recording.
ffmpeg -i source.m4a -t 3600 -ar 16000 -ac 1 -c:a pcm_s16le bench60.wav

# Time each stack on the same file.
time parakeet-mlx bench60.wav --output-format txt
time mlx_whisper bench60.wav --model mlx-community/whisper-large-v3-turbo -f txt
time ./build/bin/whisper-cli -m models/ggml-large-v3-turbo.bin -f bench60.wav -otxt

Real-time factor is just division, and it is the number to compare:

RTFx = audio duration ÷ wall-clock transcription time. A 3,600-second file finished in 60 seconds is 60x. Higher is faster.

To catch throttling rather than average it away, split the file and time the halves separately:

ffmpeg -i bench60.wav -t 1800 first30.wav
ffmpeg -i bench60.wav -ss 1800 second30.wav
time parakeet-mlx first30.wav  --output-format txt
time parakeet-mlx second30.wav --output-format txt

If the second half is materially slower on identical-length audio, that gap is your thermal ceiling, not your model. Watch it live in another terminal with sudo powermetrics --samplers smc,cpu_power -i 5000.

For word error rate, transcribe a file you have a human transcript for and compare with jiwer — backend arguments are cheap, WER on your own audio is not.


Air vs Pro/Max: What Throttling Actually Costs

We did not measure this, and we are not going to guess at a percentage. Here is what to check instead, and why the answer is probably not the one you expect.

The intuition "fanless Air throttles, so buy a Pro" assumes the work lands on the GPU. On the ANE paths it largely does not. FluidAudio's stated design goal is minimising CPU usage and avoiding GPU/MPS entirely precisely so models can run as background, always-on workloads — the profile a fanless chassis handles best. whisper.cpp's Core ML build moves the encoder to the same silicon. The MLX and default Metal paths are the ones that will heat a passively cooled machine on a long file.

So the practical guidance:

  • On an Air, prefer an ANE-resident stack (FluidAudio, or whisper.cpp built with WHISPER_COREML=1) for anything longer than a few minutes.
  • On a Pro/Max with fans, the Metal and MLX paths become more attractive because sustained GPU load is survivable.
  • Verify with the split-file test above rather than trusting either claim, including ours.

There is a second, sneakier constraint on the smallest Macs: an 8GB Air is not memory-starved by a 0.6B ASR model, but it will be if you also keep an LLM resident. If you are building a full local pipeline on Apple Silicon, our Mac local AI setup guide covers the memory budgeting, and the Apple M4 guide covers what the chip tiers actually buy you.


Memory Footprint

Rough arithmetic, so you can size before you download: parameter count times bytes per parameter gives you the weights, and activations sit on top.

ModelParamsWeights at stated precisionNotes
Parakeet TDT v30.6B~1.2GB at bf16 (arithmetic)parakeet-mlx default precision is bf16
Parakeet TDT-CTC-110M110M~0.2GB at bf16 (arithmetic)Fused preprocessor+encoder, iOS compatible
Parakeet EOU (streaming)120M~0.24GB at bf16 (arithmetic)160/320/1280ms chunk variants
WhisperKit large-v3 turbo (compressed)—626MB on disk (Argmax naming)The size is in the variant name

The 626MB figure is the only one there we can quote directly rather than derive — Argmax encodes it in the model name. The rest is params × bytes and should be treated as a floor, not a peak. Peak unified-memory use during a long transcription depends heavily on chunking: WhisperKit's --incremental-loading and parakeet-mlx's --chunk-duration both exist to keep that peak bounded, and both are the first thing to reach for on a 16GB machine.


What We Could Not Verify

Being straight about this is more useful than a fabricated table. We did not have M-series hardware on the bench for this page, so the following are open:

  • Our own RTF numbers per chip tier. Everything above is attributed to the projects' own docs or to the third-party speedtest repo.
  • WER on identical audio across all five stacks. Nobody publishes that comparison on the same file, and the LibriSpeech figures come from a clean read-speech corpus that flatters every model.
  • Sustained throughput over 60 minutes on a fanless Air. The split-file test above is the exact procedure we would have run.
  • Peak unified-memory use per stack. Derived floors only; measure with Activity Monitor during a real run.

If you run the benchmark section on your own machine, that data is more authoritative than anything on this page, because it describes your Mac.


Verdict

  1. Try Parakeet before you try to make Whisper fast. On Apple Silicon it is the faster architecture in every published comparison we found, and for English dictation the accuracy is not the compromise people assume — 2.15% WER on read speech, per FluidAudio.
  2. Use Whisper for language coverage, not for speed. 99 languages and translation is a real advantage over Parakeet v3's 25. If you need it, WhisperKit's large-v3-v20240930_626MB is the Mac-optimised way to get it.
  3. Fix your package URL before you debug anything else. WhisperKit lives in argmaxinc/argmax-oss-swift now, and the CLI target is argmax-cli.
  4. Set --model on mlx-whisper. The tiny default has generated more bad takes about MLX than any actual limitation of MLX.
  5. On a fanless Air, stay on the ANE. It is the path designed for sustained background load, and it is where the long-file win most likely is.

Once transcription is solved, the next two problems are usually who said what and how to get subtitles out of it — WhisperX handles word-level alignment and diarization, and local AI subtitles with Whisper covers the burn-in stage. For always-on dictation rather than file transcription, local voice typing compares the app layer built on top of these engines.


Sources


FAQ

🎯
AI Learning Path

Voice working locally? Build the whole pipeline.

Whisper, TTS, and voice cloning wired into real projects — hands-on courses. First chapter free, no card.

Or own it for life — Lifetime $149 $599, pay once
Once your hardware is sorted

Replace the speech-AI subscription

Local Speech Studio covers TTS, voice cloning and transcription end to end — including which licences actually let you sell what you make.

$149 once unlocks everything, forever — about $0.27/chapter for life. Prefer to spread it out? Pro is $79/year (saves 27%) or $8.99/month.
Secure checkout by Lemon Squeezy — your card never touches this siteInstant access the moment you payFirst chapter of every course is free — try before you buy

Liked this? 25 full AI courses are waiting.

From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.

Reading now
Join the discussion
TagsWhisperWhisperKitParakeetMLXApple SiliconSpeech to TextMac

LocalAimaster Research Team

Local AI Master writes hands-on courses and hardware guides for running AI on machines you own. Content is checked against current releases and corrected when readers tell us it is wrong.

Build Real AI on Your Machine

RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.

Want the structured version?

Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.

AI Learning Path
More on Local Voice & Speech
See the full Coqui TTS & Local Voice AI guide.

Comments (0)

No comments yet. Be the first to share your thoughts!

What is the fastest local speech-to-text on an M-series Mac?

On published third-party numbers, the fastest option is not Whisper — it is NVIDIA Parakeet running through a Core ML/ANE wrapper. In anvanvan/mac-whisper-speedtest on a MacBook Pro M4 24GB, FluidAudio's Core ML Parakeet returned a short utterance in 0.1935s and parakeet-mlx in 0.4995s, against 1.0230s for mlx-whisper (large-v3-turbo) and 1.2293s for whisper.cpp with Core ML enabled. That benchmark was last updated in August 2025 and measures a single short clip, so treat it as a ranking of the architectures rather than current absolute numbers. The catch is language coverage: Parakeet TDT v3 handles 25 European languages, where Whisper handles 99.

Does running Whisper on the Apple Neural Engine hurt accuracy?

Not by itself. The ANE runs the same weights, so accuracy differences come from the model variant and the quantization you chose, not from the compute unit. whisper.cpp's Core ML path only moves the encoder to the ANE and leaves the decoder on CPU/Metal — its README claims more than 3x faster encoder inference versus CPU-only. WhisperKit ships compressed Core ML variants (the recommended large-v3-v20240930_626MB is a compressed Large v3 Turbo), and compression is where quality can move. If you care about word error rate, compare variants, not backends.

Is WhisperKit still a separate repo?

No, and this is the single thing that breaks older tutorials. WhisperKit is now a library product inside argmaxinc/argmax-oss-swift (6,326 stars, MIT, last pushed August 13, 2026), alongside TTSKit and SpeakerKit. You add one Swift package — https://github.com/argmaxinc/argmax-oss-swift.git, from 0.9.0 — and select the WhisperKit product. The Homebrew formula is still brew install whisperkit-cli, but the in-repo CLI target is now argmax-cli. WhisperKit itself needs macOS 14+ and Xcode 16+.

Can you run Parakeet on Apple Silicon?

Yes, through two independent ports. parakeet-mlx (Apache-2.0, 974 stars) runs it on MLX from Python — pip install parakeet-mlx, then parakeet-mlx audio.mp3, defaulting to mlx-community/parakeet-tdt-0.6b-v3. FluidAudio (Apache-2.0, 2,655 stars) runs Core ML conversions on the Apple Neural Engine from Swift, and is the engine behind a long list of Mac dictation apps. mlx-audio also ships a Parakeet STT backend. Neither route needs CUDA or a Python-free machine to be fast.

How much does thermal throttling cost on a fanless MacBook Air?

We do not have an Air and a Max side by side, so we are not going to invent a number. What we can tell you is where the cost lands: ANE-resident models (FluidAudio, whisper.cpp with WHISPER_COREML=1) draw far less sustained power than a Metal GPU path, which is why the fanless machines favour them on long files. Measure it on your own machine with the 60-minute loop in the benchmark section below and watch sudo powermetrics --samplers smc during the run — if the first ten minutes are much faster than the last ten, you found your throttle.

Which model should I actually download first?

For English dictation and meeting audio, start with Parakeet TDT v3 (mlx-community/parakeet-tdt-0.6b-v3) — it is 0.6B parameters, which at bf16 is roughly 1.2GB of weights before activations, so it fits comfortably on an 8GB Air. For anything outside the 25 languages Parakeet v3 covers, or for translation, use WhisperKit's large-v3-v20240930_626MB. Skip Whisper tiny except for debugging — and note that mlx-whisper defaults to mlx-community/whisper-tiny, which is the most common reason someone reports "MLX Whisper is fast but useless".

Ready to Go Beyond Tutorials?

25 structured courses with hands-on chapters - build RAG chatbots, AI agents, and ML pipelines on your own hardware.

Bonus kit

Ollama Docker Templates

10 one-command Docker stacks for local models — get the LLM half of your transcription pipeline serving in minutes. Included with paid plans, or free after subscribing to both Local AI Master and Little AI Master on YouTube.

See Plans →

Was this helpful?

📅 Published: October 4, 2026🔄 Last Updated: October 4, 2026✓ Manually Reviewed
LM

Written by the Local AI Master Team

The team behind Local AI Master

We build Local AI Master around practical, testable local AI workflows: model selection, hardware planning, RAG systems, agents, and MLOps. The goal is to turn scattered tutorials into a structured learning path you can follow on your own hardware.

✓ Local AI Curriculum✓ Hands-On Projects✓ Open Source Contributor
📚
Free · no account required

Grab the AI Starter Kit — career roadmap, cheat sheet, setup guide

No spam. Unsubscribe with one click.

🎯
AI Learning Path

Voice working locally? Build the whole pipeline.

Whisper, TTS, and voice cloning wired into real projects — hands-on courses. First chapter free, no card.

Or own it for life — Lifetime $149 $599, pay once
Free Tools & Calculators