Transcribe a Meeting With Speaker Names, Locally
Want to go deeper than this article?
Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.
Go from reading about AI to building with AI 25 structured courses. Hands-on projects. Runs on your machine. Start free.
Short answer: use WhisperX with --diarize --min_speakers N --max_speakers N and pyannote's speaker-diarization-community-1 — the maintainer-published accuracy is 11.2% DER on VoxConverse and 17.0% on AMI headset audio. If you do not want Python or a HuggingFace token anywhere near this, install aTrain from the Microsoft Store or Flathub and skip to the no-Python section. Diarization on a 2-hour file is minutes, not hours — pyannote reports community-1 at ~31 seconds per hour of audio on an H100. It is the install that eats your afternoon, and there are exactly four errors responsible.
One thing to set expectations on before you start: nobody has a local pipeline that gets speaker labels perfect on a real meeting. A 17% diarization error rate means roughly one word in six lands on the wrong nameplate on conference-room audio. You are choosing between "good enough to skim and fix" and "useless", not between perfect and imperfect. The single biggest lever you control is the recording, not the model.
The 60-Second Answer
Three tools cover essentially everyone, and the deciding question is what you are willing to install.
- You have Python and a GPU, and you want the best output: WhisperX with
--diarize. It gives you word-level timestamps aligned to speaker turns, which is what makes a transcript actually readable. Costs you a HuggingFace token and a gated-model click. - You want no HuggingFace account: MahmoudAshraf97/whisper-diarization (5.6k stars). It uses NVIDIA NeMo's MarbleNet for voice activity detection and TitaNet for speaker embeddings, plus Demucs to strip music before transcription. No gated repos in the path.
- You want to double-click an app: aTrain. GUI, fully offline, no token, no Python.
Everything else on this page is why, and how to get past the errors.
Reading articles is good. Building is better.
Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.
What Accuracy to Expect
Diarization error rate (DER) on the standard benchmark sets runs 11-20% for the current open model — and the largest single error term on meeting audio is speaker confusion, not missed speech.
These are pyannote's own published numbers from the model cards, comparing the current open community-1 pipeline against the legacy 3.1 it replaces. Lower is better.
| Benchmark | Legacy 3.1 DER | community-1 DER | What it is |
|---|---|---|---|
| VoxConverse | 11.2% | 11.2% | Broadcast / YouTube-style audio |
| AISHELL-4 | 12.2% | 11.7% | Mandarin conference meetings |
| AMI (IHM) | 18.8% | 17.0% | Meetings, headset mics |
| DIHARD 3 | 21.4% | 20.2% | Deliberately hard mixed domains |
| AliMeeting | 24.5% | 20.3% | Mandarin far-field meetings |
Source: pyannote/speaker-diarization-community-1 and pyannote/speaker-diarization-3.1 model cards on HuggingFace.
Two things worth reading off that table. First, the jump from 3.1 to community-1 is real but modest — 1-4 points on most sets. If you already have a working 3.1 pipeline pinned and stable, upgrading is an improvement, not a rescue. Second, your audio type matters more than your model choice. The same pipeline scores 11.2% on clean broadcast audio and 24-ish on far-field meeting rooms. Moving from a laptop mic in the middle of the table to a per-person headset or a lapel recorder will beat any model swap you can make.
For reference on the transcription half of the job, the error rates and speed of the ASR models themselves are covered in best local speech-to-text models and the Parakeet vs Whisper comparison.
A note on what we did and did not measure
We did not run a controlled five-way bake-off on a hand-labelled 2-hour file for this page, and we are not going to pretend otherwise — a DER number is only meaningful against ground truth annotated to a defined collar, and a single in-house recording would tell you about our meeting room, not yours. Every accuracy figure above is attributed to the maintainer that published it. The runtime figures below are likewise attributed. Where we give an opinion, it is about workflow and failure modes, which is the part we have actually lived through.
Pick Your Tool
If you can only remember one rule: WhisperX for quality, whisper-diarization to dodge the HuggingFace gate, aTrain to dodge Python entirely.
| Tool | Diarizer | HF token? | Python? | Best for |
|---|---|---|---|---|
| WhisperX (23.6k stars) | pyannote community-1 | Yes (gated) | Yes | Word-level timestamps + speaker turns, batched GPU |
| whisper-diarization (5.6k) | NeMo MarbleNet + TitaNet | No | Yes | Avoiding gated repos; music/vocal separation built in |
| aTrain (1.2k) | pyannote via faster-whisper app | No | No | Researchers, journalists, anyone who wants a GUI |
| MOSS-Transcribe-Diarize (1.5k) | End-to-end, single model | No | Yes | One 0.9B model doing ASR + diarization jointly |
| sherpa-onnx | pyannote-segmentation-3.0 ONNX + embeddings | No | Optional | Embedding into an app; C/C++/Rust/Go/Java/Swift bindings |
Star counts as of August 2026.
Two of those deserve a sentence more.
MOSS-Transcribe-Diarize is the interesting outlier: a 0.9B Apache-2.0 model (Whisper-Medium encoder, Qwen3-0.6B-style decoder) that emits speaker-labelled, timestamped text in one pass — output looks like [start][S01]text[end] — instead of bolting a diarizer onto an ASR model. Its published numbers are character error rate and speaker-attributed CER rather than DER, so they do not slot into the table above: 5.97 CER / 7.37 cpCER on podcasts, 24.86 / 22.17 on AliMeeting. It supports 50+ languages. Treat it as promising rather than settled — it is a young project and its VRAM requirement is not documented.
sherpa-onnx is the one to reach for if the destination is not a transcript file but a feature in your own software. It runs the ONNX-converted pyannote segmentation model plus a separate embedding extractor (3D-Speaker or NeMo), ships int8 quantised variants for CPU, and has bindings for C, C++, C#, Go, Java, JavaScript, Kotlin, Rust and Swift. No Python runtime required at deploy time.
The WhisperX + pyannote Path
Install WhisperX, accept the gated model on HuggingFace, then always pass a speaker count.
pip install whisperx # or: uvx whisperx
On Linux and Windows with an NVIDIA card, WhisperX's docs call for a CUDA 12.8 install. Then get access to the diarization model, which is the step that trips people:
- Sign in to HuggingFace and open
huggingface.co/pyannote/speaker-diarization-community-1. - Click through and accept the user conditions. The repo is gated; without this the download fails no matter what token you use.
- Create a token at
hf.co/settings/tokens. It must start withhf_.
Then run it:
whisperx interview.wav \
--model large-v3 \
--diarize \
--min_speakers 2 --max_speakers 2 \
--hf_token hf_xxxxxxxxxxxxxxxxxxxx
Set --min_speakers and --max_speakers every single time. For a two-person interview, set both to 2. This is the highest-value flag on the command line and most people never touch it. WhisperX's own README documents it as optional; in practice, on anything longer than about twenty minutes, leaving it out is the main reason a fourth speaker appears.
If you would rather drive pyannote directly — useful when you already have a transcript and only need the speaker turns — the current API is short:
import torch
from pyannote.audio import Pipeline
pipeline = Pipeline.from_pretrained(
"pyannote/speaker-diarization-community-1",
token="hf_xxxxxxxxxxxxxxxxxxxx")
pipeline.to(torch.device("cuda"))
output = pipeline("interview.wav", num_speakers=2)
for turn, speaker in output.speaker_diarization:
print(f"start={turn.start:.1f}s stop={turn.end:.1f}s {speaker}")
WhisperX's word-level alignment is the reason to prefer it over raw pyannote for transcripts: it aligns Whisper's output with wav2vec2 so each word carries a timestamp, and speaker turns can then be mapped onto words rather than onto whole 30-second Whisper segments. That is the difference between a transcript where the speaker changes mid-sentence correctly and one where it changes at an arbitrary chunk boundary. The tool-level detail lives in our WhisperX guide.
Run this on your own machine and stop paying every month
Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.
Four Errors That Stop Everyone
Ranked by how often each one is the actual problem.
1. "Could not download ... It might be because the repository is private or gated"
The most common failure by a wide margin. pyannote's downloader raises, verbatim:
Could not download {asset_file_name} from {model_id}. It might be because the repository is private or gated:
Cause: one of three things — you never clicked accept on the model page, you are passing a token that is not a HuggingFace token, or the token lacks read access. pyannote's HuggingFace helper explicitly filters out any token that does not begin with hf_ (so a pyannoteAI API key pasted here will be silently treated as absent).
Fix, in order: open the model page while signed in and accept the conditions → regenerate a read token at hf.co/settings/tokens → confirm it starts with hf_ → pass it as token= (pyannote 4.x) or use_auth_token= (3.x). Note the two model repos are gated separately; accepting for 3.1 does not grant community-1.
2. _pickle.UnpicklingError: Weights only load failed
Symptom: the pipeline downloads fine, then dies while loading the VAD checkpoint. The error text mentions an unsupported global — commonly omegaconf.listconfig.ListConfig, and on some stacks typing.Any.
Cause: PyTorch 2.6 flipped the default of weights_only in torch.load() from False to True for security. Older pyannote/WhisperX checkpoints contain pickled objects that are not on the default allowlist, so they now refuse to deserialise.
Fix: update pyannote.audio and WhisperX to current releases first — this is patched upstream and a version bump resolves it for most people. If you are pinned to an old stack you cannot move, either allowlist the offending global before loading:
import torch, omegaconf
torch.serialization.add_safe_globals([omegaconf.listconfig.ListConfig])
…or pin torch<2.6. Prefer the upgrade. Downgrading torch usually starts a second fight with your CUDA build.
3. The pyannote 3 → 4 API break
Symptom: code copied from a 2024 tutorial throws an attribute error on the result object, or use_auth_token is rejected.
Cause: pyannote.audio 4.x changed both the keyword (use_auth_token → token) and the shape of the result. The 3.1 model card returns a diarization object you iterate directly; the community-1 card documents iterating output.speaker_diarization. Old snippets will not run against the new pipeline and vice versa.
Fix: decide which pair you are on and keep it consistent. speaker-diarization-3.1 wants pyannote.audio 3.1 or higher; speaker-diarization-community-1 is documented as compatible with pyannote.audio 4.x. Mixing a 4.x install with a 3.1-era snippet is the single most common version-compatibility complaint we see land on this cluster.
4. ffmpeg missing
Symptom: a file-not-found or decode error on an mp3/m4a/mp4 that plays fine everywhere else.
Cause: pyannote lists ffmpeg as a prerequisite and every tool here depends on it for anything that is not a plain WAV.
Fix: install ffmpeg, then — for long files especially — pre-convert once rather than making the pipeline do it repeatedly:
ffmpeg -i meeting.m4a -ac 1 -ar 16000 -c:a pcm_s16le meeting.wav
16 kHz mono PCM is what these models want anyway. Doing it up front also means a crash three hours in does not cost you the decode.
When It Merges or Invents Speakers
Work this list top-down; the first item resolves the majority of cases.
- You did not constrain the speaker count. Pass
--min_speakers/--max_speakers(WhisperX) ornum_speakers(pyannote). Clustering has to guess otherwise, and on a two-hour file it guesses wrong more often than it guesses right. This is free and takes ten seconds. - The recording is a mono downmix of a multi-mic setup. If your recorder produced separate channels per participant, diarize per channel instead — a channel is a speaker label, with zero error. This is the only route to genuinely perfect labels and it is criminally underused. Zoom, OBS and most field recorders can do it.
- Crosstalk. Look at pyannote's own benchmark breakdown: on meeting corpora the confusion term is the largest component of the error, larger than missed speech or false alarm. Two people talking over each other is a known weak point of embedding-based diarization, not a misconfiguration. If the interview is important, ask people not to interrupt — a process fix beats a model fix here.
- Similar voices. Two men of similar age and accent, or two women, will collapse into one cluster more often than a mixed pair. There is no flag for this. Passing the correct
num_speakersforces a split, which sometimes helps and sometimes splits at the wrong place — check a couple of turns before trusting it. - Bad segments getting their own label. Laughter, a cough, a door, a phone notification. These produce embeddings that match nothing and become "SPEAKER_03". Constraining the count is again the fix; whisper-diarization's Demucs source-separation step (on by default, disable with
--no-stem) also helps when there is background music.
What we cannot tell you from a desk: whether your two voices are separable. Take the first ten minutes of the file, run it with the correct speaker count, and read it. Ten minutes of checking beats two hours of a bad transcript.
The No-Python Path
aTrain is a desktop application that does transcription plus diarization fully offline, with no HuggingFace token and no Python.
Install it from the Microsoft Store on Windows or Flathub on Linux. GPU acceleration is documented for Windows and Debian-based Linux only, so a Mac runs it CPU-only. Under the hood it runs OpenAI's Whisper via faster-whisper for transcription and pyannote.audio for the speaker detection — the same engine as the Python route, wrapped so the gated-model dance and dependency pinning are handled for you. It supports 99 languages, processes everything on-device (its maintainers pitch this as the GDPR-compliant angle, which is the same reason most of our readers are here), and exports to formats that open in MAXQDA, ATLAS.ti and NVivo. Licence is AGPL-3.0.
Published speed figures from the project, on a 22-minute file with speaker detection enabled:
| Hardware | Time for 22 min of audio |
|---|---|
| RTX 2080 Ti (largest model) | 1:44 |
| Ryzen 6850U, CPU only | 13-26 min depending on model size |
Source: aTrain repository documentation.
Scale that linearly and a 2-hour recording is roughly 10 minutes on a modest GPU or 1-2.5 hours CPU-only — the linear extrapolation is ours, not theirs, but it is the right order of magnitude for planning. If you have no GPU, start it before you go to bed. For what a GPU actually buys you across voice workloads generally, see voice AI VRAM requirements by GPU.
There is a real trade-off. A GUI gives you no way to pin versions, script a batch of 40 interviews, or slot the output into a pipeline. If you transcribe one thing a month, aTrain is obviously correct. If you transcribe forty, take the Python pain once.
The 2-Hour File Problem
Diarization scales fine with length; transcription and memory do not.
pyannote reports community-1 processing audio at ~31 seconds per hour on an NVIDIA H100 (their paid precision-2 service does it in 14 seconds — noted for context, it is not something you run locally). Even scaling that down several times for a consumer GPU, the diarization pass on a 2-hour file is a coffee break, not an evening.
The parts that actually bite on long files:
- Transcription dominates the clock. WhisperX's headline claim is 70x realtime with large-v2 through batched inference, which is why it is the right ASR front-end for long audio. Without batching you are in a much slower regime — the faster-whisper guide covers the engine underneath.
- VRAM, if you run models in parallel. whisper-diarization ships a
diarize_parallel.pythat runs NeMo and Whisper simultaneously and documents a ≥10GB VRAM requirement for it. The sequentialdiarize.pyis the safer default on 8GB cards. - Crashes cost more. Convert to 16 kHz mono WAV once, up front, and if the tool supports resuming or chunking, use it. Losing 90 minutes of work to an OOM at the end is the classic way this goes wrong.
- Disk and RAM for the aligned output. Word-level timestamps on two hours of speech is a large JSON. It is not a problem, but do not be surprised by the file size.
For the wider workflow — where the transcript goes, summarisation, exporting to subtitles — see local AI meeting transcription and local AI subtitles with Whisper.
Verdict
- WhisperX +
speaker-diarization-community-1is the default. Best output quality, word-level speaker mapping, and the widest community when something breaks. Budget an hour for the install the first time. - Always pass the speaker count. It is one flag and it removes the most common failure people write in about.
- If you never want to see a HuggingFace gate, use whisper-diarization. NeMo models, no accept-the-terms step, source separation included.
- If you never want to see a terminal, use aTrain. Fully offline, no token, and its published GPU timing suggests a 2-hour file in around ten minutes.
- Fix the recording before you fix the model. The published DER gap between broadcast audio and far-field meeting audio is bigger than the gap between any two of these tools. Per-speaker channels, if you can get them, beat diarization outright.
Expect to edit the result. Local diarization in 2026 gets you a transcript you can skim and correct in fifteen minutes instead of typing for four hours. That is the honest win, and it is a large one.
Sources
- pyannote/speaker-diarization-community-1 — DER benchmarks, gating, API snippet, speaker-count arguments
- pyannote/speaker-diarization-3.1 — legacy DER benchmarks with error-component breakdown
- pyannote/pyannote-audio — install, H100 throughput figures, gated-repo error string
- m-bain/whisperX — install,
--diarize, speaker-count flags, 70x realtime claim - MahmoudAshraf97/whisper-diarization — NeMo MarbleNet/TitaNet pipeline, parallel VRAM requirement
- aTrainTranscription/aTrain — GUI, offline operation, published timings, licence
- OpenMOSS/MOSS-Transcribe-Diarize — end-to-end model, CER/cpCER results
- sherpa-onnx speaker diarization docs — ONNX segmentation + embedding models, language bindings
FAQ
Go from reading about AI to building with AI
25 structured courses. Hands-on projects. Runs on your machine. Start free.
Liked this? 25 full AI courses are waiting.
From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.
Build Real AI on Your Machine
RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.
Want the structured version?
Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.
Keep going
Comments (0)
No comments yet. Be the first to share your thoughts!