★ Reading this for free? Get 25 structured AI courses + per-chapter AI tutor — the first chapter of every course free, no card.Start free in 30 secondsOr own every course: $149 once
Voice

Transcribe a Meeting With Speaker Names, Locally

October 4, 2026
14 min read
LocalAimaster Research Team

Want to go deeper than this article?

Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.

📚AI Learning Path

Go from reading about AI to building with AI 25 structured courses. Hands-on projects. Runs on your machine. Start free.

Start free
Or own it for life — Lifetime $149, pay once

Short answer: use WhisperX with --diarize --min_speakers N --max_speakers N and pyannote's speaker-diarization-community-1 — the maintainer-published accuracy is 11.2% DER on VoxConverse and 17.0% on AMI headset audio. If you do not want Python or a HuggingFace token anywhere near this, install aTrain from the Microsoft Store or Flathub and skip to the no-Python section. Diarization on a 2-hour file is minutes, not hours — pyannote reports community-1 at ~31 seconds per hour of audio on an H100. It is the install that eats your afternoon, and there are exactly four errors responsible.

One thing to set expectations on before you start: nobody has a local pipeline that gets speaker labels perfect on a real meeting. A 17% diarization error rate means roughly one word in six lands on the wrong nameplate on conference-room audio. You are choosing between "good enough to skim and fix" and "useless", not between perfect and imperfect. The single biggest lever you control is the recording, not the model.


The 60-Second Answer

Three tools cover essentially everyone, and the deciding question is what you are willing to install.

  • You have Python and a GPU, and you want the best output: WhisperX with --diarize. It gives you word-level timestamps aligned to speaker turns, which is what makes a transcript actually readable. Costs you a HuggingFace token and a gated-model click.
  • You want no HuggingFace account: MahmoudAshraf97/whisper-diarization (5.6k stars). It uses NVIDIA NeMo's MarbleNet for voice activity detection and TitaNet for speaker embeddings, plus Demucs to strip music before transcription. No gated repos in the path.
  • You want to double-click an app: aTrain. GUI, fully offline, no token, no Python.

Everything else on this page is why, and how to get past the errors.


Reading articles is good. Building is better.

Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.

What Accuracy to Expect

Diarization error rate (DER) on the standard benchmark sets runs 11-20% for the current open model — and the largest single error term on meeting audio is speaker confusion, not missed speech.

These are pyannote's own published numbers from the model cards, comparing the current open community-1 pipeline against the legacy 3.1 it replaces. Lower is better.

BenchmarkLegacy 3.1 DERcommunity-1 DERWhat it is
VoxConverse11.2%11.2%Broadcast / YouTube-style audio
AISHELL-412.2%11.7%Mandarin conference meetings
AMI (IHM)18.8%17.0%Meetings, headset mics
DIHARD 321.4%20.2%Deliberately hard mixed domains
AliMeeting24.5%20.3%Mandarin far-field meetings

Source: pyannote/speaker-diarization-community-1 and pyannote/speaker-diarization-3.1 model cards on HuggingFace.

Two things worth reading off that table. First, the jump from 3.1 to community-1 is real but modest — 1-4 points on most sets. If you already have a working 3.1 pipeline pinned and stable, upgrading is an improvement, not a rescue. Second, your audio type matters more than your model choice. The same pipeline scores 11.2% on clean broadcast audio and 24-ish on far-field meeting rooms. Moving from a laptop mic in the middle of the table to a per-person headset or a lapel recorder will beat any model swap you can make.

For reference on the transcription half of the job, the error rates and speed of the ASR models themselves are covered in best local speech-to-text models and the Parakeet vs Whisper comparison.

A note on what we did and did not measure

We did not run a controlled five-way bake-off on a hand-labelled 2-hour file for this page, and we are not going to pretend otherwise — a DER number is only meaningful against ground truth annotated to a defined collar, and a single in-house recording would tell you about our meeting room, not yours. Every accuracy figure above is attributed to the maintainer that published it. The runtime figures below are likewise attributed. Where we give an opinion, it is about workflow and failure modes, which is the part we have actually lived through.


Pick Your Tool

If you can only remember one rule: WhisperX for quality, whisper-diarization to dodge the HuggingFace gate, aTrain to dodge Python entirely.

ToolDiarizerHF token?Python?Best for
WhisperX (23.6k stars)pyannote community-1Yes (gated)YesWord-level timestamps + speaker turns, batched GPU
whisper-diarization (5.6k)NeMo MarbleNet + TitaNetNoYesAvoiding gated repos; music/vocal separation built in
aTrain (1.2k)pyannote via faster-whisper appNoNoResearchers, journalists, anyone who wants a GUI
MOSS-Transcribe-Diarize (1.5k)End-to-end, single modelNoYesOne 0.9B model doing ASR + diarization jointly
sherpa-onnxpyannote-segmentation-3.0 ONNX + embeddingsNoOptionalEmbedding into an app; C/C++/Rust/Go/Java/Swift bindings

Star counts as of August 2026.

Two of those deserve a sentence more.

MOSS-Transcribe-Diarize is the interesting outlier: a 0.9B Apache-2.0 model (Whisper-Medium encoder, Qwen3-0.6B-style decoder) that emits speaker-labelled, timestamped text in one pass — output looks like [start][S01]text[end] — instead of bolting a diarizer onto an ASR model. Its published numbers are character error rate and speaker-attributed CER rather than DER, so they do not slot into the table above: 5.97 CER / 7.37 cpCER on podcasts, 24.86 / 22.17 on AliMeeting. It supports 50+ languages. Treat it as promising rather than settled — it is a young project and its VRAM requirement is not documented.

sherpa-onnx is the one to reach for if the destination is not a transcript file but a feature in your own software. It runs the ONNX-converted pyannote segmentation model plus a separate embedding extractor (3D-Speaker or NeMo), ships int8 quantised variants for CPU, and has bindings for C, C++, C#, Go, Java, JavaScript, Kotlin, Rust and Swift. No Python runtime required at deploy time.


The WhisperX + pyannote Path

Install WhisperX, accept the gated model on HuggingFace, then always pass a speaker count.

pip install whisperx        # or: uvx whisperx

On Linux and Windows with an NVIDIA card, WhisperX's docs call for a CUDA 12.8 install. Then get access to the diarization model, which is the step that trips people:

  1. Sign in to HuggingFace and open huggingface.co/pyannote/speaker-diarization-community-1.
  2. Click through and accept the user conditions. The repo is gated; without this the download fails no matter what token you use.
  3. Create a token at hf.co/settings/tokens. It must start with hf_.

Then run it:

whisperx interview.wav \
  --model large-v3 \
  --diarize \
  --min_speakers 2 --max_speakers 2 \
  --hf_token hf_xxxxxxxxxxxxxxxxxxxx

Set --min_speakers and --max_speakers every single time. For a two-person interview, set both to 2. This is the highest-value flag on the command line and most people never touch it. WhisperX's own README documents it as optional; in practice, on anything longer than about twenty minutes, leaving it out is the main reason a fourth speaker appears.

If you would rather drive pyannote directly — useful when you already have a transcript and only need the speaker turns — the current API is short:

import torch
from pyannote.audio import Pipeline

pipeline = Pipeline.from_pretrained(
    "pyannote/speaker-diarization-community-1",
    token="hf_xxxxxxxxxxxxxxxxxxxx")
pipeline.to(torch.device("cuda"))

output = pipeline("interview.wav", num_speakers=2)
for turn, speaker in output.speaker_diarization:
    print(f"start={turn.start:.1f}s stop={turn.end:.1f}s {speaker}")

WhisperX's word-level alignment is the reason to prefer it over raw pyannote for transcripts: it aligns Whisper's output with wav2vec2 so each word carries a timestamp, and speaker turns can then be mapped onto words rather than onto whole 30-second Whisper segments. That is the difference between a transcript where the speaker changes mid-sentence correctly and one where it changes at an arbitrary chunk boundary. The tool-level detail lives in our WhisperX guide.


Own it instead of renting it

Run this on your own machine and stop paying every month

Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.

Four Errors That Stop Everyone

Ranked by how often each one is the actual problem.

1. "Could not download ... It might be because the repository is private or gated"

The most common failure by a wide margin. pyannote's downloader raises, verbatim:

Could not download {asset_file_name} from {model_id}. It might be because the repository is private or gated:

Cause: one of three things — you never clicked accept on the model page, you are passing a token that is not a HuggingFace token, or the token lacks read access. pyannote's HuggingFace helper explicitly filters out any token that does not begin with hf_ (so a pyannoteAI API key pasted here will be silently treated as absent).

Fix, in order: open the model page while signed in and accept the conditions → regenerate a read token at hf.co/settings/tokens → confirm it starts with hf_ → pass it as token= (pyannote 4.x) or use_auth_token= (3.x). Note the two model repos are gated separately; accepting for 3.1 does not grant community-1.

2. _pickle.UnpicklingError: Weights only load failed

Symptom: the pipeline downloads fine, then dies while loading the VAD checkpoint. The error text mentions an unsupported global — commonly omegaconf.listconfig.ListConfig, and on some stacks typing.Any.

Cause: PyTorch 2.6 flipped the default of weights_only in torch.load() from False to True for security. Older pyannote/WhisperX checkpoints contain pickled objects that are not on the default allowlist, so they now refuse to deserialise.

Fix: update pyannote.audio and WhisperX to current releases first — this is patched upstream and a version bump resolves it for most people. If you are pinned to an old stack you cannot move, either allowlist the offending global before loading:

import torch, omegaconf
torch.serialization.add_safe_globals([omegaconf.listconfig.ListConfig])

…or pin torch<2.6. Prefer the upgrade. Downgrading torch usually starts a second fight with your CUDA build.

3. The pyannote 3 → 4 API break

Symptom: code copied from a 2024 tutorial throws an attribute error on the result object, or use_auth_token is rejected.

Cause: pyannote.audio 4.x changed both the keyword (use_auth_token → token) and the shape of the result. The 3.1 model card returns a diarization object you iterate directly; the community-1 card documents iterating output.speaker_diarization. Old snippets will not run against the new pipeline and vice versa.

Fix: decide which pair you are on and keep it consistent. speaker-diarization-3.1 wants pyannote.audio 3.1 or higher; speaker-diarization-community-1 is documented as compatible with pyannote.audio 4.x. Mixing a 4.x install with a 3.1-era snippet is the single most common version-compatibility complaint we see land on this cluster.

4. ffmpeg missing

Symptom: a file-not-found or decode error on an mp3/m4a/mp4 that plays fine everywhere else.

Cause: pyannote lists ffmpeg as a prerequisite and every tool here depends on it for anything that is not a plain WAV.

Fix: install ffmpeg, then — for long files especially — pre-convert once rather than making the pipeline do it repeatedly:

ffmpeg -i meeting.m4a -ac 1 -ar 16000 -c:a pcm_s16le meeting.wav

16 kHz mono PCM is what these models want anyway. Doing it up front also means a crash three hours in does not cost you the decode.


When It Merges or Invents Speakers

Work this list top-down; the first item resolves the majority of cases.

  1. You did not constrain the speaker count. Pass --min_speakers/--max_speakers (WhisperX) or num_speakers (pyannote). Clustering has to guess otherwise, and on a two-hour file it guesses wrong more often than it guesses right. This is free and takes ten seconds.
  2. The recording is a mono downmix of a multi-mic setup. If your recorder produced separate channels per participant, diarize per channel instead — a channel is a speaker label, with zero error. This is the only route to genuinely perfect labels and it is criminally underused. Zoom, OBS and most field recorders can do it.
  3. Crosstalk. Look at pyannote's own benchmark breakdown: on meeting corpora the confusion term is the largest component of the error, larger than missed speech or false alarm. Two people talking over each other is a known weak point of embedding-based diarization, not a misconfiguration. If the interview is important, ask people not to interrupt — a process fix beats a model fix here.
  4. Similar voices. Two men of similar age and accent, or two women, will collapse into one cluster more often than a mixed pair. There is no flag for this. Passing the correct num_speakers forces a split, which sometimes helps and sometimes splits at the wrong place — check a couple of turns before trusting it.
  5. Bad segments getting their own label. Laughter, a cough, a door, a phone notification. These produce embeddings that match nothing and become "SPEAKER_03". Constraining the count is again the fix; whisper-diarization's Demucs source-separation step (on by default, disable with --no-stem) also helps when there is background music.

What we cannot tell you from a desk: whether your two voices are separable. Take the first ten minutes of the file, run it with the correct speaker count, and read it. Ten minutes of checking beats two hours of a bad transcript.


The No-Python Path

aTrain is a desktop application that does transcription plus diarization fully offline, with no HuggingFace token and no Python.

Install it from the Microsoft Store on Windows or Flathub on Linux. GPU acceleration is documented for Windows and Debian-based Linux only, so a Mac runs it CPU-only. Under the hood it runs OpenAI's Whisper via faster-whisper for transcription and pyannote.audio for the speaker detection — the same engine as the Python route, wrapped so the gated-model dance and dependency pinning are handled for you. It supports 99 languages, processes everything on-device (its maintainers pitch this as the GDPR-compliant angle, which is the same reason most of our readers are here), and exports to formats that open in MAXQDA, ATLAS.ti and NVivo. Licence is AGPL-3.0.

Published speed figures from the project, on a 22-minute file with speaker detection enabled:

HardwareTime for 22 min of audio
RTX 2080 Ti (largest model)1:44
Ryzen 6850U, CPU only13-26 min depending on model size

Source: aTrain repository documentation.

Scale that linearly and a 2-hour recording is roughly 10 minutes on a modest GPU or 1-2.5 hours CPU-only — the linear extrapolation is ours, not theirs, but it is the right order of magnitude for planning. If you have no GPU, start it before you go to bed. For what a GPU actually buys you across voice workloads generally, see voice AI VRAM requirements by GPU.

There is a real trade-off. A GUI gives you no way to pin versions, script a batch of 40 interviews, or slot the output into a pipeline. If you transcribe one thing a month, aTrain is obviously correct. If you transcribe forty, take the Python pain once.


The 2-Hour File Problem

Diarization scales fine with length; transcription and memory do not.

pyannote reports community-1 processing audio at ~31 seconds per hour on an NVIDIA H100 (their paid precision-2 service does it in 14 seconds — noted for context, it is not something you run locally). Even scaling that down several times for a consumer GPU, the diarization pass on a 2-hour file is a coffee break, not an evening.

The parts that actually bite on long files:

  • Transcription dominates the clock. WhisperX's headline claim is 70x realtime with large-v2 through batched inference, which is why it is the right ASR front-end for long audio. Without batching you are in a much slower regime — the faster-whisper guide covers the engine underneath.
  • VRAM, if you run models in parallel. whisper-diarization ships a diarize_parallel.py that runs NeMo and Whisper simultaneously and documents a ≥10GB VRAM requirement for it. The sequential diarize.py is the safer default on 8GB cards.
  • Crashes cost more. Convert to 16 kHz mono WAV once, up front, and if the tool supports resuming or chunking, use it. Losing 90 minutes of work to an OOM at the end is the classic way this goes wrong.
  • Disk and RAM for the aligned output. Word-level timestamps on two hours of speech is a large JSON. It is not a problem, but do not be surprised by the file size.

For the wider workflow — where the transcript goes, summarisation, exporting to subtitles — see local AI meeting transcription and local AI subtitles with Whisper.


Verdict

  1. WhisperX + speaker-diarization-community-1 is the default. Best output quality, word-level speaker mapping, and the widest community when something breaks. Budget an hour for the install the first time.
  2. Always pass the speaker count. It is one flag and it removes the most common failure people write in about.
  3. If you never want to see a HuggingFace gate, use whisper-diarization. NeMo models, no accept-the-terms step, source separation included.
  4. If you never want to see a terminal, use aTrain. Fully offline, no token, and its published GPU timing suggests a 2-hour file in around ten minutes.
  5. Fix the recording before you fix the model. The published DER gap between broadcast audio and far-field meeting audio is bigger than the gap between any two of these tools. Per-speaker channels, if you can get them, beat diarization outright.

Expect to edit the result. Local diarization in 2026 gets you a transcript you can skim and correct in fifteen minutes instead of typing for four hours. That is the honest win, and it is a large one.


Sources


FAQ

🎯
AI Learning Path

Go from reading about AI to building with AI

25 structured courses. Hands-on projects. Runs on your machine. Start free.

Or own it for life — Lifetime $149 $599, pay once

Liked this? 25 full AI courses are waiting.

From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.

Reading now
Join the discussion
TagsSpeaker DiarizationWhisperXpyannoteTranscriptionOffline AIMeetings

LocalAimaster Research Team

Local AI Master writes hands-on courses and hardware guides for running AI on machines you own. Content is checked against current releases and corrected when readers tell us it is wrong.

Build Real AI on Your Machine

RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.

Want the structured version?

Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.

AI Learning Path

Comments (0)

No comments yet. Be the first to share your thoughts!

What is the most accurate local speaker diarization model right now?

For open weights you can run offline, pyannote's speaker-diarization-community-1 is the current reference. Its own model card reports diarization error rates of 11.2% on VoxConverse, 17.0% on AMI (IHM headset audio) and 20.2% on DIHARD 3 — improvements over the older speaker-diarization-3.1 on every dataset except VoxConverse, where they tie. Those are the maintainer's published numbers on academic benchmark sets, not a promise about your meeting recording; a phone speaker in a reverberant room will do worse than any of them.

Do I need a HuggingFace token for speaker diarization?

For anything built on pyannote, yes. Both pyannote/speaker-diarization-community-1 and the older 3.1 are gated repos: you must be signed in to HuggingFace, accept the user conditions on the model page, then pass a token that starts with hf_ to Pipeline.from_pretrained(). pyannote explicitly rejects tokens that do not start with hf_. If you want to avoid the token entirely, use whisper-diarization (NVIDIA NeMo models, no gate) or aTrain (bundles everything, no HuggingFace account needed).

Why does my transcript merge two speakers into one?

In order of how often it is the real cause: (1) you did not tell it how many speakers there are — pass --min_speakers and --max_speakers, or num_speakers to pyannote directly; (2) the two voices are genuinely similar and the clustering step collapsed them, which is a hard limit of embedding-based diarization; (3) heavy crosstalk, because pyannote reports overlap poorly and the confusion column in its own benchmark is the largest error term on meeting corpora; (4) mono downmix of a multi-channel recording, which throws away the spatial cue that separates people sitting apart.

How long does a 2-hour file take to diarize locally?

Diarization itself is the cheap half. pyannote reports community-1 running at roughly 31 seconds per hour of audio on an NVIDIA H100 — so on a consumer GPU expect single-digit minutes for a 2-hour file, not hours. Transcription dominates. As one concrete published data point, aTrain's maintainers report a 22-minute file with speaker detection taking 1:44 on an RTX 2080 Ti and 13-26 minutes CPU-only on a Ryzen 6850U, which scales to roughly 10 minutes GPU or 1-2.5 hours CPU for a 2-hour recording. Plan for overnight if you have no GPU.

Is there a way to do this without installing Python?

Yes. aTrain is a desktop GUI that runs fully offline, installs from the Microsoft Store on Windows or Flathub on Linux, and needs no Python knowledge and no HuggingFace token. It uses faster-whisper for transcription and pyannote.audio for diarization under the hood, and exports formats that open in MAXQDA, ATLAS.ti and NVivo. It is AGPL-3.0. macOS is CPU-only.

Does knowing the number of speakers actually help?

It removes an entire failure mode. Diarization pipelines estimate speaker count by clustering embeddings, and that estimate is the single most common thing they get wrong on long recordings — a cough, a laugh or a bad segment becomes "Speaker 4". Both pyannote (num_speakers, min_speakers, max_speakers) and WhisperX (--min_speakers, --max_speakers) accept the constraint. If you know it is a two-person interview, say so. It cannot fix a genuinely confused embedding, but it stops the invented speakers.

Ready to Go Beyond Tutorials?

25 structured courses with hands-on chapters - build RAG chatbots, AI agents, and ML pipelines on your own hardware.

Bonus kit

Ollama Docker Templates

10 one-command Docker stacks for local models — skip the dependency fights entirely. Included with paid plans, or free after subscribing to both Local AI Master and Little AI Master on YouTube.

See Plans →

Was this helpful?

📅 Published: October 4, 2026🔄 Last Updated: October 4, 2026✓ Manually Reviewed
LM

Written by the Local AI Master Team

The team behind Local AI Master

We build Local AI Master around practical, testable local AI workflows: model selection, hardware planning, RAG systems, agents, and MLOps. The goal is to turn scattered tutorials into a structured learning path you can follow on your own hardware.

✓ Local AI Curriculum✓ Hands-On Projects✓ Open Source Contributor
📚
Free · no account required

Grab the AI Starter Kit — career roadmap, cheat sheet, setup guide

No spam. Unsubscribe with one click.

🎯
AI Learning Path

Go from reading about AI to building with AI

25 structured courses. Hands-on projects. Runs on your machine. Start free.

Or own it for life — Lifetime $149 $599, pay once
Free Tools & Calculators