Best Local TTS on Windows: No Python, No WSL
Want to go deeper than this article?
Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.
Voice working locally? Build the whole pipeline. Whisper, TTS, and voice cloning wired into real projects — hands-on courses. First chapter free, no card.
Short answer: use sherpa-onnx. Download sherpa-onnx-v1.13.5-win-x64-shared-MT-Release.tar.bz2 (24.5 MB) and kokoro-int8-multi-lang-v1_0.tar.bz2 (131.8 MB), expand both with the tar command Windows already has, and run one .exe. No interpreter, no WSL, no Conda environment, and — the part that matters — no eSpeak NG on your PATH, because the model archive carries its own copy.
That gets you Kokoro's 53 speaker voices, with American and British English plus Chinese as the text languages sherpa-onnx actually supports for this model — its own documentation is explicit that the other speakers exist but the language support was not added. What you do not get is voice cloning, emotion control, or a graphical interface. Those live behind the Python stacks you were trying to escape, and the honest version of this page says so up front rather than at the bottom.
What We Checked, and What We Did Not
Every file size, filename and archive listing below came from downloading the actual release artifact on 18 August 2026 and listing its contents. Nothing here is copied from a project's front page.
Concretely, we pulled sherpa-onnx-v1.13.5-win-x64-shared-MT-Release.tar.bz2 (and its MD sibling), kokoro-int8-multi-lang-v1_0.tar.bz2 and piper_windows_amd64.zip, and read the file trees inside them. That is how we can tell you which executables exist, whether espeak data is bundled, and which model filename to pass on the command line — the three things tutorials get wrong most often.
What we did not do is run those Windows binaries on a clean Windows 11 machine, because we do not have one to run them on. So there are no synthesis-speed numbers on this page and no "it took four minutes to install" claims. Where the outcome depends on hardware we did not test, we tell you what to check instead of inventing a result. Treat the archive contents as verified fact and the runtime behaviour as documented behaviour.
Reading articles is good. Building is better.
Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.
Why Your Install Failed
The two failures that send people to this page are a torch/CUDA wheel that will not build, and a model that loads perfectly and then writes silence. The second one is almost always eSpeak NG.
Here is the mechanism, because knowing it saves you from repeating it across three more projects.
Most neural TTS turns text into phonemes before it turns phonemes into audio. On the Python side, that job usually goes to the phonemizer package, which does not phonemize anything itself — it shells out to an eSpeak NG binary that has to be installed separately and findable on your PATH. On Linux that is one package-manager line. On Windows it is an MSI, a PATH edit, sometimes a PHONEMIZER_ESPEAK_LIBRARY environment variable pointing at libespeak-ng.dll, and a reboot of your shell.
When it is missing, you do not always get a clean error. You get a model that initialises, a progress bar that completes, and a WAV file that is either empty or a fraction of a second long. That is the "it downloaded but produced silence" report.
The torch failure is more honest about itself but costs more time: a TTS repo pins a CUDA-specific torch build, pip resolves something that does not match your driver, and you spend an evening on wheels.
Both problems disappear the same way — by choosing a runtime that ships its phonemizer data inside the model download and its inference engine as a compiled binary. That is precisely what the ONNX-based options do, and it is the entire argument of this page.
The Four Windows Paths
Only one of these is genuinely dependency-free today, and it is not the famous one.
| Path | What you download | Needs Python? | Bundles espeak data? | Status |
|---|---|---|---|---|
| sherpa-onnx (recommended) | 24.5 MB binaries + 131.8 MB model | No | Yes — in the model archive | Active; v1.13.5 published 11 Aug 2026 |
| Piper standalone | piper_windows_amd64.zip, 22.5 MB | No | Yes — espeak-ng.dll + data | Frozen; binary dated 14 Nov 2023, repo archived |
| Piper (maintained) | piper_tts-1.7.0-cp39-abi3-win_amd64.whl, 34.1 MB | Yes | Yes, inside the wheel | Active; v1.7.0, 15 Aug 2026 |
| Kokoro-FastAPI | Docker image | No, but needs Docker Desktop | Handled in the container | Active; v0.8.0, 15 Aug 2026 |
Release tags, dates and asset sizes read from each project's GitHub releases API on 18 August 2026.
Two things in that table are worth pausing on.
First, the Piper row split into two. The project most Windows tutorials point at is not one project any more. More on that below, because the trap is subtle and expensive.
Second, sherpa-onnx is not a TTS project in the way Piper is — it is a general on-device speech toolkit from the next-gen Kaldi team (14,227 stars, Apache-2.0, last pushed 18 August 2026), and text-to-speech is one of a dozen things its binaries do. That is why the archive contains twenty-odd executables when you only want one. It is also why it has the most disciplined release engineering of anything here: it publishes Windows x64, x86 and ARM64 builds, in shared and static flavours, on every release.
sherpa-onnx, Step by Step
Two archives, two tar commands, one .exe invocation. Windows 10 build 1803 and later include tar.exe, so you do not need 7-Zip either.
1. Get the binaries
From the sherpa-onnx releases page, take the shared Windows build in its MT flavour:
sherpa-onnx-v1.13.5-win-x64-shared-MT-Release.tar.bz2 (24.5 MB)
Prefer the MT variant over the MD one. MT is the statically-linked C runtime; MD links dynamically against the Microsoft Visual C++ runtime, and on a machine with no dev tools that redistributable may simply not be installed. Choosing MT removes one thing that can go wrong on a clean box. (Both exist for every release; the MD build we inspected was 20.1 MB.) Ignore the archives whose names contain static- — those link the libraries into every executable and run to 242.7 MB for the same tools.
Inside bin/ you get sherpa-onnx-offline-tts.exe plus onnxruntime.dll and friends — we counted 21 executables in the copy we extracted. Only one of them is the one you want.
2. Get the model
From the same repo's tts-models release:
kokoro-int8-multi-lang-v1_0.tar.bz2 (131.8 MB, 6,460 downloads)
kokoro-multi-lang-v1_0.tar.bz2 (349.4 MB, 194,985 downloads)
The int8 archive is the one to start with — it is a third of the size and it is the same model quantised. The non-quantised archive is far more popular, but mostly because it is the one linked from the docs.
These two archives are not interchangeable on the command line. The int8 archive contains model.int8.onnx; the full one contains model.onnx. Passing the wrong filename is the most common first-run error, and it is the reason we listed the archives rather than trusting the docs.
Everything else the model needs is in the same folder — we counted voices.bin, tokens.txt, lexicon-us-en.txt, lexicon-gb-en.txt, lexicon-zh.txt, a dict/ folder, and an espeak-ng-data/ directory with 392 entries. That last one is why this path cannot hit the silent-phonemizer failure: the runtime is pointed at bundled data, not at your PATH.
3. Expand and run
tar -xf sherpa-onnx-v1.13.5-win-x64-shared-MT-Release.tar.bz2
tar -xf kokoro-int8-multi-lang-v1_0.tar.bz2
Then, as one line in cmd:
sherpa-onnx-v1.13.5-win-x64-shared-MT-Release\bin\sherpa-onnx-offline-tts.exe --kokoro-model=kokoro-int8-multi-lang-v1_0\model.int8.onnx --kokoro-voices=kokoro-int8-multi-lang-v1_0\voices.bin --kokoro-tokens=kokoro-int8-multi-lang-v1_0\tokens.txt --kokoro-data-dir=kokoro-int8-multi-lang-v1_0\espeak-ng-data --kokoro-lexicon=kokoro-int8-multi-lang-v1_0\lexicon-us-en.txt,kokoro-int8-multi-lang-v1_0\lexicon-zh.txt --num-threads=4 --sid=3 --output-filename=hello.wav "This is Kokoro speaking on Windows with no Python installed."
Flag names and structure follow the official sherpa-onnx Kokoro documentation; the paths are the Windows equivalents of its example, and the filenames are the ones we confirmed inside the archive.
--sid selects the speaker by integer ID. The v1.0 model has 53 speakers, IDs 0 through 52, and the full mapping is published in the sherpa-onnx docs: 0-19 are American voices (af_heart is 3, am_michael is 16), 20-27 British (bf_emma is 21, bm_george is 26), and 45-52 Chinese. The IDs in between belong to Spanish, French, Hindi, Italian, Japanese and Portuguese speakers — but sherpa-onnx's own page carries an explicit warning that although the underlying model is multilingual, only English and Chinese support was added. Do not plan a French project around ID 30.
--num-threads=4 is a starting point, not a rule. Raise it if synthesis feels slow and you have cores spare; there is no reason to guess when you can time it.
If nothing happens
Three things to check, in the order they actually go wrong:
- The model filename.
model.int8.onnxversusmodel.onnx. See above. - A missing Visual C++ runtime, if you took the MD build. Symptom is a Windows dialog about a missing DLL, not a sherpa error. Switch to the MT archive.
- SmartScreen. An unsigned
.exefrom a downloaded archive may be blocked on first run. Windows reports this clearly; unblock the file in its Properties dialog if you trust the source.
Run this on your own machine and stop paying every month
Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.
Piper: The Archived-Repo Trap
The Piper Windows binary that every tutorial links to is real, works, and is nearly three years old — and the project that maintains Piper today does not publish a standalone Windows binary at all.
The numbers, from the GitHub API on 18 August 2026:
| rhasspy/piper | OHF-Voice/piper1-gpl | |
|---|---|---|
| Stars | 11,280 | 5,162 |
| Status | Archived, last push 26 Aug 2025 | Active, last push 15 Aug 2026 |
| Licence | MIT | GPL-3.0 |
| Latest release | 2023.11.14-2 | v1.7.0 (15 Aug 2026) |
| Windows artifact | piper_windows_amd64.zip (22.5 MB, 255,399 downloads) | piper_tts-1.7.0-cp39-abi3-win_amd64.whl (34.1 MB) |
The archived repo's README is now a single line: development has moved. And the move changed the delivery model — the maintained Piper is a Python package. If your requirement is "no Python," the maintained Piper is not available to you, and pretending otherwise would be the exact mistake this page is trying to prevent.
That said, the frozen binary is a perfectly good piece of software. We expanded piper_windows_amd64.zip (21.4 MiB on disk, 363 files) and it contains piper.exe, espeak-ng.dll, piper_phonemize.dll, onnxruntime.dll and 356 espeak-ng-data entries — completely self-contained, same as sherpa-onnx. Voice files are still being maintained separately: the rhasspy/piper-voices collection on Hugging Face was last modified 15 August 2026 and holds 174 voice .onnx files across 50 language directories, MIT-licensed.
Grab a voice's .onnx and its matching .onnx.json — en_US-lessac-medium.onnx is 63.2 MB, the high variant 113.9 MB — then, per the 2023 README's usage:
echo Welcome to the world of speech synthesis! | piper\piper.exe --model en_US-lessac-medium.onnx --output_file welcome.wav
When to pick Piper over sherpa-onnx: when you want a specific Piper voice in one of those 50 languages, or when 22.5 MB plus a 63 MB voice beats 24.5 MB plus a 132 MB model on a metered connection. Piper's VITS voices are lighter and faster than Kokoro but generally less natural — if you have not heard both, the Hugging Face repo has samples/ MP3s next to every voice, which is the only honest way to choose.
Worth knowing: sherpa-onnx can run Piper voices too, and packages them itself — vits-piper-en_US-libritts_r-medium.tar.bz2 (82.0 MB) and vits-piper-en_US-amy-medium.tar.bz2 (67.2 MB) are in the same tts-models release as the Kokoro archives. So "Kokoro or Piper" is a voice choice, not a runtime choice. Our Piper setup guide covers the tool on its own terms across platforms.
Kokoro-FastAPI via Docker
If you want a browser interface and an API rather than a command line, Docker Desktop is the no-Python way to get one.
Kokoro-FastAPI (5,345 stars, Apache-2.0, v0.8.0 released 15 August 2026) wraps Kokoro in an OpenAI-compatible server. From its README, the CPU image is a single command:
docker run -p 8880:8880 ghcr.io/remsky/kokoro-fastapi-cpu:latest
That gives you a web interface at http://localhost:8880/web, API docs at /docs, and an OpenAI-shaped endpoint at /v1/audio/speech — which means anything that already speaks to OpenAI's TTS API can be pointed at your own machine by changing a base URL. There are GPU images too (kokoro-fastapi-gpu:latest, plus a CUDA 12.8 tag and a ROCm tag), all documented in the README.
Be clear-eyed about the trade. You are not installing Python, but you are installing Docker Desktop, which on Windows wants WSL2 or Hyper-V underneath. If your objection to WSL was "I do not want another Linux on my machine," this path does not respect it. If your objection was "I do not want to debug pip," it does. We cover this server on its own in self-hosting an OpenAI-compatible TTS API.
What the Easy Path Costs You
Three real losses, stated plainly, because the point of picking a constraint is knowing what it excludes.
1. No voice cloning. Kokoro gives you 53 fixed voices; Piper gives you 174 fixed voice files. Neither can take a 30-second clip of someone's speech and imitate it. Every project that does zero-shot cloning well — XTTS, Chatterbox, the GPT-SoVITS family — is a Python project with a real dependency tree, and no amount of wishing produces a Windows .exe for them. If cloning is the actual requirement, stop optimising for install ease and read Kokoro vs XTTS vs Chatterbox instead.
2. No emotion or style control. You get the voice as trained. Speed is adjustable; delivery is not. For audiobook narration that is usually fine — see our local audiobook generator walkthrough — and for anything performative it is a hard ceiling.
3. A frozen model. Kokoro-82M is Apache-2.0 with 6,703 likes and over 12.3 million downloads on Hugging Face, and its model repository was last modified in April 2025. It is not being improved. That is not a defect — a finished small model that runs everywhere is exactly what you want here — but do not expect it to move.
One thing you do not give up is hardware headroom. Kokoro is an 82M-parameter model; Piper's voices are smaller. These are rounding errors next to local image or language models, and they are the reason the CPU builds are the default rather than the fallback. If you want the full no-GPU picture across tools, the best TTS without a GPU covers it, and local voice AI VRAM requirements by GPU covers the cases where a card does start to matter.
The verdict, in one line: if you have already lost an evening to torch or eSpeak, take sherpa-onnx plus Kokoro today, and only pay the Python tax later if you find a specific thing it cannot do.
Sources
- k2-fsa/sherpa-onnx — release assets and metadata read from the GitHub API, 18 August 2026 (v1.13.5 published 11 Aug 2026; v1.13.6 was published the same day we checked). Windows archive contents verified by download and extraction. Apache-2.0, 14,227 stars.
- sherpa-onnx Kokoro documentation (
k2-fsa.github.io/sherpa/onnx/tts/pretrained_models/kokoro.html) — source of thesherpa-onnx-offline-ttsflag set and the 53-speaker ID map. - sherpa-onnx
tts-modelsrelease — asset sizes and download counts for the Kokoro and Piper-VITS model archives;espeak-ng-datapresence confirmed by extractingkokoro-int8-multi-lang-v1_0.tar.bz2. - rhasspy/piper (archived) — release
2023.11.14-2,piper_windows_amd64.zipcontents verified by download; CLI usage from the README at that tag. - OHF-Voice/piper1-gpl — v1.7.0 release assets, GPL-3.0, 15 August 2026.
- rhasspy/piper-voices — voice file counts and sizes from the Hugging Face API, MIT.
- remsky/Kokoro-FastAPI — v0.8.0 release and README Docker commands.
- hexgrad/Kokoro-82M — Apache-2.0, last modified 10 April 2025.
FAQ
Voice working locally? Build the whole pipeline.
Whisper, TTS, and voice cloning wired into real projects — hands-on courses. First chapter free, no card.
Replace the speech-AI subscription
Local Speech Studio covers TTS, voice cloning and transcription end to end — including which licences actually let you sell what you make.
Liked this? 25 full AI courses are waiting.
From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.
Build Real AI on Your Machine
RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.
Want the structured version?
Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.
Keep going
- PILLARmodels/coqui-tts
- audio.cpp: Local TTS and Speech-to-Text, No Python
- Best Local Speech-to-Text Models: 4 Tested on One File
- Best Local TTS Models 2026: 8 Open-Source Voices Tested
- Best Local TTS Without a GPU: Real-Time on CPU
- Build a $10K/Month AI Podcast: Whisper + Bark + Coqui TTS
- Build a Local Voice Assistant: Whisper + Ollama + Piper
- Chatterbox TTS Setup: Free ElevenLabs Killer (MIT, 2026)
- Coqui TTS Python Guide: pip install + XTTS API Examples
- Dub Videos Into Any Language Locally: pyVideoTrans + Whisper
Comments (0)
No comments yet. Be the first to share your thoughts!