Kitten TTS: 25MB Text-to-Speech That Runs on Any CPU (Even a Raspberry Pi)
Want to go deeper than this article?
Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.
Voice working locally? Build the whole pipeline. Whisper, TTS, and voice cloning wired into real projects — hands-on courses. First chapter free, no card.
Kitten TTS is a family of Apache-2.0 text-to-speech models from 15M to 80M parameters — the smallest ships as a 25MB ONNX file — that run entirely on CPU with 8 voices at 24kHz. One pip command installs it. No GPU, no API key, no cloud, $0.
The practical recommendation: install a nano build, which is 15M parameters and 56MB as fp32, and move up to mini (80M) only if output quality matters more than footprint. Do not assume the famous 25MB int8 file is also the fastest one — int8 buys disk space, and whether it buys speed depends on your CPU. If your target is a Raspberry Pi, this is one of the rare AI tools with an official Adafruit guide behind it.
The rest of this page is the setup: exact install commands from the project README, the model lineup, how to measure real-time factor on your own machine in four lines of Python, the Raspberry Pi picture, and the three install failures you are most likely to hit — including one that kills Python with no traceback.
What Kitten TTS Is
Kitten TTS is an open-source TTS project by KittenML that went from a single 15M-parameter preview to a four-model family within a year. It is built on the StyleTTS 2 architecture, per the acknowledgements on the official model cards.
The project's short history matters because it tells you what you are adopting:
- August 2025: launch as a "Show HN: Kitten TTS – 25MB CPU-Only, Open-Source TTS Model", which reached the Hacker News front page. A 15M-parameter model you could ship in a Docker layer without noticing was a genuinely new thing.
- March 2026: a second front-page thread, "Show HN: Three new Kitten TTS models – smallest less than 25MB", announcing the current 0.8 generation: mini (80M), micro (40M), and a retrained nano (15M).
- Along the way: Adafruit published an official Raspberry Pi learning guide, and a Rust implementation with a CLI and API server (second-state/kitten_tts_rs) appeared for people who want the model without the Python stack.
Two front-page launches a year apart plus an Adafruit guide is a decent signal that this is not abandonware — which, for a young project, is the thing to check before you build on it. Where it sits in the wider landscape: our best local TTS models roundup covers the full field; Kitten TTS occupies the "smallest hardware floor" corner of it.
Reading articles is good. Building is better.
Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.
The Four Models
There are four builds in the 0.8 generation, from a 25MB int8 nano to an 80MB mini — all output 24kHz audio and all share the same 8 voices. The lineup, as listed in the official README:
| Model | Parameters | Download size | Hugging Face ID |
|---|---|---|---|
| Mini | 80M | 80MB | KittenML/kitten-tts-mini-0.8 |
| Micro | 40M | 41MB | KittenML/kitten-tts-micro-0.8 |
| Nano (fp32) | 15M | 56MB | KittenML/kitten-tts-nano-0.8-fp32 |
| Nano (int8) | 15M | 25MB | KittenML/kitten-tts-nano-0.8-int8 |
(Source: the KittenML/KittenTTS README on GitHub, linked in full at the bottom of this page.)
Every model exposes the same eight voices, listed in the README: Bella, Jasper, Luna, Bruno, Rosie, Hugo, Kiki, and Leo. You can confirm the list on a running install by printing the loaded model's available_voices attribute. There is no voice cloning; you pick from these eight.
One thing the model cards do not claim: multilingual support. Everything documented — voices, examples, demos — is English. Treat Kitten TTS as English-only until the project says otherwise.
Install in Five Minutes
One pip command installs everything, and the official install path is the release wheel rather than PyPI. Straight from the project README:
# Use a recent-but-not-bleeding-edge Python — see gotcha #1 below
python3.12 -m venv kitten
source kitten/bin/activate
pip install https://github.com/KittenML/KittenTTS/releases/download/0.8.1/kittentts-0.8.1-py3-none-any.whl soundfile
That URL comes from the README (version 0.8.1 at the time of writing — check the repo's releases page if you are reading this later). soundfile is for writing WAV output.
Budget for the download, because the 25MB model headline does not include PyTorch, spaCy, onnxruntime and the phonemizer stack the package depends on. Commenters in the March 2026 HN thread reported roughly 756MiB in the best case, around 3GB with CPU-only torch, and 7.1GB when pip pulled CUDA builds — the last of which the author acknowledged as a bug. If pip starts downloading NVIDIA packages on a CPU-only Linux box, install the CPU wheel of torch first:
pip install torch --index-url https://download.pytorch.org/whl/cpu
Model weights download automatically from Hugging Face on first use, so the first KittenTTS(...) call needs network access. After that it is fully offline.
First Words: Quickstart
Five lines of Python produce a WAV file. This is the README's example, with the nano model swapped in:
from kittentts import KittenTTS
model = KittenTTS("KittenML/kitten-tts-nano-0.8-fp32")
audio = model.generate(
"Local text to speech has finally gotten small enough to run anywhere.",
voice="Jasper",
)
import soundfile as sf
sf.write("output.wav", audio, 24000) # 24kHz output
Play it with aplay output.wav on Linux (that is what Adafruit's Pi guide uses) or afplay output.wav on macOS.
Two practical notes. First, model loading, not generation, dominates short scripts — instantiating KittenTTS(...) reads the ONNX graph and spins up the phonemizer, and at these model sizes that setup cost is larger than synthesising a sentence. If you are building anything interactive, load once and keep the object alive rather than constructing it per request. Second, voice choice is cheap to explore: loop over model.available_voices and generate the same sentence with each.
Run this on your own machine and stop paying every month
Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.
How Fast Is It On Your CPU?
The metric that matters for TTS is real-time factor (RTF):
RTF = generation time ÷ duration of the audio produced
RTF of 0.1 means ten seconds of speech took one second to generate — "10x realtime". RTF of 1.0 means generation takes exactly as long as the audio lasts.
Anything below about 0.5 feels instant in an assistant; anything above 1.0 means you cannot stream speech in real time and have to pre-generate. We do not own the hardware to publish first-party RTF figures, so here is the four-line measurement instead — run it on the machine you actually care about:
import time, soundfile as sf
from kittentts import KittenTTS
model = KittenTTS("KittenML/kitten-tts-nano-0.8-fp32")
text = " ".join(["Local text to speech has gotten small enough to run anywhere."] * 8)
model.generate("warmup", voice="Jasper") # first call pays setup costs
start = time.perf_counter()
audio = model.generate(text, voice="Jasper")
elapsed = time.perf_counter() - start
duration = len(audio) / 24000 # 24kHz output
print(f"audio {duration:.1f}s, generated in {elapsed:.1f}s, RTF {elapsed / duration:.2f}")
Discard the first run — it includes model load and phonemizer initialisation — and take the median of a few passes.
For calibration before you install anything, two attributed third-party data points from the launch Hacker News thread: a commenter benchmarked the launch-era model at roughly 5x realtime on an Intel i9-14900HX, and another measured roughly 1x realtime on an Intel Celeron N4020 (with a several-second model load), one of the weakest x86 CPUs sold this decade. Commenters also reported no speedup from a discrete GPU, which is expected — the package has no GPU path.
One thing worth knowing before you assume the 25MB int8 file is the fast one: int8 quantization reliably reduces file size, but it only speeds up inference when the CPU has the integer instruction support to exploit it. ONNX Runtime's own quantization documentation makes this conditional on hardware (VNNI and similar). On a CPU without it, the dequantization overhead can eat the gains. Measure both nano builds with the snippet above before choosing.
Raspberry Pi Setup
Kitten TTS runs on a Raspberry Pi 4 or 5, and you do not have to take our word for it — Adafruit maintains an official learning guide, "Speech Synthesis On Raspberry Pi with KittenTTS." Their guide walks through a Pi with a Voice Bonnet speaker hat, installs the package into a venv, and ends with a genuinely useful demo: an offline weather announcer that reads National Weather Service forecasts aloud — no API keys, no cloud TTS bill.
The honest performance picture: neither Adafruit nor the KittenTTS project publishes Pi real-time-factor numbers. The best available anchor is the HN-reported ~1x realtime on a Celeron N4020; a Pi 5's Cortex-A76 cores are in a broadly similar performance class, so expect the nano models to land around realtime on a Pi 5, slower on a Pi 4 — Adafruit says only that the Pi 5 is faster. That is an inference from a third-party data point, clearly labelled as one, not a measurement. Around realtime is fine for a voice assistant that speaks a sentence or two; it is tedious for generating an hour of audio.
Three Pi-specific notes:
- Use a Pi 4/5 with standard Raspberry Pi OS. The roughly 1GB Python dependency stack (torch, spaCy) makes Pi Zero-class boards a non-starter for the official package. For truly minimal hardware, the Rust port (second-state/kitten_tts_rs — CLI plus API server) skips Python entirely.
- Watch the version skew. Adafruit's guide was written against the launch-era 0.1 wheel and old voice names (
expr-voice-2-fstyle). The concepts transfer; the model IDs and voices in this article are the current ones. This project moves fast enough that any tutorial — including ours — deserves a cross-check against the README. - A Pi that talks pairs naturally with a Pi that thinks. Small LLMs on the same board are covered in our Raspberry Pi 5 LLM guide, and the full listen-think-speak pattern in our local voice assistant build — swap Kitten TTS in for the speaking part. If your edge target is a phone instead, see running LLMs on a phone.
Three Install Gotchas
Two of these have one-line fixes; the third is a version choice you make before you start.
1. A too-new Python breaks dependency resolution. The 0.8.1 wheel depends on the misaki G2P library, and if pip cannot satisfy misaki>=0.9.4 on your interpreter you will see a "could not find a version that satisfies the requirement" failure rather than anything TTS-specific. The README says Python 3.8+; in practice, create the venv on 3.10–3.12 and this class of problem disappears.
2. A hard crash mentioning phontab. If instantiating the model kills the Python process outright — no traceback, just Error processing file '.../phontab': No such file or directory — that is espeak-ng, the phonemizer underneath, looking for its data files in the venv root instead of inside the bundled espeakng_loader package. Symlinking the data files where it looks resolves it:
ln -s "$(python -c 'import espeakng_loader; print(espeakng_loader.get_data_path())')"/* "$VIRTUAL_ENV/"
Most installs never see this one, but it is unpleasant to debug precisely because there is no stack trace.
3. Multi-gigabyte installs on Linux. Covered above: pip pulling CUDA torch for a CPU-only model. Install the CPU torch wheel first and the footprint drops to under a gigabyte.
Where It Fits vs Piper and Kokoro
Kitten TTS is the smallest-footprint option with expressive voices; Piper is the more mature Pi incumbent; Kokoro (the 82M hexgrad/Kokoro-82M) is the quality pick among small models. Choosing between them is mostly about which constraint binds:
- Choose Kitten TTS when model size is the constraint, or when you specifically want the setup Adafruit documents for the Pi. It is the youngest of the three, which cuts both ways: rapid improvement, moving APIs.
- Choose Piper when you need the boring, proven answer on embedded Linux: years of deployment in Home Assistant setups, broad language coverage, and lean native tooling rather than a torch-and-spaCy Python stack. It remains our default recommendation for production Pi voice output.
- Choose Kokoro when quality-per-parameter is the metric. At 82M parameters it plays in the same size band as Kitten's mini, and its output is the benchmark small models get compared against — the launch HN thread's consensus on Kitten TTS was "impressive for the size", which is exactly the right framing. Kokoro is also the default voice in several popular self-hosted TTS servers.
None of these three does voice cloning. If you want to narrate in a specific voice — including your own — that is a different weight class entirely; see the options in our local TTS roundup. And if the input side of your pipeline is speech too, Faster-Whisper is the matching CPU-friendly transcription piece.
Honest Limitations
Kitten TTS is genuinely impressive for its size — and "for its size" is doing real work in that sentence. What to know before building on it:
- English only. No multilingual support is documented anywhere in the repo or model cards. Piper covers dozens of languages; Kitten covers one.
- Eight fixed voices, no cloning. You cannot add voices or clone one.
- Quality trails the size classes above it. Nobody should choose Kitten over Kokoro, XTTS-class models, or cloud TTS on output quality alone. The launch-thread verdict — sounds OK, remarkable that it comes from 15M parameters — is the right expectation to set. The mini model narrows the gap; it does not close it.
- The 25MB headline is the model, not the install. Roughly 1GB of Python dependencies in practice, per the community reports above. The Rust port is the escape hatch.
- Young project, moving fast. Version 0.1 to 0.8.1 inside a year, with voice names and model IDs changing along the way — which silently broke earlier tutorials, Adafruit's included. Pin your versions in anything you deploy.
- 24kHz mono output. Fine for speech; it is not studio-grade audio, and there are no knobs for emotion or style control beyond voice choice.
None of these are dealbreakers for what the project actually is: the lowest-friction way to give any CPU-having device a voice, for free, offline.
Sources
- KittenML/KittenTTS on GitHub — model table, install command, voice list, quickstart code, Apache-2.0 license
- KittenML/kitten-tts-mini-0.8 on Hugging Face — StyleTTS 2 acknowledgement, license
- Hacker News: launch thread, August 2025 — i9-14900HX and Celeron N4020 performance reports, install-size reports, author responses
- Hacker News: "Three new Kitten TTS models" — 0.8-generation announcement and install-size discussion
- Adafruit: Speech Synthesis On Raspberry Pi with KittenTTS — Pi 4/5 support, Voice Bonnet setup, weather-announcer project
- second-state/kitten_tts_rs — Rust CLI/API-server port
- ONNX Runtime quantization documentation — why int8 speedups depend on CPU instruction support
Performance figures on this page are either attributed to the named third party who published them, or presented as a method for you to measure on your own hardware. We do not publish first-party hardware benchmarks.
FAQ
Voice working locally? Build the whole pipeline.
Whisper, TTS, and voice cloning wired into real projects — hands-on courses. First chapter free, no card.
Replace the speech-AI subscription
Local Speech Studio covers TTS, voice cloning and transcription end to end — including which licences actually let you sell what you make.
Liked this? 25 full AI courses are waiting.
From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.
Build Real AI on Your Machine
RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.
Want the structured version?
Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.
Keep going
- PILLARmodels/coqui-tts
- audio.cpp: Local TTS and Speech-to-Text, No Python
- Best Local Speech-to-Text Models: 4 Tested on One File
- Best Local TTS Models 2026: 8 Open-Source Voices Tested
- Best Local TTS Without a GPU: Real-Time on CPU
- Build a $10K/Month AI Podcast: Whisper + Bark + Coqui TTS
- Build a Local Voice Assistant: Whisper + Ollama + Piper
- Chatterbox TTS Setup: Free ElevenLabs Killer (MIT, 2026)
- Coqui TTS Python Guide: pip install + XTTS API Examples
- Dub Videos Into Any Language Locally: pyVideoTrans + Whisper
Comments (0)
No comments yet. Be the first to share your thoughts!