Open source

Open-source text to speech and voice cloning

Models and servers that turn text into spoken audio on your own hardware, including ones that copy a voice from a short reference clip. The licence on the weights matters as much as the licence on the code here, and the table shows both.

ProjectReplacesOpennessStarsLast releaseSelf-host
Fish Speech

A multilingual text-to-speech system with voice cloning from a 10 to 30 second clip and inline emotion tags, released with its weights under a research licence that requires a separate agreement for commercial use.

Source available
32,719
May 31, 2025
Yes
Chatterbox

A family of text-to-speech models from Resemble AI under MIT, with zero-shot voice cloning from a reference clip and a multilingual model covering 23 languages.

OSI
26,448
Jun 13, 2025
Yes
F5-TTS

A research text-to-speech model with zero-shot voice cloning whose code is MIT but whose pretrained weights are non-commercial, because of the dataset they were trained on.

Open weights
15,239
Jul 23, 2026
Yes
Piper

A fast, local neural text-to-speech engine maintained by the Open Home Foundation, used by Home Assistant and the NVDA screen reader, under GPL-3.0.

OSI
5,598
Sep 4, 2026
Yes
Kokoro FastAPI

A Docker image that serves the 82-million-parameter Kokoro model through an OpenAI-compatible speech endpoint, with streaming, voice mixing and nine languages, all under Apache-2.0.

OSI
5,449
Sep 10, 2026
Yes

Running one of these inside your own environment, with your own data and your own security rules, is the kind of work Reveneau does. Read how a forward deployed engagement works.