A multilingual text-to-speech system with voice cloning from a 10 to 30 second clip and inline emotion tags, released with its weights under a research licence that requires a separate agreement for commercial use.
Docker, CLI, Web UI, Server · Linux (GPU server), Docker · Runs on your own machines · Hosted version available from the maintainers · Weights: Fish Audio Research License, non-commercial without a separate agreement
Read the full page · Repository
A family of text-to-speech models from Resemble AI under MIT, with zero-shot voice cloning from a reference clip and a multilingual model covering 23 languages.
pip, Python library · Linux (tested on Debian 11), CPU or NVIDIA GPU · Runs on your own machines · Hosted version available from the maintainers
Read the full page · Repository
A research text-to-speech model with zero-shot voice cloning whose code is MIT but whose pretrained weights are non-commercial, because of the dataset they were trained on.
pip, Docker, CLI, Web UI (Gradio) · Linux (NVIDIA, AMD ROCm, Intel GPU), macOS (Apple Silicon) · Runs on your own machines · Weights: CC-BY-NC-4.0 on the pretrained models; the code is MIT
Read the full page · Repository
A runtime under Apache-2.0 from the next-generation Kaldi team that runs speech to text, text to speech, speaker diarization and voice activity detection locally on CPUs, NPUs and phones, with no internet connection.
Libraries in 12 languages, prebuilt Android and Flutter apps, WebAssembly, websocket server · Linux, macOS, Windows, Android, iOS, HarmonyOS, Node.js, WebAssembly, Raspberry Pi and named boards · Runs on your own machines
Read the full page · Repository
A fast, local neural text-to-speech engine maintained by the Open Home Foundation, used by Home Assistant and the NVDA screen reader, under GPL-3.0.
pip, CLI, Server, C/C++ library · Linux, macOS, Windows, small devices (used by Home Assistant) · Runs on your own machines
Read the full page · Repository
A Docker image that serves the 82-million-parameter Kokoro model through an OpenAI-compatible speech endpoint, with streaming, voice mixing and nine languages, all under Apache-2.0.
Docker, Server (OpenAI-compatible API) · Linux, macOS (Apple Silicon when run directly), Windows · Runs on your own machines
Read the full page · Repository
An OpenAI-compatible server under MIT for streaming transcription, translation and text to speech, loading Whisper, Kokoro and Piper models on demand, described by its maintainers as Ollama for speech models.
Docker, Server (OpenAI-compatible API) · Docker on GPU or CPU · Runs on your own machines
Read the full page · Repository