AI NewsOpen sourceAnnouncement
yovoice is a free open-source alternative to ElevenLabs for macOS and Windows
yovoice is an Apache 2.0 desktop app for macOS and Windows that runs voice cloning, emotion-controlled TTS and voice design on the user's own computer, through eleven local model builds including IndexTTS 2.5, VoxCPM2, Qwen3-TTS and Kokoro.
Image: GitHub
Why it mattersA free local voice generator with eleven built-in model builds gives a product team a way to narrate videos, demos and voiceovers without paying a per-character fee or sending the script to a vendor.
Paying for a voiceover vendor by the character means the finance team sees every narration a product team writes. A new open-source desktop app called yovoice puts the same work on the user's own computer, with voice cloning, emotion control and eleven built-in local model builds, and runs on macOS 14 or later and Windows 10 or 11.
The project is at version 0.1.7 and the GitHub README was last pushed on 2026-10-09. The repository was created on 2026-09-15, has 467 stars and 44 forks on day 25, and is released under the Apache 2.0 licence.
What ships today
yovoice bundles a desktop app, a command-line tool and an HTTP server with MCP support behind one Apache 2.0 licence. The author ships a .dmg installer for Apple silicon and a Windows -setup.exe through GitHub Releases, and models are downloaded inside the app on first use. Uninstalling leaves user data in ~/.yovoice so a reinstall picks up where the user left off.
The feature list in the README covers three jobs: speech from text with a chosen voice, cloning a voice from a reference recording, and designing a new voice from a written description. The app also handles audio project work: importing or recording a reference clip, trimming, previewing and exporting.
The models
Eleven model builds ship through yovoice, and the README lists their parameter sizes, their file sizes at Q8, BF16 and F16 precision and the languages each covers. IndexTTS 2.0 and 2.5 handle multilingual cloning with emotion control, VoxCPM2 at two billion parameters adds text-guided voice design, OmniVoice at 0.6 billion parameters covers attribute-based design with non-verbal sound tags, and four Qwen3-TTS builds at 0.6 and 1.7 billion parameters handle multilingual synthesis and nine built-in voices.
Kokoro-82M is the smallest in the set, at 82 million parameters and a 190 megabyte Q8 file size. The README lists 49 built-in voices in the official 1.0 release, 54 in the 1.0 import, and over 100 Chinese and three English voices in the experimental 1.1-zh. OmniVoice is restricted to non-commercial use under CC-BY-NC, and the README names that licence next to the model.
For an agent
The project also publishes a Claude skill at skills/yovoice and a CLI, so an agent can be asked to install the skill, download a model and then "read narration.txt using voice.wav as the reference voice, with a calm delivery, and save it as narration.wav". A separate yovoice serve command exposes an authenticated HTTP API and an MCP inference service, so another machine or another agent on the same network can call the same local models.
For a working team, the money question is simple. A paid voiceover vendor charges by the character and a desktop app with local models does not, so the test is whether the output quality on the user's own narration is close enough to the vendor's for the work it will be used in. yovoice's day-25 repository, with 467 stars, 44 forks, 11 open issues and no reported production use, is the point at which a short pilot run on real scripts is useful and a full switch is premature.
Source
Primary source: leemysw/yovoice on GitHub, Apache 2.0. Project page: yovoice.leemysw.com.
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.
