Vanta
Vanta is a free, open-source AI video engine under MIT, built on top of Remotion. Its README says it wires 40 or so open-source projects (GPT-SoVITS and OpenVoice for voice cloning, MuseTalk and LatentSync for talking-head avatars, WhisperX for animated captions, Wan and LTX-Video for text-to-video, ACE-Step for music) into one local render pipeline. It replaces ElevenLabs for the voice-cloning step and much more.
What does Vanta do?
Vanta is described in its README as "a programmatic AI video engine built on Remotion that combines 40+ open source repositories into a single creation pipeline." Its list of parts covers voice cloning (GPT-SoVITS and OpenVoice), talking-head avatars (MuseTalk, LatentSync, InfiniteTalk), word-accurate animated captions (WhisperX), text-to-video generation (Wan 2.2, LTX-Video), music generation (ACE-Step), background removal, image editing (Sharp), vector graphics (SVG.js), a drag-and-drop editor, a video timeline, "100+ GPU-accelerated transitions", and motion graphics, all running locally.
Each part has a TypeScript module in src/integrations/ that connects to Remotion's render pipeline through standard React components. The README shows the pattern: call cloneVoice() with a WAV sample, then drop the returned audio URL into Remotion's <Audio> component.
Key facts
- Licence: MIT on the code, per the LICENSE file. GitHub's license API returns NOASSERTION for the repository metadata, but the LICENSE is standard MIT text.
- Built on Remotion, a React video framework. Remotion's own licence is free for individuals, non-profits and companies with up to 3 employees; a larger for-profit company needs a paid Remotion company licence.
- 40+ open-source repositories integrated, per the README's own count.
- Voice cloning models named: GPT-SoVITS (MIT), OpenVoice (MIT), Chatterbox (MIT), F5-TTS (MIT), VoxCPM (Apache-2.0), and kokoro-js (Apache-2.0).
- Talking-head avatars named: MuseTalk (MIT), LatentSync (Apache-2.0), InfiniteTalk (Apache-2.0), EchoMimicV3 (Apache-2.0).
- Auto-captions use WhisperX (BSD-2), which the README says runs at 70x real time with batching and gives word-level timestamps.
- Install:
git clone,npm install,npm startfor Remotion Studio,npm run renderfor the showcase. - The README says every integration was audited for commercial-safe licensing (Apache-2.0, MIT, BSD), and lists what was removed for licence problems (SadTalker as abandoned, Wav2Lip as non-commercial, V-Express as research-only).
What does it replace, and where does it fall short?
Vanta replaces ElevenLabs for the voice-cloning step, since GPT-SoVITS and OpenVoice run locally with the models the README names. It also replaces parts of Adobe Creative Cloud, Synthesia, HeyGen, Descript, CapCut Pro, Topaz Video AI, and the Remotion Pro Store, per the README's own claim. It is a candidate on the open-source alternatives to ElevenLabs page.
Where it falls short: this is a wiring project on top of 40 other projects. Each of those projects has its own install, its own model download, and its own hardware needs, and the README's code examples connect to a separate local server for each one, for example a GPT-SoVITS server at http://localhost:9880. You are also using models whose weights carry their own separate licences, which the README acknowledges by naming what was removed. For voice cloning alone, Chatterbox or F5-TTS are simpler starting points.
How does Vanta run?
The README quick start is git clone, npm install, npm start to open Remotion Studio in the browser, and npm run render to render the showcase to out/vanta-showcase.mp4. Each integration needs the corresponding upstream model server running locally (for example, a GPT-SoVITS api_v2.py server for voice cloning, or a MuseTalk server for avatars). The README describes each integration's server URL as configuration you pass in.
There is no hosted paid version. Vanta is code that renders on the machine it runs on, and no per-video charges or API keys are named in the README.
Who is Vanta for?
A developer who is comfortable running local model servers and wants one Remotion pipeline that ties them together into a video render. A person who wants a one-click video app should look at the individual voice and video projects instead, since Vanta is the plumbing on top of them.
What limits does the README state?
Each integration expects its own upstream model server running locally, on the ports and URLs the README names, and the README acknowledges hardware demands vary: EchoMimicV3 is called out as running on 12 GB VRAM, and other avatar and video models need more. The README says SadTalker, Wav2Lip and V-Express were removed for licence and abandonment reasons. Vanta renders through Remotion, so a for-profit company with more than 3 employees needs a paid Remotion company licence, under Remotion's own licence file.
Questions people ask
Is Vanta open source?
Yes. The LICENSE file is standard MIT, an OSI approved licence. GitHub's license API returns NOASSERTION for the repository metadata, but reading the LICENSE file itself shows the MIT grant with no additional restriction. Each upstream project Vanta integrates has its own licence, which the README lists per integration.
Can I use Vanta commercially?
The Vanta code itself is MIT. Vanta renders through Remotion, whose licence is free for individuals and companies with up to 3 employees and requires a paid company licence above that. The README says every integration was audited for commercial-safe licensing (Apache-2.0, MIT, or BSD) and names what was removed for non-commercial or research-only clauses. That covers the code; the individual model weights you download for voice cloning or video generation carry their own licences, and the README says to check each one before selling what you make.
Does Vanta run without an internet connection?
Yes, once the models are downloaded, per the README. Every integration it shows connects to a server on your own machine, and the README names no API keys or subscriptions.
How does Vanta compare with ElevenLabs?
For voice cloning, Vanta uses the local GPT-SoVITS and OpenVoice models the README names, so audio never leaves your machine and there is no per-character bill. ElevenLabs is a hosted service with an account. Vanta is a pipeline you assemble and run yourself. The open-source alternatives to ElevenLabs page has narrower options if all you want is TTS.
Do I need a GPU?
For the model servers, plan on one. Vanta's README says EchoMimicV3, one of its avatar models, runs on 12 GB of GPU memory, and each other model sets its own hardware needs in its own README. Vanta itself is Node code on top of Remotion. Check each model's requirements before you install it.
Sources
- Vanta README and LICENSE: github.com/itsjwill/vanta, read 2026-09-28. GitHub's license API returns NOASSERTION for the repository metadata; the LICENSE file is standard MIT.
More text to speech and voice cloning
An AGPL-3.0 desktop workflow engine for voice cloning, voice design, video dubbing, dictation and audiobooks, running 16 speech engines and 10 transcription engines locally, with a local API and an MCP server for agents.
A multilingual text-to-speech system with voice cloning from a 10 to 30 second clip and inline emotion tags, released with its weights under a research licence that requires a separate agreement for commercial use.
A family of text-to-speech models from Resemble AI under MIT, with zero-shot voice cloning from a reference clip and a multilingual model covering 23 languages.
A research text-to-speech model with zero-shot voice cloning whose code is MIT but whose pretrained weights are non-commercial, because of the dataset they were trained on.
A fast, local neural text-to-speech engine maintained by the Open Home Foundation, used by Home Assistant and the NVDA screen reader, under GPL-3.0.
A Docker image that serves the 82-million-parameter Kokoro model through an OpenAI-compatible speech endpoint, with streaming, voice mixing and nine languages, all under Apache-2.0.
Added September 28, 2026. Every claim above comes from the project's README, LICENSE or model card, read on September 28, 2026, or from the GitHub API on the date shown in the panel. Found an error? Write to reveneau@licheo.com and it is fixed in the next weekly pass. Repository: github.com/itsjwill/vanta.
Running one of these inside your own environment, with your own data and your own security rules, is the kind of work Reveneau does. Read how a forward deployed engagement works.