Reference

Open-source alternatives to AI tools

Every project here replaces a named paid tool for a real use. The licence, the release and the star count are read from GitHub and dated. The feature claims were checked against the README on the date shown on each page.

Alternatives to

ProjectReplacesOpennessStarsLast releaseSelf-host
Whisper

OpenAI's general-purpose speech recognition model, released with code and weights under MIT, in six sizes from 39 million to 1.55 billion parameters, with multilingual transcription, translation to English and language detection.

Speech-to-text engines
OSI
109,222
Jun 26, 2025
Yes
whisper.cpp

A dependency-free C and C++ port of Whisper under MIT that runs on CPUs, Apple Silicon and phones, with quantised models, an HTTP server and bindings for a dozen languages.

Speech-to-text engines
OSI
53,711
Sep 11, 2026
Yes
Fish Speech

A multilingual text-to-speech system with voice cloning from a 10 to 30 second clip and inline emotion tags, released with its weights under a research licence that requires a separate agreement for commercial use.

Text to speech and voice cloning
Source available
32,719
May 31, 2025
Yes
Handy

A free, offline push-to-talk dictation app for macOS, Windows and Linux that runs Whisper or NVIDIA Parakeet on your own machine and pastes the text at your cursor, under MIT.

Voice dictation apps
OSI
31,790
Aug 24, 2026
Yes
Meetily

A meeting assistant under MIT that records your microphone and system audio together, transcribes live on your machine with Whisper or Parakeet, and summarises with a language model you choose.

Meeting transcription
OSI
30,859
Sep 12, 2026
Yes
Chatterbox

A family of text-to-speech models from Resemble AI under MIT, with zero-shot voice cloning from a reference clip and a multilingual model covering 23 languages.

Text to speech and voice cloning
OSI
26,448
Jun 13, 2025
Yes
WhisperX

Batched Whisper transcription under BSD-2-Clause with word-level timestamps from forced alignment and speaker labels from pyannote, the closest open-source match to a hosted transcription API's output.

Speech-to-text engines
OSI
24,069
May 25, 2026
Yes
Buzz

A desktop app under MIT that transcribes and translates audio and video files, YouTube links and a live microphone offline with Whisper, with speaker identification and SRT, VTT and TXT export.

Meeting transcription
OSI
21,529
Aug 23, 2026
Yes
Parakeet (NeMo Speech)

NVIDIA's speech models and toolkit: Apache-2.0 code, and the Parakeet TDT 0.6B v3 model under CC-BY-4.0 with 25 European languages, automatic language detection, punctuation and word timestamps.

Speech-to-text engines
Open weights
18,465
Aug 7, 2026
Yes
Pipecat

A Python framework under BSD-2-Clause, maintained by Daily, for real-time voice and multimodal agents, with the widest provider list on this site and telephony serialisers for Twilio, Telnyx, Plivo, Vonage, Exotel and Genesys.

Voice agent frameworks
OSI
15,594
Sep 12, 2026
Yes
F5-TTS

A research text-to-speech model with zero-shot voice cloning whose code is MIT but whose pretrained weights are non-commercial, because of the dataset they were trained on.

Text to speech and voice cloning
Open weights
15,239
Jul 23, 2026
Yes
Vosk

An offline speech recognition toolkit under Apache-2.0 with streaming, models of about 50 MB, 20+ languages, a reconfigurable vocabulary and speaker identification, from small boards to clusters.

Speech-to-text engines
OSI
15,131
Apr 22, 2024
Yes
sherpa-onnx

A runtime under Apache-2.0 from the next-generation Kaldi team that runs speech to text, text to speech, speaker diarization and voice activity detection locally on CPUs, NPUs and phones, with no internet connection.

Speech-to-text engines
OSI
14,808
Sep 10, 2026
Yes
LiveKit Agents

An Apache-2.0 framework for real-time voice agents that connects any speech, language and voice provider, with semantic turn detection, native MCP tool support, a test framework, and an open-source media and SIP stack it can run on entirely.

Voice agent frameworks
OSI
14,230
Sep 15, 2026
Yes
TEN Framework

A real-time conversational AI framework from Agora with a visual designer, a SIP extension and turn detection, released under Apache-2.0 plus conditions that forbid competing deployments, which makes it source available rather than open source.

Voice agent frameworks
Source available
11,130
Jul 31, 2026
Yes
Moonshine

An on-device speech toolkit under MIT built for live streaming, with speech-to-text models trained from scratch in sizes down to 1 MB, one library across Python, the browser, phones and desktops.

Speech-to-text engines
OSI
11,089
Aug 24, 2026
Yes
OpenWhispr

A dictation and meeting-notes app under MIT that runs Whisper or Parakeet locally or cloud models with your own key, detects Zoom, Teams and FaceTime calls, and labels speakers on device.

Voice dictation apps
OSI
8,286
Sep 15, 2026
Yes
Vibe

An offline transcription app under MIT that records system audio and the microphone, supports Whisper, Parakeet and Nemotron models, labels speakers, and exports SRT, VTT, TXT, HTML, PDF, JSON and DOCX.

Meeting transcription
OSI
7,510
Sep 5, 2026
Yes
VoiceInk

A macOS dictation app under GPL-3.0 whose README names Superwhisper and Wispr Flow as the tools it replaces, with local models, per-app modes, a personal dictionary and an AI assistant mode.

Voice dictation apps
OSI
6,435
Aug 27, 2026
Yes
Piper

A fast, local neural text-to-speech engine maintained by the Open Home Foundation, used by Home Assistant and the NVDA screen reader, under GPL-3.0.

Text to speech and voice cloning
OSI
5,598
Sep 4, 2026
Yes
Kokoro FastAPI

A Docker image that serves the 82-million-parameter Kokoro model through an OpenAI-compatible speech endpoint, with streaming, voice mixing and nine languages, all under Apache-2.0.

Text to speech and voice cloning
OSI
5,449
Sep 10, 2026
Yes
Whispering

A speech-to-text app inside the Epicenter monorepo under AGPL-3.0 that records, transcribes with a provider you choose, cloud or local, optionally polishes the text, and pastes it at the cursor.

Voice dictation apps
OSI
4,801
Dec 27, 2025
Yes
Speaches

An OpenAI-compatible server under MIT for streaming transcription, translation and text to speech, loading Whisper, Kokoro and Piper models on demand, described by its maintainers as Ollama for speech models.

Speech-to-text engines
OSI
3,665
Dec 27, 2025
Yes
Bolna

An MIT voice-agent platform that defines an agent in a JSON file and orchestrates speech to text, a language model and text to speech over websockets, with Twilio and Plivo for phone calls, whose maintainers say they are looking for help.

Voice agent frameworks
OSI
763
Sep 16, 2026
Yes

Stars and release dates read from GitHub on September 16, 2026. Openness labels: The code, and the weights where the project is a model, carry a licence on the OSI approved list. The code is public, and the licence limits what you may do with it. Read the LICENSE file before commercial use. The model weights can be downloaded and run. The training data or code is closed, or the weights carry a restriction the code does not.

How a project gets in, and how it leaves

What counts

A public repository, a README, a commit in the last 90 days, and one paid tool it replaces for a use a reader can test. Star counts are shown as a dated fact and never used as a bar.

What the label means

Open source means an OSI licence on the code and the weights. Source available means the licence limits use. Open weights means a model you can download whose training data or code is closed.

What is never said

That a project is faster, better or cheaper than the tool it replaces. Prices for anyone. User counts. A roadmap. Every feature claim is the project's own, from its README.

Categories: Text to speech and voice cloning, Voice dictation apps, Meeting transcription, Speech-to-text engines, Voice agent frameworks.

Running one of these inside your own environment, with your own data and your own security rules, is the kind of work Reveneau does. Read how a forward deployed engagement works.