Open-source alternatives to AI tools
Every project here replaces a named paid tool for a real use. The licence, the release and the star count are read from GitHub and dated. The feature claims were checked against the README on the date shown on each page.
Alternatives to
OpenAI's general-purpose speech recognition model, released with code and weights under MIT, in six sizes from 39 million to 1.55 billion parameters, with multilingual transcription, translation to English and language detection.
Speech-to-text enginesA dependency-free C and C++ port of Whisper under MIT that runs on CPUs, Apple Silicon and phones, with quantised models, an HTTP server and bindings for a dozen languages.
Speech-to-text enginesA multilingual text-to-speech system with voice cloning from a 10 to 30 second clip and inline emotion tags, released with its weights under a research licence that requires a separate agreement for commercial use.
Text to speech and voice cloningA free, offline push-to-talk dictation app for macOS, Windows and Linux that runs Whisper or NVIDIA Parakeet on your own machine and pastes the text at your cursor, under MIT.
Voice dictation appsA meeting assistant under MIT that records your microphone and system audio together, transcribes live on your machine with Whisper or Parakeet, and summarises with a language model you choose.
Meeting transcriptionA family of text-to-speech models from Resemble AI under MIT, with zero-shot voice cloning from a reference clip and a multilingual model covering 23 languages.
Text to speech and voice cloningBatched Whisper transcription under BSD-2-Clause with word-level timestamps from forced alignment and speaker labels from pyannote, the closest open-source match to a hosted transcription API's output.
Speech-to-text enginesA desktop app under MIT that transcribes and translates audio and video files, YouTube links and a live microphone offline with Whisper, with speaker identification and SRT, VTT and TXT export.
Meeting transcriptionNVIDIA's speech models and toolkit: Apache-2.0 code, and the Parakeet TDT 0.6B v3 model under CC-BY-4.0 with 25 European languages, automatic language detection, punctuation and word timestamps.
Speech-to-text enginesA Python framework under BSD-2-Clause, maintained by Daily, for real-time voice and multimodal agents, with the widest provider list on this site and telephony serialisers for Twilio, Telnyx, Plivo, Vonage, Exotel and Genesys.
Voice agent frameworksA research text-to-speech model with zero-shot voice cloning whose code is MIT but whose pretrained weights are non-commercial, because of the dataset they were trained on.
Text to speech and voice cloningAn offline speech recognition toolkit under Apache-2.0 with streaming, models of about 50 MB, 20+ languages, a reconfigurable vocabulary and speaker identification, from small boards to clusters.
Speech-to-text enginesA runtime under Apache-2.0 from the next-generation Kaldi team that runs speech to text, text to speech, speaker diarization and voice activity detection locally on CPUs, NPUs and phones, with no internet connection.
Speech-to-text enginesAn Apache-2.0 framework for real-time voice agents that connects any speech, language and voice provider, with semantic turn detection, native MCP tool support, a test framework, and an open-source media and SIP stack it can run on entirely.
Voice agent frameworksA real-time conversational AI framework from Agora with a visual designer, a SIP extension and turn detection, released under Apache-2.0 plus conditions that forbid competing deployments, which makes it source available rather than open source.
Voice agent frameworksAn on-device speech toolkit under MIT built for live streaming, with speech-to-text models trained from scratch in sizes down to 1 MB, one library across Python, the browser, phones and desktops.
Speech-to-text enginesA dictation and meeting-notes app under MIT that runs Whisper or Parakeet locally or cloud models with your own key, detects Zoom, Teams and FaceTime calls, and labels speakers on device.
Voice dictation appsAn offline transcription app under MIT that records system audio and the microphone, supports Whisper, Parakeet and Nemotron models, labels speakers, and exports SRT, VTT, TXT, HTML, PDF, JSON and DOCX.
Meeting transcriptionA macOS dictation app under GPL-3.0 whose README names Superwhisper and Wispr Flow as the tools it replaces, with local models, per-app modes, a personal dictionary and an AI assistant mode.
Voice dictation appsA fast, local neural text-to-speech engine maintained by the Open Home Foundation, used by Home Assistant and the NVDA screen reader, under GPL-3.0.
Text to speech and voice cloningA Docker image that serves the 82-million-parameter Kokoro model through an OpenAI-compatible speech endpoint, with streaming, voice mixing and nine languages, all under Apache-2.0.
Text to speech and voice cloningA speech-to-text app inside the Epicenter monorepo under AGPL-3.0 that records, transcribes with a provider you choose, cloud or local, optionally polishes the text, and pastes it at the cursor.
Voice dictation appsAn OpenAI-compatible server under MIT for streaming transcription, translation and text to speech, loading Whisper, Kokoro and Piper models on demand, described by its maintainers as Ollama for speech models.
Speech-to-text enginesAn MIT voice-agent platform that defines an agent in a JSON file and orchestrates speech to text, a language model and text to speech over websockets, with Twilio and Plivo for phone calls, whose maintainers say they are looking for help.
Voice agent frameworksStars and release dates read from GitHub on September 16, 2026. Openness labels: The code, and the weights where the project is a model, carry a licence on the OSI approved list. The code is public, and the licence limits what you may do with it. Read the LICENSE file before commercial use. The model weights can be downloaded and run. The training data or code is closed, or the weights carry a restriction the code does not.
How a project gets in, and how it leaves
What counts
A public repository, a README, a commit in the last 90 days, and one paid tool it replaces for a use a reader can test. Star counts are shown as a dated fact and never used as a bar.
What the label means
Open source means an OSI licence on the code and the weights. Source available means the licence limits use. Open weights means a model you can download whose training data or code is closed.
What is never said
That a project is faster, better or cheaper than the tool it replaces. Prices for anyone. User counts. A roadmap. Every feature claim is the project's own, from its README.
Categories: Text to speech and voice cloning, Voice dictation apps, Meeting transcription, Speech-to-text engines, Voice agent frameworks.
Running one of these inside your own environment, with your own data and your own security rules, is the kind of work Reveneau does. Read how a forward deployed engagement works.