AI NewsOpen sourceAnnouncement

Oído puts open-vocabulary speech-to-text on a $5 Wi-Fi chip, and reports 3.7% word error rate on LibriSpeech

Oído, a GPL-3.0 open-source speech-to-text engine from Lokutor that runs on a $5 ESP32-S3 microcontroller with no cloud and no neural accelerator, reports 3.7% word error rate on LibriSpeech test-clean, and the repository has 84 stars after four days.

AI News

Editorial2 min read

LinkedInX
GitHub repository card for lokutor-ai/oido, offline speech recognition for the ESP32-S3 microcontroller

Image: GitHub

Why it mattersA team building a voice-enabled device no longer has to send audio to a cloud API or buy a specialty neural chip, so a battery-powered gadget can understand speech on hardware the size of a thumb and keep the recording on the device.

Adding voice-to-text to a battery-powered gadget usually means either sending audio to a cloud API or buying a specialty neural accelerator chip. Oído, released by Lokutor on 30 September 2026 as a GPL-3.0 project, puts open-vocabulary speech-to-text on an ESP32-S3, a $5 Wi-Fi microcontroller with 8 MB of PSRAM, and reports a word error rate within a tenth of a point of the fp32 model it was quantized from.

Oído is a port of NVIDIA's Conformer-CTC Small to int8, running on a new C engine the authors wrote for the ESP32-S3's PIE vector unit. The repository is at github.com/lokutor-ai/oido and has 84 stars and 4 forks per the GitHub API on 4 October 2026. The models are on Hugging Face under the lokutor-ai organization.

The numbers Lokutor publishes

On LibriSpeech, Lokutor reports that the on-chip int8 engine scores 3.7 percent word error rate on test-clean and 8.2 percent on test-other, measured on the host build of the same C code the firmware compiles. With a 1.3 M-parameter GRU language model and CTC prefix beam search, those numbers fall to 3.3 and 7.2. The int4 model is 8.3 MB and leaves a 6 MB app partition on the chip for the rest of the device's code, at 4.6 and 10.0 percent WER.

Lokutor's own comparison in the repository puts those figures next to Whisper tiny.en (6.3 and 15.9 on a laptop in fp32), Moonshine tiny (5.0 and 12.1) and Espressif MultiNet7 (8.5 and 21.3 on the same chip, with the constraint that MultiNet7 takes fixed command lists rather than open speech). The repository says this is the most accurate LibriSpeech result Lokutor is aware of for any microcontroller, and acknowledges that Arm has shown Conformer models on Cortex-M55 chips paired with Ethos-U neural processors.

What is still estimated

Accuracy is measured, speed is not. The status note dated 2 October 2026 says real-time factor is estimated from exact Espressif QEMU emulator instruction counts, between 0.7 and 0.95 at 240 MHz, and that measurements on physical boards will be added in the next days. Twenty-two of 27 random clips give identical transcripts under the emulator and the laptop build; on five, last-bit rounding in the math library changes a word.

There is a Spanish model trained on 2,492 hours of audio from Common Voice, VoxPopuli, Multilingual LibriSpeech and FLEURS, reported at 13.3 WER on Common Voice with the language model. Lokutor flags that Oído was fine-tuned on the training splits of those corpora while Whisper is zero-shot, so the comparison favours Oído on those domains.

A streaming mode, trained with the same model as a starting point, chunks audio into 1.28 s blocks so partial text can appear while someone is still speaking. Lokutor estimates the time between the end of a sentence and the final text at 1.1 to 1.4 seconds in that mode.

Source

SourceGitHub

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX