We tested every open Amharic speech recognition model.
Nobody had measured them on data the models couldn't have already seen. So we did, on 1,548 recordings where we knew exactly what was said.
Three things surprised us:
• The smaller model beat the bigger one
• A monolingual model beat the multilingual one
• Two fair ways of counting mistakes disagreed about the winner
Then we built a 538 MB Amharic language model that fixes three more words in every hundred, and put it online so you can talk to it.
🎙️ Try it (speak or upload audio):
https://huggingface.co/spaces/Chapimenge/amharic-asr-demo
📊 Benchmark, language model, code, and everything that went wrong:
https://huggingface.co/datasets/snapwre/amharic-asr-benchmark
📝 The full story:
https://dataset.et/blog/which-amharic-speech-model
Free to use. If you're building anything that listens to Amharic, take it.
Dataset.ET ·
Forwarded fromChapi Dev Talks
3
1September 6, 2026 215 1 2