Bytephilosopher: post #1338 — TG.ME

Forwarded fromCHChapi Dev Talks
We're releasing the first batch of Dataset.ET open speech data for Amharic.

22.7 hours. 7,405 recordings. 320 speakers. Free, CC BY 4.0, on Hugging Face.

Amharic has close to no open speech data. That's the gap we're trying to close, and this is the first step rather than the finished thing.

Some honesty about what this is: it's a first batch. The audio isn't preprocessed yet. We're not going to tell you it's the best quality out there. We don't think it is. What we want is for people to actually use it and tell us where it falls short.

The next release will be different: properly preprocessed audio, a real data pipeline with testing and validation built in. This exists because 320 people in Ethiopia gave their voices to it, and because contributors reviewed each other's recordings. Thank you to all of them.

If you're working on Ethiopian languages, low-resource ASR, or you just want to talk about the data, DM us. We'd genuinely like to hear from you.

🔗 https://huggingface.co/datasets/snapwre/amharic-speech
huggingface.co
snapwre/amharic-speech · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
🔥5
August 26, 2026 117