Google Releases 'Waxal' Open-Source African Speech Dataset
Google has officially released the Waxal Dataset, a massive open-source collection comprising nearly 2 million speech records across key African languages including Amharic, Oromo, Tigrinya, and Swahili. Licensed under the permissive CC BY 4.0, this initiative allows startups and researchers to freely leverage the data to build and commercially deploy AI products in local tongues. The project was executed in collaboration with African institutions such as Addis Ababa University, aiming to bridge the language gap in AI.
ጎግል (Google) አማርኛ፣ ኦሮምኛ እና ትግርኛን ጨምሮ በተለያዩ የአፍሪካ ቋንቋዎች የተቀረጹ ወደ 2 ሚሊዮን የሚጠጉ የንግግር መረጃዎችን የያዘ "ዋክሳል" የተሰኘ ለህዝብ ክፍት የሆነ ዳታሴት ይፋ አደረገ። ይህ ስብስብ በ"CC BY 4.0" የፈቃድ አይነት የተለቀቀ በመሆኑ፣ ጀማሪ የንግድ ተቋማት እና አልሚዎች መረጃውን በነጻ በመጠቀም በሀገር በቀል ቋንቋዎች የሚሰሩ የኤአይ (AI) ምርቶችን ማልማት እና ለገበያ ማቅረብ ይችላሉ። እንደ አዲስ አበባ ዩኒቨርሲቲ ካሉ የአፍሪካ ተቋማት ጋር በመተባበር የተሰራው ይህ ፕሮጀክት፣ ለድምፅ ቴክኖሎጂ ወሳኝ የሆነ የመረጃ ምንጭ በማቅረብ በአፍሪካ ቋንቋዎች ላይ ያለውን የቴክኖሎጂ ክፍተት ለመሙላት ያለመ ነው።
@webthreeth
Forwarded fromWeb 3.0 Ethiopia - DeFi & AI

2
2
2February 5, 2026 390 1