WorldSpeech: A New Multilingual Speech Corpus for Improved ASR Accuracy
Researchers have introduced WorldSpeech, a comprehensive multilingual speech corpus designed to address the scarcity of aligned audio-transcript data for low-resource languages in Automatic Speech Recognition (ASR). While ASR systems perform well for high-resource languages, their accuracy significantly declines for others due to data limitations. WorldSpeech comprises 65,000 hours of 24 kHz aligned data across 76 languages, sourced from parliamentary proceedings, international broadcasts, and public-domain audiobooks. The dataset provides substantial coverage, with 37 languages having over 200 hours of data, 28 exceeding 500 hours, and 24 surpassing 1,000 hours. Experimental results demonstrate that fine-tuning existing ASR models on this corpus yields an average relative Word-Error-Rate reduction of 63.5% across 11 typologically diverse languages. This development marks a significant step forward in making speech recognition technology more inclusive and accurate for a broader range of global languages, potentially bridging the digital divide in voice technology accessibility.
Wire timeline
WorldSpeech: A New Multilingual Speech Corpus for Improved ASR Accuracy
Researchers have introduced WorldSpeech, a comprehensive multilingual speech corpus designed to address the scarcity of aligned audio-transcript data for low-resource languages in Automatic Speech Recognition (ASR). While ASR systems perform well for high-resource languages, their accuracy significantly declines for others due to data limitations. WorldSpeech comprises 65,000 hours of 24 kHz aligned data across 76 languages, sourced from parliamentary proceedings, international broadcasts, and public-domain audiobooks. The dataset provides substantial coverage, with 37 languages having over 200 hours of data, 28 exceeding 500 hours, and 24 surpassing 1,000 hours. Experimental results demonstrate that fine-tuning existing ASR models on this corpus yields an average relative Word-Error-Rate reduction of 63.5% across 11 typologically diverse languages. This development marks a significant step forward in making speech recognition technology more inclusive and accurate for a broader range of global languages, potentially bridging the digital divide in voice technology accessibility.
cs.AI updates on arXiv.org