Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Speech Recognition
Researchers have introduced Bangla-WhisperDiar, a novel system designed to enhance Automatic Speech Recognition (ASR) and speaker diarization for the Bangla language. Addressing challenges such as long-form recordings, diverse acoustic conditions, and speaker variability, the team fine-tuned the tugstugi bengaliai regional ASR Whisper medium model using a custom dataset of approximately 15,000 audio segments. They employed extensive data augmentation techniques, including noise injection and reverb simulation. For speaker diarization, they fine-tuned the pyannote/segmentation-3.0 model within the PyTorch Lightning framework, integrating it into the pyannote/speaker-diarization-community-1 pipeline. The resulting ASR system achieved a Word Error Rate (WER) of 0.2441, while the diarization system recorded a Diarization Error Rate (DER) of 0.2392. These results demonstrate significant improvements over pretrained baselines. The study details a complete pipeline covering data preprocessing, text normalization, audio augmentation, training strategies, and inference optimization, offering a robust solution for Bangla spoken language understanding in complex acoustic environments.
Wire timeline
Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Speech Recognition
Researchers have introduced Bangla-WhisperDiar, a novel system designed to enhance Automatic Speech Recognition (ASR) and speaker diarization for the Bangla language. Addressing challenges such as long-form recordings, diverse acoustic conditions, and speaker variability, the team fine-tuned the tugstugi bengaliai regional ASR Whisper medium model using a custom dataset of approximately 15,000 audio segments. They employed extensive data augmentation techniques, including noise injection and reverb simulation. For speaker diarization, they fine-tuned the pyannote/segmentation-3.0 model within the PyTorch Lightning framework, integrating it into the pyannote/speaker-diarization-community-1 pipeline. The resulting ASR system achieved a Word Error Rate (WER) of 0.2441, while the diarization system recorded a Diarization Error Rate (DER) of 0.2392. These results demonstrate significant improvements over pretrained baselines. The study details a complete pipeline covering data preprocessing, text normalization, audio augmentation, training strategies, and inference optimization, offering a robust solution for Bangla spoken language understanding in complex acoustic environments.
cs.AI updates on arXiv.org