SiNFluD: Creating and Evaluating Figurative Language Dataset for Sindhi
Researchers have introduced SiNFluD, a novel benchmark dataset designed specifically for the classification of figurative language in the Sindhi language. Addressing the scarcity of resources for low-resource languages, the team collected raw text from diverse sources, including blogs, social media platforms, and literary works. The corpus was subsequently prepared and annotated by two native speakers using the Doccano text annotation tool, achieving a robust inter-annotator agreement score of 0.81. To establish performance baselines, the study utilized 5-fold and 10-fold cross-validation methods. The researchers evaluated several advanced pre-trained language models, including mBERT, XLM-RoBERTa, and XLM-RoBERTa-XL, alongside SetFit for few-shot fine-tuning of sentence transformers. Among the tested architectures, the pre-trained XLM-RoBERTa-XL model demonstrated superior performance. This contribution significantly advances natural language processing capabilities for Sindhi, providing essential tools for future linguistic analysis and AI development in South Asian languages.
Wire timeline
SiNFluD: Creating and Evaluating Figurative Language Dataset for Sindhi
Researchers have introduced SiNFluD, a novel benchmark dataset designed specifically for the classification of figurative language in the Sindhi language. Addressing the scarcity of resources for low-resource languages, the team collected raw text from diverse sources, including blogs, social media platforms, and literary works. The corpus was subsequently prepared and annotated by two native speakers using the Doccano text annotation tool, achieving a robust inter-annotator agreement score of 0.81. To establish performance baselines, the study utilized 5-fold and 10-fold cross-validation methods. The researchers evaluated several advanced pre-trained language models, including mBERT, XLM-RoBERTa, and XLM-RoBERTa-XL, alongside SetFit for few-shot fine-tuning of sentence transformers. Among the tested architectures, the pre-trained XLM-RoBERTa-XL model demonstrated superior performance. This contribution significantly advances natural language processing capabilities for Sindhi, providing essential tools for future linguistic analysis and AI development in South Asian languages.
cs.AI updates on arXiv.org