EverydayMMQA: A Multilingual and Multimodal Framework for Culturally Grounded Spoken Visual QA
Researchers have introduced EverydayMMQA, a scalable semi-automatic framework designed to create localized spoken and visual question-answering resources. This initiative addresses the limitations of large-scale multimodal models in handling cultural context, everyday knowledge, and underrepresented languages. The framework produced OASIS, a large-scale dataset featuring approximately 0.92 million real images and 14.8 million question-answer pairs. OASIS includes 3.7 million spoken questions, comprising 383 hours of human-recorded speech and 20,000 hours of voice-cloned speech from 42 speakers. The dataset supports four input modalities: text-only, speech-only, text-plus-image, and speech-plus-image. It specifically targets English and various Arabic dialects across 18 countries, covering both Modern Standard Arabic and regional variations. By focusing on pragmatic, commonsense, and culturally grounded reasoning rather than simple object recognition, the project aims to improve model performance in real-world scenarios. The team benchmarked several closed-source, open-source, and fine-tuned models against this new standard. Both the EverydayMMQA framework and the OASIS dataset are being made publicly available to the research community via Hugging Face to foster further development in multilingual and multimodal AI systems.
Wire timeline
EverydayMMQA: A Multilingual and Multimodal Framework for Culturally Grounded Spoken Visual QA
Researchers have introduced EverydayMMQA, a scalable semi-automatic framework designed to create localized spoken and visual question-answering resources. This initiative addresses the limitations of large-scale multimodal models in handling cultural context, everyday knowledge, and underrepresented languages. The framework produced OASIS, a large-scale dataset featuring approximately 0.92 million real images and 14.8 million question-answer pairs. OASIS includes 3.7 million spoken questions, comprising 383 hours of human-recorded speech and 20,000 hours of voice-cloned speech from 42 speakers. The dataset supports four input modalities: text-only, speech-only, text-plus-image, and speech-plus-image. It specifically targets English and various Arabic dialects across 18 countries, covering both Modern Standard Arabic and regional variations. By focusing on pragmatic, commonsense, and culturally grounded reasoning rather than simple object recognition, the project aims to improve model performance in real-world scenarios. The team benchmarked several closed-source, open-source, and fine-tuned models against this new standard. Both the EverydayMMQA framework and the OASIS dataset are being made publicly available to the research community via Hugging Face to foster further development in multilingual and multimodal AI systems.
cs.AI updates on arXiv.org