Soul App Open-Sources SoulX-Podcast Model for Human-Like AI Audio
Soul AI Lab, the research division behind the popular social networking platform Soul App, has officially open-sourced its advanced voice podcast generation model, known as SoulX-Podcast. This new artificial intelligence tool is designed to create highly natural, human-like audio content, supporting multi-speaker and multi-turn dialogues in Mandarin, English, and various Chinese dialects, including Cantonese and Sichuanese. The model is capable of generating fluent conversations lasting over 60 minutes while maintaining consistent tone, rhythm, and emotional nuances such as laughter and sighs. A standout feature of SoulX-Podcast is its ability to perform zero-shot cross-dialect voice cloning, allowing for versatile and realistic voice synthesis without extensive training data. Upon its release, the model quickly gained traction within the developer community, briefly topping the trending Text-to-Speech (TTS) models list on Hugging Face, a prominent platform for machine learning resources. This move highlights Soul App's strategic expansion into generative AI technologies, aiming to enhance user engagement through innovative audio experiences. The open-source release allows developers and researchers worldwide to utilize and further refine the technology, potentially accelerating advancements in natural language processing and synthetic media creation within the tech industry.
Wire timeline
Soul App Open-Sources SoulX-Podcast Model for Human-Like AI Audio
Soul AI Lab, the research division behind the popular social networking platform Soul App, has officially open-sourced its advanced voice podcast generation model, known as SoulX-Podcast. This new artificial intelligence tool is designed to create highly natural, human-like audio content, supporting multi-speaker and multi-turn dialogues in Mandarin, English, and various Chinese dialects, including Cantonese and Sichuanese. The model is capable of generating fluent conversations lasting over 60 minutes while maintaining consistent tone, rhythm, and emotional nuances such as laughter and sighs. A standout feature of SoulX-Podcast is its ability to perform zero-shot cross-dialect voice cloning, allowing for versatile and realistic voice synthesis without extensive training data. Upon its release, the model quickly gained traction within the developer community, briefly topping the trending Text-to-Speech (TTS) models list on Hugging Face, a prominent platform for machine learning resources. This move highlights Soul App's strategic expansion into generative AI technologies, aiming to enhance user engagement through innovative audio experiences. The open-source release allows developers and researchers worldwide to utilize and further refine the technology, potentially accelerating advancements in natural language processing and synthetic media creation within the tech industry.
TechNode