MixtureTT: Diffusion-Based Style Transfer for Polyphonic Music Stems
Researchers have introduced MixtureTT, a novel system designed for flexible per-stem timbre transfer directly from polyphonic musical mixtures. Unlike existing methods that rely on separate-then-transfer pipelines, which often propagate source separation artifacts and produce incoherent timbres, MixtureTT employs a joint stem diffusion transformer. This approach models dependencies across per-stem content and cross-stem harmonics, allowing for the simultaneous transfer of all stems to specified instruments through a shared diffusion process. The system effectively eliminates cascaded separation errors and significantly reduces inference costs by a factor equal to the number of stems. Evaluations conducted on the SATB choral dataset demonstrate that MixtureTT outperforms single-instrument baselines in both objective and subjective metrics, even under stricter input conditions. The study confirms that cross-stem modeling is essential for achieving coherent mixture-level timbre transfer, marking a significant advancement over naive pipeline approaches in audio processing and artificial intelligence.
Wire timeline
MixtureTT: Diffusion-Based Style Transfer for Polyphonic Music Stems
Researchers have introduced MixtureTT, a novel system designed for flexible per-stem timbre transfer directly from polyphonic musical mixtures. Unlike existing methods that rely on separate-then-transfer pipelines, which often propagate source separation artifacts and produce incoherent timbres, MixtureTT employs a joint stem diffusion transformer. This approach models dependencies across per-stem content and cross-stem harmonics, allowing for the simultaneous transfer of all stems to specified instruments through a shared diffusion process. The system effectively eliminates cascaded separation errors and significantly reduces inference costs by a factor equal to the number of stems. Evaluations conducted on the SATB choral dataset demonstrate that MixtureTT outperforms single-instrument baselines in both objective and subjective metrics, even under stricter input conditions. The study confirms that cross-stem modeling is essential for achieving coherent mixture-level timbre transfer, marking a significant advancement over naive pipeline approaches in audio processing and artificial intelligence.
cs.AI updates on arXiv.org