PoDAR: Power-Disentangled Audio Representation for Generative Modeling
Researchers have introduced PoDAR (Power-Disentangled Audio Representation), a novel framework designed to enhance the performance of audio latent diffusion models. The study addresses the challenge of latent space modelability by explicitly decoupling signal power from invariant semantic content through randomized power augmentation and a latent consistency objective. This factorization simplifies the latent space, leading to faster convergence and improved overall performance in downstream generative tasks. When integrated with a Stable Audio 1.0 VAE and an F5-TTS generator, PoDAR demonstrated a twofold acceleration in convergence speed compared to baselines. Additionally, it achieved significant improvements in quality metrics on the LibriSpeech-PC dataset, including a 0.055 increase in speaker similarity and a 0.22 rise in UTMOS scores. The approach also enables the exclusive application of Classifier-Free Guidance (CFG) to power-invariant content, extending stable guidance regimes to higher scales. This advancement highlights the importance of factor disentanglement in improving audio codec reconstruction fidelity and generator expressivity, offering a promising direction for future developments in high-quality audio synthesis and speech processing technologies.
Wire timeline
PoDAR: Power-Disentangled Audio Representation for Generative Modeling
Researchers have introduced PoDAR (Power-Disentangled Audio Representation), a novel framework designed to enhance the performance of audio latent diffusion models. The study addresses the challenge of latent space modelability by explicitly decoupling signal power from invariant semantic content through randomized power augmentation and a latent consistency objective. This factorization simplifies the latent space, leading to faster convergence and improved overall performance in downstream generative tasks. When integrated with a Stable Audio 1.0 VAE and an F5-TTS generator, PoDAR demonstrated a twofold acceleration in convergence speed compared to baselines. Additionally, it achieved significant improvements in quality metrics on the LibriSpeech-PC dataset, including a 0.055 increase in speaker similarity and a 0.22 rise in UTMOS scores. The approach also enables the exclusive application of Classifier-Free Guidance (CFG) to power-invariant content, extending stable guidance regimes to higher scales. This advancement highlights the importance of factor disentanglement in improving audio codec reconstruction fidelity and generator expressivity, offering a promising direction for future developments in high-quality audio synthesis and speech processing technologies.
cs.AI updates on arXiv.org