GibbsTTS: Kinetic-Optimal Scheduling for Zero-Shot Text-to-Speech
Researchers have introduced GibbsTTS, a novel method for zero-shot text-to-speech (TTS) generation that addresses limitations in Metric-induced Discrete Flow Matching (MI-DFM). The study tackles two primary issues: the reliance on heuristic schedulers requiring hyperparameter search and finite-step path-tracking errors from first-order continuous-time Markov chain solvers. To resolve these, the team derived a kinetic-optimal scheduler that traverses probability paths at constant Fisher-Rao speed without training. Additionally, they implemented a finite-step moment correction to adjust jump probabilities while preserving destination distributions. Validated on codec-based zero-shot TTS using a unified architecture and large-scale dataset, GibbsTTS demonstrated superior objective naturalness compared to masked discrete generative baselines. Subjective evaluations favored GibbsTTS, and it exhibited strong speaker similarity, ranking highest on three of four test sets against state-of-the-art systems. This advancement significantly improves the efficiency and quality of discrete token generation in speech synthesis, offering a training-free numerical schedule that enhances both naturalness and speaker fidelity in zero-shot scenarios.
Wire timeline
GibbsTTS: Kinetic-Optimal Scheduling for Zero-Shot Text-to-Speech
Researchers have introduced GibbsTTS, a novel method for zero-shot text-to-speech (TTS) generation that addresses limitations in Metric-induced Discrete Flow Matching (MI-DFM). The study tackles two primary issues: the reliance on heuristic schedulers requiring hyperparameter search and finite-step path-tracking errors from first-order continuous-time Markov chain solvers. To resolve these, the team derived a kinetic-optimal scheduler that traverses probability paths at constant Fisher-Rao speed without training. Additionally, they implemented a finite-step moment correction to adjust jump probabilities while preserving destination distributions. Validated on codec-based zero-shot TTS using a unified architecture and large-scale dataset, GibbsTTS demonstrated superior objective naturalness compared to masked discrete generative baselines. Subjective evaluations favored GibbsTTS, and it exhibited strong speaker similarity, ranking highest on three of four test sets against state-of-the-art systems. This advancement significantly improves the efficiency and quality of discrete token generation in speech synthesis, offering a training-free numerical schedule that enhances both naturalness and speaker fidelity in zero-shot scenarios.
cs.AI updates on arXiv.org