Drum Synthesis from Expressive Drum Grids via Neural Audio Codecs
Researchers have proposed a novel system for generating realistic drum audio directly from symbolic representations, addressing a complex challenge at the intersection of music perception and machine learning. The method transforms an expressive drum grid, which is a time-aligned MIDI representation containing microtiming and velocity data, into high-quality drum audio. This is achieved by predicting discrete codes of a neural audio codec using a Transformer-based model. These predicted codec tokens are subsequently converted into waveform audio through a pre-trained codec decoder. The study experiments with several state-of-the-art neural codecs, including EnCodec, DAC, and X-Codec, to evaluate how different audio representations impact synthesis quality. Training and evaluation were conducted using the Expanded Groove MIDI Dataset (E-GMD), a extensive collection of human drum performances with paired MIDI and audio files. Objective metrics were employed to assess the fidelity and musical alignment of the generated outputs. The results demonstrate that codec-token prediction is an effective approach for drum grid-to-audio generation, offering valuable insights for selecting audio tokenizers in percussive synthesis applications.
Wire timeline
Drum Synthesis from Expressive Drum Grids via Neural Audio Codecs
Researchers have proposed a novel system for generating realistic drum audio directly from symbolic representations, addressing a complex challenge at the intersection of music perception and machine learning. The method transforms an expressive drum grid, which is a time-aligned MIDI representation containing microtiming and velocity data, into high-quality drum audio. This is achieved by predicting discrete codes of a neural audio codec using a Transformer-based model. These predicted codec tokens are subsequently converted into waveform audio through a pre-trained codec decoder. The study experiments with several state-of-the-art neural codecs, including EnCodec, DAC, and X-Codec, to evaluate how different audio representations impact synthesis quality. Training and evaluation were conducted using the Expanded Groove MIDI Dataset (E-GMD), a extensive collection of human drum performances with paired MIDI and audio files. Objective metrics were employed to assess the fidelity and musical alignment of the generated outputs. The results demonstrate that codec-token prediction is an effective approach for drum grid-to-audio generation, offering valuable insights for selecting audio tokenizers in percussive synthesis applications.
cs.AI updates on arXiv.org