Vector Faculty Present TimbreTron Musical Style Transfer Model at ICLR
Faculty members from the Vector Institute for Artificial Intelligence, including Sageev Oore and Roger Grosse, along with their research team, presented a new musical style transfer model called TimbreTron at the International Conference on Learning Representations (ICLR). The project explores the intersection of artificial intelligence and music, aiming to transform audio recordings from one instrument to another while preserving pitch, rhythm, and expressiveness. The team utilized a novel pipeline combining Constant-Q Transform (CQT) spectrograms, CycleGAN for image-based style transfer, and Google DeepMind’s WaveNet for audio reconstruction. This approach allows users to change an instrument's timbre, such as converting a piano piece to sound like a harpsichord, without altering other musical elements. Additionally, the system enables independent adjustment of tempo and pitch, avoiding common audio artifacts. The research highlights both technical advancements in controlling audio generation through neural networks and creative possibilities for artists. By treating audio waveforms as images for processing, the team circumvented difficulties in modeling timbre directly. The project underscores the growing integration of deep learning in creative fields, offering new tools for musical expression and experimentation.
Wire timeline
Vector Faculty Present TimbreTron Musical Style Transfer Model at ICLR
Faculty members from the Vector Institute for Artificial Intelligence, including Sageev Oore and Roger Grosse, along with their research team, presented a new musical style transfer model called TimbreTron at the International Conference on Learning Representations (ICLR). The project explores the intersection of artificial intelligence and music, aiming to transform audio recordings from one instrument to another while preserving pitch, rhythm, and expressiveness. The team utilized a novel pipeline combining Constant-Q Transform (CQT) spectrograms, CycleGAN for image-based style transfer, and Google DeepMind’s WaveNet for audio reconstruction. This approach allows users to change an instrument's timbre, such as converting a piano piece to sound like a harpsichord, without altering other musical elements. Additionally, the system enables independent adjustment of tempo and pitch, avoiding common audio artifacts. The research highlights both technical advancements in controlling audio generation through neural networks and creative possibilities for artists. By treating audio waveforms as images for processing, the team circumvented difficulties in modeling timbre directly. The project underscores the growing integration of deep learning in creative fields, offering new tools for musical expression and experimentation.
Vector Institute for Artificial Intelligence