Lexical Acoustic Coding: Transmitting Sound via Natural Language
Researchers Emanuele Rossi and Emanuele Rodolà have introduced Lexical Acoustic Coding (LAC), a novel framework enabling the transmission of audio data through natural language. Published on arXiv, this study addresses the limitation of current systems where text describes but does not carry audio. In the LAC framework, pre-trained Large Language Model (LLM) agents act as sender and receiver. The sender analyzes input waveforms into interpretable acoustic descriptors, quantizes them using a specific vocabulary, and converts them into English sentences. The receiver parses these lexical codes to reconstruct the waveform through closed-loop refinement. This method frames text as both a descriptive caption and the actual transport representation for sound. Experiments involving short sounds and symbolic music demonstrate that plain text can preserve measurable acoustic structure while remaining editable and native to LLM communication. The research highlights trade-offs between vocabulary size, transmission rate, and fidelity, positioning LAC as a finite-rate lossy quantizer. This development signifies a potential shift in how audio systems are controlled and represented, leveraging the interpretability of natural language for efficient, human-readable audio transmission.
Wire timeline
Lexical Acoustic Coding: Transmitting Sound via Natural Language
Researchers Emanuele Rossi and Emanuele Rodolà have introduced Lexical Acoustic Coding (LAC), a novel framework enabling the transmission of audio data through natural language. Published on arXiv, this study addresses the limitation of current systems where text describes but does not carry audio. In the LAC framework, pre-trained Large Language Model (LLM) agents act as sender and receiver. The sender analyzes input waveforms into interpretable acoustic descriptors, quantizes them using a specific vocabulary, and converts them into English sentences. The receiver parses these lexical codes to reconstruct the waveform through closed-loop refinement. This method frames text as both a descriptive caption and the actual transport representation for sound. Experiments involving short sounds and symbolic music demonstrate that plain text can preserve measurable acoustic structure while remaining editable and native to LLM communication. The research highlights trade-offs between vocabulary size, transmission rate, and fidelity, positioning LAC as a finite-rate lossy quantizer. This development signifies a potential shift in how audio systems are controlled and represented, leveraging the interpretability of natural language for efficient, human-readable audio transmission.
cs.AI updates on arXiv.org