Cola DLM: A Hierarchical Latent Diffusion Language Model for Text Generation
Researchers have introduced Cola DLM, a novel hierarchical latent diffusion language model designed to overcome the limitations of traditional autoregressive text generation. Unlike standard models that rely on fixed left-to-right token prediction, Cola DLM employs a hierarchical information decomposition approach. It first establishes a stable text-to-latent mapping using a Text VAE, then models global semantic priors in continuous latent space via a block-causal Diffusion Transformer (DiT), and finally generates text through conditional decoding. This architecture separates global semantic organization from local textual realization, offering a flexible non-autoregressive inductive bias. Extensive experiments across eight benchmarks, involving strictly matched ~2B-parameter baselines and scaling curves up to 2000 EFLOPs, demonstrate the model's strong scaling behavior and efficiency. The findings suggest that hierarchical continuous latent prior modeling serves as a principled alternative to token-level language modeling, potentially enabling unified modeling across discrete text and continuous modalities while improving generation quality and representation learning.
Wire timeline
Cola DLM: A Hierarchical Latent Diffusion Language Model for Text Generation
Researchers have introduced Cola DLM, a novel hierarchical latent diffusion language model designed to overcome the limitations of traditional autoregressive text generation. Unlike standard models that rely on fixed left-to-right token prediction, Cola DLM employs a hierarchical information decomposition approach. It first establishes a stable text-to-latent mapping using a Text VAE, then models global semantic priors in continuous latent space via a block-causal Diffusion Transformer (DiT), and finally generates text through conditional decoding. This architecture separates global semantic organization from local textual realization, offering a flexible non-autoregressive inductive bias. Extensive experiments across eight benchmarks, involving strictly matched ~2B-parameter baselines and scaling curves up to 2000 EFLOPs, demonstrate the model's strong scaling behavior and efficiency. The findings suggest that hierarchical continuous latent prior modeling serves as a principled alternative to token-level language modeling, potentially enabling unified modeling across discrete text and continuous modalities while improving generation quality and representation learning.
cs.AI updates on arXiv.org