Yeti: A Compact Protein Structure Tokenizer for Multimodal Generation
Researchers have introduced Yeti, a novel and compact protein structure tokenizer designed to enhance multimodal models in computational biology. Published on arXiv, this tool addresses the limitations of existing tokenizers that often prioritize reconstruction over generative capabilities. Yeti utilizes lookup-free quantization and is trained end-to-end with a flow matching objective, converting continuous atomic coordinates into discrete representations suitable for transformer architectures. The model demonstrates superior codebook utilization and token diversity while achieving second-best reconstruction accuracy with ten times fewer parameters than comparable models like ESM3. To validate its effectiveness, the team trained a compact multimodal model from scratch using Yeti’s structure tokens and amino acid sequences. This model successfully generated plausible protein structures and sequences through unconditional cogeneration, matching the performance of significantly larger models. Yeti represents a significant advancement in creating efficient, expressive tools for integrating multimodal data, facilitating the design of new proteins with specific functional properties without requiring extensive pretrained initialization.
Wire timeline
Yeti: A Compact Protein Structure Tokenizer for Multimodal Generation
Researchers have introduced Yeti, a novel and compact protein structure tokenizer designed to enhance multimodal models in computational biology. Published on arXiv, this tool addresses the limitations of existing tokenizers that often prioritize reconstruction over generative capabilities. Yeti utilizes lookup-free quantization and is trained end-to-end with a flow matching objective, converting continuous atomic coordinates into discrete representations suitable for transformer architectures. The model demonstrates superior codebook utilization and token diversity while achieving second-best reconstruction accuracy with ten times fewer parameters than comparable models like ESM3. To validate its effectiveness, the team trained a compact multimodal model from scratch using Yeti’s structure tokens and amino acid sequences. This model successfully generated plausible protein structures and sequences through unconditional cogeneration, matching the performance of significantly larger models. Yeti represents a significant advancement in creating efficient, expressive tools for integrating multimodal data, facilitating the design of new proteins with specific functional properties without requiring extensive pretrained initialization.
cs.AI updates on arXiv.org