ProteinOPD: Efficient Multi-Objective Preference Alignment for Protein Design
Researchers have introduced ProteinOPD, a novel framework designed to enhance protein design through effective and efficient preference alignment. While protein language models (PLMs) can generate designable sequences, aligning them with specific functional preferences often leads to catastrophic forgetting of pretrained knowledge and struggles with balancing competing objectives. ProteinOPD addresses these challenges by adapting a pretrained PLM into preference-specific teachers and distilling their knowledge into a shared student model using token-level On-Policy Distillation (OPD). This method aligns the student to a normalized geometric consensus of weighted teachers, ensuring bounded optimization even when objectives conflict. Extensive experiments demonstrate that ProteinOPD significantly improves target preference objectives without compromising the inherent designability of the proteins. Notably, the framework offers an eightfold training speedup compared to traditional reinforcement learning-based alignment methods. This advancement holds significant potential for synthetic biology and drug discovery, enabling more precise control over protein functions and properties while maintaining computational efficiency.
Wire timeline
ProteinOPD: Efficient Multi-Objective Preference Alignment for Protein Design
Researchers have introduced ProteinOPD, a novel framework designed to enhance protein design through effective and efficient preference alignment. While protein language models (PLMs) can generate designable sequences, aligning them with specific functional preferences often leads to catastrophic forgetting of pretrained knowledge and struggles with balancing competing objectives. ProteinOPD addresses these challenges by adapting a pretrained PLM into preference-specific teachers and distilling their knowledge into a shared student model using token-level On-Policy Distillation (OPD). This method aligns the student to a normalized geometric consensus of weighted teachers, ensuring bounded optimization even when objectives conflict. Extensive experiments demonstrate that ProteinOPD significantly improves target preference objectives without compromising the inherent designability of the proteins. Notably, the framework offers an eightfold training speedup compared to traditional reinforcement learning-based alignment methods. This advancement holds significant potential for synthetic biology and drug discovery, enabling more precise control over protein functions and properties while maintaining computational efficiency.
cs.AI updates on arXiv.org