AI Improves Antibody Design with Preference-Based Expression Ranking

Josh Qixuan Sun, Morteza Babaie, Wenyang Hou, Mark Crowley, David Young· July 21, 2026 View original

Summary

Researchers propose a preference-based learning framework that combines scarce quantitative data with large-scale weak supervision to effectively rank antibody expression, overcoming data scarcity in antibody design. This method adapts Direct Preference Optimization to protein language models, showing consistent improvements over baselines.

A significant challenge in antibody design is the scarcity of labeled experimental data, which hinders the development of effective models for ranking antibody expression. To address this, a new unified preference-based learning framework has been introduced. This framework intelligently integrates the limited available quantitative expression data with a vast amount of weakly positive supervision derived from immunization data. The core of this approach involves adapting Direct Preference Optimization (DPO) for use with protein language models. This adaptation includes a union-masked log-likelihood approximation and IMGT-based alignment, which together enable efficient training on antibody sequences of varying lengths. The method was evaluated on a diverse internal dataset of 1,254 labeled sequences and 4 million unlabeled camelid-derived antibodies. Results consistently showed that this preference-based learning framework outperformed existing baseline methods across most metrics, demonstrating its scalability and effectiveness in optimizing antibody expressibility even in data-constrained environments.

Why it matters

Professionals in biotechnology and pharmaceutical R&D can leverage this AI framework to accelerate antibody discovery and development by more accurately predicting antibody expression, reducing experimental costs and timelines.

How to implement this in your domain

  1. 1Explore integrating preference-based learning techniques into existing antibody design pipelines.
  2. 2Assess the availability and utility of large-scale weak supervision data within internal datasets.
  3. 3Collaborate with AI/ML teams to adapt protein language models for antibody expression ranking.
  4. 4Pilot the framework on a specific antibody development project to validate its performance and efficiency gains.
  5. 5Develop strategies for collecting and curating diverse data sources, including both quantitative and weak supervision.

Who benefits

BiotechnologyPharmaceuticalsHealthcareLife Sciences

Key takeaways

  • A new AI framework improves antibody expression ranking despite data scarcity.
  • It combines limited quantitative data with large-scale weak supervision.
  • The method adapts Direct Preference Optimization for protein language models.
  • This scalable solution accelerates antibody discovery and development.

Original post by Josh Qixuan Sun, Morteza Babaie, Wenyang Hou, Mark Crowley, David Young

"arXiv:2607.16263v1 Announce Type: new Abstract: Antibody expression ranking is a critical task in antibody design, yet its modelling is severely hindered by the scarcity of labeled experimental data. To address this, we propose a unified preference-based learning framework that i…"

View on X

Originally posted by Josh Qixuan Sun, Morteza Babaie, Wenyang Hou, Mark Crowley, David Young on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses