Domain Adaptation Boosts Molecular Language Models for Discovery

Henrik Wille, Luis-Finley Sch\"utz, Felix Strieth-Kalthoff· August 19, 2026 View original

Key takeaways

  • Pretrained molecular language models (MLMs) show variable performance across different molecular discovery domains.
  • Molecular fingerprints can be a robust baseline, highlighting MLM domain-representation mismatch.
  • Explicit domain adaptation significantly improves MLM performance for specific target libraries.
  • Fine-tuning MLMs on target domain data enhances sample efficiency in virtual screening.

Who benefits

PharmaceuticalsBiotechnologyMaterials ScienceChemical EngineeringDrug Discovery

Summary

This study demonstrates that explicit domain adaptation significantly improves the performance of molecular language models (MLMs) for molecular discovery within specific virtual libraries. While native MLMs vary, fine-tuning them on target library structures consistently enhances sample efficiency, often outperforming molecular fingerprints.

Pretrained molecular language models (MLMs) are increasingly used as encoders to understand structure-property relationships in molecules. However, their effectiveness for molecular discovery, especially when moving beyond their initial pretraining domain, has been an open question. This research systematically benchmarks several MLMs across diverse virtual molecular libraries relevant to drug discovery, organic materials, and catalysis. The study found that the performance of native MLM embeddings varied substantially across different libraries, and surprisingly, traditional molecular fingerprints often provided a more consistently robust baseline. This variability suggested a potential mismatch between the MLM's pretraining domain and the target discovery domain. Crucially, the research shows that explicit domain adaptation significantly enhances representation performance. Fine-tuning MLM encoders using structures from the specific target virtual library consistently improved sample efficiency. In many benchmark tasks, these adapted encoders emerged as the top-performing representations, demonstrating that the quality of molecular representations is highly dependent on the target domain and that tailored adaptation is a promising strategy for efficient adaptive decision-making in virtual screening and self-driving laboratories.

Why it matters

Professionals in drug discovery, materials science, and chemical engineering can significantly accelerate their research and development processes by leveraging domain-adapted molecular language models. This leads to more efficient virtual screening and faster identification of promising compounds.

How to implement this in your domain

  1. 1Assess current molecular discovery workflows for opportunities to integrate AI-driven virtual screening.
  2. 2Explore the use of pretrained molecular language models as molecular encoders.
  3. 3Implement domain adaptation techniques by fine-tuning MLMs on data specific to your target molecular libraries.
  4. 4Benchmark the performance of domain-adapted MLMs against traditional molecular fingerprints and unadapted MLMs.
  5. 5Integrate the best-performing adapted models into your virtual screening or self-driving laboratory platforms.

Original post by Henrik Wille, Luis-Finley Sch\"utz, Felix Strieth-Kalthoff

"arXiv:2608.17567v1 Announce Type: new Abstract: Pretrained molecular language models are increasingly used as molecular encoders for learning structure-property relationships. However, their practical suitability for molecular discovery within and beyond their pretraining domain…"

View on X

Originally posted by Henrik Wille, Luis-Finley Sch\"utz, Felix Strieth-Kalthoff on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI Research