Reward Transport Controls Molecular Properties in Flow Matching

Kehan Guo, Yili Shen, Yujun Zhou, Yue Huang, Chujie Gao, Shiyi Du, Xiangliang Zhang· July 13, 2026 View original

Key takeaways

  • Reward Transport enables direct control over molecular properties in flow matching models.
  • It aligns noise-space coordinates with molecular rewards during training.
  • Control is achieved during inference by varying a scalar coordinate, without extra computation.
  • The method shows effective and targeted control over properties like logP and QED.

Who benefits

PharmaceuticalsBiotechnologyMaterials ScienceChemical EngineeringAI/ML Engineering

Summary

Reward Transport is a new method that uses optimal transport coupling during flow matching training to align a noise-space coordinate with molecular rewards. This allows continuous, oracle-free control over generated molecular properties like logP and QED during inference.

This paper introduces "Reward Transport," a novel technique for controlling specific properties of generated molecules within flow matching models. Traditionally, the coupling mechanism in flow matching, which pairs noise vectors with data points, is viewed primarily as a computational choice. However, this research re-frames it as an alignment interface. By matching noise and data according to a desired molecular property, the method embeds controllable structure directly into the learned flow field. Reward Transport leverages optimal transport coupling during the training phase to align a scalar coordinate in the noise space with molecular rewards. This alignment enables a powerful capability during inference: by simply varying this noise-space coordinate, the generated molecular distribution can be steered to exhibit desired properties, such as logP or QED. Crucially, this control is achieved without needing an external oracle, a separate reward model, gradient guidance, or additional computational overhead during generation. The approach offers a principled, continuously adjustable "distribution-level control knob." Empirical results on standard datasets like ZINC-250K and GuacaMol demonstrate monotone control over logP and consistent control over QED. Interestingly, the same control knob produces opposite structural responses for different targets (e.g., growing molecules for logP but shrinking for QED), confirming it's not merely a generic size bias. This method is also complementary to existing techniques like classifier-free guidance.

Why it matters

For drug discovery and materials science, this provides a highly efficient and controllable method for generating molecules with specific desired properties, accelerating the design and optimization process.

How to implement this in your domain

  1. 1Explore integrating Reward Transport into your generative molecular design pipelines for targeted property control.
  2. 2Experiment with aligning different scalar noise-space coordinates to various molecular properties relevant to your research.
  3. 3Compare the efficiency and control capabilities of Reward Transport against existing conditional generation or reward-guided methods.
  4. 4Consider adapting the core concept of noise-space alignment for property control in other generative AI applications beyond molecules.

Original post by Kehan Guo, Yili Shen, Yujun Zhou, Yue Huang, Chujie Gao, Shiyi Du, Xiangliang Zhang

"arXiv:2607.08781v1 Announce Type: new Abstract: The coupling in flow matching -- the rule pairing noise vectors with data points -- is typically treated as a computational choice. We show that this coupling can instead serve as an alignment interface: by matching noise and data a…"

View on X

Originally posted by Kehan Guo, Yili Shen, Yujun Zhou, Yue Huang, Chujie Gao, Shiyi Du, Xiangliang Zhang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI Research

AI Engineering & DevToolsAI Research

Resilient Decentralized Federated Learning for Wireless IoT Networks

This paper introduces QEF-GT-AdamW, a communication-efficient and outage-resilient algorithm for decentralized federated learning over wireless IoT networks. It combines gradient tracking, AdamW optimization, and dual-stream biased quantization with error feedback to improve robustness and convergence under heterogeneous data and unreliable communication.

Nguyen Van Thieu, Ti Ti Nguyen, Ons Aouedi, Vu Nguyen Ha, Symeon ChatzinotasAug 27, 2026
AI Engineering & DevToolsAI Research

FedQoS Predicts QoS Risk for Wireless Access Selection

This paper proposes FedQoS, a federated QoS-risk learning framework that predicts future QoS degradation for reliable access selection in heterogeneous indoor-outdoor wireless environments. It enables access nodes to locally learn from network logs and collaboratively train a global predictor without centralizing user data, significantly reducing QoS failure rates.

Nguyen Van Thieu, Ti Ti Nguyen, Ons Aouedi, Zerihun Huruy, Vu Nguyen Ha, Symeon ChatzinotasAug 27, 2026
AI ResearchAI Engineering & DevTools

Parametric Knowledge Graphs Show Storage-Retrieval Gap

This paper explores compiling knowledge graphs into LoRA adapters for parametric memory, finding that while adapters effectively store factual knowledge, retrieving it via semantic similarity or weight-space geometry is ineffective. This highlights a "storage-retrieval gap" and the need for new query-conditioned composition mechanisms.

Martino M. L. Pulici, Cuong Xuan Chu, Evgeny Kharlamov, Volker TrespAug 27, 2026