Reward Transport Controls Molecular Properties in Flow Matching.

Kehan Guo, Yili Shen, Yujun Zhou, Yue Huang, Chujie Gao, Shiyi Du, Xiangliang Zhang· July 13, 2026 View original

Key takeaways

  • Reward Transport enables direct control of molecular properties in flow matching models.
  • It aligns noise-space coordinates with molecular rewards during training.
  • This allows for steering generated distributions at inference without external guidance.
  • The method shows effective control over properties like logP and QED in molecular generation.

Who benefits

PharmaceuticalsBiotechnologyMaterials ScienceChemical EngineeringAI/ML Development

Summary

Reward Transport introduces a novel method to control molecular properties in flow matching models by aligning a scalar noise-space coordinate with molecular rewards during training. This allows for steering generated distributions at inference time without needing an oracle or gradient guidance.

Researchers have developed "Reward Transport," an innovative technique that leverages the coupling mechanism in flow matching models to embed controllable structure directly into learned flow fields. Traditionally, this coupling, which pairs noise vectors with data points, is seen as a computational detail. Reward Transport redefines this coupling as an alignment interface. By using optimal transport coupling during training, it aligns a scalar coordinate in the noise space with specific molecular rewards. This alignment enables precise control over generated molecular properties during inference. At inference, simply varying this noise-space coordinate allows for steering the generated distribution towards desired properties, eliminating the need for external oracles, reward models, or gradient guidance. The method has been empirically validated on ZINC-250K and GuacaMol datasets, demonstrating monotone control of logP and consistent QED control, with the same control knob producing opposite structural responses for different targets, confirming its specific property-steering capability.

Why it matters

This breakthrough offers a more direct and efficient way to design molecules with desired properties, accelerating drug discovery, materials science, and chemical engineering processes.

How to implement this in your domain

  1. 1Explore integrating Reward Transport into existing generative AI pipelines for molecular design or materials discovery.
  2. 2Identify specific molecular properties that are critical for current research or product development and define corresponding reward functions.
  3. 3Experiment with varying the noise-space coordinate during inference to generate molecules with a controlled range of desired properties.
  4. 4Validate the generated molecules through simulations or experimental synthesis to confirm the efficacy of the property control.

Original post by Kehan Guo, Yili Shen, Yujun Zhou, Yue Huang, Chujie Gao, Shiyi Du, Xiangliang Zhang

"arXiv:2607.08781v1 Announce Type: cross Abstract: The coupling in flow matching -- the rule pairing noise vectors with data points -- is typically treated as a computational choice. We show that this coupling can instead serve as an alignment interface: by matching noise and data…"

View on X

Originally posted by Kehan Guo, Yili Shen, Yujun Zhou, Yue Huang, Chujie Gao, Shiyi Du, Xiangliang Zhang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI Research

AI Engineering & DevToolsAI Research

Resilient Decentralized Federated Learning for Wireless IoT Networks

This paper introduces QEF-GT-AdamW, a communication-efficient and outage-resilient algorithm for decentralized federated learning over wireless IoT networks. It combines gradient tracking, AdamW optimization, and dual-stream biased quantization with error feedback to improve robustness and convergence under heterogeneous data and unreliable communication.

Nguyen Van Thieu, Ti Ti Nguyen, Ons Aouedi, Vu Nguyen Ha, Symeon ChatzinotasAug 27, 2026
AI Engineering & DevToolsAI Research

FedQoS Predicts QoS Risk for Wireless Access Selection

This paper proposes FedQoS, a federated QoS-risk learning framework that predicts future QoS degradation for reliable access selection in heterogeneous indoor-outdoor wireless environments. It enables access nodes to locally learn from network logs and collaboratively train a global predictor without centralizing user data, significantly reducing QoS failure rates.

Nguyen Van Thieu, Ti Ti Nguyen, Ons Aouedi, Zerihun Huruy, Vu Nguyen Ha, Symeon ChatzinotasAug 27, 2026
AI ResearchAI Engineering & DevTools

Parametric Knowledge Graphs Show Storage-Retrieval Gap

This paper explores compiling knowledge graphs into LoRA adapters for parametric memory, finding that while adapters effectively store factual knowledge, retrieving it via semantic similarity or weight-space geometry is ineffective. This highlights a "storage-retrieval gap" and the need for new query-conditioned composition mechanisms.

Martino M. L. Pulici, Cuong Xuan Chu, Evgeny Kharlamov, Volker TrespAug 27, 2026