WMLLM: Self-Evolving AI Agents Optimize Black-Box Problems

Zhongzheng Li, Qingsong Ran, Shikun Feng, Nian Ran, Wenhao Li, Xiaoyuan Zhang, Yue Wang, Xiaoguang Zhao· September 3, 2026 View original

Key takeaways

  • WMLLM uses LLMs for "predict-then-act" world modeling to enhance black-box optimization.
  • The framework improves sample efficiency and final performance in complex search spaces.
  • It combines multi-turn refinement, population-based search, and reinforcement learning.
  • WMLLM achieved state-of-the-art results in multi-objective molecular optimization.

Who benefits

PharmaceuticalsMaterials ScienceChemical EngineeringManufacturing

Summary

WMLLM is a new self-evolving optimization agent framework that uses large language models for "predict-then-act" world modeling to improve black-box optimization. It refines its implicit world model and strategy through multi-turn refinement, population-based search, and reinforcement learning, achieving state-of-the-art results in molecular optimization.

Black-box optimization problems, characterized by vast and complex search spaces, often struggle with poor sample efficiency due to reliance on direct candidate generation or trial-and-error. This research introduces WMLLM, a self-evolving optimization agent framework that leverages large language models (LLMs) for a "predict-then-act" world modeling approach. The core idea is that LLMs can predict promising optimization directions before costly evaluations, significantly improving search efficiency. WMLLM operates by first predicting potential outcomes and then acting to generate candidates. This process is enhanced through agentic multi-turn refinement, population-based search, and reinforcement learning, allowing the agent to continuously refine both its internal world model and its optimization strategy during the search. Experimental results on black-box optimization tasks, particularly multi-objective molecular optimization, demonstrate that WMLLM achieves state-of-the-art performance and superior sample efficiency under limited evaluation budgets.

Why it matters

Professionals in fields requiring complex optimization, such as drug discovery or materials science, can leverage WMLLM to accelerate research and development by finding optimal solutions more efficiently.

How to implement this in your domain

  1. 1Explore integrating WMLLM's predict-then-act world modeling into existing black-box optimization workflows.
  2. 2Pilot WMLLM for specific multi-objective optimization challenges within R&D departments.
  3. 3Investigate how large language models can be fine-tuned to predict outcomes for domain-specific optimization problems.
  4. 4Collaborate with AI researchers to adapt the self-evolving agent framework to proprietary datasets and objectives.

Original post by Zhongzheng Li, Qingsong Ran, Shikun Feng, Nian Ran, Wenhao Li, Xiaoyuan Zhang, Yue Wang, Xiaoguang Zhao

"arXiv:2609.01608v1 Announce Type: new Abstract: Black-box optimization problems remain challenging because of large, weakly structured, and high-dimensional search spaces. Existing methods often suffer from poor sample efficiency because they rely on direct candidate generation o…"

View on X

Originally posted by Zhongzheng Li, Qingsong Ran, Shikun Feng, Nian Ran, Wenhao Li, Xiaoyuan Zhang, Yue Wang, Xiaoguang Zhao on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses