MEMENTO Evolves Robot Policies as Code for Complex Tasks.

Alkis Sygkounas, Victor Aregbede, Amy Loutfi, Andreas Persson· July 28, 2026 View original

Summary

MEMENTO is a memory-guided memetic framework that evolves "code-as-policy" for long-horizon embodied tasks, outperforming existing methods in robot manipulation and household interaction. It uses an evolved rollout evaluator and feedback-conditioned policy proposals to achieve higher task success and generalization, even transferring policies to physical robots.

Developing policies for robots to perform complex, long-horizon tasks, where success is only observable after many dependent actions, is a significant challenge. Representing these policies as executable control programs ("code-as-policy") allows for inspection and revision of their decision logic based on rollout evaluations. This research introduces MEMENTO, a memory-guided single-elite memetic framework designed for code-as-policy evolution. Unlike methods that select from independently generated variants, MEMENTO incorporates a sequential local improvement phase. It first evolves a rollout evaluator that maps policy executions to scalar fitness and structured feedback metrics. This fitness guides the selection of candidates and the next "elite" policy, while the feedback metrics inform memory-guided hill-climbing, macro-mutation, and crossover operations to generate new policy proposals. MEMENTO was evaluated on challenging embodied domains, including Robosuite Franka Tower-of-Hanoi manipulation and AI2-THOR household interaction, where it significantly outperformed baseline evolutionary methods. The best-evolved Robosuite policy was successfully deployed on a physical Franka robot, demonstrating the potential for sim-to-real transfer.

Why it matters

Robotics engineers and AI researchers can leverage MEMENTO to develop more robust, generalizable, and interpretable policies for complex robotic tasks, accelerating the deployment of autonomous systems in real-world environments.

How to implement this in your domain

  1. 1Evaluate current robot policy generation methods for long-horizon tasks and their interpretability.
  2. 2Explore the "code-as-policy" paradigm for developing more transparent and revisable robot behaviors.
  3. 3Investigate integrating memetic algorithms and memory-guided search into policy evolution frameworks.
  4. 4Pilot MEMENTO-like approaches in simulation environments for complex manipulation or interaction tasks.
  5. 5Develop strategies for sim-to-real transfer of evolved code-as-policy, including robust testing and validation.

Who benefits

RoboticsManufacturingLogisticsHealthcareSmart Homes

Key takeaways

  • MEMENTO evolves robot policies as executable code for long-horizon embodied tasks.
  • It uses a memory-guided memetic framework with an evolved rollout evaluator and feedback.
  • The approach significantly outperforms baselines in task success and generalization.
  • Evolved policies can be successfully transferred from simulation to physical robots.

Original post by Alkis Sygkounas, Victor Aregbede, Amy Loutfi, Andreas Persson

"arXiv:2607.22832v1 Announce Type: new Abstract: Long-horizon embodied tasks require policies that execute many dependent actions before task success can be observed. Representing policies as executable control pro- grams (code-as-policy) enables their decision logic to be inspect…"

View on X

Originally posted by Alkis Sygkounas, Victor Aregbede, Amy Loutfi, Andreas Persson on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses