EXIMO Improves Robot Policy Finetuning with VLM-Guided Exploration.

Bhavya Sukhija, Oliver Groth, Mohit Shridhar, Tim Hertweck, Michael Bloesch, Markus Wulfmeier, Abbas Abdolmaleki, Martin Riedmiller· August 21, 2026 View original

Key takeaways

  • EXIMO offers a novel three-stage algorithm for efficient finetuning of VLA robot policies.
  • It uses a VLM for intelligent planning and orchestrated data collection, improving sample efficiency.
  • The method combines exploration, imitation, and residual reinforcement learning.
  • EXIMO significantly outperforms prior approaches in learning new robotic tasks.

Who benefits

RoboticsManufacturingLogisticsHealthcareAgriculture

Summary

This paper introduces EXIMO, an algorithm that significantly enhances the efficiency of finetuning Vision-Language-Action (VLA) robot policies by integrating a Vision Language Model (VLM) for planning and data collection. It combines exploration, imitation, and optimization stages to overcome challenges in learning new robotic tasks.

Current robotic manipulation policies rely heavily on behavior cloning from large VLA models, but finetuning these for new tasks remains difficult due to the high cost of data collection and the sample inefficiency of reinforcement learning (RL). EXIMO addresses this by proposing a three-stage approach. First, an exploration phase uses a VLM as a planner to break down complex tasks and guide data collection for the VLA. Second, an imitation phase finetunes the VLA using this newly orchestrated dataset. Finally, an optimization stage applies residual off-policy RL for further refinement. Experiments demonstrate that EXIMO significantly outperforms existing methods in both sample efficiency and final performance, making it a promising advancement for teaching robots new skills more quickly.

Why it matters

Professionals in robotics and automation can leverage this research to develop more agile and adaptable robotic systems, reducing the time and resources needed to deploy robots for novel tasks.

How to implement this in your domain

  1. 1Investigate integrating VLM-guided exploration into existing robot learning pipelines.
  2. 2Experiment with EXIMO's three-stage finetuning process on specific robotic manipulation challenges.
  3. 3Evaluate the sample efficiency gains and performance improvements for new task acquisition.
  4. 4Consider how to adapt VLM planning capabilities for complex, long-horizon robotic operations.

Original post by Bhavya Sukhija, Oliver Groth, Mohit Shridhar, Tim Hertweck, Michael Bloesch, Markus Wulfmeier, Abbas Abdolmaleki, Martin Riedmiller

"arXiv:2608.19891v1 Announce Type: new Abstract: How to efficiently finetune robot policies to learn new tasks on the fly? State of the art robotic manipulation policies are based on behaviour cloning of large vision-language-action (VLA) models with billions of parameters on huge…"

View on X

Originally posted by Bhavya Sukhija, Oliver Groth, Mohit Shridhar, Tim Hertweck, Michael Bloesch, Markus Wulfmeier, Abbas Abdolmaleki, Martin Riedmiller on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI Research

AI ResearchAI Engineering & DevTools

Decoding Silent Reading from Non-Invasive EEG

This research demonstrates that open-vocabulary word-level and semantic information can be reliably decoded from non-invasive EEG during silent reading. Using a contrastive decoder and a large dataset from a single participant, the study shows decoding scales log-linearly with training data and extends to rare words.

Ingo Marquardt, Anthilia Alchanat, Priyanka JainAug 21, 2026
AI ResearchAI Engineering & DevTools

Exact Learning Coefficients for Singular Models

This paper presents the first deterministic algorithm for exactly computing local learning coefficients (Real Log Canonical Thresholds) for two-dimensional singular models. This breakthrough provides ground truth for calibrating sampling-based estimators and reveals algebraic structure in learning coefficients, outperforming sampling in shallow regimes.

Gr\'egoire Sergeant-Perthuis (CQSB, Sorbonne Universit\'e), Elias Tsigaridas (Ouragan Team, INRIA), Jules Tsukahara (Ouragan Team, INRIA)Aug 21, 2026
AI Engineering & DevToolsAI Research

Standardized ML Evaluation for Power System Protection

This paper proposes a standardized framework for evaluating machine learning applications in power system protection, addressing inconsistencies in current research. It defines seven critical study dimensions and instantiates the framework with a case study on fault classification and localization using a public benchmark.

Julian Oelhaf, Georg Kordowich, Paula Andrea P\'erez-Toro, Christian Bergler, Johann J\"ager, Andreas Maier, Siming BayerAug 21, 2026