LMMs Struggle with Spatial Modality Transfer for GIS

Ivan Majic, Zexian Huang, Franziska H\"ubl, Krzysztof Janowicz, Meilin Shi, Mina Karimi, Zilong Liu, Alexandra Fortacz-Lazan· August 10, 2026 View original

Key takeaways

  • Autonomous GIS agents require seamless spatial information transfer between image and text modalities.
  • Current LMMs struggle significantly with this modality transfer task.
  • Robust geospatial understanding in LMMs depends on rigorous multi-modal alignment.
  • This limitation is a critical bottleneck for fully automated GIS workflows.

Who benefits

GISUrban PlanningEnvironmental ScienceLogisticsAgriculture

Summary

This research introduces a modality transfer task for Large Multimodal Models (LMMs) in GIS workflows, revealing that current LMMs struggle to seamlessly transfer spatial information between image and text modalities. This limitation is a critical bottleneck for achieving autonomous GIS agents.

The development of autonomous Geographic Information System (GIS) agents requires AI models to effectively process and understand spatial information across different modalities, much like humans do. While current AI models are improving in spatial reasoning, most research has focused on text-based inputs and outputs, contrasting with human GIS workflows that blend visual and textual information. This study proposes a "modality transfer" task to evaluate Large Multimodal Models (LMMs) on their ability to convert spatial information between images and text. In this task, an LMM describes an image of colored squares, and a second LMM then attempts to regenerate the original image from that textual description. The results indicate that even recent LMMs, such as those from OpenAI, struggle significantly with this transfer, highlighting a critical need for more robust multi-modal alignment to achieve strong geospatial understanding and truly autonomous GIS agents.

Why it matters

Professionals in GIS, urban planning, environmental science, and logistics relying on AI for spatial analysis need to be aware of the current limitations in LMMs' ability to seamlessly integrate visual and textual spatial data.

How to implement this in your domain

  1. 1Prioritize human oversight and manual verification for LMM-generated spatial data or analyses that involve modality transfer.
  2. 2Develop specialized datasets for LMM training that explicitly focus on aligning spatial information across image and text.
  3. 3Explore hybrid AI approaches that combine LMMs with traditional GIS tools for robust spatial processing.
  4. 4Advocate for and invest in research aimed at improving multi-modal alignment in LMMs for geospatial applications.

Original post by Ivan Majic, Zexian Huang, Franziska H\"ubl, Krzysztof Janowicz, Meilin Shi, Mina Karimi, Zilong Liu, Alexandra Fortacz-Lazan

"arXiv:2608.06948v1 Announce Type: new Abstract: AI models are becoming increasingly adept at understanding and processing spatial information, thereby facilitating agentic problem-solving in spatial tasks and workflows. However, most of the research on their spatial capabilities…"

View on X

Originally posted by Ivan Majic, Zexian Huang, Franziska H\"ubl, Krzysztof Janowicz, Meilin Shi, Mina Karimi, Zilong Liu, Alexandra Fortacz-Lazan on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses