LLMs Show Limits in Predicting Neighborhood Mobility Patterns

Saad Mohammad Abrar, Eesha Kurella, Arnav Dadarya, Naman Awasthi, Kazi Tasnim Zinat, Vanessa Frias-Martinez· September 2, 2026 View original

Key takeaways

  • LLMs currently underperform supervised models in predicting neighborhood-level human mobility.
  • LLMs rely on coarse, stable priors that can introduce biases, especially concerning protected groups.
  • Auditing LLM predictions for empirical alignment and bias is crucial before deployment.
  • Combining LLMs with traditional methods might offer a more robust approach for urban analytics.

Who benefits

Urban PlanningTransportationPublic HealthGovernment

Summary

A study evaluates zero-shot LLMs for predicting aggregate neighborhood-level human mobility, finding them less accurate than supervised models and often relying on coarse, biased priors. While LLMs can partially recover patterns, their predictions lack structural grounding and require careful auditing for bias.

Researchers investigated the capability of large language models to predict human mobility patterns at a neighborhood level, a critical area for urban planning and public health. They compared zero-shot LLM performance against supervised baselines across four U.S. metropolitan areas, using anonymized mobility data alongside sociodemographic and built-environment factors. The study revealed that LLMs achieved significantly lower accuracy than supervised models, particularly struggling with spatial extent outcomes. A key finding was that LLMs tend to rely on broad, stable prior assumptions about predictor effects, which often remain consistent across different outcomes and cities. This includes an observed asymmetric treatment of protected-group predictors, indicating potential biases. Ultimately, while LLMs can identify some aggregate mobility patterns from urban context, their predictions are not inherently structurally sound. The research emphasizes the necessity of auditing LLM outputs for empirical alignment and potential biases before relying on them for sensitive applications.

Why it matters

Professionals in urban planning, transportation, and public health need to understand the capabilities and limitations of AI for critical decision-making, especially regarding sensitive data like human mobility. This research highlights that LLMs, while promising, are not yet reliable for fine-grained, unbiased predictions in these domains without rigorous auditing.

How to implement this in your domain

  1. 1Conduct thorough bias audits on LLM outputs before deploying them for sensitive urban planning or public health applications.
  2. 2Integrate LLM predictions with traditional supervised models, using LLMs for initial pattern identification and supervised models for refinement and accuracy.
  3. 3Develop domain-specific benchmarks and evaluation metrics to assess LLM performance in urban mobility prediction beyond general accuracy.
  4. 4Prioritize data privacy and ethical considerations when using any AI model for human mobility analysis, ensuring anonymization and compliance.

Original post by Saad Mohammad Abrar, Eesha Kurella, Arnav Dadarya, Naman Awasthi, Kazi Tasnim Zinat, Vanessa Frias-Martinez

"arXiv:2609.00345v1 Announce Type: new Abstract: Human mobility is central to urban planning, transportation, public health, and emergency response, yet fine-grained trajectory data are often proprietary, restricted, and privacy-sensitive. Large language models (LLMs) offer a pote…"

View on X

Originally posted by Saad Mohammad Abrar, Eesha Kurella, Arnav Dadarya, Naman Awasthi, Kazi Tasnim Zinat, Vanessa Frias-Martinez on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses