Interpreting AI Weather Models for Atmospheric Chemistry Predictions

Jason Y. Hu, Ivan Higuera-Mendieta, Patrick Obin Sturm, Makoto M. Kelp· July 24, 2026 View original

Summary

This study investigates Microsoft's Aurora model, an AI foundation model fine-tuned for atmospheric chemistry, to understand if it learns physical mechanisms or merely statistical regularities. It finds the model captures some chemical responses but lacks explicit physical constraints and internal representations remain meteorology-focused.

Weather forecasting foundation models (FMs) are increasingly being adapted to predict air quality, offering rapid global pollution forecasts with lower computational demands compared to traditional chemical transport models. These FMs are typically trained on reanalysis data and generate forecasts through an autoregressive process, without explicitly encoding physical or chemical governing processes. Consequently, high forecast accuracy alone doesn't reveal whether the model has genuinely learned underlying physical mechanisms or is simply exploiting statistical patterns within its training data. This research presents the first in-depth study aimed at understanding what an FM fine-tuned for atmospheric chemistry has actually learned, specifically examining Microsoft's Aurora model. The methodology involves imposing controlled chemical perturbations on its forecasts and comparing them against established photochemical relationships. Additionally, the internal representations that drive these forecasts are scrutinized. The findings indicate that Aurora can capture a first-order ozone response to reactive nitrogen, but it fails to enforce the explicit chemical constraints that a process-based model would. It produces chemically inconsistent combinations of related species and tends to smooth out localized emission features, such as wildfire plumes, towards background levels. Internally, the model's representations largely retain the meteorological organization inherited during its pretraining phase, showing little structure specifically related to chemistry. Using sparse autoencoders, the study identified internal components that causally influence chemical forecasts, but these components do not cleanly map to individual atmospheric processes. This work establishes a framework for evaluating whether AI forecasting systems truly learn atmospheric chemistry from reanalysis data, emphasizing that composition forecasts should be judged by their internal mechanisms, not just benchmark skill, especially as these models increasingly inform environmental policy.

Why it matters

Professionals relying on AI for critical environmental predictions need to understand the underlying mechanisms, not just the output, to ensure policy decisions are based on robust, physically consistent models rather than statistical correlations.

How to implement this in your domain

  1. 1Demand transparency and mechanistic interpretability from AI models used for critical environmental or scientific predictions.
  2. 2Implement rigorous testing frameworks that go beyond benchmark skill, including controlled perturbations and consistency checks against known physical laws.
  3. 3Invest in tools and techniques for model interpretability, such as sparse autoencoders, to understand internal representations of AI systems.
  4. 4Collaborate with domain experts (e.g., atmospheric chemists) to define and validate physical constraints that AI models should adhere to.
  5. 5Develop hybrid modeling approaches that combine the speed of AI with the physical consistency of process-based models for high-stakes applications.

Who benefits

Environmental MonitoringClimate SciencePublic HealthGovernment PolicyEnergy

Key takeaways

  • AI foundation models for atmospheric chemistry may achieve high skill without learning true physical mechanisms.
  • Microsoft's Aurora model captures some chemical responses but lacks explicit physical constraints and consistency.
  • Internal representations of such models often remain tied to pretraining data (meteorology) rather than chemistry.
  • Mechanistic interpretability is crucial for AI models informing environmental policy, beyond just forecast skill.

Original post by Jason Y. Hu, Ivan Higuera-Mendieta, Patrick Obin Sturm, Makoto M. Kelp

"arXiv:2607.20778v1 Announce Type: new Abstract: Weather forecasting foundation models (FMs) are increasingly fine-tuned to predict air quality, offering fast global pollution forecasts at lower computational cost than conventional chemical transport models. These FMs are typicall…"

View on X

Originally posted by Jason Y. Hu, Ivan Higuera-Mendieta, Patrick Obin Sturm, Makoto M. Kelp on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses