Data Scaling Boosts High-Resolution Weather Forecasting Accuracy

Yang Zhao, Peisong Niu, Tian Zhou, Ziqing Ma, Guanlong Ma, Rong Jin, Huiling Yuan, Liang Sun· August 18, 2026 View original

Key takeaways

  • High-resolution weather forecasting is primarily limited by data availability, not model architecture.
  • Super-resolution is an effective method for synthesizing high-resolution data from coarser sources.
  • BaguanHR significantly improves ML-based weather forecast accuracy, outperforming traditional models.
  • Data scaling exhibits a power-law effect, leading to consistent RMSE reduction.

Who benefits

AgricultureLogisticsEnergyInsuranceDisaster Management

Summary

Researchers developed BaguanHR, a framework that synthesizes extensive 0.1° high-resolution weather data from coarser ERA5 reanalysis using super-resolution. This data-centric approach significantly improves ML-based weather forecasting, outperforming existing methods and IFS-HRES across most lead times.

The development of high-resolution global weather forecasting models, particularly those based on machine learning, has been significantly hampered by the scarcity of fine-grained data. While reanalysis data exists, it's typically at a coarser 0.25° resolution, and attempts to fine-tune models on limited 0.1° samples have proven insufficient due to irreversible information loss. This research introduces BaguanHR, a novel framework that shifts focus from model transfer to data transfer, addressing this bottleneck directly. BaguanHR leverages variable-wise super-resolution (SR) to synthesize vast amounts of 0.1° data from the widely available ERA5 reanalysis. The core insight is that SR is a more robust mechanism for resolution transfer than direct forecasting, exhibiting lower conditional entropy and input amplification. By generating this extensive synthetic-plus-real dataset, BaguanHR enables ML-based models to achieve unprecedented accuracy. The framework's performance is remarkable, surpassing both current ML-based methods and the traditional IFS-HRES model for over 85% of lead times within 72 hours. Furthermore, the study identified a clear power-law scaling effect, where doubling the data reduces RMSE by approximately 4.6-4.9% for 72-hour and 120-hour forecasts, respectively. These findings strongly suggest that data availability, rather than model architecture, is the primary constraint for high-resolution ML-based weather forecasting, and super-resolution offers a general solution.

Why it matters

This breakthrough significantly enhances the accuracy and resolution of weather forecasting, providing more precise predictions crucial for various industries and disaster preparedness.

How to implement this in your domain

  1. 1Explore integrating super-resolution techniques to augment existing low-resolution datasets for ML model training.
  2. 2Investigate BaguanHR's variable-wise SR approach for generating high-resolution data in other scientific domains.
  3. 3Evaluate the potential of data scaling strategies to improve predictive models in areas with data scarcity.
  4. 4Collaborate with meteorological experts to validate and deploy high-resolution ML forecasts.

Original post by Yang Zhao, Peisong Niu, Tian Zhou, Ziqing Ma, Guanlong Ma, Rong Jin, Huiling Yuan, Liang Sun

"arXiv:2608.14652v1 Announce Type: new Abstract: The development of 0.1$^{\circ}$ global weather forecasting models based on machine learning (ML) is constrained by the limited availability of high-resolution data, as decades of reanalysis are only available at 0.25$^{\circ}$ reso…"

View on X

Originally posted by Yang Zhao, Peisong Niu, Tian Zhou, Ziqing Ma, Guanlong Ma, Rong Jin, Huiling Yuan, Liang Sun on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses