LiveHouse-TS: New Benchmark for Time Series Foundation Models

Haomin Wen, Ziyu Zhou, Qingxiang Liu, Siru Zhong, Yuxuan Liang· August 19, 2026 View original

Key takeaways

  • LiveHouse-TS is a new benchmark for evaluating Time Series Foundation Models in real-world, evolving environments.
  • Static benchmarks fail to capture model performance under distribution shifts and unexpected events.
  • Model rankings can dramatically change when evaluated continuously on live data.
  • Continuous temporal validity is crucial for reliable time series forecasting in dynamic settings.

Who benefits

FinanceRetailLogisticsEnergyManufacturing

Summary

LiveHouse-TS is introduced as the first open-world living benchmark infrastructure for Time Series Foundation Models (TSFMs), shifting evaluation from static snapshots to continuous temporal validity. It assesses model behavior on real future data, revealing that static model rankings dramatically reshuffle under live conditions due to evolving real-world environments.

Time Series Foundation Models (TSFMs) are emerging as a promising solution for cross-domain zero-shot forecasting. However, existing evaluation methods typically rely on static benchmarks with fixed historical data, which only provide a snapshot of average performance. These benchmarks fail to capture how models perform in dynamic, real-world environments characterized by seasonal changes, distribution shifts, and unexpected events. To address this critical gap, researchers have developed LiveHouse-TS, an innovative open-world living benchmark infrastructure specifically for TSFMs. Unlike static leaderboards, LiveHouse-TS continuously evaluates models prequentially on real future data, emphasizing continuous temporal validity over snapshot accuracy. This infrastructure is designed to answer crucial long-term scientific questions, such as the stability of model rankings and robustness under distribution shifts. Extensive streaming evaluations across 11 domains and 17 datasets using LiveHouse-TS demonstrated a dramatic reshuffling of static model rankings under live protocols. This highlights that models performing well on historical data may not maintain their superiority in continuously evolving real-world scenarios, underscoring the need for dynamic evaluation.

Why it matters

For professionals relying on time series forecasting, LiveHouse-TS provides a more realistic and robust way to evaluate models, ensuring they perform reliably in dynamic, real-world conditions rather than just on historical data. This can lead to more accurate predictions and better decision-making.

How to implement this in your domain

  1. 1Adopt dynamic, continuous evaluation protocols like LiveHouse-TS for your time series forecasting models.
  2. 2Re-evaluate your current TSFMs using real-time data streams to assess their robustness under distribution shifts.
  3. 3Prioritize TSFMs that demonstrate consistent performance in evolving environments, not just high accuracy on static benchmarks.
  4. 4Invest in monitoring systems that track model performance drift and trigger re-training or model switching in response to real-world changes.

Original post by Haomin Wen, Ziyu Zhou, Qingxiang Liu, Siru Zhong, Yuxuan Liang

"arXiv:2608.17299v1 Announce Type: new Abstract: Time Series Foundation Models (TSFMs) have recently emerged as a highly promising paradigm for cross-domain zero-shot forecasting. However, existing evaluation protocols predominantly rely on static benchmarks with fixed historical…"

View on X

Originally posted by Haomin Wen, Ziyu Zhou, Qingxiang Liu, Siru Zhong, Yuxuan Liang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI Engineering & DevTools