New Protocol Evaluates Temporal Fidelity of Synthetic Sequential Data

Kiwan Kwon, Kangmin Kim, Hojin Lee, Yeseong Jung, Hyeongwoo Kong, Vamsi K. Potluru, Saerom Park, Yongjae Lee· July 20, 2026 View original

Summary

Researchers introduce a taxonomy-guided evaluation protocol to assess the temporal fidelity of synthetic sequential tabular data, addressing limitations of conventional methods that fail to detect issues like backward timestamps or unrealistic entity trajectories. The protocol measures timestamp validity, cross-sectional structure, within-entity dynamics, and time-varying relational structure.

This paper highlights a critical flaw in how synthetic sequential tabular data is often evaluated: conventional methods, which treat records as static distributions, fail to capture temporal inconsistencies. Generative models might produce data that looks statistically similar but contains timestamps running backward, repeated entries, or entity paths that are entirely unrealistic. To address this, the authors propose a new, taxonomy-guided evaluation protocol. The protocol begins by characterizing each dataset based on four properties: time representation, sampling regularity, trajectory dependencies, and schema links. These properties then dictate which specific evaluation dimensions are meaningful. The framework measures several aspects of temporal fidelity, including the validity of timestamps, the cross-sectional structure at aligned time points, the dynamics within individual entities, and how relational structures evolve over time. It also redefines utility and privacy evaluation to focus on trajectories rather than isolated rows. Applying this protocol to eight generative models across thirteen datasets revealed significant discrepancies between rankings based on conventional evaluation and those based on temporal evaluation. The observed failures were systematic, not random, underscoring that temporal fidelity must be directly measured along the time axis.

Why it matters

For professionals using synthetic data, especially in privacy-preserving scenarios, ensuring temporal fidelity is crucial for the data's utility and reliability in downstream analytical tasks and model training.

How to implement this in your domain

  1. 1Adopt the proposed time-aware evaluation protocol when generating or using synthetic sequential data.
  2. 2Review your current synthetic data generation pipelines to ensure they account for temporal consistency.
  3. 3Benchmark your generative models using the new protocol to identify and address temporal fidelity gaps.
  4. 4Prioritize generative models that demonstrate strong temporal fidelity for privacy-preserving data sharing.
  5. 5Educate your data science and engineering teams on the importance of temporal evaluation for sequential data.

Who benefits

HealthcareFinanceRetailLogisticsAI/ML Research

Key takeaways

  • Conventional synthetic data evaluation often misses critical temporal inconsistencies.
  • A new taxonomy-guided protocol measures timestamp validity, dynamics, and relational structure.
  • Temporal fidelity must be directly measured, not inferred from static distributions.
  • Generative models perform differently when evaluated with a time-aware approach.

Original post by Kiwan Kwon, Kangmin Kim, Hojin Lee, Yeseong Jung, Hyeongwoo Kong, Vamsi K. Potluru, Saerom Park, Yongjae Lee

"arXiv:2607.15606v1 Announce Type: new Abstract: Synthetic sequential tabular data are increasingly used for privacy-preserving data sharing, yet a generator can reproduce every marginal and every foreign-key relationship while emitting timestamps that run backwards or repeat, and…"

View on X

Originally posted by Kiwan Kwon, Kangmin Kim, Hojin Lee, Yeseong Jung, Hyeongwoo Kong, Vamsi K. Potluru, Saerom Park, Yongjae Lee on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses