SenWorld Creates Privacy-Safe Digital-Twin Evaluation Data

Zenghui Zhou, Xiaoyang Li, Xiaoxuan Qiao, Zhilang Wei, Tianming Lei· July 23, 2026 View original

Summary

SenWorld is a digital-twin simulation that generates context-rich, privacy-safe evaluation data with ground truth for smartphone personal assistants. It simulates personas living a day, archiving all signals, and exposing assistant failures without LLM judges.

This paper introduces SenWorld, a novel digital-twin simulation designed to generate context-rich evaluation data for smartphone personal assistants. The primary motivation is the need for privacy-safe data with known ground truth, as real device traces are too sensitive to share. SenWorld creates a physically grounded, deterministic, event-sourced simulation where personas navigate a full day in a world built from real-world map, weather, holiday, and network data. Every observable signal is meticulously archived in full-system snapshots, and evaluation cases are labeled by direct pointers to existing records, eliminating the need for post-hoc annotation or Large Language Model (LLM) judges. The method was evaluated with 16 personas in Beijing, demonstrating that the generated data closely matches real-user benchmarks in category distribution and daily communication rhythms. The simulation successfully exposed 78 failures in a production smartphone assistant, primarily related to call and SMS records, with the snapshot pointers confirming these as assistant-side retrieval errors. SenWorld offers a reproducible, privacy-preserving, and distribution-checked approach to generating high-quality evaluation data with inherent ground truth.

Why it matters

For developers of AI assistants and other context-aware systems, SenWorld provides a crucial method to generate realistic, privacy-safe evaluation data with verifiable ground truth, accelerating development and improving reliability.

How to implement this in your domain

  1. 1Explore using digital-twin simulations like SenWorld to generate synthetic, privacy-safe evaluation data for your AI products.
  2. 2Investigate integrating real-world environmental data (maps, weather, events) into your simulation environments for richer context.
  3. 3Design evaluation frameworks where ground truth is fixed by construction within the simulation, reducing reliance on manual annotation or LLM judges.
  4. 4Prioritize testing for context-aware retrieval errors in personal assistants, as highlighted by SenWorld's findings.

Who benefits

Software DevelopmentAI/ML PlatformsConsumer ElectronicsTelecommunicationsAutomotive (for in-car assistants)

Key takeaways

  • SenWorld generates privacy-safe, context-rich evaluation data via digital-twin simulation.
  • It provides ground truth by construction, eliminating manual annotation or LLM judges.
  • The simulated data closely matches real-user behavior patterns.
  • SenWorld effectively exposes failures in production smartphone assistants, particularly retrieval errors.

Original post by Zenghui Zhou, Xiaoyang Li, Xiaoxuan Qiao, Zhilang Wei, Tianming Lei

"arXiv:2607.19949v1 Announce Type: new Abstract: Smartphone personal assistants reason over longitudinal personal data, yet evaluating them requires context-rich evaluation data whose correct answers are known, and real device traces are too privacy-sensitive to share. To address…"

View on X

Originally posted by Zenghui Zhou, Xiaoyang Li, Xiaoxuan Qiao, Zhilang Wei, Tianming Lei on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses