Shapley Values Enhance Data Masking for Privacy and Utility

Xinxue (Shawn), Qu, Francis Bilson Darku, Hong Guo· August 3, 2026 View original

Key takeaways

  • A new framework uses Shapley values for feature attribution in data masking.
  • It balances disclosure risk and data utility at the feature level.
  • The method is agnostic to specific masking techniques and evaluation metrics.
  • Experimental results show effective risk reduction while preserving utility.

Who benefits

BFSIHealthcareGovernmentMarketingResearch

Summary

This study proposes a novel framework using Shapley-value-based feature attribution to holistically manage the trade-off between disclosure risk and data utility in data masking, operating at the feature level.

The research introduces a new framework that applies Shapley-value-based feature attribution to the domain of data privacy, specifically for data masking. This approach aims to address the critical balance between minimizing disclosure risk and preserving data utility. Unlike existing literature that often focuses on this trade-off at the dataset level, this framework tackles it at the individual feature level. The proposed method is designed to be agnostic to the specific data masking techniques, statistical or machine learning methods, and evaluation metrics for utility and risk. Experimental results indicate that this Shapley-value-based framework effectively reduces disclosure risk while maintaining data utility. By providing a fair feature attribution, it offers a more granular and adaptable approach to privacy-preserving data transformations.

Why it matters

Professionals handling sensitive data can use this framework to make more informed decisions about data masking, ensuring better privacy protection without unduly sacrificing the analytical value of their datasets.

How to implement this in your domain

  1. 1Adopt the Shapley-value-based framework to assess the privacy-utility trade-off for individual features in sensitive datasets.
  2. 2Integrate this feature attribution method into existing data masking pipelines to guide masking strategy.
  3. 3Develop tools or scripts to calculate Shapley values for features in datasets requiring anonymization.
  4. 4Use the insights from feature-level attribution to prioritize which data elements to mask and to what extent.

Original post by Xinxue (Shawn), Qu, Francis Bilson Darku, Hong Guo

"arXiv:2607.28946v1 Announce Type: new Abstract: Despite its many benefits, widespread access to individuals' personal data also causes severe privacy concerns for consumers, companies, and policymakers. This study proposes a novel framework that adapts the Shapley-value-based fea…"

View on X

Originally posted by Xinxue (Shawn), Qu, Francis Bilson Darku, Hong Guo on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses