PopFS Selects Robust Features for Diverse Populations

Ruiqi Lyu, Alistair Turcan, Bryan Wilder· August 5, 2026 View original

Key takeaways

  • PopFS enables robust feature selection for heterogeneous populations using a single shared feature set.
  • The method balances overall predictive benefit with protection for less-benefited groups.
  • It scales effectively to thousands of candidate features by combining sparse learning and direct search.
  • PopFS improves fairness and performance across diverse demographics, as shown in COVID-19 nowcasting.

Who benefits

HealthcarePublic HealthFinancial ServicesGovernmentSocial Media

Summary

Researchers introduce PopFS, a method for selecting a single, shared feature set that remains robust across heterogeneous populations while allowing each population to train its own model. It uses a tunable welfare objective to balance overall predictive benefit with protection for less-benefited populations, demonstrating strong performance across various datasets.

A new research paper presents PopFS, a novel feature selection method designed to address the challenge of deploying models across diverse populations. Unlike standard approaches that optimize for a single large population or robust methods that force a shared model, PopFS learns one universal feature set that works effectively for multiple heterogeneous groups, allowing each group to develop its own specific predictive model. PopFS incorporates a customizable welfare objective, enabling practitioners to weigh the overall predictive accuracy against ensuring adequate performance for populations that might otherwise be underserved. To scale this objective, the method first uses multitask sparse learning to narrow down potential features, then directly searches for optimal feature sets by ranking additions and swaps, only fully refitting a small shortlist. The approach has shown consistent strong performance on both average and worst-case population metrics across eight population splits from six prediction tasks, including a COVID-19 nowcasting study where adjusting the welfare objective improved outcomes for the least-served states with minimal impact on overall performance.

Why it matters

This method allows organizations to build more equitable and robust AI systems by ensuring that a single set of collected features can serve diverse user groups effectively, improving fairness and utility across different demographics.

How to implement this in your domain

  1. 1Identify use cases where a single feature set must serve multiple distinct user populations.
  2. 2Evaluate current feature selection processes for potential biases or performance disparities across different groups.
  3. 3Explore integrating PopFS or similar generalized welfare optimization techniques into model development pipelines.
  4. 4Define and quantify "welfare" objectives relevant to your organization's fairness and performance goals for diverse populations.
  5. 5Pilot PopFS on a specific project to assess its ability to improve both average and worst-case performance metrics.

Original post by Ruiqi Lyu, Alistair Turcan, Bryan Wilder

"arXiv:2608.02887v1 Announce Type: new Abstract: Choosing which features to collect is a deployment decision: the same limited questionnaire, test panel, or sensor set may need to serve several heterogeneous populations. Standard feature-selection methods typically optimize for on…"

View on X

Originally posted by Ruiqi Lyu, Alistair Turcan, Bryan Wilder on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses