66th ISI World Statistics Congress

66th ISI World Statistics Congress

Trustworthy Population Inference in the Era of AI: Privacy, Granularity, and Generalizability in Survey and Data Science

Organiser

S
Dr Yajuan Si

Participants

  • L
    Dr Parthasarathi Lahiri
    (Chair)

  • S
    Prof. Yajuan Si
    (Presenter/Speaker)
  • Where Granularity Matters: Calibrating Subdomain Inference for Binary Outcomes

  • D
    Dr Diana Dilshanie Deepawansa
    (Presenter/Speaker)
  • Multidimensional Poverty Mapping

  • MW
    Matt Williams
    (Presenter/Speaker)
  • Privacy Amplification for Synthetic Data Using Range Restriction

  • JD
    Jörg Drechsler
    (Presenter/Speaker)
  • Synthetic Data in the Era of AI

  • E
    Prof. Michael R. Elliott
    (Presenter/Speaker)
  • Generalizing Causal Inference from Probability Samples with Modified Doubly Robust Estimators

  • Category: International Association of Survey Statisticians (IASS)

    Proposal Description

    AI-era data science has intensified a central statistical challenge: how to draw valid population conclusions from complex, incomplete, confidential, and integrated data sources. This invited session presents innovative methods for trustworthy population inference, emphasizing granular estimation, privacy protection, causal generalizability, uncertainty quantification, and policy relevance. These issues are increasingly urgent as governments, researchers, and organizations seek to use large-scale surveys, administrative records, synthetic data, and AI-enabled tools while maintaining scientific validity, public trust, and confidentiality.

    The session is important because many high-stakes decisions require reliable inference for populations and subpopulations that are difficult to measure directly. National averages can hide substantial local or subgroup variation; confidential data must be protected before they can be shared; observational studies must address confounding and selection; and AI-driven data synthesis can create attractive but potentially misleading data products if not evaluated statistically. By bringing these challenges together, the session highlights the essential role of statistical science in ensuring that AI-era data products remain interpretable, representative, private, and useful for decision-making.

    The session reflects the international breadth of ISI, with participants from Asia, Europe, and North America, and from government statistics, academia, and research organizations. It also supports diversity in gender, geography, career perspective, and statistical interests, spanning official statistics, survey methodology, Bayesian modeling, synthetic data, differential privacy, causal inference, public health, and poverty measurement.

    Chair: Partha Lahiri, Professor at the University of Maryland, is a leading expert in small area estimation and survey methodology, a Fellow of ASA and IMS, an elected member of ISI, and current President of IASS.

    Dilshanie Deepawansa, Director of the Sample Survey Division, Department of Census and Statistics, Sri Lanka, develops methods for poverty measurement and official statistics. Her talk proposes hierarchical Bayesian area-level models for estimating multidimensional poverty in small geographic areas, using census auxiliary data and complex survey design adjustments, with application to Uva Province, Sri Lanka.

    Jörg Drechsler, Head of Statistical Methods at the Institute for Employment Research, Germany, is an expert in confidentiality protection and synthetic data. His talk examines the rapid growth of synthetic data driven by generative AI and critiques the “three black boxes” problem: insufficient attention to source data, intended use, and downstream analysis validity.

    Matt Williams, Senior Research Statistician at RTI International, develops methods for privacy, survey inference, and synthetic data. His talk introduces range-restricted formal privacy standards for synthetic data that incorporate data owners’ beliefs about sensitive value ranges, providing privacy amplification while balancing utility.

    Michael Elliott, Professor of Biostatistics at the University of Michigan and Research Professor at the Institute for Social Research, studies missing data, survey inference, and causal inference. His talk develops sampling-weighted doubly robust estimators for causal inference using complex probability samples, with application to food assistance and household food insecurity.

    Yajuan Si, Research Associate Professor at the University of Michigan, develops Bayesian methods for survey inference, data integration, nonresponse adjustment, and confidentiality protection. She will present calibrated intervals for binary subdomain estimation, with application to COVID-19 subgroup infection rates.