66th ISI World Statistics Congress

66th ISI World Statistics Congress

Novel Methods for Inference with Deep Learning and Complex Biomedical Data

Organiser

SS
Stephen Salerno

Participants

  • YL
    PROF. DR. Yi Li
    (Presenter/Speaker)
  • Deep Learning of Semi-Competing Risk Data via a New Neural Expectation-Maximization Algorithm

  • SS
    Dr Stephen Salerno
    (Presenter/Speaker)
  • A Pseudo-Value Approach to Causal Deep Learning of Semi-Competing Risks

  • A
    PROF. DR. Syed Ejaz Ahmed
    (Presenter/Speaker)
  • Leveraging Weak Signals for Better Predictions: A Post-Shrinkage Approach in High Dimensions

  • M
    Prof. Shuangge Ma
    (Presenter/Speaker)
  • Multi-source treatment effect estimation via semiparametric sparse deep neural networks

  • Proposal Description

    This invited session will bring together four methodological talks at the intersection of deep learning, high-dimensional statistical inference, and complex biomedical data analysis. Rather than focusing on deep learning solely as a predictive tool, the session emphasizes how modern machine learning architectures can be embedded within principled statistical frameworks to answer inferential questions in biomedical research. The talks collectively address a central challenge facing biostatistics and data science: how to retain valid estimation, interpretation, and uncertainty quantification when the data structure, covariate dimension, and outcome processes exceed the assumptions of classical models.

    The session is organized around several recurring themes. First, multiple presentations focus on time-to-event and semi-competing risk outcomes, where non-terminal events such as disease progression may be censored or precluded by terminal events such as death. These settings are common in cancer research and clinical epidemiology, yet they remain difficult to analyze when treatment effects, disease trajectories, and covariate effects are nonlinear or heterogeneous. Talks by Stephen Salerno and Yi Li will present complementary deep learning-based approaches for semi-competing risk data, including pseudo-value methods for causal estimation and neural expectation-maximization methods for multistate survival prediction. Both talks are motivated by lung cancer applications and illustrate how deep learning can be used not merely to improve prediction, but to estimate scientifically meaningful quantities in the presence of censoring, competing event processes, and complex covariate relationships.

    Second, the session highlights the role of regularization, sparsity, and weak signals in high-dimensional biomedical modeling. S. Ejaz Ahmed’s talk addresses the practical reality that biomedical datasets often contain a mixture of strong predictors, weak predictors, and noise variables. Standard variable selection procedures may discard weak but collectively informative signals, limiting predictive performance and interpretability. The proposed post-shrinkage framework offers a strategy for improving prediction after model selection while retaining theoretical guarantees and practical applicability in high-dimensional regression settings.

    Third, the session considers how deep neural networks can be integrated into semiparametric models to balance flexibility with interpretability. Shuangge Ma’s talk on multi-source treatment effect estimation introduces sparse deep neural networks within semiparametric accelerated failure time models, allowing a common treatment effect to be estimated across multiple data sources while flexibly modeling source-specific nuisance effects. This framework is particularly relevant for large clinical and population-based studies, where data heterogeneity across locations or cohorts must be accommodated without sacrificing the interpretability of the primary treatment effect.

    Together, the four presentations will appeal to statisticians, biostatisticians, machine learning researchers, and applied biomedical scientists interested in rigorous methods for analyzing modern biomedical data. Attendees will gain exposure to new approaches for causal inference, survival analysis, high-dimensional shrinkage, semiparametric modeling, and deep learning-based representation of complex risk processes. The session will also encourage discussion about broader methodological questions: when deep learning improves inference, how statistical assumptions should be encoded in neural architectures, how uncertainty should be quantified, and how flexible models can remain interpretable enough to support biomedical decision-making.