Advanced Statistical techniques for high throughput data integration analysis
Conference
Category: Special Interest Group on Data Science
Proposal Description
Session Overview
The rapid evolution of high-throughput technologies—such as single-cell multiomics, spatial transcriptomics, and high-resolution neuroimaging—has revolutionized biomedical research. However, these technologies generate massive, heterogeneous, and multi-view datasets that defy traditional statistical methods. To translate this wealth of data into actionable biological insights, the scientific community requires robust, scalable, and sophisticated statistical frameworks.
Significance and Innovation:
This session brings together pioneering methodological developments designed to integrate multi-source, high-throughput biomedical data. By moving beyond simple correlation and pooling techniques, the featured methodologies address the critical challenges of causality, spatial architectures, multi-scale clustering, and temporal dynamics.
Significance and Innovation
The innovation of this session lies in its multi-faceted approach to data integration, spanning four cutting-edge statistical domains:
1. High-Dimensional Mediation: Moving past standard low-dimensional mediation to connect complex multiomics layers with phenotypic imaging data, providing a clearer causal pathway for complex diseases like Alzheimer's.
2. Causal Spatial Networks: Shifting from descriptive cell-clustering to inferring directed, causal cell-cell communication networks by utilizing the precise physical coordinates provided by spatial transcriptomics.
3. Bayesian Multi-Scale Integration: Overcoming sample-to-sample variation and resolution differences in spatial transcriptomics through a unified Bayesian framework that performs simultaneous clustering and feature selection.
4. Multi-Dimensional Clustering: Advancing traditional clustering into bi-clustering and tri-clustering models capable of capturing synchronized, heterogeneous patterns across both longitudinal (temporal) and cross-sectional study designs.
Core Thematic Focus Areas
A. Multiview Multivariate Mediation for Multiomics and Imaging
To understand complex diseases, researchers must trace how genetic and molecular variations translate into structural biological changes. This focus area introduces an advanced framework for multiview multivariate mediation analysis. Engineered for high-dimensional data, this method models the intricate intermediary pathways connecting multiomics profiles to neuroimaging outcomes, with a specific, validated application to uncovering the mechanisms of Alzheimer’s disease progression.
B. Causal Cell-Cell Communication in Spatial Transcriptomics
Cells do not operate in isolation; their functions are heavily dictated by their local microenvironment. This segment focuses on novel statistical modeling that leverages spatial transcriptomics data to map cell-cell communication. By integrating spatial orientation with gene expression profiles, this method moves beyond association to infer causal signaling networks, offering a deeper look into tissue architecture and cellular dynamics.
C. Bayesian Multi-Scale Clustering and Feature Selection
Integrating spatial transcriptomics data across multiple samples or different technical scales often introduces severe batch effects and noise. We introduce BayesClint, an innovative Bayesian multi-scale clustering method that simultaneously performs factor analysis and spatial clustering on multiple samples, where the clustering is done jointly at the single-cell and tissue regional scale. It allows for seamless multi-sample integration while simultaneously performing automated feature selection, ensuring that only the most biologically relevant spatial transcripts drive the multi-scale tissue alignment.
D. Integrative Bi- and Tri-Clustering for Complex Study Designs
Modern clinical trials and cohort studies frequently mix longitudinal follow-ups with cross-sectional observations. Standard matrix factorization often falls short here. This focus area presents next-generation integrative bi-clustering and tri-clustering techniques. These algorithms are uniquely capable of isolating local, highly correlated feature subsets across multiple dimensions simultaneously, providing a robust toolkit for analyzing complex, multi-dimensional timeline data.