IPS 400 - Recent advances of large-scale data integration and meta-analysis
Category: IPSParticipants
With the advances of technology and increases in computational speed, the need to analyze large-scale data has emerged. Multi-cohort, multi-source and multi-modal datasets often need to be combined for integrated clustering, increased statistical power, reduced biases of estimation of treatment or causal effects, among many other analytical purposes. Therefore, we propose to organize an invited session to bring leading researchers together to share their recent research advances and discuss ideas and important issues in the fields. In this session, speakers will be invited to present their latest work in data integrative analysis and meta-analysis, providing an interaction and brainstorming opportunity for researchers in these two fast evolving and growing fields.
Presenters are from United States, Hong Kong and Taiwan.
The issue of combining individual p-values to aggregate multiple small effects is a longstanding statistical topic. Many classical methods are designed for combining independent and frequent signals using the sum of transformed p-values with the transformation of light-tailed distributions, in which Fisher’s method and Stouffer’s method are the most well-known. In recent years, advances in big data promoted methods to aggregate correlated, sparse and weak signals; among them, Cauchy and harmonic mean combination tests were proposed to robustly combine p-values under unspecified dependency structure. Both of the proposed tests are the transformation of heavy-tailed distributions for improved power with the sparse signal. Motivated by this observation, we investigate the transformation of regularly varying distributions, which is a rich family of heavy-tailed distribution, to explore the conditions for a method to possess robustness to dependency and optimality of power for sparse signals. We show that only an equivalent class of Cauchy and harmonic mean tests has sufficient robustness to dependency in a practical sense. Moreover, a practical guideline to adjust significance level under dependency is provided based on our theorem and simulation. We also show an issue caused by large negative penalty in the Cauchy method and propose a simple, yet practical modification with fast computation. Finally, we present simulations and apply to a neuroticism GWAS application to verify the discovered theoretical insights.
Organiser: Dr Chung Chang
Chair: Dr Chung Chang
Speaker: Dr. Fangda Song
Speaker: Prof. George C. Tseng
Speaker: Dr Chung Chang
Speaker: Prof. Ming-Chieh Shih