Symbolic Data Analysis for Challenging and Complex Big Data
Conference
Category: International Association for Statistical Computing (IASC)
Proposal Description
Symbolic data analysis (SDA) aggregates massively large individual-level datasets into a small number of distributional summaries, such as intervals or histograms. Unfortunately, statistical methodology is struggling to keep up with the development of analytical procedures to handle these distributional summary data. Traditional data are point observations in p-dimensional space; whereas symbolic data are hypercubes in the p-dimensional space.
Our session aims to present and discuss new methodologies in SDA as follows:
The dilemma with SDA occurs when trying to analyze the new second-level summary data formats. To date, often, analyses on these new data formats proceed by using classical surrogates; these surrogates can take various forms, but typically they are some functions of the mean (of the values within the aggregated set), or a range or variance measure, or extreme values (such as interval end-points). Our session aims to review some available methodologies which use all the data information contrasting the results with those from classical surrogates, such as regression, principal components, clustering, and time series. Hence, prospects for new directions in SDA will be highlighted and discussed.
If inference is carried out using the summary data in place of the original data, it results in computational gains at a loss of some information. In likelihood-based SDA, the likelihood function is characterized by an integral with a large exponent and in some circumstances the likelihood function is known to produce biased parameter estimates. A hybrid Gibbs sampler algorithm with Open-Faced Sandwich post-sampling adjustment is developed to enable robust posterior inference. Simulation studies exhibit that this approach yields unbiased estimates and maintains coverage probabilities close to the nominal levels. The study demonstrates a solid computational performance and improved asymptotic efficiency in comparison to previous approaches for more complex symbolic data.
Extracting information using symbolic principal component analysis (PCA) is well established in SDA. An extension of the PCA biplot methodology to the more general framework of SDA where each subject has multiple observations will be presented and discussed. This new PCA biplot integrates the information of between subject variation, within interval variation and within multiple observations variation of the interval-valued data.
Commonly, survey datasets are large and complex. Data collected via mobile phones, including ecological momentary assessment (EMA) and passive mobile sensing related to the music adolescents listened to during a one-week monitoring period are analyzed using SDA. EMA involved responses to short questionnaires delivered seven times per day over the observed week. Data were collected from high school students in two waves: the first wave in 2024, while the second wave in 2025. SDA is used to cluster EMA and passive sensing data to define adolescent profiles that are subsequently analyzed using survey variables. The motivation for applying SDA to psychological data lies in the assumption that preserving information about variability in EMA self-evaluations, together with analyzing emotions contained in listened music as distributions rather than as separate emotions, may provide additional insight into adolescents’ behavior, mental health, and everyday habits.