Interdisciplinary Statistical Learning Applications

Created with ChatGPT Images 2.0.

This work presents eight studies that apply machine learning and statistical methods to address challenges in diverse real-world contexts.

The first study investigated factors influencing the popularity of TED talks. Using network analysis, principal component analysis, and LASSO regression, a predictive model is developed. Results show that the number of available languages is the most influential factor, contributing at least three times more than other variables, while publishing talks on weekends (especially Saturdays) is associated with higher audience reach.

The second study evaluated the effectiveness of Diclectin versus placebo for treating nausea and vomiting during pregnancy in a double-blind randomized controlled trial with substantial missing longitudinal data. Results show that statistical inferences about the treatment effect vary markedly depending on the missing-data method and modeling approach. Moreover, the estimated effect does not reach the pre-specified minimal clinically important difference of 3 points, suggesting limited clinical relevance. These findings highlight the sensitivity of conclusions to missing-data assumptions and underscore the importance of robust inference in clinical trials.

The third study examined the impact of geographical and meteorological conditions on pipeline failures. Neural networks and random forests are used to identify key predictors. Meteorological factors dominate, with snow on ground, total snow, and total rainfall emerging as the most important variables. Snow on ground consistently ranks as the top predictor, while geographical characteristics contribute relatively little.

The fourth study developed a data-driven framework to evaluate physician performance in intensive care units (ICUs). Using tree-based ensemble models (XGBoost, random forests, and tree boosting mixed models) together with explainable AI tools such as TreeSHAP, the study quantifies contributions of physician practice patterns and ICU departments to patient outcomes. Propensity weighting is used to adjust for patient heterogeneity, enabling more fair comparisons across physicians, while super learner–based approaches improve robustness under model misspecification.

The fifth study examines COVID-19 incubation times under multiple statistical models, showing that the commonly adopted 14-day quarantine period may not sufficiently reduce the probability of releasing infectious individuals.

The sixth study estimated the causal effects of myocardial infarction and ischemic stroke on self-rated health using propensity weighting and matching methods. Results indicate that both events significantly reduce the probability of reporting very good or excellent health.

The seventh study introduced a cost-aware conformal cascade triage model for the interpretable diagnosis of Parkinson’s disease, augmented with large language model (LLM)-guided reasoning. The framework employs a staged cascade architecture that routes patients through progressively more expensive diagnostic stages. Early stages rely on low-cost features (e.g., Level 1: self-reported data; Level 2: summaries of Level 1 data), and more complex and costly data, requiring professional assistance (Level 3) or additional professional resources (Level 4), are incorporated adaptively when necessary. An LLM component can further generate natural-language explanations, contextualizes predictions with clinical knowledge, and supports differential diagnostic reasoning. Overall, the approach demonstrates superior performance in accuracy, calibration, cost reduction, and interpretability compared to traditional end-to-end models. An early version of this work was presented at the Statistical Society of Canada Annual Meeting 2026, where it won the case study competition among 23 teams. We are currently preparing the full paper for submission.

The eighth study developed conformal prediction methods for online chess integrity, proposing calibrated cheating and sandbagging detection, as well as robust benchmark calibration. We are currently preparing the papers for submission.

Overall, these eight studies demonstrate how machine learning and statistical modeling can generate actionable insights across domains ranging from social media analytics and infrastructure reliability to clinical decision support and causal inference in health outcomes. Collectively, the work emphasizes robustness, reliability, interpretability, uncertainty quantification, and practical efficiency in real-world applications.

Yuan Bian
Yuan Bian
Postdoctoral Research Scientist