Statistical Learning with Imperfect Data
Created with ChatGPT Images 2.0.Modern biomedical studies often collect complex longitudinal biomarkers that are subject to censoring and measurement error. We investigated statistical design considerations for genetic studies and derived a closed-form sample size formula for testing the overall effect of a single nucleotide polymorphism (SNP) within a joint longitudinal and time-to-event modeling framework. The practical utility of the proposed approach was demonstrated using data from the Diabetes Control and Complications Trial. To improve robustness against potential model misspecification arising from nonlinear longitudinal trajectories, we incorporated spline-based representations to capture subject-specific nonlinear evolution. Empirical studies demonstrate that misspecification of longitudinal trajectories has limited impact on SNP association testing but can substantially influence risk-factor and survival associations.
Incomplete data represent another fundamental challenge in modern data analysis. Focusing on missing responses, we developed a unified regularized framework that provides flexible inference without requiring specification of a particular missing-data mechanism, overcoming a key limitation of many existing approaches. Despite the rapid advancement of machine learning methods, theoretical understanding of these approaches remains limited, particularly in the presence of incomplete observations. To address this gap, we developed statistically justified boosting methods for datasets with missing responses using semiparametric estimation techniques. The proposed approaches employ functional gradient descent algorithms, and we establish theoretical guarantees for the resulting estimators. Extending these ideas to the survival setting, we developed boosting methods for interval-censored outcomes in both regression and classification problems, where response values are only known to lie within specified intervals. The theoretical properties, including estimation error bounds and optimality results, are rigorously established. We also validate our methods through extensive experiments on both synthetic and real-world datasets. These frameworks provide general principles for extending modern machine learning methods to complex data settings with imperfect observations.

























