Degree

Doctor of Philosophy (PhD)

Department

Mathematics

Document Type

Dissertation

Abstract

Reliable feature selection is crucial for building interpretable and practical risk models, particularly when decisions depend on identifying the highest-risk groups rather than relying on average trends. This dissertation presents and tests a supervised filter that ranks predictors by their upper-tail concordance with the outcome, using a Gumbel-implied upper-tail concordance score $\lambda_U$. The score is calculated from pseudo-observations and Kendall’s $\tau$ (via the Gumbel $\tau\mapsto\lambda_U$ mapping), does not require model fitting during selection, and focuses directly on joint extreme events, such as when both a predictor and the positive class are high. The method is compared to Mutual Information, mRMR, ReliefF, and L1/Elastic-Net, combined with four common tabular classifiers: logistic regression, random forests, gradient boosting, and XGBoost. Validation is performed on two medical datasets: a large public health survey (CDC Diabetes Health Indicators (CDC)), comprising 253,680 samples and 21 features, and a classic clinical cohort (Pima Indians Diabetes Database (PIMA)), consisting of 768 samples and 8 features. All methods use a single stratified train, validation, and test split. Class imbalance is addressed through built-in weighting, non-linear ensembles are probability-calibrated, and operating thresholds are set based on validation (F1) scores, which are then applied to the test set. The primary metric is the ROC-AUC, reported alongside accuracy, precision, recall, and F1. Interpretability is assessed using permutation importance (ROC-AUC scorer), and paired comparisons use DeLong’s test (AUC) and McNemar’s test (error discordance). On the CDC dataset, the proposed Gumbel-$\lambda_U$ filter reduces the feature set by approximately half (from 21 to 10) while maintaining nearly the same level of discrimination as using all features (test AUCs of about 0.823 vs. about 0.827). It performs better than standard filters such as Mutual Information and mRMR in paired AUC tests, matches ReliefF statistically, and is the fastest among the selectors tested in wall-clock time on the CDC dataset, while on the PIMA dataset, times are negligible, and L1/Elastic-Net is slightly faster. On the PIMA dataset, where reduction is less important, the Gumbel ranking achieves competitive (numerically highest) ROC-AUC among strong alternatives, and paired tests show no significant differences, serving as a ranking-focused clinical check in a low-dimensional setting. Stability analyses using bootstrap resampling reveal high set overlap and strong rank agreement, with bootstrap confidence intervals for selected $\lambda_U$ scores being tight, indicating a clear dependence signal. Under stressors such as label flips, small feature changes, missing data with median imputation, and reduced training size, discrimination decreases gradually and stays in the same range, consistent with the main ROC-AUC results. In summary, learning from joint extremes with copula-implied upper-tail concordance scoring yields a simple, fast, and interpretable feature selector that is competitive with standard filters. Although tested here on medical datasets, the framework can be applied to machine learning risk prediction in any domain where tail behavior is crucial for decision-making.

Date

2-6-2026

DOI

https://proquest.com/docview/3347858031

First Committee Chair

Bruce Wade

First Committee Member

Amanda Mayeaux

Second Committee Member

Xiang-Sheng Wang

Third Committee Member

Yongli Sang

Share

COinS