Methodology
How I ran the benchmark
I use the same benchmark in the explorer and the paper. In my primary comparison, I hold the eligible feature set constant across MCAR, MAR, and MNAR and report an all-feature MCAR stress test separately in the paper.
Datasets
WDBC classification
The clean classes are highly separable in this dataset. Under target-set missingness, probability error changed more consistently than ROC-AUC.
Statlog Heart classification
This dataset is smaller and less separable, with wider variation across outer splits.
Models
Logistic Regression
I fit this regularized linear classifier after imputation and feature scaling.
Random Forest
I evaluate this decision-tree ensemble with the same masks and outer splits.
Gradient Boosting
I evaluate this histogram-based boosted-tree classifier with the same masks and outer splits.
Rate definition
A target-feature missingness rate is not a whole-matrix missingness rate.
The selector shows a requested conditional rate within five WDBC target features or three Statlog target features. Every scenario separately reports the achieved rate across the complete feature matrix.
Missingness mechanisms
Missing Completely At Random
I mask each eligible cell independently at the requested probability.
Missing At Random
I make masking probability depend on a prespecified observed anchor variable.
Missing Not At Random
I make masking probability depend on the value being hidden in this simulation.
Evaluation
- ROC-AUC and accuracy summarize discrimination and classification.
- Brier score and expected calibration error measure probability error and calibration discrepancy, which can change even when ranking remains similar.
- I use two repeats of stratified five-fold outer cross-validation for 10 outer splits and three-fold inner cross-validation to select hyperparameters by ROC-AUC.
- I average five mask seeds within each outer split and bootstrap the 10 split estimates for descriptive intervals rather than counting seeds as independent samples.
- I build reliability views from saved predictions. Counts are repeated prediction records across cross-validation repeats and mask seeds, not unique patients.
Limitations
Controlled, not clinical deployment
I use public benchmark datasets and controlled missingness constructions, not hospital production systems or patient-specific predictions.
Educational, not diagnostic
I use the robustness index only as an illustration, not a scientific endpoint or validated medical score. The explorer should never be used for personal medical decisions.