Summary
Higher metric retention
Missing Completely At Random at a requested 20% rate applies to 5 prespecified target features, not the full record.
Mild information loss: the achieved rate is 19.8% within the target features and 3.3% across all 30 features. A limited share of the complete feature matrix is missing.
I mask each eligible cell independently at the requested probability. The clean classes are highly separable in this dataset. Under target-set missingness, probability error changed more consistently than ROC-AUC.
ROC-AUC changes little, so Brier score and ECE provide useful complementary information.
Probability error increases modestly despite the limited change in discrimination.
Logistic Regression has the highest heuristic index among the three models in this matched slice. I use this ordering descriptively, not as a clinical ranking.
The heuristic index is higher in this setting. I recommend checking the component metrics before interpreting the result.