Research question
Measure more than discrimination.
In this benchmark, I ask what happens to classification, ranking, probability error, and calibration when selected inputs are unavailable. These outcomes can move in different directions, so ROC-AUC alone is not enough to describe the result.
I also separate the rate inside a target feature set from the rate across the full matrix. Without that denominator, differences can be attributed to a mechanism when the real difference is how much information was removed.