What Medical AI Doesn't Know

Explorer

Explore how scope-matched missingness changes model metrics.

Compare datasets, model families, mechanisms, and requested target-feature rates. Each selection reports achieved target and whole-matrix missingness, discrimination, probability error, and reliability records from my benchmark.

Benchmark condition and results

Choose a benchmark condition

Showing WDBC classification, Logistic Regression, Missing Completely At Random, requested target-feature rate 20%. General view.

Choose a scope-matched target-feature condition and compare measured changes in ranking and probability error. The 0 to 100 index is illustrative; use the individual metrics and reliability records for scientific interpretation.

Summary

Higher metric retention

Missing Completely At Random at a requested 20% rate applies to 5 prespecified target features, not the full record.

Mild information loss: the achieved rate is 19.8% within the target features and 3.3% across all 30 features. A limited share of the complete feature matrix is missing.

I mask each eligible cell independently at the requested probability. The clean classes are highly separable in this dataset. Under target-set missingness, probability error changed more consistently than ROC-AUC.

ROC-AUC changes little, so Brier score and ECE provide useful complementary information.

Probability error increases modestly despite the limited change in discrimination.

Logistic Regression has the highest heuristic index among the three models in this matched slice. I use this ordering descriptively, not as a clinical ranking.

The heuristic index is higher in this setting. I recommend checking the component metrics before interpreting the result.

Example inputs

Example variables in this dataset

These are dataset fields, not information entered by a user. Only the prespecified target features are eligible for masking in the explorer.

  • 01mean radiussignal
  • 02mean texturesignal
  • 03mean perimetersignal
  • 04worst texturesignal
  • 05worst perimetersignal

Scenario snapshot

Dataset
WDBC classification
Model
Logistic Regression
Mechanism
Missing Completely At Random
Whole-matrix missingness
3.3%
Comparison rank
#1 of 3
Clean baseline ROC-AUC
99.5%
Reliability bins
10

Missingness severity

Mild information loss

A limited share of the complete feature matrix is missing.

Ranking signal

Near the clean reference

ROC-AUC remains close to its clean-reference value.

Calibration discrepancy

Near the clean reference

ECE remains comparatively close to the clean reference.

Reliability view

Predicted probabilities and observed outcomes

The solid line shows the selected condition, and the diagonal marks perfect calibration. The darker dashed line shows the clean reference when one is available.

Reliability diagrampredicted vs observed
Predicted probabilities and observed outcomesMean predicted probability on the horizontal axis; observed fraction with target value 1 on the vertical axis. The diagonal marks agreement. Bin counts include repeated prediction records, not unique patients. The bin values are available in the table below.000.250.250.50.50.750.7511MEAN PREDICTED PROBABILITY
  • Dashed diagonal: perfect calibration
  • Solid line: selected missing-data condition
  • Dashed dark line: clean reference
Pooled predictions
5690
Mean abs gap
0.023
Max bin gap
0.228
View reliability bin values

Probabilities and observed target-1 fractions are shown on a 0 to 1 scale. Counts are repeated prediction records, not unique patients. Empty bins are marked n/a.

Predicted probabilities and observed outcomes: selected-condition and clean-reference bin values
Bin rangeSelected countSelected mean probabilitySelected observed fractionClean countClean mean probabilityClean observed fraction
0.0 to 0.117420.0090.0003570.0080.000
0.1 to 0.2940.1430.000170.1470.000
0.2 to 0.3520.2500.058110.2460.091
0.3 to 0.4790.3550.127140.3480.143
0.4 to 0.5690.4510.246140.4580.143
0.5 to 0.6870.5550.609150.5510.867
0.6 to 0.7840.6490.786140.6410.786
0.7 to 0.81110.7550.919230.7570.957
0.8 to 0.92400.8510.933470.8510.936
0.9 to 1.031320.9860.9886260.9860.989

Model comparison

Model results under the same benchmark condition

Each model uses the same dataset, target-feature mechanism, requested rate, outer splits, and saved masks. The heuristic ordering is descriptive and has no clinical meaning.

01

Logistic Regression

stableSelected model

I fit this regularized linear classifier after imputation and feature scaling.

Illustrative index
98/100
ROC-AUC
99.4%
Brier
0.024
ECE
0.042

02

Gradient Boosting

stable

I evaluate this histogram-based boosted-tree classifier with the same masks and outer splits.

Illustrative index
95/100
ROC-AUC
99.1%
Brier
0.037
ECE
0.042

03

Random Forest

stable

I evaluate this decision-tree ensemble with the same masks and outer splits.

Illustrative index
95/100
ROC-AUC
99.0%
Brier
0.036
ECE
0.062