What Medical AI Doesn't Know

Research paper

Read the research report.

I use this website for an interactive view of selected results. In the report, I document the study design, simulated missingness, model comparisons, probability metrics, figures, limitations, references, and reproducibility records.

PDF report

Target-Set and All-Feature Missingness Stress Tests for Two Tabular Clinical Benchmarks

Use the report to examine the statistical methods, numeric results, uncertainty, and limits of the benchmark.

What it covers

  • 01Scope-matched MCAR, MAR, and MNAR experiments plus a separate MCAR-all stress test
  • 02Logistic regression, random forest, and gradient boosting comparisons
  • 03ROC-AUC, accuracy, Brier score, ECE, and reliability diagrams
  • 04Descriptive intervals, limitations, reproducibility, and claim-linked artifacts

Focus

Calibration under missing data

I distinguish ranking performance from probability error and calibration.

Evidence

Saved benchmark outputs

I link numeric claims to saved metrics, predictions, figures, and summary tables.

Scope

Research, not diagnosis

I study model behavior on public benchmarks. I do not provide medical advice.