What Medical AI Doesn't Know

About

Why I built this benchmark

I became interested in uncertainty in medicine after experiencing seizures in high school. Decisions had to be made from incomplete signals and imperfect measurements, and that stayed with me.

That experience moved my interest beyond a general idea of combining health and technology. I wanted to understand the statistics behind uncertain predictions, how models behave when information is missing, and how those limits should be communicated.

How should a model communicate uncertainty when some of its usual evidence is absent?

Research question

Measure more than discrimination.

In this benchmark, I ask what happens to classification, ranking, probability error, and calibration when selected inputs are unavailable. These outcomes can move in different directions, so ROC-AUC alone is not enough to describe the result.

I also separate the rate inside a target feature set from the rate across the full matrix. Without that denominator, differences can be attributed to a mechanism when the real difference is how much information was removed.

Project boundary

A reproducible study, not a medical product.

I use this site to present controlled experiments on two public benchmark datasets. It does not accept medical records, produce patient-specific predictions, or support diagnosis or treatment decisions.

My aim is to make the methods and limitations inspectable. I include the report, saved predictions, masks, tables, and provenance records so readers can check the analysis rather than rely on the interface alone.