International Journal of Research and Innovation in Applied Science (IJRIAS)
Evaluating Uncertainty Quantification in Clinical Machine Learning: Calibration, Robustness, and Decision Utility under Distribution Shift
Published August 4, 2026 • Vol. 11, Issue 7, pp. 1241–1252Open Access
DOI: 10.51584/IJRIAS.2026.11070085
Abstract
Machine learning models deployed in clinical decision support systems almost universally produce point predictions without any accompanying measure of uncertainty. In high-stakes healthcare settings, this is not merely a technical limitation: a miscalibrated prediction can directly influence patient management decisions with real consequences for safety and outcomes. Several uncertainty quantification (UQ) methods have been proposed to address this gap, including conformal prediction, Bayesian neural networks (BNNs), and Monte Carlo (MC) dropout; however, their comparative evaluation has predominantly been conducted under idealised conditions that do not reflect clinical deployment, where patient populations, treatment practices, and data recording procedures change over time. We present a rigorous empirical framework for comparing these three UQ approaches on two clinical prediction tasks, in-hospital mortality and 30-day readmission, using 74,829 ICU admissions from the MIMIC-IV database. Methods are assessed across three dimensions: calibration quality (ECE, ACE, Brier score), robustness under temporal distribution shift, and clinical decision utility via net benefit analysis. All experiments are replicated across five independent seeds, with comparisons made using Wilcoxon signed-rank tests with Holm-Bonferroni correction. Under standard evaluation conditions, all three methods achieve similar discriminative performance (AUROC 0.836-0.844 for mortality; 0.637-0.641 for readmission). Under temporal shift, BNN calibration degrades most sharply on the readmission task (ΔECE = 0.011 ± 0.002) compared with MC Dropout (ΔECE = 0.002 ± 0.003), while AUROC paradoxically improves for all methods, demonstrating that discriminative and calibration performance can decouple under distribution shift. Conformal prediction maintains near-nominal empirical coverage on the mortality task (0.886 ± 0.002) but shows notable violations on readmission, raising practical concerns about exchangeability assumptions in deployed systems. These findings support a more demanding evaluation standard for UQ in clinical machine learning, one that moves beyond static i.i.d. benchmarks toward temporally robust, decision-aware assessment.
Keywords: uncertainty quantification; conformal prediction; Bayesian neural networks; Monte Carlo dropout
| Journal | International Journal of Research and Innovation in Applied Science (IJRIAS) |
|---|---|
| ISSN | 2454-6194 |
| Volume / Issue | Volume 11, Issue 7 |
| Pages | 1241–1252 |
| Publication date | August 4, 2026 |
| DOI | 10.51584/IJRIAS.2026.11070085 |
| Publisher | RSIS International |
| License | Open Access |
How to cite this article
Isaac Tosin Adisa, Francis Mawutor Amuyao, & Ezekiel Olaoluwa Joaquim (2026). Evaluating Uncertainty Quantification in Clinical Machine Learning: Calibration, Robustness, and Decision Utility under Distribution Shift. International Journal of Research and Innovation in Applied Science (IJRIAS), 11(7), 1241-1252. https://doi.org/10.51584/IJRIAS.2026.11070085
BibTeX
@article{Isaac2026,
title = {Evaluating Uncertainty Quantification in Clinical Machine Learning: Calibration, Robustness, and Decision Utility under Distribution Shift},
author = {Isaac Tosin Adisa and Francis Mawutor Amuyao and Ezekiel Olaoluwa Joaquim},
journal = {International Journal of Research and Innovation in Applied Science (IJRIAS)},
volume = {11},
number = {7},
pages = {1241--1252},
year = {2026},
doi = {10.51584/IJRIAS.2026.11070085},
publisher = {RSIS International}
}