RSIS Repository Open-access research from RSIS International journals

International Journal of Research and Innovation in Applied Science (IJRIAS)

Evaluating Uncertainty Quantification in Clinical Machine Learning: Calibration, Robustness, and Decision Utility under Distribution Shift

byIsaac Tosin Adisa; Francis Mawutor Amuyao; Ezekiel Olaoluwa Joaquim

Published August 4, 2026  •  Vol. 11, Issue 7, pp. 1241–1252Open Access
DOI: 10.51584/IJRIAS.2026.11070085

Abstract

Machine learning models deployed in clinical decision support systems almost universally produce point predictions without any accompanying measure of uncertainty. In high-stakes healthcare settings, this is not merely a technical limitation: a miscalibrated prediction can directly influence patient management decisions with real consequences for safety and outcomes. Several uncertainty quantification (UQ) methods have been proposed to address this gap, including conformal prediction, Bayesian neural networks (BNNs), and Monte Carlo (MC) dropout; however, their comparative evaluation has predominantly been conducted under idealised conditions that do not reflect clinical deployment, where patient populations, treatment practices, and data recording procedures change over time. We present a rigorous empirical framework for comparing these three UQ approaches on two clinical prediction tasks, in-hospital mortality and 30-day readmission, using 74,829 ICU admissions from the MIMIC-IV database. Methods are assessed across three dimensions: calibration quality (ECE, ACE, Brier score), robustness under temporal distribution shift, and clinical decision utility via net benefit analysis. All experiments are replicated across five independent seeds, with comparisons made using Wilcoxon signed-rank tests with Holm-Bonferroni correction. Under standard evaluation conditions, all three methods achieve similar discriminative performance (AUROC 0.836-0.844 for mortality; 0.637-0.641 for readmission). Under temporal shift, BNN calibration degrades most sharply on the readmission task (ΔECE = 0.011 ± 0.002) compared with MC Dropout (ΔECE = 0.002 ± 0.003), while AUROC paradoxically improves for all methods, demonstrating that discriminative and calibration performance can decouple under distribution shift. Conformal prediction maintains near-nominal empirical coverage on the mortality task (0.886 ± 0.002) but shows notable violations on readmission, raising practical concerns about exchangeability assumptions in deployed systems. These findings support a more demanding evaluation standard for UQ in clinical machine learning, one that moves beyond static i.i.d. benchmarks toward temporally robust, decision-aware assessment.

Keywords: uncertainty quantification; conformal prediction; Bayesian neural networks; Monte Carlo dropout

JournalInternational Journal of Research and Innovation in Applied Science (IJRIAS)
ISSN2454-6194
Volume / IssueVolume 11, Issue 7
Pages1241–1252
Publication dateAugust 4, 2026
DOI10.51584/IJRIAS.2026.11070085
PublisherRSIS International
LicenseOpen Access

How to cite this article

Isaac Tosin Adisa, Francis Mawutor Amuyao, & Ezekiel Olaoluwa Joaquim (2026). Evaluating Uncertainty Quantification in Clinical Machine Learning: Calibration, Robustness, and Decision Utility under Distribution Shift. International Journal of Research and Innovation in Applied Science (IJRIAS), 11(7), 1241-1252. https://doi.org/10.51584/IJRIAS.2026.11070085

BibTeX

@article{Isaac2026,
  title   = {Evaluating Uncertainty Quantification in Clinical Machine Learning: Calibration, Robustness, and Decision Utility under Distribution Shift},
  author  = {Isaac Tosin Adisa and Francis Mawutor Amuyao and Ezekiel Olaoluwa Joaquim},
  journal = {International Journal of Research and Innovation in Applied Science (IJRIAS)},
  volume  = {11},
  number  = {7},
  pages   = {1241--1252},
  year    = {2026},
  doi     = {10.51584/IJRIAS.2026.11070085},
  publisher = {RSIS International}
}