Background and Motivation
Clinical risk models are typically judged on a single, aggregate calibration score — a summary of how closely predicted probabilities track observed outcomes across an entire cohort. That summary can look excellent even when the model is badly miscalibrated exactly where it matters most: the small, high-acuity subgroup of patients closest to the decision boundary between survival and death. A model that is well-calibrated on average but overconfident or underconfident for the sickest patients can drive incorrect triage and resourcing decisions while its overall calibration metrics stay reassuringly flat.
Proposed Metric
This paper introduces Severity-Weighted Calibration Error (SWCE), a metric that re-weights calibration error by patient acuity rather than treating every patient as equally important to the average. By stratifying calibration into severity bands and weighting errors accordingly, SWCE surfaces miscalibration in high-acuity subgroups that aggregate metrics such as Expected Calibration Error (ECE) systematically hide.
The paper further evaluates two correction strategies: a single global isotonic recalibration fitted across the whole cohort, and a per-band isotonic recalibration fitted separately within each severity stratum. Across the evaluated models and cohorts, per-band recalibration consistently corrected the high-acuity miscalibration that global recalibration missed — global recalibration improved the aggregate score without fixing the subgroup failure it was meant to catch.
Significance in the Broader Research Portfolio
SWCE extends the same underlying argument as the author's work on Risk-Weighted Error Metrics (RWEM) and Clinical Risk-Weighted evaluation (CRWS): aggregate performance numbers for clinical AI systems routinely obscure failures concentrated in the patients who can least afford them. Where RWEM and CRWS reweight classification error by severity, SWCE applies the same lens to calibration — closing the gap between how confident a model claims to be and how confident it should be, specifically for the sickest patients.
Publication Details
- Venue
- IEEE Access
- Quartile
- Q1
- Impact Factor
- 4.2
- H-Index
- 338
Authors: Aman Chandra H, K.C. Narendra · 2026
Cite this Work
@ARTICLE{10717951,
author = "Chandra, H. Aman and Narendra, K. C.",
journal = "IEEE Access",
title = "Severity-Weighted Calibration Error for Reliability Assessment
Under Outcome-Severity Imbalance",
year = "2026",
publisher = "IEEE",
doi = "10.1109/ACCESS.2026.3717951"
}