Back to Research

Published — IEEE Access

Severity-Weighted Calibration Error for Reliability Assessment Under Outcome-Severity Imbalance

Aggregate calibration scores mask failures in high-acuity subgroups. Per-band isotonic recalibration consistently fixed this; global recalibration did not.

Background and Motivation

Clinical risk models are typically judged on a single, aggregate calibration score — a summary of how closely predicted probabilities track observed outcomes across an entire cohort. That summary can look excellent even when the model is badly miscalibrated exactly where it matters most: the small, high-acuity subgroup of patients closest to the decision boundary between survival and death. A model that is well-calibrated on average but overconfident or underconfident for the sickest patients can drive incorrect triage and resourcing decisions while its overall calibration metrics stay reassuringly flat.

Proposed Metric

This paper introduces Severity-Weighted Calibration Error (SWCE), a metric that re-weights calibration error by patient acuity rather than treating every patient as equally important to the average. By stratifying calibration into severity bands and weighting errors accordingly, SWCE surfaces miscalibration in high-acuity subgroups that aggregate metrics such as Expected Calibration Error (ECE) systematically hide.

The paper further evaluates two correction strategies: a single global isotonic recalibration fitted across the whole cohort, and a per-band isotonic recalibration fitted separately within each severity stratum. Across the evaluated models and cohorts, per-band recalibration consistently corrected the high-acuity miscalibration that global recalibration missed — global recalibration improved the aggregate score without fixing the subgroup failure it was meant to catch.

Significance in the Broader Research Portfolio

SWCE extends the same underlying argument as the author's work on Risk-Weighted Error Metrics (RWEM) and Clinical Risk-Weighted evaluation (CRWS): aggregate performance numbers for clinical AI systems routinely obscure failures concentrated in the patients who can least afford them. Where RWEM and CRWS reweight classification error by severity, SWCE applies the same lens to calibration — closing the gap between how confident a model claims to be and how confident it should be, specifically for the sickest patients.

Publication Details

Venue
IEEE Access
Quartile
Q1
Impact Factor
4.2
H-Index
338

Authors: Aman Chandra H, K.C. Narendra · 2026

Cite this Work

@ARTICLE{10717951,
  author    = "Chandra, H. Aman and Narendra, K. C.",
  journal   = "IEEE Access",
  title     = "Severity-Weighted Calibration Error for Reliability Assessment
               Under Outcome-Severity Imbalance",
  year      = "2026",
  publisher = "IEEE",
  doi       = "10.1109/ACCESS.2026.3717951"
}