Explainability Versus Predictive Accuracy in Machine-Learning Models for 30-Day Hospital Readmission Among Patients With Heart Failure: A Systematic Evidence Synthesis
DOI:
https://doi.org/10.66687/JMRISKeywords:
ML, Predictive Accuracy, Heart FailureAbstract
Background: Thirty-day hospital readmission after heart-failure hospitalization remains difficult to predict because readmission reflects heterogeneous clinical, social, behavioral, and healthcare-utilization factors. Machine-learning models can capture nonlinear relationships and high-dimensional electronic health-record data, but increasingly complex models may sacrifice interpretability for only modest gains in discrimination.
Objective: To synthesize high-quality evidence available through 2022 concerning the predictive performance and explainability of machine-learning models for 30-day hospital readmission among patients with heart failure and to determine whether increases in model complexity translate into clinically meaningful predictive benefit.
Methods: A structured evidence synthesis was conducted using peer-reviewed Q1 literature published no later than 01 December 2022. Studies developing or comparing models for 30-day heart-failure readmission, readmission-or-death, or closely related early post-discharge outcomes were prioritized. Logistic regression, tree-based models, ensemble methods, neural networks, and deep-learning approaches were compared with respect to discrimination, class imbalance, calibration, interpretability, validation, and potential clinical utility. Methodological literature concerning explainable artificial intelligence and clinical prediction was incorporated.
Results: Prediction of 30-day readmission remained only moderately discriminative. Mortazavi et al. reported improvement of machine-learning approaches over logistic regression, but absolute discrimination remained limited. Golas et al. reported c-statistics of 0.664 for logistic regression, 0.650 for gradient boosting, 0.695 for maxout networks, and 0.705 for a deep unified network. The deep model therefore gained only 0.041 in c-statistic over logistic regression while losing direct feature-level interpretability. In a 2022 independent validation study, XGBoost achieved AUROC 0.65 compared with 0.58 for a neural network and 0.57 for the modified LaCE score. Feature-importance approaches improved global interpretation, but patient-level post-hoc explanations introduce additional questions concerning fidelity and stability.
Conclusion: More complex machine-learning models can modestly improve prediction of 30-day heart-failure readmission, but the gain in accuracy is frequently smaller than expected and may not justify reduced transparency. Model selection should consider calibration, precision-recall performance, external validity, explanation fidelity, and clinical actionability in addition to AUROC. For clinical deployment, the preferred model is the simplest approach that preserves clinically meaningful predictive performance while providing reliable individual-level explanations.
References
Dharmarajan K, Hsieh AF, Lin Z, Bueno H, Ross JS, Horwitz LI, et al. Diagnoses and timing of 30-day readmissions after hospitalization for heart failure, acute myocardial infarction, or pneumonia. JAMA. 2013;309(4):355-363. doi:10.1001/jama.2012.216476.
Mortazavi BJ, Downing NS, Bucholz EM, Dharmarajan K, Manhapra A, Li SX, et al. Analysis of machine learning techniques for heart failure readmissions. Circ Cardiovasc Qual Outcomes. 2016;9(6):629-640. doi:10.1161/CIRCOUTCOMES.116.003039.
Frizzell JD, Liang L, Schulte PJ, Yancy CW, Heidenreich PA, Hernandez AF, et al. Prediction of 30-day all-cause readmissions in patients hospitalized for heart failure: comparison of machine learning and other statistical approaches. JAMA Cardiol. 2017;2(2):204-209. doi:10.1001/jamacardio.2016.3956.
Golas SB, Shibahara T, Agboola S, Otaki H, Sato J, Nakae T, et al. A machine learning model to predict the risk of 30-day readmissions in patients with heart failure: a retrospective analysis of electronic medical records data. BMC Med Inform Decis Mak. 2018;18:44. doi:10.1186/s12911-018-0620-z.
Awan SE, Bennamoun M, Sohel F, Sanfilippo FM, Dwivedi G. Machine learning-based prediction of heart failure readmission or death: implications of choosing the right model and the right metrics. ESC Heart Fail. 2019;6(2):428-435. doi:10.1002/ehf2.12419.
Shin S, Austin PC, Ross HJ, Abdel-Qadir H, Freitas C, Tomlinson G, et al. Machine learning vs. conventional statistical models for predicting heart failure readmission and mortality. ESC Heart Fail. 2021;8(1):106-115. doi:10.1002/ehf2.13073.
Sharma V, Kulkarni V, McAlister F, Eurich D, Keshwani S, Simpson SH, et al. Predicting 30-day readmissions in patients with heart failure using administrative data: a machine learning approach. J Card Fail. 2022;28(5):710-722. doi:10.1016/j.cardfail.2021.12.004.
Christodoulou E, Ma J, Collins GS, Steyerberg EW, Verbakel JY, Van Calster B. A systematic review shows no performance benefit of machine learning over logistic regression for clinical prediction models. J Clin Epidemiol. 2019;110:12-22. doi:10.1016/j.jclinepi.2019.02.004.
Ouwerkerk W, Voors AA, Zwinderman AH. Factors influencing the predictive power of models for predicting mortality and/or heart failure hospitalization in patients with heart failure. JACC Heart Fail. 2014;2(5):429-436. doi:10.1016/j.jchf.2014.04.006.
Alba AC, Agoritsas T, Walsh M, Hanna S, Iorio A, Devereaux PJ, et al. Discrimination and calibration of clinical prediction models: users’ guides to the medical literature. JAMA. 2017;318(14):1377-1384. doi:10.1001/jama.2017.12126.
Van Calster B, McLernon DJ, van Smeden M, Wynants L, Steyerberg EW. Calibration: the Achilles heel of predictive analytics. BMC Med. 2019;17:230. doi:10.1186/s12916-019-1466-7.
Rudin C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nat Mach Intell. 2019;1:206-215. doi:10.1038/s42256-019-0048-x.
Arrieta AB, Díaz-Rodríguez N, Del Ser J, Bennetot A, Tabik S, Barbado A, et al. Explainable artificial intelligence (XAI): concepts, taxonomies, opportunities and challenges toward responsible AI. Inf Fusion. 2020;58:82-115. doi:10.1016/j.inffus.2019.12.012.
Liu Y, Chen PHC, Krause J, Peng L. How to read articles that use machine learning: users’ guides to the medical literature. JAMA. 2019;322(18):1806-1816. doi:10.1001/jama.2019.16489.
Collins GS, Reitsma JB, Altman DG, Moons KGM. Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD). Ann Intern Med. 2015;162(1):55-63. doi:10.7326/M14-0697.
Wolff RF, Moons KGM, Riley RD, Whiting PF, Westwood M, Collins GS, et al. PROBAST: a tool to assess the risk of bias and applicability of prediction model studies. Ann Intern Med. 2019;170(1):51-58. doi:10.7326/M18-1376.