RELIABILITY-AWARE ARTIFICIAL INTELLIGENCE FOR MULTILABEL ELECTROCARDIOGRAM DIAGNOSIS
DOI:
https://doi.org/10.37943/WWSQ922%20Keywords:
artificial intelligence, electrocardiogram, multilabel classification, calibration, selective prediction, robustness, attribution stability, decision supportAbstract
Artificial intelligence can classify electrocardiograms, but high discrimination alone does not show whether probabilities, decisions, and explanations remain reliable under technical signal failures. This study evaluates an integrated reliability-aware framework for multi-label electrocardiogram classification on PTB-XL with five diagnostic superclasses: normal, myocardial infarction, ST and T wave change, conduction disturbance, and hypertrophy. Patient-disjoint folds 1–8, 9, and 10 were used for training, validation, and held-out testing, respectively. XGBoost and three one-dimensional deep-learning architectures were compared, and ResNet1D was selected by validation macro area under the receiver operating characteristic curve. On the held-out test fold, ResNet1D achieved a macro area under the receiver operating characteristic curve of 0.9110, macro area under the precision-recall curve of 0.7746, and macro F1 score of 0.7051. Isotonic calibration, fitted on validation predictions only, produced a held-out macro Brier score of 0.08523 and adaptive calibration error of 0.00053. Selective prediction showed a coverage-risk trade-off: higher coverage reduced abstention but increased accepted-case risk and expected harm. Robustness tests covered Gaussian noise, baseline wander, amplitude scaling, lead dropout, and partial waveform failures; grouped limb- and precordial-lead dropout caused the largest discrimination losses. A complete-missing-lead input gate rejected zeroed or flat leads and reduced expected harm under complete lead loss, but did not detect the tested partial waveform failures. Attribution stability was high under Gaussian noise and lower under lead dropout. External analysis on Georgia and Chapman-Shaoxing data was limited to a binary abnormal-versus-normal sanity check and showed source-dependent calibration transfer. The contribution is a reproducible reliability evaluation framework, not a new architecture, full external validation, or a clinical deployment claim.
References
Wagner, P., Strodthoff, N., Bousseljot, R.-D., Kreiseler, D., Lunze, F. I., Samek, W., & Schaeffter, T. (2020). PTB-XL, a large publicly available electrocardiography dataset. Scientific Data, 7(1), 154. https://doi.org/10.1038/s41597-020-0495-6
Strodthoff, N., Wagner, P., Schaeffter, T., & Samek, W. (2021). Deep learning for ECG analysis: Benchmarks and insights from PTB-XL. IEEE Journal of Biomedical and Health Informatics, 25(5), 1519–1528. https://doi.org/10.1109/JBHI.2020.3022989
Siontis, K. C., Noseworthy, P. A., Attia, Z. I., & Friedman, P. A. (2021). Artificial intelligence-enhanced electrocardiography in cardiovascular disease management. Nature Reviews Cardiology, 18(7), 465–478. https://doi.org/10.1038/s41569-020-00503-2
Ose, B., Sattar, Z., Gupta, A., Toquica, C., Harvey, C., & Noheria, A. (2024). Artificial intelligence interpretation of the electrocardiogram: A state-of-the-art review. Current Cardiology Reports, 26(6), 561–580. https://doi.org/10.1007/s11886-024-02062-1
Van Calster, B., Collins, G. S., Vickers, A. J., Wynants, L., Kerr, K. F., Barreñada, L., Varoquaux, G., Singh, K., Moons, K. G. M., Hernandez-Boussard, T., Timmerman, D., McLernon, D. J., van Smeden, M., & Steyerberg, E. W. (2025). Evaluation of performance measures in predictive artificial intelligence models to support medical decisions: Overview and guidance. The Lancet Digital Health, 7(12), 100916. https://doi.org/10.1016/j.landig.2025.100916
Banerji, C. R. S., Chakraborti, T., Harbron, C., & MacArthur, B. D. (2023). Clinical AI tools must convey predictive uncertainty for each individual patient. Nature Medicine, 29(12), 2996–2998. https://doi.org/10.1038/s41591-023-02562-7
Seoni, S., Jahmunah, V., Salvi, M., Datta Barua, P., Molinari, F., & Acharya, U. R. (2023). Application of uncertainty quantification to artificial intelligence in healthcare: A review of last decade (2013–2023). Computers in Biology and Medicine, 165, 107441. https://doi.org/10.1016/j.compbiomed.2023.107441
Liu, X., Glocker, B., McCradden, M. M., Ghassemi, M., Denniston, A. K., & Oakden-Rayner, L. (2022). The medical algorithmic audit. The Lancet Digital Health, 4(5), e384–e397. https://doi.org/10.1016/S2589-7500(22)00003-6
Kwong, J. C. C., Erdman, L., Khondker, A., Skreta, M., Goldenberg, A., McCradden, M. D., Lorenzo, A. J., & Rickard, M. (2022). The silent trial: The bridge between bench-to-bedside clinical AI applications. Frontiers in Digital Health, 4, 929508. https://doi.org/10.3389/fdgth.2022.929508
Vazquez, J., & Facelli, J. C. (2022). Conformal prediction in clinical medical sciences. Journal of Healthcare Informatics Research, 6(3), 241–252. https://doi.org/10.1007/s41666-021-00113-8
Kwon, H., & Kim, D.-J. (2026). Conformal selective prediction with cost aware deferral for safe clinical triage under distribution shift. Scientific Reports, 16, 10016. https://doi.org/10.1038/s41598-026-40637-w
Romero, F. P., Piñol, D. C., & Vázquez-Seisdedos, C. R. (2021). DeepFilter: An ECG baseline wander removal filter using deep learning techniques. Biomedical Signal Processing and Control, 70, 102992. https://doi.org/10.1016/j.bspc.2021.102992
Petmezas, G., Stefanopoulos, L., Kilintzis, V., Tzavelis, A., Rogers, J. A., Katsaggelos, A. K., & Maglaveras, N. (2022). State-of-the-art deep learning methods on electrocardiogram data: Systematic review. JMIR Medical Informatics, 10(8), e38454. https://doi.org/10.2196/38454
Mongan, J., Moy, L., & Kahn, C. E., Jr. (2020). Checklist for Artificial Intelligence in Medical Imaging (CLAIM): A guide for authors and reviewers. Radiology: Artificial Intelligence, 2(2), e200029. https://doi.org/10.1148/ryai.2020200029
Collins, G. S., Moons, K. G. M., Dhiman, P., Riley, R. D., Beam, A. L., Van Calster, B., Ghassemi, M., Liu, X., Reitsma, J. B., van Smeden, M., Boulesteix, A.-L., Camaradou, J. C., Celi, L. A., Denaxas, S., Denniston, A. K., Glocker, B., Golub, R. M., Harvey, H., Heinze, G., … Logullo, P. (2024). TRIPOD+AI statement: Updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ, 385, e078378. https://doi.org/10.1136/bmj-2023-078378
Vasey, B., Nagendran, M., Campbell, B., Clifton, D. A., Collins, G. S., Denaxas, S., Denniston, A. K., Faes, L., Geerts, B., Ibrahim, M., Liu, X., Mateen, B. A., Mathur, P., McCradden, M. D., Morgan, L., Ordish, J., Rogers, C., Saria, S., Ting, D. S. W., … Perkins, Z. B. (2022). Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nature Medicine, 28(5), 924–933. https://doi.org/10.1038/s41591-022-01772-9
Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 785–794. https://doi.org/10.1145/2939672.2939785
Ismail Fawaz, H., Lucas, B., Forestier, G., Pelletier, C., Schmidt, D. F., Weber, J., Webb, G. I., Idoumghar, L., Muller, P.-A., & Petitjean, F. (2020). InceptionTime: Finding AlexNet for time series classification. Data Mining and Knowledge Discovery, 34(6), 1936–1962. https://doi.org/10.1007/s10618-020-00710-y
Lea, C., Flynn, M. D., Vidal, R., Reiter, A., & Hager, G. D. (2017). Temporal convolutional networks for action segmentation and detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 156–165. https://doi.org/10.1109/CVPR.2017.113
Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. Proceedings of the 34th International Conference on Machine Learning, 70, 1321–1330. https://proceedings.mlr.press/v70/guo17a.html
Sundararajan, M., Taly, A., & Yan, Q. (2017). Axiomatic attribution for deep networks. Proceedings of the 34th International Conference on Machine Learning, 70, 3319–3328. https://proceedings.mlr.press/v70/sundararajan17a.html
Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., & Kim, B. (2018). Sanity checks for saliency maps. Advances in Neural Information Processing Systems, 31, 9525–9536. https://proceedings.neurips.cc/paper/2018/hash/294a8ed24b1ad22ec2e7efea049b8737-Abstract.html
Perez Alday, E. A., Gu, A., Shah, A. J., Robichaux, C., Wong, A.-K. I., Liu, C., Liu, F., Rad, A. B., Elola, A., Seyedi, S., Li, Q., Sharma, A., Clifford, G. D., & Reyna, M. A. (2021). Classification of 12-lead ECGs: The PhysioNet/Computing in Cardiology Challenge 2020. Physiological Measurement, 41(12), 124003. https://doi.org/10.1088/1361-6579/abc960
Zheng, J., Zhang, J., Danioko, S., Yao, H., Guo, H., & Rakovski, C. (2020). A 12-lead electrocardiogram database for arrhythmia research covering more than 10,000 patients. Scientific Data, 7(1), 48. https://doi.org/10.1038/s41597-020-0386-x
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Articles are open access under the Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Authors who publish a manuscript in this journal agree to the following terms:
- The authors reserve the right to authorship of their work and transfer to the journal the right of first publication under the terms of the Creative Commons Attribution License, which allows others to freely distribute the published work with a mandatory link to the the original work and the first publication of the work in this journal.
- Authors have the right to conclude independent additional agreements that relate to the non-exclusive distribution of the work in the form in which it was published by this journal (for example, to post the work in the electronic repository of the institution or publish as part of a monograph), providing the link to the first publication of the work in this journal.
- Other terms stated in the Copyright Agreement.