RELIABILITY-AWARE ARTIFICIAL INTELLIGENCE FOR MULTILABEL ELECTROCARDIOGRAM DIAGNOSIS

Authors

DOI:

https://doi.org/10.37943/WWSQ922%20

Keywords:

artificial intelligence, electrocardiogram, multilabel classification, calibration, selective prediction, robustness, attribution stability, decision support

Abstract

Artificial intelligence can classify electrocardiograms, but high discrimination alone does not show whether probabilities, decisions, and explanations remain reliable under technical signal failures. This study evaluates an integrated reliability-aware framework for multi-label electrocardiogram classification on PTB-XL with five diagnostic superclasses: normal, myocardial infarction, ST and T wave change, conduction disturbance, and hypertrophy. Patient-disjoint folds 1–8, 9, and 10 were used for training, validation, and held-out testing, respectively. XGBoost and three one-dimensional deep-learning architectures were compared, and ResNet1D was selected by validation macro area under the receiver operating characteristic curve. On the held-out test fold, ResNet1D achieved a macro area under the receiver operating characteristic curve of 0.9110, macro area under the precision-recall curve of 0.7746, and macro F1 score of 0.7051. Isotonic calibration, fitted on validation predictions only, produced a held-out macro Brier score of 0.08523 and adaptive calibration error of 0.00053. Selective prediction showed a coverage-risk trade-off: higher coverage reduced abstention but increased accepted-case risk and expected harm. Robustness tests covered Gaussian noise, baseline wander, amplitude scaling, lead dropout, and partial waveform failures; grouped limb- and precordial-lead dropout caused the largest discrimination losses. A complete-missing-lead input gate rejected zeroed or flat leads and reduced expected harm under complete lead loss, but did not detect the tested partial waveform failures. Attribution stability was high under Gaussian noise and lower under lead dropout. External analysis on Georgia and Chapman-Shaoxing data was limited to a binary abnormal-versus-normal sanity check and showed source-dependent calibration transfer. The contribution is a reproducible reliability evaluation framework, not a new architecture, full external validation, or a clinical deployment claim.

References

Wagner, P., Strodthoff, N., Bousseljot, R.-D., Kreiseler, D., Lunze, F. I., Samek, W., & Schaeffter, T. (2020). PTB-XL, a large publicly available electrocardiography dataset. Scientific Data, 7(1), 154. https://doi.org/10.1038/s41597-020-0495-6

Strodthoff, N., Wagner, P., Schaeffter, T., & Samek, W. (2021). Deep learning for ECG analysis: Benchmarks and insights from PTB-XL. IEEE Journal of Biomedical and Health Informatics, 25(5), 1519–1528. https://doi.org/10.1109/JBHI.2020.3022989

Siontis, K. C., Noseworthy, P. A., Attia, Z. I., & Friedman, P. A. (2021). Artificial intelligence-enhanced electrocardiography in cardiovascular disease management. Nature Reviews Cardiology, 18(7), 465–478. https://doi.org/10.1038/s41569-020-00503-2

Ose, B., Sattar, Z., Gupta, A., Toquica, C., Harvey, C., & Noheria, A. (2024). Artificial intelligence interpretation of the electrocardiogram: A state-of-the-art review. Current Cardiology Reports, 26(6), 561–580. https://doi.org/10.1007/s11886-024-02062-1

Van Calster, B., Collins, G. S., Vickers, A. J., Wynants, L., Kerr, K. F., Barreñada, L., Varoquaux, G., Singh, K., Moons, K. G. M., Hernandez-Boussard, T., Timmerman, D., McLernon, D. J., van Smeden, M., & Steyerberg, E. W. (2025). Evaluation of performance measures in predictive artificial intelligence models to support medical decisions: Overview and guidance. The Lancet Digital Health, 7(12), 100916. https://doi.org/10.1016/j.landig.2025.100916

Banerji, C. R. S., Chakraborti, T., Harbron, C., & MacArthur, B. D. (2023). Clinical AI tools must convey predictive uncertainty for each individual patient. Nature Medicine, 29(12), 2996–2998. https://doi.org/10.1038/s41591-023-02562-7

Seoni, S., Jahmunah, V., Salvi, M., Datta Barua, P., Molinari, F., & Acharya, U. R. (2023). Application of uncertainty quantification to artificial intelligence in healthcare: A review of last decade (2013–2023). Computers in Biology and Medicine, 165, 107441. https://doi.org/10.1016/j.compbiomed.2023.107441

Liu, X., Glocker, B., McCradden, M. M., Ghassemi, M., Denniston, A. K., & Oakden-Rayner, L. (2022). The medical algorithmic audit. The Lancet Digital Health, 4(5), e384–e397. https://doi.org/10.1016/S2589-7500(22)00003-6

Kwong, J. C. C., Erdman, L., Khondker, A., Skreta, M., Goldenberg, A., McCradden, M. D., Lorenzo, A. J., & Rickard, M. (2022). The silent trial: The bridge between bench-to-bedside clinical AI applications. Frontiers in Digital Health, 4, 929508. https://doi.org/10.3389/fdgth.2022.929508

Vazquez, J., & Facelli, J. C. (2022). Conformal prediction in clinical medical sciences. Journal of Healthcare Informatics Research, 6(3), 241–252. https://doi.org/10.1007/s41666-021-00113-8

Kwon, H., & Kim, D.-J. (2026). Conformal selective prediction with cost aware deferral for safe clinical triage under distribution shift. Scientific Reports, 16, 10016. https://doi.org/10.1038/s41598-026-40637-w

Romero, F. P., Piñol, D. C., & Vázquez-Seisdedos, C. R. (2021). DeepFilter: An ECG baseline wander removal filter using deep learning techniques. Biomedical Signal Processing and Control, 70, 102992. https://doi.org/10.1016/j.bspc.2021.102992

Petmezas, G., Stefanopoulos, L., Kilintzis, V., Tzavelis, A., Rogers, J. A., Katsaggelos, A. K., & Maglaveras, N. (2022). State-of-the-art deep learning methods on electrocardiogram data: Systematic review. JMIR Medical Informatics, 10(8), e38454. https://doi.org/10.2196/38454

Mongan, J., Moy, L., & Kahn, C. E., Jr. (2020). Checklist for Artificial Intelligence in Medical Imaging (CLAIM): A guide for authors and reviewers. Radiology: Artificial Intelligence, 2(2), e200029. https://doi.org/10.1148/ryai.2020200029

Collins, G. S., Moons, K. G. M., Dhiman, P., Riley, R. D., Beam, A. L., Van Calster, B., Ghassemi, M., Liu, X., Reitsma, J. B., van Smeden, M., Boulesteix, A.-L., Camaradou, J. C., Celi, L. A., Denaxas, S., Denniston, A. K., Glocker, B., Golub, R. M., Harvey, H., Heinze, G., … Logullo, P. (2024). TRIPOD+AI statement: Updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ, 385, e078378. https://doi.org/10.1136/bmj-2023-078378

Vasey, B., Nagendran, M., Campbell, B., Clifton, D. A., Collins, G. S., Denaxas, S., Denniston, A. K., Faes, L., Geerts, B., Ibrahim, M., Liu, X., Mateen, B. A., Mathur, P., McCradden, M. D., Morgan, L., Ordish, J., Rogers, C., Saria, S., Ting, D. S. W., … Perkins, Z. B. (2022). Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nature Medicine, 28(5), 924–933. https://doi.org/10.1038/s41591-022-01772-9

Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 785–794. https://doi.org/10.1145/2939672.2939785

Ismail Fawaz, H., Lucas, B., Forestier, G., Pelletier, C., Schmidt, D. F., Weber, J., Webb, G. I., Idoumghar, L., Muller, P.-A., & Petitjean, F. (2020). InceptionTime: Finding AlexNet for time series classification. Data Mining and Knowledge Discovery, 34(6), 1936–1962. https://doi.org/10.1007/s10618-020-00710-y

Lea, C., Flynn, M. D., Vidal, R., Reiter, A., & Hager, G. D. (2017). Temporal convolutional networks for action segmentation and detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 156–165. https://doi.org/10.1109/CVPR.2017.113

Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. Proceedings of the 34th International Conference on Machine Learning, 70, 1321–1330. https://proceedings.mlr.press/v70/guo17a.html

Sundararajan, M., Taly, A., & Yan, Q. (2017). Axiomatic attribution for deep networks. Proceedings of the 34th International Conference on Machine Learning, 70, 3319–3328. https://proceedings.mlr.press/v70/sundararajan17a.html

Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., & Kim, B. (2018). Sanity checks for saliency maps. Advances in Neural Information Processing Systems, 31, 9525–9536. https://proceedings.neurips.cc/paper/2018/hash/294a8ed24b1ad22ec2e7efea049b8737-Abstract.html

Perez Alday, E. A., Gu, A., Shah, A. J., Robichaux, C., Wong, A.-K. I., Liu, C., Liu, F., Rad, A. B., Elola, A., Seyedi, S., Li, Q., Sharma, A., Clifford, G. D., & Reyna, M. A. (2021). Classification of 12-lead ECGs: The PhysioNet/Computing in Cardiology Challenge 2020. Physiological Measurement, 41(12), 124003. https://doi.org/10.1088/1361-6579/abc960

Zheng, J., Zhang, J., Danioko, S., Yao, H., Guo, H., & Rakovski, C. (2020). A 12-lead electrocardiogram database for arrhythmia research covering more than 10,000 patients. Scientific Data, 7(1), 48. https://doi.org/10.1038/s41597-020-0386-x

Downloads

Published

2026-09-30

How to Cite

Konyspayev, D., Kuzenbayev, B., Abatov, N., & Alippayeva, D. (2026). RELIABILITY-AWARE ARTIFICIAL INTELLIGENCE FOR MULTILABEL ELECTROCARDIOGRAM DIAGNOSIS. Scientific Journal of Astana IT University, 27(3), 66–78. https://doi.org/10.37943/WWSQ922

Issue

Section

Information Technologies