Comparative Evaluation of LSTM, GRU, and Bi-LSTM for 30-Day Heart Failure Readmission and Incident Atrial Fibrillation Prediction from Longitudinal Electronic Health Records
DOI:
https://doi.org/10.71222/0fv3gm77Keywords:
Recurrent Neural Networks, Heart Failure Readmission, Atrial Fibrillation, Electronic Health RecordsAbstract
Recurrent neural networks have become the default choice for sequence modeling in longitudinal electronic health records, yet practitioners have limited empirical guidance on which RNN variant to select for cardiovascular early-warning tasks. This study reports a controlled comparison of long short-term memory (LSTM), gated recurrent unit (GRU), and bidirectional LSTM (Bi-LSTM) on two clinically meaningful endpoints: 30-day all-cause readmission among heart failure patients evaluated at discharge, and 6-month incident atrial fibrillation among at-risk adults evaluated from a fixed retrospective horizon anchored to the index hospitalization. The primary cohort is drawn from MIMIC-IV v3.1 (49,283 heart failure index admissions and 38,712 at-risk patients after exclusions), with eICU-CRD v2.0 used as a multi-center external validation cohort. Each variant is trained under a shared architectural search space, identical preprocessing pipelines, missing-value handling strategies, and class-rebalancing schemes. On the heart failure task, Bi-LSTM achieves the highest internal AUROC of 0.764 (95% CI: 0.752--0.776), exceeding LSTM by 1.6 percentage points and GRU by 1.3 percentage points, with consistent gains in DeLong testing. On the atrial fibrillation task, GRU attains an AUROC of 0.798 with roughly 25% fewer parameters and 24% shorter training time than LSTM. Performance gaps among the three variants narrow under external validation, with all models losing approximately 3-4 AUROC points (range: 3.4--4.2). Findings support task-conditioned variant selection rather than a universal recommendation, and we caution that the observed absolute AUROC differences, although statistically significant, are modest and require prospective decision-curve evaluation before any clinical deployment claim is justified.References
1. S. Hochreiter and J. Schmidhuber, "Long short-term memory," Neural Computation, vol. 9, no. 8, pp. 1735–1780, 1997.
2. K. Cho, B. van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, "Learning phrase representations using RNN encoder–decoder for statistical machine translation," in *Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing*, pp. 1724–1734, 2014.
3. M. Schuster and K. K. Paliwal, "Bidirectional recurrent neural networks," IEEE Transactions on Signal Processing, vol. 45, no. 11, pp. 2673–2681, 1997.
4. J. Chung, C. Gulcehre, K. Cho, and Y. Bengio, "Empirical evaluation of gated recurrent neural networks on sequence modeling," arXiv preprint arXiv:1412.3555, 2014.
5. E. Choi, A. Schuetz, W. F. Stewart, and J. Sun, "Using recurrent neural network models for early detection of heart failure onset," Journal of the American Medical Informatics Association, vol. 24, no. 2, pp. 361–370, 2017.
6. Z. C. Lipton, D. C. Kale, C. Elkan, and R. Wetzel, "Learning to diagnose with LSTM recurrent neural networks," in Proceedings of the 4th International Conference on Learning Representations, 2016.
7. E. Choi, M. T. Bahadori, J. A. Kulas, A. Schuetz, W. F. Stewart, and J. Sun, "RETAIN: An interpretable predictive model for healthcare using reverse time attention mechanism," in Advances in Neural Information Processing Systems 29, pp. 3504–3512, 2016.
8. F. Ma, R. Chitta, J. Zhou, Q. You, T. Sun, and J. Gao, "Dipole: Diagnosis prediction in healthcare via attention-based bidirectional recurrent neural networks," in *Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining*, pp. 1903–1911, 2017.
9. A. Rajkomar et al., "Scalable and accurate deep learning with electronic health records," npj Digital Medicine, vol. 1, p. 18, 2018.
10. P. Tiwari, K. L. Colborn, D. E. Smith, F. Xing, D. Ghosh, and M. A. Rosenberg, "Assessment of a machine learning model applied to harmonized electronic health record data for the prediction of incident atrial fibrillation," JAMA Network Open, vol. 3, no. 1, p. e1919396, 2020.
11. A. Guo, S. Smith, Y. M. Khan, J. R. Langabeer II, and R. E. Foraker, "Application of a time-series deep learning model to predict cardiac dysrhythmias in electronic health records," PLOS ONE, vol. 16, no. 9, p. e0239007, 2021.
12. Z. I. Attia et al., "An artificial intelligence-enabled ECG algorithm for the identification of patients with atrial fibrillation during sinus rhythm: A retrospective analysis of outcome prediction," The Lancet, vol. 394, no. 10201, pp. 861–867, 2019.
13. A. E. W. Johnson et al., "MIMIC-IV, a freely accessible electronic health record dataset," Scientific Data, vol. 10, p. 1, 2023.
14. T. J. Pollard et al., "The eICU Collaborative Research Database, a freely available multi-center database for critical care research," Scientific Data, vol. 5, p. 180178, 2018.
15. H. Harutyunyan, H. Khachatrian, D. C. Kale, G. Ver Steeg, and A. Galstyan, "Multitask learning and benchmarking with clinical time series data," Scientific Data, vol. 6, p. 96, 2019.
16. E. Choi, M. T. Bahadori, A. Schuetz, W. F. Stewart, and J. Sun, "Doctor AI: Predicting clinical events via recurrent neural networks," in Proceedings of the 1st Machine Learning for Healthcare Conference (PMLR Vol. 56), pp. 301–318, 2016.
17. Z. Che, S. Purushotham, K. Cho, D. Sontag, and Y. Liu, "Recurrent neural networks for multivariate time series with missing values," Scientific Reports, vol. 8, p. 6085, 2018.
18. E. Choi et al., "Multi-layer representation learning for medical concepts," in *Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining*, pp. 1495–1504, 2016.
19. I. M. Baytas et al., "Patient subtyping via time-aware LSTM networks," in *Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining*, pp. 65–74, 2017.
20. T. Pham, T. Tran, D. Phung, and S. Venkatesh, "DeepCare: A deep dynamic memory model for predictive medicine," in Pacific-Asia Conference on Knowledge Discovery and Data Mining (LNCS Vol. 9652), pp. 30–41, 2016.
21. S. Purushotham, C. Meng, Z. Che, and Y. Liu, "Benchmarking deep learning models on large healthcare datasets," Journal of Biomedical Informatics, vol. 83, pp. 112–134, 2018.
22. B. Shickel, P. J. Tighe, A. Bihorac, and P. Rashidi, "Deep EHR: A survey of recent advances in deep learning techniques for electronic health record (EHR) analysis," IEEE Journal of Biomedical and Health Informatics, vol. 22, no. 5, pp. 1589–1604, 2018.

