Machine Learning-Based Prediction of Culture-Confirmed Neonatal Sepsis in a Tertiary Neonatal Intensive Care Unit: Retrospective Cohort Study
Machine Learning-Based Prediction of Culture-Confirmed Neonatal Sepsis in a Tertiary Neonatal Intensive Care Unit: Retrospective Cohort Study

Machine Learning-Based Prediction of Culture-Confirmed Neonatal Sepsis in a Tertiary Neonatal Intensive Care Unit: Retrospective Cohort Study

JMIR Med Inform. 2026 Sep 14;14:e88732. doi: 10.2196/88732.

ABSTRACT

BACKGROUND: Neonatal sepsis remains a major cause of neonatal morbidity and mortality in low- and middle-income countries (LMICs). Early diagnosis is challenging because of nonspecific clinical manifestations and delays in laboratory confirmation. Machine learning (ML) approaches using structured electronic health record (EHR) data may improve early risk stratification in neonatal intensive care units (NICUs).

OBJECTIVE: This study aimed to evaluate ML models for predicting culture-confirmed neonatal sepsis among neonates admitted to a tertiary NICU in Jordan, with the objective of addressing diagnostic gaps in resource-limited settings. Specifically, we aimed to identify key predictors through feature importance analysis, evaluate model performance with class-imbalanced data, and propose strategies to improve interpretability and generalizability in LMICs.

METHODS: A retrospective cohort study was conducted using structured EHRs of 3274 neonates admitted to a tertiary NICU in Jordan between 2018 and 2024. Neonates who underwent blood culture testing were included. The dataset was divided into training (n=2619, 80%) and testing (n=655, 20%) subsets using stratified sampling. Three ML models-Extreme Gradient Boosting (XGBoost), decision trees, and neural networks-were trained using clinical, laboratory, and demographic variables. Class imbalance was addressed using the synthetic minority oversampling technique (SMOTE) applied to the training dataset. Model performance was evaluated using accuracy, sensitivity, specificity, and the area under the receiver operating characteristic curve (AUC).

RESULTS: Among 3274 neonates included in the study, the XGBoost model demonstrated the best predictive performance on the independent test set (n=655, 20%), achieving an accuracy of 94% (616/655 correct predictions, 95% CI 92% to 96%), sensitivity of 98% (95/97 sepsis cases correctly identified, 95% CI 96% to 99%), and an AUC of 0.98 (95% CI 0.97 to 0.99). Decision trees provided interpretable classification rules with moderate performance, whereas neural networks showed lower discriminative ability, with an AUC of 0.81 (95% CI 0.78 to 0.84). Important predictive features included C-reactive protein, platelet count, and gestational age.

CONCLUSIONS: XGBoost demonstrated strong predictive performance in this retrospective cohort, supporting its potential as a foundation for future prospective clinical decision support tools. External validation and prospective studies are required before clinical implementation.

PMID:42735002 | DOI:10.2196/88732