Early Sepsis Prediction Dataset
Early identification of sepsis is critical, as delayed diagnosis significantly increases morbidity and mortality. Sepsis is defined as life-threatening organ dysfunction caused by a dysregulated host response to infection and remains a major global health concern. According to the Global Burden of Disease Study, sepsis affects an estimated 49 million people annually and contributes to approximately 11 million deaths, representing nearly 20 percent of global mortality.
This project develops and validates high-performance AI and machine learning models for early sepsis prediction using only clinical information available within the first hour of hospital arrival. Structured electronic health record data, continuous waveform vital signs, and a combination of both data types are evaluated. The goal is to support real-time clinical screening in emergency and low-resource settings where complete laboratory panels or continuous monitoring may not yet be available.
Dataset and Cohort
The study uses the AIM-AHEAD60 subset of the CHoRUS dataset, formatted according to the Observational Medical Outcomes Partnership (OMOP) Common Data Model. Adult patients aged 18 years and older were included. Patients with a final clinical diagnosis of sepsis formed the positive class, while patients without a sepsis diagnosis served as controls. Pediatric patients and records without documented diagnoses were excluded.
A total of 11,312 unique patients met the inclusion criteria. Among them, 2,245 individuals (19.85 percent) were diagnosed with sepsis at least once. Features included demographics, initial vital signs, laboratory results, and continuous vital-sign waveforms recorded during the first hour after hospital arrival. Extreme outliers were capped using predefined clinical thresholds (for example, heart rate above 250, respiratory rate above 40, systolic blood pressure above 300, and temperature outside 25–50 degrees Celsius after conversion to Celsius).
Public dataset references and related resources:
• CHoRUS / AIM-AHEAD60 (OMOP CDM) — primary source used in the study; data availability upon request to corresponding authors pending public release.
• PhysioNet — public repository of clinical waveform and EHR datasets commonly used for sepsis research.
• MIMIC-III / MIMIC-IV — widely used critical-care databases supporting related sepsis prediction work.
• IHI Sepsis Resources — clinical guidelines and early recognition frameworks.
• CDC Sepsis — public health guidance on sepsis awareness and prevention.
Figure 1. AI/ML Model Development Pipeline
The end-to-end pipeline used for early sepsis prediction follows a structured sequence from cohort definition through model evaluation. First-hour structured EHR features and waveform-derived statistics are extracted, cleaned, and fed into gradient boosting algorithms optimized for high recall.
Adults ≥18, sepsis vs control
EHR + Waveform vitals
Cap outliers, retain missing
Labs, RR, SBP, Temp trends
XGBoost · LightGBM · HistGB
AUROC, Recall, F1, Accuracy
Figure 1. End-to-end development pipeline from cohort selection to performance evaluation using only first-hour clinical information.
Modeling Approach and Algorithms
Three gradient boosting algorithms were developed with a primary focus on maximizing recall, which is clinically important for early sepsis detection. The algorithms evaluated were XGBoost (extreme gradient boosting), LightGBM (Light Gradient Boosting Machine), and HistGB (histogram-based gradient boosting). Models were trained separately on structured EHR data alone, waveform-flattened data alone, and the combination of both data types.
Structured EHR inputs included demographics, initial vital signs, and laboratory results such as lactate and leukocyte count. Waveform inputs captured continuous vital signs over the first hour, allowing the models to learn trends in respiratory rate, systolic blood pressure, and temperature. Combined models integrated both static laboratory values and dynamic vital-sign trajectories.
Performance was assessed using accuracy, precision, recall, F1 score, and area under the receiver operating characteristic curve (AUROC). Across configurations, XGBoost consistently achieved the strongest discrimination, reaching an AUROC of 0.922 on the combined data setting with recall above 80 percent, demonstrating robust behavior even in the presence of substantial missing data.
Performance Summary
Representative performance metrics for XGBoost, LightGBM, and HistGB under different data configurations are summarized below. Values reflect the strong discriminative ability of gradient boosting methods when restricted to first-hour clinical information only.
| Configuration | Model | Accuracy | Recall | F1 | AUROC |
|---|---|---|---|---|---|
| EHR Structured | XGBoost | 0.87+ | >0.80 | ~0.47 | 0.915–0.93 |
| Waveform Only | XGBoost | 0.85+ | >0.80 | ~0.36 | ~0.87–0.90 |
| EHR + Waveform | XGBoost | 0.87–0.95 | 0.72–0.91 | 0.39–0.55 | 0.922–0.957 |
| EHR + Waveform | LightGBM | ~0.91 | ~0.56–0.81 | ~0.37 | ~0.90 |
| EHR + Waveform | HistGB | ~0.91 | ~0.56 | ~0.37 | ~0.90 |
The combination of structured EHR and waveform data produced the highest AUROC values. XGBoost remained the top-performing algorithm across all three data settings, supporting its suitability for real-time clinical screening applications.
Figure 2. Feature Importance of XGBoost Prediction of Early Sepsis
Feature importance analysis reveals distinct predictive patterns depending on the data source. With structured EHR data alone, laboratory variables such as lactate and leukocyte count ranked highest. Waveform-only models emphasized respiratory rate, systolic blood pressure, and temperature trends measured toward the end of the first monitoring hour. In the combined setting, mean temperature and mean systolic blood pressure emerged among the leading predictors.
(A) Structured EHR Data
(B) Waveform (Vitals) Data
(C) Combined EHR + Waveform
Figure 2. Relative feature importance for XGBoost models trained on (A) structured EHR data, (B) waveform-flattened vital signs, and (C) combined EHR and waveform data. Longer bars indicate higher contribution to early sepsis prediction.
Clinical Relevance of First-Hour Information
A distinctive aspect of this approach is its exclusive reliance on health information available within the first hour of a patient’s arrival. This constraint mirrors real-world emergency and low-resource settings where complete laboratory panels, imaging, or specialist evaluations may not be immediately accessible. Early triage using vital signs plays a central role in identifying disease severity.
By definition, sepsis involves a confirmed or suspected source of infection accompanied by systemic inflammatory response syndrome (SIRS). SIRS criteria encompass abnormalities in temperature, heart rate, and respiratory rate, underscoring the importance of vital signs in initial risk stratification. When patients meet SIRS criteria, the risk of sepsis is substantially elevated. Lactate measurement is a commonly employed diagnostic tool; a level of 2 mmol/L or higher is often considered indicative of early tissue hypoperfusion in the setting of infection and is associated with worse outcomes. Sensitivities for elevated lactate in sepsis typically range from 66 to 83 percent, with specificities of 80 to 85 percent depending on the chosen cutoff.
Timely recognition remains essential because mortality risk increases with each hour of delay in initiating appropriate antimicrobial therapy. Prior work has reported approximately a 7.6 percent increase in mortality for every hour antimicrobial administration is delayed after the onset of hypotension in septic shock. Consequently, prediction models that can flag sepsis risk within one to two hours of presentation have clear potential to improve patient outcomes through earlier intervention.
Key Findings and Practical Implications
High-performing AI/ML models for early sepsis prediction can be constructed from publicly available datasets using only first-hour clinical information. XGBoost models demonstrated the strongest overall performance, with AUROC values around 0.90 when using structured or waveform data independently and reaching 0.922 in the combined configuration. Laboratory markers such as lactate and leukocyte count, together with dynamic trends in respiratory rate, systolic blood pressure, and temperature, consistently ranked among the most informative predictors.
These results support the feasibility of deploying gradient boosting based screening tools in emergency departments and resource-limited environments where continuous monitoring or extensive laboratory results may not yet be available. The models show strong potential for integration into real-time clinical workflows for early sepsis risk stratification.
Tools and Frameworks
Frequently Asked Questions
Additional clinical context emphasizes that gradient boosting models remain effective even when laboratory panels are incomplete. Missingness is retained rather than imputed aggressively, allowing the tree-based learners to treat absence of a value as informative. This design choice is particularly relevant for emergency triage, where certain assays may not be ordered until sepsis is already suspected.
Waveform aggregation over the first hour includes summary statistics such as mean, minimum, maximum, and late-window values of respiratory rate, systolic blood pressure, and temperature. These engineered features capture both the absolute level and the short-term trajectory of vital signs, which aligns with the physiologic progression of systemic inflammatory response.
Model calibration and operating-point selection can be adjusted according to local prevalence and resource constraints. Higher recall settings favor sensitivity for screening, while higher precision settings reduce false alarms in crowded emergency departments. The reported AUROC of 0.922 for the best XGBoost configuration indicates strong ranking ability across decision thresholds.