Deep Learning's Impact on Sepsis Care
Deep Learning's Impact on Sepsis Care
com/npjdigitalmed
ARTICLE OPEN
Sepsis remains a major cause of mortality and morbidity worldwide. Algorithms that assist with the early recognition of sepsis may
improve outcomes, but relatively few studies have examined their impact on real-world patient outcomes. Our objective was to
assess the impact of a deep-learning model (COMPOSER) for the early prediction of sepsis on patient outcomes. We completed a
before-and-after quasi-experimental study at two distinct Emergency Departments (EDs) within the UC San Diego Health System.
We included 6217 adult septic patients from 1/1/2021 through 4/30/2023. The exposure tested was a nurse-facing Best Practice
Advisory (BPA) triggered by COMPOSER. In-hospital mortality, sepsis bundle compliance, 72-h change in sequential organ failure
assessment (SOFA) score following sepsis onset, ICU-free days, and the number of ICU encounters were evaluated in the pre-
intervention period (705 days) and the post-intervention period (145 days). The causal impact analysis was performed using a
Bayesian structural time-series approach with confounder adjustments to assess the significance of the exposure at the 95%
1234567890():,;
confidence level. The deployment of COMPOSER was significantly associated with a 1.9% absolute reduction (17% relative
decrease) in in-hospital sepsis mortality (95% CI, 0.3%–3.5%), a 5.0% absolute increase (10% relative increase) in sepsis bundle
compliance (95% CI, 2.4%–8.0%), and a 4% (95% CI, 1.1%–7.1%) reduction in 72-h SOFA change after sepsis onset in causal
inference analysis. This study suggests that the deployment of COMPOSER for early prediction of sepsis was associated with a
significant reduction in mortality and a significant increase in sepsis bundle compliance.
npj Digital Medicine (2024)7:14; [Link]
1
Department of Medicine, University of California San Diego, San Diego, CA, USA. 2Department of Quality, University of California San Diego, San Diego, CA, USA. 3Department of
Emergency Medicine, University of California San Diego, San Diego, CA, USA. 4These authors contributed equally: Aaron Boussina, Supreeth P. Shashikumar. 5These authors jointly
supervised this work: Shamim Nemati, Gabriel Wardi. ✉email: gwardi@[Link]
Characteristic
Number of patients, N (%) 6217 (100%) 5065 (81.5%) 1152 (18.5%) -
Age, mean (SD) 63 (17.1) 63 (17.0) 64 (17.3) 0.08
Sex, N (%)
Male 3592 (57.8%) 2966 (58.6%) 626 (54.3%) -
Female 2625 (42.2%) 2099 (41.4%) 526 (45.7%) -
Race
Asian 530 (8.5%) 404 (8%) 126 (10.9%) -
Black or African American 639 (10.3%) 519 (10.2%) 120 (10.4%) -
White 2983 (48%) 2440 (48.2%) 543 (47.1%) -
Otherb 2065 (33.2%) 1702 (33.6%) 363 (31.5%) -
Ethnic group -
Hispanic/Latino 1756 (28.2%) 1449 (28.6%) 307 (26.6%) -
Not Hispanic/Latino 4461 (71.8%) 3616 (71.4%) 845 (73.4%) -
Organ dysfunction
Elixhauser Comorbidity Index, Median (IQR) 5 (0–13) 5 (0–13) 5 (0–14) 0.64
SOFA Score at Time of Sepsis, Median (IQR) 2 (1–3) 2 (1–3) 2 (1–3) 0.99
1234567890():,;
Lab values
Lactate at the time of sepsis 2.4 (1.6–4.3) 2.4 (1.6–4.3) 2.4 (1.6–4.3) 0.76
Interventions
Mechanical Ventilation, N (%)c 1035 (16.6%) 849 (16.8%) 186 (16.1%) 0.64
Administration of Vasoactive Medications, N (%)c 424 (6.8%) 345 (6.8%) 79 (6.9%) 1.0
a
P-values for continuous variables are based on Kruskal–Wallis rank sum tests. P-values for categorical variables are based on Pearson’s chi-squared tests.
b
Other race corresponds to Native Hawaiian or Other Pacific Islander, American Indian or Alaska Native, Other Race or Mixed Race, or Unknown.
c
Within 72-h of ED arrival.
Fig. 1 Acknowledgements to Each COMPOSER Best Practice Advisory alert from December, 2022 until April, 2023.
month. Alerts by acknowledgement reason are visualized in Fig. 1. Interventions and patient outcomes
The most common acknowledgement reason was “Will Notify MD The results from causal impact analysis on our primary and
Immediately” which comprised over half of all acknowledgement secondary outcomes are summarized in Table 2. The observed in-
reasons. Only about 5.9% of BPAs were exited without acknowl- hospital mortality rate and the corresponding predictions from the
edgement and responses to the BPA remained consistent across Bayesian structural time-series model are shown in Fig. 2a. The
the 5-month intervention period. residual quantile-quantile and autocorrelation plots are described
npj Digital Medicine (2024) 14 Published in partnership with Seoul National University Bundang Hospital
A. Boussina et al.
3
Table 2. Observed outcomes in the pre-intervention period, the expected counterfactual values from causal impact analysis, and the actual post-
intervention values.
Fig. 2 Causal impact analysis of COMPOSER Best Practice Advisory on patient outcomes. Plots of the causal impact analysis using a
Bayesian structural time-series model. The top subpanel (“original”) shows the actual outcome (black) and the average model predictions
(dashed blue) and 95% confidence limits (shaded blue) during the pre-intervention and post-intervention periods, indicated by the solid gray
vertical line. The middle subpanel (“pointwise”) shows the difference between the model predictions and the observed outcome. The bottom
subpanel (“cumulative”) shows the sum of the pointwise differences during the post-intervention period. Preparation for the implementation
of COMPOSER began in May 2022 approximately 6 months prior to the go-live date of the model. a The cumulative post-intervention in-
hospital sepsis mortality rate is below the 95% confidence limit. b The cumulative post-intervention 72-h change in SOFA score is below the
95% confidence limit.
Published in partnership with Seoul National University Bundang Hospital npj Digital Medicine (2024) 14
A. Boussina et al.
4
Fig. 3 Causal impact analysis of COMPOSER Best Practice Advisory (BPA) on sepsis bundle compliance rate. Implementation of the
COMPOSER BPA was significantly associated with an increase in sepsis bundle compliance.
in Supplementary Figs. 4 and 5. The average sepsis mortality rate immediately, there was a significant reduction in time to antibiotic
during the post-intervention period was 9.49%. If the COMPOSER administration (p = 0.002; two-sided t-test with adjustments
algorithm had not been deployed, the expected counterfactual for ED volume, sex, baseline SOFA, Elixhauser comorbidity score,
mortality rate would have been 11.39% with a 95% confidence and age).
interval of [9.79%, 13.00%], corresponding to a 1.9% absolute
decrease in sepsis-related in-hospital mortality. This value
corresponds to a 17% relative decrease in in-hospital mortality DISCUSSION
among patients with sepsis and 22 additional patients who In this before-and-after quasi-experimental study, we demon-
survived during the 5-month intervention period. The probability strated that the implementation of a real-time deep-learning
of this occurring by chance is determined from the Bayesian one- model to predict sepsis in two EDs was associated with a 5.0%
sided tail-area probability, p = 0.014. Additional data regarding the absolute increase in sepsis bundle compliance and a 1.9%
difference in mortality at our two hospitals are provided in absolute decrease in-hospital sepsis-related mortality. This finding
Supplementary Figs. 6–9. We found at one site (the “safety net” represents, to our knowledge, the first instance of prospective use
hospital) we had a significant decrease in mortality in the post- of a deep-learning model demonstrating an association with
intervention period, but we did not observe a significant change improved patient-centered outcomes in sepsis. Our findings also
at the other clinical site (the quaternary care facility). suggest that the utilization of such models in clinical care was also
The average compliance rate during the post-intervention associated with improvements in intermediate outcomes, such as
period was 53.42% while, in the absence of the COMPOSER less organ injury at 72 h from the time of sepsis and improve-
intervention, the expected compliance rate would have been ments in elements of sepsis bundles which may explain the
48.38% (95% CI, 45.46%–51.01%; Fig. 3), or a 5.0% (95% CI, mortality benefit described. Importantly, we show in scenarios
2.4%–8.0%) increase in compliance. This corresponds to a 10% where nursing staff reported notification of the provider with
(95% CI, 5%–16%) relative increase in sepsis bundle compliance concern for sepsis (approximately 55% of cases) that antibiotics
following the implementation of COMPOSER. As shown in were administered sooner, providing a plausible mechanism for
Supplementary Figs. 10 and 11, compliance with our sepsis the lower-than-expected in-hospital mortality we report.
bundle increased at both EDs during the intervention period. Despite major interest in strategies to relieve the morbidity and
Compliance with specific bundle elements is shown in Table 2 and mortality of sepsis, novel therapeutics have failed to translate into
Supplementary Figs. 12–17. We observe significant improvements meaningful patient-centered outcomes. The potential to improve
in antibiotic compliance, repeat lactate compliance, and admin- care through the use of artificial intelligence is attractive,
istration of fluids compliance. particularly with advances in machine learning in the past
We also observed a reduction in the 72-h change in SOFA score decade22,23. Unfortunately, the majority of algorithms designed
following sepsis onset (Fig. 2b). The average change in SOFA score to predict sepsis never make it to the bedside24. Older models
during the post-intervention period was 3.56. In the absence of designed to detect sepsis were largely based on clinical criteria
the COMPOSER intervention, however, the expected counter- (i.e., SIRS criteria, hypotension, or a combination of these). These
factual average change in SOFA score would have been about models were associated with occasional improvement in quality
3.71 with a 95% confidence interval of [3.58, 3.83]. This metrics (i.e., increased rates of lactate orders or time to antibiotics),
corresponds to a 4% decrease in the average 72-h change in but did not improve patient-centered outcomes and had poor
SOFA score following sepsis onset. The probability of this PPV25–27.
occurring by chance is p = 0.013. Additional data on the change More recently, several studies have implemented sophisticated
in SOFA score at each emergency department are provided in models at various hospitals showing benefits to patients.
Supplementary Figs. 18 and 19. We further observe a downward Shimaburuko et al. conducted a small randomized trial of 142
trend in our secondary endpoint of ICU admissions (Supplemen- patients in the ICU using a machine-learning algorithm to predict
tary Fig. 20) and an upward trend in ICU-free days (Supplementary severe sepsis and found a decrease in in-hospital mortality and
Fig. 21) although neither reaches statistical significance. The length of stay in the intervention group, although this study was
temporal trends of all covariates used in the Bayesian structural limited to patients either in the hospital wards or intensive care
time-series models are provided in Supplementary Fig. 22. units18. Adams et al. recently provided a prospective analysis of
Associations between time-to-antibiotics in septic patients and the TREWS model at five hospital systems in which they
acknowledgement reasons are shown in Table 3. We observe that demonstrated a significant decrease in mortality, organ failure,
in cases where nurses indicated that they would notify physicians and length of stay in hospitalized patients when the sepsis alert
npj Digital Medicine (2024) 14 Published in partnership with Seoul National University Bundang Hospital
A. Boussina et al.
5
to the training samples, it will flag the case as ‘indeterminate’. The
Table 3. Associations Between BPA Response and Antibiotics Timing.
resulting reduction in false alarms, previously reported to be 75%,
Variable Coefficient [95% CI] P greatly reduces the burden of resources or time spent on false
Valuea diagnoses.
There are various potential reasons that may explain the
(Intercept) 22.42 [−45.27–90.11] 0.513 reduction in mortality described above. First, we noted a high
Age 0.01 [−0.15–0.18] 0.875 percentage (~55%) of alerts were transmitted by nursing staff to
Male −0.53 [−5.99–4.93] 0.848 physicians. In this scenario, we found that these patients were
more likely to receive timely antibiotics, thus providing a potential
Elixhauser Comorbidity Index 0.21 [−0.10–0.52] 0.180
mechanism to decrease mortality and mitigate organ dysfunction
Baseline SOFA 1.12 [−2.89–5.13] 0.581 at 72 h. The use of artificial intelligence to facilitate a shared
Monthly ED Volume (in thousands) 0.62 [−13.33–14.58] 0.930 mental model of risk between nursing staff and providers has
BPA Acknowledgement: −13.24 [−26.60–0.127] 0.052 demonstrated good acceptance and improved use of these
No Infection Suspected models in other clinical areas36. In our system, for instance, we
BPA Acknowledgement: −19.16 [−33.73–−4.60] 0.010 chose to have the nurses receive the alert and determine if
Sepsis Treatment/Workup in escalation to the provider was appropriate. While the ideal target
Progress population for such an intervention is unclear, we felt that our
BPA Acknowledgement: −19.95 [−32.39–−7.51] 0.002 nurses would be the ideal candidate for this alert because of the
Will Notify MD Immediately high frequency of nurses opening patients’ charts. In the author’s
a
collective experience, physicians in the ED may have up to 15–20
P values are based on two-sided t-tests with adjustments for ED volume,
patients at a time and may not receive a BPA that requires a chart
sex, baseline SOFA, Elixhauser comorbidity score, and age.
Coefficients and P values are tabulated for a linear regression model with
to be open to receive the notification. Given the high rate of
adjustments. If the nurse selected “No Infection Suspected” on an 80-year- provider notification, we suspect that this approach was beneficial
old male with an Elixhauser of 4 and a baseline SOFA of 1, the expected to patient care while additionally minimizing unnecessary alerts.
“time from ED triage to antibiotics administration” would be 14.5 hours. Finally, although speculative, it is possible that the implementa-
However, if the nurse had selected “Will Notify MD Immediately” the tion of the alert improved situational awareness of sepsis care
expected “time from ED triage to antibiotics administration” would have within our ED staff. This finding has been reported in other sepsis
been 7.7 hours. clinical decision tools as well37.
Despite our study’s strengths, we acknowledge several limita-
tions. First, our study was not randomized and thus our findings
was confirmed by a provider16. While this study was not do not allow definitive causal inferences or mechanistic insights.
randomized, the data are compelling that proper attention to We performed a causal impact analysis with common confounders
implementation may improve patient-centered outcomes in which revealed that the implementation was significantly
sepsis. associated with positive outcomes. Regardless, we view the
However, a commonly used predictive model, the Epic Sepsis findings as important and believe that they provide a strong
Score (ESS), has not demonstrated consistent improvement in rationale for further research. Second, our study was conducted at
patient-centered outcomes. Although a small randomized quality two EDs in a large academic center that has a major interest in
improvement initiative from a single center found an improve- sepsis and clinical informatics. Although we had a large sample
ment of the composite clinical outcome measure of days alive and size and a diverse population of patients (racial, ethnic, socio-
out of hospital at 28 days was greater in the ESS care group, these economic status, etc.), we acknowledge the need for external
results have not been generalized thus far. Importantly, research- validation in other healthcare settings (e.g., community hospitals,
ers at the University of Michigan highlighted a substantial drop in different demographics, hospitals without robust IT infrastructure,
test characteristics (sensitivity, specificity, PPV) of the ESS at their etc.). Third, one could argue that an abrupt intervention has
institution from what was reported by Epic, as well as an important immediate benefits raising awareness and helping to
unacceptably high rate of false positives20. prioritize the care of a specific group of patients. Conversely, the
To the best of our knowledge, the only deep-learning model sustainability of the intervention could be questioned, emphasiz-
previously tested in an ED setting is the Sepsis Watch by ing the need for longer-term follow-up. Although human
investigators at Duke; however, no patient-centered outcomes interventions are subject to fatigue and complacency, we
have been reported thus far28. As such, the present study is the anticipate our automated algorithms will improve over time with
first reporting of improvement in patient-centered outcomes increasing experience and larger data sets, which will likely result
attributable to the deployment of a deep-learning-based sepsis in improvements in end-user satisfaction. However, we certainly
prediction model. recognize the importance of continuous education as a compo-
The use of deep learning for early prediction of sepsis is nent of care optimization. Finally, we did not evaluate the impact
significant since such models are capable of modeling temporal, of this alert on patients who ultimately did not have sepsis, such
nonlinear, and complex correlations among risk factors, thus as the potential adverse effects of inappropriate use of antibiotics
enabling them to solve more difficult problems. Moreover, deep- and healthcare costs associated with this. We also acknowledge
learning models are capable of handling large quantities of that we did not have any comparison data from the same time
multimodal data from radiology imaging, clinical notes, and period as all of our EDs used this model. However, we did not have
wearable sensors, among other29–32. Additionally, this class of any other quality improvement initiatives during the same time
models provides a flexible framework for transfer learning and period. Despite these limitations, we view our new findings as
continual learning to enable the adoption of such models to local actionable and important.
healthcare settings33–35. In the before-and-after quasi-experimental design study con-
Importantly, the COMPOSER deep-learning model was designed ducted at two EDs, we demonstrate that the implementation of a
to minimize false alarms via the conformal prediction framework. real-time deep-learning model to predict sepsis was associated
This approach imposes a boundary around the algorithm, which with a significant increase in bundle compliance, a significant
enables the model to identify whether it has enough prior reduction in in-hospital mortality, less organ dysfunction at 72 h,
knowledge of similar cases to determine reliably whether a patient and improved timeliness to antibiotics when nurses notified the
is at risk for sepsis. If the algorithm finds the data non-conformant physician of the BPA. To our knowledge, this is the first time that
Published in partnership with Seoul National University Bundang Hospital npj Digital Medicine (2024) 14
A. Boussina et al.
6
the improvement of patient outcomes due to the use of a deep- (2) if an antibiotics order occurred first, then a blood culture draw
learning model for sepsis prediction has been reported. Future had to occur within the next 24 h. Evidence of organ dysfunction
multicenter randomized trials are indicated to validate these was defined as an increase in the Sequential Organ Failure
findings across a diverse hospital and patient population. Assessment (SOFA) score by two or more points. In particular,
evidence of organ dysfunction occurring 48 h before to 24 h after
the time of suspected infection was considered, as suggested in
METHODS Seymour et al.1. Finally, the time of onset of sepsis was taken as
Study design and cohort the time of clinical suspicion of infection. The inclusion of 4 days
We conducted a prospective before-and-after quasi-experimental of non-prophylactic antibiotics to improve the specificity of sepsis
study to evaluate the impact of a sepsis Best Practice Advisory is similar to what Rhee et al. proposed as a surveillance approach
(BPA; Fig. 4) powered by the COMPOSER deep-learning model on for identifying sepsis from electronic health records which have
patient outcomes and process measures. The University of outperformed reliance on administrative coding of sepsis39. We
California San Diego Institutional review board (IRB) approval included all adult patients (age ≥18 years old) who met the criteria
was obtained with the waiver of informed consent (#805726) and for the above-described Sepsis-3 definition within the first 12 h of
additional approval was obtained from the Aligning and their ED stay. We excluded patients who were transitioned to
Coordinating QUality Improvement, Research, and Evaluation comfort measures prior to their time of sepsis and patients who
(ACQUIRE) Committee (project #609). Our study was completed developed sepsis after 12 h of hospital admission. All data used to
in accordance with STROBE guidelines38. A completed checklist is derive the onset time of sepsis and patient outcomes were
provided in Supplementary Note 1. These EDs have a total volume extracted via SQL queries on Epic Clarity.
of approximately 100,000 patients annually with one serving at a
quaternary academic center and the other at an urban “safety net” Sepsis algorithm and platform
hospital. The COMPOSER algorithm for the early prediction of sepsis is
Patients were identified as septic according to the latest described in Shashikumar et al.15. It is a feed-forward neural
international consensus definitions for sepsis (“Sepsis-3”)1,3. The network model that incorporates routinely collected laboratory
onset time of sepsis was established by following previously and vital signs as well as patient demographics (age and sex),
published methodology, using evidence of organ dysfunction and comorbidities, and concomitant medications to output a risk score
suspicion of clinical infection1,12,15. Clinical suspicion of infection for the onset of sepsis within the next 4 h. Importantly, the model
was defined by a blood culture draw and at least 4 days of non- utilizes the conformal prediction method to reject out-of-
prophylactic intravenous antibiotic therapy satisfying either of the distribution samples that may arise due to data entry error or
following conditions: (1) if a blood culture draw was ordered first, unfamiliar cases. The model achieves an area under the receiver
then an antibiotics order had to occur within the following 72 h, or operating characteristic curve (AUROC) of 0.938–0.945 within ED
npj Digital Medicine (2024) 14 Published in partnership with Seoul National University Bundang Hospital
A. Boussina et al.
7
settings15. We fixed the score threshold to achieve an 80% during their stay. Nurses would have to have the patient’s chart
sensitivity level. Prior work demonstrated that at this sensitivity, open for it to fire. The lockout periods for each acknowledgement
the PPV was 20.1%. reason were: “No infection suspected” 8 h; “Will notify MD
The COMPOSER algorithm is hosted on a cloud-based immediately” 12 h; Sepsis treatment/work-up in progress” 12 h.
healthcare analytics platform that enables access to data elements We defined our pre-intervention time period from January 1st,
in real-time by leveraging the FHIR and HL7v2 standards 2021 to December 6th, 2022. COMPOSER went live December 7th,
(Supplementary Fig. 1)40. Specifically, the Amazon Web Services 2022. Our post-implementation phase was from December 7th,
(AWS)-hosted infrastructure receives a continuous stream of 2022 until April 30th, 2023.
Admit, Discharge, Transfer (ADT) messages from the hospital’s
integration engine to determine the active patients and map their Primary and secondary outcomes
journey through care units. The platform extracts data at an hourly Our primary outcome was in-hospital mortality. Secondary
resolution for these patients using FHIR APIs with OAuth2.0 outcomes included: compliance with our sepsis bundle (initial
authentication and passes the feature set to COMPOSER. Hourly and repeat serum lactate if initial lactate >2 mmol/L, initiation and
frequency was selected to ensure adequate data availability for a completion of a 30 mL/kg crystalloid fluid bolus, checking of blood
prediction. The resulting sepsis risk score and the top features cultures prior to antibiotics, and initiation of intravenous
driving the recommendation are then written to a flowsheet antibiotics within 3 h of time of sepsis), 72-h change in sequential
within the EHR using an HL7v2 outbound message. This flowsheet organ failure assessment (SOFA) score following sepsis onset, ICU
triggers a nurse-facing BPA (Fig. 4) on the chart open which alerts admission, and ICU-free days. ICU-free days were calculated as 30
the caregiver that the patient is at risk of developing severe sepsis less the number of ICU days with in-hospital death and stays
and provides the model’s top reasons. The nurse could acknowl- longer than 30 days were fixed at 0. For patients who died in the
edge the alert by selecting one of the four options: (i) no infection ICU, this value is 0. For example, a patient who is in the ICU for
suspected, (ii) sepsis treatment/workup in progress, or (iii) will 4 days and survives would have a value of 26. Patients who either
notify MD immediately. If a nurse exited the patient’s chart die in the ICU or in the ICU for > 29 days have a value of 0.
without selecting an option, we recorded this as “no acknowl-
edgement”. If the “will notify MD immediately” option is selected,
the nurse can use a ‘secure chat’ feature to contact the provider Statistical methods
from within the BPA to discuss the care of the patient. Descriptive statistics were provided as indicated. Differences
Prospectively deployed algorithms are susceptible to model between the pre-intervention and post-intervention cohort were
drift in which their performance degrades overtime due to assessed with Kruskal–Wallis rank sum tests on continuous
changes in the patient population or treatment practices41. variables and Pearson’s chi-squared tests on categorical variables
To detect this possibility, we implemented a data quality and significance was assessed at a P-value of 0.05. All statistical
dashboard that tracks the median values of all input features to analyses were performed using the R statistical software version
ensure they are within their upper and lower process control limits 4.0.4 and the CausalImpact package version 1.2.743,44.
(based on the upper and lower quantiles from the training cohort). To estimate the causal effect of the COMPOSER BPA interven-
We further evaluate model performance, such as sensitivity and tion, we performed causal inference using a Bayesian structural
positive predictive value (PPV), biweekly to ensure there is no time-series model44. This approach, pioneered by Brodersen et al.
degradation in COMPOSER performance (Supplementary Note 3). from Google Inc., has been widely used to assess the impact of
We established a Predetermined Change Control Plan (PCCP) to advertisement campaigns on product sales and the effect of
trigger model retraining if the performance drops below economic changes on markets45–47. Here, we apply it to patient
predetermined thresholds, although this has not been required outcomes data to assess the impact of the COMPOSER algorithm
as of the time of this reporting. adjusted for confounders. Briefly, a state-space model is trained on
the control time-series prior to the intervention of interest. The
observed outcome is modeled as a function of the latent state and
Implementation of COMPOSER into our electronic Gaussian noise, with the latent state modeled by a local linear
health record trend, in addition to a linear regression on the model covariates. In
Implementation was completed in various stages according to the this work, we assume the effect of the regression coefficients on
EPIS (exploration, preparation, implementation, sustainment) the outcome of interest is independent of time and the static
framework, with frequent feedback provided to nursing staff regression parameters are sampled from a spike-and-slab prior
during the implementation and sustainment periods42. Early distribution. Posterior inference is then performed on the post-
stages in the Exploration stage began 2–3 years prior to model intervention time-series to estimate the counterfactual outcome if
deployment with significant institutional support at the depart- the intervention had not been introduced. Under the assumption
mental and health system level. Preparation began approximately that the response variable in the control time-series is indepen-
6 months prior to the go-live date. Included in this was the dent of the intervention, the difference between the model
creation of a multidisciplinary team to guide implementation, prediction and the observed value is a probability density of the
surveys of the nursing staff to identify specific needs, educational causal impact of the intervention over time (See Supplementary
sessions, and iterative changes to the BPA from end-users. We Note 3 for more details). Additionally, model residuals were
employed a “silent mode trial” in which COMPOSER outputs were evaluated using quantile-quantile and autocorrelation plots to
reviewed in real-time by a team of physicians to assess the ensure adequate modeling of the time-series information.
accuracy and usefulness of the alerts. Adjustments to the Emergency department volume, sex, baseline SOFA, comorbid-
algorithm were iteratively made based on these reviews to ity burden via the Elixhauser comorbidity score, age, COVID-19
improve the timeliness and appropriateness of the alerts. During infection status, ED location (La Jolla or Hillcrest), local trends, and
the Implementation phase, frequent feedback and education were season were included as covariates in the Bayesian structural time-
provided to nursing staff on the COMPOSER model. In final form, series model and 1000 samples of Markov Chain Monte Carlo were
our BPA would fire for all adult (at least 18 years old) patients who used for posterior inference. No imputation was performed since
were receiving care in the ED with a score above the threshold all covariates were fully observed. There were no missing values
and the following exclusions: patient discharged or deceased, present in the covariates. We included seasons and local trends
comfort care measures initiated, patient no longer under the care (e.g., ED volume) as prior data suggest outcomes of sepsis patients
of ED nurses, or a sepsis bundle had previously been instituted are worse in the winter and may be impacted by high patient
Published in partnership with Seoul National University Bundang Hospital npj Digital Medicine (2024) 14
A. Boussina et al.
8
volumes48,49. Outcomes were predicted at a monthly resolution to 16. Adams, R. et al. Prospective, multi-site study of patient outcomes after imple-
reduce the influence of random fluctuations on the outcome mentation of the TREWS machine learning-based early warning system for sepsis.
variable. We plotted the model predictions and true outcomes Nat. Med. 28, 1455–1460 (2022).
data, the pointwise difference between the two, and the 17. Giannini, H. M. et al. A machine learning algorithm to predict severe sepsis and
septic shock: development, implementation, and impact on clinical practice. Crit.
cumulative difference across the post-intervention period and
Care Med. 47, 1485–1492 (2019).
assessed significance against the 95% confidence intervals. 18. Shimabukuro, D. W., Barton, C. W., Feldman, M. D., Mataraso, S. J. & Das, R. Effect
We further evaluated the association of alert acknowledgement of a machine learning-based severe sepsis prediction algorithm on patient sur-
and sepsis intervention, measured by the time from ED triage to vival and hospital length of stay: a randomised clinical trial. BMJ Open Respir. Res.
the administration of antibiotics. We adjusted for the aforemen- 4, e000234 (2017).
tioned confounders and performed a two-sided t-test on time-to- 19. McCoy, A. & Das, R. Reducing patient mortality, length of stay and readmissions
antibiotics as a function of acknowledgement reason. through machine learning-based sepsis prediction in the emergency depart-
ment, intensive care unit and hospital floor units. BMJ Open Qual. 6, e000158
(2017).
Reporting summary 20. Wong, A. et al. External Validation of a Widely Implemented Proprietary Sepsis
Further information on research design is available in the Nature Prediction Model in Hospitalized Patients. JAMA Intern Med. Published online
Research Reporting Summary linked to this article. June 21. [Link] (2021).
21. Lyons, P. G. et al. Factors associated with variability in the performance of a
proprietary sepsis prediction model across 9 networked hospitals in the US. JAMA
DATA AVAILABILITY Intern. Med. 183, 611–612 (2023).
22. Wardi, G. et al. Bringing the promise of artificial intelligence to critical care: what
Access to the de-identified UCSD cohort can be made available by contacting the
the experience with sepsis analytics can teach us. Crit. Care Med. 51, 985–991
corresponding author and via approval from the UCSD Institutional Review Boards
(2023).
(IRB) and Health Data Oversight Committee (HDOC).
23. Classen, D.C., Longhurst, C. & Thomas, E.J. Bending the patient safety curve: how
much can AI help?. npj Digit. Med. 6, 2 (2023). [Link]
022-00731-5.
CODE AVAILABILITY 24. Bektaş, M., Tuynman, J. B., Costa Pereira, J., Burchell, G. L. & van der Peet, D. L.
Access to the R code used in this research is available upon request to the Machine learning algorithms for predicting surgical outcomes after colorectal
corresponding author. Details on the Causal Impact implementation can be found surgery: a systematic review. World J. Surg. 46, 3100–3110 (2022).
here: [Link] 25. Narayanan, N., Gross, A. K., Pintens, M., Fee, C. & MacDougall, C. Effect of an
electronic medical record alert for severe sepsis among ED patients. Am. J. Emerg.
Received: 19 June 2023; Accepted: 6 December 2023; Med. 34, 185–188 (2016).
Published online: 23 January 2024 26. Berger, T., Birnbaum, A., Bijur, P., Kuperman, G. & Gennis, P. A computerized alert
screening for severe sepsis in emergency department patients increases lactate
testing but does not improve inpatient mortality. Appl. Clin. Inf. 1, 394–407
(2010).
REFERENCES
27. Makam, A. N., Nguyen, O. K. & Auerbach, A. D. Diagnostic accuracy and effec-
1. Singer, M. et al. The Third International Consensus Definitions for sepsis and tiveness of automated electronic sepsis alert systems: a systematic review. J.
septic shock (sepsis-3). JAMA 315, 801–810 (2016). Hosp. Med. 10, 396–402 (2015).
2. Rudd, K. E. et al. Global, regional, and national sepsis incidence and mortality, 28. Sendak, M. P. et al. Real-world integration of a sepsis deep learning technology
1990–2017: analysis for the Global Burden of Disease Study. Lancet 395, 200–211 into routine clinical care: implementation study. JMIR Med. Inf. 8, e15182 (2020).
(2020). 29. Carlile, M. et al. Deployment of artificial intelligence for radiographic diagnosis of
3. Rhodes, A. et al. Surviving sepsis campaign: International Guidelines for Man- COVID-19 pneumonia in the emergency department. J. Am. Coll. Emerg. Physi-
agement of Sepsis and Septic Shock: 2016. Intensive Care Med 43, 304–377 cians Open 1, 1459–1464 (2020).
(2017). 30. Amrollahi, F., Shashikumar, S. P., Razmi, F. & Nemati, S. Contextual embeddings
4. Kumar, A. et al. Duration of hypotension before initiation of effective anti- from clinical notes improves prediction of sepsis. AMIA Annu Symp. Proc. AMIA
microbial therapy is the critical determinant of survival in human septic shock. Symp. 2020, 197–202 (2020).
Crit. Care Med 34, 1589–1596 (2006). 31. Goh, K. H. et al. Artificial intelligence in sepsis early prediction and diagnosis
5. Ferrer, R. et al. Empiric antibiotic treatment reduces mortality in severe sepsis and using unstructured data in healthcare. Nat. Commun. 12, 711 (2021).
septic shock from the first hour: results from a guideline-based performance 32. Amrollahi, F. et al. Predicting Hospital Readmission among Patients with Sepsis
improvement program. Crit. Care Med 42, 1749–1755 (2014). Using Clinical and Wearable Data. Annu Int Conf IEEE Eng Med Biol Soc. 2023, 1–4
6. Liu, V. X. et al. The timing of early antibiotics and hospital mortality in sepsis. Am. (2023).
J. Respir. Crit. Care Med 196, 856–863 (2017). 33. Wardi, G. et al. Predicting progression to septic shock in the emergency
7. Peltan, I. D. et al. ED door-to-antibiotic time and long-term mortality in sepsis. department using an externally generalizable machine-learning algorithm. Ann.
Chest 155, 938–946 (2019). Emerg. Med 77, 395–406 (2021).
8. Chamberlain, D. J., Willis, E. M. & Bersten, A. B. The severe sepsis bundles as 34. Holder, A. L., Shashikumar, S. P., Wardi, G., Buchman, T. G. & Nemati, S. A locally
processes of care: a meta-analysis. Aust. Crit. Care J. Confed. Aust. Crit. Care Nurses optimized data-driven tool to predict sepsis-associated vasopressor use in the
24, 229–243 (2011). ICU. Crit. Care Med. 49, e1196–e1205 (2021).
9. Centers for Medicare & Medicaid Services. QualityNet—inpatient hospitals spe- 35. Amrollahi, F., Shashikumar, S. P., Holder, A. L. & Nemati, S. Leveraging clinical data
cifications manual Version 5.13 (2023). [Link] across healthcare institutions for continual learning of predictive risk models. Sci.
specifications-manuals. Rep. 12, 8380 (2022).
10. Reyna, M. A. et al. Early prediction of sepsis from clinical data: The PhysioNet/ 36. Li R. C., et al. Using AI to empower collaborative team workflows: Two imple-
Computing in Cardiology Challenge 2019. Crit. Care Med. 48, 210–217 (2020). mentations for advance care planning and care escalation. NEJM Catal. 2022;3.
11. Shashikumar, S. P., Li, Q., Clifford, G. D. & Nemati, S. Multiscale network repre- [Link]
sentation of physiological time series for early prediction of sepsis. Physiol. Meas. 37. Gibbs, K. D. et al. Evaluation of a sepsis alert in the pediatric acute care setting.
38, 2235–2248 (2017). Appl. Clin. Inf. 12, 469–478 (2021).
12. Lauritsen, S. M. et al. Explainable artificial intelligence model to predict acute 38. von Elm, E. et al. The strengthening the reporting of observational studies in
critical illness from electronic health records. Nat. Commun. 11, 3852 (2020). epidemiology (STROBE) statement: guidelines for reporting observational studies.
13. Nemati, S. et al. An interpretable machine learning model for accurate prediction Lancet 370, 1453–1457 (2007).
of sepsis in the ICU. Crit. Care Med. 46, 547–553 (2018). 39. Rhee, C. et al. Incidence and trends of sepsis in US hospitals using clinical vs
14. Henry, K. E., Hager, D. N., Pronovost, P. J. & Saria, S. A targeted real-time early claims data, 2009-2014. JAMA 318, 1241–1249 (2017).
warning score (TREWScore) for septic shock. Sci. Transl. Med. 7, 299ra122 40. Boussina, A. et al. "Development & Deployment of a Real-time Healthcare Pre-
(2015). dictive Analytics Platform," 2023 45th Annual International Conference of the
15. Shashikumar, S. P., Wardi, G., Malhotra, A. & Nemati, S. Artificial intelligence sepsis IEEE Engineering in Medicine & Biology Society (EMBC), Sydney, Australia, 2023,
prediction algorithm learns to say “I don’t know. NPJ Digit Med. 4, 134 (2021). 1-4, [Link]
npj Digital Medicine (2024) 14 Published in partnership with Seoul National University Bundang Hospital
A. Boussina et al.
9
41. Davis, S. E., Greevy, R. A., Lasko, T. A., Walsh, C. G. & Matheny, M. E. Detection of joint first authors with equal contributions to this work. G.W. and S.N. jointly
calibration drift in clinical prediction models to inform model updating. J. Biomed. supervised this work. All authors contributed to manuscript preparation, critical
Inf. 112, 103611 (2020). revisions, and have read and approved the manuscript.
42. Moullin, J. C., Dickson, K. S., Stadnick, N. A., Rabin, B. & Aarons, G. A. Systematic
review of the exploration, preparation, implementation, sustainment (EPIS) fra-
mework. Implement Sci. IS 14, 1 (2019). COMPETING INTERESTS
43. R Core Team. A Language and Environment for Statistical Computing. R Founda- S.N., A.B., S.S., and A.M. are co-founders of a UCSD start-up, Healcisio Inc., which is
tion for Statistical Computing [Link] (2021) focused on commercialization of advanced analytical decision support tools, and
44. Brodersen, K. H., Gallusser F., Koehler J., Remy N., Scott S. L. Inferring causal impact formed in compliance with UCSD conflict of interest policies. The remaining authors
using Bayesian structural time-series models. Ann. Appl. Stat. 9, 247–274 (2015). declare no competing interests.
45. Takyi, P. O. & Bentum-Ennin, I. The impact of COVID-19 on stock market perfor-
mance in Africa: A Bayesian structural time series approach. J. Econ. Bus. 115,
105968 (2021).
ADDITIONAL INFORMATION
46. Jalan, A., Matkovskyy, R. & Urquhart, A. What effect did the introduction of Bitcoin
futures have on the Bitcoin spot market? Eur. J. Financ. 27, 1251–1281 (2021). Supplementary information The online version contains supplementary material
47. Martin W., Sarro F. & Harman M. Causal impact analysis for app releases in Google available at [Link]
Play. In: Proceedings of the 2016 24th ACM SIGSOFT International Symposium on
Foundations of Software Engineering. 435–446 (ACM, 2016). Correspondence and requests for materials should be addressed to Gabriel Wardi.
48. Danai, P. A., Sinha, S., Moss, M., Haber, M. J. & Martin, G. S. Seasonal variation in
the epidemiology of sepsis. Crit. Care Med. 35, 410–415 (2007). Reprints and permission information is available at [Link]
49. Woodworth, L. Swamped: emergency department crowding and patient mor- reprints
tality. J. Health Econ. 70, 102279 (2020).
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims
in published maps and institutional affiliations.
ACKNOWLEDGEMENTS
G.W. has been supported by the National Foundation of Emergency Medicine and
the National Institutes of Health (#K23GM146092). A.B. is funded by the National Open Access This article is licensed under a Creative Commons
Library of Medicine (#2T15LM011271-11). S.N. is funded by the National Institutes of Attribution 4.0 International License, which permits use, sharing,
Health (#R01LM013998, #R01HL157985, #R35GM143121). S.S. has no sources of adaptation, distribution and reproduction in any medium or format, as long as you give
funding to declare. The opinions or assertions contained herein are the private ones appropriate credit to the original author(s) and the source, provide a link to the Creative
of the author and are not to be construed as official or reflecting the views of the NIH Commons licence, and indicate if changes were made. The images or other third party
or any other agency of the US Government. The authors would like to acknowledge material in this article are included in the article’s Creative Commons licence, unless
the support provided by the Joan & Irwin Jacobs Center for Health Innovation at UC indicated otherwise in a credit line to the material. If material is not included in the
San Diego Health. article’s Creative Commons licence and your intended use is not permitted by statutory
regulation or exceeds the permitted use, you will need to obtain permission directly
from the copyright holder. To view a copy of this licence, visit http://
AUTHOR CONTRIBUTIONS [Link]/licenses/by/4.0/.
A.B., S.S., and S.N. designed the implementation architecture. G.W. led the clinical
preparation efforts. A.B. analyzed the data and synthesized the results. A.M., K.Q., R.O.,
R.E., C.L., A.D., T.C., and G.W. provided clinical interpretation of results. A.B. and S.S. are © The Author(s) 2024, corrected publication 2024
Published in partnership with Seoul National University Bundang Hospital npj Digital Medicine (2024) 14