Machine Learning in Early Warning Systems
Machine Learning in Early Warning Systems
Review
Sankavi Muralitharan1,2, MPharm, MSc; Walter Nelson1*, BSc; Shuang Di1,3*, BSc, MEd, MSc; Michael McGillion4,5,
BScN, PhD; PJ Devereaux5,6, MD, PhD, FRCPC; Neil Grant Barr7, BA, MSc, PhD; Jeremy Petch1,5,8,9, HBA, MA,
PhD
1
Centre for Data Science and Digital Health, Hamilton Health Sciences, Hamilton, ON, Canada
2
DeGroote School of Business, McMaster University, Hamilton, ON, Canada
3
Dalla Lana School of Public Health, University of Toronto, Toronto, ON, Canada
4
School of Nursing, McMaster University, Hamilton, ON, Canada
5
Population Health Research Institute, Hamilton, ON, Canada
6
Departments of Health Evidence and Impact and Medicine, McMaster University, Hamilton, ON, Canada
7
Health Policy and Management, DeGroote School of Business, McMaster University, Hamilton, ON, Canada
8
Institute of Health Policy, Management and Evaluation, University of Toronto, Toronto, ON, Canada
9
Department of Medicine, Faculty of Health Sciences, McMaster University, Hamilton, ON, Canada
*
these authors contributed equally
Corresponding Author:
Sankavi Muralitharan, MPharm, MSc
Centre for Data Science and Digital Health
Hamilton Health Sciences
293 Wellington St. N
Hamilton, ON, L8L 8E7
Canada
Phone: 1 2897882965
Email: sankavi_22@[Link]
Abstract
Background: Timely identification of patients at a high risk of clinical deterioration is key to prioritizing care, allocating
resources effectively, and preventing adverse outcomes. Vital signs–based, aggregate-weighted early warning systems are
commonly used to predict the risk of outcomes related to cardiorespiratory instability and sepsis, which are strong predictors of
poor outcomes and mortality. Machine learning models, which can incorporate trends and capture relationships among parameters
that aggregate-weighted models cannot, have recently been showing promising results.
Objective: This study aimed to identify, summarize, and evaluate the available research, current state of utility, and challenges
with machine learning–based early warning systems using vital signs to predict the risk of physiological deterioration in acutely
ill patients, across acute and ambulatory care settings.
Methods: PubMed, CINAHL, Cochrane Library, Web of Science, Embase, and Google Scholar were searched for peer-reviewed,
original studies with keywords related to “vital signs,” “clinical deterioration,” and “machine learning.” Included studies used
patient vital signs along with demographics and described a machine learning model for predicting an outcome in acute and
ambulatory care settings. Data were extracted following PRISMA, TRIPOD, and Cochrane Collaboration guidelines.
Results: We identified 24 peer-reviewed studies from 417 articles for inclusion; 23 studies were retrospective, while 1 was
prospective in nature. Care settings included general wards, intensive care units, emergency departments, step-down units, medical
assessment units, postanesthetic wards, and home care. Machine learning models including logistic regression, tree-based methods,
kernel-based methods, and neural networks were most commonly used to predict the risk of deterioration. The area under the
curve for models ranged from 0.57 to 0.97.
Conclusions: In studies that compared performance, reported results suggest that machine learning–based early warning systems
can achieve greater accuracy than aggregate-weighted early warning systems but several areas for further research were identified.
While these models have the potential to provide clinical decision support, there is a need for standardized outcome measures to
allow for rigorous evaluation of performance across models. Further research needs to address the interpretability of model outputs
by clinicians, clinical efficacy of these systems through prospective study design, and their potential impact in different clinical
settings.
KEYWORDS
machine learning; early warning systems; clinical deterioration; ambulatory care; acute care; remote patient monitoring; vital
signs; sepsis; cardiorespiratory instability; risk prediction
full-text articles, inclusion in the review, and extraction of study • Qualitative studies, reviews, preprints, case reports,
data were carried out by a single author. commentaries, or conference proceedings.
Search Strategy Study Selection
We searched PubMed, CINAHL, Cochrane Library, Web of References from the preliminary searches were handled using
Science, Embase, and Google Scholar for peer-reviewed studies Mendeley reference management software. After duplicates
without using any filters for study design and language. Searches were removed, titles and abstracts were screened to assess
were also conducted without any date restrictions. The reference preliminary eligibility. Eligible studies were then read in full
lists of all studies that met the inclusion criteria were screened length to be assessed against the inclusion and exclusion criteria.
for additional articles. The search strategy involved a series of
searches using a combination of relevant keywords and
Data Extraction
synonyms, including “vital signs,” “clinical deterioration,” and Data were extracted from eligible studies using an extraction
“machine learning.” See Multimedia Appendix 1 for search sheet that followed the PRISMA [18] and Cochrane
terms. Collaboration guidelines for systematic reviews [19] and the
Transparent Reporting of a Multivariable Prediction Model for
Eligibility Criteria Individual Prognosis or Diagnosis (TRIPOD) guidelines [20]
The inclusion criteria covered the following: for the reporting of predictive models. Study characteristics,
setting, demographics, patient outcomes, ML model
• Peer-reviewed studies evaluating continuous or intermittent
characteristics, and model performance data were extracted.
vital sign monitoring in adult patients so that all data
The model performance results were extracted from the
collection or sampling frequencies (eg, 1 measurement per
validation data set rather than from the model derivation or
minute vs 1 measurement every 2 hours) wedre taken into
training data set to decrease the potential for model overfitting.
consideration;
When studies explored multiple ML models, the model with
• Studies conducted using data gathered from all acute and
the best performance was selected for reporting and comparison.
ambulatory care settings including medical or surgical
If studies compared the performance of ML models to
hospital wards, ICUs, step-down units, ED, and in-home
aggregate-weighted EWS, then the performance data of these
care;
warning systems were also extracted.
• Quantitative, observational, retrospective, and prospective
cohort studies and randomized controlled trials;
• Studies that involved ML or multivariable statistical or ML
Results
models and reported some model performance measure (eg, Search Results and Study Selection
area under the curve) [17];
• Studies that reported mortality or any outcomes related to The search for “vital signs” AND “clinical deterioration” AND
clinical deterioration so that EWS models and performance “machine learning” using the same query terms and filters
can be examined for all explored outcomes. identified 417 studies after duplicate removal. During the title
and abstract screening process, 386 studies were excluded. Of
The exclusion criteria included the following: the 31 full-text articles that were assessed, 7 studies were
• Studies that used any laboratory values as predictors for excluded for not meeting the eligibility criteria: 2 studies did
the ML-based EWS, as this review focuses on examining not use ML models to predict deterioration, 3 studies included
time-sensitive predictions of clinical deterioration using vital sign measurements in addition to laboratory values as
patient parameters that are readily available across all care predictors, 1 study focused on a cohort of pregnant women, and
settings; 1 study did not meet the criteria for model performance
• Studies involving pediatric or obstetric populations due to measures. A review of the reference lists of the 24 selected
these patients having different or altered physiologies that studies did not yield any additional studies fulfilling the
cannot be compared to standard adult patients; eligibility criteria (refer to Figure 1).
Authors, Setting(s) Data collec- Cohort descrip- Event rate Study purpose Predictors Measurement Outcome
year tion tion frequency
Desautels Beth Israel ICU bedside 22,853 ICU 2577 (11.28%) Validate a sepsis GCS, HR, RR, At least 1 Onset of sepsis
et al, 2016 Deaconess monitors and stays stays with con- prediction SpO2, temper- measurement
[37] Medical medical firmed sepsis method, InSight, ature, invasive per hour
Center ICU records (MIM- for the new Sep- and noninva-
ICm-III) sis-3 definitions sive SBP and
and make predic- DBP
tions using a min-
imal set of vari-
ables
Forkan et Beth Israel ICU bedside 1023 patients Not specified Develop a proba- HR, SBP, All samples Abnormal clini-
al, 2017 Deaconess monitors and bilistic model for DBP, mean converted to cal events
[28] Medical medical predicting the fu- BP, RR, SpO2 per-minute
Center ICU records (MIM- ture clinical sampling
IC-II) episodes of a pa-
tient using ob-
served vital sign
values prior to
the clinical event
Forkan et Beth Israel ICU bedside 85 patients Not specified Develop an intel- HR, SBP, Per-minute Patient-specific
al, 2017 Deaconess monitors and ligent method for DBP, mean sampling anomalies, dis-
[27] Medical medical personalized BP, RR, SpO2 ease symptoms,
Center ICU records (MIM- monitoring and and emergen-
IC & MIMIC- clinical decision cies
II) support through
early estimation
of patient-specif-
ic vital sign val-
ues
Forkan et Beth Israel ICU bedside 4893 patients Not specified Build a prognos- HR, SBP, Per-minute Dangerous clini-
al, 2017 Deaconess monitors and tic model, ViSi- DBP, mean sampling cal events
[29] Medical medical BiD, that can ac- BP, RR, SpO2
Center ICU records (MIM- curately identify
IC-II) dangerous clini-
cal events of a
home-monitored
patient in ad-
vance
Guillame- Step-down Bedside moni- 297 admis- 127 patients Forecast CRI uti- HR, RR, Every 20 sec- At least 1 event
Bert et al, unit tor measure- sions (43%) exhibited lizing data from SPO2, SBP, onds (HR, threshold limit
2017 [43] ments over 8 at least 1 real continuous moni- DBP, mean RR, SPO2), criteria exceed-
weeks event during toring of physio- BP every 2 hours ed for >80% of
their stay logic vital sign (SBP, DBP, last 3 minutes
measurements and mean BP)
Ho et al, Beth Israel ICU bedside 763 patients 197 patients Build a cardiac Temperature, 1 reading per Cardiac arrest
2017 [38] Deaconess monitors and (25.8%) experi- arrest risk predic- SpO2, HR, hour
Medical medical enced a cardiac tion model capa- RR, DBP,
Center ICU records (MIM- arrest event ble of early notifi- SBP, pulse
IC-II) cation at time z (z pressure index
≥5 hours prior to
the event)
Jang et al, ED visits to EHR data Nontraumatic 374,605 eligible Develop and test Age, sex, Not specified Development of
2019 [35] a tertiary from ED visits ED visits ED visits of artificial neural chief com- cardiac arrest
academic 233,763 pa- network classi- plaint, SBP, within 24 hours
hospital tients; 1097 fiers for early de- DBP, HR, RR, after prediction
(0.3%) patients tection of patients temperature,
with cardiac ar- at risk of cardiac AVPU
rest arrest in EDs
Authors, Setting(s) Data collec- Cohort descrip- Event rate Study purpose Predictors Measurement Outcome
year tion tion frequency
Kwon et al, Cardiovascu- Data collected 52,131 pa- 419 patients Predict whether SBP, HR, RR, 3 times a day Primary out-
2018 [26] lar teaching manually by tients (0.8%) with car- an input vector temperature on general come: first car-
hospital and staff on gener- diac arrest; 814 belonged within wards, every diac arrest; sec-
community al wards, by (1.56%) deaths the prediction 10 minutes in ondary out-
general hos- bedside moni- without attempt- time window ICUs come: death
pital tors in ICUs ed resuscitation (0.5-24 hours be- without attempt-
fore the outcome) ed resuscitation
Kwon et al, 151 EDs in Korean Nation- 10,967,518 153,217 (1.4%) Validate that a Age, sex, At ED admis- Primary out-
2018 [11] Korea al Emergency ED visits in-hospital DTASn identifies chief com- sion come: in-hospi-
Department deaths; 625,117 high-risk patients plaint, time tal mortality;
Information (5.7%) critical more accurately from symptom secondary out-
System care admis- than existing onset to ED come: critical
(NEDIS) sions; triage and acuity visit, arrival care; tertiary
2,964,367 scores mode, trauma, outcome: hospi-
(27.0%) hospi- initial vital talization
talizations signs (SBP,
DBP, HR, RR,
temperature),
mental status
Larburu et OSI Bilbao- Collected 242 patients 202 predictable Prevent mobile SBP, DBP, At diagnosis Heart failure
al, 2018 Basurto (Os- manually by decompensa- heart failure pa- HR, SaO2, and 3-7 times decompensation
[22] akidetza) clinicians and tions tients’decompen- weight per week in
Hospital and patients sation using pre- ambulatory
ED admis- dictive models patients
sions, ambu-
latory
Li et al, Beth Israel ICU bedside 12 patients Not specified Adaptive online HR, SBP, At least 1 Signs of deterio-
2016 [39] Deaconess monitors and monitoring of pa- DBP, MAPo, measurement ration
Medical medical tients in ICUs RR per hour
Center ICU records (MIM-
IC-II)
Liu et al, ED of a ter- Manual vital 702 patients 29 (4.13%) pa- Discover the SBP, RR, HR Not specified Composite of
2014 [36] tiary hospital measurements with undiffer- tients met prima- most relevant events such as
in Singapore by nurses or entiated, non- ry outcome variables for risk death and car-
physicians traumatic prediction of ma- diac arrest with-
chest pain jor adverse car- in 72 hours of
diac events using arrival at the
clinical signs and ED
HR variability
Mao et al, ICU, inpa- UCSFp UCSF: 90,353 UCSF: 1179 Sepsis prediction SBP, DBP, Hourly Sepsis, severe
2018 [34] tient wards, dataset:inpa- patients; (1.3%) sepsis, HR, RR, sepsis, septic
outpatient tient and outpa- MIMIC-III: 349 (0.39%) se- SpO2, temper- shock
visits tient visits; 21,604 pa- vere sepsis, 614 ature
MIMIC-III: tients (0.68%) septic
ICU bedside shock; MIMIC-
monitors III: sepsis
(1.91%), severe
sepsis (2.82%),
septic shock
(4.36%)
Olsen et al, PACUq, IntelliVue 178 patients 160 (89.9%) Develop a predic- SpO2, SBP, Every minute Signs of deterio-
2018 [46] Rigshospi- MP5, BM- had ≥1 mi- tive algorithm de- HR, MAP (SpO2, SBP, ration
talet, Univer- EYE Nexfin croevent occur- tecting early HR), every 15
sity of bedside moni- ring during ad- signs of deteriora- minutes
Copenhagen, tors during ad- mission; 116 tion in the PACU (MAP)
Denmark mission to patients using continuous-
post anesthetic (65.2%) had ≥1 ly collected car-
care unit microevent with diopulmonary vi-
a duration >15 tal signs
minutes
Authors, Setting(s) Data collec- Cohort descrip- Event rate Study purpose Predictors Measurement Outcome
year tion tion frequency
Shashiku- Adult ICU ICU bedside Patients with 242 sepsis cases Predict onset of MAP, HR, ≥1 measure- Onset of sepsis
mar et al, units monitors, Bed- unselected sepsis 4 hours SpO2, SBP, ment per hour
2017 [40] master sys- mixed surgical ahead of time, us- DBP, RR,
tem; up to 24 procedures ing commonly GCS, tempera-
hours of moni- measured vital ture, comorbid-
toring signs ity, clinical
context, admis-
sion unit, sur-
gical special-
ty, wound
type, age, gen-
der, weight,
race
Tarassenko General Bedside moni- 150 general- Not specified A real-time auto- HR, RR, Every 30 min- Signs of deterio-
et al, 2006 wards at tors for at ward patients mated system, SpO2, skin utes (BP), ev- ration
[32] John Rad- least 24 hours BioSign, which temperature, ery 5 seconds
cliffe Hospi- per patient tracks patient sta- average SBP - (other vitals)
tal in Ox- tus by combining average DBP
ford, United information from
Kingdom vital signs
Van Wyk Methodist Bedside moni- 2995 patients 343 patients Classify patients HR, MAP, Every minute Sepsis detection
et al, 2017 LeBonheur tors: Cerner (11.5%) diag- into sepsis and DBP, SBP,
[33] Hospital, CareAware nosed with sep- nonsepsis groups SpO2, age,
Memphis, iBus system sis using data collect- race, gender,
TN ed at various fre- fraction of in-
quencies from the spired oxygen
first 12 hours af-
ter admission
Yoon et al, Beth Israel ICU bedside 2809 subjects 787 tachycardia Predicting tachy- Arterial DBP, 1/60 Hz or 1 Tachycardia
2019 [41] Deaconess monitors and episodes cardia as a surro- arterial SBP, Hz episode
Medical medical gate for instabili- HR, RR,
Center ICU records (MIM- ty SpO2, MAP
IC-II)
a
ICU: intensive care unit.
b
NEWS: National Early Warning Score.
c
HR: heart rate.
d
RR: respiratory rate.
e
SBP: systolic blood pressure.
f
AVPU: alert, verbal, pain, unresponsive.
g
CRI: cardiorespiratory instability.
h
DBP: diastolic blood pressure.
i
ED: emergency department.
j
EHR: electronic health record.
k
GCS: Glasgow Coma Score.
l
NHS: National Health Service.
m
MIMIC: Medical Information Mart for Intensive Care.
n
DTAS: Deep learning–based Triage and Acuity Score.
o
MAP: mean arterial pressure.
p
UCSF: University of California, San Francisco.
q
PACU: postanesthesia care unit.
Table 2. Machine learning (ML) models and comparisons used for outcome prediction.
Study Cohort Event rate ML model(s) Missing data Best ML model ML model com- Predic- Aggregate
handling performance parisons tion win- weighted
dow EWSa compar-
isons
Badriyah et 35,585 ad- 199 (0.56%), car- Decision tree Not speci- Decision tree pre- Not specified Within 24 NEWSd AU-
al, 2014 [45] missions diac arrest; analysis fied dicted cardiac ar- hours pre- ROC: cardiac
1161 (3.26%) rest: AU- ceding arrest, 0.722;
unanticipated ROCc=0.708; events unanticipated
b
ICU admissions; unanticipated ICU ICU admis-
1789 (5.02%) admission: AU- sion, 0.857;
deaths; 3149 ROC=0.862; death, 0.894;
(8.85%) any out- death: AU- any outcomes,
come ROC=0.899; any 0.873
outcomes: AU-
ROC=0.877
Chen at al, 1880 pa- 997 patients Variant of the Not speci- Random forest Logistic regres- Within 4 No compari-
2017 [44] tients (1971 (53%) or 1056 random forest fied f
AUC initially re- sion: AUC=0.7; hours pre- son
admissions) admissions classification mained constant lasso logistic re- ceding
(53.6%) who ex- model using (0.58-0.60), fol- gression: events
perienced CRIe nonrandom lowed by an in- AUC=0.82
events splits creasing trend,
with AUCs rising
from 0.57 to 0.89
during the 4 hours
immediately pre-
ceding events
Churpek et 269,999 ad- 16,452 outcomes Univariate Forward im- Trends increased Not specified Within 4 No compari-
al, 2016 [24] missions (6.09%) analysis, bi- putation, me- model accuracy hours pre- son
variate analy- dian value compared to a ceding
sis imputation model containing events
only current vital
signs (AUC 0.78
vs 0.74); vital sign
slope improved
AUC by 0.013
Chiew et al, 214 patients 40 patients K-nearest Not speci- Gradient boosting K-nearest neigh- Within 30 SEDSi:
2019 [23] (18.7%) met out- neighbor, ran- fied predicted 30-day bor: F1 days pre- F1=0.40,
come dom forest, sepsis-related mor- score=0.10, ceding AUPRC=0.22;
adaptive tality: F1 AUPRC=0.10, event
qSOFAj:
boosting, gra- score=0.50, precision
F1=0.32,
dient boost- AUPRC=0.35, pre- (PPV)=0.33, re-
AUPRC=0.21;
ing, support cision call=0.6; random
NEWS;
vector ma- (PPVg)=0.62, re- forest: F1
F1=0.38,
chine call=0.5 score=0.35,
AUPRC=0.28;
AUPRC=0.27,
precision MEWSk:
(PPV)=0.26, re- F1=0.30,
call=0.56; adap- AUPRC=0.25
tive boosting: F1
score=0.40,
AUPRC=0.31,
precision
(PPV)=0.43, re-
call=0.38; SVMh:
F1 score=0.43,
AUPRC=0.29,
precision
(PPV)=0.33, re-
call=0.63
Study Cohort Event rate ML model(s) Missing data Best ML model ML model com- Predic- Aggregate
handling performance parisons tion win- weighted
dow EWSa compar-
isons
Chiu et al, Adults under- 578 patients Logistic re- Observations Logistic regression Not specified Within NEWS: 24
2019 [42] going risk- (4.2%) with an gression with missing predicted the event 24, 12, hours before
stratified ma- outcome; 499 pa- values were 24 hours in ad- and 6 event,
jor cardiac tients (3.66%) excluded vance: AU- hours pre- AU-
surgery with unplanned ROC=0.779; 12 ceding ROC=0.754;
(n=13,631) ICU readmissions hours in advance: event 12 hours be-
AUROC=0.815; 6 fore event,
hours in advance: AU-
AUROC=0.841 ROC=0.789; 6
hours before
event, AU-
ROC=0.813
Clifton et al, 200 patients Not specified Classifiers, Missing SVM predicted de- Conventional Not speci- No compari-
2014 [25] in the postop- Gaussian pro- channels re- terioration: accura- SVM: accura- fied son
erative ward cess, one-class placed by cy=0.94, partial cy=0.90, partial
following support vector mean of that AUC=0.28, sensi- AUC=0.26, sensi-
upper gas- machine, ker- channel tivity=0.96, speci- tivity=0.92,
trointestinal nel estimate ficity=0.93 specificity=0.87;
cancer Gaussian mixture
surgery models: accura-
cy=0.9, partial
AUC=0.24, sensi-
tivity=0.97,
specificity=0.84;
Gaussian process-
es: accura-
cy=0.90, partial
AUC=0.26, sensi-
tivity=0.91,
specificity=0.89;
kernel density es-
timate: accura-
cy=0.91, partial
AUC=0.26, sensi-
tivity=0.94,
specificity=0.87
Desautels et 22,853 ICU 2577 (11.28%) Insight classifi- Carry for- Classifier predicts Not specified Within 4 SIRSm: AU-
al, 2016 [37] stays stays with con- er ward imputa- sepsis at onset: hours pre- ROC= 0.609,
firmed sepsis tion AUROC=0.880, ceding APR= 0.160;
APRl=0.6, accura- event and qSOFA: AU-
cy=0.8; classifier at time of ROC= 0.772,
predicts sepsis 4 event on- APR=0.277;
hours before onset: set MEWS: AU-
AUROC=0.74, ROC=0.803,
APR=0.28, accura- APR=0.327;
cy=0.57 SAPSn II:
AU-
ROC=0.700,
APR=0.225;
SOFA: AU-
ROC=0.725,
APR=0.284
Study Cohort Event rate ML model(s) Missing data Best ML model ML model com- Predic- Aggregate
handling performance parisons tion win- weighted
dow EWSa compar-
isons
Forkan et al, 1023 pa- Not specified PCAo used to Data with Hidden Markov Neural network: Within 30 No compari-
2017 [28] tients separate pa- consecutive Model event predic- accuracy=93% minutes son
tients into missing val- tion: accura- preceding
multiple cate- ues over a cy=97.8%, preci- event
gories; hidden long period sion=92.3, sensitiv-
Markov Mod- are eliminat- ity=97.7, specifici-
el adopted for ed ty=98, F-
probabilistic score=95%
classification
and future pre-
diction
Forkan et al, 85 patients Not specified Multilabel Where ≥1 vi- Predictions across Not specified Within 1 No compari-
2017 [27] classification tal signs data 24 classifier combi- hour pre- son
algorithms are are missing nations yielded a ceding
applied in while clean Hamming score of event
classifier de- values of 90%-95%; F1-mi-
sign; result others are cro average of
analysis with available, 70.1%-84%; accu-
J48 decision considered racy of 60.5%-
tree, random as recover- 77.7%
tree and se- able and im-
quential mini- puted using
mal optimiza- median-pass
tion (SMO, a and k-near-
simplified ver- est neighbor
sion of SVM) filter
Forkan et al, 4893 pa- Not specified J48 decision Data with Event prediction J48 decision tree: 1 hour No compari-
2017 [29] tients tree, random consecutive by random forest: within a 60- preceding son
forest, sequen- missing val- within a 60-minute minute forecast event
tial minimal ues over a forecast horizon, F horizon, F
optimization, long period score=0.96, accura- score=0.93, accu-
MapReduce are eliminat- cy=95.86; within a racy=92.46; with-
random forest ed 90-minute forecast in a 90-minute
horizon, F- forecast horizon,
score=0.95, accura- F score=0.92, ac-
cy=95.35; within a curacy=91.59;
120-minute fore- within a 120-
cast horizon, F- minute forecast
score=0.95, accura- horizon, F
cy=95.18 score=0.91, accu-
racy=91.30;
Event prediction
with sequential
minimal optimiza-
tion: within a 60-
minute forecast
horizon, F
score=0.91, accu-
racy=90.72; with-
in a 90-minute
forecast horizon,
F score=0.90, ac-
curacy=90.08;
within a 120-
minute forecast
horizon,
F score=0.89, ac-
curacy=89.23
Study Cohort Event rate ML model(s) Missing data Best ML model ML model com- Predic- Aggregate
handling performance parisons tion win- weighted
dow EWSa compar-
isons
Guillame- 297 admis- 127 patients TITAp rules, Not speci- Event forecast alert Random forest: Within 17 No compari-
Bert et al, sions (43%) exhibited rule fusion al- fied within 17 minutes, event forecast minutes, son
2017 [43] at least 1 real gorithm; map- 51 seconds before alert within 11 51 sec-
CRI event during ping function onset of CRI (false minutes, 25 sec- onds pre-
their stay in the from rule- alert every 12 onds before onset ceding
step-down unit based features hours); event fore- of CRI (false CRI onset
to forecast cast alert within 10 alert every 12
model learned minutes, 58 sec- hours); event
using random onds before onset forecast alert
forest classifi- of CRI (false alert within 5 minutes,
er every 24 hours) 52 seconds be-
fore onset of CRI
(false alert every
24 hours)
Ho et al, 763 patients 197 patients Temporal Imputed val- TTL-Reg predicts Not specified Within 6 No compari-
2017 [38] (25.8%) experi- transfer learn- ues based on events with an hours pre- son
enced a cardiac ing-based the median AUC of 0.63 ceding
arrest event model (TTL- from patients event
Reg) of the same
gender and
similar ages
Jang et al, Non-traumat- 374,605 eligible ANNq with Not speci- Event prediction: Random forest, Within 24 MEWS: AU-
2019 [35] ic ED visits ED visits of multilayer per- fied ANN with multilay- AUROC=0.923; hours pre- ROC=0.886
233,763 patients; ceptron, ANN er perceptron, AU- logistic regres- ceding
1097 (0.3%) pa- ROC=0.929; ANN sion, AU- event
with LSTMr,
tients with car- with LSTM, AU- ROC=0.914
hybrid ANN;
diac arrest ROC=0.933; hy-
comparison
brid ANN, AU-
with random
ROC=0.936
forest and lo-
gistic regres-
sion
Kwon et al, 52,131 pa- 419 patients 3 RNNs layers Most recent Event prediction: Random forest, 30 min- MEWS: AU-
2018 [26] tients (0.8%) with car- with LSTM to value was RNNs, AU- AUROC=0.78, utes to 24 ROC=0.603,
diac arrest; 814 deal with time used; if no ROC=0.85, AUPRC=0.014; hours pre- AUPRC=0.003
(1.56%) deaths series data; value avail- AUPRCt=0.044 logistic regres- ceding
without attempt- compared to able, then sion, AU- event
ed resuscitation random forest median val- ROC=0.613,
and logistic re- ue used AUPRC=0.007
gression
Kwon et al, 10,967,518 153,217 (1.4%) DTASu using Excluded Event prediction: Random forest: Not speci- Korean triage
2018 [11] ED visits in-hospital multilayer per- DTAS using multi- AUROC= 0.89, fied and acuity
deaths; 625,117 ceptron with 5 layer perceptron, AUPRC= 0.14; score: AU-
(5.7%) critical hidden AUROC=0.935, logistic regres- ROC =0.785,
care admissions; AUPRC=0.264 sion: AUROC= AUPRC=0.192;
layers
2,964,367 0.89, MEWS: AU-
(27.0%) hospital- AUPRC=0.16 ROC=0.810,
izations AUPRC=0.116;
Larburu et 242 patients 202 predictable Naïve Bayes, Not speci- Decompensation Decision tree, Not speci- No compari-
al, 2018 [22] decompensations decision tree, fied event prediction: neural network, fied son
random forest, naïve Bayes, random forest,
SVM AUC=67% support vector
machine, stochas-
tic gradient de-
scent
Study Cohort Event rate ML model(s) Missing data Best ML model ML model com- Predic- Aggregate
handling performance parisons tion win- weighted
dow EWSa compar-
isons
Li et al, 12 patients Not specified L-PCA (com- Not speci- Fault detection rate Not specified Not speci- No compari-
2016 [39] bination of fied with L-PCA: 20% fied son
just-in-time higher than with
learning and PCA; 47% higher
PCA) than with fast
moving-window
PCA; best detec-
tion rate achieved
was 99.8%
Liu et al, 702 patients 29 (4.13%) pa- Novel variable Not speci- Event prediction Not specified Within 72 TIMIv:
2014 [36] with undiffer- tients met prima- selection fied with ensemble hours of AUC=0.637;
entiated, ry outcome framework learning model: arrival at MEWS:
non-traumat- based on en- AUC=0.812, cut- ED AUC=0.622
ic chest pain semble learn- off score=43, sensi-
ing; random tivity=82.8%,
forest was the specificity=63.4%
independent
variable selec-
tor for creat-
ing the deci-
sion ensemble
Mao et al, UCSFw: UCSF: 1179 Gradient tree Carry for- Detection with gra- Not specified At onset MEWS: AU-
2018 [34] 90,353 pa- (1.3%) sepsis, boosting + ward imputa- dient tree boosting: of sepsis ROC=0.76;
tients; MIM- 349 (0.39%) se- transfer learn- tion AUROC=0.92 for and se- SOFA: AU-
vere sepsis, 614 ing using sepsis; AU- vere sep- ROC=0.65;
ICx-III:
(0.68%) septic MIMIC-III as ROC=0.87 for se- sis; with- SIRS: AU-
21,604 pa-
shock; MIMIC- source and vere sepsis at on- in 4 hours ROC=0.72
tients
III: sepsis UCSF as tar- set; AUROC=0.96 preceding
(1.91%), severe get for septic shock 4 septic
sepsis (2.82%), hours before; AU- shock and
septic shock ROC=0.85 for se- severe
(4.36%) vere sepsis predic- sepsis
tion 4 hours before
Olsen et al, 178 patients 160 (89.9%) had Random forest Not speci- Detection of early Not specified Not speci- Compared
2018 [46] ≥1 microevent classifier fied signs of deteriora- fied with hospital's
occurring during tion with random current alarm
admission; 116 forest: accura- system: num-
patients (65.2%) cy=92.2%, sensitiv- ber of false
had ≥1 mi- ity=90.6%, speci- alarms de-
croevent with a ficity=93.0%, AU- creased by
duration >15 ROC=96.9% 85%, number
minutes of missed ear-
ly signs of de-
terioration de-
creased by
73%
Study Cohort Event rate ML model(s) Missing data Best ML model ML model com- Predic- Aggregate
handling performance parisons tion win- weighted
dow EWSa compar-
isons
Shashikumar Patients with 242 sepsis cases Elastic net lo- Median val- Event prediction: Not specified 4 hours No compari-
et al, 2017 unselected gistic classifi- ues (if multi- elastic net logistic prior to son
[40] mixed surgi- er ple measure- classifier using en- onset
cal proce- ment were tropy features
dures available); alone, AU-
otherwise, ROC=0.67, accura-
the old val- cy=47%; elastic
ues were net logistic classifi-
kept (sam- er using social de-
ple-and-hold mographics +
interpola- EMRy features,
tion); mean AUROC=0.7, accu-
imputation racy=50%; elastic
for replacing net logistic classifi-
all remaining er using all fea-
missing val- tures, AU-
ues ROC=0.78, accura-
cy=61%
Tarassenko 150 general- Not specified Biosign; data Historic, me- 95% of Biosign Not specified Within No compari-
et al, 2006 ward pa- fusion dian filtering alerts were classi- 120 min- son
[32] tients method: proba- fied as “True” by utes of
bilistic model clinical experts event
of normality
in five dimen-
sions
Van Wyk et 2995 pa- 343 patients CNNz (con- Not speci- Event classifica- Event classifica- Not speci- No compari-
al, 2017 [33] tients (11.5%) diag- structed im- fied tion with a 1- tion with a 1- fied son
nosed with sepsis ages using raw minute observation minute observa-
patient data) frequency: CNN, tion frequency:
with random accuracy=86.1%; multilayer percep-
dropout to re- event classification tron, accura-
duce overfit- with a 10-minute cy=76.5%;
ting; multilay- observation fre- event classifica-
er perceptron quency: CNN, ac- tion with a 10-
with random curacy=78.2% minute observa-
dropout be- tion frequency:
tween layers multilayer percep-
to avoid over- tron, accura-
fitting cy=71%
Study Cohort Event rate ML model(s) Missing data Best ML model ML model com- Predic- Aggregate
handling performance parisons tion win- weighted
dow EWSa compar-
isons
Yoon et al, 2809 sub- 787 tachycardia Regularized Discrete Event prediction: Logistic regres- Within 3 No compari-
2019 [41] jects episodes logistic regres- Fourier random forest, sion with L1 regu- hours pre- son
sion and ran- transform, AUC=0.869, accu- larization, ceding
dom forest cubic-spline racy=0.806 AUC=0.8284, ac- onset
classifiers interpolation curacy=0.7668
of heart rate
and respirato-
ry rate data
for missing
data as long
as ≥20% of
the data
were avail-
able
a
EWS: early warning system.
b
ICU: intensive care unit.
c
AUROC: area under the receiver operator characteristic.
d
NEWS: National Early Warning Score.
e
CRI: cardiorespiratory instability.
f
AUC: area under the curve.
g
PPV: positive predictive value.
h
SVM: support vector machine.
i
SEDS: Singapore Emergency Department Sepsis.
j
qSOFA: quick Sequential Organ Failure Assessment.
k
MEWS: Modified Early Warning Score.
l
APR: area under the precision-recall curve.
m
SIRS: systemic inflammatory response syndrome.
n
SAPS II: simplified acute physiology score.
o
PCA: principal component analysis.
p
TITA: temporal interval tree association.
q
ANN: artificial neural network.
r
LSTM: long short-term memory.
s
RNN: recurrent neural network.
t
AUPRC: area under the precision-recall curve.
u
DTAS: Deep learning–based Triage and Acuity Score.
v
TIMI: Thrombolysis in Myocardial Infarction.
w
UCSF: University of California, San Francisco.
x
MIMIC: Medical Information Mart for Intensive Care.
y
EMR: electronic medical record.
z
CNN: convolutional neural network.
an AUROC of 0.779 using logistic regression, compared to ambulatory setting, particularly postdischarge. For example,
0.754 using MEWS for the same 24-hour prediction window. the VISION study [54] found that 1.8% of all patients die within
A full side-by-side comparison of ML vs aggregate-weighted 30 days postsurgery and 29.4% of all deaths occurred after
EWS is presented in Multimedia Appendix 3. patients were discharged from hospital. Patients often receive
postoperative monitoring only 3-4 weeks [54] after discharge
Discussion during a follow-up visit with their surgeon. During this period,
it has been shown that many patients suffer from prolonged
Based on this scoping review, ML-based EWS models show unidentified hypoxemia [55] and hypotension [56], which are
considerable promise, but there exist several important avenues precursors to serious postoperative complications. While EWS
for future research if these models are to be effectively research has historically focused on inpatient settings due to
implemented in clinical practice. the availability of continuous vital signs data, the increasing
Prediction Window availability of remote patient monitoring and wearable
technologies offer the opportunity to direct future EWS research
A model’s prediction window refers to how far in advance a to the ambulatory setting to address a significant clinical need.
model is predicting an adverse event. Most studies included in
our review used a prediction window between 30 minutes [26] Retrospective Versus Prospective Evaluation
and 72 hours [36] before the clinical deterioration took place. All but one study [21] included in this review were retrospective
The length of a model’s prediction window is important because in nature, leaving open the possibility that algorithm
a prediction window that is too short will not yield any real performance in a clinical environment may be lower than the
clinical benefit (it would not give a clinical team sufficient time performance achieved in a controlled retrospective setting [34].
to intervene), but a number of studies [29,34,37,42] showed a It is also unclear how often these EWS were able to identify
decrease in model performance when the prediction window clinical deterioration that had not already been detected by a
was longer (eg, AUROC drops from 0.88 at the time of onset care team. Further, alerts for clinical deterioration may be easily
to 0.74 at 4 hours before the event). Future research seeking to disregarded by clinicians due to alert fatigue, even when the
maximize the clinical benefit of ML EWS should strive to risk of deterioration has been correctly identified [43]. In the
achieve an optimum balance between a clinically relevant single case where an ML-based EWS was studied prospectively,
prediction window and clinically acceptable model performance, Olsen et al [21] found that the random forest classifier decreased
rather than simply maximizing a model performance metric, false alarm rates by 85% and the rate of missed alerts by 73%
such as AUROC. when compared to the existing aggregate-weighted alarm
Clinically Actionable Explanations system. While the predictions were independently scored for
severity by 2 clinician experts, the interpretation of the clinical
The studies included in this review focused on ML model impact of these alerts was not explored any further, leaving the
development and did not explore how the output of these models question of clinical benefit unanswered. Future research into
would be communicated to clinicians. Since many ML models ML-based EWS should begin to include prospective evaluation,
are “black boxes” [46,47], it may not be immediately clear to both of model accuracy (to understand how model performance
clinicians what the likely reason for an alert might be until the is affected when faced with real-world data) and of clinical
patient is assessed, which can cause further delays in outcomes (to understand whether alerts in fact produce clinical
time-sensitive scenarios. However, in the broader ML field, benefits).
there has been significant recent progress in explainable ML
techniques, and it has been pointed out that these approaches Standardizations of Performance Metrics
may be preferred by the medical community and regulators A key observation from this review is the lack of an agreed-upon
[48,49]. Several explanation methods take specific, previously standard among the research community for reporting
black-box methods, such as convolutional neural networks [50], performance measures across studies. This makes meaningful
and allow for post-hoc explanation of their decision-making comparison between the outcomes of these studies difficult,
process. Other explainability algorithms are model-agnostic, and where there is overlap, it is not clear that the most clinically
meaning they can be applied to any type of model, regardless relevant metrics have been chosen. The majority of the studies
of its mathematical basis [51]. In the study by Lauritsen et al in this review report the AUROC as the main performance
[52], an explainable EWS was developed based on a temporal metric, reflecting a common practice in the ML literature.
convolutional network, using a separate module for explanations. However, AUROC may not be adequate for evaluating the
These methodologies are promising, but their application to performance of the EWS in a clinical setting [57].
health care, including to EWS, has been limited. Objective
evaluation of the utility of explanation methods is a difficult, As Romero-Brufau et al [58] discussed in their article, AUROC
ongoing problem, but is an important direction for future does not incorporate information about the prevalence of
research in the area of ML-based EWS if they are to be physiological deterioration, which can be lower than 0.02 daily
effectively deployed in clinical practice [53]. in a general inpatient setting. This can make AUROC a
misleading metric, leading to overestimation of clinical benefit
Expanded Study Settings and underestimation of clinical workload and resources. [58]
Nearly all the studies included in this review were conducted When the prevalence is low (<0.1), even a model with high
in inpatient settings. While EWS are highly valuable in an sensitivity and specificity may not yield a high posttest
inpatient context, there is also considerable need in the probability for a positive prediction [15]. Therefore, reporting
metrics that incorporate the prevalence would be more the original search, indicating our search strategy was
appropriate. comprehensive. Unlike previous reviews, inclusion criteria for
the review supported the examination of findings from studies
The performance of an EWS depends on the tradeoff between
conducted across a variety of clinical settings including specialty
2 goals: early detection of outcomes versus issuance of fewer
units or wards and ambulatory care. This helped in
false-positive alerts to prevent alarm fatigue [43]. Sensitivity
characterizing the use of ML-based prediction models in
can be a good metric to evaluate the first goal as it would
different patient-care environments with varying clinical
provide the percentage of true-positive predictions within a
endpoints. Wherever the original studies provided the data,
certain time period. To evaluate the clinical burden of
comparisons were drawn between the performance of the ML
false-positive alerts, the positive predictive value, which
models and that of aggregate-weighted EWS. This gives an
incorporates prevalence, can be used as it gives a percentage of
indication of the differences in accuracy of the models in
useful alerts that lead to a clinical outcome. The number needed
predicting clinical deterioration.
to evaluate can be a useful measure of clinical utility and
cost-efficiency of each alert as it provides the number of patients Limitations
that need to be evaluated further to detect one outcome. Using The findings within this review are subject to some limitations.
these metrics to evaluate tradeoffs between outcome detection First, the literature search, assessment of eligibility of full-text
and workload would be essential for determining the clinical articles, inclusion in the review, and extraction of study data
utility of the EWS [58]. Additionally, the F1 score can also be were carried out by only 1 author. Second, only the findings
a useful metric as it provides a measure of the model’s overall from published studies were included in this scoping review,
accuracy through the calculation of the harmonic mean of the which may affect the results due to publication bias. While
precision and recall (sensitivity). Balancing the use of these 2 studies from a variety of settings were included, the
metrics could yield a more realistic measure of the model’s generalizability of our findings may be limited due to the
performance [58]. heterogeneity of patient populations, clinical practices, and
Comparison to “Gold Standard” EWS study methodologies. Sampling procedures and frequencies
varied across studies from single to multiple observations of
On a related note, only 9 of the studies included in our review
patient vital signs, and clinical outcome definitions were based
made comparisons between their ML-based models and a “gold
on different criteria or aggregate-weighted EWS. Finally, due
standard” aggregate-weighted EWS, such as MEWS or NEWS.
to this variation in ML methods, prediction windows, and
Future research in the area should report a commonly used
outcome reporting, a meta-analysis was not feasible.
aggregate-weighted EWS as a baseline model, which would aid
in making effective comparisons between them. NEWS may Conclusion
be particularly well suited to this area of research as its input Our findings suggest that ML-based EWS models incorporating
variables can all be measured automatically and continuously easily accessible vital sign measurements are effective in
via devices. predicting physiological deterioration in patients. Improved
Strengths of the Review prediction performance was also observed with these models
when compared to traditional aggregate-based risk stratification
The search strategy was comprehensive while not being too
tools. The clinical impact of these ML-based EWS could be
focused on specific clinical outcomes, sampling frequencies,
significant for clinical staff and patients due to decreased false
or filtering for time. This allowed for the identification of as
alerts and increased early detection of warning signs for timely
many studies as possible that examined the use of ML models
intervention, though further development of these models is
and vital signs to predict the risk of patient deterioration. No
needed and the necessary prospective research to establish actual
additional studies were identified through citation tracking after
clinical utility does not yet exist.
Authors' Contributions
SM contributed to conceptualization, data collection, data analysis, and manuscript writing. JP contributed to conceptualization,
manuscript writing, and manuscript review. WN and SD contributed equally to manuscript writing and review. PD contributed
to manuscript writing and review. MM and NB contributed to manuscript review.
Conflicts of Interest
PJD is a member of a research group with a policy of not accepting honorariums or other payments from industry for their own
personal financial gain. They do accept honorariums/payments from industry to support research endeavours and costs to participate
in meetings.
Based on study questions PJD has originated and grants he has written, he has received grants from Abbott Diagnostics, AstraZeneca,
Bayer, Boehringer Ingelheim, Bristol-Myers-Squibb, Coviden, Octapharma, Philips Healthcare, Roche Diagnostics, Siemens,
and Stryker.
PJD has participated in advisory board meetings for GlaxoSmithKline, Boehringer Ingelheim, Bayer, and Quidel Canada. He
also attended an expert panel meeting with AstraZeneca and Boehringer Ingelheim.
The other authors declare no conflicts of interest.
Multimedia Appendix 1
Search terms.
[DOCX File , 12 KB-Multimedia Appendix 1]
Multimedia Appendix 2
Description of ML methods and relevant terms.
[DOCX File , 15 KB-Multimedia Appendix 2]
Multimedia Appendix 3
Comparison between performance of ML based EWS and aggregate EWS.
[DOCX File , 20 KB-Multimedia Appendix 3]
References
1. Barfod C, Lauritzen MMP, Danker JK, Sölétormos G, Forberg JL, Berlac PA, et al. Abnormal vital signs are strong predictors
for intensive care unit admission and in-hospital mortality in adults triaged in the emergency department - a prospective
cohort study. Scand J Trauma Resusc Emerg Med 2012 Apr 10;20:28 [FREE Full text] [doi: 10.1186/1757-7241-20-28]
[Medline: 22490208]
2. Hillman KM, Bristow PJ, Chey T, Daffurn K, Jacques T, Norman SL, et al. Antecedents to hospital deaths. Intern Med J
2001 Aug;31(6):343-348. [doi: 10.1046/j.1445-5994.2001.00077.x] [Medline: 11529588]
3. McGaughey J, Alderdice F, Fowler R, Kapila A, Mayhew A, Moutray M. Outreach and Early Warning Systems (EWS)
for the prevention of intensive care admission and death of critically ill adult patients on general hospital wards. Cochrane
Database Syst Rev 2007 Jul 18(3):CD005529. [doi: 10.1002/14651858.CD005529.pub2] [Medline: 17636805]
4. Smith GB, Prytherch DR, Schmidt PE, Featherstone PI, Higgins B. A review, and performance evaluation, of single-parameter
"track and trigger" systems. Resuscitation 2008 Oct;79(1):11-21. [doi: 10.1016/[Link].2008.05.004] [Medline:
18620794]
5. Gardner-Thorpe J, Love N, Wrightson J, Walsh S, Keeling N. The value of Modified Early Warning Score (MEWS) in
surgical in-patients: a prospective observational study. Ann R Coll Surg Engl 2006 Oct;88(6):571-575 [FREE Full text]
[doi: 10.1308/003588406X130615] [Medline: 17059720]
6. Gao H, McDonnell A, Harrison DA, Moore T, Adam S, Daly K, et al. Systematic review and evaluation of physiological
track and trigger warning systems for identifying at-risk patients on the ward. Intensive Care Med 2007 Apr;33(4):667-679.
[doi: 10.1007/s00134-007-0532-3] [Medline: 17318499]
7. Subbe CP, Slater A, Menon D, Gemmell L. Validation of physiological scoring systems in the accident and emergency
department. Emerg Med J 2006 Nov;23(11):841-845 [FREE Full text] [doi: 10.1136/emj.2006.035816] [Medline: 17057134]
8. Smith GB, Prytherch DR, Meredith P, Schmidt PE, Featherstone PI. The ability of the National Early Warning Score
(NEWS) to discriminate patients at risk of early cardiac arrest, unanticipated intensive care unit admission, and death.
Resuscitation 2013 Apr;84(4):465-470. [doi: 10.1016/[Link].2012.12.016] [Medline: 23295778]
9. Tam B, Xu M, Kwong M, Wardell C, Kwong A, Fox-Robichaud A. The Admission Hamilton Early Warning Score (HEWS)
Predicts the Risk of Critical Event during Hospitalization. Can Journ Gen Int Med 2017 Feb 24;11(4):1. [doi:
10.22374/cjgim.v11i4.190]
10. Prytherch DR, Smith GB, Schmidt PE, Featherstone PI. ViEWS--Towards a national early warning score for detecting
adult inpatient deterioration. Resuscitation 2010 Aug;81(8):932-937. [doi: 10.1016/[Link].2010.04.014] [Medline:
20637974]
11. Kwon J, Lee Y, Lee Y, Lee S, Park H, Park J. Validation of deep-learning-based triage and acuity score using a large
national dataset. PLoS One 2018;13(10):e0205836 [FREE Full text] [doi: 10.1371/[Link].0205836] [Medline:
30321231]
12. Bates DW, Saria S, Ohno-Machado L, Shah A, Escobar G. Big data in health care: using analytics to identify and manage
high-risk and high-cost patients. Health Aff (Millwood) 2014 Jul;33(7):1123-1131. [doi: 10.1377/hlthaff.2014.0041]
[Medline: 25006137]
13. Churpek MM, Yuen TC, Park SY, Gibbons R, Edelson DP. Using electronic health record data to develop and validate a
prediction model for adverse outcomes in the wards*. Crit Care Med 2014 Apr;42(4):841-848 [FREE Full text] [doi:
10.1097/CCM.0000000000000038] [Medline: 24247472]
14. Linnen DT, Escobar GJ, Hu X, Scruth E, Liu V, Stephens C. Statistical Modeling and Aggregate-Weighted Scoring Systems
in Prediction of Mortality and ICU Transfer: A Systematic Review. J Hosp Med 2019 Mar;14(3):161-169 [FREE Full text]
[doi: 10.12788/jhm.3151] [Medline: 30811322]
15. Brekke IJ, Puntervoll LH, Pedersen PB, Kellett J, Brabrand M. The value of vital sign trends in predicting and monitoring
clinical deterioration: A systematic review. PLoS One 2019;14(1):e0210875 [FREE Full text] [doi:
10.1371/[Link].0210875] [Medline: 30645637]
16. Tricco AC, Lillie E, Zarin W, O'Brien KK, Colquhoun H, Levac D, et al. PRISMA Extension for Scoping Reviews
(PRISMA-ScR): Checklist and Explanation. Ann Intern Med 2018 Oct 02;169(7):467-473 [FREE Full text] [doi:
10.7326/M18-0850] [Medline: 30178033]
17. Zweig MH, Campbell G. Receiver-operating characteristic (ROC) plots: a fundamental evaluation tool in clinical medicine.
Clin Chem 1993 Apr;39(4):561-577. [Medline: 8472349]
18. Moher D, Liberati A, Tetzlaff J, Altman DG, PRISMA Group. Preferred reporting items for systematic reviews and
meta-analyses: the PRISMA statement. PLoS Med 2009 Jul 21;6(7):e1000097 [FREE Full text] [doi:
10.1371/[Link].1000097] [Medline: 19621072]
19. Higgins JPT, Green S. Cochrane Handbook for Systematic Reviews of Interventions Version 5.1.0. The Cochrane
Collaboration. 2011. URL: [Link] [accessed 2021-01-09]
20. Collins GS, Reitsma JB, Altman DG, Moons KGM. Transparent Reporting of a multivariable prediction model for Individual
Prognosis Or Diagnosis (TRIPOD): the TRIPOD Statement. Br J Surg 2015 Feb;102(3):148-158. [doi: 10.1002/bjs.9736]
[Medline: 25627261]
21. Olsen RM, Aasvang EK, Meyhoff CS, Dissing Sorensen HB. Towards an automated multimodal clinical decision support
system at the post anesthesia care unit. Comput Biol Med 2018 Oct 01;101:15-21. [doi: 10.1016/[Link].2018.07.018]
[Medline: 30092398]
22. Larburu N, Artetxe A, Escolar V, Lozano A, Kerexeta J. Artificial Intelligence to Prevent Mobile Heart Failure Patients
Decompensation in Real Time: Monitoring-Based Predictive Model. Mobile Information Systems 2018 Nov 05;2018:1-11.
[doi: 10.1155/2018/1546210]
23. Chiew CJ, Liu N, Tagami T, Wong TH, Koh ZX, Ong MEH. Heart rate variability based machine learning models for risk
prediction of suspected sepsis patients in the emergency department. Medicine (Baltimore) 2019 Feb;98(6):e14197 [FREE
Full text] [doi: 10.1097/MD.0000000000014197] [Medline: 30732136]
24. Churpek MM, Adhikari R, Edelson DP. The value of vital sign trends for detecting clinical deterioration on the wards.
Resuscitation 2016 May;102:1-5 [FREE Full text] [doi: 10.1016/[Link].2016.02.005] [Medline: 26898412]
25. Clifton L, Clifton DA, Pimentel MAF, Watkinson PJ, Tarassenko L. Predictive monitoring of mobile patients by combining
clinical observations with data from wearable sensors. IEEE J Biomed Health Inform 2014 May;18(3):722-730. [doi:
10.1109/JBHI.2013.2293059] [Medline: 24808218]
26. Kwon J, Lee Y, Lee Y, Lee S, Park J. An Algorithm Based on Deep Learning for Predicting In-Hospital Cardiac Arrest. J
Am Heart Assoc 2018 Jun 26;7(13):1 [FREE Full text] [doi: 10.1161/JAHA.118.008678] [Medline: 29945914]
27. Forkan ARM, Khalil I. PEACE-Home: Probabilistic estimation of abnormal clinical events using vital sign correlations
for reliable home-based monitoring. Pervasive and Mobile Computing 2017 Jul;38:296-311. [doi: 10.1016/[Link].2016.12.009]
28. Forkan ARM, Khalil I. A clinical decision-making mechanism for context-aware and patient-specific remote monitoring
systems using the correlations of multiple vital signs. Comput Methods Programs Biomed 2017 Feb;139:1-16. [doi:
10.1016/[Link].2016.10.018] [Medline: 28187881]
29. Forkan ARM, Khalil I, Atiquzzaman M. ViSiBiD: A learning model for early discovery and real-time prediction of severe
clinical events using vital signs as big data. Computer Networks 2017 Feb;113:244-257. [doi: 10.1016/[Link].2016.12.019]
30. Lee J, Scott D, Villarroel M, Clifford G, Saeed M, Mark R. Open-access MIMIC-II database for intensive care research.
Conf Proc IEEE Eng Med Biol Soc 2011;2011:8315-8318. [doi: 10.1109/iembs.2011.6092050] [Medline: 22256274]
31. Johnson AEW, Pollard TJ, Shen L, Lehman LH, Feng M, Ghassemi M, et al. MIMIC-III, a freely accessible critical care
database. Sci Data 2016 May 24;3:160035 [FREE Full text] [doi: 10.1038/sdata.2016.35] [Medline: 27219127]
32. Tarassenko L, Hann A, Young D. Integrated monitoring and analysis for early warning of patient deterioration. Br J Anaesth
2006 Jul;97(1):64-68 [FREE Full text] [doi: 10.1093/bja/ael113] [Medline: 16707529]
33. van Wyk F, Khojandi A, Kamaleswaran R, Akbilgic O, Nemati S, Davis RL. How much data should we collect? A case
study in sepsis detection using deep learning. 2017 Presented at: IEEE Healthcare Innovations and Point of Care Technologies
(HI-POCT); November 6-8, 2017; Bethesda, MD. [doi: 10.1109/hic.2017.8227596]
34. Mao Q, Jay M, Hoffman JL, Calvert J, Barton C, Shimabukuro D, et al. Multicentre validation of a sepsis prediction
algorithm using only vital sign data in the emergency department, general ward and ICU. BMJ Open 2018 Jan 26;8(1):e017833
[FREE Full text] [doi: 10.1136/bmjopen-2017-017833] [Medline: 29374661]
35. Jang D, Kim J, Jo YH, Lee JH, Hwang JE, Park SM, et al. Developing neural network models for early detection of cardiac
arrest in emergency department. Am J Emerg Med 2020 Jan;38(1):43-49. [doi: 10.1016/[Link].2019.04.006] [Medline:
30982559]
36. Liu N, Koh ZX, Goh J, Lin Z, Haaland B, Ting BP, et al. Prediction of adverse cardiac events in emergency department
patients with chest pain using machine learning for variable selection. BMC Med Inform Decis Mak 2014 Aug 23;14:75
[FREE Full text] [doi: 10.1186/1472-6947-14-75] [Medline: 25150702]
37. Desautels T, Calvert J, Hoffman J, Jay M, Kerem Y, Shieh L, et al. Prediction of Sepsis in the Intensive Care Unit With
Minimal Electronic Health Record Data: A Machine Learning Approach. JMIR Med Inform 2016 Sep 30;4(3):e28 [FREE
Full text] [doi: 10.2196/medinform.5909] [Medline: 27694098]
38. Ho J, Park Y. Learning from different perspectives: Robust cardiac arrest prediction via temporal transfer learning. Annu
Int Conf IEEE Eng Med Biol Soc 2017 Jul;2017:1672-1675. [doi: 10.1109/EMBC.2017.8037162] [Medline: 29060206]
39. Li X, Wang Y. Adaptive online monitoring for ICU patients by combining just-in-time learning and principal component
analysis. J Clin Monit Comput 2016 Dec;30(6):807-820. [doi: 10.1007/s10877-015-9778-4] [Medline: 26392184]
40. Shashikumar SP, Stanley MD, Sadiq I, Li Q, Holder A, Clifford GD, et al. Early sepsis detection in critical care patients
using multiscale blood pressure and heart rate dynamics. J Electrocardiol 2017;50(6):739-743 [FREE Full text] [doi:
10.1016/[Link].2017.08.013] [Medline: 28916175]
41. Yoon JH, Mu L, Chen L, Dubrawski A, Hravnak M, Pinsky MR, et al. Predicting tachycardia as a surrogate for instability
in the intensive care unit. J Clin Monit Comput 2019 Dec;33(6):973-985 [FREE Full text] [doi: 10.1007/s10877-019-00277-0]
[Medline: 30767136]
42. Chiu Y, Villar SS, Brand JW, Patteril MV, Morrice DJ, Clayton J, et al. Logistic early warning scores to predict death,
cardiac arrest or unplanned intensive care unit re-admission after cardiac surgery. Anaesthesia 2020 Feb;75(2):162-170
[FREE Full text] [doi: 10.1111/anae.14755] [Medline: 31270799]
43. Guillame-Bert M, Dubrawski A, Wang D, Hravnak M, Clermont G, Pinsky MR. Learning temporal rules to forecast
instability in continuously monitored patients. J Am Med Inform Assoc 2017 Jan;24(1):47-53 [FREE Full text] [doi:
10.1093/jamia/ocw048] [Medline: 27274020]
44. Chen L, Ogundele O, Clermont G, Hravnak M, Pinsky MR, Dubrawski AW. Dynamic and Personalized Risk Forecast in
Step-Down Units. Implications for Monitoring Paradigms. Ann Am Thorac Soc 2017 Mar;14(3):384-391 [FREE Full text]
[doi: 10.1513/AnnalsATS.201611-905OC] [Medline: 28033032]
45. Badriyah T, Briggs JS, Meredith P, Jarvis SW, Schmidt PE, Featherstone PI, et al. Decision-tree early warning score
(DTEWS) validates the design of the National Early Warning Score (NEWS). Resuscitation 2014 Mar;85(3):418-423. [doi:
10.1016/[Link].2013.12.011] [Medline: 24361673]
46. Cabitza F, Rasoini R, Gensini GF. Unintended Consequences of Machine Learning in Medicine. JAMA 2017 Aug
08;318(6):517-518. [doi: 10.1001/jama.2017.7797] [Medline: 28727867]
47. Stead WW. Clinical Implications and Challenges of Artificial Intelligence and Deep Learning. JAMA 2018 Sep
18;320(11):1107-1108. [doi: 10.1001/jama.2018.11029] [Medline: 30178025]
48. Tonekaboni S, Joshi S, McCradden MD, Goldenberg A. What Clinicians Want: Contextualizing Explainable Machine
Learning for Clinical End Use. Proceedings of the 4th Machine Learning for Healthcare Conference 2019;106:359-380.
49. Wiens J, Saria S, Sendak M, Ghassemi M, Liu VX, Doshi-Velez F, et al. Do no harm: a roadmap for responsible machine
learning for health care. Nat Med 2019 Sep;25(9):1337-1340. [doi: 10.1038/s41591-019-0548-6] [Medline: 31427808]
50. Wang ZJ, Turko R, Shaikh O, Park H, Das N, Hohman F, et al. CNN EXPLAINER: Learning Convolutional Neural
Networks with Interactive Visualization. IEEE Trans. Visual. Comput. Graphics 2020:1. [doi: 10.1109/tvcg.2020.3030418]
51. Ribeiro MT, Singh S, Guestrin C. "Why Should I Trust You?": Explaining the Predictions of Any Classifier. Cornell
University. 2016. URL: [Link] [accessed 2021-01-09]
52. Lauritsen SM, Kristensen M, Olsen MV, Larsen MS, Lauritsen KM, Jørgensen MJ, et al. Explainable artificial intelligence
model to predict acute critical illness from electronic health records. Nat Commun 2020 Jul 31;11(1):3852 [FREE Full
text] [doi: 10.1038/s41467-020-17431-x] [Medline: 32737308]
53. Sokol K, Flach P. Explainability Fact Sheets: A Framework for Systematic Assessment of Explainable Approaches. Cornell
University. 2019. URL: [Link] [accessed 2021-01-09]
54. Vascular Events in Noncardiac Surgery Patients Cohort Evaluation (VISION) Study Investigators, Spence J, LeManach
Y, Chan MT, Wang CY, Sigamani A, et al. Association between complications and death within 30 days after noncardiac
surgery. CMAJ 2019 Jul 29;191(30):E830-E837 [FREE Full text] [doi: 10.1503/cmaj.190221] [Medline: 31358597]
55. Sun Z, Sessler DI, Dalton JE, Devereaux PJ, Shahinyan A, Naylor AJ, et al. Postoperative Hypoxemia Is Common and
Persistent: A Prospective Blinded Observational Study. Anesth Analg 2015 Sep;121(3):709-715 [FREE Full text] [doi:
10.1213/ANE.0000000000000836] [Medline: 26287299]
56. Turan A, Chang C, Cohen B, Saasouh W, Essber H, Yang D, et al. Incidence, Severity, and Detection of Blood Pressure
Perturbations after Abdominal Surgery: A Prospective Blinded Observational Study. Anesthesiology 2019 Apr;130(4):550-559
[FREE Full text] [doi: 10.1097/ALN.0000000000002626] [Medline: 30875354]
57. Cook NR. Use and misuse of the receiver operating characteristic curve in risk prediction. Circulation 2007 Feb
20;115(7):928-935 [FREE Full text] [doi: 10.1161/CIRCULATIONAHA.106.672402] [Medline: 17309939]
58. Romero-Brufau S, Huddleston JM, Escobar GJ, Liebow M. Why the C-statistic is not informative to evaluate early warning
scores and what metrics to use. Crit Care 2015 Aug 13;19:285 [FREE Full text] [doi: 10.1186/s13054-015-0999-1] [Medline:
26268570]
Abbreviations
AUROC: area under the receiver operating characteristic
AVPU: alert, verbal, pain, unresponsive
BP: blood pressure
ED: emergency department
EWS: early warning system
Edited by G Eysenbach; submitted 21.10.20; peer-reviewed by N Liu, J Kellett; comments to author 07.11.20; revised version received
19.12.20; accepted 20.12.20; published 04.02.21
Please cite as:
Muralitharan S, Nelson W, Di S, McGillion M, Devereaux PJ, Barr NG, Petch J
Machine Learning–Based Early Warning Systems for Clinical Deterioration: Systematic Scoping Review
J Med Internet Res 2021;23(2):e25187
URL: [Link]
doi: 10.2196/25187
PMID: 33538696
©Sankavi Muralitharan, Walter Nelson, Shuang Di, Michael McGillion, PJ Devereaux, Neil Grant Barr, Jeremy Petch. Originally
published in the Journal of Medical Internet Research ([Link] 04.02.2021. This is an open-access article distributed
under the terms of the Creative Commons Attribution License ([Link] which permits
unrestricted use, distribution, and reproduction in any medium, provided the original work, first published in the Journal of
Medical Internet Research, is properly cited. The complete bibliographic information, a link to the original publication on
[Link] as well as this copyright and license information must be included.
The clinical implementation of ML-based Early Warning Systems (EWS) faces several challenges. First, these models are often computationally intensive, which can result in increased and unsustainable clinical workloads . There is also the challenge of variability in model performance due to differences in outcome measures, prediction windows, and data quality . Implementing these systems requires alignment with existing clinical practices and workflows, necessitating a careful balance between predictive accuracy and clinical applicability . Additionally, there is a need for overcoming the lack of standardized performance metrics and clear definitions of clinical deterioration, which complicates model validation and generalization across different clinical settings .
Current studies on ML-based Early Warning Systems (EWS) reveal several methodological trends and limitations. One trend is the use of ML models as classification tasks to predict patient deterioration, employing a variety of models, such as tree-based, linear, and neural networks, with diverse performance metrics like AUROC . A significant limitation is the variability in performance outcomes based on chosen prediction windows, the ML method employed, and the outcome measures being predicted . There is also a call for standardized performance metrics and clearer definitions of deterioration outcomes across different clinical contexts . Highlighting these limitations points towards a need for more extensive controlled trials and a consensus on methodological approaches to enhance the efficacy and comparability of studies in this field .
Standardization in the development and evaluation of ML-based Early Warning Systems (EWS) plays a crucial role in improving their reliability and efficacy. Standardized performance metrics are essential to objectively compare different EWS models across studies and ensure consistency in evaluating their predictive capabilities . In addition, consistent definitions of clinical deterioration outcomes are necessary to accurately assess model effectiveness across various clinical settings . Standardization aids researchers in identifying best practices and common methodological trends, which is critical for translating research into practical, clinically applicable systems. This uniformity also facilitates the aggregation of evidence, making it easier to compare results and refine approaches based on broader datasets .
The demand for new evidence and systematic reviews in the state of ML-based EWS research stems from the need to identify common methodological trends, assess the existing body of evidence, and address the limitations of current approaches. While existing systematic reviews have highlighted that ML models outperform aggregate-weighted EWS in accuracy, they also underscore the lack of standardized metrics and definitions of clinical deterioration outcomes . As ML continues to evolve rapidly, ongoing studies can benefit from synthesized evidence to improve model validity, applicability, and integration into clinical practice. Furthermore, systematic reviews can help guide future research directions by identifying gaps, such as the efficacy of intermittently monitored vital sign trends, and provide recommendations to optimize model development .
Prediction accuracy among different ML models used in Early Warning Systems (EWS) varies significantly based on the model type, the outcome being predicted, and the chosen prediction window. For instance, tree-based models, such as decision trees, and linear models, like logistic regression, have different performance metrics, typically evaluated by the AUROC . For cardiac-related events, decision trees showed varied AUROCs; for predicting cardiac arrest, the AUROC was 0.708, whereas unanticipated ICU admissions had a higher AUROC of 0.862 . Random forest models have shown increasing accuracy during specific prediction windows, illustrating how model choice and temporal factors affect predictive performance. These differences highlight the necessity of selecting models based on specific clinical objectives and data characteristics .
Integration with Electronic Health Records (EHR) significantly enhances the functionality of machine learning-based Early Warning Systems (EWS) by facilitating continuous analysis of vital sign measurements and patient data to provide real-time predictions of patient outcomes . This integration allows ML models to utilize comprehensive data sets, including varying clinical covariates and historical patient data, to produce personalized risk assessments, which can improve early detection of clinical deterioration . Additionally, EHR integration supports the seamless workflow of healthcare providers by embedding decision support directly into the systems they use daily, increasing the likelihood of timely interventions .
Trends in vital sign measurements are crucial for detecting clinical deterioration because they provide dynamic insight into a patient's physiological changes over time. A 2019 systematic review found that despite the value of intermittent vital sign trends, there is a lack of extensive research on this approach, emphasizing the need for controlled trials to better understand their utility in clinical settings . These trends allow for the early identification of deterioration by highlighting deviations from baseline measurements, which is particularly valuable in intermittently monitored setups like hospital wards . Such trends are integral to ML-based EWS, enhancing the prediction of adverse events by closely monitoring fluctuations in key physiological indicators .
Controlled trials are necessary in research concerning intermittently monitored vital sign trends to establish definitive evidence of their efficacy in detecting clinical deterioration. A systematic review by Brekke et al identified that although vital sign trends have shown value in noticing clinical deterioration, the existing evidence is limited to just a few retrospective studies . Controlled trials can provide high-quality, robust data that can confirm these preliminary findings and help in developing standardized protocols for utilizing trends in vital sign measurements efficiently. These trials can also address existing gaps and variabilities in current methodologies, leading to better-informed and reliable EWS models .
Machine learning-based Early Warning Systems (EWS) offer several advantages over aggregate-weighted EWS. Firstly, they can process and analyze large volumes of data, adjusting for varying numbers of clinical covariates, which allows them to be tailored for different care settings and populations . ML models can also continuously incorporate trends in risk scores, enhancing prediction accuracy, and they have shown consistently better performance metrics compared to aggregate-weighted models . Additionally, ML models integrated into electronic health records can provide real-time predictions of patient outcomes, improving clinical decision support . However, ML models are computationally intensive, which may increase clinical workload, a factor highlighted in systematic reviews calling for standardized performance metrics to fully realize ML models' potential .
Trends in the frequencies of vital sign measurements are utilized to enhance the predictive accuracy of patient outcomes by capturing both acute and prolonged physiological changes that may indicate deterioration. Frequent and real-time monitoring, such as every 5 seconds for vitals or every 30 minutes for blood pressure, allows for the identification of subtle changes and trends that can precede adverse events . This frequent data collection provides a detailed dynamic profile that machine learning models can analyze to predict outcomes like cardiovascular instability or sepsis. Variability in measurement frequencies, adapted to different care settings, ensures the models are responsive to the specific clinical context, thereby increasing the reliability of predictions .