0% found this document useful (0 votes)
3 views41 pages

Machine Learning Deep Learning and Data Preprocessing

This review paper examines the application of machine learning (ML) techniques for detecting, predicting, and monitoring mental stress and related disorders, highlighting the significant public health impact of mental health issues. It synthesizes findings from 98 peer-reviewed studies, identifying key ML algorithms such as Support Vector Machines and Neural Networks as effective tools for analyzing physiological data related to stress. The paper also discusses research gaps and future directions, emphasizing the need for improved model interpretability and real-time processing capabilities in mental health applications.

Uploaded by

Harish
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views41 pages

Machine Learning Deep Learning and Data Preprocessing

This review paper examines the application of machine learning (ML) techniques for detecting, predicting, and monitoring mental stress and related disorders, highlighting the significant public health impact of mental health issues. It synthesizes findings from 98 peer-reviewed studies, identifying key ML algorithms such as Support Vector Machines and Neural Networks as effective tools for analyzing physiological data related to stress. The paper also discusses research gaps and future directions, emphasizing the need for improved model interpretability and real-time processing capabilities in mental health applications.

Uploaded by

Harish
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Review Paper

Moein Razavia,b,f, Samira Ziyadidegana, Ahmadreza Mahmoudzadehe,f, Saber Kazeminasabc,


Elaheh Baharloueid, Vahid Janfazab, Reza Jahromia,b, Farzan Sasangohara,*
a
Department of Industrial and Systems Engineering, Texas A&M University
b
Department of Computer Science and Engineering, Texas A&M University
c
Harvard Medical School, Harvard University
d
Department of Computer Science, University of Houston
e
Engineering Academic and Student Affairs, College of Engineering, Texas A&M University
f
Ford Motor Company, Global Data Insight & Analytics

Corresponding author. 3131 TAMU, College Station, TX, 77843, USA. E-mail address: sasangohar@[Link]

Machine Learning, Deep Learning and Data Preprocessing


Techniques for Detection, Prediction, and Monitoring of Stress and
Stress-related Mental Disorders: A Scoping Review
Abstract
Background: Mental stress and its consequent mental disorders (MDs) constitute a significant
public health issue. With the advent of machine learning (ML), there's potential to harness
computational techniques for better understanding and addressing mental stress and MDs. This
comprehensive review seeks to elucidate the current ML methodologies employed in this domain
to pave the way for enhanced detection, prediction, and analysis of mental stress and its subsequent
mental disorders.
Objective: This review aims to investigate the scope of Machine Learning (ML) methodologies
employed in the detection, prediction, and analysis of mental stress and its consequent mental
disorders (MDs).
Methods: Utilizing a rigorous scoping review process with Preferred Reporting Items for
Systematic Reviews and Meta-Analyses extention for Scoping Reviews (PRISMA-ScR)
guidelines, this investigation delves into the latest ML algorithms, preprocessing techniques, and
data types employed in the context of stress and stress-related MDs.
Results and Discussion: Total of 98 peer-reviewed publication were examined for this review.
The findings highlight that Support Vector Machine (SVM), Neural Network (NN), and Random
Forest (RF) models consistently exhibit superior accuracy and robustness among all machine
learning algorithms examined. Physiological parameters such as heart rate measurements and skin
response are prevalently used as stress predictors due to their rich explanatory information
concerning stress and stress-related MDs, as well as the relative ease of data acquisition. The
application of dimensionality reduction techniques, including mappings, feature selection,
filtering, and noise reduction, is frequently observed as a crucial step preceding the training of ML
algorithms.
Conclusion: The synthesis of this review identifies significant research gaps and outlines future
directions for the field. These encompass areas such as model interpretability, model
personalization, the incorporation of naturalistic settings, and real-time processing capabilities
for the detection and prediction of stress and stress-related MDs.

Keywords: Machine Learning; Deep Learning; Data Preprocessing; Stress Detection; Stress
Prediction; Stress Monitoring; Mental Disorders

Introduction
Mental health has become a public health concern. According to Institute of Health Metrics and
Evaluation (IHME), in 2019, about 53 million people in the United States and about one in eight
individuals worldwide (about 1 billion people) suffer from at least one mental health disorder
(MD) [1]. MD is defined as an impairment in a person's cognition, emotional control, or behavior
pattern, which has clinical significance and is often linked to distress or functional impairment [2].
MDs severely limit people’s daily functioning and can be fatal [3], [4]. In 2019, mental health
(MH) problems accounted for 6.6% of all disability-adjusted life years in the US, making it the
fifth most significant cause of disability overall [5], [6].

Some of the more prevalent MDs are anxiety disorders, depression or mood disorders, bipolar
disorders, psychotic disorders (including schizophrenia), eating disorders, social disorders and
disruptive behavior and addictive behaviors [2]. In 2019, anxiety and depression have been the
most prevalent forms of MDs (301 and 280 million people affected worldwide, respectively).
Anxiety disorder encompasses emotions of concern, anxiety, excessive fear, or associated
behavioral problems that are severe enough to affect everyday activities [2]. Symptoms include an
unproportionate level of stress compared to the significance of the triggering event, difficulty in
putting worries out of one's mind, and nervousness [7], [8]. Generalized anxiety disorder, panic
attacks, social anxiety disorder, and post-traumatic stress disorder are all examples of different
types of anxiety disorders [2], [9]. Depression is characterized by a long-lasting sadness and a lack
of desire to be active. One of the main symptoms of depression is the inability to enjoy or find
pleasure in most of one's daily activities as well as felling sadness, anger, or emptiness [2], [10].
A depressive episode typically lasts for at least two weeks. Additionally, a loss of self-worth,
feelings of hopelessness for the future and suicidal thoughts are indicators and symptoms of
depression. People who are depressed are more prone to commit suicide [2], [10], [11].

Stress is categorized into distress, which typically has chronic negative effects on health, and
eustress, which is short-term and positively influences motivation and development [12].
Throughout this paper, the term stress is specifically used to denote distress, rather than eustress.
Mental stress has shown to significantly contribute to developing and worsening anxiety and
depression disorders [13], [14], [15]. Mental stress is the body’s natural response to various events
in which a person feels that the demands of their external environment exceed their psychological
and physiological resources for dealing with those demands [16]. Mental stress leads to an
asynchrony between the sympathetic and parasympathetic nervous systems (SNS and PNS) which
are the main divisions of the autonomic nervous system (ANS) [17] and serve an important role in
regulating vital biological activities [18], [19]. The sympathetic nervous system is an integrative
system that responds to potentially dangerous circumstances. Activation of the sympathetic
nervous system is part of the system responsible for controlling ‘fight-or-flight’ responses. The
parasympathetic nervous system is responsible for the body's "rest and digest" processes.

Given the import role and impact of stress in MDs, previous research has investigated various
qualitative and quantitative methods to measure and monitor stress to inform effective stress
mitigation approaches. While majority of stress literature relies on self-reported measures, recent
literature has used physiological variables such as heart rate, heart rate variability [20], [21], [22],
[23], [24], and behavioral data (e.g., speech, movement, facial expressions) [25] to understand
changes to SNS and PNS associated with stress. The recent advances in sensor and mobile health
technologies has resulted in the emergence of “big data” related to mental health as well as
advanced bioinformatics methods, tools, or techniques to use such data for modeling or inference.
One such tool that has recently emerged as a robust, rapid, objective, reliable, and cost-efficient
technique for studying chronic illnesses and MDs is Machine Learning (ML). ML uses advanced
statistical and probabilistic techniques to construct systems that can automatically learn from data.
Several characteristics of ML makes it suitable for applications in MH monitoring including
significant pattern recognition and forecasting capabilities [26], capacity to extract crucial
information from various data resources and opportunity to create personalized experiences [26],
and ability to analyze large amounts of data in a short time [27]. As such, ML has gained popularity
and has been applied to MH data to enable detection, monitoring, and treatment [28]. The objective
of this research is to review the literature to summarize and synthesize the application of ML in
the detection, monitoring, or prediction of stress and stress-related MDs, in particular anxiety and
depression. This paper documents methods-specific findings such as data types, preprocessing
methods, and different algorithms used as well as type and characteristics of studies that used ML.
Traditional statistical methods, such as linear regression, logistic regression, t-tests, and ANOVA
[29], have been widely employed in the past to detect and analyze stress and stress-related MDs.
These methods have proven useful in specific contexts, such as comparing means of different
groups, or modeling linear relationships between variables. As demonstrated by [22], [23], [24],
[25] and [26], these methods have provided valuable insights in situations where the data is
relatively simple and adheres to the underlying assumptions of the statistical techniques. However,
when faced with complex, high-dimensional mental health data, which has become increasingly
available thanks to advancements in technology and data collection techniques, these traditional
statistical methods might not be sufficient. The limitations of these methods stem from their
inherent simplicity and the assumptions they rely on, which might not hold true in the context of
MH data. For example, linear and logistic regression assume linear relationships between
variables, while t-tests and ANOVA require specific assumptions about the data distribution.
These assumptions may not be applicable in the case of intricate and heterogeneous MH data,
potentially leading to inaccurate or incomplete conclusions.

Advanced data analytics methods, such as machine learning (ML), offer a more powerful and
flexible alternative to traditional statistical methods. ML algorithms, with their significant pattern
recognition and forecasting capabilities [26], are capable of capturing complex, nonlinear
relationships between variables and can adapt to various data distributions. These capabilities
enable ML techniques to provide more accurate and insightful predictions, classifications, and
associations in the context of MH data [34]. Additionally, ML algorithms can handle large-scale,
high-dimensional data more efficiently than traditional methods, allowing researchers to analyze
vast amounts of information from diverse sources, such as electronic health records, wearable
devices, and online platforms [27]. This capacity for handling big data is crucial for understanding
the multifaceted nature of mental health disorders and developing tailored interventions. ML
techniques also offer the advantage of automation and adaptability, allowing them to continuously
learn and improve as new data becomes available [26]. This iterative learning process enables the
development of more sophisticated and accurate models for detecting, monitoring, and predicting
stress and stress-related MDs over time.

While traditional statistical methods have contributed significantly to our understanding of stress
and stress related MDs in specific contexts, the growing complexity and volume of MH data
necessitate the adoption of advanced data analytics methods like ML. By leveraging the power of
ML, researchers can gain deeper insights into the underlying patterns and relationships between
stress and MDs [34], ultimately leading to the development of more effective stress mitigation
approaches and improved care for individuals suffering from anxiety, depression, and other MDs.

Acknowledging the substantial contributions of traditional statistical methods, it becomes evident


that the escalating complexity and scale of mental health data demands the adoption of more
sophisticated approaches such as ML. This advancement stands not as a replacement but as an
essential evolution in the analytical toolbox available to researchers. As this paper delves into the
myriad ways ML has been applied to mental health, particularly in the realms of stress, anxiety,
and depression, it seeks to consolidate the current knowledge on the subject. By examining the
types of data, preprocessing methods, and the algorithms used in existing studies, this review
aspires to offer a detailed synthesis of the field. It aims to provide a clearer understanding of ML's
effectiveness in the detection, monitoring, and prediction of mental health disorders, setting a
foundation for future research and the enhancement of therapeutic strategies for those impacted by
these conditions.

Methods

Protocol and Registration


This scoping review adhered to the Preferred Reporting Items for Systematic reviews and Meta-
Analyses extension for Scoping Reviews (PRISMA-ScR) guidelines [35]. No formal review
protocol was registered due to the exploratory nature of this study, which aimed to map out existing
research rather than address a prespecified hypothesis. This approach aligns with the
methodological flexibility often required in emergent areas of research.
Eligibility Criteria
We included studies published in English from 2017 to 2022 that utilized machine learning (ML)
techniques to evaluate mental health disorders, specifically focusing on stress and stress-related
conditions. Studies were excluded if they did not use ML as the primary analysis method or if they
were published in languages other than English.
Information Sources
The literature search involved databases such as EI Engineering Village, Web of Science, ACM
Digital Library, and IEEE Xplore. Additional sources were identified through contact with experts
and review of references in relevant articles.

Search Strategy
A comprehensive search was conducted using a combination of keywords related to ML and
mental health disorders (Table 1). The search strategy was designed to capture a broad spectrum
of ML applications within this field. The full search list from all databases is available in the
Multimedia Appendix.

Table 1. Keywords and search strategy for articles since 2017 (last 5 years)
First keyword Second keyword Third keyword
predict OR mental health OR machine learning
detect mental disorder OR OR
depression OR deep learning OR
anxiety OR data mining OR
AND AND
stress pattern
classification OR
artificial
intelligence OR
neural networks

Study Selection, Inclusion, and Exclusion Criteria


Articles that did not fully use ML for stress or stress related MDs evaluations were excluded from
the research. Studies published in languages other than English were also excluded. The initial
search yielded 1241 results. After duplicate articles were deleted and eligibility was confirmed
using Rayyan QCRI [36], 1204 articles remained. After applying the exclusion criteria, 98 papers
were selected for full review (Figure 1).
Identification
Peer-reviewed published
records identified through
the initial search (n=1241)
Duplicated records
removed (n=37)

Records screened (title,


Screening

abstract and keyword Records excluded (n=719)


screening) (n=1204)

Excluded articles:
Eligibility

Full-text articles assessed • Not related to


for eligibility (n=485) stress/anxiety/depression
(n=297)
• Not related to ML/DL
(n=76)
• Similar or duplicate
content(n=11)
Included

Studies included in the • Non-English (n=3)


synthesis (n=98)

Figure 1. Preferred items for scoping literature review and meta-analysis flowchart [35]

Data Charting Process


Data charting was conducted by two reviewers independently using a standardized form, which
had been pretested on a subset of included studies. Discrepancies were resolved through discussion
or consultation with a third reviewer. Study authors were contacted for clarification or additional
data where necessary.

Data Items
Data extracted included publication year, study design, population characteristics, ML techniques
used, outcomes measured, and key findings. Other variables sought included data preprocessing
methods and performance metrics of the ML models. Simplifying assumptions, such as
considering different ML algorithms within the same family as a single technique, were made to
facilitate synthesis.

Synthesis of Results
Data were synthesized descriptively, grouping findings by ML techniques, data type and
preprocessing techniques. Where possible, quantitative performance metrics were extracted or
derived. Results were analyzed in the context of the overall study designs and populations to
highlight trends and identify gaps in the current research landscape. No formal critical appraisal
or quantitative meta-analysis was conducted due to the diversity of the included studies and the
scoping nature of this review.

Results and Discussion


In this section types of data, preprosessing techniques, and ML techniques used on the data in the
literature have been reviewed, and compared with the existing literature.
Types of Data
Various data types were used in the studies that used ML algorithms for stress and stress-related
MDs. Studies used questionnaire (n=31), heart rate variability (HRV, n=25), skin response (e.g.,
skin temperature, skin conductance, etc., n=24), photoplethysmogram (PPG, n=21),
electrocardiogram (ECG, n=19), heart rate (HR, n=17), electroencephalogram (EEG, n=9),
acceleration/body movement (n=8), text data (n=7), respiratory signals (n=7), electromyogram
(EMG, n=3), eye-tracking (n=3), speech signals (n=3) and others (n=4) including audio signals
(n=2), blood pressure (BP, n=1), and hormones (n=1). Figure 2 shows the distribution of the type
of data used for stress detection using ML techniques.

Distribution of Articles by Type of Data


31
25 24

17

9 8 7 7
3 3 3

Figure 2. Number of articles by types of data

Heart Measures
Heart metrics are primarily utilized for stress detection and are typically gathered through two
main methods: ECG (Electrocardiography) and PPG (Photoplethysmography). ECG is a non-
invasive diagnostic test that records the heart's electrical activity, while PPG is a non-invasive
optical technique that detects changes in blood volume within the tissue's microvascular bed. By
employing these methods, it is possible to measure various heart-related parameters, including
heart rate (HR), as well as time and frequency domain features of heart rate variability (HRV), and
blood pressure (BP).

• Heart Rate Variability (HRV) (n=25): Heart Rate Variability (HRV) has been used to assess
mental health issues, such as stress, anxiety, and depression, due to its rich time and
frequency domain features [37]. The Blood Volume Pulse (BVP) signal is another effective
method for capturing HRV features, as it represents the heart's beat-to-beat volume changes.
From the BVP signal, time domain measures like the Root Mean Square of Successive RR
Interval Differences (RMSSD), Standard Deviation of NN intervals (SDNN), and Standard
Deviation of RR intervals (SDRR) can be derived. Additionally, the frequency domain
aspects of HRV, including Total Power (TP, frequencies below 0.4 Hz), Low Frequencies
(LF, ranging from 0.04 Hz to 0.15 Hz), and High Frequencies (HF, between 0.15 Hz and 0.4
Hz), reflect the autonomic nervous system's dynamics during beat-to-beat measurements of
the heart rate (Figure 3) [38], [39]. These HRV measures, both in the time and frequency
domains, provide a nuanced view of the physiological underpinnings associated with various
mental health conditions.

• Heart Rate (HR) (n=17): One of the most important indicators of stress is an abrupt increase
in HR. Among the physiological signals, HR is among the top measures that explains stress
in ML models and it has been used in different studies with almost all ML algorithms [40],
[41], [42].

• Blood Pressure (BP) (n=1): BP can be obtained by pulse transit time (PTT) or by pressure
cuffs [43]. Stressful conditions create an influx of hormones that increase HR and constrict
blood vessels leading to a temporary BP elevation [44]. In most cases, BP recovers to its pre-
stress level after the stress response diminishes [45]. Schultebraucks et al. used systolic BP
as one of the measures in predicting one’s level of susceptibility to Post-Traumatic Stress
Disorder (PTSD) [46].

(a)
(b)

Figure 3. (a) Depiction of heart’s beat-to-beat measurements using Blood Volume Pulse (BVP) signal
(b) Power Spectral Density (PSD) of RR intervals (the signal is bandpass filtered with cut-off frequencies
of 0.04 Hz and 0.4 Hz)

Electroencephalogram (EEG) (n=9): EEG detects brain electrical activity. Compared to other
brain mapping techniques, for stress detection, it is more practical due to several factors including
affordability, non-invasiveness, non-intrusiveness and most importantly its high temporal
resolution [47]. The high temporal resolution of EEG makes it appropriate for real-time stress
detection, as well as DL approaches which require large dataset for training [47], [48], [49], [50],
[51].
The most commonly used EEG features for detection of stress are power of different frequency
bands, Alpha (8-13 Hz), Beta (12.5-30 Hz), Theta (4-7.5 Hz), Gamma (30-40 Hz), average and
standard deviation of a specific time window of EEG signal, and time-frequency features obtained
by Discrete Wavelet Transform (DWT) algorithm [51], [52], [53]. It has also been shown that
statistical features of EEG signal such as Kurtosis and Entropy are useful features in stress
prediction using ML algorithms [50]. Moreover, Power Spectral Density (PSD), correlation (C),
divisional asymmetry (DASM), rational asymmetry (RASM), and power spectrum (PS) are other
EEG features that have been used in different studies for stress detection [54].

Since EEG signals are collected from the scalp, they include excessive noise and so they have high
uncertainty. Therefore, signal processing and feature selection/extraction is a very important step
while dealing with EEG data. Several well-developed methods are available for treating the EEG
data. Among them, using latent space derived from auto-encoders and signal reconstruction
techniques such as Artifact Subspace Reconstruction (ASR) are well-known methods that can be
applied on EEG data to significantly reduce the artifacts [49]. These methods are also fast enough
that can make online detection feasible.

Amygdala and hippocampus are the parts of the brain that have the major responsibility for human
reactions to stress [55]. Brain activity caused by stress in those regions would affect the prefrontal
cortex. Studies collecting data from prefrontal cortex have also verified that EEG data from this
brain region can be used for stress detection [56]. EEG can be collected from the prefrontal cortex
using off-the-shelf EEG recording products such as MUSE and Neurosky Mindwave [50], [53],
[54], [56].

Eye Tracking (n=3): Eye-tracking features can be indicators of stress. For example, to diagnose
the level of stress, the changes in the striations of muscle material in the iris as a response to stress
can be used as features for ML algorithms. In other words, pupil diameter, which would be
controlled by iris sphincter muscles can be used as a feature [57]. Other eye-tracking features that
have been for stress detection are visual fixations, saccade movements, pupil size, micro saccades
and number of eye-blinks in specific time window during a certain task [58], [59], [60].

Skin Response (n=24): A skin response can be defined as a stimulus-regulated electrodermal


response and is typically measured using electrodes placed on the fingertips or hands. Skin
response is usually associated with increase in sympathetic activity upon inducing stress events
[61]. The skin becomes a better conductor of electricity when it is stimulated either externally or
internally by physiologically stimulating factors, including stressful conditions [62].

Respiratory Signals (n=7): Mental stress can affect different respiratory cycle phases and
breathing patterns [63], [64]. For example, It is discovered that stress had no impact on overall
breath duration (respiration rate), but that exhalation periods were longer and pause periods were
shorter in the stress experiment compared to the neutral condition [65].
Based on the findings of several studies, it can be concluded that respiratory signal is one of the
top contributing factors in explanation of stress in ML models. The most common time domain
respiratory signal features that are extracted for stress detection are: Root Mean Square (RMS),
Interquartile range (IQR), Mean of squared Differences between Adjacent elements (MDA) of
breathing rate and blood oxygenation levels. The most commonly used frequency domain features
of the respiratory signal are the power of low frequencies (LF, under 2 Hz), the power of high
frequencies (HF, above 2 Hz) and the ratio of power of low frequencies over the power of high
frequencies (LF/HF) [42], [46], [47], [66], [67], [68].

Electromyogram (EMG) (n=3): EMG detects the electrical activity of muscles at rest, during a
modest contraction, and during a strong contraction [69]. Similar to acceleration data, several
studies have shown that, using EMG data can help increasing the performance of ML models
trained on ECG data. The action potential intrigued in the EMG during stress can reduce the
variance for decision making of classification models that use ECG [42], [70], [71].

Hormones (n=1): It has been shown that stress can alter the levels of glucocorticoids,
catecholamines, growth hormones, and prolactin in the bloodstream. Therefore in ML models,
level of hormones such as cortisol, dehydroepiandrosterone sulfate (DHEAS), thyroid-stimulating
hormone (TSH), free triiodothyronine (FT3), and free thyroxine (FT4) can be used as predictors
for detection of stress-related disorders [46].

Acceleration/Body Movement (n=8): Mental Stress may cause a broad variety of behavioral/body
movement symptoms such as shaking hands and feet which can be measured by the acceleration
data [72]. Moreover, research has shown that people with a greater stress score had less variance
in their activity level and body movements [73], [74], [75]. For example, In the elderly, stressful
life events can be related to a reduced rate of regular physical exercise [76]. Time and frequency
features such as mean absolute deviation from mean (MAD), total power of acceleration, standard
deviation, mean norm of acceleration, absolute integral, peak frequency of each axis are the
features of hand/body acceleration used for stress detection [41], [77], [78]. One practical
characteristic of motion/acceleration data would be the fact that it can be used to identify” sources
of noise in other signals . For example, motion data can help distinguishing stress from physical
activity (e.g., exercise) when other physiological measures such as ECG have uncertainty in
prediction [79], [80].

Audio and Speech Signals


• Speech Signals (n=3): Using speech signals, it is feasible to diagnose and assess neurological
and MDs [81]. Moreover, studies have shown that, like body acceleration and EMG, features
of speech signal can make stress predictions of heart measurements more robust. The best
explanatory parameters of speech signal are frequency domain parameters (e.g., PSD,
strongest frequency from FFT transform) and time-frequency features such as Mel-
Frequency Cepstral Coefficient (MFCC) [40], [82], [83]. Since time-frequency measures are
2-dimentional measurements with high number of samples, they make this signal suitable for
using in convolutional neural network models (CNNs) of stress and depression detection
[84].

• Audio Signals (n=2): For lab based studies, audio signals (e.g. beep sounds) can be used for
stimulating stress events in participants [85], [86].

Text data (n=7): Social media content is frequently subjected to reviews, opinions, and influence,
as well as sentiment analysis. Natural language processing methods may be used to evaluate social
networking posts and comments for mood and emotion to detect whether a user is stressed [87],
[88], [89], [90], [91], [91], [92], [93].

Questionnaire (n=31): There are different questionnaires that are used for diagnosis of stress and
different MDs including anxiety and depression. The scores from different items on these
questionnaires can be used as dependent/independent variables in ML studies. The questionnaires
mentioned here were selected based on their prevalence in the literature as well as their relevance
to the ML outcomes being predicted. For instance, some studies have successfully leveraged scores
from multiple questionnaires, such as the Diagnostic and Statistical Manual of Mental Disorders
(DSM), Depression Anxiety and Stress Scale (DASS), Edinburgh Perinatal/Postnatal Depression
Scale (EPDS) Center for Epidemiological Studies-Depression (CES-D) survey, Mean Opinion
Score (MOS), Hamilton Depression Rating Scale (HAM-D), State-Trait Anxiety Inventory
(STAI), Posttraumatic Stress Disorder Checklist for DSM (PCL), Beck Depression Inventory
(BDI, Beck Anxiety Inventory (BAI), Hospital Anxiety and Depression Scale (HADS),
Goldberg’s Depression Scale (GDS), self-reports and clinician reports, [94], [95], [96], [97], [98],
[99], [100], [101], [102], [103], [104], [105], [106], [107], [108], [109], [110], [111].

Preprocessing Techniques
In this section, important preprocessing techniques that have yielded significant findings and how
they are used to help the detection of stress and its related MDs have been reviewed.
Synthetic Minority Oversampling Technique (SMOTE) (n=3): In detection of stress and its related
MDs, usually the number of samples for the stress or MD class is significantly lower than the non-
stress or non-MD class. This imbalance in the number of samples for each class leads to a bias in
prediction (towards the majority class). To correct for data bias, it is possible to oversample the
underrepresented group. In stress detection studies using ML models, SMOTE is one of the most
common approaches to boost the minority class using, which creates new samples by synthesizing
those already available in the data (by combining their features) [77], [95], [112].

Early Modality Fusion (n=1): In ML models used for prediction of stress with a multimodal
approach, it has been shown that early fusion of multimodal data before feature extraction is more
effective and archives a better performance. This is due to the fact that early modality fusion
catches better the important characteristics that are in coherence with each other. For example a
study showed that combining different measures including skin response, skin temperature and
body acceleration before feature extraction outperforms the approach that extracts the features for
each measure separately and combines them afterwards (Figure 4) [113].

Feature Extraction
Fusion

Sensor Stress
Preprocessing (e.g., using
Data Prediction
Autoencoder)

(a)

Feature Extraction
Fusion
Sensor Stress
Preprocessing (e.g., using
Data Prediction
Autoencoder)

(b)

Figure 4. (a) Early Modality Fusion (b) Late Modality Fusion

Power Spectral Density (PSD) (n=13): In physiological signals for stress detection, usually power
of the signal changes during the moments of stress. PSD explains the frequency-based power
distribution of a time series and reveals the locations of strong and weak frequency variation.
Welch’s method is one of the most common approaches to calculate PSD [49]. PSD is often used
in the studies that include frequency domain HRV features for stress detection such as total HF or
LF power [66], [67], [68], [86], [114], [114], [115], [116], [117], [118], [119], [120], [121].

ILIOU (n=1): In detection of MDs such as depression and anxiety using machine learning
techniques, having the least error rate is significantly important so that the person can take further
actions appropriately. In this matter, data preprocessing step has an important role to minimize the
noise and bias towards the false prediction. Iliou et al. proposed ILIOU, a data mapping and
transformation method, that identifies useful information for detection of MDs, especially for
depression. This method outperforms common data preprocessing techniques such as Principal
component analysis (PCA), Evolutionary Search Algorithm (ESA) and Isomap for detection of
depression [99].

Principal Component Analysis (PCA) (n=3): Principal component analysis (PCA) is a method for
lowering the dimensionality of such datasets while maximizing interpretability and minimizing
loss of information. It does this by generating new variables that are uncorrelated and progressively
optimize variance [42], [108], [122].

Independent component analysis (ICA) (n=4): Independent component analysis (ICA) is a


computational and statistical method for uncovering hidden elements underlying random
variables, observations, or signals. This method is mostly used for removing artifacts from
stationary signal noises of the multi-channel data. ICA optimizes higher-order statistics such as
kurtosis, while PCA optimizes the covariance matrix of the data, which reflects second-order
statistics. In stress detection using physiological signals that contain stationary noises (e.g.
eyeblink noise in EEG) it is recommended to remove noises using ICA [47], [48], [49], [51].

Artifact subspace reconstruction (ASR) (n=1): ASR is an adaptive approach for removing artifacts
from signal recordings online or offline, mostly non-stationary signal noises. To identify artifacts
based on their statistical qualities in the component subspace, it repeatedly computes a PCA on
covariance matrices [123]. Since there are usually lots of non-stationary noises in the EEG data,
in order to classify stress in multiple levels using EEG data, using ASR before classification is
highly recommended [49].

Latent Growth Mixture Modeling (LGMM) (n=1): Growth mixture modeling (GMM) is to
discover numerous hidden subpopulations, describe longitudinal development within each hidden
subpopulation, and investigate variation in hidden subpopulations' rates of change. Latent growth
mixture models are gaining popularity as a statistical tool for estimating individual development
over time and for probing the presence of latent trajectories, in which people belong to trajectories
that are not directly observable [46], [124], [125].

Dynamic Time Warping (DTW) (n=1): It is common practice to transform data from two time
series into vectors and then compute the Euclidean distance between the resulting points in vector
space to determine the degree of similarity or dissimilarity between the series, regardless of if they
vary in time or velocity. DTW method can be applied to find such similarities that may exist
between people in terms of their mood series. As an example, one may compare time-series to find
whether they match for stress, depression, or anxiety. Moreover, it can be utilized to forecast the
mental condition of persons with substantially comparable series patterns [115], [126].The
difference between DTW and Euclidian matching is that unlike Euclidean matching, DTW
considers distance of each point in one sequence, to every point in the other sequence to determine
the similarity between them (Figure 5).

Figure 5. Dynamic Time Warping Vs. Euclidian Matching [127]

Kalman Filter (n=2): The Kalman filter is a technique for making predictions about unknown
variables (e.g., missing data) based on observable data. Kalman filters include two iterative steps—
predict and update—that are used to estimate states using linear dynamical systems in state space
format. Iterative cycles of predict and update are performed until convergence is achieved [128].
Kalman filter has been used to handle the missing data for stress detection in some studies [129],
[130].

Autoencoders (n=3): Autoencoders are a type of Neural Networks that learn a representation of
the data in lower dimensions than the original data (encoding) by regenerating the input from the
encodings (decoding). For data with very high dimensionality, usually clustering is not optimized
because of the noise present in the original data. Hence, it is an appropriate practice to use the
encoded representation of the data, obtained by autoencoders, to have lower and more optimized
dimensions for clustering [49], [93], [131].

Self-Organizing Map (SOM) (n=3): In ML, a self-organizing map (SOM) produces a low-
dimensional – typically two-dimensional – representation of a high-dimensional dataset while
preserving its topology by creating clusters. It is therefore possible to visualize and analyze high-
dimensional data more easily (Figure 6) [92], [118], [132].

Figure 6. Representation of SOM before (left) and after mapping (right) [118].

Wrapper Feature Selection Methods: Wrapper methods try to use a subset of features while
training a model. Changes will be made to the feature subset based on the performance about the
prior model (Figure 7). Therefore, finding the best features using wrapper method is a search
problem. These methods often have high computing costs [133]. Some most common wrapper
methods are: Naïve search, Sequential Forward Feature Selection (SFFS), Sequential Backward
Feature Selection (SBFS), and Generalized Sequential Search (GSS) [134]. Some studies used this
approach as their feature selection technique [56], [59].
Finding the subset of features
with the best model performance

Train Machine Model with


All Features Pick a Subset Learning Model Best
of Features Performance
Algorithm Performance

Figure 7. Steps of a wrapper feature selection method

Filter Feature Selection Methods: In general, filter methods are used as a preprocessing step
without regard to any ML algorithms. Statistic tests are used instead to select features based on
their correlation with dependent variables (Figure 8). The filter feature selection methods used in
the literature are mentioned below.
Select Best Train Machine Model
All Features Features by one Learning Performance
time filtering Algorithm

Figure 8. Steps of a filter feature selection method

• Chi-square test (n=3): This test checks for independence between categorical features and
the target variable. Features with high Chi-square scores are selected, implying a strong
association with the target variable, which may be valuable for the model [40], [120], [135].

• Pearson Correlation (n=2): Pearson linear correlation coefficient is a way to quantify how
closely two sets of data are correlated linearly. It indicates how different measures are related
to each other by a number between -1 to 1. Therefore, among highly correlated variables
some them can be removed as they don’t add useful information to ML models [98], [136].

• Minimum Redundancy Maximum Relevance (mRMR) (n=2): mRMR technique chooses


characteristics having a high correlation to output (relevance) and a low correlation to one
another (redundancy). F-statistic is used to determine the correlation between features and
the output, whilst Pearson correlation coefficient (for non-time series features) and Dynamic
Time Warping (DTW for time series features) may be used to calculate the correlation
between features (Figure 9). The objective function, which is a function of relevance and
redundancy, is then maximized by selecting features one at a time using a greedy search.
Mutual Information Difference (MID) and Mutual Information Quotient (MIQ) criteria are
both frequently employed objective functions that depict the difference or quotient between
relevance and redundancy [137], [138]. Using this feature selection method, Giannakakis et
al. have ranked ECG measurements in the order of importance as mean HR , LF, NN50,
standard deviation of HR,pNN50, LF/HF, RMSSD, HF, and total power [115].

(a)

(b)
Figure 9. calculation of relevance and redundancy for (a) non-time series features (b) time-series
features
(DTW: Dynamic Time Warping)

Machine Learning (ML) Techniques


The ML algorithms used for stress and MD detection have been reviewed in this section. The
papers used DL approach or Neural Network (NN, n=39) Logistic Regression (LR, n=26) Naive
Bayes (NB, n=22), Decision Tree (DT, n=23), Boosting (e.g., Adaptive Boosting (AdaBoost),
extreme Gradient Boosting (XGBoost), etc., n=22), Random Forest (RF, n=36), Discriminant
Analysis (e.g., Linear Discriminant Analysis (LDA), Quadratic Discriminant Analysis (QDA),
n=6), Fuzzy C-means (n=2), K-nearest neighbors (KNN, n=22) and Support Vector Machines
(SVM, n=48). Figure 10 shows the distribution of articles by ML model. Refer to table A1 to find
which papers have used each ML technique.

Distribution of Articles by ML Model


60
48
50
39
40 36

30 26
22 23 22 22
19
20 15
8 6
10 2 4
0

Figure 10. Number of articles for each ML model

Logistic Regression (LR) (n=26): LR is a supervised parametric ML technique in which multiple


independent variables will be utilized to detect the occurrence of stress or normal condition [56],
[102]. Some studies utilized the numerical independent variables (e.g., HRV time-domain features:
RMSSD, HR, pNN50) [79], [139] or categorical data (e.g., answer to multiple choice questions)
obtained from questionnaires [92], [99], [100].

Naïve Bayes (NB) (n=22): Naïve Bayes algorithm is a supervised, generally parametric,
classification method that uses the Bayes Theorem as its foundation and has the naïve assumption
of predictor independence. In other words, Naïve Bayes classifier assumes that the existence of a
given independent variable to predict the dependent variable is independent of the presence of any
other independent variable that predicts the dependent variable.

Decision Tree (DT) (n=23): Decision Tree is a supervised non-parametric ML algorithm used in
classification and regression applications. It comprises a root node, branches, internal nodes, and
leaf nodes in a hierarchical, tree-like structure (Figure 11).
Decision Node Root Node

Decision Node Decision Node

Leaf Node Leaf Node Leaf Node Decision Node

Leaf Node Leaf Node

Figure 11. Structure of a Decision Tree

Boosting (n=22): Boosting is an ensemble learning for reducing training errors by combining a
group of weak learners. When using Boosting algorithm, models are fitted on random samples of
data, and then models are trained repeatedly in a sequence. When each model starts being trained
in that sequence, it attempts to make up for the flaws of the one that came before it. The most
commonly used Boosting algorithms are: Adaptive boosting (AdaBoost), Gradient boosting, and
Extreme gradient boosting (XGBoost).

Random Forest (RF) (n=36): Random Forest is a supervised non-parametric ensemble learning
algorithm that uses many Decision Trees built during the training process. Random Forest
algorithm is used for both classification and regression problems. When it comes to classification,
the Random Forest’s output is the class that the majority of the Decision Trees choose. For
regression purposes, an individual tree's predicted mean or average is returned as the output. Using
Random Forests, we can overcome the tendency of decision trees to overfit to their training data.
Discriminant Analysis (n=6): Discriminant Analysis is a supervised parametric classification
algorithm that works with data including a dependent variable and independent variables and
mostly used to classify the observation into a certain group based on the independent variables in
the data. Linear Discriminant Analysis (LDA) and Quadratic Discriminant Analysis (QDA) are
the two forms of Discriminant Analysis.

K-nearest neighbors (K-NN) (n=22): K-nearest neighbors (K-NN) is a non-parametric supervised


ML algorithm that is used for both classification and regression purposes. In classification, the
algorithm determines the label of a new sample not available in the training data by assigning the
label of the majority of k-nearest training data points to that new sample (Figure 12). In regression,
the output for each sample, is the average of the values of k-nearest neighbors to that sample (not
including the sample itself). In this literature K-NN has been only used for classification.
Class A

Class B

Class C

K=7
Figure 12. Example of K-NN classification with K = 7. In this example, the label of “Class C”
is assigned to the new (black) datapoint since the majority of the 7-nearest datapoints to the new
datapoint are from “Class C”.

Support Vector Machines (SVM) (n=48): Support Vector Machine is a parametric supervised ML
algorithm used for both classification and regression problems. It can solve both linear and non-
linear problems using non-linear kernels. For classification, the SVM algorithm finds a line (or a
hyperplane for non-linear kernels) between each pair of classes of the training data in a way that
the margin distance of that line or hyperplane to the closest point of each of those two classes is
maximized (Figure 13). This is repeated for all pairs of classes in the dataset. Then the obtained
lines are used as boundaries for classes. In regression, the SVM tries to find the line/hyperplane
that within a very small margin of 𝜀 (epsilon) has maximum number of datapoints. That
line/hyperplane used for regression.

Support Vectors
Class A

Class B

Ma
xim
ize Line/Hyperplane
dM
arg
in

Figure 13. Visual representation of Support Vector Machine algorithm

K-means clustering (n=4): K-means clustering is an unsupervised ML algorithm that aims to


arrange objects into groups based on their similarity. To find those similarities, it calculates the
distance of data points into K random cluster centroids and assigns each data point to its closest
centroid. Then location of each centroid is then updated by average value of all datapoints
associated with that centroid. This process is repeated until there is no change in the location of
the centroids. In ML models for stress detection, K-means clustering has been used in the literature
for personalization of the ML models [42], [51], and for labeling the dataset [131], [140].

Neural Network (NN) (n=39): DL methods are a subset of ML methods that. NNs are the heart of
the DL algorithms. The neural network is a method for implementing ML that utilizes
interconnected nodes or neurons arranged in a layered structure resembling the human brain. There
are different types of NNs have been explained below:

• Artificial Neural Network (ANN): It is possible to think of a single perceptron (or neuron)
as an abstract Logistic Regression. In each layer of ANNs, a group of multiple perceptron
or artificial neurons is used. Figure 14 shows an ANN with one layer and its working
mechanism.
(a)

(b)

(1) (2)
Figure 14. (a) Representation of an ANN with one hidden layer. 𝑊𝑖𝑗 and 𝑊𝑖𝑗 denote the weights of the
links connecting the first layer (input layer) to the hidden layer and weights of the links connecting the
second layer to the next layer (output layer), respectively. (b) Representation of how a single neuron works.
First, all the outputs of the previous layer are multiplied by the weights associated with the links connecting
them to the 𝑗𝑡ℎ neuron of the next layer and summed by a bias (summation and bias step). The result is
then passed through an activation function (activation step).

• Convolutional Neural Network (CNN): CNNs are a form of neural network that are
especially adept at handling data structures with a grid-like layout, such as images/objects.
Classification and computer vision applications are common uses for convolutional neural
networks (ConvNets or CNNs) (Figure 15).
Feature Maps
Stress
Feature Maps
Feature Maps
Output

Normal

Physiological Signal
Preprocessing

Convolutions Subsampling Convolutions Subsampling Fully Connected

Figure 15. Representation of CNN for physiological signal

• Recurrent Neural Network (RNN): An RNN is a subset of artificial neural networks


designed specifically for use with time series data and other sequence-based data. Long
Short-Term Memory (LSTM) networks are the most common type of RNNs. In RNNs,
Attention mechanism is a method that simulates cognitive attention in neural networks.
The purpose of the impact is to encourage the network to give greater attention to the small
but significant portions of the input data by enhancing some and reducing others. Since
stress may alter a small portion of physiological data (e.g. ECG), attention mechanism can
be used to detect stress using RNNs when large datasets are available [141].
Cong et al. introduced X-A-BiLSTM, which is a DL model that includes XGBoost (to filter
data and handle imbalanced data) and Attention Bi-LSTM (LSTM with forward and
backward memory and Attention mechanism) Neural Network used for stress classification
using text data [87].

Other ML techniques (n=19):


• Voting ensemble classifier: The classification is decided based on weighted voting, which
is determined by using a voting ensemble approach. The voting classifier allows for voting
in which the final class labels are determined either by the class chosen most frequently by
the classification models, or by the average of the output probabilities from each
classification model. In the literature, this method has been utilized for PTSD detection
[112], stress and stress related MDs [68], [80], [103], [140], [142].

• Fuzzy C-means (FCM) clustering: Fuzzy C-means clustering (FCM) is a clustering


approach that assigns every data point to all the clusters with a certain probability instead
of assigning each point to only one cluster. A data point that is near to the cluster's center,
for instance, will have a high degree of membership there, while a data point that is distant
from the cluster's center would have a low degree of membership [143]. Since depression,
anxiety are not discrete measures, some studies have used FCM as an alternative to other
clustering techniques for detection of these MDs [99], [101].

In this article, the recent ML algorithms, preprocessing techniques, and data (e.g., physiological
data, questionnaire data, etc.) used in detection, prediction and monitoring of stress and the most
common MDs (i.e., depression, anxiety, other stress-related MDs) have been reviewed.
Based on this review, it is concluded that among classic ML algorithms (excluding DL
approaches), supervised models of Support Vector Machines (SVMs) and Random Forest (RF),
have been used more often and achieved better performance in terms of model accuracy and
robustness (measured by parameters like Area Under the Receiver Operating Characteristic curve
(AUROC)). The accuracy of ML models is a critical indicator of their utility in real-world
applications. The review demonstrates that SVM consistently achieves high accuracy across
various data types, including HR, HRV, and skin response. For instance, SVM achieved 93%
accuracy with HR, PPG, and skin response data in study [34], and 96% with skin response data in
study [140]. These results underscore SVM's robustness in handling complex, non-linear data.
Random Forest also shows commendable performance, with an accuracy of 99.88% in study [144],
reflecting its strength in ensemble learning to mitigate overfitting and noise.
Moreover, among the predicting measures for stress and stress-related MDs, HR, HRV and skin
response have been used the most often (Figure 16). These measures were the major explaining
factors in the ML algorithms to predict stress and stress-related MDs. It is noticeable that DL
approaches are becoming more popular as these techniques provide unique specifications that
classic ML algorithms cannot provide.

Since stress is a time dependent event, the relationship between different lags of time can be
important for detection of stress. Recurrent Neural Networks (RNNs) and Convolutional Neural
Networks (CNNs) will take into account the relationship between datapoints in different time-
series for their decision making and they have the potential to enhance the detections. Deep
learning models, specifically CNNs and LSTMs, show promising results, with CNNs achieving
92.8% accuracy in HRV and ECG data in study [139], indicating their potential in feature-rich
physiological data. However, it is worth noting that deep learning models require substantial data
for training, which may limit their applicability in studies with smaller datasets. Attention
mechanism in RNNs is a new technique that is becoming popular for finding animalities in
physiological signals. However, based on the review of literature, this mechanism has only been
used on text data (not on physiological signals) to detect stress. Therefore, Attention mechanism
is technology that can be further utilized for physiological signals to detect stress.
Unsupervised ML (and DL algorithms) such as clustering techniques have been used mostly for
the preprocessing step to label the data (if labels are not available) and also for finding a
representation of the data that achieves the best performance in detection algorithms.

For data preprocessing, feature selection (i.e., filter and wrapper methods) and extraction
techniques have been commonly used. In feature extraction approaches, latent representations of
data by transformations such as output of encoder in autoencoders have been useful to remove data
noises and to make the data more compact, making further computations more efficient. PCA and
ICA are other most common feature extraction approaches used in the literature.

Among the selected features, statistical indicators of heart measurements such as mean, standard
deviation of HR, along with time and frequency representations of HRV such as RMSSD and total
LF and HF power were most widely used. Heart measurements also have been more often than
other measurements as they are unobtrusive, non-invasive, affordable and easier to measure and
also describing a big portion of stress events. After those, skin response measure has been found
as one of the most important factors in detection of stress and its related disorders. The time-
frequency approaches to analyze time series data are getting more popular in this area as they are
proper representations of data for DL approaches which can be more accurate and robust. As an
example, for DL algorithms, RNNs with attention mechanisms can help to find portions of data
related to stress and its related disorders with higher confidence.

Most of the studies models do not interpret the ML models and look at them as black box. This
limits the contribution to the body of science. SHapley Additive exPlanations (SHAP) is a
technique used by some studies to interpret the models such as evaluation of features to find the
most important ones and also how in what direction each feature affects the predictions. SHAP
correlation plot provides insight into the distribution of the features themselves, as well as the
relationship between their influence on the model. In other words, it provides the importance of
each feature on prediction of the dependent variable by taking into account both the main effect as
and the interaction effect of that feature with other features in the data [46], [77], [105], [144],
[145], [146].

Despite progress in stress detection methodologies, the exploration of personalized models has
been limited. Most studies have not gone beyond basic normalization techniques, overlooking the
fact that physiological measures are as distinct to individuals as biometric identifiers. A notable
exception can be found in a select few studies [51], [113], [147], which have employed more
sophisticated personalization techniques, integrating complex data transformations to account for
individual variability.
Data Type Vs. ML Model
25

20

15

10

0
Heart Measures Skin Response Other Activity Sentiment Perceived Measures
Psychophysiological
Measures
SVM NN RF LR DT KNN NB Boosting LDA/QDA K-means Fuzzy

Figure 16. Distribution of ML models used for each type of data. In this figure, skin response and heart
measures (including HR, HRV and blood pressure) have been shown separately due to their high usage
and importance in the literature. Other psychophysiological measures include EEG, EMG, Eye-
tracking and respiratory signals. Activity includes body movement. Sentiment data includes speech and
text data. Finally, perceived measures include questionnaire and self-report data.

Strengths of the review

In undertaking this scoping review, we have embarked on a rich exploration of the applications of
machine learning (ML) in the field of stress detection, articulating a narrative that is both
comprehensive and detailed. The review lays out a landscape where diverse data types are not
merely cataloged but deeply analyzed for their roles and interconnections within the broader
context of methodological approaches. This provides a robust understanding of the field’s current
state and its complexities.
This review has documented a comprehensive assessment on various physiological measurement
techniques, including heart rate variability (HRV), electroencephalograms (EEG), and
electrocardiograms (ECG), etc. This assessment is not just a recounting of the types of data
employed in the literature but a thoughtful consideration of how each contributes to a multifaceted
understanding of stress indicators. It is an acknowledgment that the signals of stress are as complex
as the condition itself, necessitating a rich palette of investigative tools.
The review also examines a range of advanced preprocessing techniques such as mRMR, SOM,
SMOTE and PCA. This examination sheds light on how different studies leverage these methods
to refine the quality of data fed into ML models, thereby potentially enhancing the models'
accuracy and reliability in detecting stress. It is an illustration of how sophisticated data treatment
can lead to more nuanced insights, even if our own methodology did not directly employ these
techniques.
Limitations
Our scoping review acknowledges its inherent constraints, including a possible selection bias due
to potential omissions of pertinent studies. It serves as a contemporary cross-section of the rapidly
evolving domains of machine learning and mental health, underscoring the imperative for periodic
scholarly review to sustain its relevance and precision. While we survey a broad spectrum of
machine learning techniques applied to stress detection, we do not extensively assess their efficacy,
suggesting a fertile ground for future empirical investigations to assess these methods across
diverse data cohorts and settings. Additionally, while we address the preprocessing techniques and
their impact on model performance, our discussion does not delve into detailed technical analysis.
Finally, the crucial issue of model interpretability is touched upon but not explored in depth,
presenting an opportunity for further scholarly explorations.

Conclusions and Future Directions

The pivotal insights from this review underscore the potential of ML to redefine the approach to
mental health care, particularly in the diagnosis and management of stress-related conditions and
MDs. As we have discerned, there is an expansive field ripe for further exploration, with research
gaps suggesting a number of promising directions. Guided by these insights, we can now chart a
course for future research that not only expands the boundaries of our scientific understanding but
also translates into tangible improvements in clinical practice.

Real-time and Naturalistic ML Applications


The scarcity of real-time studies in naturalistic settings has highlighted the importance of
developing ML models that accurately reflect and respond to the complexities of real life. Future
research must prioritize the creation of algorithms capable of operating amidst the unpredictability
of daily life, providing immediate insights and adaptable interventions. These models hold the
potential to transform practice by offering tools that can preemptively identify stress and MD
symptoms, enabling clinicians to intervene before conditions worsen.

Temporal Data and Deep Learning


Our review illuminates the untapped potential of time series data in capturing the evolution of
stress and MDs. Deep learning techniques, specifically designed to interpret complex, sequential
data, could lead to breakthroughs in how we understand and predict mental health trajectories. For
practice, this means more sophisticated diagnostic tools that can provide a nuanced picture of a
patient's mental health over time, enabling personalized treatment plans that are responsive to the
patient’s changing condition.

Personalization in ML Models
The need for individualized care in mental health cannot be overstated. The heterogeneity of stress
responses and MD symptoms calls for personalized ML models tailored to individual
physiological and behavioral patterns. Future research should focus on leveraging multi-task
learning to refine algorithms that adapt to individual baselines, enhancing the personalization of
care. For clinicians, this means access to tools that can more accurately reflect and respond to the
unique needs of each patient, reducing the risk of misdiagnosis and improving treatment efficacy.
Predictive analytics can be instrumental in identifying key factors that contribute to misdiagnosis
and delayed help-seeking. Future studies should look to build on this knowledge to inform the
creation of interventions that encourage timely and accurate diagnosis. In practice, this could lead
to the development of targeted screening tools that assist clinicians in recognizing at-risk
individuals more effectively. The integration of clinical expertise with ML innovation is crucial
for the development of tools that are both advanced and clinically relevant. Collaboration between
healthcare professionals, patients, and AI developers will be essential in creating user-centered
tools that address real-world needs. This collaborative approach will likely result in the
development of AI applications that are more intuitive and effective in clinical settings.

Conflict of Interest
There is no conflict of interest.
References
[1] IHME, “GBD Results,” Institute for Health Metrics and Evaluation. Accessed: Mar. 31,
2023. [Online]. Available: [Link]
2019-permalink/b9a68deae92b3b8772ae33689f5c7fce
[2] WHO, “Mental disorders.” Accessed: Aug. 09, 2022. [Online]. Available:
[Link]
[3] U.S. Department of Health and Human Services, “Key Substance Use and Mental Health
Indicators in the United States: Results from the 2019 National Survey on Drug Use and
Health,” Security Research Hub Reports, Jan. 2020, [Online]. Available:
[Link]
[4] R. Mojtabai and M. Olfson, “National Trends in Mental Health Care for US Adolescents,”
JAMA Psychiatry, vol. 77, no. 7, pp. 703–714, Jul. 2020, doi:
10.1001/jamapsychiatry.2020.0279.
[5] Institute for Health Metrics and Evaluation, “GBD Results,” Institute for Health Metrics
and Evaluation. Accessed: Aug. 30, 2022. [Online]. Available:
[Link]
[6] National Institute of Mental Health, “Mental Illness,” National Institute of Mental Health
(NIMH). Accessed: Aug. 30, 2022. [Online]. Available:
[Link]
[7] Dan Brennan, “Anxiety Disorders: Types, Causes, Symptoms, Diagnosis, Treatment.”
Accessed: Aug. 09, 2022. [Online]. Available: [Link]
panic/guide/anxiety-disorders
[8] Mayo Clinic, “Anxiety disorders - Symptoms and causes,” Mayo Clinic. Accessed: Aug.
09, 2022. [Online]. Available: [Link]
conditions/anxiety/symptoms-causes/syc-20350961
[9] National Institute of Mental Health, “Anxiety Disorders,” National Institute of Mental
Health (NIMH). Accessed: Aug. 09, 2022. [Online]. Available:
[Link]
[10] National Institute of Mental Health, “Depression,” National Institute of Mental Health
(NIMH). Accessed: Aug. 09, 2022. [Online]. Available:
[Link]
[11] Mayo Clinic, “Depression (major depressive disorder) - Symptoms and causes,” Mayo
Clinic. Accessed: Aug. 09, 2022. [Online]. Available:
[Link]
20356007
[12] J. Bienertova-Vasku, P. Lenart, and M. Scheringer, “Eustress and Distress: Neither Good
Nor Bad, but Rather the Same?,” Bioessays, vol. 42, no. 7, p. e1900238, Jul. 2020, doi:
10.1002/bies.201900238.
[13] CAMH, “20131 Stress,” CAMH. Accessed: Jun. 18, 2022. [Online]. Available:
[Link]
[14] J. Cooper, “The Link Between Stress and Depression,” WebMD. Accessed: Aug. 21, 2022.
[Online]. Available: [Link]
[15] NIMH, “I’m So Stressed Out! Fact Sheet,” National Institute of Mental Health (NIMH).
Accessed: Aug. 21, 2022. [Online]. Available:
[Link]
[16] S. Cohen, P. J. Gianaros, and S. B. Manuck, “A Stage Model of Stress and Disease,”
Perspect Psychol Sci, vol. 11, no. 4, pp. 456–463, Jul. 2016, doi:
10.1177/1745691616646305.
[17] E. Won and Y.-K. Kim, “Stress, the Autonomic Nervous System, and the Immune-
kynurenine Pathway in the Etiology of Depression,” Curr Neuropharmacol, vol. 14, no. 7,
pp. 665–673, Oct. 2016, doi: 10.2174/1570159X14666151208113006.
[18] L. K. McCorry, “Physiology of the Autonomic Nervous System,” Am J Pharm Educ, vol.
71, no. 4, p. 78, Aug. 2007.
[19] G. Pongratz and R. H. Straub, “The sympathetic nervous response in inflammation,”
Arthritis Research & Therapy, vol. 16, no. 6, p. 504, Dec. 2014, doi: 10.1186/s13075-014-
0504-2.
[20] P. Schmidt, A. Reiss, R. Dürichen, and K. V. Laerhoven, “Wearable-Based Affect
Recognition—A Review,” Sensors, vol. 19, no. 19, Art. no. 19, Jan. 2019, doi:
10.3390/s19194079.
[21] H.-G. Kim, E.-J. Cheon, D.-S. Bai, Y. H. Lee, and B.-H. Koo, “Stress and Heart Rate
Variability: A Meta-Analysis and Review of the Literature,” Psychiatry Investig, vol. 15,
no. 3, pp. 235–245, Mar. 2018, doi: 10.30773/pi.2017.08.17.
[22] A. V. Machado et al., “Association between distinct coping styles and heart rate variability
changes to an acute psychosocial stress task,” Sci Rep, vol. 11, no. 1, Art. no. 1, Dec.
2021, doi: 10.1038/s41598-021-03386-6.
[23] T. Adjei, W. von Rosenberg, T. Nakamura, T. Chanwimalueang, and D. P. Mandic, “The
ClassA Framework: HRV Based Assessment of SNS and PNS Dynamics Without LF-HF
Controversies,” Frontiers in Physiology, vol. 10, 2019, Accessed: Feb. 23, 2024. [Online].
Available:
[Link]
[24] H. H. Yoo, S. J. Yune, S. J. Im, B. S. Kam, and S. Y. Lee, “Heart Rate Variability-
Measured Stress and Academic Achievement in Medical Students,” Med Princ Pract, vol.
30, no. 2, pp. 193–200, 2021, doi: 10.1159/000513781.
[25] J. Chen, M. Abbod, and J.-S. Shieh, “Pain and Stress Detection Using Wearable Sensors
and Devices—A Review,” Sensors, vol. 21, no. 4, Art. no. 4, Jan. 2021, doi:
10.3390/s21041030.
[26] M. I. Jordan and T. M. Mitchell, “Machine learning: Trends, perspectives, and prospects,”
Science, vol. 349, no. 6245, pp. 255–260, Jul. 2015, doi: 10.1126/science.aaa8415.
[27] J. Luo, M. Wu, D. Gopukumar, and Y. Zhao, “Big Data Application in Biomedical
Research and Health Care: A Literature Review,” Biomed Inform Insights, vol. 8, pp. 1–
10, Jan. 2016, doi: 10.4137/BII.S31559.
[28] A. B. R. Shatte, D. M. Hutchinson, and S. J. Teague, “Machine learning in mental health: a
scoping review of methods and applications,” Psychol. Med., vol. 49, no. 09, pp. 1426–
1448, Jul. 2019, doi: 10.1017/S0033291719000151.
[29] O. Kofman, N. Meiran, E. Greenberg, M. Balas, and H. Cohen, “Enhanced performance on
executive functions associated with examination stress: Evidence from task-switching and
Stroop paradigms,” Cognition and Emotion, vol. 20, no. 5, pp. 577–595, Aug. 2006, doi:
10.1080/02699930500270913.
[30] S. Mayya, V. Jilla, V. N. Tiwari, M. M. Nayak, and R. Narayanan, “Continuous monitoring
of stress on smartphone using heart rate variability,” in 2015 IEEE 15th International
Conference on Bioinformatics and Bioengineering (BIBE), Nov. 2015, pp. 1–5. doi:
10.1109/BIBE.2015.7367627.
[31] D. A. Dimitriev, E. V. Saperova, and A. D. Dimitriev, “State Anxiety and Nonlinear
Dynamics of Heart Rate Variability in Students,” PLOS ONE, vol. 11, no. 1, p. e0146131,
Jan. 2016, doi: 10.1371/[Link].0146131.
[32] M. R. Huerta-Franco, F. M. Vargas-Luna, and I. Delgadillo-Holtfort, “Effects of
psychological stress test on the cardiac response of public safety workers: alternative
parameters to autonomic balance,” J. Phys.: Conf. Ser., vol. 582, p. 012040, Jan. 2015,
doi: 10.1088/1742-6596/582/1/012040.
[33] F. Wilhelm, P. Grossman, and W. Roth, “Assessment of heart rate variability during
alterations in stress: Complex demodulation vs. spectral analysis,” Biomedical sciences
instrumentation, vol. 41, pp. 346–51, Feb. 2005.
[34] R. K. Nath, H. Thapliyal, A. Caban-Holt, and S. P. Mohanty, “Machine Learning Based
Solutions for Real-Time Stress Monitoring,” IEEE Consumer Electron. Mag., vol. 9, no.
5, pp. 34–41, Sep. 2020, doi: 10.1109/MCE.2020.2993427.
[35] A. C. Tricco et al., “PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist
and Explanation,” Ann Intern Med, vol. 169, no. 7, pp. 467–473, Oct. 2018, doi:
10.7326/M18-0850.
[36] M. Ouzzani, H. Hammady, Z. Fedorowicz, and A. Elmagarmid, “Rayyan—a web and
mobile app for systematic reviews,” Systematic Reviews, vol. 5, no. 1, p. 210, Dec. 2016,
doi: 10.1186/s13643-016-0384-4.
[37] Cleveland Clinic, “Heart Rate Variability (HRV): What It Is and How You Can Track It,”
Cleveland Clinic. Accessed: Aug. 08, 2022. [Online]. Available:
[Link]
[38] F. Shaffer and J. P. Ginsberg, “An Overview of Heart Rate Variability Metrics and Norms,”
Front Public Health, vol. 5, p. 258, Sep. 2017, doi: 10.3389/fpubh.2017.00258.
[39] Medicore, “HRV.” Accessed: Aug. 21, 2022. [Online]. Available: [Link]
[Link]/en/technology/[Link]
[40] Y. Choi, Y.-M. Jeon, L. Wang, and K. Kim, “A Biological Signal-Based Stress Monitoring
Framework for Children Using Wearable Devices,” Sensors, vol. 17, no. 9, Art. no. 9, Sep.
2017, doi: 10.3390/s17091936.
[41] O. Sakri, C. Godin, G. Vila, E. Labyt, S. Charbonnier, and A. Campagne, “A Multi-User
Multi-Task Model For Stress Monitoring From Wearable Sensors,” in 2018 21st
International Conference on Information Fusion (FUSION), Jul. 2018, pp. 761–766. doi:
10.23919/ICIF.2018.8455378.
[42] M. Elgendi and C. Menon, “Machine Learning Ranks ECG as an Optimal Wearable
Biosignal for Assessing Driving Stress,” IEEE Access, vol. 8, pp. 34362–34374, 2020,
doi: 10.1109/ACCESS.2020.2974933.
[43] R. Wang, W. Jia, Z.-H. Mao, R. J. Sclabassi, and M. Sun, “Cuff-Free Blood Pressure
Estimation Using Pulse Transit Time and Heart Rate,” Int Conf Signal Process Proc, vol.
2014, pp. 115–118, Oct. 2014, doi: 10.1109/ICOSP.2014.7014980.
[44] Mayo Clinic, “Stress and high blood pressure: What’s the connection?,” Mayo Clinic.
Accessed: Aug. 09, 2022. [Online]. Available: [Link]
conditions/high-blood-pressure/in-depth/stress-and-high-blood-pressure/art-20044190
[45] AHA, “Managing Stress to Control High Blood Pressure,” [Link]. Accessed: Aug.
09, 2022. [Online]. Available: [Link]
pressure/changes-you-can-make-to-manage-high-blood-pressure/managing-stress-to-
control-high-blood-pressure
[46] K. Schultebraucks, M. Sijbrandij, I. Galatzer-Levy, J. Mouthaan, M. Olff, and M. van
Zuiden, “Forecasting individual risk for long-term Posttraumatic Stress Disorder in
emergency medical settings using biomedical data: A machine learning multicenter cohort
study,” Neurobiology of Stress, vol. 14, p. 100297, May 2021, doi:
10.1016/[Link].2021.100297.
[47] K. Masood and M. A. Alghamdi, “Modeling Mental Stress Using a Deep Learning
Framework,” IEEE Access, vol. 7, pp. 68446–68454, 2019, doi:
10.1109/ACCESS.2019.2917718.
[48] D. Shah, G. Y. Wang, M. Doborjeh, Z. Doborjeh, and N. Kasabov, “Deep Learning of EEG
Data in the NeuCube Brain-Inspired Spiking Neural Network Architecture for a Better
Understanding of Depression,” in Neural Information Processing, vol. 11955, T. Gedeon,
K. W. Wong, and M. Lee, Eds., in Lecture Notes in Computer Science, vol. 11955. ,
Cham: Springer International Publishing, 2019, pp. 195–206. doi: 10.1007/978-3-030-
36718-3_17.
[49] A. Akella et al., “Classifying Multi-Level Stress Responses From Brain Cortical EEG in
Nurses and Non-Health Professionals Using Machine Learning Auto Encoder,” IEEE J.
Transl. Eng. Health Med., vol. 9, pp. 1–9, 2021, doi: 10.1109/JTEHM.2021.3077760.
[50] A. Arsalan, M. Majid, S. M. Anwar, and U. Bagci, “Classification of Perceived Human
Stress using Physiological Signals,” in 2019 41st Annual International Conference of the
IEEE Engineering in Medicine and Biology Society (EMBC), Jul. 2019, pp. 1247–1250.
doi: 10.1109/EMBC.2019.8856377.
[51] L. Gonzalez-Carabarin, E. A. Castellanos-Alvarado, P. Castro-Garcia, and M. A. Garcia-
Ramirez, “Machine Learning for personalised stress detection: Inter-individual variability
of EEG-ECG markers for acute-stress response,” Computer Methods and Programs in
Biomedicine, vol. 209, p. 106314, Sep. 2021, doi: 10.1016/[Link].2021.106314.
[52] A. R. Subhani, W. Mumtaz, M. N. B. M. Saad, N. Kamel, and A. S. Malik, “Machine
Learning Framework for the Detection of Mental Stress at Multiple Levels,” IEEE Access,
vol. 5, pp. 13545–13556, 2017, doi: 10.1109/ACCESS.2017.2723622.
[53] P. Nagar and D. Sethia, “Brain Mapping Based Stress Identification Using Portable EEG
Based Device,” in 2019 11th International Conference on Communication Systems &
Networks (COMSNETS), Bengaluru, India: IEEE, Jan. 2019, pp. 601–606. doi:
10.1109/COMSNETS.2019.8711009.
[54] A. Arsalan, M. Majid, A. R. Butt, and S. M. Anwar, “Classification of Perceived Mental
Stress Using A Commercially Available EEG Headband,” IEEE Journal of Biomedical
and Health Informatics, vol. 23, no. 6, pp. 2257–2264, Nov. 2019, doi:
10.1109/JBHI.2019.2926407.
[55] B. S. McEwen and P. J. Gianaros, “Central role of the brain in stress and adaptation: Links
to socioeconomic status, health, and disease,” Ann N Y Acad Sci, vol. 1186, pp. 190–222,
Feb. 2010, doi: 10.1111/j.1749-6632.2009.05331.x.
[56] A. Arsalan, M. Majid, and S. M. Anwar, “Electroencephalography Based Machine
Learning Framework for Anxiety Classification,” in Intelligent Technologies and
Applications, vol. 1198, I. S. Bajwa, T. Sibalija, and D. N. A. Jawawi, Eds., in
Communications in Computer and Information Science, vol. 1198. , Singapore: Springer
Singapore, 2020, pp. 187–197. doi: 10.1007/978-981-15-5232-8_17.
[57] A. Haider et al., “An Iris based Smart System for Stress Identification,” in 2019
International Conference on Electrical, Communication, and Computer Engineering
(ICECCE), Jul. 2019, pp. 1–5. doi: 10.1109/ICECCE47252.2019.8940707.
[58] V. Skaramagkas et al., “Review of eye tracking metrics involved in emotional and
cognitive processes,” IEEE Reviews in Biomedical Engineering, pp. 1–1, 2021, doi:
10.1109/RBME.2021.3066072.
[59] G. Giannakakis et al., “Stress and anxiety detection using facial cues from videos,”
Biomedical Signal Processing and Control, vol. 31, pp. 89–101, Jan. 2017, doi:
10.1016/[Link].2016.06.020.
[60] A. I. Korda et al., “Recognition of Blinks Activity Patterns during Stress Conditions Using
CNN and Markovian Analysis,” Signals, vol. 2, no. 1, Art. no. 1, Mar. 2021, doi:
10.3390/signals2010006.
[61] E. A. Bauer, K. A. Wilson, and A. MacNamara, “3.03 - Cognitive and Affective
Psychophysiology,” in Comprehensive Clinical Psychology (Second Edition), G. J. G.
Asmundson, Ed., Oxford: Elsevier, 2022, pp. 49–61. doi: 10.1016/B978-0-12-818697-
8.00013-3.
[62] M. E. Dawson, A. M. Schell, and D. L. Filion, “The electrodermal system,” in Handbook of
psychophysiology, 4th ed, in Cambridge handbooks in psychology. , New York, NY, US:
Cambridge University Press, 2017, pp. 217–243.
[63] P. Philippot, G. Chapelle, and S. Blairy, “Respiratory feedback in the generation of
emotion,” Cognition & Emotion, vol. 16, no. 5, pp. 605–627, Aug. 2002, doi:
10.1080/02699930143000392.
[64] W. M. Suess, A. B. Alexander, D. D. Smith, H. W. Sweeney, and R. J. Marion, “The effects
of psychological stress on respiration: a preliminary study of anxiety and
hyperventilation,” Psychophysiology, vol. 17, no. 6, pp. 535–540, Nov. 1980, doi:
10.1111/j.1469-8986.1980.tb02293.x.
[65] H. D. Cohen, D. R. Goodenough, H. A. Witkin, P. Oltman, H. Gould, and E. Shulman,
“The Effects of Stress on Components of the Respiration Cycle,” Psychophysiology, vol.
12, no. 4, pp. 377–380, 1975, doi: 10.1111/j.1469-8986.1975.tb00005.x.
[66] W. Seo, N. Kim, S. Kim, C. Lee, and S.-M. Park, “Deep ECG-Respiration Network
(DeepER Net) for Recognizing Mental Stress,” Sensors, vol. 19, no. 13, p. 3021, Jul.
2019, doi: 10.3390/s19133021.
[67] F. Kong, W. Wen, G. Liu, R. Xiong, and X. Yang, “Autonomic nervous pattern analysis of
trait anxiety,” Biomedical Signal Processing and Control, vol. 71, p. 103129, Jan. 2022,
doi: 10.1016/[Link].2021.103129.
[68] P. Cipresso, D. Colombo, and G. Riva, “Computational Psychometrics Using
Psychophysiological Measures for the Assessment of Acute Mental Stress,” Sensors, vol.
19, no. 4, Art. no. 4, Jan. 2019, doi: 10.3390/s19040781.
[69] “Electromyography (EMG).” Accessed: Aug. 09, 2022. [Online]. Available:
[Link]
emg
[70] A. A. Al-Jumaily, N. Matin, and A. N. Hoshyar, “Machine Learning Based Biosignals
Mental Stress Detection,” in Soft Computing in Data Science, A. Mohamed, B. W. Yap, J.
M. Zain, and M. W. Berry, Eds., in Communications in Computer and Information
Science. Singapore: Springer, 2021, pp. 28–41. doi: 10.1007/978-981-16-7334-4_3.
[71] S. DHAOUADI and M. M. BEN KHELIFA, “A multimodal Physiological-Based Stress
Recognition: Deep Learning Models’ Evaluation in Gamers’ Monitoring Application,” in
2020 5th International Conference on Advanced Technologies for Signal and Image
Processing (ATSIP), Sep. 2020, pp. 1–6. doi: 10.1109/ATSIP49331.2020.9231666.
[72] W. C. Liang, J. Yuan, D. C. Sun, M. H. Lin, and S.-D. City, “Variation in Physiological
Parameters Before and After an In- door Simulated Driving Task: Effect of Exercise
Break,” p. 14, 2007.
[73] A. Sano and R. W. Picard, “Stress Recognition Using Wearable Sensors and Mobile
Phones,” in 2013 Humaine Association Conference on Affective Computing and Intelligent
Interaction, Geneva, Switzerland: IEEE, Sep. 2013, pp. 671–676. doi:
10.1109/ACII.2013.117.
[74] A. Ghandeharioun et al., “Objective Assessment of Depressive Symptoms with Machine
Learning and Wearable Sensors Data,” p. 8, 2017.
[75] Y. Tazawa et al., “Evaluating depression with multimodal wristband-type wearable device:
screening and assessing patient severity utilizing machine-learning,” Heliyon, vol. 6, no. 2,
p. e03274, Feb. 2020, doi: 10.1016/[Link].2020.e03274.
[76] K. Yoshiuchi et al., “Stressful life events and habitual physical activity in older adults: 1-
year accelerometer data from the Nakanojo Study,” Mental Health and Physical Activity,
vol. 3, no. 1, pp. 23–25, Jun. 2010, doi: 10.1016/[Link].2010.02.001.
[77] M. Sadeghi, A. D. McDonald, and F. Sasangohar, “Posttraumatic Stress Disorder
Hyperarousal Event Detection Using Smartwatch Physiological and Activity Data,”
arXiv:2109.14743 [cs], Sep. 2021, Accessed: Apr. 25, 2022. [Online]. Available:
[Link]
[78] P. Schmidt, A. Reiss, R. Duerichen, C. Marberger, and K. Van Laerhoven, “Introducing
WESAD, a Multimodal Dataset for Wearable Stress and Affect Detection,” in
Proceedings of the 20th ACM International Conference on Multimodal Interaction, in
ICMI ’18. New York, NY, USA: Association for Computing Machinery, Oct. 2018, pp.
400–408. doi: 10.1145/3242969.3242985.
[79] R. Dai, C. Lu, L. Yun, E. Lenze, M. Avidan, and T. Kannampallil, “Comparing stress
prediction models using smartwatch physiological signals and participant self-reports,”
Computer Methods and Programs in Biomedicine, vol. 208, p. 106207, Sep. 2021, doi:
10.1016/[Link].2021.106207.
[80] P. Siirtola and J. Röning, “Comparison of Regression and Classification Models for User-
Independent and Personal Stress Detection,” Sensors, vol. 20, no. 16, Art. no. 16, Jan.
2020, doi: 10.3390/s20164402.
[81] Frontiers, “Speech signal analysis with applications in biomedicine and the life sciences |
Frontiers Research Topic.” Accessed: Aug. 09, 2022. [Online]. Available:
[Link]
applications-in-biomedicine-and-the-life-sciences
[82] K. Nkurikiyeyezu, A. Yokokubo, and G. Lopez, “The Effect of Person-Specific Biometrics
in Improving Generic Stress Predictive Models.” arXiv, Dec. 31, 2019. Accessed: Aug.
21, 2022. [Online]. Available: [Link]
[83] E. Garcia-Ceja et al., “HTAD: A Home-Tasks Activities Dataset with Wrist-accelerometer
and Audio Features,” p. 10.
[84] S. Zhao, Q. Li, C. Li, Y. Li, and K. Lu, “A CNN-Based Method for Depression Detecting
Form Audio,” in Digital Health and Medical Analytics, Y. Wang, W. Y. C. Wang, Z. Yan,
and D. Zhang, Eds., in Communications in Computer and Information Science. Singapore:
Springer, 2021, pp. 1–10. doi: 10.1007/978-981-16-3631-8_1.
[85] W. C. de Melo, E. Granger, and A. Hadid, “Depression Detection Based on Deep
Distribution Learning,” in 2019 IEEE International Conference on Image Processing
(ICIP), Taipei, Taiwan: IEEE, Sep. 2019, pp. 4544–4548. doi:
10.1109/ICIP.2019.8803467.
[86] K. Awasthi, P. Nanda, and K. V. Suma, “Performance analysis of Machine Learning
techniques for classification of stress levels using PPG signals,” in 2020 IEEE
International Conference on Electronics, Computing and Communication Technologies
(CONECCT), Jul. 2020, pp. 1–6. doi: 10.1109/CONECCT50063.2020.9198481.
[87] Q. Cong, Z. Feng, F. Li, Y. Xiang, G. Rao, and C. Tao, “X-A-BiLSTM: a Deep Learning
Approach for Depression Detection in Imbalanced Data,” in 2018 IEEE International
Conference on Bioinformatics and Biomedicine (BIBM), Madrid, Spain: IEEE, Dec. 2018,
pp. 1624–1627. doi: 10.1109/BIBM.2018.8621230.
[88] Y. Ding, X. Chen, Q. Fu, and S. Zhong, “A Depression Recognition Method for College
Students Using Deep Integrated Support Vector Algorithm,” IEEE Access, vol. 8, pp.
75616–75629, 2020, doi: 10.1109/ACCESS.2020.2987523.
[89] Md. R. Islam, M. A. Kabir, A. Ahmed, A. R. M. Kamal, H. Wang, and A. Ulhaq,
“Depression detection from social network data using machine learning techniques,”
Health Inf Sci Syst, vol. 6, no. 1, p. 8, Aug. 2018, doi: 10.1007/s13755-018-0046-0.
[90] S. Liao, Q. zhang, and R. Gan, “Construction of real-time mental health early warning
system based on machine learning,” J. Phys.: Conf. Ser., vol. 1812, no. 1, p. 012032, Feb.
2021, doi: 10.1088/1742-6596/1812/1/012032.
[91] M. Pandey, B. Jha, and R. Thakur, “An Exploratory Analysis Pertaining to Stress Detection
in Adolescents,” in Soft Computing: Theories and Applications, M. Pant, T. Kumar
Sharma, R. Arya, B. C. Sahana, and H. Zolfagharinia, Eds., in Advances in Intelligent
Systems and Computing. Singapore: Springer, 2020, pp. 413–421. doi: 10.1007/978-981-
15-4032-5_38.
[92] T. Simms, C. Ramstedt, M. Rich, M. Richards, T. Martinez, and C. Giraud-Carrier,
“Detecting Cognitive Distortions Through Machine Learning Text Analytics,” in 2017
IEEE International Conference on Healthcare Informatics (ICHI), Park City, UT, USA:
IEEE, Aug. 2017, pp. 508–512. doi: 10.1109/ICHI.2017.39.
[93] X. Wang et al., “Assessing depression risk in Chinese microblogs: a corpus and machine
learning methods,” in 2019 IEEE International Conference on Healthcare Informatics
(ICHI), Jun. 2019, pp. 1–5. doi: 10.1109/ICHI.2019.8904506.
[94] “[Link] - DSM.” Accessed: Aug. 15, 2022. [Online]. Available:
[Link]
[95] M. U. Salma and K. A. A. Ann, “Active Learning from an Imbalanced Dataset: A Study
Conducted on the Depression, Anxiety, and Stress Dataset,” in Handbook of Machine
Learning for Computational Optimization, CRC Press, 2021.
[96] M.-S. Mushtaq and A. Mellouk, “4 - QoE-based Power Efficient LTE Downlink
Scheduler,” in Quality of Experience Paradigm in Multimedia Services, M.-S. Mushtaq
and A. Mellouk, Eds., Elsevier, 2017, pp. 91–125. doi: 10.1016/B978-1-78548-109-
3.50004-7.
[97] V. Arun, P. V., M. Krishna, A. B.V., P. S.K., and S. V., “A Boosted Machine Learning
Approach For Detection of Depression,” in 2018 IEEE Symposium Series on
Computational Intelligence (SSCI), Nov. 2018, pp. 41–47. doi:
10.1109/SSCI.2018.8628945.
[98] S. Casaccia et al., “Measurement of Users’ Well-Being Through Domotic Sensors and
Machine Learning Algorithms,” IEEE Sensors J., vol. 20, no. 14, pp. 8029–8038, Jul.
2020, doi: 10.1109/JSEN.2020.2981209.
[99] T. Iliou et al., “ILIOU machine learning preprocessing method for depression type
prediction,” Evolving Systems, vol. 10, no. 1, pp. 29–39, Mar. 2019, doi: 10.1007/s12530-
017-9205-9.
[100] R. Joseph, S. Udupa, S. Jangale, K. Kotkar, and P. Pawar, “Employee Attrition Using
Machine Learning And Depression Analysis,” in 2021 5th International Conference on
Intelligent Computing and Control Systems (ICICCS), Madurai, India: IEEE, May 2021,
pp. 1000–1005. doi: 10.1109/ICICCS51141.2021.9432259.
[101] R. M. Khalil and A. Al-Jumaily, “Machine learning based prediction of depression among
type 2 diabetic patients,” in 2017 12th International Conference on Intelligent Systems
and Knowledge Engineering (ISKE), Nanjing: IEEE, Nov. 2017, pp. 1–5. doi:
10.1109/ISKE.2017.8258766.
[102] A. Kumar, “Machine learning for psychological disorder prediction in Indians during
COVID-19 nationwide lockdown,” IDT, vol. 15, no. 1, pp. 161–172, Mar. 2021, doi:
10.3233/IDT-200061.
[103] P. Kumar, R. Chauhan, T. Stephan, A. Shankar, and S. Thakur, “A Machine Learning
Implementation for Mental Health Care. Application: Smart Watch for Depression
Detection,” in 2021 11th International Conference on Cloud Computing, Data Science &
Engineering (Confluence), Jan. 2021, pp. 568–574. doi:
10.1109/Confluence51648.2021.9377199.
[104] S. T. A. Mary and L. Jabasheela, “ANT COLONY OPTIMIZATION BASED FEATURE
SELECTION AND DATA CLASSIFICATION FOR DEPRESSION ANXIETY AND
STRESS,” COMPUTER SCIENCE, vol. 9, p. 8, 2018.
[105] M. D. Nemesure, M. V. Heinz, R. Huang, and N. C. Jacobson, “Predictive modeling of
depression and anxiety using electronic health records and a novel machine learning
approach with artificial intelligence,” Sci Rep, vol. 11, no. 1, p. 1980, Dec. 2021, doi:
10.1038/s41598-021-81368-4.
[106] J. S. Obeid et al., “Automated detection of altered mental status in emergency department
clinical notes: a deep learning approach,” BMC Med Inform Decis Mak, vol. 19, no. 1, p.
164, Aug. 2019, doi: 10.1186/s12911-019-0894-9.
[107] U. S. Reddy, A. V. Thota, and A. Dharun, “Machine Learning Techniques for Stress
Prediction in Working Employees,” in 2018 IEEE International Conference on
Computational Intelligence and Computing Research (ICCIC), Madurai, India: IEEE,
Dec. 2018, pp. 1–4. doi: 10.1109/ICCIC.2018.8782395.
[108] A. Sau and I. Bhakta, “Predicting anxiety and depression in elderly patients using machine
learning technology,” Healthc. technol. lett., vol. 4, no. 6, pp. 238–243, Dec. 2017, doi:
10.1049/htl.2016.0096.
[109] A. Sharma and W. J. M. I. Verbeke, “Improving Diagnosis of Depression With
XGBOOST Machine Learning Model and a Large Biomarkers Dutch Dataset (n =
11,081),” Front. Big Data, vol. 3, p. 15, Apr. 2020, doi: 10.3389/fdata.2020.00015.
[110] H. Yang and P. A. Bath, “Automatic Prediction of Depression in Older Age,” in
Proceedings of the third International Conference on Medical and Health Informatics
2019, in ICMHI 2019. New York, NY, USA: Association for Computing Machinery, May
2019, pp. 36–44. doi: 10.1145/3340037.3340042.
[111] W. Zhang, H. Liu, V. M. B. Silenzio, P. Qiu, and W. Gong, “Machine Learning Models
for the Prediction of Postpartum Depression: Application and Comparison Based on a
Cohort Study,” JMIR Med Inform, vol. 8, no. 4, p. e15516, Apr. 2020, doi: 10.2196/15516.
[112] P. Annapureddy et al., “Predicting PTSD Severity in Veterans from Self-reports for Early
Intervention: A Machine Learning Approach,” in 2020 IEEE 21st International
Conference on Information Reuse and Integration for Data Science (IRI), Las Vegas, NV,
USA: IEEE, Aug. 2020, pp. 201–208. doi: 10.1109/IRI49571.2020.00036.
[113] B. Li and A. Sano, “Early versus Late Modality Fusion of Deep Wearable Sensor Features
for Personalized Prediction of Tomorrow’s Mood, Health, and Stress,” in 2020 42nd
Annual International Conference of the IEEE Engineering in Medicine & Biology Society
(EMBC), Montreal, QC, Canada: IEEE, Jul. 2020, pp. 5896–5899. doi:
10.1109/EMBC44109.2020.9175463.
[114] N. P. Novani, L. Arief, R. Anjasmara, and A. S. Prihatmanto, “Heart Rate Variability
Frequency Domain for Detection of Mental Stress Using Support Vector Machine,” in
2018 International Conference on Information Technology Systems and Innovation
(ICITSI), Oct. 2018, pp. 520–525. doi: 10.1109/ICITSI.2018.8695938.
[115] G. Giannakakis, K. Marias, and M. Tsiknakis, “A stress recognition system using HRV
parameters and machine learning techniques,” in 2019 8th International Conference on
Affective Computing and Intelligent Interaction Workshops and Demos (ACIIW), Sep.
2019, pp. 269–272. doi: 10.1109/ACIIW.2019.8925142.
[116] Y. S. Can, N. Chalabianloo, D. Ekiz, and C. Ersoy, “Continuous Stress Detection Using
Wearable Sensors in Real Life: Algorithmic Programming Contest Case Study,” Sensors,
vol. 19, no. 8, p. 1849, Apr. 2019, doi: 10.3390/s19081849.
[117] L. V. Coutts, D. Plans, A. W. Brown, and J. Collomosse, “Deep learning with wearable
based heart rate variability for prediction of mental and general health,” Journal of
Biomedical Informatics, vol. 112, p. 103610, Dec. 2020, doi: 10.1016/[Link].2020.103610.
[118] D. Cho et al., “Detection of Stress Levels from Biosignals Measured in Virtual Reality
Environments Using a Kernel-Based Extreme Learning Machine,” Sensors, vol. 17, no.
10, Art. no. 10, Oct. 2017, doi: 10.3390/s17102435.
[119] U. Pluntke, S. Gerke, A. Sridhar, J. Weiss, and B. Michel, “Evaluation and Classification
of Physical and Psychological Stress in Firefighters using Heart Rate Variability,” Annu
Int Conf IEEE Eng Med Biol Soc, vol. 2019, pp. 2207–2212, Jul. 2019, doi:
10.1109/EMBC.2019.8856596.
[120] R. Lima, D. Osório, and H. Gamboa, “Heart Rate Variability and Electrodermal Activity
Biosignal Processing: Predicting the Autonomous Nervous System Response in Mental
Stress,” in Biomedical Engineering Systems and Technologies, A. Roque, A. Tomczyk, E.
De Maria, F. Putze, R. Moucek, A. Fred, and H. Gamboa, Eds., in Communications in
Computer and Information Science. Cham: Springer International Publishing, 2020, pp.
328–351. doi: 10.1007/978-3-030-46970-2_16.
[121] V. Montesinos, F. Dell’Agnola, A. Arza, A. Aminifar, and D. Atienza, “Multi-Modal
Acute Stress Recognition Using Off-the-Shelf Wearable Devices,” in 2019 41st Annual
International Conference of the IEEE Engineering in Medicine and Biology Society
(EMBC), Berlin, Germany: IEEE, Jul. 2019, pp. 2196–2201. doi:
10.1109/EMBC.2019.8857130.
[122] A. Ghandeharioun et al., “Objective assessment of depressive symptoms with machine
learning and wearable sensors data,” in 2017 Seventh International Conference on
Affective Computing and Intelligent Interaction (ACII), Oct. 2017, pp. 325–332. doi:
10.1109/ACII.2017.8273620.
[123] S. Blum, N. S. J. Jacobsen, M. G. Bleichner, and S. Debener, “A Riemannian Modification
of Artifact Subspace Reconstruction for EEG Artifact Handling,” Frontiers in Human
Neuroscience, vol. 13, 2019, Accessed: Aug. 22, 2022. [Online]. Available:
[Link]
[124] R. Van de Schoot, “Latent Growth Mixture Models to estimate PTSD trajectories,”
European Journal of Psychotraumatology, vol. 6, no. 1, p. 27503, Dec. 2015, doi:
10.3402/ejpt.v6.27503.
[125] N. Ram and K. J. Grimm, “Methods and Measures: Growth mixture modeling: A method
for identifying differences in longitudinal change among unobserved groups,”
International Journal of Behavioral Development, vol. 33, no. 6, pp. 565–576, Nov. 2009,
doi: 10.1177/0165025409343765.
[126] W. van Breda, J. Pastor, M. Hoogendoorn, J. Ruwaard, J. Asselbergs, and H. Riper,
“Exploring and Comparing Machine Learning Approaches for Predicting Mood Over
Time,” in Innovation in Medicine and Healthcare 2016, vol. 60, Y.-W. Chen, S. Tanaka,
R. J. Howlett, and L. C. Jain, Eds., in Smart Innovation, Systems and Technologies, vol.
60. , Cham: Springer International Publishing, 2016, pp. 37–47. doi: 10.1007/978-3-319-
39687-3_4.
[127] “Understanding Dynamic Time Warping - The Databricks Blog,” Databricks. Accessed:
Aug. 22, 2022. [Online]. Available:
[Link]
[128] Q. Gui, Z. Jin, and W. Xu, “Exploring missing data prediction in medical monitoring: A
performance analysis approach,” in 2014 IEEE Signal Processing in Medicine and
Biology Symposium (SPMB), Dec. 2014, pp. 1–6. doi: 10.1109/SPMB.2014.7002968.
[129] A. D. McDonald, F. Sasangohar, A. Jatav, and A. H. Rao, “Continuous monitoring and
detection of post-traumatic stress disorder (PTSD) triggers among veterans: A supervised
machine learning approach,” IISE Transactions on Healthcare Systems Engineering, vol.
9, no. 3, pp. 201–211, Jul. 2019, doi: 10.1080/24725579.2019.1583703.
[130] M. A. Adheena, N. Sindhu, and S. Jerritta, “Physiological Detection of Anxiety,” in 2018
International Conference on Circuits and Systems in Digital Enterprise Technology
(ICCSDET), Kottayam, India: IEEE, Dec. 2018, pp. 1–5. doi:
10.1109/ICCSDET.2018.8821162.
[131] W. Gerych, E. Agu, and E. Rundensteiner, “Classifying Depression in Imbalanced
Datasets Using an Autoencoder- Based Anomaly Detection Approach,” in 2019 IEEE 13th
International Conference on Semantic Computing (ICSC), Newport Beach, CA, USA:
IEEE, Jan. 2019, pp. 124–127. doi: 10.1109/ICOSC.2019.8665535.
[132] D. Huysmans et al., “Unsupervised Learning for Mental Stress Detection - Exploration of
Self-Organizing Maps,” in Proc. of Biosignals 2018, Scitepress, Jan. 2018, pp. 26–35. doi:
10.5220/0006541100260035.
[133] R. Jain and W. Xu, “Artificial Intelligence based wrapper for high dimensional feature
selection,” BMC Bioinformatics, vol. 24, no. 1, p. 392, Oct. 2023, doi: 10.1186/s12859-
023-05502-x.
[134] U. Braga-Neto, Fundamentals of Pattern Recognition and Machine Learning. Cham:
Springer International Publishing, 2020. doi: 10.1007/978-3-030-27656-0.
[135] Y. Zhang, S. Wang, A. Hermann, R. Joly, and J. Pathak, “Development and validation of a
machine learning algorithm for predicting the risk of postpartum depression among
pregnant women,” Journal of Affective Disorders, vol. 279, pp. 1–8, Jan. 2021, doi:
10.1016/[Link].2020.09.113.
[136] D. Deka and B. Deka, “Characterization of heart rate variability signal for distinction of
meditative and pre-meditative states,” Biomedical Signal Processing and Control, vol. 66,
p. 102414, Apr. 2021, doi: 10.1016/[Link].2021.102414.
[137] M. Radovic, M. Ghalwash, N. Filipovic, and Z. Obradovic, “Minimum redundancy
maximum relevance feature selection approach for temporal gene expression data,” BMC
Bioinformatics, vol. 18, no. 1, p. 9, Jan. 2017, doi: 10.1186/s12859-016-1423-9.
[138] L. Xia, A. S. Malik, and A. R. Subhani, “A physiological signal-based method for early
mental-stress detection,” Biomedical Signal Processing and Control, vol. 46, pp. 18–32,
Sep. 2018, doi: 10.1016/[Link].2018.06.004.
[139] L. Quintero, P. Papapetrou, J. E. Munoz, and U. Fors, “Implementation of Mobile-Based
Real-Time Heart Rate Variability Detection for Personalized Healthcare,” in 2019
International Conference on Data Mining Workshops (ICDMW), Beijing, China: IEEE,
Nov. 2019, pp. 838–846. doi: 10.1109/ICDMW.2019.00123.
[140] M. Srividya, S. Mohanavalli, and N. Bhalaji, “Behavioral Modeling for Mental Health
using Machine Learning Algorithms,” J Med Syst, vol. 42, no. 5, p. 88, Apr. 2018, doi:
10.1007/s10916-018-0934-5.
[141] P. Zhang et al., “Psychological Stress Detection According to ECG Using a Deep
Learning Model with Attention Mechanism,” Applied Sciences, vol. 11, no. 6, Art. no. 6,
Jan. 2021, doi: 10.3390/app11062848.
[142] Y. E. Alharahsheh and M. A. Abdullah, “Predicting Individuals Mental Health Status in
Kenya using Machine Learning Methods,” in 2021 12th International Conference on
Information and Communication Systems (ICICS), Valencia, Spain: IEEE, May 2021, pp.
94–98. doi: 10.1109/ICICS52457.2021.9464608.
[143] “Fuzzy C-Means Clustering - MATLAB & Simulink.” Accessed: Aug. 14, 2022. [Online].
Available: [Link]
[144] V. Trevisan, “Using SHAP Values to Explain How Your Machine Learning Model
Works,” Medium. Accessed: Aug. 22, 2022. [Online]. Available:
[Link]
model-works-732b3f40e137
[145] E. U. P. Vieira, “You are underutilizing shap values — feature groups and correlations,”
Medium. Accessed: Aug. 22, 2022. [Online]. Available:
[Link]
correlations-8df1b136e2c2
[146] C. O’Sullivan, “Analysing Interactions with SHAP,” Medium. Accessed: Aug. 22, 2022.
[Online]. Available: [Link]
8c4a2bc11c2a
[147] A. Saeed, T. Ozcelebi, J. Lukkien, J. B. F. van Erp, and S. Trajanovski, “Model
Adaptation and Personalization for Physiological Stress Detection,” in 2018 IEEE 5th
International Conference on Data Science and Advanced Analytics (DSAA), Oct. 2018,
pp. 209–216. doi: 10.1109/DSAA.2018.00031.
[148] A. Mallol-Ragolta, S. Dhamija, and T. E. Boult, “A Multimodal Approach for Predicting
Changes in PTSD Symptom Severity,” in Proceedings of the 20th ACM International
Conference on Multimodal Interaction, Boulder CO USA: ACM, Oct. 2018, pp. 324–333.
doi: 10.1145/3242969.3242981.
[149] J. Huang, X. Luo, and X. Peng, “A Novel Classification Method for a Driver’s Cognitive
Stress Level by Transferring Interbeat Intervals of the ECG Signal to Pictures,” Sensors,
vol. 20, no. 5, Art. no. 5, Jan. 2020, doi: 10.3390/s20051340.
[150] C. M. Durán Acevedo, J. K. Carrillo Gómez, and C. A. Albarracín Rojas, “Academic
stress detection on university students during COVID-19 outbreak by using an electronic
nose and the galvanic skin response,” Biomedical Signal Processing and Control, vol. 68,
p. 102756, Jul. 2021, doi: 10.1016/[Link].2021.102756.
[151] J. C. Y. Wong, J. Wang, E. Y. Fu, H. V. Leong, and G. Ngai, “Activity Recognition and
Stress Detection via Wristband,” in Proceedings of the 17th International Conference on
Advances in Mobile Computing & Multimedia, in MoMM2019. New York, NY, USA:
Association for Computing Machinery, Dec. 2019, pp. 102–106. doi:
10.1145/3365921.3365950.
[152] J. Šalkevicius, R. Damaševičius, R. Maskeliunas, and I. Laukienė, “Anxiety Level
Recognition for Virtual Reality Therapy System Using Physiological Signals,”
Electronics, vol. 8, no. 9, Art. no. 9, Sep. 2019, doi: 10.3390/electronics8091039.
[153] A. Wongkoblap, M. A. Vadillo, and V. Curcin, “Classifying Depressed Users With
Multiple Instance Learning from Social Network Data,” in 2018 IEEE International
Conference on Healthcare Informatics (ICHI), New York, NY: IEEE, Jun. 2018, pp. 436–
436. doi: 10.1109/ICHI.2018.00088.
[154] V. Chandra, A. Priyarup, and D. Sethia, “Comparative Study of Physiological Signals
from Empatica E4 Wristband for Stress Classification,” in Advances in Computing and
Data Sciences, M. Singh, V. Tyagi, P. K. Gupta, J. Flusser, T. Ören, and V. R. Sonawane,
Eds., in Communications in Computer and Information Science. Cham: Springer
International Publishing, 2021, pp. 218–229. doi: 10.1007/978-3-030-88244-0_21.
[155] J.-W. Baek and K. Chung, “Context Deep Neural Network Model for Predicting
Depression Risk Using Multiple Regression,” IEEE Access, vol. 8, pp. 18171–18181,
2020, doi: 10.1109/ACCESS.2020.2968393.
[156] I. Bichindaritz, C. Breen, E. Cole, N. Keshan, and P. Parimi, “Feature Selection and
Machine Learning Based Multilevel Stress Detection from ECG Signals,” in Innovation in
Medicine and Healthcare 2017, Y.-W. Chen, S. Tanaka, R. J. Howlett, and L. C. Jain,
Eds., in Smart Innovation, Systems and Technologies. Cham: Springer International
Publishing, 2018, pp. 202–213. doi: 10.1007/978-3-319-59397-5_22.
[157] M. Maritsch et al., “Improving heart rate variability measurements from consumer
smartwatches with machine learning,” in Adjunct Proceedings of the 2019 ACM
International Joint Conference on Pervasive and Ubiquitous Computing and Proceedings
of the 2019 ACM International Symposium on Wearable Computers, London United
Kingdom: ACM, Sep. 2019, pp. 934–938. doi: 10.1145/3341162.3346276.
[158] N. E. J. Asha, Ehtesum-Ul-Islam, and R. Khan, “Low-Cost Heart Rate Sensor and Mental
Stress Detection Using Machine Learning,” in 2021 5th International Conference on
Trends in Electronics and Informatics (ICOEI), Tirunelveli, India: IEEE, Jun. 2021, pp.
1369–1374. doi: 10.1109/ICOEI51242.2021.9452873.
[159] A. Ghandeharioun et al., “Objective assessment of depressive symptoms with machine
learning and wearable sensors data,” in 2017 Seventh International Conference on
Affective Computing and Intelligent Interaction (ACII), Oct. 2017, pp. 325–332. doi:
10.1109/ACII.2017.8273620.
[160] R. Katarya and S. Maan, “Predicting Mental health disorders using Machine Learning for
employees in technical and non-technical companies,” in 2020 IEEE International
Conference on Advances and Developments in Electrical and Electronics Engineering
(ICADEE), Coimbatore, India: IEEE, Dec. 2020, pp. 1–5. doi:
10.1109/ICADEE51157.2020.9368923.
[161] A. E. Tate, R. C. McCabe, H. Larsson, S. Lundström, P. Lichtenstein, and R. Kuja-
Halkola, “Predicting mental health problems in adolescence using machine learning
techniques,” PLoS ONE, vol. 15, no. 4, p. e0230389, Apr. 2020, doi:
10.1371/[Link].0230389.
[162] S. Wshah, C. Skalka, and M. Price, “Predicting Posttraumatic Stress Disorder Risk: A
Machine Learning Approach,” JMIR Ment Health, vol. 6, no. 7, p. e13946, Jul. 2019, doi:
10.2196/13946.
[163] W. A. van Eeden et al., “Predicting the 9-year course of mood and anxiety disorders with
automated machine learning: A comparison between auto-sklearn, naïve Bayes classifier,
and traditional logistic regression,” Psychiatry Research, vol. 299, p. 113823, May 2021,
doi: 10.1016/[Link].2021.113823.
[164] H. Alharthi, “Predicting the level of generalized anxiety disorder of the coronavirus
pandemic among college age students using artificial intelligence technology,” in 2020
19th International Symposium on Distributed Computing and Applications for Business
Engineering and Science (DCABES), Xuzhou, China: IEEE, Oct. 2020, pp. 218–221. doi:
10.1109/DCABES50732.2020.00064.
[165] J. He, K. Li, X. Liao, P. Zhang, and N. Jiang, “Real-Time Detection of Acute Cognitive
Stress Using a Convolutional Neural Network From Electrocardiographic Signal,” IEEE
Access, vol. 7, pp. 42710–42717, 2019, doi: 10.1109/ACCESS.2019.2907076.
[166] R. K. Sah and H. Ghasemzadeh, “Stress Classification and Personalization: Getting the
most out of the least.” arXiv, Jul. 12, 2021. Accessed: Sep. 04, 2022. [Online]. Available:
[Link]
[167] Z. Zainudin, S. Hasan, S. M. Shamsuddin, and S. Argawal, “Stress Detection using
Machine Learning and Deep Learning,” J. Phys.: Conf. Ser., vol. 1997, no. 1, p. 012019,
Aug. 2021, doi: 10.1088/1742-6596/1997/1/012019.
[168] M. Sadeghi, A. D. McDonald, and F. Sasangohar, “Posttraumatic stress disorder
hyperarousal event detection using smartwatch physiological and activity data,” PLOS
ONE, vol. 17, no. 5, p. e0267749, May 2022, doi: 10.1371/[Link].0267749.

Appendix

Table A1. Summary of all articles included in the synthesis

Population Data
Population Personalization Real- Performance
Article Year Data Type Demographic ML Model Collection Study Type ML Type Ground-Truth
Type Method Time? (Best Model)
Data Tool
Wearable
HR, PPG, Audio Accuracy:
[40] 2017 Children Yes NB, DT, SVM Sensor, Naturalistic Classification
Signals 83% (DT)
Microphone
Pre-
collected Accuracy:
Boosting, DSM scores for
[97] 2018 Questionnaire Patients N=270 No Questionnaire data Classification 97.8%
XGBoost depression
without (XGBoost)
experiment
NN,
N=693 (108 frequency of Precision:
College Boosting, Weibo Sina
2020 Text Depressed, No Naturalistic Regression possitive and 0.88
[88] Students Adaboost, Application
585 Normal) negative words (Adaboost)
SVM
Indian Accuracy:
Ant Colony DASS-21
[104] 2018 Questionnaire College N=938 No Lab Classification DASS-21 scores 93.57%
Optimization Questionnaire
Students (ACO)
Depressed
N=142 (42 Depression
Patients and Accuracy:
[84] 2021 Audio Signals Patient, 100 No NN Naturalistic Classification Diagnosis
Healthy 84% (CNN)
Healthy) (RSDD) dataset
People
N=334 (86 Questionnaire Accuracy:
Unemployed LR, NB, DT,
[103] 2021 Questionnaire Depressed, No Naturalistic Classification scores for 89.6%
people KNN, SVM
248 Normal) depression (Ensemble)
HR, PPG,
Acceleration/Body Stress-inducing Accuracy:
[41] 2018 N=6 No NB, DT, SVM Empatica E4 Lab Classification
Movement, Skin task 93% (SVM)
Response
PCL-5
Skin Response, PCL-5 MSE: 138
[148] 2018 N=110 No SVM Lab Regression questionnaire
Questionnaire Questionnaire (SVR)
for PTSD
EMG, Skin Young Wearable Questionnaire Accuracy:
[71] 2020 N=15 Yes NN Lab Classification
Response Gamers Sensor scores 95% (LSTM)
Three phases
BIOPAC
of driving task Accuracy:
[149] 2020 HRV, ECG No NN MP150- Lab Classification
in driving 92.8% (CNN)
BioNomadix
simulator
Baseline NB, RF,
Removal, LDA/QDA, Ag/AgCl Stress-inducing Accuracy:
[115] 2019 HRV, ECG N=24 No Lab Classification
Pairwise KNN, SVM, Electrodes task 84.4% (SVM)
Transformation Other
Standard Scale KNN, SVM, Stress-inducing Accuracy:
[150] 2021 Skin Response Students N=25 No Gas sensor Lab Classification
on E-nose data. LDA task 96% (SVM)

HR, PPG, Skin College N=1 (22 years Stress-inducing Accuracy:


[151] 2021 No SVM Empatica E4 Lab Classification
Response Students old) task 80% (SVM)
Labeling by
LR, NB, RF, Accuracy:
[91] 2020 Text ~11k Tweets No Twitter Naturalistic Classification sentiment
SVM 90% (SVM)
analysis
N=100 (50
Number of Accuracy:
[57] 2019 Eye Tracking Healthy, 50 No NN Nvidia GPU Lab Classification
rings in eye Iris 98% (NN)
Non-healthy)
Patients Min-max
Skin Response, during normalization, Stress-inducing Accuracy:
[152] 2019 N=30 No SVM Empatica E4 Lab Classification
PPG, HR, HRV virtual Zero means task 86.3% (SVM)
speaking normalization
Hamilton
~14k posts
NN, SVM, Weibo Sina Depression F1 score:
[93] 2019 Text from Weibo No Naturalistic Classification
BERT Application Rating Scale for 0.538 (BERT)
Sina
depression
Normalization
by Term
Frequency– Accuracy:
Adult Labels by 94.5%,
[106] 2019 Questionnaire Inverse No NN, Other Naturalistic Classification
patients clinician AUROC:
Document
98.5%(NN)
Frequency (TF-
IDF)
Pre- Log-loss:
NN, LR, collected 0.241,
Adults above
[110] 2019 Questionnaire No Boosting, Questionnaire data Classification CES-D scores AUROC:
50
XGBoost, RF without 0.886
experiment (XGBoost)
Resting state
Respiratory, ECG, NB, DT,
ECG removed E-prime Accuracy:
[67] 2022 Questionnaire, Students N=99 No LDA/QDA, Lab Classification STAI scores
from data as software 73.68% (DT)
HRV KNN, SVM
baseline
Students and LR, NB, DT,
Clustering Accuracy:
[140] 2018 Questionnaire Working N=656 No K-means, Questionarate Naturalistic Classification
labels 90% (RF)
professionals KNN, SVM
PSS-14
Questionnaire, Accuracy:
[53] 2019 EEG Students N=63 No KNN, SVM Neurosky Lab Classification PSS-14 scores 74.43%
Mindwave (KNN)
EEG
Pre-
Precollected collected Accuracy:
Yoga and Chi Stress-inducing
[136] 2021 HRV, ECG N=12 No KNN dataset (from data Classification 95.31%
practitioners task
PhysioNet) without (SVM)
experiment
EEG, Skin College
NN, NB, MUSE EEG, Accuracy:
[50] 2019 Response, PPG, Students and N=28 No Lab Classification PSS scores
SVM Shimmer GSR+ 75% (NN)
HRV Instructors
College
NN, NB, Accuracy:
[54] 2019 EEG Students and N=28 No MUSE EEG Lab Classification PSS scores
SVM 92.85% (NN)
Instructors
N=431 (319 Facebook,
Questionnaire, Facebook Accuracy:
[153] 2018 Depressed, No NN, Other CES-D Naturalistic Classification CES-D scores
Text users 72% (CNN)
162 Normal) Questionnaire
NN,
Boosting,
College AUROC: 0.92
[131] 2019 Questionnaire N=48 No XGBoost, RF, Naturalistic Classification PHQ-9 scores
Students (SVM)
K-means,
SVM
Adaboost,
N=80 (30 RF, KNN,
Nurses and NeuroScan
Nurses, 50 SVM, LDA, Accuracy:
[49] 2021 EEG non-health No EEG, SynAmps Lab Classification Self-report
non-health Ridge, Deep 91% (SVM)
professionals 2 amplifier
professionals) Bilief
Network
Boosting,
Engineering
HR, Skin Adaboost, Stress-inducing Accuracy:
[154] 2021 College N=21 No Empatica E4 Lab Classification
Response, PPG XGBoost, RF, task 99.88% (RF)
Students
KNN
PSS scores Accuracy:
were used to LR, Boosting, 82.6%,
Acceleration/Body Both lab Stress-inducing
Students and change the Adaboost, Fossil Gen4 AUROC:
[79] 2021 Movement, PPG, N=32 Yes and Classification task, Self-
Staff threshold for XGBoost, RF, Explorist 0.790, F-1
Questionnaire naturalistic report
stress SVM score: 0.623
detection (SVM)
One model was
Pre-
Acceleration/Body trained for
collected Baseline and
Movement, Skin each subject RMSE: 0.03
[80] 2020 Drivers N=9 No RF Empatica E4 data Regression driving
Response, PPG, (with different (Regression)
without condition
HR, HRV thersholds for
experiment
each subject)
thoracic
respiration
Skin Response, LR, NB, DT, belt, skin
Accuracy:
[68] 2019 Respiratory, ECG, Students N=60 No RF, SVM, conductance Lab
74% (LR)
HRV Other adhesive
patches, BVP
sensor
NB, RF, KNN, web page
[90] 2021 Text Yes Naturalistic Classification Self-report
SVM, Other crawler
Pre-
Accuracy:
collected Precollected
Korean Context- 94.57%
[155] 2020 Questionnaire N=39,225 Yes data Classification Depression
people DNN (Context-
without Scores
DNN)
experiment
NN, NB, RF, Self-reported AUROC: 0.67
[129] 2019 HR, PPG Veterans N=100 Yes iPhone Naturalistic Classification
SVM, Other PTSD triggers (SVM)
HRV, PPG, PSS, NASA-TLX,
One model was
Acceleration/Body RF, KNN, LR, Samsung Gear STAI and other Accuracy:
[116] 2019 Students N=21 trained for Yes Naturalistic Classification
Movement, Skin NN S, Empatica E4 questionnaire 97.92% (RF)
each subject
Response scores
Resting state
removed from Zephyr Accuracy:
HRV, Respiratory, Stress-inducing
[66] 2019 Students N=18 ECG and RESP No NN BioHarness Lab Classification 83.9%
ECG task
data as 3.0 (DeepER)
baseline
Spiking
Healthy and SynAmps Accuracy:
Neural BDI scores for
[48] 2019 EEG Mild- N=22 No amplifier, 61- Lab Classification 72.13%
Network depression
depressed channel EEG (SNN)
(SNN)
HRV, PPG, Participants BioBeats PSS, STAI and Accuracy:
[117] 2020 N=632 No NN Naturalistic Classification
Questionnaire taking exams Biobeam band DASS scores 83% (LSTM)
Pre-
AVEC2013 and collected
BDI scores for RMSE: 8.5
[85] 2019 Video N=82 No NN, Other AVEC2014 data Regression
depression (CNN)
datasets without
experiment
459 posts Pre-
(207 NN, LR, NB, collected
Hand-labeled Accuracy:
[92] 2017 Text distorted, No DT, KNN, data Classification
data 73% (LR)
252 SOM without
undistorted) experiment
Skin Response, Healthy Samsung Gear Stress-inducing Accuracy:
[118] 2017 N=12 No CNN, SOM Lab Classification
PPG, HRV subjects VR task 95% (K-ELM)
Pre-
Depression
collected
NN, LR, DT, from electronic AUROC:
[135] 2020 Other Women N=69,169 No data Classification
XGBoost, RF health records 0.937 (LR)
without
(EHRs)
experiment
Pre-
Multi-task
Acceleration/Body collected
College learning with SNAP-SHOT
[113] 2020 Movement, Skin N=239 No NN data Classification Self-report
Students participants as dataset
Response without
tasks in NN
experiment
Healthy Accuracy:
[56] 2020 EEG N=28 No LR, RF, NN MUSE EEG Lab Classification STAI Scores
subjects 78.5% (RF)
Pre-
LR, NB, DT, Goldberg’s collected Goldberg
Accuracy:
[100] 2021 Questionnaire No RF, KNN, Depression data Classification questionnaire
86% (RF)
SVM Questionnaire without scores
experiment
Accuracy:
76%,
HR, PPG, Depressed
N=86 (45 Sensitivity:
Acceleration/Body and Healthy Silmee W20 healthy/patient
[75] 2019 Depressed, No XGBoost Naturalistic Classification 73%,
Movement, Skin Japanese Wristband participants
41 Healthy) Specificity:
Response, Other people
79%
(XGBoost)
Polar H7 chest Stress-inducing Accuracy:
[119] 2019 HRV, ECG Firefighters N=26 Yes DT, SVM Lab Classification
strap ECG task 88% (DT)
eMate EMA
Normalizing application, Mood self- MSE: 0.410
[126] 2016 Other N=270 No RF, SVM Naturalistic Regression
the variables iYouVU report (SVM)
application
Pre-
collected Accuracy: Up
Wearable Stress-inducing
[156] 2018 ECG, HRV Drivers N=17 No NN, DT, RF data Classification to 100%,
Sensors task
without AUC: 1 (RF)
experiment
HR, Hormones, Normalizing AUC: 0.89
[46] 2021 Patients N=417 No XGBoost Naturalistic Classification Self-report
BP, Respiratory the variables (XGBoost)
PLUX BITalino
HRV, PPG, Skin LR, RF, SVM, Stress-inducing Accuracy:
[120] 2020 N=15 No Wearable Lab Classification
Response Other task 80% (RF)
Sensor
Accuracy:
[114] 2018 HRV, PPG No SVM Polar H7 HRM Lab Classification Self-report
81% (SVM)
Pre-
NN, RF, collected
Stressed Accuracy:
[99] 2017 Questionnaire N=249 No Fuzzy, SVM, data Classification BDI scores
Students 100% (NN)
Other without
experiment
Accurary:
97.65%,
Pre-
Depressed N=11,081 Precision:
collected
and Healthy (570 self- 95.48%,
[109] 2020 Questionnaire No XGBoost data Classification Self-report
Dutch reported Recall:
without
citizens depression) 99.87%, F1
experiment
score: 0.98
(XGBoost)
Firstbeat Heart Rate
RMSE: 28.5
[157] 2019 HRV, PPG No NN Bodyguard 2, Lab Regression Monitor (HRM)
(Regression)
Smartwatch measurements
PPG Sensor, Stress-inducing Accuracy:
[158] 2021 HR, PPG No SVM Lab Classification
Arduino UNO task 62% (SVM)

Healthy ECG, EMG, Stress-inducing Accuracy:


[70] 2021 EMG, ECG, HRV N=16 No SVM Lab Classification
Drivers Volvo S70 task 93.7% (SVM)
Pre-
NN, Fuzzy, collected Pre-collected
Accuracy:
[101] 2017 Questionnaire No K-means, data Classification dataset for
97% (SVM)
SVM without depression
experiment
NN, DT, RF, 19-channel
Using K-means Stress-inducing Accuracy:
[51] 2021 EEG, ECG, HR Students N=24 No K-means, EEG, 12- Lab Classification
on EEG energy task 82% (ANN)
KNN, SVM channel ECG
Indian
people with LR, NB, RF, Accuracy:
[102] 2021 Questionnaire N=395 No Google Forms Naturalistic Classification Self-report
mental KNN, SVM 92.15%
disorder
EEG 128
channels,
Healthy Electrical Stress-inducing Accuracy:
[52] 2017 EEG N=42 No LR, NB, SVM Lab Classification
subjects Geodesic Net task 94.6% (NB)
Amps 300
amplifier
Pregnant Accuracy:
[111] 2020 Questionnaire N=508 No RF, SVM Questionnaire Naturalistic Classification EPDS scores
Women 80% (SVM)
HR, EMG, Skin LR, DT,
Stress markers Accuracy:
[42] 2020 Response, Drivers N=17 No LDA/QDA, Lab Classification
in driving task 75.02%
Respiratory, ECG KNN, SVM
LR, DT, Accuracy:
[107] 2018 Questionnaire Employees N=750 No Boosting, RF, Survey Naturalistic Classification 75.13%
KNN (Boosting)
Thermostat,
and Passive MOS (SF36) MSE: 0.17
[98] 2020 Questionnaire N=8 No DT, RF Naturalistic Regression
InfraRed Questionnaire (RF)
Sensors
Pre-
Multi-task
collected
ECG, HR, Skin learning with Stress-inducing AUROC: 0.91
[147] 2018 No NN, LR, SVM data Classification
Response participants as task (NN)
without
tasks in NN
experiment
HRV, EEG, Skin Polar Wearlink
Healthy Stress-inducing Accuracy:
[47] 2019 Response, N=24 No NN HRM, E243 Lab Classification
subjects task 90% (NN)
Respiratory electrodes
ECG, PPG, HRV, Shimmer 3
Healthy Stress-inducing Accuracy:
[121] 2019 Skin Response, N=30 No RF, DT, KNN ECG, Empatica Lab Classification
subjects task 84.13 (RF)
Respiratory E4
Acceleration/Body LR, Empatica E4,
Patients with Clinical score, RMSE: 2.8
[159] 2017 Movement, Skin N=12 No Adaboost, Android Naturalistic Regression
MDD self-report (Ridge)
Response, Other RF, Other Smartphones
Accuracy:
Audio-visual
HR, HRV, PPG, Healthy LR, DT, RF, Self-made PPG 91% (RF),
[86] 2020 No Lab Classification stress inducing
Questionnaire subjects SVM sensor AUROC: 0.96
stimulus
(RF)
SVM,
Accuracy:
[130] 2018 HR, ECG N=15 No Kalman Lab Classification Anxiety stimuli
71% (SVM)
Filter
Accuracy:
Geriatric NB, RF, 90%,
[108] 2017 Questionnaire N=520 No Lab Classification HADS scale
patients Other AUROC:
94.3% (RF)
Accuracy:
85%, F1
LR, NB, DT, Pre-
score= 0.78
Boosting, collected
Kenya (SVM, RF,
[142] 2021 Questionnaire N=800 No Adaboost, data Classification Survey
people Ada
XGBoost, RF, without
Boosting,
SVM experiment
Voting
Ensemble)
Pre-
People from Accuracy:
collected
tech and LR, DT, RF, 84%, F1
[160] 2020 Questionnaire No Survey data Classification Survey
non-tech KNN, SVM score: 0.87
without
companies (LR, DT)
experiment
Pre-
Child and
NN, LR, collected
Adolescent Survey, AUROC:
[161] 2020 Questionnaire N=7,638 No XGBoost, RF, data Classification Survey
Twins in Reports 0.739 (RF)
SVM without
Sweden
experiment
LR, NB, RF,
SVM,
PTSD Metricwire AUROC: 0.85
[162] 2019 Questionnaire N=90 No Ensemble Naturalistic Classification DSM-5 scores
patients Mobile app (Ensemble)
Methods,
Hard Voting
F1 score:
Veterans
[112] 2020 Questionnaire N=305 No LR, Boosting Naturalistic Classification Self report 0.69 (Voting
with PTSD
Classifier)
Pre-
NESDA LR, NB, collected Accuracy:
[163] 2021 Questionnaire cohort No AUTO- data Classification Self report 79% (auto-
participants SKLEARN without sklearn)
experiment
NN, Pre-
Boosting, collected
Accuracy:
[164] 2020 Questionnaire Students N=917 No Adaboost, WhatsApp data Classification GAD-7 score
75.4% (NN)
KNN, DT, without
SVM, RF experiment
Pre-
NN, LR, AUC: 0.73
collected MDD/GAD
Boosting, GAD, 0.67
[105] 2021 Questionnaire Students N=4184 No data Classification patients and
XGBoost, RF, MDD
without normal people
KNN, SVM (XGBoost)
experiment
Accuracy:
Chinese Stress-inducing
Sticker-Type 86.8%,
[141] 2021 ECG Academy of N=34 No NN Lab Classification task, self-
ECG Specificity:
Sciences report
0.93 (LSTM)
LDA/QDA, Stress-inducing Accuracy:
[165] 2019 HRV, ECG No Lab Classification
SVM task 82.7% (CNN)
Pre-
Only showed
collected Accuracy:
WESAD the importance WESAD Stress-inducing
[166] 2021 Skin Response N=15 No NN data Classification 92.85%
dataset of Dataset task
without (CNN)
personalization
experiment
HRV, ECG, Skin Accuracy:
[167] 2021 No DT Lab Classification Self-report
Response 79% (SVM)
One model was
trained for
Audio Signals, WESAD and SWELL and Questionnaire
each subject Accuracy:
[82] 2020 HRV, Skin SWELL N=50 No DT WESAD Naturalistic Classification scores, self-
(from ground- 95.2% (RF)
Response, ECG datasets Datasets report
up and transfer
learning)
Employees MindMedia Accuracy:
SOMs were
Skin Response, with NeXus-10 79%,
[132] 2018 N=12 No SOM Lab Classification used to
ECG, PPG, HRV reported MKII, imec Sensitivity:
generates
stress Health Patch 75.6%
Precision:
16.9%,
Pre- Reddit Self-
Recall:
NN, collected reported
Depressed 17.8%, F1
[87] 2018 Text N~9000 No Boosting, data Classification Depression
people score: 17.6%
XGBoost without Diagnosis
(increased
experiment (RSDD) dataset
comapred to
CNN)
Pre-
SVM, RF, collected
DASS AUROC: 0.98
[95] 2021 Questionnaire N=39,975 No XGBoost, data Classification DASS scores
Questionnaire (SVM)
DT, NB without
experiment
Accuracy:
PTSD XGBoost, RF, Apple Watch, 83%,
[168] 2021 HR N=99 No Naturalistic Classification Self-report
patients GLM, SVM MOTO 360 AUROC: 0.7
(XGBoost)
Eye Tracking,
Adaboost, Accuracy:
Acceleration/Body Stress-inducing
[59] 2017 Adults N=23 No KNN, SVM, Camera Lab Classification 88.32%
Movement, HR, task
NB (KNN)
PPG
Accuracy:
NN (CNN, Stress-inducing
[60] 2021 Eye Tracking Adults N=23 No Camera Lab Classification 86.1%
LSTM) task
(LSTM)

You might also like