0% found this document useful (0 votes)
2 views26 pages

Embedded Machine Learning Using Microcontrollers I

This article reviews the integration of embedded machine learning (TinyML) with microcontrollers in wearable health and care applications, highlighting the challenges and specifications of existing hardware and software tools. It emphasizes the potential of TinyML to enhance privacy, reduce power consumption, and improve usability in wearable devices. The work serves as a comprehensive guide for researchers and designers in the development of healthcare-focused TinyML systems.

Uploaded by

Rahul R
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views26 pages

Embedded Machine Learning Using Microcontrollers I

This article reviews the integration of embedded machine learning (TinyML) with microcontrollers in wearable health and care applications, highlighting the challenges and specifications of existing hardware and software tools. It emphasizes the potential of TinyML to enhance privacy, reduce power consumption, and improve usability in wearable devices. The work serves as a comprehensive guide for researchers and designers in the development of healthcare-focused TinyML systems.

Uploaded by

Rahul R
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

This article has been accepted for publication in IEEE Access.

This is the author's version which has not been fully edited and
content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2022.3206782

Date of publication xxxx 00, 0000, date of current version xxxx 00, 0000.
Digital Object Identifier 10.1109/[Link]

Embedded Machine Learning Using


Microcontrollers in Wearable and
Ambulatory Systems for Health and Care
Applications: A Review
MAHA S. DIAB, (Graduate Student Member, IEEE), and ESTHER RODRIGUEZ-VILLEGAS
Wearable Technologies Lab, Department of Electrical and Electronic Engineering, Imperial College London, London SW7 2AZ,
United Kingdom
Corresponding author: Maha S. Diab (e-mail: m.diab21@[Link]).
This work was supported in part by the European Research Council (ERC) for the NOSUDEP project grant no 724334
and in part by the Engineering and Physical Sciences Research Council (EPSRC), UK / grant agreement no.
EP/P009794/1.

ABSTRACT The use of machine learning in medical and assistive applications is receiving
significant attention thanks to the unique potential it offers to solve complex healthcare problems
for which no other solutions had been found. Particularly promising in this field is the combination
of machine learning with novel wearable devices. Machine learning models, however, suffer from
being computationally demanding, which typically has resulted on the acquired data having to be
transmitted to remote cloud servers for inference. This is not ideal from the system’s requirements
point of view. Recently, efforts to replace the cloud servers with an alternative inference device closer
to the sensing platform, has given rise to a new area of research Tiny Machine Learning (TinyML).
In this work, we investigate the different challenges and specifications trade-offs associated to
existing hardware options, as well as recently developed software tools, when trying to use
microcontroller units (MCUs) as inference devices for health and care applications. The paper
also reviews existing wearable systems incorporating MCUs for monitoring, and management, in
the context of different health and care intended uses. Overall, this work can be used as a kick-start
for embedding machine learning models on MCUs, focusing on healthcare wearables.

INDEX TERMS Edge ML, embedded machine learning, healthcare, microcontroller, TinyML,
wearable.

I. INTRODUCTION different challenges along the way. These challenges are


EARABLE devices have witnessed a notable summarized in Fig. 1. Although there is an evident
W increase in their use the past years. Recent
statistics show the worldwide shipment of wearable units
increase in the use of wearables, this reflects the accep-
tance of only part of the population which voluntarily
reached 533.6 million in 2021, a 20% growth from pre- uses them. But, when wearables are intended to be
vious year [1]. These devices include smart watches, fit- used for monitoring health status and providing care
ness trackers, hearables, and others. This users’ interest and assistance, users might not show the same enthusi-
in wearables trend together with advances in machine asm, especially the elderly or disabled. Therefore, user
learning have led to researchers starting to look into acceptability is a major challenge that must be taken
combining the benefits of the two- i.e., usability and into account during the system design. Devices should,
high ability to extract information of interest- within for example, be compact, easy to use, comfortable to
the context of health and care applications. wear and have minimum maintenance [2]. Moreover,
Designing wearable devices for these applications, continuous long-term use of wearable devices can result
however, is not a straightforward process. As techno- in large amounts of data. This data will need to be
logical advances provide advantages, they also present transmitted, stored, and processed. All of this comes

VOLUME X, 2022 1

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and
content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2022.3206782

[Link] et al.: Embedded ML Using MCUs in Wearable and Ambulatory Systems for Health & Care Applications: A Review

FIGURE 1. Design considerations and challenges of wearable health and care systems.

with its own set of constraints, in terms for example algorithm will handle the collected big data, dictate the
of data storage, and power consumption. system’s performance (which in the context of health-
Another practical aspect to take into account in the care will be linked to both safety and intended use),
design of a wearable device is that this will be collecting and its implementation location will have an effect in
physiological data, and, in some cases, this might be all the above-mentioned challenges: power consumption,
linked to also personal identifiable information. This privacy and security, and indirectly usability. Therefore,
data might need to be transmitted for processing and/or a closer look on the application of machine learning
storage, which can make it susceptible to security at- algorithms in healthcare, and its integration in wearable
tacks during transmission and, in some cases, it might devices is a must.
pose a threat to the user’s privacy if the information ex-
tracted can somehow be linked to personal information. A. MACHINE LEARNING IN HEALTHCARE
Hence, securing the user’s private information, via for The application of machine learning algorithms in var-
example encryption, differential privacy, and others, is ious medical scenarios is a relatively new, but expo-
one of the challenges that require careful attention in nentially growing, field of research. Machine learning
the system design [3]. (ML), including deep learning (DL), algorithms have
Power consumption is another important challenge the ability to handle and interpret large amount of
that needs to be taken into consideration during the data, identifying patterns and trends not usually no-
design process since power will have a knock down ticeable by physicians. This has been demonstrated in
effect on aspects that can, not only impact performance, a variety of medical areas; namely, oncology [4]–[8],
but also acceptability. Power consumption cannot be pulmonology [9]–[11], cardiovascular diseases [12]–[15],
looked up in isolation, and neither can the many dif- neurology [16]–[18], orthopaedics [19], [20], histopathol-
ferent design factors that will have an effect on it. ogy [21], [22], and others. Different types of medical data
These range from the different components of the sys- that have been successfully used in the context of ML for
tem to the chosen algorithm and its distribution in healthcare applications include [23]: (1) medical images,
different system’s parts, or the decisions on data man- such as X-rays, MRI, CT scans; (2) physiological signals,
agement/transmission/storage. None of these should be such as electrocardiogram (ECG), electroencephalogram
optimized individually, since in the context of severely (EEG), electromyography (EMG), and Photoplethys-
constrained wearables, optimization in one aspect is mogram (PPG); (3) omics data such as genomics, tran-
most likely going to lead to critical usability bottlenecks. scriptomics, and others [24]; and (4) electronic health
When translating this to usability, the challenge is to records. These various medical data are used in health-
implement a system with an overall power consumption care applications including diagnosis, prognosis, mon-
that leads to an acceptable form factor and reducing itoring of disease and physical fitness, as well as in
the charging cycles to a point that this will not have a body-machine interfaces for ambient assisted living and
significant effect on users’ adherence, where the meaning patient rehabilitation.
of significant is determined by the specific intended use. The incorporation of machine learning (ML) into
Lastly, an additional challenge in the design of ML healthcare applications has given rise to three different
based healthcare wearables is the choice of ML al- architectural levels of algorithm computation: cloud,
gorithm and its actual physical implementation. The fog, and edge [25]. This is shown in Fig.2. The data
2 VOLUME X, 2022

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and
content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2022.3206782

[Link] et al.: Embedded ML Using MCUs in Wearable and Ambulatory Systems for Health & Care Applications: A Review

FIGURE 2. Wearable healthcare system framework demonstrating the three architectural levels for computation of machine learning algorithms: edge
computation runs the ML algorithm close to sensor level; fog computation receives sensor data for processing on local network; cloud computation receives sensor
data for processing at remote servers.

collected from different wearable sensors are sent for controller units (MCUs) [28]–[32]. GPUs provide high
processing at one of the three computational levels. computational ability by utilizing parallel processing.
First, computation at the cloud level depends on cloud This enables fast inference but at the expense of high-
services. This is generally needed as a result of high power consumption [30]. FPGAs consume less power
computational and data storage demands. Sending the and can fit ML algorithms with an acceptable perfor-
data to remote servers, although advantageous in many mance and programmability on hardware, but are less
aspects, also poses some potential drawbacks, such as area and energy efficient than ASICs [33] (and this, in
the need of network connectivity, privacy and security turn can affect the size of the overall healthcare device).
increased risks, and, sometimes, power consumption ASICs, however, have huge development costs, which
limiting usability. Fog computation, similarly, to cloud are not always affordable when taking into account the
computation, requires sending medical data for process- commercial route to market of the devices. MCUs, on
ing. However, the server is a local one, such as a PC on the other hand, although also requiring higher power
the local network. Edge computing, of ML algorithms, consumption than ASICs, can be more area and cost
on the other hand, performs computation close to the efficient than FPGAs. In addition, they are also easier to
sensor level using the collected data. This eliminates reprogram requiring embedded C rather than specialist
the need for transmission of raw data, increasing privacy knowledge in VHDL; this allowing for quick updates.
and security, and allowing interpretability directly in the The performance of ML algorithms on MCUs is depen-
device’s front end [26]. But it is restricted by the edge dent on multiple variables, related to both the MCUs
device computational capability. Ultimately, the choice specification, and the specific application. In general,
of architecture depends on the application, device at weighting the pros and cons of the hardware options,
hand, and the overall system features and specifications. MCUs are the most appealing option for use in wear-
In order to utilize the advantages of edge computation, able devices, because of their low power consumption,
several research works addressed the implementation of latency, size, flexibility, and cost. Therefore, this review
machine learning algorithms on edge, using different explores the use of MCUs as edge device for inferences
edge devices, algorithms, and deployment tools, aiming in wearable devices for healthcare applications.
to achieve as close as possible performance to what
would be obtained with the algorithms running on a An example design for a future health and care
PC. This has been referred to as edge inference, Edge wearable device using TinyML technology is presented
ML, and TinyML in the case of using microcontrollers in Fig. 3. Data collected from wearable sensors and
[27]. related public datasets are used for building the ML
model. The process of building the ML model follows
Different hardware options are available for Edge the conventional process requiring pre-processing, data
ML implementation, namely graphic processing units splitting for training followed by evaluation of the cho-
(GPUs), field programmable gate arrays (FPGAs), ap- sen model. Development of the model is then achieved
plication specific integrated circuits (ASICs), and micro- through error calculations on validation set and chang-
VOLUME X, 2022 3

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and
content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2022.3206782

[Link] et al.: Embedded ML Using MCUs in Wearable and Ambulatory Systems for Health & Care Applications: A Review

FIGURE 3. Design flow of wearable health and care device as a TinyML technology.

ing of hyperparameters as needed. Once a final optimum This work fills the gap found in literature in terms
model is reached, the actual work related to TinyML of review and analysis of the use of TinyML in wear-
begins. The model is optimized through different com- able healthcare systems, covering also the existing tools
pression techniques, and then converted to an MCU needed for implementation of TinyML based systems,
compatible format. The process of model optimization both hardware and software. This includes the different
through different software tools and frameworks is later types of MCUs used thus far analyzing their different
covered within the paper. Once the model is converted, specifications and design limitations; and the software
it is deployed on the microcontroller for inference using tools and frameworks used in optimization and deploy-
direct data from the wearable sensor to output the ment of ML model on MCU, also analyzing their use
inference result. and compatibility with specific hardware. The paper
also covers how these hardware and software tools have
B. RESEARCH GAP AND PAPER CONTRIBUTION been used in the design of the reported healthcare aimed
A gap was identified within the literature in relation TinyML systems. Overall, this work can serve as a guide
to the use of TinyML technology in healthcare systems. for researchers and system designers when trying to
Although there have been other reviews targeting in one make system architecture decisions in the initial phases
way or another the use of ML/DL algorithms, a number of the design process.
of those reviews focused on the specific application of
these algorithms categorizing their use based on: type C. RESEARCH METHODOLOGY AND PAPER
of medical data used [23], medical application [34], ORGANIZATION
[35], diagnosis from medical images [36], and from the Peer-reviewed articles written in English reporting
privacy and security point of view [37]. Other reviews TinyML wearable systems for healthcare applications
targeted the edge implementation of ML/DL algorithms were included in this review. The articles were searched
reviewing the different edge hardware [26], [28], [29], for using Google Scholar and IEEExplore. Further arti-
[32], [33]. Some focused on a specific platform such as cles and data repositories were obtained from the cita-
FPGAs [28], [38], or ASICs [28], while others focused tions of found articles. Search terms used included: edge
on the application [39], [40]. However, even within the machine learning, embedded machine learning, TinyML,
reviews of Edge ML in biomedical applications, the inference on microcontroller, machine learning on micro-
use of MCUs as the edge device was hardly addressed, controllers combined with health and care (healthcare)
and the focus was instead on FPGAs, ASICs, mobile applications, wearable devices/systems. 1240 records
phones, or microprocessors such as the Raspberry Pi. were found in this search. Article titles and abstracts
As for reviews targeting MCU implementations, they were used for initial screening to identify articles pre-
provided a general overview of the TinyML field without senting health and care systems with embedded ML
a specific application [41]. To the author’s knowledge, algorithms. This led to the exclusion of 1207 records.
there has not been any work discussing the full pipeline Selected articles (33 studies) were further screened based
for TinyML technology applied on wearable health and on eligibility criteria: articles with missing informa-
care devices, despite the fact that TinyML does indeed tion related to on-board performance metrics, algorithm
provide several advantages over cloud or fog computing training and deployment, not using MCUs, and are non-
which could be capitalized on, depending on the specific wearable systems were excluded. This resulted in further
application. This work addresses this gap. exclusion of 16 studies. The remaining 17 studies were
4 VOLUME X, 2022

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and
content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2022.3206782

[Link] et al.: Embedded ML Using MCUs in Wearable and Ambulatory Systems for Health & Care Applications: A Review

included in the review. development which results in few commercially available


Data extraction was undertaken on eligible articles hardware implementations using it. Two RISC-V based
using a customized data extraction form. Information MCUs, [Link] (research SoC) [43], and GAP8 (com-
related to the system application, components, sensor mercially available by GreenWaves) [44] have been used
placement, dataset used for training, ML algorithm and for TinyML devices to achieve low power consumption
architecture, software tools for deployment, choice of [45]–[47]. The choice of ISA alone does not dictate the
MCU, and performance achieved for each TinyML sys- overall performance of the processor, as the hardware
tem was extracted; together with performance metrics integration/architecture and available peripherals play
focusing on memory usage, time per inference, and important role too. Therefore, MCUs with similar core
power per inference. processor ISA can still perform differently and other
The rest of the paper is organized as follows. Section characteristics of the MCU should be investigated before
II will cover features and specifications of the most a decision.
commonly used MCUs in healthcare devices for on- Another key aspect when choosing a microcontroller
device inference. Then, software tools including frame- for wearable devices is the power consumption since this
works and libraries used in deployment of ML algorithms plays a vital role in usability (linked to both size and de-
onto MCU are discussed in section III. The various vice maintenance). This on its turn is linked to minimum
wearable systems using on-device inference are reviewed accuracy of inference required by the application, as it
afterwards in section IV. Finally, a discussion of the is going to be one of the factors determining the infor-
reviewed MCU options, software tools, and their ap- mation loss that can be afforded when transferring ML
plication in the implementation of TinyML healthcare algorithms from high computational servers to the edge.
wearable devices is presented in section V followed by The latter is also going to be affected by the memory
concluding remarks in section VI. constraints of the chosen MCU. The available memory
in an MCU both Flash and RAM generates further
II. MICROCONTROLLERS AS EDGE INFERENCE constraints on the size and architecture of the deployed
DEVICE ML models. There should be enough storage for both
Although using MCUs for inference in the context of collected data and ML model parameters (weights and
wearable healthcare devices has clear advantages, there activation functions), as well as enough RAM to process
are also important hardware limitations that need to be data and run inference. Speed of the processor is also a
considered during design and deployment of ML algo- factor to take into account, which can be crucial in the
rithms. Those limitations are linked to the specification context of certain applications since it will determine
of the MCU. Hence, aspects that need to be considered how fast an inference and corresponding decision can
when choosing the latter include the type of processor, be made. And the size is also important because wear-
available memory (both RAM and flash), speed, power, able devices are heavily volume and area constrained.
and size. Overall, considering the limitations of MCUs, it is im-
The type of processor used in a chosen MCU deter- portant in the design process to carefully establish the
mines the performance of the board to a certain level. Its acceptable performance trade-offs as a function of all
architecture and bus width relay information about how these different variables, taking into account the device
instructions for certain functions are executed and the intended use. Table 1 summarizes the specifications of
number of cycles required. A 32-bit processor can handle the different ARM and AVR MCUs that have been used
more data in comparison to an 8-bit processor at any in the context of wearable healthcare devices, including
specific time, thus requiring fewer number of instruction type of processor, its speed, available memory Flash and
cycles. As for the instruction set architecture (ISA) of RAM, as well as operating voltage range with current
the processor, within the reviewed papers, three main consumption when MCU is in run mode.
ISAs were utilized in the chosen MCU boards, namely: From the table, it can be seen how ARM is the most
Advanced RISC Machine (ARM), Advanced Virtual dominant provider for various Vendors, including STMi-
RISC (AVR), and RISC-V. The three architectures are croelectronics [48]–[52], and Nordic Semiconductors [53],
somewhat similar, being based on a reduced instruction [54]. ARM provides low power Cortex-M processors suit-
set computer (RISC) architecture. The major difference able for use in embedding neural network architectures
to highlight is that both AVR and ARM (commonly for on-device inference [57]. Within the different Cortex-
used) are license-based architectures, requiring mem- M processors, Cortex-M4, and Cortex-M4F (a version
bership from vendors for hardware incorporation into with an addition of floating-point unit (FPU)) are
their designed SoC. While RISC-V- a fairly new ISA- is mostly used in the realization of edge inference systems
an open-source instruction set that provides designers having been the core of a number of reported wear-
freedom and flexibility of use; this leading to an increas- able medical systems [45], [58]–[60]. Another Cortex-
ing level of interest from the research community [42]. M processor that has also been used is the Cortex-
However, due to its recent emergence it is still under M7(F), which provides higher performance capabilities
VOLUME X, 2022 5

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and
content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2022.3206782

[Link] et al.: Embedded ML Using MCUs in Wearable and Ambulatory Systems for Health & Care Applications: A Review

TABLE 1. Summary of specifications for ARM and AVR microcontrollers used for edge inference in healthcare applications.

Processor Core (F in the


Microcontroller Clock Speed Flash RAM Voltage range Current
name refers to floating point)
STM32L476JG/RG [48] Arm 32-bit Cortex-M4F 80 MHz 1 MB 128 KB 1.71 V - 3.6 V 117 µA/MHz
STM32L475VG [49] Arm 32-bit Cortex-M4F 80 MHz 1 MB 128 KB 1.71 V - 3.6 V 117 µA/MHz
STM32F303RE [50] Arm 32-bit Cortex-M4F 72 MHz 512 KB 64 KB 2.0 V - 3.6 V 379 µA/MHz
STM32F756ZG [51] Arm 32-bit Cortex-M7F 216 MHz 1 MB 340 KB 1.7 V - 3.6 V 602 µA/MHz
STM32F769NI [52] Arm 32-bit Cortex-M7F 216 MHz 2 MB 532 KB 1.7 V - 3.6 V 546 µA/MHz
nRF52832 [53] Arm 32-bit Cortex-M4F 64 MHz 512 KB 64 KB 1.7 V - 3.6 V 58 µA/MHz
nRF52840 [54] Arm 32-bit Cortex-M4F 64 MHz 1 MB 256 KB 1.7 V - 5.5 V 52 µA/MHz
ATmega2560 [55] 8-bit AVR 16 MHz 256 KB 8 KB 2.7 V - 5.5 V 875 µA/MHz
ATmega328p [56] 8-bit AVR 8 MHz 32 KB 2 KB 2.7 V - 5.5 V 437 µA/MHz

at cost of power consumption when compared to Cortex- are a feasible choice.


M4(F) processors. As for the AVR processors (ATmega) In order to deploy machine learning algorithms onto
with 8-bit bus width by Microchip [61], they are mostly MCUs, the model needs to be compressed to a size
incorporated in Arduino boards and Tiny circuits. that fits the memory limitation of the chosen MCU.
In regard to the RISC-V microcontrollers, the GAP8 Two main methods for compressing a model are quan-
by GreenWaves Technology [44] and Mr. Wolf [43] were tization and pruning [63], [64]. Quantization converts
used in TinyML healthcare applications. They are both 32-bit floating point values of weight and activation
based on the parallel ultra-low power platform (PULP) functions into a less precise representation compati-
[62], which is an open-source platform utilizing the open- ble with the MCU, an 8-bit fixed point value that
source instructions of RISC-V. The architecture is a bit will occupy less memory space. This process can be
different from that of ARM and AVR. It has two set done after training the model for optimum performance
of cores to achieve fast processing with ultra-low power. (post-quantization) to provide up to 4x smaller model.
The first is a fabric controller (FC) operating as a normal Pruning, on the other hand, eliminates neurons and
MCU for control, security and communication functions. connections in the model’s architecture that do not
The second is a compute cluster of 8 32-bit RISC-V affect the overall model performance as much as major
cores for parallel operation and computationally inten- connections do. The use of quantization and pruning
sive operation providing speedup of operation. GAP8 provide an MCU compatible model that reduces both
and Mr. Wolf share similar PULP design architectures power consumption and latency as a result of reduced
with few differences. GAP8 is based on TSMC 55nm memory usage, but at the expense of model’s accuracy.
technology, while [Link] on TSMC 40nm technology. The challenge is to find an acceptable trade-off between
The maximum clock frequency provided by GAP8 is performance and power/latency for the embedded ML
250 MHz while [Link] is 450 MHz. Moreover, GAP8 model. It is possible to use a single technique alone
does not support floating point operation as Mr. Wolf or both combined. This was demonstrated in the work
does. GAP8 integrates a CNN accelerator, the Hardware of [64], where both techniques were used to compress
Convolution Engine (HWCE) in the cluster core for a 7-layers convolutional neural network (CNN) and a
running CNN algorithms. As for the memory, the PULP ResNet-50 model. Both classification models were tested
architecture provides three levels: L1, L2, and L3 which using the CIFAR-10 dataset [65] to classify images to one
is an optional external memory. L2 provides 512 KB of of the corresponding 10 classes. The dataset composed
memory arranged in multi-banks that are accessible by of 60,000 images (6000/class) divided into six batches.
all processors. While L1 has a smaller capacity of 80 The sixth batch containing 1000 images from each class
KB for GAP8 and 64 KB for [Link] shared between was kept for testing. First, the trained model’s weights
all cores of cluster processor. The power range for GAP8 were extracted, and then pruned using different sparsity
is 3.6 µW - 75 mW operating on voltage between 0.8 V - levels. The pruned models were retrained using the
0.2 158 V, while Mr. Wolf is 72 µW - 153 mW operating same dataset to conclude with the best performing
in the range of 0.8 V and 1.1 V. pruned model. The chosen model was then quantized,
transforming the 32-bit floating point values into 8-bit
III. SOFTWARE TOOLS AND FRAMEWORKS FOR integers. The pruning of models decreased the number
EDGE INFERENCE of parameters. In the case of the CNN model, it was
Choosing the right MCU for a specific application is reduced from 0.95376 million to 0.19118 million param-
just one part of the decision in the design process of a eters. The combined techniques presented models with
TinyML healthcare device. The software tools used for smaller number of parameters represented using lower
deployment of the model onto the MCU, their features number of bits. The resulted compressed CNN model
capabilities and the compatibility of the latter with the achieved reduction in size from 15.03 MB to 0.97 MB.
existing MCUs can play a major role into which MCUs The accuracy of the model dropped from 84.17% to
6 VOLUME X, 2022

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and
content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2022.3206782

[Link] et al.: Embedded ML Using MCUs in Wearable and Ambulatory Systems for Health & Care Applications: A Review

83.73%, an acceptable 0.44% loss. While the ResNet-50 fixed point representation compatible for deployment
model presented a 1.32% decrease in accuracy for 20% onto Cortex-M MCUs [27]. The use of CMSIS-NN ker-
reduction in number of parameters. nels improves the throughput by 4.6x and the energy
In addition to quantization and pruning, the deploy- efficiency by 4.9x [68]. The library kernels are divided
ment of ML algorithms onto MCU requires a certain into two functions: NNFunctions and NNSupportFunc-
level of optimization for embedding. To achieve this, tions. The first is concerned with the implementation of
several software tools, libraries, and frameworks are NN layers, and the second with utility functions for data
used to facilitate it, some which include quantization conversion and activation functions. The kernel APIs
within their process of optimization. The ones that have are simplified to allow compatibility with multiple ML
been used within the context of the reviewed healthcare frameworks, TensorFlow, PyTorch, or Caffe.
systems are covered in the following.
C. X-CUBE-AI
A. TENSORFLOW LITE X-CUBE-AI is an artificial intelligence (AI) expansion
TensorFlow Lite (TFLite) is an open-source framework package of [Link] specific for STM32 micro-
developed by Google that enables the deployment of controllers [70]. The latter is an ecosystem which con-
machine learning models on mobile and embedded plat- verts and then optimizes pre-trained NN models for in-
forms [66]. The key features of this framework are: tegration on board. The package offers validation of NNs
its optimized ML model for on-device deployment, its on PC and MCU, as well as evaluation of performance
compatibility with several platforms including mobiles on STM32 MCU. The package can be simply added
(iOS and Android) and microcontrollers (using TFLite to the STM32CubeMX tool for use, helping optimize
for microcontrollers), and its support for multiple lan- pre-trained models and offering help in choosing most
guages including Python, C++, Objective-C, Java, and suitable STM32 board in terms of memory and compu-
Swift. TFLite models are represented in a FlatBuffers tation. For the implementation, data can be collected,
format, which provides reduced size and faster inference cleaned and processed for model training using any of
compared to the TensorFlow’s Buffer format, due to its the major frameworks (TFLite, PyTorch, MATLAB,
smaller code footprint and directly accessible data. The Keras, ONNX, and others), then X-CUBE-AI package
generation of the TFLite model can be done through can be used to automatically convert the model into a
three different methods: (1) using existing TFLite mod- computationally optimized version with optimum mem-
els from available examples; (2) designing an own model ory usage ready for integration. The package provides
through TFLite Maker; or (3) converting TF models to three main features: dimensionality through assessment
TFLite using the TFLite converter and applying quan- of model architecture and MCU needs; optimization
tization as an optimization method through the process. by conversion of the pre-trained model to C-code; and
Multiple optimization techniques are available through support of quantization for optimum performance, and
the optimization toolkit, including quantization, prun- fine tuning by allocating optimum memory usage.
ing, and clustering. Once a TFLite model is finalized it
is converted to a C source file for running on the MCU. D. MICROSOFT EDGEML
Running the generated C code on MCU requires an Apart from compression of pre-trained ML models for
MCU specific library version of TFLite, the “TFLite for deployment on MCU, a ready compact algorithm can
microcontrollers” to run/interpret the deployed model. be used for direct integration on resource-constrained
TFLite for microcontrollers was specifically designed for devices. With this notion, Microsoft developed a library
MCU deployment written in C++ 11, requiring a 32-bit of algorithms (EdgeML) that can be used for direct in-
platform. ARM Cortex-M-series based processors were ference on edge devices [71]. The algorithms are written
tested for TFLite for microcontroller; and some of the in Python using TensorFlow and PyTorch and provided
supported development boards are: Arduino Nano 33 as C++ implementation. The four algorithms provided
BLE Sense, SparkFun Edge, STM32F746 Discovery Kit, by EdgeML are:
and others [67]. Overall, the use of TFLite generated • ProtoNN: an algorithm based on k-nearest neigh-
model addresses the design constraints of embedded bor (kNN) that can be used for classification and
ML namely, power consumption, memory, latency, and regression problems. The model occupies around 2
privacy. KB of memory.
• Bonsai: a tree-based algorithm for classification
B. CMSIS-NN and regression problems with complex logarithmic
CMSIS-NN is an open-source library for ARM Cortex- based prediction. The model occupies around 2 KB
M processor cores, which maximizes the performance of of memory.
neural networks (NNs) through a collection of optimized • EMI-RNN: recurrent neural network (RNN) based
kernels producing minimum memory footprint [68], [69]. algorithm for time-series predictions with faster
It allows the conversion of floating-point models into inference than traditional RNN.
VOLUME X, 2022 7

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and
content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2022.3206782

[Link] et al.: Embedded ML Using MCUs in Wearable and Ambulatory Systems for Health & Care Applications: A Review

• Fast cells: a model for smaller and faster RNN cells, G. NEMO/DORY
including FastRNN and FastGRNN algorithms for NEMO (NEural Minimizer for pytOrch) is an open-
time-series classification in place of LSTM and source python library for compressing deep neural net-
GRU. The model size is less than 10 KB. work (DNN) models developed in PyTorch [74]. It tar-
SeeDot Embedded Learning Library (ELL) is a frame- gets deployment of DNN models on constrained MCUs
work provided by Microsoft EdgeML for deployment of specifically PULP based MCUs with different features
these algorithms to IoT devices [71]. SeeDot provides for model quantization. DORY (Deployment Oriented to
compilation of the model, and quantization from floating memoRY) is a tool for automatic deployment of DNNs
point to fixed point representation to run efficiently. on memory constrained MCU [75]. It uses constraint
programming to solve the tiling problem, maximizing
E. GAP FLOW memory usage based on constrained introduced by in-
dividual DNN layer. It provides memory management
Deployment of trained models onto RISC-V based
through three optimization tasks: loop tiling; memory
MCUs of the GAP family require specific set of neural
access optimization; and memory fragmentation. Three
network tools provided by GAPflow [72]. GAPflow is
steps are performed through DORY framework before
composed of a multiple set of tools that can auto-
deployment: (1) ONNX decoding of quantized DNN
matically intake a trained NN model and produce an
model in Open Neural Network Exchange (ONNX for-
MCU compatible algorithm for deployment and on-
mat); (2) layer analyzer, which produces an optimized
board inference. The tools used are, NNTool; AutoTiler;
code for tiling loop and calls back-end APIs for individ-
and the GCC. The NNtool translates a TFLite model
ual execution of layers; and (3) network parser, which
(quantized or unquantized) to an “AutoTiler Model”,
collects information from the architecture and allocates
which is a .c file describing the NN topology and the
memory buffers accordingly, generating a C file ready for
quantization policy of the different NN layers. The main
MCU deployment. DORY is currently only supported by
tool in GAPflow is the AutoTiler, which is responsible
the GAP8 MCU.
for optimizing the execution of convolutional layers and
minimizing access to memory, as well as use of parallel
H. FANN-ON-MCU
convolutional kernels that leverage the available multi-
FANN-on-MCU is an open-source framework for the
core clusters of GAP8 MCU. The optimized model files
implementation and deployment of optimized multi-
produced by the Autotiler tool defining the application
layer neural network based on fast artificial neural net-
code and memory allocation of constant parameters are
work (FANN) library [76]. The framework generates an
then compiled using the GCC tool. Finally, the GAP8
optimized C code for compilation on chosen platform
executable file and flash file are used to run inference on
from a pre-trained model in FANN’s format. The toolkit
GAP8 [72].
provides support for deployment on both ARM Cortex-
M and PULP based MCUs for on-board inference of
F. PULP-NN
trained FANN model in either fixed or floating point
PULP-NN is an open-source library similar to the representation. As for the FANN library, it is an open-
CMSIS-NN library designed for use with PULP based source library for implementation of artificial neural
MCUs [73]. It adopts the data flow and layout of networks (ANNs) [77]. It implements multi-layer neural
the CMSIS-NN library [69]. The library uses a set of architecture for both fully connected and sparsely con-
kernels to facilitate the deployment of deep learning nected networks in C and provides bindings to multiple
models for inference on edge devices. Quantized Neural languages such as MATLAB, python, and others. FANN
Network (QNN) models (8-bits, 4-bits, 2-bits, and 1- library optimizes the number of neurons and hidden
bit) are supported by the library using DSP extensions layers in an ANN architecture through cascade training
and multi-core architecture of the RISC-V processor [77]. A graphical interface, the FANNTool was developed
to speed up it performance. For comparison, a classi- to simplify the use of FANN library [78]. The tool allows
fication QNN model trained on CIFAR-10 dataset was changes to the model’s architecture, the activation func-
used for inference on GAP8 using PULP-NN library, tion, the weight initialization, the training method and
and 2 Cortex-M MCUs, STM32H743 (Cortex-M7), and allows monitoring of training process [78].
STM32L467 (Cortex-M4) using CMSIS-NN library [73]. The choice of a software tool for optimum deploy-
The quantized model (8-bits) had 3 convolutional layers ment is not limited to one library or framework. It is
and a 1 fully connected layer. The PULP-NN based possible to use a combination of compatible tools and
implementation resulted in speed up of inference, requir- libraries. An example using both TFLite and X-CUBE-
ing less clock cycles by a factor of 19.6x (STM32H7) AI for the implementation of “Hello World” ML model
and 30x (STM32L4) in comparison to the CMSIS-NN on NUCLEO-F74ZG board was presented in [63]. The
implementation. CNN model for the recognition of handwritten numbers
(MNIST dataset [81] of 70,000 samples for the digits (0-
8 VOLUME X, 2022

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and
content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2022.3206782

[Link] et al.: Embedded ML Using MCUs in Wearable and Ambulatory Systems for Health & Care Applications: A Review

9) was initially trained (60,000 samples for training and toNN. The standard ML models required larger memory
10,000 for testing) in TensorFlow using Keras, having > 100 KB and were more computationally intensive
a size of 7.17 MB. It was then converted into a lighter compared to Bonsai and ProtoNN. Even with additional
version (2.4 MB) using TFLite and TFLite Converter compression, 8 KB memory was required. The final
tool. Further compression using the X-CUBE-AI STM32 system using ProtoNN with optimized feature extrac-
package applied multiple actions for model reduction, tion model and a window size of 2s was implemented
including compression of weights of the fully connected using 1.4 KB of memory after pruning, with an average
layers, fusion of layers by combining two layers, and recall of 93.58% (1.3% less than standard ML). The
optimization of activation function leading to model size system required 747 ms for feature extraction before
reduction by a factor of 4 (668.97 KB). optimization and 49.3 ms for the optimized feature set,
with 20 ms for classification time.
IV. TINYML HEALTH AND CARE SYSTEMS Other works reported in literature focused on the
The use of ML in the context of health and care applica- detection of cardiac arrhythmia from a single lead ECG
tions has been so far mostly dependent on cloud and fog [45], [58]. In [58], a convolutional-recurrent neural net-
computation. The ability to perform edge inference on work (C-RNN) with gated recurrent unit (GRU) in-
MCUs has only become recently feasible due to the tech- stead of long short-term memory (LSTM) was trained
nological advances, allowing deployment of compressed using dataset from the computing in cardiology 2017
ML algorithms with minimum loss in model perfor- competition [94]. The model was trained to classify the
mance. In this work, various proof-of-concept TinyML input data into 4 output classes (normal rhythm, atrial
systems are reviewed, focusing on application in health fibrillation, noises, and other rhythms) using Keras with
and care including medical use, ambient assistant liv- TensorFlow. The trained model’s weights and activation
ing, and physical health for rehabilitation and fitness were quantized to 8-bit fixed point representation, us-
tracking. Related information for each system is summa- ing the CMSIS-NN library, for deployment on Nordic’s
rized in Tables 2-5. System components and prototype nRF52832 MCU. The model performance was tested
placement (if available) are summarized in Table 2, before and after quantization. The latter caused a drop
while datasets used in training and testing of the ML in accuracy of 0.66%, as well as 2% drop in F1 score.
algorithms are given in Table 3 with corresponding ML The final system occupied 195.6 KB of flash memory,
architecture in Table 4. Finally, a summary for the and 6.8 KB of RAM, whilst being able to generate an
embedded ML implementation- including accuracy of inference in 94.8 388 ms with an accuracy of 85.44%,
running algorithms on board, the occupied memory, and F1 score of 78%.
time and power consumption per single inference- are The reported work, also within the context of detec-
tabulated in Table 5. Most of the reviewed works re- tion of cardiac arrhythmia [45] used temporal convolu-
ported only the accuracy of the embedded algorithm. tional network (TCN) embedded on edge. The model
A proof-of-concept wearable device for Parkinson’s was trained using TensorFlow and Pytorch with the
patients was presented in [79], focusing on recovery of dataset ECG5000 [95] which is based on the BIDMC
patients from Freeze of Gait (FoG), this is a “brief, Congestive Heart Failure Database (chfdb) [96] consid-
episodic absence or marked reduction of forward pro- ering 5 output classes. The work compared the deploy-
gression of the feet despite the intention to walk” [109]. ment of the classification model onto two platforms, an
The subject-dependent device aimed to monitor the pa- ARM Cortex M4F (STM32L475), and a RISC-V PULP
tient and provided rhythmic auditory stimulation (RAS) (GAP8) using different deployment/quantization tools.
when a FoG event was detected. It was composed of The GAP8 was tested using 2 deployment methods,
four main parts: (1) wearable sensors, which are 3 tri- the GAPflow and NEMO/DORY (NEMO for quan-
axial accelerometers to be positioned at the ankle, leg, tization and DORY for deployment using PULP-NN
and torso providing a total of 9 readings; (2) a feature backend). The STM32L475 was tested under three de-
extraction model based on the time domain features for ployment methods: (1) using TFLite alone; (2) using
less computation; (3) a classification model for detection X-CUBE-AI with TFLite; and (3) using X-CUBE-AI
of a binary problem; and (4) a RAS module placed in the with Keras. Among the 5 implementations, the use of
patient’s ear for stimulating patient by metronome click- GAP8 outperformed STM32L475 in terms of memory
embedded music whenever a FoG event was detected. A footprint, time for inference, and power per inference.
16 MHz ArduinoMega (ATMega 2560) MCU with 8 KB The accuracy of the deployed quantized model was the
internal SRAM was used in the implementation. Several same across both platforms (94%) with a negligible
options were tested as classification models: standard drop of 0.2% in accuracy for GAP8 with NEMO/DORY
ML algorithms including Decision Tree (DT), Random which had the best performance overall. For the ARM
Forest (RF), AdaBoost (AB), k-Nearest Neighbor (k- Cortex implementations, the use of X-CUBE-AI with
NN), and Support Vector Machine (SVM)); and Mi- TFLite provided the best performance out of the three.
crosoft developed Edge ML algorithms, Bonsai and Pro- Comparing the best performance of both platforms, the
VOLUME X, 2022 9

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and
content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2022.3206782

[Link] et al.: Embedded ML Using MCUs in Wearable and Ambulatory Systems for Health & Care Applications: A Review

TABLE 2. Summary of TinyML healthcare systems components and device placements.

Ref Healthcare application System components Placement


Detection and recovery from FoG 3 tri-axial accelerometers, RAS module, & Arduino
[79] Ankle, leg, torso, ear
in Parkinson’s patients Mega
[58] Detection of cardiac arrhythmia nRF52832 No prototype
[45] Detection of cardiac arrhythmia GAP8 & STM32L475 No prototype
No prototype, future imple-
Continuous glucose monitor (CGM), insulin pump, sen-
[59] Prediction of blood glucose level mentation of MCU in wear-
sor band, STM32F303RE
able device
2 bipolar EEG channels (F7-T7 and F8-T8), ADS1299 No prototype but referred
EEG front-end, tri-axial accelerometer, BlueNRG- to the use of similar setup
[60] Detection of epileptic seizure
MS Bluetooth low energy (BLE) network processor, as previous work (e-Glasses
STM32L476 [80])
Patch-like form (Biowolf)
[47] Detection of epileptic seizures Biowolf ExG wearable platform [81]
on head
No prototype, future wrist-
[82] Heart rate prediction Nucleo STM32L476RG development board
worn device
Pressure sensor, 9-axis motion sensor, microphone,
ECG/EMG & bioimpedance AFE (Maxim
MAX30001), galvanic skin response (GSR) front-
[46] Stress detection Wrist bracelet
end, 120 mAh LiPo battery, dual-source energy
harvester, & two processors (Nordic nRF52832 & Mr.
Wolf)
TinyLily mini processor (ATMega328p),TinyLily
[83] Fall detection ASL2002 module with tri-axial accelerometer (Bosch Hip-level belt
BMA250), & piezo buzzer
SensorTile (STM32L476JGY MCU, 2 tri-axial ac-
[84] Fall detection celerometers, gyroscope, magnetometer, & a barome- Belt buckle
ter)
SensorTile (STM32L476JGY MCU, 2 tri-axial ac-
[85] Fall detection celerometers, gyroscope, magnetometer, & a barome- No prototype
ter)
SensorTile (STM32L476JGY MCU, 2 tri-axial ac-
[86] Fitness tracker celerometers, gyroscope, magnetometer, & a barome- Wrist band
ter)
No prototype, future wrist-
[87] Exercise tracker CLOUD-JAM L4 board (STM32L476RG)
worn device
Motor-Imagery Brain Computer
[88] STM32L475VG & STM32F756ZG No prototype
interface
s-EMG based hand gesture recog- Ring configuration around
[89] 8-channel AFE (ADS1298), GAP8
nition forearm
s-EMG based hand kinematics fin-
[90] GAP8 No prototype
ger position decoding
Conductive polymer composite (CPC) based low re-
Tactile sensing for prosthetic &
[91] sistive tactile sensor [92], AFE, IMU (MPU-9250), Hand glove
robotics
STM32F769NI discovery board

GAP8 and STM32L475, the GAP8 with NEMO/DORY using TensorFlow and Keras using the “OhioT1DM
demonstrated energy efficiency (9.91 GMAC/s/W), 23x dataset” [98], having a single LSTM layer and two dense
higher than STM32 using X-CUBE-AI with TFLite, as layers with ReLU activation before the output layer.
well as 46.8x faster inference (2.7 ms). The final trained model was then deployed on an ARM
Edge inference was also proposed in [59] in the context Cortex-M4F STM32F303RE MCU for edge inference,
of a wearable artificial pancreas systems for patients occupying 34.69 KB of flash memory and 1 KB of RAM.
with Type 1 Diabetes (T1D). The system input read- The deployment of the model onto the STM32 MCU
ings, acquired from a continuous glucose monitoring made use of the available X-CUBE-AI library for an 8-
(CGM) sensor, were used to predict blood glucose bit fixed point representation. As a regression problem,
through the use of an RNN model, based on LSTM the performance was evaluated by calculating root mean
layers. The LSTM-RNN regression model was trained square error (RMSE) and mean absolute error (MAE)

10 VOLUME X, 2022

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and
content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2022.3206782

[Link] et al.: Embedded ML Using MCUs in Wearable and Ambulatory Systems for Health & Care Applications: A Review

TABLE 3. Summary of datasets used in training and evaluation of the algorithms used in the reviewed TinyML healthcare systems.

Ref Dataset used Description Data sensors Part of dataset used Train Test Validation
Accelerometer
3 tri-axial 237 FoG events. Pa-
readings from 10
accelerometers tients 4 & 10 ex-
DAPHNet [93] patients during daily 70 % 30 % 10-folds
[79] (ankle, leg, torso) cluded (No FoG) (2
life activities (8 hrs &
@ 64 Hz classes)
20 mins)
Computing in ECG readings (30s- ALivCor device sin-
8,528 samples (4
Cardiology 2017 60s long) from 8,528 gle lead ECG @ 300 82 % 18 %(1) -
[58] classes)
Challenge [94] subjects Hz
20 hrs long ECG
recordings from single
2 ECG lead @ 250 5,000 samples (5
ECG500 [95] subject (chf07) of 10 % 90 % -
[45] Hz classes) (2)
original dataset [96],
[97]
CGM sensor
(Medtronic Enlite),
insulin pump
(Medtronic 530G or only CGM blood
Data collected from
630G), sensor band glucose readings
OhioT1DM [98] 12 T1D subjects over 64.8 % 19 % 16.2 %
[59] for physiological every 5 mins (166461
8 weeks clinical trial
data (Basis Peak samples)
fitness band or
Empatica Embrace
band)
F7-T7 & F8-T8
CHB-MIT scalp Data collected from 23 channels EEG channel readings (2 70 % 30 %(3) 10-folds
[60] EEG [99] 23 subjects. Total of electrodes, 10-20 classes)
664 edf (most are 1 bipolar montage
hr lonf, some are 2 No splitting ratio given.
@ 256 Hz F7-T7, T7-P7, F8-
CHB-MIT scalp & 4 hrs long) with Classes were given weight
T8, T8-P8 channel
[47] EEG [99] 182 annotated inverse to their occurrence
readings (2 classes)
seizures. frequency.
PPG sensor @
64 Hz & tri-axial
accelerometer @
Data collected from 32 Hz (wrist worn PPG, accelerometer
15 subjects doing 8 device Empatica readings, & golden Leave-one-subject-out (LOSO)
PPGDalia [100]
[82] different activities To- E4), & chest worn HR values. 64697 to- cross validation
tal of 37.5 hrs of data device RespiBAN tal samples (1 class)
Professional @ 700
Hz for golden HR
values
Data collected from
ECG, EMG (right
17 drives lasting for ECG, GSR readings
Stress recognition trapezius), GSR
(65-93 mins) driving only. (Number of
in automobile measured on the subsets of equal stress levels
[46] in city & highways samples not given).
drivers [97], [101] hand & foot, &
(number of subjects (3 classes)
respiration
not clear)
Data from 38 subjects
in 2 groups (23 adults 3 hip level sensors
real-
( 3.5 hrs/subj.) (2 accelerometers
Accelerometer read- time
& 15 elderly( 1.5 (ADXL345 &
SisFall [102] ings only, 4510 sam- SisFall data -
[83] hrs/subj.)) for 19 MMA8451Q))
ples. (2 classes) (2 sub-
types of activities of & a gyroscope
jects)
daily life (ADL) & 15 (ITG3200)
types of falls
Accelerometer read-
SisFall dataset 3 hip level sensors ings of 38 subjects.
SisFall Enhanced 30 subj. 8 subj.
[102] with temporal (2 accelerometers Total samples after -
[84] [103] (12 (3
annotation for 3 (ADXL345 & sub annotation not
elders) elders)
classes (fall, alert, MMA8451Q)) & given. (3 classes)
background) a gyroscope Accelerometer &
(ITG3200) gyroscope readings
SisFall Enhanced of 38 subjects. Total 20% of
[85] [103] samples after sub training
annotation not
given. (3 classes)

VOLUME X, 2022 11

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and
content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2022.3206782

[Link] et al.: Embedded ML Using MCUs in Wearable and Ambulatory Systems for Health & Care Applications: A Review

Ref Dataset used Description Data sensors Part of dataset used Train Test Validation
15 subjects repeating
3 exercises (squat, Tri-axial
curl, push-up). Data accelerometer & Total of 700 reps
Exercise for fit-
collected between tri-axial gyroscope (samples) for the 3 80 % 20 % -
[86] ness tracking [86]
exercises labelled (LSM6DSM) @ 20 exercises. (4 classes)
as not an exercise Hz
(NAE))
PPG & accelerome-
PPG & tri-axial
ter readings from 7
accelerometer
subjects. Each with 5
PPG_ACC (maxim integrated 210 recording ses- 5
series of 3 activities 2 subj. -
[87] Dataset [104] MAXREFDES100 sions. (3 classes) subj.(4)
(squats, stepper, rest-
health sensor) @
ing). Total of 17,201s
400 Hz
recorded data.
EEG recordings
for both real &
BCI2000 systems Only imagery
EEG Motor imagery motor task
using 64 recordings from 105
Movement from 109 subjects,
electrodes of 10- subjects with 21 80 % 20 % 5 folds
[88] Imagery [97], each performing 14
10 international trials per class per
[105] experiments (2 rest,
system @ 160 Hz subject. (4 classes)
& 3 reps of 4 tasks (2
real, 2 imagery))
sEMG readings
from 10 healthy 14 Delsys Trigno
10 sessions. Total 10
subjects. Each sEMG wireless
NinarPro DB6 sessions. (12x7 = 84 Sessions Sessions
subject performed electrodes on higher
[106] samples/session). (8 1-5 (5) 6-10 %
[89] 12 reps of 7 grasps. half of forearm @
classes)
Each grasp lasted for 2kHz
6s with 2s rest. 2-fold
stratified
sEMG readings from
cross
3 subjects, total of 20
validation
sessions. 6 reps of 8 8 channel AFE
20 sessions, each ses-
hand gestures. Each (ADS1298) around Sessions Sessions
20-sessions [89] sion with 8 gestures
gesture lasted 3s with middle forearm @ 4 1-10 (5) 11-20
& rest. (9 classes)
3s rest between reps kHz
& 5s rest between ges-
tures.
12 subjects (10 able-
bodied & 2 right-
2 rings of 8 ac-
hand transradial
tive double differen-
amputees). Each
tial sensors (Delsys
subject recorded
Trigno IM Wireless 3 sessions/subject;
3 sessions for 9
NinaPro DB8 EMG system) on total of 198 Session Session
movements in both 2-folds
[90] [107] the right forearm & samples/subj. (5 1&2 3
hands lasting (6s-9s)
18-Degree of Free- classes - 5 DoA (6) )
each with 3s rest.
dom (DoF) (Cyber-
Session 1 & 2 (10
glove 2) on the left
reps/movement)
hand @ 2 kHz
& session 3 (2
reps/movement)
Conductive
polymer composite
Tactile & inertial data (CPC) based Total of 340,000
collected in 5 sessions. low resistive tactile frames
Each session using 16 tactile senor [92], (samples). 320,000
ETHZ-STAG 4 1
different objects ma- AFE, inertial (from the 5 sessions) -
[91] Dataset [91] sessions session
nipulated for 40 sec measurement unit + 20,000 empty
each. (Number of sub- (IMU) (MPU- hand frames. (17
jects not mentioned) 9250) (@100 Hz for classes)
accelerometer &
gyroscope)
(1) Data with balanced classes; (2) unbalanced dataset with 2 samples for one of the classes; (3) test data had ratio of 1:300 for
ictal:no-seizure; (4) data augmentation used (oversampling) to produce balanced classes; (5) incremental training protocol; (6) the 5
DoAs are defined as linear combination of the 18 DoF.

12 VOLUME X, 2022

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and
content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2022.3206782

[Link] et al.: Embedded ML Using MCUs in Wearable and Ambulatory Systems for Health & Care Applications: A Review

TABLE 4. Tabulated summary of the architecture of used ML algorithms in the reviewed TinyML healthcare systems.

Ref ML algorithm Input ML Architecture


(128 x 45) with w = 2s @ 64 Hz; 5
Projection dimension d = 5. Tuned to binary implementation
[79] ProtoNN TD feature for each of the 9 ACC
of ProtoNN by Gupta et al. [108]
inputs
Input (256, 1) – Conv1(128, 8) – Conv2(64, 16) – Conv3(32,
32) – Conv4(16, 64) – Conv5(8, 64) – Conv6(4, 128) – Conv7(2,
(256 x 1) with w = 256 samples 128) – Global Average Pooling (128) – GRU (64) – Dense +
[58] CNN + GRU
with 50 % overlap Softmax (4).
Each Conv layer followed by average pooling (size = 2, stride
=2); all Conv layers have filter size = 5.
Input (1 x 140) – Conv(1 x 1) – TCN1 – TCN2 – TCN3 – FC
(140 x 1) with 140 samples input
[45] TCN (5). Conv layer (2 filters); TCN 1–3 (filters = 11, kernel length
(0.56 s)
= 11, and d = 1, 2, 4).
LSTM (1, 32) – Dense (1, 64) – ReLU (1, 64) – Dense (1, 32)
[59] RNN-LSTM 5 mins blood glucose readings
– ReLu (1, 32) – Dense (1, 1)
54 features from each 4s EEG
epoch with data fusion of 2 EEG Bagging method using bootstrap samples of training data for
[60] Random Forest(RF)
epochs (time separation 1/4 of av- each tree. No further information provided.
erage seizure)
SVM
General description of algorithms given with no details
RF 4 level DWT for w = 8s of about parameters. Classification output was post
[47]
Extra Trees EEG epoch processed (smoothed) using moving average of 4s
window over 3 successive classifications
AdaBoost
3 convolution blocks (2 dilated Conv, 1 strided Conv, 1 pooling
(256 x 4) with w = 8s @ 32 Hz with
[82] TEMPONet layer) output channel of each block (32, 64, 128) – FC (1).
75 % overlap
All layers use ReLU activation & Batch Normalization.
5 features from overlapping win-
[46] MLP [5 50 50 3]
dows (3 ECG and 2 GSR)
(240 x 1) with w = 3s sliding win-
Single decision tree using the standard hyperparameters val-
[83] Bonsai dow @ 20 Hz of 4 calculated values
ues.
from accelerometer readings
Input (n x 3) – FC (n x 32) – Batch Normalization (n x 32)
[84] RNN-LSTM (100 x 3) with w = 1s @ 100 Hz – dropout – LSTM1 (n x 32) – dropout – LSTM2 (n x 32) –
dropout – FC (1 x 3) – softmax. n = 100
Input (n x 6) – FC (n x 16) – Batch Normalization (n x 16) –
dropout – LSTM (n x 16) – dropout – FC (1 x 3) – softwax.
[85] RNN-LSTM (256 x 6) with w = 1.28s @ 200 Hz
n = 256. (Used weighted cross entropy loss function during
training to compensate for class imbalance)
Input (68 x 6) – Conv1 – ReLu – Conv2 – sigmoid – MaxPool-
[86] CNN-1D (68 x 6) with w = 3.4s @ 20 Hz ing - output (1 x 4).
Conv1 and Conv2 have kernel sizes = 3.
Input (n x 4) – Dense 1 (n x 32) – Batch Normalization (n x
(30 x 4) with w = 3s @ 10 Hz with 32) – LSTM1 (n x 32)- dropout 1 (n x 32)– LSTM 2 (n x 32)
[87] RNN
50 % overlap – dropout 2 (n x 32) – LSTM 3 (1 x 32) – dropout 3 (1 x 32)
– Dense 2 (1 x 3). n = 30
Input (n x 38) – Conv (n x 38 x 8) – Batch Normalization
– DepthConv (1 x n//Nf x 16) – Batch Normalization –
(n x 38) with w = 1s and 2s @ 160
Exponential linear unit (ELU) – Average Pooling – Separable
[88] EEGNet Hz but down sampled by factor 3
Conv (1 x n//Np//8 x 16) – Batch Normalization – ELU –
(160/3) for 38 EEG channels
Average Pooling – FC (4) - Softmax.
Nf = 128/3; Np = 8/3; n = w * freq
3 Convolutional blocks – 2 FC blocks – FC – Softmax.
(300 x 14) for NinaPro DB6 and Each Convolution block: (2 x (Dilated Conv – Batch Normal-
[89] TEMPONet (300 x 8) for 20-sessions; with w = ization – ReLu) – 1 x (Strided Conv – Average Pooling – Batch
150ms @ 2 kHz Normalization - ReLu)).
FC block: (FC- Batch Normalization – ReLu - Dropout).
Table 4 continues on next page

VOLUME X, 2022 13

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and
content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2022.3206782

[Link] et al.: Embedded ML Using MCUs in Wearable and Ambulatory Systems for Health & Care Applications: A Review

Ref ML algorithm Input ML Architecture


3 Convolutional blocks – 2 FC blocks (256 & 32)– FC (5).
Convolution block: (2 x (Dilated Conv – Batch Normalization
(256 x 16) with w = 128ms @ 2
[90] TEMPONet – ReLu) – 1 x (Strided Conv – Average Pooling – Batch
kHz
Normalization - ReLu)).
FC block: (FC - Batch Normalization – ReLu - Dropout).
Modified ResNet-18
Conv (32 x 32 x 16) -Batch Normalization – ReLU – MaxPool
[91] CNN (32 x 32) Single tactile frame
(16 x 16 x 16) – ResNet – Dropout – ResNet (8 x 8 x 32) –
Conv (8 x 8 x 32) – Average Pooling (32) – FC (17)
w: input window size; TD: time-domain; ACC: accelerometer.

at two prediction horizon (PH), 30-min PH (MAE = minimize the reported false negatives, and finally the
13.59 ± 1.47 mg/dL and RMSE = 19.10 ± 2.04 mg/dL) training/testing of ML models on global or subject
and 60-min PH (MAE = 24.25 ± 2.8 mg/dL and RMSE specific use cases. The investigation was applied us-
= 32.61 ± 3.45 mg/dL). The reported error between ing four different supervised models: SVM, RF, Extra
inference on edge and local server was reported to be Trees, and AdaBoost, for final implementation on em-
0.0029 mg/dL and 0.0025 mg/dL for RMSE and MSE bedde physiological platform the Biowolf [81]. BioWolf
respectively. In relation to power consumption, individ- is an ExG wearable device, composed of 8-channel AFE
ual blood glucose level predictions were made within ((ADS1298) for signal acquisition, Mr. Wolf for process-
22.2 ms for every new CGM reading (every 5 minutes), ing and embedding ML algorithms, and an nRF52832
with the system being in sleep mode otherwise. This for communication. The CHB-MIT public dataset [99]
led to an ultra-low average power of 8 µW for running was used for evaluation of the different algorithms in
the algorithm on MCU. The authors did not present a different use cases. The pre-processing of the data used a
final device prototype. However, they demonstrated the 4 level DWT comparing window sizes 2s, 4s, and 8s. The
feasibility of using edge inference in predicting blood dataset had lower number of seizure epochs compared to
glucose level for future incorporation with CGM and non-seizure epochs. Therefore, in order to “penalize the
insulin pumps for diabetes care device. false alarm rate”, each class was given a weight equal to
the inverse of its occurrence frequency. The four seizure
An epileptic seizure detection wearable system was
classification models were then trained/tested under
proposed in [60]. Using only two EEG channels for signal
four scenarios based on the global/subject-specific, and
acquisition (F7-T7, F8-T8), 54 features were extracted
23 channels/4 temporal channels (F7-T7, T7-P7, F8-T8,
and data fusion was then used to enhance variability of
T8-P8). Moreover, in order to smooth-out the detection
data. The publicly available dataset “CHB-MIT” [99]
of a seizure, a post processing technique, the moving
was used for the training and testing of a subject-
average with 4s window was used. The four models using
based RF ML model. The implemented system used
the four temporal EEG channels with an 8s window size
an ADS1299 EEG AFE for data conditioning, and the
were finally deployed onto the MCU (Mr. Wolf) for per-
trained ML model was deployed on STM32L476 MCU
formance comparison. For subject-specific inference, RF,
for on-board inference. An inference was made every 4
ET, and AB achieved 100% sensitivity and specificity
seconds (length of EEG epoch) and the algorithm for
with zero false positives, while SVM reported a 99.4%
processing and classifying the readings required 27.9 ms
specificity with 2.7 false positives per hour. The RF,
per EEG epoch for both channels, consuming 7.34 mA
ET, and AB models had better performance than SVM
to run the algorithm, the power consumption was not
in terms of time and energy per inference. The reported
reported nor the energy or operating voltage. Operating
results are summarized in the given tables.
the system on a 300 mAh battery, it was reported the
system could continuously monitor epilepsy for 40.87
Heart rate monitoring system based on PPG readings
hours. The embedded ML model resulted in an average
was presented in [82]. The work used a publicly available
sensitivity of 96.6%, and specificity of 92.5% across all
dataset, the PPG-DaLiA: “a PPG dataset for motion
subjects. The presented wearable system referred to the
compensation and heart rate estimation in Daily Life
use of a previously designed e-Glasss wearable [80].
Activities” [100] to train and analyse regression models
The work in [47] presented another implementation based on temporal convolutional network (TCN), the
of a seizure detection algorithm on MCU for use in TEMPONet. Raw data from PPG sensor and tri-axial
future wearable system. The authors investigated the accelerometer (to mitigate for motion artifacts) collected
seizure detection algorithm considering different number from wrist, were fed to the regression model for heart
of factors, such as the number of EEG channels used, the rate estimation. Models were trained using Python 3.6,
pre-processing window size for discrete wavelet trans- TensorFlow 1.14. The authors used an automatic Neural
form(DWT), post-processing of classification results to Architecture Search (NAS) tool starting with a single
14 VOLUME X, 2022

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and
content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2022.3206782

[Link] et al.: Embedded ML Using MCUs in Wearable and Ambulatory Systems for Health & Care Applications: A Review

TABLE 5. Performance summary of reviewed TinyML healthcare systems.

Flash RAM Time/inf. Power/inf.


Ref Model SW tool Accuracy MCU
(KB) (KB) (ms) (mW)
Arduino Mega
ProtoNN Edge ML 93.58 % 1.4 NG 20 NG
[79] (ATMega 2560)

CNN+GRU CMSIS-NN 85.4 % nRF52832 195.6 6.8 94.8 20.65


[58]
NEMO/DORY 93.8 % 26.63 NG 2.7 38.52
GAP8 (2)
TFLite+GAPflow 94.0 % 25.03 NG 20.2 41.67
TCN TFLite 94.0 % 35.86 NG 188 45.54
[45]
[Link]+TFLite 94.0 % STM32L475 35.86 NG 126.5 42.75
[Link]+Keras(1) 94.2 % 113.4 NG 374.3 47.68
13.59 ± 1.47
RNN X-CUBE-AI STM32F303RE 34.69 1.0 22.22 0.008
[59] mg/dL (3)
7.34 mA
RF NA 94.5 % (4) STM32L476 NG NG 27.9
[60] (7)

SVM NA 99.6 % (4) NG NG 0.2 25.41


RF NA 100 % (4)
[Link] NG NG 0.052 21.89
[47] Extra Trees NA 100 % (4) (2)
NG NG 0.052 21.89
Adaboost NA 100 % (4) NG NG 0.057 22.44
TEMPONet 5.64 BPM(1) 160 84.0 427 12.10 (5)

(Best MCU) 7.06 BPM 94.7 24.9 376 12.10 (5)


X-CUBE-AI STM32L476RG
[82] 6.29 BPM(1) 16.5 11.5 17.1 12.28 (5)
TEMPONet
(Best Size) 7.55 BPM 8.07 7.21 19.0 12.11 (5)

NG [Link] (2) 14 NG 0.061 (5) 19.58 (5)


MLP FANN-on-MCU
[46] NG nRF52832 14 NG 0.472 (5) 10.8 (5)

TinyLily
Bonsai EdgeML 94.2 % 15 NG 18 NG
[83] (ATMega328p)

RNN(1) CMSIS-NN 94.13 % (6) STM32L476JGY NG 82 300 5 mA (7)


[84]

RNN(1) CMSIS-NN 93.5 % (6) STM32L476JGY NG 18.5 51 NG


[85]

CNN X-CUBE-AI 100 % STM32F756 66.53 16.25 80 NG


[86]

RNN(1) X-CUBE-AI 95.54 % STM32L476RG 94.7 12.03 150.1 NG


[87]
STM32L475VG
100.84 42.44
(@ 80 MHz)
EEGNet(1) STM32F756ZG
54.99 131.41
(w = 1s) 62.51 % (@ 80 MHz) 6.61 70.27
[88] X-CUBE-AI STM32F756ZG
20.4 412.76
(@ 216 MHz)
EEGNet(1) STM32F756ZG
64.76 % 7.12 139.2 43.81 413.06
(w = 2s) (@ 216 MHz)
93.3 % (HG) GAP8 (@ 170
TEMPONet PULP-NN 460 NG 12.84 70 (5)
[89] 61.0 % (G) MHz)

TEMPONet NEMO/DORY 6.89° GAP8 (2) 70.9 NG 4.76 51


[90]

CNN X-CUBE-AI 77.84 % (8) STM32F769NI 177 52 100 NG


[91]

NG: not given; NA: not applicable HG: hand gesture; G: grasp; (1) floating point representation (un-quantized deployment); (2) clock
frequency set to 100 MHz; (3) MAE for PH = 30 minutes; (4) geometric mean of sensitivity & specificity for subject-specific
classification; (5) calculated values; (6) calculated average of accuracy for the 3 classes; (7) operating voltage not reported; (8) top-3
inter-session accuracy.

VOLUME X, 2022 15

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and
content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2022.3206782

[Link] et al.: Embedded ML Using MCUs in Wearable and Ambulatory Systems for Health & Care Applications: A Review

“seed” (TEMPONet) to produce a family of models place of cloud inference. Specific intended uses included
for analysis. The MorphNet [110] was chosen as the fall detection in elderly [83]–[85]. The fall detection
NAS algorithm to help explore the search space for system developed in [83] provided an offline real-time
models with reduced complexity. The MorphNet reduces wearable hip-level belt with embedded ML for inference.
either model memory footprint, its number of Multiply- It used the Bonsai algorithm trained in TensorFlow
and-Accumulate (MAC) operations, or both together using the SisFall dataset [102], and the SeeDot library
through optimizing the number of channels in each layer. for model quantization and deployment on MCU. The
Three models were highlighted for analysis within the SisFall dataset is not labeled, therefore authors had to
work namely, BestMAE, BestSize, BestMCU. BestMAE manually label events into fall or activities of daily living
model achieved lowest MAE of 5.30 beats per minutes (ADL). The original data was recorded for 10-15 seconds
(BPM) with 232k trainable parameters which exceeded long, which can not be evaluated on MCU. Therefore,
available memory. BestSize reduced the number of pa- the authors focused on training the model using the
rameters down to 5k with acceptable increase of MAE to actual event at center of the input window. Choosing the
6.29 BPM. As for the BestMCU, the model presented a right frame of the event (fall or ADL) was done based
compromise between the first two to fit on the resource on maximum total variance of acceleration and was
constrained MCU, achieving MAE of 5.64 BPM using validated by recorded videos. The model was deployed
41.7k parameters. Both BestSize and BestMCU models on a TinyLily processor incorporating the ATMega328p
were deployed on STM32L476RG development board for MCU chosen for its sewable compact design (coin size).
performance evaluation in both 32-bit floating point and A compatible TinyLily ASL2002 module with tri-axial
8-bit quantized format. Performance results are reported Bosch BMA250 accelerometer provided input for infer-
in Table 5 achieving an average power consumption of ence and a piezo buzzer was used for acoustic output
12.1 mW per inference. when fall is detected. The prototype sewed the three
In addition to the above mentioned works focusing on components into the wearable belt. The inference model
systems in the context of physical health, work has also on MCU was tested on real-time acquired data from two
been reported in the context of mental health. A wear- people doing 9 different activities, 3 of which were fall.
able smart bracelet, named InfiniWolf, for stress detec- The Bonsai model was tested using different sampling
tion was presented in [46]. InfiniWolf (size of a coin) was frequencies, window sizes, and number of features taking
assembled to contain two MCUs, the Nordic nRF52832 into account memory constraint and model performance.
and Mr. Wolf, a pressure sensor, a 9-axis motion sensor, The best performing model reported achieved an accu-
a microphone, an ECG/EMG and bioimpedance analog racy of 94.2% for window size of 3s using 4 features
front-end (AFE) (Maxim MAX30001), a low power extracted from absolute acceleration and variance of the
galvanic skin response (GSR) front-end, and 120 mAh tri-axial accelerometer at a sampling frequency of 20 Hz.
LiPo battery. The bracelet also provided two energy The two other systems reported for fall detection
harvesting sources based on thermal and solar energy, [84], [85] chose a deep learning model for inference.
which eliminated the need for recharging of the LiPo Training using RNN was carried out with an enhanced
battery. The system was tested in an application for version of the SisFall dataset [103], which was manually
stress detection based on collected readings from ECG labeled to provide 3 output classes (fall, alert, and back-
and GSR, implementing a multi-layer perceptron (MLP) ground). The work in [84] used an LSTM based RNN
using fast artificial neural network (FANN) library [76]. deployed on a STMicroelectonic device, the SensorTile.
The model was given five features as input to classify The latter has a STM32L476JGY MCU operating at
into one of the three output classes (stress, medium 80 MHz. Beside the MCU, the tile includes two tri-
stress, no stress). The estimated size of the model was axial accelerometers, gyroscope, magnetometer, and a
reported to be 14 KB. The trained model was deployed barometer. Initially, the LSTM model was trained using
on both MCUs to demonstrate the energy efficiency of TensorFlow, achieving an accuracy of 98% for fall de-
Mr. Wolf over nRF52832. It was shown that Mr. Wolf tection, 90.93% for Alert, and 93.12% for background.
performs better. Hence final inference was done on Mr. The model was then deployed on the MCU without
Wolf, while the nRF52832 was used for the purpose of quantization, making use of the CMSIS-library of the
communication, monitoring of battery level, and sup- ARM Cortex-M4 MCUs occupying 82 KB of the total
port of data processing when needed. Mr. Wolf used in 128 KB memory, and providing an inference within
parallel computation using its 8 RI5CY cores provided 300 ms for 1 second input window sampled at 100
a 4.9x speedup compared to a Cortex-M4F MCU during Hz. Although the use of floating point representation
an inference, which required 6126 cycles at frequency of of the LSTM model provided the same classification
100 MHz (equivalent to 61.26 µs) consuming 1.2 µJ of performance as on workstation (in terms of accuracy,
energy (around 20 mW). sensitivity, and specificity), it was at the cost of higher
Wearable systems for ambient assisted living (AAL) memory occupancy and longer time for inference. The
have also been reported employing edge inference in calculated consumed current for operating the 32-bit
16 VOLUME X, 2022

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and
content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2022.3206782

[Link] et al.: Embedded ML Using MCUs in Wearable and Ambulatory Systems for Health & Care Applications: A Review

floating point representation of the trained LSTM model The NAE was used as an indication for start/end of
reached 5 mA. This allows 20 hours of operation using an exercise, which helped in counting the number of
100 mAh battery as reported. This number of hours repetitions. A window size of 3.4s (68 samples) was
would unavoidably be reduced as other indispensable used as input with 6 readings (3 accelerometer, 3 gy-
electronic blocks were accounted in the system architec- roscope). A total of 700 repetitions were collected and
ture. split into 80% and 20% for training and testing. The
As for the fall detection presented in [85], the work in- trained CNN was tested in Keras at first (50 cases)
vestigated the implementation of the best classification reporting 100% accuracy. Then using X-CUBE-AI, the
model on three levels: (1) the inputs (3D accelerometer, model was translated into MCU compatible format for
3D gyroscope, or both); (2) the number of LSTM layers another testing. Testing on the MCU was done using a
(1 or 2); and (3) the size of inner cells (4, 8, 16, 32). total of 70 repetitions which were all correctly classified.
The use of accelerometer readings alone showed better However, the testing set was imbalanced having 6 curls,
performance than the gyroscope alone, while the use of 15 push-ups, 13 squats, and 36 NAE. Having more NAE
both provided a noticeable increase in accuracy. The events was expected as resting between each exercise
number of LSTM layers (1, 2) had very little influence on was classified as NAE. The model on MCU was also
the results. Therefore, a single layer was chosen taking compressed by factor of 8 reporting no effect on model
into consideration memory needs. As for the cell size, accuracy. Results of embedding are reported in tables.
both 16 and 32 presented nearly similar results for best Another embedded system for exercise recognition
performance. Therefore, both inputs were used and the was presented in [87]. Data collected from PPG and
RNN with single LSTM layer and 16 cells was trained tri-axial accelerometer (to compensate for motion arti-
in TensorFlow. To minimize the effect of imbalanced facts) were fed to a RNN model for classification. The
classes, the authors used a weighted cross-entropy loss RNN model was trained and tested using the publicly
function. Each window contributing to the gradient is available PPG-ACC dataset [104]. Data records of 5
weighted by the inverse of its class size in training set. out of the total 7 subjects were used for training, while
For the embedded design, the SensorTile was also chosen data for the remaining 2 subjects were used for testing
with an ARM Cortex M4F MCU, and the RNN model of the model. Data pre-processing including cleaning
was deployed using CMSIS library without quantization and normalization, data down-sampling using simple
keeping its 32 floating point format. Hence, the clas- decimation, and data augmentation, were all applied on
sification performance of the model was very close to the training data. The training data had an imbalanced
that of the workstation with a reported mean squared record for the three classes (squats, stepper, and rest-
numerical error in the order of 10−7 . The embedded ing). Hence, authors used data augmentation, specifi-
model achieved an average accuracy of 93.52% across the cally oversampling method (duplicating records of less
three classes, occupying less than 18.5 KB of memory occurring classes) to balance the classes for improved
and requiring a processing time of 51 ms/window. It training process. The performance of the classification
was also reported that a wearable device with minimal model after down sampling using different decimation
architecture is estimated to run for 132 hours using a factors was reported. The best performing model was
100 mAh battery. Moreover, an expected decrease in achieved using decimation factor of 40 i.e., sampling
memory usage by 2.5 KB and of 1 ms in inference frequency of 10 Hz (original sampling frequency was
time was reported to be possible if only accelerometer 400 Hz) having 30 samples per window size of 3s.
readings were used, translating into 4 hours increase in Information related to implementation and performance
battery life. of final model on STM32L476RG are summarized in the
Other application of embedded ML in healthcare was given tables.
demonstrated in fitness tracking. In [86], a wrist worn In the context of assistive devices for rehabilitation,
fitness tracker was designed to monitor the type of multiple works have been proposed. The authors in [88]
exercise and number of its repetitions. Three exercises presented a wearable device to monitor EEG signals for
were targeted (squat, curl, and push-ups). The wearable motor imagery commands. A compact CNN algorithm,
was designed using the STMicroelectronics SensorTile for use in EEG based brain computer interface (BCI)
module, which incorporates the STM32L476 MCU, a application, the EEGNet [111] was used. The EEGNet
Bluetooth low energy (BLE), LSM6DSM 581 (tri-axial model was trained and validated on the Physionet EEG
accelerometer and tri-axial gyroscope) and other unused Motor Movement/Imagery dataset [97], [105]. From the
sensors. The device was used to collect real-time data for original dataset, only motor imagery (MI) recordings of
use in training and testing of an ML model. As summa- 105 subjects were used, each subject with 21 trials per
rized in Table 3, data was collected from 15 subjects at class for each of the 4 motor imagery classes: left fist
a sampling frequency of 20 Hz. The collected data was (L), right fist (R), both feet (F), and rest (0). The model
classified into four classes, having the fourth as not an was trained and tested using 5-fold cross validation for
exercise (NAE) for resting or the time between exercise. a global model and 4-fold cross validation for a subject
VOLUME X, 2022 17

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and
content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2022.3206782

[Link] et al.: Embedded ML Using MCUs in Wearable and Ambulatory Systems for Health & Care Applications: A Review

specific transfer learning (SS-TL) model with an input placed around the forearm and GAP8 MCU for hand
window of 3s. The subject specific model made use of gesture recognition. A novel deep learning model based
transfer learning from the global model. Initially, the on TCN topology, Temporal Embedded Muscular Pro-
global model is trained, then each subject within the test cessing Online Network (TEMPONet) was trained and
set has its trials split for a 4-fold cross validation and tested on two datasets. The first was a publicly available
the model is then trained. Both presented models were dataset, the Non-Invasive Adaptive hand Prosthetics
compared to the baseline state of the art CNN model Database 6 (NinaPro DB6) [106]. The second was their
presented in [112] using same datasets, validation meth- own 20-sessions dataset collected using the wearable
ods, and input size window. Performance under three prototype. Relevant information of both datasets are
cases was checked: 2 classes (L/R), 3 classes (L/R/0), summarized in Table 3. Incremental training protocol
and 4 classes (L/R/F/0). The global EEGNet model was used, splitting the first half of the datasets for
performed better than global CNN model in all three training, and the second for testing, with 2-fold stratified
cases, with an increase in accuracy of 2.05%. 5.25%, cross validation. The resulting average accuracy of test-
and 6.49% for the 2, 3, 4 classes cases respectively. ing sessions are reported in Table 5. It was noticed the
As for the SS-TL model, compared to the CNN SS- performance of TEMPONet on the NinaPro DB6 was
TL model, the 2-classes model presented a decrease much lower than the 20-sessions dataset. This was due
in accuracy of 2.17%, while an increase by 0.82% and to the fact the NinaPro dataset differentiates between
2.32% for 3, and 4-class models. While comparison different grasps, while the 20-session datasets between
between global and SS-TL, the improvement provided different gestures. Nevertheless, it was reported that the
by the SS-TL over global for the EEGNet model was performance of TEMPONet on NinaPro DB6 presented
not as significant as the improvement noticed in the an increase in accuracy of 7.8% compared to the state-of-
CNN model. Therefore, the authors chose to proceed the-art models. The trained TEMPONet was quantized
implementation using their global model. The model to 8-bit representation after training using PULP-NN
trained on Keras and TensorFlow was deployed in its library for deployment on GAP8. An accuracy loss
32-bit floating point representation onto 2 STM32 MCU of 0.4% and 4.2% was reported due to quantization
boards using X-CUBE-AI package. These MCUs have in comparison to the floating-point model for the 20-
an ARM Cortex-M processor; the first with Cortex- session dataset and NinaPro DB6 dataset respectively.
M4F (STM32L475VG) for lower power operation, while While an improvement in memory footprint resulted
the other Cortex-M7 core (STM32F756 Nucleo 144) from quantization, resulting in its decrease by a factor
for higher processing capability. Due to memory con- of 4.
straints, the classification model was reduced in size Unlike the conventional approach of gesture recogni-
by investigating the effect of temporal down sampling, tion through classification of sEMG signals into lim-
reduced input window, and reduced channel numbers ited set of static predefined gestures, the work in [90]
on model accuracy. Down sampling factors (ds) of 2 approached the problem with a regression model. The
and 3 reported maximum decrease in accuracy of 0.32% target of using regression model was to decode sEMG
and 1.25% respectively. Channel reduction from original data into hand kinematic (joint angle), allowing a more
64 channels to 32, 19, and 8 channels were examined natural control which can be applied in controlling
reporting a decrease in accuracy of 0.95%, 2.66% and prosthetic. The purpose of the work was to produce a
6.52% respectively. As for window size, 2s and 1s win- regression model to fit embedded systems and maintain
dow resulted in accuracy reduction of 1.62% and 3.6% close performance to the state-of-the-art (SoA) hand
respectively. Two final models were deployed on the kinematic regression models which cannot fit embedded
two MCUs. Down sampling factor of 3 and input of 32 systems. A TCN based model, the TEMPONet was
EEG channels were used in both models. However, the trained and tested using the publicly available Non-
input window for the first model was 1s (deployed on Invasive Adaptive hand Prosthetics Database 8 (Ni-
both MCUs), while the second model used 2s window naPro DB8) dataset [107]. The dataset provided sEMG
and deployed only on the Cortex-M7 MCU due to signals for decoding of finger position to be used in
limited memory of the Cortex-M4F MCU. The trade- the estimation of hand kinematics rather than classi-
off between speed and power consumption was clearly fication of gestures. The TCN model was implemented
demonstrated in this work, where the choice of MCU using PyTorch 1.6, using raw sEMG signals as input to
depends on the system’s intended use which would result the TEMPONet and using exponential moving average
in different design criteria prioritizing either speed or method for post processing of output. The final trained
power. model deployed on GAP8 provided a MAE of 6.89°
Another proposed system focused on human-machine (0.15° better than SoA), while occupying 70.9 KB of
interaction for application in prosthetic control was memory and providing an inference within 4.76 ms
presented in [89]. The wearable system was composed of consuming 0.243 mJ per inference.
eight-electrode surface-EMG (s-EMG) AFE (ADS1298) Further application of embedded ML in prosthetic
18 VOLUME X, 2022

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and
content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2022.3206782

[Link] et al.: Embedded ML Using MCUs in Wearable and Ambulatory Systems for Health & Care Applications: A Review

was presented in the work of [91]. The proposed Edge prosthetics for rehabilitation. This demonstrated the ca-
ML system utilized embedded machine learning in the pability of utilizing TinyML for a wide range of intended
acquisition and processing of tactile information from uses. However, due to this difference in target applica-
sensor arrays arranged in hand shape. The proposed tion, it is difficult and unfair to compare systems against
embedded system named SmartHand was designed for each other. Each application targeted a different prob-
application in prosthetic and robotics that lack tactile lem, be it classification or regression, requiring the use
feedback, for addition of smart sense of touch. The work of different algorithms, architectures, hyperparameters,
reproduced a hand-shaped multi-array tactile sensor and datasets which resulted in application-dependent
presented in [92] using a scalable tactile glove (STAG) algorithms. This defined the performance of the base
for acquisition of tactile information. The presented model which then went through post-processing before
system in [91] provides three modes of operation: (i) data deployment on MCU. Hence, the model’s performance
collection; (2) real-time visualization through a graphic was initially influenced by the original floating point
user interface (GUI); and (3) SmartHand embedded base model before deployment.
system. The data collection mode was used to collect In the context of software tools and frameworks,
340,000 tactile frames (32 x 32 matrix) for 16 objects the reviewed tools can be categorized based on their
and empty hand during 5 sessions at 100 Hz with 40 MCU compatibility as shown in Fig.5. TFLite, Microsoft
seconds of data collection per object. The visualization EdgeML, and FANN-on-MCU support a wide range of
mode provided an interface showing the collected frame MCUs. TFLite limits the MCU options to 32-bit plat-
and real-time classification result if required. The Smart- forms, while EdgeML provides compressed algorithms
Hand system embedded a convolutional neural network for use on any MCU. FANN-on-MCU provides current
based on modified ResNet-18 model for the on-board support for ARM and PULP based MCUs [46]. The
classification of acquired data. The model was trained other tools are more specific, targeting a certain family
in PyTorch before deployment. Results are reported in of MCU; such as CMSIS-NN for ARM based MCUs
the given tables. The power consumption per inference [58], [84], [85], while PULP-NN and NEMO for PULP
was not reported, however the full system measured 505 based MCUs [45], [89], [90]. More specific targeted MCU
mW including the IMU and MCU supplied at 3.3 V and software tools are X-CUBE-AI for use with only STM32
a readout circuit at 5V. MCUs [45], [59], [82], [86]–[88], [91], GAPflow for GAP
The use of MCUs as the choice for edge inference MCUs, and DORY with current support for only GAP8.
as mentioned earlier is relatively new. Hence, most Due to these compatibility concerns, the choice of MCU
systems for health and care applications are based on and software tools was co-dependent.
academic research so far. However, one commercially It was also noticed the target of most of these
available system that has successfully implemented its tools/frameworks is compression of neural network
edge computation using Cortex-M4 MCU is the Amiko based architectures, while classical machine learning
Respiro [113]. Although the device is not wearable, algorithms where not supported, this resulting on them
it is worth mentioning. Amiko Respiro uses sensors not being used when classical ML algorithms were
embedded with ML algorithms for smart monitoring deployed on MCU [47], [60]. The reason behind this
of use of inhalers. It is an add-on sensor which fits exclusion could be related to the architecture complex-
to multiple commercially available inhalers. The smart ity and intensive computational requirements of the
sensor collects data related to vibrations from inhaler neural networks in comparison to classical algorithms.
during use, information related to patient’s inhalation Nevertheless, some works still used the classical ML
frequency and time, and breathing pattern providing algorithms through deployment of the compact ver-
real time feedback. Furthermore, information related sion provided by Microsoft EdgeML ProtoNN [79] and
to lung capacity and inhalation techniques and other Bonsai [83], in place of the conventional one. The use
important parameters are calculated on device using the of these compacted algorithms achieved low memory
embedded ML algorithm. Moreover, the sensor comes footprint down to 1.4 KB [79]. Concurrently, it was
with a smartphone app that patients can use to monitor noticed that AVR microcontrollers were only used with
their use of inhalers, and a professional dashboard for EdgeML algorithms due to incompatibility with other
clinicians is also available if required for remote moni- tools and frameworks. Other systems relying on ARM
toring of patient’s use. or PULP MCUs used other compatible tools. The effect
of using different tools can be observed in the work of
V. DISCUSSION [45], where the use of NEMO/DORY against TFLite
The TinyML systems covered in this review had been and GAPflow for deployment on GAP8 provided a speed
designed for a variety of health and care applications. up in inference time by a factor of 7.48, and 3.15 mW
As shown in Fig. 4, some targeted the detection and decrease in power/inference at the negligible cost of
management of medical conditions, while others focused 1.6 KB increase in Flash occupancy and a 0.2% drop
on elderly fall detection, as well as fitness tracking and in accuracy. As for Cortex-M4 deployment, the use of
VOLUME X, 2022 19

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and
content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2022.3206782

[Link] et al.: Embedded ML Using MCUs in Wearable and Ambulatory Systems for Health & Care Applications: A Review

FIGURE 4. The different application areas of TinyML systems for health and care.

X-CUBE-AI for STM32 MCU with TFLite versus the In terms of ARM processors, it was observed that only
Keras floating-point model provided memory savings by Cortex-M4 and Cortex-M7 were utilized in the TinyML
a factor of x3.16 and speed up in inference x2.95 with healthcare applications. In some cases, this was also
a drop of 4.93 mW/inference at the cost of only 0.2% likely to be the result of the factors described above,
decrease in accuracy. since a wider range of ARM processors are available
which in certain situations could also have been used.
In terms of targeted MCUs, their percentage usage More specifically, ARM provides three main processor
in Tiny ML health and care systems based on the core families, Cortex-A, Cortex-R, and Cortex-M (which is
processor is shown in Fig. 6. More than half of the dedicated for microcontrollers). Within the Cortex-M
reviewed systems used ARM based MCUs, followed by processors, other than the M4 and M7 processors, there
RISC-V MCUs, and only tenth used AVR MCUs. 8-bit is also the M0, M0+, M1, M3, M23, M33, M35 and
MCUs are more suitable for use in simpler operations M55. A comparison of the different processors and
as opposed to 32-bit MCUs which can handle more their specifications is summarized in Table 6. Based on
complex computations. This is demonstrated by the lack the ARM ISA, the processors can be divided into 3
of their use when deeper networks were deployed, as well groups: (a) Armv6-M for M0, M0+, M1; (b) Armv7-
as the lack of targeted software tools. As for the 32-bit M for M3, M4 and M7; and (c) Armv8-M (Baseline
MCUs, both ARM and RISC-V MCUs were used for and Mainline) for M23, M33, M35P, and M55. The
deployment of complex networks for on-board inference. Armv6-M processors provide less computational capa-
Although the RISC-V is fairly new, it was able to present bility compared to M4 and M7, and among the three
itself as a competitor to the well-established ARM pro- processor, only M0+ might be considered for use in
cessors. One of the main reasons for this is likely to TinyML applications. This is because, M1 processors
be the fact it is open source, allowing developers and were optimized especially for FPGA use, while M0+
vendors to use and develop it without royalty charges or is an optimized version of M0 providing better per-
license as opposed to the propriety ARM. The drawback formance, and lower energy consumption, as well as
of this, though, is the limited support that is offered in an optional Memory Protection Unit (MPU) for task
software and environment development when compared isolation. The M7, similar to the M4, has optional
to the larger support community of ARM. On that note, floating-point unit (FPU) and a digital signal processing
25% of MCUs in the reviewed work which used RISC-V (DSP) unit, but its performance is superior with the use
was noticed to be authored by a research group which of 6-stages instruction pipeline and the optional addition
is a developer of the RISC-V MCU PULP architecture. of instruction and data cache/TCM [114]. The M3 has
Therefore, the distribution in Fig. 6 can be regarded as a very close structure to the M4 but without the FPU
biased to user’s familiarity and preference rather than a and DSP extensions. In the case where neither the FPU
representation of processor capabilities.
20 VOLUME X, 2022

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and
content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2022.3206782

[Link] et al.: Embedded ML Using MCUs in Wearable and Ambulatory Systems for Health & Care Applications: A Review

FIGURE 5. Categorization of software tools and frameworks based on their hardware MCU compatibility.

Apart from the MCU core processor, the memory


limitations and clock speed of the MCU also played
important roles in the overall system performance. One
common goal pursued by all researchers was to maintain
a balance between on-board performance, in terms of
time/inference and power/inference, and memory con-
straints. The memory footprint was decided by the algo-
rithm and its architecture, as well as the post-processing
techniques used through the software tools. The effect
of doubling the input window size was reflected in an
approximate doubling of the RAM occupancy and the
time/inference, with slight increase of Flash (0.51 KB)
and an accuracy improvement of 2.25% [88]. The pur-
pose of the algorithm also affected the memory footprint
and the overall performance. This was observed in the
FIGURE 6. Percentage distribution of MCUs used in TinyML healthcare implementation of the same algorithm (TEMPONet) for
systems based on core processor type.
a 9 classes classification problem versus a regression
problem, where the classification problem was more
nor the DSP units are needed the M3 processor could complex requiring x6.5 of Flash memory and x2.7 more
be used instead of the M4 , as it occupies less area and time/inference although running at x1.7 higher clock
consumes lower dynamic power. speed than the regression problem [89], [90]. The effect
The M23 with Armv8-M Mainline ISA, and M33, of the algorithm’s architecture and input size is fur-
M35P, and M55 with Armv8-M Baseline ISA, advanced ther observed in the deployment of an RNN model on
Cortex-M processors provide an optional security exten- STM32L476 [84], [85], [87]. Although the input size was
sion for software isolation, the ARM TrustZone technol- increased in [85], the removal of one of the LSTM layers
ogy, which can be beneficial in certain TinyML based and the decrease in hyperparameter values achieved a
healthcare applications where security is key . In addi- drop of factor 4.43 in RAM footprint with negligible
tion to the software protection layer, the M35 further loss in accuracy of 0.63% and speeded up inference time
provides built-in physical protection against invasive by almost a factor of 6. An intermediate compromise
and non-invasive physical attacks. Amongst the four for the implementation of the RNN decreased the input
processors, the M23 does not provide optional FPU size but with the addition of an extra LSTM layer to
nor DSP extension, making it the smallest and lowest achieve x6.8 drop in memory requirement and two times
power one with TrustZone security. Similar to the M7, increase in inference speed [87] compared to [84]. Other
the M55 further provides optional instruction and data than the algorithm, the clock speed directly affected
cache/TCM, making it “Arm’s most AI-capable Cortex- the performance, where a factor of 2.7 increase in clock
M processor” [114], [115]. With this wide range of speed achieved a speedup of x2.69 in inference, with
options, it is clear that depending on the specific in- a cost of x3.14 increase in power [88]. The role of the
tended use, as well as performance and safety constraints processor in the overall performance was noted in the
different processors might be the most suitable design use of Cortex-M4 versus M7 with higher performance
choice. capabilities, where a speed up in time/inference of x1.83

VOLUME X, 2022 21

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and
content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2022.3206782

[Link] et al.: Embedded ML Using MCUs in Wearable and Ambulatory Systems for Health & Care Applications: A Review

TABLE 6. Comparison table for the ARM Cortex-M processors [114], [116]

Cortex- Cortex- Cortex- Cortex- Cortex- Cortex- Cortex- Cortex- Cortex- Cortex-
Feature
M0 M0+ M1 M3 M4 M7 M23 M33 M35P M55
Armv8- Armv8- Armv8- Armv8-
Armv6- Armv6- Armv6- Armv7- Armv7- Armv7-
ISA M M M M
M M M M M M
Baseline Mainline Mainline Mainline
TrustZone for
✗ ✗ ✗ ✗ ✗ ✗ ✓(option) ✓(option) ✓(option) ✓(option)
Armv8-M
HP,
SP (op- SP, DP SP (op- SP (op-
FPU ✗ ✗ ✗ ✗ ✗ SP, DP
tion) (option) tion) tion)
(option)
DSP ✗ ✗ ✗ ✗ ✓ ✓ ✗ ✓(option) ✓(option) ✓
Pipeline 3-stage 2-stage 3-stage 3-stage 3-stage 6-stage 2-stage 3-stage 3-stage 4-stage
CoreMark®/MHz 2.33 2.46 1.83 3.45 3.54 5.29 2.64 4.10 4.10 4.40
Maximum MPU
0 8 0 8 8 16 16 16 16 16
regions
Instruction cache ✗ ✗ ✗ ✗ ✗ 0-64 kB ✗ ✗ 2-16 kB 0-64 kB
Data cache ✗ ✗ ✗ ✗ ✗ 0-64 kB ✗ ✗ ✗ 0-64 kB
Instruction TCM ✗ ✗ 0-1 MB ✗ ✗ 0-16 MB ✗ ✗ ✗ 0-16 MB
Data TCM ✗ ✗ 0-1 MB ✗ ✗ 0-16 MB ✗ ✗ ✗ 0-16 MB
Dynamic power
66∗ 47.4∗ - 141∗ 151∗ 58.5∗∗ 3.86∗∗ 12.0∗∗ 14.7∗∗ -
(µW/MHz)
Floor plan area
0.11∗ 0.098∗ - 0.35∗ 0.44∗ 0.105∗∗ 0.0088∗∗ 0.028∗∗ 0.091∗∗ -
(mm2 )

SP: single precision; DP: double precision; HP: half precision.* 180ULL (7-track, typical 1.8V, 25°C). ** 40LP (9-track, typical 0.99V,
-40°C).

was achieved but at a cost of tripled power consumption non-trivial decision, considering the number of choices
[88]. The use of RISC-V MCU with x1.56 clock speed available. In order to find a balance between perfor-
against a Cortex-M4 MCU, provided a faster inference mance, memory, speed, and power a list of prioritized
by a factor of 7.7 at a power cost increase of x1.8. specifications needs to be set at first.
As TinyML is still a growing field, constant develop-
The overall performance of the TinyML system in ments and advances are surfacing each day to address
the context of health and care applications is affected these challenges and provide an easier path for the im-
by different factors. A few points can be summa- plementation of TinyML systems. One of these advances
rized and deducted from the above observations. First, are the TinyML development platforms which provide
the algorithm’s architecture and application (classifica- a bridge between the development and training of the
tion/regression) will directly affect the memory foot- ML algorithm and its deployment on embedded systems.
print and accuracy of the model. The input size also These platforms incorporate some of the libraries and
affects the model’s performance, where enough samples tools previously reviewed such as TFLite, and CMSIS-
are needed for an inference, but too many will affect NN, and provide compatibility with multiple commer-
the RAM occupancy and the time/inference. Hence, the cially available boards; allowing designers to collect
first challenge is to train and develop a well performing data, train, test, optimize, and deploy their model on
model, finding a balance between the input size and different MCUs using a single stop. Some of these newly
architecture’s hyperparameters. This, in its turn, brings developed platforms are Edge Impulse, imagimob, Qe-
with it challenges related to the use of ML in healthcare exo, NanoEdge AI, OctoML Apache TVM, and others.
such as the availability and reliability of datasets to These platforms are worth exploring as they provide a
train the base model. The following challenge would be smooth starting point for beginners and can be used to
the compression of the model whilst keeping a balance compare different implementation scenarios and assess
between accuracy loss and required memory footprint. their trade-offs.
This is conditioned to the post processing techniques
and software tools used which are also constrained by VI. CONCLUSION
the choice of MCU. Hence, the choice of MCU and The emerging field of TinyML technology, deploying
software tools is a co-dependent process. Choosing the machine learning algorithms on resource-constraints em-
most suitable MCU for an application is on its own a bedded devices, can provide several advantages over
22 VOLUME X, 2022

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and
content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2022.3206782

[Link] et al.: Embedded ML Using MCUs in Wearable and Ambulatory Systems for Health & Care Applications: A Review

cloud computing related to latency, power, and privacy. mal mammograms for predicting breast cancer risk,” Medical
In this work, we explored the application of TinyML in physics, vol. 47, no. 1, pp. 110–118, 2020.
[9] M. Anthimopoulos, S. Christodoulidis, L. Ebner, A. Christe,
wearable health and care systems. Several health and and S. Mougiakakou, “Lung pattern classification for interstitial
care applications including detection and management lung diseases using a deep convolutional neural network,” IEEE
of medical conditions, fitness tracking, elderly fall de- transactions on medical imaging, vol. 35, no. 5, pp. 1207–1216,
2016.
tection, and rehabilitation using prosthetics were found [10] L. Y. Tang, H. O. Coxson, S. Lam, J. Leipsic, R. C. Tam,
in this review, which demonstrate the feasibility of and D. D. Sin, “Towards large-scale case-finding: training
TinyML applications. The TinyML pipeline, covering and validation of residual networks for detection of chronic
obstructive pulmonary disease using low-dose ct,” The Lancet
the chosen MCUs and software tools and frameworks Digital Health, vol. 2, no. 5, pp. e259–e267, 2020.
were also covered and discussed within the review. This [11] J. Park, J. Yun, N. Kim, B. Park, Y. Cho, H. J. Park, M. Song,
included the specifications and limitations of available M. Lee, and J. B. Seo, “Fully automated lung lobe segmentation
in volumetric chest ct with 3d u-net: validation with intra-and
hardware options, related to MCU core processor, clock extra-datasets,” Journal of digital imaging, vol. 33, no. 1, pp.
speed, and available memory, as well as the features and 221–230, 2020.
hardware compatibility constraints of the software tools [12] H.-H. Chen, C.-M. Liu, S.-L. Chang, P. Y.-C. Chang, W.-S.
Chen, Y.-M. Pan, S.-T. Fang, S.-Q. Zhan, C.-M. Chuang, Y.-J.
and frameworks. A comprehensive review of the TinyML Lin et al., “Automated extraction of left atrial volumes from
wearable health and care systems was later presented two-dimensional computer tomography images using a deep
and summarized for further analysis and discussion, learning technique,” International Journal of Cardiology, vol.
316, pp. 272–278, 2020.
in terms of system implementation, design constraints, [13] A. Ghorbani, D. Ouyang, A. Abid, B. He, J. H. Chen,
and performance metrics related to memory occupancy, R. A. Harrington, D. H. Liang, E. A. Ashley, and J. Y. Zou,
accuracy, speed, and power. This showed that the over- “Deep learning interpretation of echocardiograms,” NPJ digital
medicine, vol. 3, no. 1, pp. 1–10, 2020.
all performance of TinyML systems depends on the [14] S. W. Baalman, F. E. Schroevers, A. J. Oakley, T. F. Brouwer,
software tools and choice of MCU, which was on its W. van der Stuijt, H. Bleijendaal, L. A. Ramos, R. R. Lopes,
turn also conditioned by the target application and ML H. A. Marquering, R. E. Knops et al., “A morphology based
deep learning model for atrial fibrillation detection using single
problem at hand. Overall, this review covers a gap in cycle electrocardiographic samples,” International journal of
literature in terms of the application of TinyML in cardiology, vol. 316, pp. 130–136, 2020.
wearable health and care devices and can help system [15] U. R. Acharya, H. Fujita, S. L. Oh, Y. Hagiwara, J. H. Tan, and
M. Adam, “Application of deep convolutional neural network
designers by guiding them through their initial system for automated detection of myocardial infarction using ecg
design decisions and trade-offs. signals,” Information Sciences, vol. 415, pp. 190–198, 2017.
[16] K. T. Oh, S. Lee, H. Lee, M. Yun, and S. K. Yoo, “Semantic
segmentation of white matter in fdg-pet using generative ad-
DATA ACCESS STATEMENT
versarial network,” Journal of digital imaging, pp. 1–10, 2020.
Not applicable. There is no data associated with the [17] S. Saha, A. Pagnozzi, P. Bourgeat, J. M. George, D. Bradford,
article. P. B. Colditz, R. N. Boyd, S. E. Rose, J. Fripp, and K. Pannek,
“Predicting motor outcome in preterm infants from very early
brain diffusion mri using a deep learning convolutional neural
REFERENCES network (cnn) model,” Neuroimage, vol. 215, p. 116807, 2020.
[1] Statista, “Wearables | statista,” Available online: https: [18] Q. Dou, H. Chen, L. Yu, L. Zhao, J. Qin, D. Wang, V. C.
//[Link]/study/15607/wearables-statista-dossier/. Mok, L. Shi, and P.-A. Heng, “Automatic detection of cerebral
[Accessed: 8 June 2022]. microbleeds from mr images via 3d convolutional neural net-
[2] M. Wu and J. Luo, “Wearable technology applications in works,” IEEE transactions on medical imaging, vol. 35, no. 5,
healthcare: a literature review,” Online J. Nurs. Inform, vol. 23, pp. 1182–1195, 2016.
no. 3, 2019. [19] J. Xu, D. Xu, Q. Wei, and Y. Zhou, “Automatic classification
[3] F. Sabry, T. Eltaras, W. Labda, K. Alzoubi, and Q. Malluhi, of male and female skeletal muscles using ultrasound imaging,”
“Machine learning for healthcare wearable devices: The big Biomedical Signal Processing and Control, vol. 57, p. 101731,
picture,” Journal of Healthcare Engineering, vol. 2022, 2022. 2020.
[4] A. Esteva, B. Kuprel, R. A. Novoa, J. Ko, S. M. Swetter, H. M. [20] A. Jodeiri, R. A. Zoroofi, Y. Hiasa, M. Takao, N. Sugano,
Blau, and S. Thrun, “Dermatologist-level classification of skin S. Yoshinobu, and Y. Otake, “Fully automatic estimation of
cancer with deep neural networks,” Nature, vol. 542, no. 7639, pelvic sagittal inclination from anterior-posterior radiography
pp. 115–118, 2017. image using deep learning framework,” Computer methods and
[5] Z. Jiang, F.-F. Yin, Y. Ge, and L. Ren, “A multi-scale programs in biomedicine, vol. 184, p. 105282, 2020.
framework with unsupervised joint training of convolutional [21] Y. Song, L. Zhang, S. Chen, D. Ni, B. Lei, and T. Wang,
neural networks for pulmonary deformable image registration,” “Accurate segmentation of cervical cytoplasm and nuclei based
Physics in Medicine & Biology, vol. 65, no. 1, p. 015011, 2020. on multiscale convolutional network and graph partitioning,”
[6] T. Apiparakoon, N. Rakratchatakul, M. Chantadisai, U. Vu- IEEE Transactions on Biomedical Engineering, vol. 62, no. 10,
trapongwatana, K. Kingpetch, S. Sirisalipoch, Y. Rakvongthai, pp. 2421–2433, 2015.
T. Chaiwatanarat, and E. Chuangsuwanich, “Malignet: Semisu- [22] K. Sirinukunwattana, S. E. A. Raza, Y.-W. Tsang, D. R.
pervised learning for bone lesion instance segmentation using Snead, I. A. Cree, and N. M. Rajpoot, “Locality sensitive deep
bone scintigraphy,” IEEE Access, vol. 8, pp. 27 047–27 066, learning for detection and classification of nuclei in routine
2020. colon cancer histology images,” IEEE transactions on medical
[7] Z. Zhao, J. Zhao, K. Song, A. Hussain, Q. Du, Y. Dong, J. Liu, imaging, vol. 35, no. 5, pp. 1196–1206, 2016.
and X. Yang, “Joint dbn and fuzzy c-means unsupervised deep [23] F. Piccialli, V. D. Somma, F. Giampaolo, S. Cuomo, and
clustering for lung cancer patient stratification,” Engineering G. Fortino, “A survey on deep learning in medicine: Why,
Applications of Artificial Intelligence, vol. 91, p. 103571, 2020. how and when?” Information Fusion, vol. 66, pp. 111–
[8] D. Arefan, A. A. Mohamed, W. A. Berg, M. L. Zuley, J. H. 137, 2021. [Online]. Available: [Link]
Sumkin, and S. Wu, “Deep learning modeling using nor- science/article/pii/S1566253520303651

VOLUME X, 2022 23

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and
content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2022.3206782

[Link] et al.: Embedded ML Using MCUs in Wearable and Ambulatory Systems for Health & Care Applications: A Review

[24] C. M. Micheel, S. J. Nass, G. S. Omenn et al., “Omics-based [45] T. M. Ingolfsson, X. W. M. Hersche, A. Burrello, L. Cavigelli,
clinical discovery: Science, technology, and applications,” in and L. Benini, “Ecg-tcn: Wearable cardiac arrhythmia detection
Evolution of Translational Omics: Lessons Learned and the with a temporal convolutional network,” in 2021 IEEE 3rd
Path Forward. National Academies Press (US), 2012, pp. 34– International Conference on Artificial Intelligence Circuits and
37. Systems (AICAS). IEEE, 2021, pp. 1–4.
[25] L. Greco, G. Percannella, P. Ritrovato, F. Tortorella, and [46] M. Magno, X. Wang, M. Eggimann, L. Cavigelli, and L. Benini,
M. Vento, “Trends in iot based solutions for health care: Moving “Infiniwolf: Energy efficient smart bracelet for edge computing
ai to the edge,” Pattern recognition letters, vol. 135, pp. 346– with dual source energy harvesting,” in 2020 Design, Automa-
353, 2020. tion & Test in Europe Conference & Exhibition (DATE).
[26] M. P. Véstias, R. P. Duarte, J. T. de Sousa, and H. C. Neto, IEEE, 2020, pp. 342–345.
“Moving deep learning to the edge,” Algorithms, vol. 13, no. 5, [47] T. M. Ingolfsson, A. Cossettini, X. Wang, E. Tabanelli,
p. 125, 2020. G. Tagliavini, P. Ryvlin, L. Benini, and S. Benatti, “Towards
[27] R. Sanchez-Iborra and A. F. Skarmeta, “Tinyml-enabled frugal long-term non-invasive monitoring for epilepsy via wearable
smart objects: Challenges and opportunities,” IEEE Circuits eeg devices,” in 2021 IEEE Biomedical Circuits and Systems
and Systems Magazine, vol. 20, no. 3, pp. 4–18, 2020. Conference (BioCAS). IEEE, 2021, pp. 01–04.
[28] M. Rahimiazghadi, C. Lammie, J. K. Eshraghian, M. Payvand, [48] STMicroelectronics, “Stm32l476xx,” DS10198 datasheet Rev 8,
E. Donati, B. Linares-Barranco, and G. Indiveri, “Hardware im- 2019.
plementation of deep network accelerators towards healthcare [49] ——, “Stm32l475xx,” DS10969 Rev 5, 2019.
and biomedical applications,” IEEE Transactions on Biomedi- [50] ——, “Stm32f303xd stm32f303xe,” DocID026415 Rev 5, 2016.
cal Circuits and Systems, 2020. [51] ——, “Stm32f756xx,” DocID027589 Rev 4, 2016.
[29] P. Jawandhiya, “Hardware design for machine learning,” Int. J. [52] ——, “Stm32f765xx stm32f767xx stm32f768xx stm32f769xx,”
Artif. Intell. Appl, vol. 9, no. 1, pp. 63–84, 2018. DS11532 Rev 7, 2021.
[30] L. Du and Y. Du, “Hardware accelerator design for ma- [53] Nordic-Semiconductor, “nrf52832 product specification v1.8,”
chine learning,” in Machine Learning-Advanced Techniques and 2021.
Emerging Applications. IntechOpen, 2017. [54] ——, “nrf52840 product specification v1.7,” 2021.
[31] M. P. Véstias, “Processing systems for deep learning inference [55] Microchip, “Atmega640/v-1280/v-1281/v-2560/v-2561/v
on edge devices,” in Convergence of Artificial Intelligence and datasheet,” DS40002211A, 2020.
the Internet of Things. Springer, 2020. [56] ——, “Atmega48a/pa/88a/pa/168a/pa/328/p datasheet,”
[32] T. S. Ajani, A. L. Imoize, and A. A. Atayero, “An overview DS40002061B, 2020.
of machine learning within embedded and mobile devices– [57] R. Chéour, S. Khriji, O. Kanoun et al., “Microcontrollers for iot:
optimizations and applications,” Sensors, vol. 21, no. 13, p. Optimizations, computing paradigms, and future directions,” in
4412, 2021. 2020 IEEE 6th World Forum on Internet of Things (WF-IoT).
[33] M. Capra, B. Bussolino, A. Marchisio, G. Masera, M. Martina, IEEE, 2020, pp. 1–7.
and M. Shafique, “Hardware and software optimizations for [58] A. Faraone and R. Delgado-Gonzalo, “Convolutional-recurrent
accelerating deep neural networks: Survey of current trends, neural networks on low-power wearable platforms for cardiac
challenges, and the road ahead,” IEEE Access, vol. 8, pp. arrhythmia detection,” in 2020 2nd IEEE International Confer-
225 134–225 180, 2020. ence on Artificial Intelligence Circuits and Systems (AICAS).
[34] G. Rong, A. Mendez, E. B. Assi, B. Zhao, and M. Sawan, IEEE, 2020, pp. 153–157.
“Artificial intelligence in healthcare: review and prediction case [59] T. Zhu, L. Kuang, K. Li, J. Zeng, P. Herrero, and P. Georgiou,
studies,” Engineering, vol. 6, no. 3, pp. 291–301, 2020. “Blood glucose prediction in type 1 diabetes using deep learning
[35] R. Zemouri, N. Zerhouni, and D. Racoceanu, “Deep learning in on the edge,” in 2021 IEEE International Symposium on
the biomedical applications: Recent and future status,” Applied Circuits and Systems (ISCAS). IEEE, 2021, pp. 1–5.
Sciences, vol. 9, no. 8, p. 1526, 2019. [60] R. Zanetti, A. Aminifar, and D. Atienza, “Robust epileptic
[36] A. M. Hafiz and G. M. Bhat, “A survey of deep learning seizure detection on wearable systems with reduced false-alarm
techniques for medical diagnosis,” in Information and Commu- rate,” in 2020 42nd Annual International Conference of the
nication Technology for Sustainable Development. Springer, IEEE Engineering in Medicine & Biology Society (EMBC).
2020, pp. 161–170. IEEE, 2020, pp. 4248–4251.
[37] A. Qayyum, J. Qadir, M. Bilal, and A. Al-Fuqaha, “Secure [61] Microchip, “8-bit avr® mcus,” Available on-
and robust machine learning for healthcare: A survey,” IEEE line: [Link]
Reviews in Biomedical Engineering, vol. 14, pp. 156–180, 2020. microcontrollers-and-microprocessors/8-bit-mcus/avr-mcus
[38] A. Shawahna, S. M. Sait, and A. El-Maleh, “Fpga-based accel- [Accessed: 20 November 2021].
erators of deep learning networks for learning and classification: [62] PULP, “Pulp hardware reference manual, version 1.6,”
A review,” IEEE Access, vol. 7, pp. 7823–7859, 2018. Available online: [Link]
[39] M. Hartmann, U. S. Hashmi, and A. Imran, “Edge computing blob/master/doc/[Link] [Accessed: 8 July 2021], 2019.
in smart health care systems: Review, challenges, and research [63] M. Merenda, C. Porcaro, and D. Iero, “Edge machine learning
directions,” Transactions on Emerging Telecommunications for ai-enabled iot devices: a review,” Sensors, vol. 20, no. 9, p.
Technologies, p. e3710, 2019. 2533, 2020.
[40] Y. Wei, J. Zhou, Y. Wang, Y. Liu, Q. Liu, J. Luo, C. Wang, [64] U. Kulkarni, S. M. Meena, S. V. Gurlahosur, P. Benagi,
F. Ren, and L. Huang, “A review of algorithm & hardware A. Kashyap, A. Ansari, and V. Karnam, “Ai model
design for ai-based biomedical applications,” IEEE transactions compression for edge devices using optimization techniques,”
on biomedical circuits and systems, vol. 14, no. 2, pp. 145–163, in Modern Approaches in Machine Learning and Cognitive
2020. Science: A Walkthrough: Latest Trends in AI, Volume
[41] L. Dutta and S. Bharali, “Tinyml meets iot: A comprehensive 2, V. K. Gunjan and J. M. Zurada, Eds. Springer
survey,” Internet of Things, vol. 16, p. 100461, 2021. International Publishing, 2021, pp. 227–240. [Online]. Available:
[42] RISC-V, Available online: [Link] [Accessed: 20 [Link]
November 2021]. [65] A. Krizhevsky, G. Hinton et al., “Learning multiple layers of
[43] A. Pullini, D. Rossi, I. Loi, G. Tagliavini, and L. Benini, “Mr. features from tiny images,” Citeseer, 2009.
wolf: An energy-precision scalable parallel ultra low power soc [66] TensorFlow, “Tensorflow lite,” Available online: [Link]
for iot edge processing,” IEEE Journal of Solid-State Circuits, [Link]/lite/guide [Accessed: 22 June 2021].
vol. 54, no. 7, pp. 1970–1981, 2019. [67] ——, “Tensorflow lite for microcontrollers,” Available online:
[44] “Greenwaves technologies. gap8 next generation [Link] [Accessed: 22
processor for smart sensors,” Available online: June 2021].
[Link] [68] N. Suda, “Machine learning on arm cortex-m microcontrollers,”
[Accessed: 8 July 2021]. ARM White paper, pp. 1–10, 2019.

24 VOLUME X, 2022

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and
content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2022.3206782

[Link] et al.: Embedded ML Using MCUs in Wearable and Ambulatory Systems for Health & Care Applications: A Review

[69] L. Lai, N. Suda, and V. Chandra, “Cmsis-nn: Efficient neu- [88] X. Wang, M. Hersche, B. Tömekce, B. Kaya, M. Magno, and
ral network kernels for arm cortex-m cpus,” arXiv preprint L. Benini, “An accurate eegnet-based motor-imagery brain–
arXiv:1801.06601, 2018. computer interface for low-power edge computing,” in 2020
[70] STMicroelectronics, “X-cube-ai, ai expansion pack for IEEE International Symposium on Medical Measurements and
stm32cubemx,” Available online: [Link] Applications (MeMeA). IEEE, 2020, pp. 1–6.
embedded-software/[Link] [Accessed: 22 June 2021]. [89] M. Zanghieri, S. Benatti, A. Burrello, V. Kartsch, F. Conti,
[71] Dennis, Don Kurian and Gopinath, Sridhar and Gupta, Chirag and L. Benini, “Robust real-time embedded emg recognition
and Kumar, Ashish and Kusupati, Aditya and Patil, Shishir G framework using temporal convolutional networks on a multi-
and Simhadri, Harsha Vardhan, “EdgeML: Machine Learning core iot processor,” IEEE transactions on biomedical circuits
for resource-constrained edge devices.” [Online]. Available: and systems, vol. 14, no. 2, pp. 244–256, 2019.
[Link] [90] M. Zanghieri, S. Benatti, A. Burrello, V. J. K. Morinigo,
[72] GreenWavesTechnologies, “Nn quick start guide,” Available R. Meattini, G. Palli, C. Melchiorri, and L. Benini, “semg-based
Online: [Link] regression of hand kinematics with temporal convolutional
nn_quick_start_guide/ [Accessed: 30 Nov. 2021]. networks on a low-power edge microcontroller,” in 2021 IEEE
[73] A. Garofalo, M. Rusci, F. Conti, D. Rossi, and L. Benini, “Pulp- International Conference on Omni-Layer Intelligent Systems
nn: A computing library for quantized neural network inference (COINS). IEEE, 2021, pp. 1–6.
at the edge on risc-v based parallel ultra low power clusters,” [91] X. Wang, F. Geiger, V. Niculescu, M. Magno, and L. Benini,
in 2019 26th IEEE International Conference on Electronics, “Smarthand: Towards embedded smart hands for prosthetic
Circuits and Systems (ICECS). IEEE, 2019, pp. 33–36. and robotic applications,” in 2021 IEEE Sensors Applications
[74] F. Conti, “Technical report: Nemo dnn quantization for deploy- Symposium (SAS). IEEE, 2021, pp. 1–6.
ment model,” arXiv preprint arXiv:2004.05930, 2020. [92] S. Sundaram, P. Kellnhofer, Y. Li, J.-Y. Zhu, A. Torralba, and
[75] A. Burrello, A. Garofalo, N. Bruschi, G. Tagliavini, D. Rossi, W. Matusik, “Learning the signatures of the human grasp using
and F. Conti, “Dory: Automatic end-to-end deployment of a scalable tactile glove,” Nature, vol. 569, no. 7758, pp. 698–702,
real-world dnns on low-cost iot mcus,” IEEE Transactions on 2019.
Computers, 2021. [93] M. Bachlin, M. Plotnik, D. Roggen, I. Maidan, J. M. Hausdorff,
[76] X. Wang, M. Magno, L. Cavigelli, and L. Benini, “Fann-on- N. Giladi, and G. Troster, “Wearable assistant for parkin-
mcu: An open-source toolkit for energy-efficient neural network son’s disease patients with the freezing of gait symptom,”
inference at the edge of the internet of things,” IEEE Internet IEEE Transactions on Information Technology in Biomedicine,
of Things Journal, vol. 7, no. 5, pp. 4403–4417, 2020. vol. 14, no. 2, pp. 436–446, 2009.
[77] S. Nissen et al., “Implementation of a fast artificial neural [94] G. D. Clifford, C. Liu, B. Moody, H. L. Li-wei, I. Silva,
network library (fann),” Report, Department of Computer Q. Li, A. Johnson, and R. G. Mark, “Af classification from
Science University of Copenhagen (DIKU), vol. 31, p. 29, 2003. a short single lead ecg recording: The physionet/computing in
[78] B. Kuyumcu, “Fann tool.” [Online]. Available: [Link] cardiology challenge 2017,” in 2017 Computing in Cardiology
[Link]/archive/p/fanntool/ (CinC). IEEE, 2017, pp. 1–4.
[79] H. Gokul, P. Suresh, B. H. Vignesh, R. P. Kumaar, and V. Vi- [95] Y. Chen, Y. Hao, T. Rakthanmanon, J. Zakaria, B. Hu, and
jayaraghavan, “Gait recovery system for parkinson’s disease E. Keogh, “A general framework for never-ending learning from
using machine learning on embedded platforms,” in 2020 IEEE time series streams,” Data mining and knowledge discovery,
International Systems Conference (SysCon). IEEE, 2020, pp. vol. 29, no. 6, pp. 1622–1664, 2015.
1–8. [96] D. S. Baim, W. S. Colucci, E. S. Monrad, H. S. Smith, R. F.
[80] D. Sopic, A. Aminifar, and D. Atienza, “e-glass: A wearable sys- Wright, A. Lanoue, D. F. Gauthier, B. J. Ransil, W. Grossman,
tem for real-time detection of epileptic seizures,” in 2018 IEEE and E. Braunwald, “Survival of patients with severe congestive
International Symposium on Circuits and Systems (ISCAS). heart failure treated with oral milrinone,” Journal of the
IEEE, 2018, pp. 1–5. American College of Cardiology, vol. 7, no. 3, pp. 661–670, 1986.
[81] V. Kartsch, G. Tagliavini, M. Guermandi, S. Benatti, D. Rossi, [97] A. L. Goldberger, L. A. Amaral, L. Glass, J. M. Hausdorff, P. C.
and L. Benini, “Biowolf: A sub-10-mw 8-channel advanced Ivanov, R. G. Mark, J. E. Mietus, G. B. Moody, C.-K. Peng,
brain–computer interface platform with a nine-core processor and H. E. Stanley, “Physiobank, physiotoolkit, and physionet:
and ble connectivity,” IEEE transactions on biomedical circuits components of a new research resource for complex physiologic
and systems, vol. 13, no. 5, pp. 893–906, 2019. signals,” circulation, vol. 101, no. 23, pp. e215–e220, 2000.
[82] M. Risso, A. Burrello, D. J. Pagliari, S. Benatti, E. Macii, [98] C. Marling and R. Bunescu, “The ohiot1dm dataset for blood
L. Benini, and M. Poncino, “Robust and energy-efficient ppg- glucose level prediction: Update 2020,” in CEUR workshop
based heart-rate monitoring,” in 2021 IEEE International Sym- proceedings, vol. 2675. NIH Public Access, 2020, p. 71.
posium on Circuits and Systems (ISCAS). IEEE, 2021, pp. [99] A. H. Shoeb, “Application of machine learning to epileptic
1–5. seizure onset detection and treatment,” PhD Thesis. Mas-
[83] L. Oden and T. Witt, “Fall-detection on a wearable micro sachusetts Institute of Technology, 2009.
controller using machine learning algorithms,” in 2020 IEEE In- [100] A. Reiss, I. Indlekofer, P. Schmidt, and K. Van Laerhoven,
ternational Conference on Smart Computing (SMARTCOMP). “Deep ppg: large-scale heart rate estimation with convolutional
IEEE, 2020, pp. 296–301. neural networks,” Sensors, vol. 19, no. 14, p. 3079, 2019.
[84] E. Torti, A. Fontanella, M. Musci, N. Blago, D. Pau, F. Lepo- [101] J. A. Healey and R. W. Picard, “Detecting stress during
rati, and M. Piastra, “Embedded real-time fall detection with real-world driving tasks using physiological sensors,” IEEE
deep learning on wearable devices,” in 2018 21st euromicro Transactions on intelligent transportation systems, vol. 6, no. 2,
conference on digital system design (DSD). IEEE, 2018, pp. pp. 156–166, 2005.
405–412. [102] A. Sucerquia, J. D. López, and J. F. Vargas-Bonilla, “Sisfall:
[85] M. Musci, D. De Martini, N. Blago, T. Facchinetti, and M. Pi- A fall and movement dataset,” Sensors, vol. 17, no. 1, p. 198,
astra, “Online fall detection using recurrent neural networks 2017.
on smart wearable devices,” IEEE Transactions on Emerging [103] “Sisfall temporally annotated.” [Online]. Available: https:
Topics in Computing, 2020. //[Link]/unipv_cvmlab/sisfalltemporallyannotated/
[86] M. Merenda, M. Astrologo, D. Laurendi, V. Romeo, and src/master/
F. G. Della Corte, “A novel fitness tracker using edge machine [104] G. Biagetti, P. Crippa, L. Falaschetti, L. Saraceni, A. Tiranti,
learning,” in 2020 IEEE 20th Mediterranean Electrotechnical and C. Turchetti, “Dataset from ppg wireless sensor for activity
Conference (MELECON). IEEE, 2020, pp. 212–215. monitoring,” Data in brief, vol. 29, p. 105044, 2020.
[87] M. Alessandrini, G. Biagetti, P. Crippa, L. Falaschetti, and [105] G. Schalk, D. J. McFarland, T. Hinterberger, N. Birbaumer,
C. Turchetti, “Recurrent neural network for human activity and J. R. Wolpaw, “Bci2000: a general-purpose brain-computer
recognition in embedded systems using ppg and accelerometer interface (bci) system,” IEEE Transactions on biomedical engi-
data,” Electronics, vol. 10, no. 14, p. 1715, 2021. neering, vol. 51, no. 6, pp. 1034–1043, 2004.

VOLUME X, 2022 25

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and
content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2022.3206782

[Link] et al.: Embedded ML Using MCUs in Wearable and Ambulatory Systems for Health & Care Applications: A Review

[106] F. Palermo, M. Cognolato, A. Gijsberts, H. Müller, B. Caputo,


and M. Atzori, “Repeatability of grasp recognition for robotic
hand prosthesis control based on semg data,” in 2017 Interna-
tional Conference on Rehabilitation Robotics (ICORR). IEEE,
2017, pp. 1154–1159.
[107] A. Krasoulis, S. Vijayakumar, and K. Nazarpour, “Effect of
user practice on prosthetic finger control with an intuitive
myoelectric decoder,” Frontiers in neuroscience, p. 891, 2019.
[108] C. Gupta, A. S. Suggala, A. Goyal, H. V. Simhadri, B. Paran-
jape, A. Kumar, S. Goyal, R. Udupa, M. Varma, and P. Jain,
“Protonn: Compressed and accurate knn for resource-scarce
devices,” in International Conference on Machine Learning.
PMLR, 2017, pp. 1331–1340.
[109] J. G. Nutt, B. R. Bloem, N. Giladi, M. Hallett, F. B. Horak,
and A. Nieuwboer, “Freezing of gait: moving forward on a mys-
terious clinical phenomenon,” The Lancet Neurology, vol. 10,
no. 8, pp. 734–744, 2011.
[110] A. Gordon, E. Eban, O. Nachum, B. Chen, H. Wu, T.-J. Yang,
and E. Choi, “Morphnet: Fast & simple resource-constrained
structure learning of deep networks,” in Proceedings of the
IEEE conference on computer vision and pattern recognition,
2018, pp. 1586–1595.
[111] V. J. Lawhern, A. J. Solon, N. R. Waytowich, S. M. Gordon,
C. P. Hung, and B. J. Lance, “Eegnet: a compact convolutional
neural network for eeg-based brain–computer interfaces,” Jour-
nal of neural engineering, vol. 15, no. 5, p. 056013, 2018.
[112] H. Dose, J. S. Møller, H. K. Iversen, and S. Puthusserypady,
“An end-to-end deep learning approach to mi-eeg signal classi-
fication for bcis,” Expert Systems with Applications, vol. 114,
pp. 532–542, 2018.
[113] P. Ferguson, “This ai-powered asthma inhaler keeps tabs on
your dose so you don’t have to,” Available online: [Link]
[Link]/blogs/blueprint/ai-asthma-inhaler-keeps-tabs-dose,
2018.
[114] ARM, “Arm cortex-m processor comparison table,” Version
2022.
[115] arm Developer, “Cortex-m55,” Available online: https://
[Link]/Processors/Cortex-M55 [Accessed: 8 June
2022].
[116] ——, “Processors - microcontrollers,” Available
online: [Link]
aq=%40navigationhierarchiescategories%3D%
3D%22Processor%20products%22%20AND%20%
40navigationhierarchiescontenttype%3D%3D%
22Product%20Information%22&numberOfResults=48&
f[navigationhierarchiesprocessortype]=Microcontrollers
[Accessed: 8 June 2022].

26 VOLUME X, 2022

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]

You might also like