Transforming Unstructured Voice and Text Data Into Insight For Paramedic Emergency Service Using Recurrent and Convolutional Neural Networks
Transforming Unstructured Voice and Text Data Into Insight For Paramedic Emergency Service Using Recurrent and Convolutional Neural Networks
ABSTRACT
Paramedics often have to make lifesaving decisions within a limited time in an ambulance. They sometimes ask the
doctor for additional medical instructions, during which valuable time passes for the patient. This study aims to
automatically fuse voice and text data to provide tailored situational awareness information to paramedics. To train and
test speech recognition models, we built a bidirectional deep recurrent neural network (long short-term memory
(LSTM)). Then we used convolutional neural networks on top of custom-trained word vectors for sentence-level
classification tasks. Each sentence is automatically categorized into four classes, including patient status, medical
history, treatment plan, and medication reminder. Subsequently, incident reports were automatically generated to extract
keywords and assist paramedics and physicians in making decisions. The proposed system found that it could provide
timely medication notifications based on unstructured voice and text data, which was not possible in paramedic
emergencies at present. In addition, the automatic incident report generation provided by the proposed system improves
the routine but error-prone tasks of paramedics and doctors, helping them focus on patient care.
Keywords: deep learning, convolutional neural networks, recurrent neural networks, natural language understanding
1. INTRODUCTION
Paramedics must make numerous life-saving decisions, often in the back of an ambulance with limited time. They
sometimes ask the doctor for additional medical instructions, during which the patient's precious time is wasted. The
Department of Homeland Security (DHS) Science and Technology Directorate (S&T) partnered with Canada’s
Department of National Defence Science and Technology Organization, Defence Research and Development Canada
Centre for Security Science (DRDC CSS) to support this experiment to examine whether artificial intelligence could be
used to improve the information overload. The experiment was part of the Next Generation First Responder Program by
.
DHS 1–5
Transcribing voice communications in paramedic emergency response environments is the first step for
providing faster, more accurate first response to emergency patients. However, automatic speech recognition in this
environment is particularly challenging due to the lack of training data, unfamiliar medical terms in acronyms, and noisy
environment in emergency vehicles. We used bidirectional deep recurrent neural networks to train and test speech
recognition performance. We showed that data augmentation and custom language models can improve speech
recognition accuracy 6 . Then we used keywords extraction techniques to analyze spoken sentences and to summarize the
next steps for treatment and for communication with doctors for the medical orders. The proposed system will help the
machine analyze medical information in paramedic services and accelerate emergency medical responses.
Automatic speech recognition has three main applications, including input/output devices, communication aids,
and information retrieval. The performance of automatic speech recognition has been greatly improved by using deep
neural networks mainly for input / output devices 7 . Industry leaders such as Amazon, Apple, Baidu, Google, Microsoft,
and IBM are evolving to help the public easily use their products through automatic speech recognition. Conversational
interfaces based on Amazon’s Alexa and Google’s Home have been enhanced to naturally acquire information, access
User Paramedic
Short term
Lemmatization, parsing, value extraction, speech recognition (offline)
improvement
Long term
Background listening, streaming, real-time processing, wearable
improvement
The log was introduced to handle very small values of the probability, P (y(t) | x, y(1), ... , y(t − 1)) , especially when
the decoding sentence is long.
We found that bidirectional recurrent neural networks with data augmentation and custom language models
achieved the best performance in a highly specialized language environment with specialized abbreviations and grammar
structures. The result is significant in that even perfect human hearing perception cannot fully comprehend
communication in the paramedic emergency response without prior knowledge. A vast amount of training and
knowledge of specific situations is important for transcription of professional voice communications.
The speech recognition problem is to map the audio input x to the script y. The human ear converts a
one-dimensional audio input to the intensity of the frequency component. This can be thought of as a preprocessing step
to generate a spectrogram that maps 2D information of the time and frequency of the audio input. The human brain
and social cues
utilizes a variety of contextual information to fully understand and perceive speech, including attention 23
24,25
, as well as sound itself.
By using attention models in the future, we can bias the state of recurrent neural networks to improve
. The attention model must be
context-based speech recognition by paying more attention to specific words and phrases 26
trained as a separate RNN in the previous state to determine how much attention should be paid to adjacent inputs to bias
the current state of the primary RNN. Naturally, the processing time to train the network will be at least doubled.
Because of its inefficiency, the attention model should be carefully considered for speech recognition solutions. The
and image caption with visual attention 28
attention model can also be used for machine translation 27 .
Automatic speech recognition, supported by computer vision, can be useful for improving speech recognition
accuracy in noisy environments. Previous studies have shown that lip reading computer vision is far superior to
traditional noise reduction methods 29–31 . Another possibility that can improve accuracy is to use background information
such as location information 32 and history information 33 . This information has been shown to improve speech
recognition accuracy by reducing possible word combinations in certain situations.
Table 2. Performance comparisons among the existing system, proposed system (current), and estimated ideal system.
Improvements are highlighted in yellow.
Paramedic Existing
Proposed Estimated Ideal
Emergency Paramedic Work Notes
Model (min) Model (min)
Response Steps (min)
Paramedic call Patient data was automatically shared with the doctor, so
3 0 0
physician there was no need to call the doctor
Patch form 5 7 0
The proposed model took longer due to the inconvenient
Request to user interface for fixing errors in the patch form
1 5 0
physician
The doctor reviewed the patch form and ordered
Physician order 1 1 1
treatment
ePCR (ACP) The proposed model improved the post-incident report
10 5 0
data input generation process
Total (min) 38.75 50.75 16.75
* ePCR: Electronic Patient Care Record
* ACP: Advanced Care Protocol
Finally, in the traditional paramedic work, the entire process took 38 minutes, 50 minutes from the proposed
model, and 16 minutes from the expected ideal model. We believe that the estimated ideal model performance can be
achieved after improving the user interface and real-time data processing capabilities of the proposed model in the
future.
Figure 1. Convolutional neural networks to categorize sentences. The networks were trained on top of custom-trained word
vectors for sentence-level classifications.
For speech understanding to support emergency response, converting speech to text was the first step, and to
help further understand the sentence, each sentence had to be classified into three major classes. We organized three
classes, including basic medical information, vital signs, and medication doses (Figure 1). The given sentence is
transformed into a multidimensional vector matrix, and the input matrix goes through a convolutional layer with multiple
.
filter widths and features, eventually sending three classes to the output 37
Figure 2. Complete patch form generation process 1. Generate a ground truth voice transcript through communication
between paramedics and doctors. 2. Generate an expert knowledge model (doctor's mental logic) to translate from
paramedic voice (ground truth transcript) to patch form. 3. Train the automatic voice transcription model. 4. Train the
language model to translate the script into a patch using the expert model. 5. Generate an automatic patch form.
Lemmatization was implemented to improve text understanding using Python’s NLTK (Natural Language
. The goal of lemmatization was to reduce inflectional forms and sometimes derivationally related forms of a
Toolkit) 38
word to a common base form. For instance, “am, are, is” are changed to “be”, and “car, cars, car's, cars'” to “car”. Using
Python’s NLTK, we implemented a parser that can identify unnecessary words to remove them and key nouns to extract
them. Patch form is the physician’s summary of the paramedic’s verbal report about the patient on scene. We
implemented automatic generation of patch form information (audio summary) using paramedic speeches. The entire
patch form generation process is illustrated in Figure 2.
A domain-specific language model that transforms voice scripts into patch forms used by emergency personnel
and doctors was key to improving the accuracy of natural language understanding. Graph-based expert knowledge
models were used to generate domain-specific language models. For example, if a patient is identified as having a heart
problem, the relevant database is loaded and ECG (electrocardiogram) vital signs are analyzed. Depending on the ECG
results, the likelihood of relevant keywords (e.g. sinus rhythm, sinus tach) appearing in paramedic and physician
conversations increases, thus maximizing the accuracy of natural language understanding in the semantic language
model. This process is very similar to the process of understanding human language in that it not only analyzes sounds of
speeches, but also understands the context of a particular situation and the topic area of the mentioned subject.
Figure 3. Graph database of terminology used in patching (paramedic-doctor communication). The database can be used to
improve parsing and extracting information from unstructured text (speech transcript).
We achieved 93.3% test accuracy in speech recognition and understanding. Examples of voice transcription and
keyword extraction are listed in Table 3 below. Keywords were extracted and categorized for further understanding,
including previous nitroglycerin doses for treatment, abdominal physical findings, skin condition, past medical history,
and medication reminders. In addition, basic medical information and vital signs were extracted for treatment
recommendations and data transfer to the hospital. Basic medical information includes age, gender, Canadian Triage and
Acuity Scale (CTAS), blood pressure, medication, pupil examination, temperature, pulse, allergies, physical examination
and skin color.
Figure 4. Inception V3 architecture for image classification, and flooded road detection dashboard showing real-time
Ontario public CCTV website, extracted image from Ontario public CCTV website, and CCTV location and automatic
flood detection confidence (video demo: [Link]
Figure 5. Example of improvement in optical character recognition by custom language model (CLM)
Using the same OCR technology as used in medicine bottle recognition, we added a hazard placard recognition function
to allow paramedics to quickly identify each hazard placard and send safety-related warnings (Figure 6). The hazard
placard number identified in the Emergency Response Guidebook (ERG) is automatically entered to extract relevant
information, and the information is then sent to paramedics.
We developed a real-time classifier that automatically retrieves and categorizes information from the database
to assist in the rapid patient diagnosis of emergency personnel. The classifier was built on a historical record of
generating recommendations for standardized medical orders using information such as patient age, patient gender, and
patient's reported telephone calls. Paramedics refer to standing orders to determine treatment options according to the
restrictions outlined in the document.
Paramedics' comments are automatically sent to the description field along with patient data collected by the
initial dispatcher, and recommendations are generated when the minimum amount of information is met. The
information collected through the call (between the dispatcher and patient) is stored and updated, and the standing order
is automatically corrected if necessary.
Example input data: {'problem_nature_type': 'CHEST', 'problem_nature': 'Ischemic Chest Pain-(51)', 'gender':
'M', 'comment': '50YOM, SOB, pale diaphoretic, history of cardiac’}
Example output standing order: {'confidence_levels': [{'order': 'ACPE-ACP-2019', 'confidence': '0.086'},
{'order': 'CSMD-2019', 'confidence': '0.110'}, {'order': 'CIMD-ACP-2019', 'confidence': '0.804'}], 'timestamp':
'20190101T010101-000000'}
Figure 6. Hazard placard recognition using OCR and matching information from Emergency Response Guidebook (ERG)
7. CONCLUSION
In this study, natural language understanding and optical character recognition technology were developed to provide
paramedics with the most important situational awareness at an appropriate timing. Specific use cases, including hazard
placard recognition, medicine bottle recognition, medication dosage notifications, automatic medical documentation
generation (patch form), and automatic first treatment recommendations (standing orders), were especially helpful for
paramedics to focus on patient care and reduce errors. However, the technology provided has not been optimized yet,
with room for improvement, especially in the user interface and real-time processing capabilities. These design and
engineering issues may be addressed in the future commercialization of technology.
ACKNOWLEDGMENT
The research was carried out at the Jet Propulsion Laboratory, California Institute of Technology, under a contract with
the National Aeronautics and Space Administration. The research was funded by the U.S. Department of Homeland
Security Science and Technology Directorate Next Generation First Responders Apex Program (DHS S&T NGFR)
under NASA prime contract NAS7-03001, Task Plan Number 82-106095.
REFERENCES
[1] Yun, K., Lu, T. and Chow, E., “Occluded object reconstruction for first responders with augmented reality
glasses using conditional generative adversarial networks,” Pattern Recognition and Tracking XXIX 10649,
106490T, International Society for Optics and Photonics (2018).
[2] Lu, T., Huyen, A., Nguyen, L., Osborne, J., Eldin, S. and Yun, K., “Optimized training of deep neural network
for image analysis using synthetic objects and augmented reality,” Pattern Recognition and Tracking XXX
10995, 109950I, International Society for Optics and Photonics (2019).
[3] Yun, K., Yu, K., Osborne, J., Eldin, S., Nguyen, L., Huyen, A. and Lu, T., “Improved visible to IR image
transformation using synthetic data augmentation with cycle-consistent adversarial networks,” Pattern
Recognition and Tracking XXX 10995, 1099502, International Society for Optics and Photonics (2019).
[4] Yun, K., Nguyen, L., Nguyen, T., Kim, D., Eldin, S., Huyen, A., Lu, T. and Chow, E., “Small target detection
for search and rescue operations using distributed deep learning and synthetic data generation,” Pattern
Recognition and Tracking XXX 10995, 1099507, International Society for Optics and Photonics (2019).
[5] Yun, K., Bustos, J. and Lu, T., “Predicting rapid fire growth (flashover) using conditional generative adversarial
networks,” Electronic Imaging 2018(9), 127–1 (2018).
[6] Yun, K., Osborne, J., Lee, M., Lu, T. and Chow, E., “Automatic speech recognition for launch control center
communication using recurrent neural networks with data augmentation and custom language model,”
Disruptive Technologies in Information Sciences 10652, 1065202, International Society for Optics and
Photonics (2018).
[7] Hinton, G., Deng, L., Yu, D., Dahl, G., Mohamed, A., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P. and
Kingsbury, B., “Deep neural networks for acoustic modeling in speech recognition,” IEEE Signal processing
magazine 29 (2012).
[8] McTear, M., Callejas, Z. and Griol, D., “The dawn of the conversational interface,” [The Conversational
Interface], Springer, 11–24 (2016).
[9] Xiong, W., Droppo, J., Huang, X., Seide, F., Seltzer, M., Stolcke, A., Yu, D. and Zweig, G., “Achieving human
parity in conversational speech recognition,” arXiv preprint arXiv:1610.05256 (2016).
[10] Anusuya, M. A. and Katti, S. K., “Speech Recognition by Machine, a Review (Department of Computer Science
and Engineering Sri Jayachamarajendra College of Engineering Mysore, India, 2010),” arXiv preprint
arXiv:1001.2267.
[11] Dahl, G. E., Yu, D., Deng, L. and Acero, A., “Context-dependent pre-trained deep neural networks for
large-vocabulary speech recognition,” IEEE Transactions on audio, speech, and language processing 20(1),
30–42 (2011).
[12] Yu, D. and Li, J., “Recent progresses in deep learning based acoustic models,” IEEE/CAA Journal of
Automatica Sinica 4(3), 396–409 (2017).
[13] Graves, A. and Jaitly, N., “Towards end-to-end speech recognition with recurrent neural networks,” International
conference on machine learning, 1764–1772 (2014).
[14] Panayotov, V., Chen, G., Povey, D. and Khudanpur, S., “Librispeech: an ASR corpus based on public domain
audio books,” 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP),
5206–5210, IEEE (2015).
[15] Amodei, D., Ananthanarayanan, S., Anubhai, R., Bai, J., Battenberg, E., Case, C., Casper, J., Catanzaro, B.,
Cheng, Q. and Chen, G., “Deep speech 2: End-to-end speech recognition in english and mandarin,” International
conference on machine learning, 173–182 (2016).
[16] Hannun, A., Case, C., Casper, J., Catanzaro, B., Diamos, G., Elsen, E., Prenger, R., Satheesh, S., Sengupta, S.
and Coates, A., “Deep speech: Scaling up end-to-end speech recognition,” arXiv preprint arXiv:1412.5567
(2014).
[17] Hochreiter, S., Bengio, Y., Frasconi, P. and Schmidhuber, J., “A field guide to dynamical recurrent neural
networks,” [chapter Gradient Flow in Recurrent Nets: The Difficulty of Learning Long-Term Dependencies],
Wiley-IEEE Press, 237–243 (2001).
[18] LeCun, Y., Bengio, Y. and Hinton, G., “Deep learning,” nature 521(7553), 436–444 (2015).
[19] Schuster, M. and Paliwal, K. K., “Bidirectional recurrent neural networks,” IEEE Transactions on Signal
Processing 45(11), 2673–2681 (1997).
[20] Hochreiter, S. and Schmidhuber, J., “Simplifying neural nets by discovering flat minima,” Advances in neural
information processing systems, 529–536 (1995).
[21] Graves, A., Fernández, S., Gomez, F. and Schmidhuber, J., “Connectionist temporal classification: labelling
unsegmented sequence data with recurrent neural networks,” Proceedings of the 23rd international conference on
Machine learning, 369–376, ACM (2006).
[22] Antoniol, G., Brugnara, F., Cettolo, M. and Federico, M., “Language model representations for beam-search
decoding,” 1995 International Conference on Acoustics, Speech, and Signal Processing 1, 588–591, IEEE
(1995).
[23] Yun, K. and Stoica, A., “Improved target recognition response using collaborative brain-computer interfaces,”
presented at Systems, Man, and Cybernetics (SMC), 2016 IEEE International Conference on, 2016,
002220–002223, IEEE.
[24] Yun, K., Watanabe, K. and Shimojo, S., “Interpersonal body and neural synchronization as a marker of implicit
social interaction,” Scientific Reports 2(959), 959 (2012).
[25] Yun, K., “On the same wavelength: Face-to-face communication increases interpersonal neural
synchronization,” The Journal of Neuroscience 33(12), 5081–5082 (2013).
[26] Bahdanau, D., Chorowski, J., Serdyuk, D., Brakel, P. and Bengio, Y., “End-to-end attention-based large
vocabulary speech recognition,” 2016 IEEE international conference on acoustics, speech and signal processing
(ICASSP), 4945–4949, IEEE (2016).
[27] Bahdanau, D., Cho, K. and Bengio, Y., “Neural machine translation by jointly learning to align and translate,”
arXiv preprint arXiv:1409.0473 (2014).
[28] Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R. and Bengio, Y., “Show, attend and
tell: Neural image caption generation with visual attention,” International conference on machine learning,
2048–2057 (2015).
[29] Thangthai, K., Harvey, R. W., Cox, S. J. and Theobald, B.-J., “Improving lip-reading performance for robust
audiovisual speech recognition using DNNs.,” AVSP, 127–131 (2015).
[30] Chung, J. S., Senior, A., Vinyals, O. and Zisserman, A., “Lip reading sentences in the wild,” 2017 IEEE
Conference on Computer Vision and Pattern Recognition (CVPR), 3444–3453, IEEE (2017).
[31] Petajan, E., “Automatic lipreading to enhance speech recognition,” Proc. CVPR’85 (1985).
[32] Endo, N., Reaves, B. K., Prieto, R. E. and Brookes, J. R., “Method and system for speech recognition using
grammar weighted based upon location information” (2008).
[33] Mozer, T. F. and Mozer, F. S., “Systems and methods of performing speech recognition using historical
information” (2011).
[34] Buckle, P., Clarkson, P. J., Coleman, R., Bound, J., Ward, J. and Brown, J., “Systems mapping workshops and
their role in understanding medication errors in healthcare,” Applied ergonomics 41(5), 645–656 (2010).
[35] Chan, E. W., Taylor, S. E., Marriott, J. L. and Barger, B., “Bringing patients’ own medications into an
emergency department by ambulance: effect on prescribing accuracy when these patients are admitted to
hospital,” Medical journal of Australia 191(7), 374–377 (2009).
[36] Maurin Söderholm, H., “Emergency visualized: exploring visual technology for paramedic-physician
collaboration in emergency care,” PhD Thesis (2013).
[37] Kim, Y., “Convolutional neural networks for sentence classification,” arXiv preprint arXiv:1408.5882 (2014).
[38] Perkins, J., [Python 3 text processing with NLTK 3 cookbook], Packt Publishing Ltd (2014).
[39] Xia, X., Xu, C. and Nan, B., “Inception-v3 for flower classification,” 2017 2nd International Conference on
Image, Vision and Computing (ICIVC), 783–787, IEEE (2017).
[40] Breuel, T. M., “High performance text recognition using a hybrid convolutional-lstm implementation,” 2017
14th IAPR International Conference on Document Analysis and Recognition (ICDAR) 1, 11–16, IEEE (2017).
[41] Breuel, T. M., Ul-Hasan, A., Al-Azawi, M. A. and Shafait, F., “High-performance OCR for printed English and
Fraktur using LSTM networks,” 2013 12th International Conference on Document Analysis and Recognition,
683–687, IEEE (2013).