0% found this document useful (0 votes)
44 views8 pages

Filipino Sign Language Recognition Using Deep Learning

The paper discusses the development of a Filipino Sign Language (FSL) recognition system using deep learning and computer vision techniques. Utilizing a Convolutional Neural Network (CNN) ResNet architecture, the study achieved a validation accuracy of 86.7% for recognizing Filipino number signs (0-9) from static images. Future work aims to enhance the system for real-time recognition and expand its capabilities to include Filipino alphabets and common phrases.

Uploaded by

naynaveran
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
44 views8 pages

Filipino Sign Language Recognition Using Deep Learning

The paper discusses the development of a Filipino Sign Language (FSL) recognition system using deep learning and computer vision techniques. Utilizing a Convolutional Neural Network (CNN) ResNet architecture, the study achieved a validation accuracy of 86.7% for recognizing Filipino number signs (0-9) from static images. Future work aims to enhance the system for real-time recognition and expand its capabilities to include Filipino alphabets and common phrases.

Uploaded by

naynaveran
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

See discussions, stats, and author profiles for this publication at: [Link]

net/publication/357306934

Filipino Sign Language Recognition using Deep Learning

Conference Paper · August 2021


DOI: 10.1145/3485768.3485783

CITATIONS READS
9 6,063

3 authors:

Myron Darrel Layug Montefalcon Jay Rhald Padilla


National University National University
13 PUBLICATIONS 31 CITATIONS 12 PUBLICATIONS 26 CITATIONS

SEE PROFILE SEE PROFILE

Ramon Rodriguez
National University
43 PUBLICATIONS 191 CITATIONS

SEE PROFILE

All content following this page was uploaded by Myron Darrel Layug Montefalcon on 03 August 2023.

The user has requested enhancement of the downloaded file.


Filipino Sign Language Recognition using Deep Learning
Myron Darrel L. Montefalcon Jay Rhald C. Padilla Ramon L. Rodriguez
College of Computing and College of Computing and College of Computing and
Information Technologies, National Information Technologies, National Information Technologies, National
University-Manila, Philippines University-Manila, Philippines University-Manila, Philippines,
montefalconml@[Link]- padillaj1@[Link]- [Link]
[Link] [Link]

ABSTRACT is through the use of spoken languages or verbal communication. In


The Filipino deaf community continues to lag behind the fast-paced the Philippines alone, the total number of the deaf, mute or hearing-
and technology-driven society in the Philippines. The use of Fil- impaired individuals is equivalent to 1.23% of the entire population
ipino Sign Language (FSL) has contributed to the improvement of [1]. With the help of sign language, it has contributed in bridging
communication of deaf people, however, the majority of the popula- the gap between the deaf community and the hearing majorities.
tion in the Philippines do not understand FSL. This project utilized Philippines has its own sign language called Filipino Sign Language
computer vision in obtaining the images and Convolutional Neural or FSL, it was declared as the national sign language in the year
Network (CNN) ResNet architecture in building the automated FSL 2012. Despite the implementation of RA 11106, which mandates
recognition model, with the goal of bridging the communication FSL to be used in all transactions involving the deaf, and to be use
gap between the deaf community and the hearing majorities. In the in schools, broadcast media, and workplaces [2]. Majority of the
experimentation, the dataset used are static images generated from hearing Filipinos do not understand FSL and learning it requires
a signer which gestured Filipino number signs which range from training [3]. As a result, it has led to an undeniable communication
(0-9). Based on experimentation, the best-achieved performance is gap between the Deaf community and the hearing majorities in the
on fine-tuned ResNet-50 model which obtained a validation accu- Philippines.
racy as high as 86.7% when the epoch value equals 15. For future With the continuous advancement of technology and research,
work, real-time FSL recognition will be implemented and more data it has helped the deaf community to have more opportunities in
will be collected to enable recognition of Filipino alphabets, basic life [4]. Innovations in automated Sign Language Recognition (SLR)
phrases, and common greetings. attempt to minimize the communication barrier deaf community
faced. Sign Language recognition is a collaborative research field
CCS CONCEPTS wherein its objective is to build various methods and algorithms to
identify visual signs and to translate its meaning [5]. For the past
• Computing methodologies → Neural network.
decades of exploration on SLR, the research has achieved significant
KEYWORDS milestones. In a study of Martinez et al. [6] in the year 2008, presents
a research road map which envisions the Philippines to make an
Sign Language Recognition, Filipino Sign Language, Convolutional impact on sign language recognition and computational linguistics
Neural Network, ResNet Architecture in the next 20 years. However, in the present time research on
ACM Reference Format: Filipino Sign Language recognition remains limited, as developing
Myron Darrel L. Montefalcon, Jay Rhald C. Padilla, and Ramon L. Rodriguez. successful FSL recognition system requires an extensive amount of
2021. Filipino Sign Language Recognition using Deep Learning. In 2021 5th data and expertise in different fields, including image processing,
International Conference on E-Society, E-Education and E-Technology (ICSET
computer vision, natural language processing, human-computer
2021), August 21–23, 2021, Taipei, Taiwan. ACM, New York, NY, USA, 7 pages.
interaction, linguistics, and knowledge on Filipino deaf culture [7].
[Link]
Through the years, hundreds of approaches have been devel-
1 INTRODUCTION oped in building SLR systems. In [8] investigated 240 approaches in
SLR and discussed its strengths and limitations. Nevertheless, the
Communication is the foundation of building a progressive nation;
more popular approaches can be classified into three main architec-
an effective communication leads to not only understanding but
tures; glove-based [9], computer vision [10] and hybrid architecture
development to different sectors of the society. However, there
[11]. Glove-based approach utilizes gloves equipped with various
are certain groups of speople like the deaf community who are
kinds of sensors to acquire gesture-related data. On the other hand,
struggling or incapable of the conventional communication which
visual-based systems utilize cameras as the main device in obtain-
Permission to make digital or hard copies of all or part of this work for personal or ing the needed input data. While, hybrid system is the combination
classroom use is granted without fee provided that copies are not made or distributed
for profit or commercial advantage and that copies bear this notice and the full citation
of both architectures. Although each of architecture mentioned
on the first page. Copyrights for components of this work owned by others than ACM has its own strength, the main disadvantage of the glove-based
must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, and hybrid architecture is its costly proponents and the computa-
to post on servers or to redistribute to lists, requires prior specific permission and/or a
fee. Request permissions from permissions@[Link]. tional challenges of the entire setup [12]. On contrary, vision-based
ICSET 2021, August 21–23, 2021, Taipei, Taiwan approaches only requires a camera in obtaining data which is a
© 2021 Association for Computing Machinery. cheaper proponent compared to commercially made gloves; not
ACM ISBN 978-1-4503-9015-6/21/08. . . $15.00
[Link]

219
ICSET 2021, August 21–23, 2021, Taipei, Taiwan Myron Darrel Montefalcon et al.

only removing the need for gloves equipped with sensors but also 2.2 Visual-based Approach
reducing the development cost of the system. Sign language is a system of communication which uses visual
The proposed study aims to build a system capable of recogniz- gestures and signs, because of this many has explored the use of
ing static FSL gesture and translate it to its corresponding texts. computer vision in building SLR systems. Computer vision has been
In the proposed approach in building the FSL recognition model, remarkable in the recent years and this technology allows computer
the use of computer vision techniques powered by deep learning to extract useful insights from videos and images [23]. In SLR, the
algorithm will be explored. Whereas the main proponents of the visual-based approach utilize camera as its main proponent to ob-
approach include; a web camera and specific spot where hands tain necessary image features which includes (hands configuration,
should be placed. The chosen deep learning algorithm is Convo- movements, articulation joint, hands orientation, and even facial
lutional Neural Network (CNN) as it provides high accuracy on expression) to recognize sign languages and convey its meaning.
image classification and recognition [13]. The scope of the study Visual-based system has not always been impressive [24], it heavily
includes Filipino number signs (0-9). The purpose of this study is relies on human-curated rules to detect features and in classifica-
to establish a benchmark score for FSL recognition model for static tion of images. The mechanisms were considered inflexible, time
signs and later on progress to real-time SLR system. The project consuming to produce and challenging to improve [25]. However,
focuses on advocating inclusiveness and equal opportunities for thanks to advances in machine learning, deep learning in particular
the deaf community. made computer vision techniques extremely effective in different
applications.
In this part, related works on SLR using computer vision and deep
2 REVIEW OF RELATED LITERATURE learning is briefly reviewed. In Tolentino et al. [26], implemented a
In this section of the paper, the three main established approach in skin-color modeling technique to recognize ASL numbers, alphabets
building SLR systems is discussed. Including glove-based, visual- and static word, and trained the model using CNN. The model has
based and hybrid architecture. yielded an average of 93.67% on validation accuracy. In Sandjaja
and Marcos [27], implemented Hidden Markov Model (HMM) in
building the model capable of recognizing Filipino number signs,
the dataset used is 5000 Filipino number video files. The result
2.1 Glove-based Approach
of their proposed approach has yielded an average accuracy of
In this framework, gloves equipped with various kinds of sensors 85.52% and could still be improved by incorporating more advanced
are utilized to acquire gesture-related data. Over the years, many color tracking algorithm. While Cabalfin et al. [28] implemented
researchers have explored the use of glove-based systems in SLR a Manifold Projection learning model as a recognizer for only 72
in different languages but majority in American Sign Language common Filipino signs. The model has achieved a recognition rate
(ASL) and has achieved significant milestones in the field. Sensory that exceeds 80% and an accuracy as high as 89% in its peak, however,
devices like electromyography (EMG) sensors [14], RGB cameras the system can be considered inadequate for general Filipino sign
[15], Kinect sensors [16], leap motion controllers [17] or their com- recognition system.
binations [18] have been used in SLR development. In Chouhan et
al [19] made use of bend sensor and accelerator to obtain gesture
parameters like hand and finger orientation. Based on their exper- 2.3 Hybrid Approach
imentations from their acquired data, the results have yielded a Hybrid architecture or the combination of both sensor-based and vi-
recognition accuracy as high as 96%. The glove-based architecture sual based approach is also possible. In the works of Ong et al. [29],
has also been applied in FSL recognition. In the paper of Oliva et al introduced a system prototype which combines data gloves with
[20], made use of Kinect V2 to obtain gesture data and trained it computer vision approach to translate FSL and be used in health-
using Dynamic Time Warping and Support Vector Machine algo- care conversations. Viterbi algorithm has been used to find the best
rithm to recognize basic Filipino words. The proposed method has gesture fit for the Hidden Markov Model (HMM). The model was
achieved peak accuracy of 95%. Its limitation, however, is it did not evaluated using the FSL alphabet, numbers and 30 commonly used
tackle finger and facial expression, which is an important factor words in health-care related conversations. The results yielded an
in sign languages. On the other hand, in the works of Rivera and 80.55% recognition rate and could still be improved by making mod-
Ong [21], also made use of Kinect to gather 3D gesture to recognize ification on the sensor prototype and incorporating more robust ML
FSL signs with facial expression. Based on their series of exper- techniques. In Balbin et al [30] also made use of hybrid approach
imentations, the performance is below average and researchers to convert Filipino hand gestures to words; the proponents include
suggests the use of deep learning approaches like Convolutional a colored glove and a web camera. The captured hand images were
Neural Network to improve its performance. used as input for the Kohonen Self-organizing map, a type of neural
In the survey study of Ahmed et al [22] on state-of-the-art sensor- network used for recognizing patterns. Based on their tests on 5
based approach in SLR from 2007 to 2017, mentioned that glove- respondents, the system has achieved 97.6% recognition rate. In the
based approach offers a lot of advantages specifically is its capability paper of Ariesta et al. [31], have conducted a comparative analysis
to acquire gesture data accurately. On the other hand, the challenges on the various methods employed in developing an SLR, based on
on this approach is its infeasibility on practical implementation, the analysis of the results, Hybrid CNN-HMM and fully deep learn-
costly proponents, and inability to capture non-manual signals like ing methods have shown promising outcomes and provide more
facial expression and eye movements. research opportunities. Although hybrid framework can potentially

220
Filipino Sign Language Recognition using Deep Learning ICSET 2021, August 21–23, 2021, Taipei, Taiwan

Figure 2: Hand Gesture being captured using CV-technique


Figure 1: Hand Gesture being captured using CV-technique

deal with the gaps in SLR systems, the entire setup is challenging • Color Image Conversion: The entire dataset has been con-
and the development cost can be expensive. verted to grayscale to restrict it to one layer and lessen the
complexity of the images
3 METHODOLOGY • Gaussian Blur: This filter has been applied to help extract
This section of the paper discusses the development of the Filipino useful features on the hand gesture images
Sign Language (FSL) recognition model using the computer vision • Image Augmentation: The image augmentation technique
approach and deep learning. This section includes the data collec- used is horizontal flipping to consider left hand signs. Also,
tion, data description, data preprocessing, the ResNet architecture, image augmentation expands the dataset and helps avoid
experimental setup and evaluation metrics used. overfitting the model.
• Data Shuffling: The image data was shuffled to ensure that
each image data creates an independent effect on model
3.1 Data Collection
creation. Figure 4 presents the image data after undergoing
The image data used in this study is generated from an individual preprocessing.
signer using computer vision technique and Python implemen-
tation. In this process, a 1080P full-HD web camera is used to
continuously capture the hand gesture images of Filipino number 3.4 Residual Network Architecture
signs from (0-9) patterned from video tutorial of De La-Salle Col- After undergoing preprocessing, a deep convolutional neural net-
lege of Saint Benilde School of Deaf Education and Applied Studies work based on the ResNet architecture is implemented in building
(SDEAS) [32]. The image generated from this process has been the FSL number recognition model. ResNet architecture is proposed
validated by an FSL trained signer. Figure 1 presents the process of by Zhang et al [35], a breakthrough architecture in the field of com-
collecting and generating hand gestures images of FSL numbers. puter vision [36]. It has outperformed previous state-of-the-art
CNN-based architecture like AlexNet [37], GoogleNet [38], and
3.2 Data Description VGG network [39] and tackles degradation problem efficiently even
on higher number of layers.
The total count of data generated is 10,000 images whereas each
As shown in Figure 5, in this architecture, it follows the basic
class of Filipino number signs from 0-9 contained 1000 images,
idea of skipping blocks using the shortcut connections and complies
respectively. The training size per class is set to 1000 images as
with a couple of straightforward design rule: if the same output
this number can be a good representation of the class [33]. Also,
feature map size then the number of filter will remain the same
in the study of Shahinfar et al [34] indicates that in case of limited
and if the size of the feature map is halved then the number of
resources, this amount of data instance per class is sufficient and
filters will be doubled. In case of same input and output dimensions,
can achieve reasonable recognition accuracy. Figure 2 presents the
identity shortcut is used otherwise projection shortcuts are used.
sample image data generated.
The basic building block of ResNet architecture is summarized by
the following equation:
3.3 Data Preprocessing
To prepare the image dataset before undergoing the ResNet architec- y = F (x, {W i} + x
ture, the following preprocessing techniques has been implemented.
Figure 3 presents all the preprocessing techniques utilized.
where x and y are input and output vectors of the convolution layer
• Region of Interest: In capturing the hand gesture images under consideration. The function F (x, Wi) represents the residual
the ROI or the spot where hand is placed is automatically mapping learned. The dimensions of x and F should be equal in
cropped to 128x128 pixels using the python CV technique. equation.

221
ICSET 2021, August 21–23, 2021, Taipei, Taiwan Myron Darrel Montefalcon et al.

Figure 3: Preprocessing Techniques Used

Figure 4: FSL numbers (0-9) image data after preprocessing

Figure 5: Residual Network

3.5 Experimental Setup 4 RESULTS AND DISCUSSIONS


In the initial experimentation, the goal is to establish benchmark 4.1 Experimental Results
scores of FSL number recognition model on FSL numbers dataset.
The performance of the FSL number recognition model using
For this goal, the model will be evaluated based on its recognition
ResNet architecture has been evaluated using the mentioned perfor-
rate and generalization ability. The chosen cross validation tech-
mance metrics. The described dataset has been trained using cross
nique is a train test split of (80:20) wherein n in training set is equal
entropy loss function and Adam Optimizer with a learning rate
to 8000 and n in validation set is 2000.
of 0.001, on one machine with GPU enabled using Google Colab.
The performance of the model has been analyzed by comparison
of models with and without Gaussian blur and by varying epoch
value. The table 1. 2 and 3 shows the result in terms of training and
3.6 Evaluation Metrics validation accuracy. While, Table 4 presents the confusion matrix
The purpose of the proposed study is to recognize FSL numbers of the fine-tuned ResNet-50 model.
signs and convey its corresponding numbers accurately. Thus, for
the purpose of setting the benchmark scores of the FSL recognition
system on FSL numbers, accuracy on both training and validation 4.2 Discussion
are considered in evaluating the model. The equations employ the As shown in Table 1, the best achieved validation accuracy on
following terms: TP, True Positive; TN, True Negative; FP, False the ResNet-18 model without the preprocessing Gaussian Blur is
Positive; and FN, False Negative. 57.40% when the epoch value is equal to 6. In comparison with
the model with Gaussian Blur filter presented in Table 2, there
is a significant improvement on the performance of the model as
the validation accuracy is 72.20% on first epoch, a 14.8% difference.
The ResNet-18 model has achieved a peak validation accuracy of
TP + TN 83.10% and 93.36% on training accuracy on epoch 16. However, the
Accuracy = results from the ResNet-18 model is clear case of overfitting, as
TP + FP + TN + FN

222
Filipino Sign Language Recognition using Deep Learning ICSET 2021, August 21–23, 2021, Taipei, Taiwan

Table 1: ResNet-18 without Gaussian Blur Filter on FSL numbers (0-9)

Epoch Training Accuracy Validation Accuracy


1 0.1916 0.4080
2 0.3675 0.4680
3 0.4384 0.4980
4 0.4960 0.5320
5 0.5048 0.5250
6 0.5263 0.5740

Table 2: ResNet-18 with Gaussian Blur Filter on FSL numbers (0-9)

Epoch Training Accuracy Validation Accuracy


1 0.6203 0.7220
2 0.7979 0.7255
3 0.8349 0.7795
4 0.8521 0.7890
5 0.8690 0.7685
6 0.8689 0.7880
7 0.8812 0.7940
8 0.9219 0.7880
9 0.9234 0.8085
10 0.9235 0.7985
11 0.9260 0.8185
12 0.9234 0.7855
13 0.9300 0.8085
14 0.9278 0.8035
15 0.9256 0.8070
16 0.9303 0.8310
17 0.9326 0.7990
18 0.9329 0.8055
19 0.9320 0.8185
20 0.9336 0.8140

the model performed too well on training data and failed to make
consistent generalization on the validation set. By increasing the
layers and fine-tuning the Residual network architecture, the model
has handled the overfitting problem. In the results shown in Table 3,
the best achieved performance of the fine-tuned ResNet-50 model
on FSL numbers is 92% accuracy on training and 86.7% accuracy on
validation set on the epoch value 15. Even so, the model resulted
to some misclassification specifically on Filipino number class 2
and 6 as seen in the confusion matrix in Table 4. The accuracy is
only 17.5% and this could be due to the similarity on features like
edges and finger orientation of both 2 and 6 Filipino number hand
gestures shown in Figure 6. The misclassification can be solved
by obtaining more training samples and by further adjusting the Figure 6: Misclassified Classes
hyperparameters of the model.
In comparison to other mentioned studies on FSL number recog-
nition, the model is slightly worse compared to the works San-
and data glove prototype. Their proposed method has only achieved
daja and Marcos [27] which implemented an HMM model and has
71.8% on Filipino numbers. However, on the studies which utilized
achieved 85% accuracy on Filipino numbers. On the other hand, the
commercially available gloves like Kinect on FSL recognition [20],
model has achieved better accuracy to the works of Ong et al [38]
their glove-based approach has produced significantly better accu-
which implemented a hybrid approach combining computer vision
racy on recognizing Filipino numbers. The proposed approach in

223
ICSET 2021, August 21–23, 2021, Taipei, Taiwan Myron Darrel Montefalcon et al.

Table 3: ResNet-50 with Gaussian Blur Filter on FSL numbers (0-9)

Epoch Training Accuracy Validation Accuracy


1 0.5229 0.6625
2 0.7396 0.7245
3 0.7933 0.7830
4 0.8029 0.7745
5 0.8316 0.7965
6 0.8210 0.7808
7 0.8340 0.7829
8 0.8744 0.8310
9 0.8988 0.8219
10 0.8985 0.8450
11 0.9036 0.8392
12 0.8979 0.8419
13 0.9011 0.8463
14 0.9160 0.8532
15 0.9200 0.8670
16 0.9201 0.8290
17 0.9189 0.8115
18 0.9233 0.8040
19 0.9214 0.8005
20 0.9336 0.8140

Table 4: Confusion Matrix of the Fine-tuned ResNet-59 Model

0 1 2 3 4 5 6 7 8 9
0 200 0 0 0 0 0 0 0 0 0
1 16 174 0 0 0 0 7 0 3 0
2 0 39 152 0 0 0 1 0 5 3
3 0 0 0 193 0 7 0 0 0 0
4 0 0 1 0 196 0 0 3 0 0
5 0 0 0 3 0 195 0 0 0 2
6 0 0 154 0 0 0 35 10 0 1
7 0 0 0 0 3 0 12 185 0 0
8 0 0 0 0 0 0 0 23 167 10
9 0 1 2 3 4 1 1 19 6 163

this study using visual-based techniques and CNN ResNet architec- epoch value is equals to 15. The performance of the model is com-
ture was able to yield at best an 83.10% validation accuracy. With parable to other mentioned literature, even so, there remains an
only utilization of minimal resources like web camera as its main opportunity to enhance the accuracy of the model. The results of
proponent, the model was able to achieve decent accuracy without this study will be used as baselines score for further improvement
that aid of data gloves of the system. For future direction, other deep learning architecture
like HMM or the hybrid HMM-CNN will be explored for continuous
Filipino signs. Also, more data will be collected to enable recog-
nition on sign language of Filipino alphabets, basic phrases and
5 CONCLUSION AND FUTURE WORK common greetings. Lastly, optimization of the model is important
In this paper, the development of FSL number recognition system for faster computations to progress from static to the implementa-
using computer vision and deep learning has been explored. The tion of real-time FSL recognition system.
model has been trained on CNN Resnet architecture using the image
data generated from an individual signer. The FSL number recog-
nition model developed is capable of recognizing static Filipino ACKNOWLEDGMENTS
number signs which ranges from 0-9. Based on experimentation, We would like to express our deepest gratitude to Ms. Angelica De
the best achieved performance is on fine-tuned Resnet-50 model La Cruz, for her continues guidance in both research writing and
which obtained a validation accuracy as high as 86.7% when the implementation. Also, huge appreciation to National University

224
Filipino Sign Language Recognition using Deep Learning ICSET 2021, August 21–23, 2021, Taipei, Taiwan

CCIT department for providing the necessary resources through 2014 IEEE Global Humanitarian Technology Conference-South Asia Satellite
the HI-DSP laboratory. (GHTC-SAS) (pp. 105-110). IEEE.
[20] Oliva, K. E., Ortaliz, L. L., Tobias, M. A., & Vea, L. Filipino Sign Language Recog-
nition for Beginners using Kinect. In 2018 IEEE 10th International Conference
REFERENCES on Humanoid, Nanotechnology, Information Technology, Communication and
[1] F. R. Session, A. Pacific, and S. Africa, Senat’2117, 2014, pp. 1–3, Control, Environment and Management (HNICEM) (pp. 1-6). IEEE.
[2] RA 11106 – An Act Declaring the Filipino Sign Language as The National Sign [21] Rivera, J. P., & Ong, C. (2018). Facial expression recognition in filipino sign
Language of The Filipino Deaf And The Official Sign Language Of Government In language: Classification using 3D Animation units. In Proc. the 18th Philippine
All Transactions Involving The Deaf, And Mandating Its Use In Schools, Broadcast Computing Science Congress (PCSC 2018) (pp. 1-8).
Media, And Workplaces [Link] [22] Ahmed, M. A., Zaidan, B. B., Zaidan, A. A., Salih, M. M., & Lakulu, M. M. B. (2018).
ra-11106/ A review on systems-based sensory gloves for sign language recognition state of
[3] Mendoza, A. (2018, October 18). The sign language unique to Deaf Filipinos. the art between 2007 and 2017. Sensors, 18(7), 2208
Retrieved from [Link] [23] Babich, N. (2020, July 28). What Is Computer Vision; How Does it Work? An
[Link]?fbcid Introduction. Retrieved from [Link]
technology/what-is-computer vision-how-does-it-work/
[4] Mairona-Basas, M., & Pagliario, C. (2014, March 14). Technology Use Among [24] Editor, P. (2018, August 15). A brief history of Computer Vision and AI Image
Adults Who Are Deaf and Hard of Hearing: A National Survey, The Journal of Recognition. Retrieved from [Link]
Deaf Studies and Deaf Education. (Vol. 19, pp. 400-410). Retrieved from https: history-computer-vision vertical-ai-image-recognition/
//[Link]/jdsde/article/19/3/400/2937196. [25] Zeller, M. (n.d.). What Is Computer Vision? Why Deep Learning Changed It All.
[5] Wadhawan, A., Kumar, P. Sign Language Recognition Systems: A Decade Sys- Retrieved from [Link]
tematic Literature Review. Arch Computat Methods Eng (2019). [Link] [26] Tolentino, L. K. S., Juan, R. O. S., Thio-ac, A. C., Pamahoy, M. A. B., Forteza, J.
10.1007/s11831-019-09384-2 R. R., & Garcia, X. J. O. (2019). Static Sign Language Recognition Using Deep
[6] Martinez, L., & Cabalfin, E. P. (2008, November). Sign language and computing Learning. International Journal of Machine Learning and Computing, 9(6).
in a developing country: A research roadmap for the next two decades in the [27] Sandjaja, I. N., & Marcos, N. (2009, August). Sign language number recognition.
Philippines. In Proceedings of the 22nd Pacific Asia Conference on Language, In 2009 Fifth International Joint Conference on INC, IMS and IDC (pp. 1503-1508).
Information and Computation (pp. 438-444). IEEE.
[7] Bragg, D., Koller, O., Bellard, M., Berke, L., Boudreault, P., Braffort, A., ... & Vogler, [28] Cabalfin, E. P., Martinez, L. B., Guevara, R. C. L., & Naval, P. C. (2012, Novem-
C. (2019, October). Sign Language Recognition, Generation, and Translation: ber). Filipino sign language recognition using manifold projection learning. In
An Interdisciplinary Perspective. In The 21st International ACM SIGACCESS TENCON 2012 IEEE Region 10 Conference (pp. 1-5). IEEE.
Conference on Computers and Accessibility (pp. 16-31). [29] Ong, C., Lim, I., Lu, J., Ng, C., & Ong, T. (2018). Sign-Language Recognition
[8] Elakkiya, R. (2020). Machine learning based sign language recognition: a re- Through Gesture & Movement Analysis (SIGMA). In Mechatronics and Machine
view and its research frontier. Journal of Ambient Intelligence and Humanized Vision in Practice 3 (pp. 235-245). Springer, Cham.
Computing, 1-20. [30] Balbin, J. R., Padilla, D. A., Caluyo, F. S., Fausto, J. C., Hortinela, C. C., Manlises,
[9] Al-Ahdal, M. E., & Nooritawati, M. T. (2012, March). Review in sign language C. O., ... & Ventura, L. T. (2016, November). Sign language word translator using
recognition systems. In 2012 IEEE Symposium on Computers & Informatics (ISCI) Neural Networks for the Aurally Impaired as a tool for communication. In 2016 6th
(pp. 52-57). IEEE. IEEE International Conference on Control System, Computing and Engineering
[10] Bantupalli, K., & Xie, Y. (2018, December). American sign language recognition (ICCSCE) (pp. 425-429). IEEE.
using deep learning and computer vision. In 2018 IEEE International Conference [31] Ariesta, M. C., Wiryana, F., & Kusuma, G. P. (2018). A Survey of Hand Gesture
on Big Data (Big Data) (pp. 4896-4899). IEEE. Recognition Methods in Sign Language Recognition. Pertanika Journal of Science
[11] Culver, V. R. (2004, October). A hybrid sign language recognition system. In & Technology, 26(4).
Eighth International Symposium on Wearable Computers (Vol. 1, pp. 30-33). [32] SDEAS, B. (Director). (2013, May 7). FSL Numbers [Video file]. Retrieved from
IEEE. [Link] =
[12] LaViola, J. (1999). A survey of hand posture and gesture recognition techniques BenildeSDEAS
and technology. Brown university, providence, ri, 29. [33] Warden, P. (2017, December 14). How many images do you need to train a
[13] Maladkar, K. (2018, January 25). Overview Of Convolutional Neural Net- neural network? Retrieved from [Link]
work In Image Classification. Retrieved from [Link] images-do-you-need-to-train-a-neural-network/
convolutional-neural-network-image-classification-overview/ [34] Shahinfar, S., Meek, P., & Falzon, G. (2020). “How many images do I need?”
[14] Wu, J., Tian, Z., Sun, L., Estevez, L., & Jafari, R. (2015, June). Real-time American Understanding how sample size per class affects deep learning model performance
sign language recognition using wrist-worn motion and surface EMG sensors. metrics for balanced designs in autonomous wildlife monitoring. Ecological
In 2015 IEEE 12th International Conference on Wearable and Implantable Body Informatics, 57, 101085.
Sensor Networks (BSN) (pp. 1-6). IEEE. [35] He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image
[15] Martinez, D. C. (2012). Sign language translator using Microsoft Kinect XBOX 360. recognition. In Proceedings of the IEEE conference on computer vision and
In Master tesis ERASMUS MUNDUS Masters Vibot. Departamento de ingeniería pattern recognition (pp. 770- 778)
electrónica y ciencias de la computación [36] Mishra, V. (2020, November 16). CNN Architecture: How ResNet works and why?
[16] Dong, C., Leu, M. C., & Yin, Z. (2015). American sign language alphabet recogni- Retrieved from [Link]
tion using microsoft kinect. In Proceedings of the IEEE conference on computer resnet-works-andwhy-1c197b8eba34#
vision and pattern [37] Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). Imagenet classification with
[17] Toghiani-Rizi, B., Lind, C., Svensson, M., & Windmark, M. (2017). Static gesture deep convolutional neural networks. Advances in neural information processing
recognition using leap motion. arXiv preprint arXiv:1705.05884. systems, 25, 1097- 1105
[18] Vogler, C., & Metaxas, D. (2003, April). Handshapes and movements: Multiple- [38] Simonyan, K., & Zisserman, A. (2014). Very deep convolutional networks for
channel american sign language recognition. In International Gesture Workshop large-scale image recognition. arXiv preprint arXiv:1409.1556.
(pp. 247-258). Springer, Berlin, Heidelberg. [39] Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., ... & Rabinovich,
[19] Chouhan, T., Panse, A., Voona, A. K., & Sameer, S. M. (2014, September). Smart A. (2015). Going deeper with convolutions. In Proceedings of the IEEE conference
glove with gesture recognition ability for the hearing and speech impaired. In on computer vision and pattern recognition (pp. 1-9).

225

View publication stats

Common questions

Powered by AI

Advanced learning techniques like CNN-HMM hybrids improve sign language recognition by combining the strengths of both models. The Convolutional Neural Network (CNN) processes and identifies intricate patterns in images, while the Hidden Markov Model (HMM) manages temporal dynamics of gesture sequences. This combination allows for better recognition of complex sign language features compared to traditional methods, leading to higher accuracy and a more robust model capable of dealing with the variation inherent in human sign language .

Convolutional Neural Networks (CNN) significantly enhance static Filipino Sign Language recognition by providing a robust architecture for image classification. CNNs can learn patterns and features from images, allowing them to achieve high accuracy in translating static gestures like Filipino number signs into text. This deep learning model is particularly well-suited for tasks involving complex patterns in image data, making it a crucial tool for effective sign language recognition .

The main advantage of the glove-based approach in sign language recognition systems is its ability to acquire gesture data accurately due to sensors like electromyography (EMG), which provide detailed information about hand and finger movements . However, the challenges include its high cost, infeasibility for practical implementation, and inability to capture non-manual signals such as facial expressions and eye movements, which are crucial in sign languages .

The computer vision approach significantly contributes to inclusivity and equal opportunities for the deaf community by lowering the barriers to communication through cost-effective and accessible sign language recognition systems. Unlike glove-based methods, vision-based systems require only cameras, reducing the overall development and implementation cost. This affordability and accessibility enable wider adoption, which supports inclusive communication by translating sign language into text, thereby facilitating interactions and reducing the isolation experienced by deaf individuals .

Methods reliant on Kinect sensors for Filipino Sign Language recognition face the limitation of not adequately capturing non-manual signals such as facial expressions, which are important in sign languages. Although Kinect can gather 3D gesture data, the performance can be suboptimal, as demonstrated by below-average results when used for FSL signs recognition . Adopting deep learning approaches, such as CNNs, is recommended to overcome these limitations and improve recognition performance .

Deploying sign language recognition systems relying on computer vision technology requires addressing several ethical considerations. These include ensuring privacy for users whose video data is captured and processed, avoiding potential biases in AI models which may result in inaccuracies across different sign language features or skin tones, and the ethical use of data gathered from users, particularly in maintaining confidentiality and securing informed consent. Consideration must also be given to how these systems can be accessible and equitably available to different communities, thus ensuring that they support rather than hinder inclusivity for individuals they are designed to assist .

Using pre-trained deep learning models like ResNet in developing sign language recognition systems provides significant benefits, including reduced training time and improved model accuracy. ResNet, known for its depth and the ability to mitigate vanishing gradient problems, can enhance feature extraction from complex image data due to prior training on large datasets . This transfer learning allows developers to leverage established architectures and pre-trained weights for efficient and effective SLR model development, particularly beneficial when working with limited data resources .

Deep learning enhances the effectiveness of visual-based sign language recognition by providing flexibility and improving accuracy in image recognition tasks. While traditional visual systems relied on human-curated rules that were often inflexible and time-consuming, deep learning techniques, such as Convolutional Neural Networks (CNN), have made it possible to automatically extract and classify visual features with high accuracy and efficiency .

Relying solely on visual-based sign language recognition systems presents limitations related to inflexible rule-based image interpretation and classification processes. Before the integration of deep learning, visual-based methods heavily depended on predefined human-curated rules, making them cumbersome and time-consuming to adapt or improve . These systems can struggle to manage real-world variations in sign gestures unless enhanced by more flexible models such as deep learning techniques .

The hybrid approach combines the advantages of both glove-based and visual-based systems to address their respective limitations. It can capture detailed gesture data through sensors like data gloves while also utilizing computer vision to recognize non-manual signals, such as facial expressions. This approach provides a more comprehensive recognition system capable of dealing with the gaps left by each method alone, such as the high cost of gloves and their inability to capture facial expressions . However, the hybrid method often involves complex setups and can be expensive to develop .

You might also like