Sign Language Identifier Project Report
Sign Language Identifier Project Report
Submitted by
101703545 Simran
CPG No. 76
Professor
December, 2020
ABSTRACT
Deaf people do not have many options for communicating with a hearing person, and
all of the alternatives do have major flaws. Interpreters aren't usually available, and also
could be expensive. The pen and paper approach is also a bad idea as it is not practical.
So a framework is required that provides a helping-hand for deaf and dumb people to
communicate using sign language.
Deaf and Dumb people simply just have to perform the actions of Indian Sign Language
and in front of the webcam which will be converted to text; and speech to ISL. Hence,
the communication gap between the normal people and Deaf and Dumb people will be
reduced. They will be also able to express themselves through this, hence reducing the
feeling of isolation felt by them.
i
DECLARATION
We hereby declare that the design principles and working prototype model of the
project entitled Sign Language Identifier is an authentic record of our own work carried
out in the Computer Science and Engineering Department, TIET, Patiala, under the
guidance of Dr. Seema Bawa, Professor, T.I.E.T. during 6th semester (2020).
101703545 Simran
ii
ACKNOWLEDGEMENT
We would like to express our thanks to our mentor “Dr. Seema Bawa Mam”. She has
been of great help in our venture, and an indispensable resource of technical knowledge.
She is truly an amazing mentor to have. We would also like to express our thanks to
“Ms. Sawinder Kaur”, PHD student T.I.E.T. Her guidance was of immense help to us.
We are also thankful to, “Dr. Maninder Singh”, Head, Computer Science and
Engineering Department, entire faculty and staff of Computer Science and Engineering
Department, and also our friends who devoted their valuable time and helped us in all
possible ways towards successful completion of this project. We thank all those who
have contributed either directly or indirectly towards this project.
Lastly, we would also like to thank our families for their unyielding love and
encouragement. They always wanted the best for us and we admire their determination
and sacrifice.
101703545 Simran
iii
TABLE OF CONTENTS
ABSTRACT..................................................................................................................i
DECLARATION..........................................................................................................ii
ACKNOWLEDGEMENT..........................................................................................iii
LIST OF TABLES......................................................................................................viii
LIST OF FIGURES......................................................................................................x
LIST OF ABBREVIATIONS......................................................................................xii
CHAPTER.......................................................................................................Page No.
1. Introduction
1.1 Project Overview..............................................................................1
1.1.1 Technical Terminology………....……………………………….1
1.1.2 Problem Statement .......................................................................1
1.1.3 Goal……….. …………………………………………………...3
1.1.4 Solution…….…………………………………………………...3
1.2 Need Analysis ……………………………………….…………………...4
1.3 Research Gaps …………………………………………………………....4
1.4 Problem Definition and Scope……………………………………………8
1.5 Assumptions and Constraints …………………………………………....8
1.6 Standards………………………………………………………………....9
1.7 Approved Objectives …………………………………………………….9
1.8 Methodology …………………………………………………………….10
1.9 Project Outcomes and Deliverables….……...…………………………..13
1.10 Novelty of Work………………………………………………………..14
2. Requirement Analysis
2.1 Literature Survey………………………………………………………...15
2.1.1 Theory Associated With Problem Area………………………..15
2.1.2 Existing Systems and Solutions……………….…....................16
2.1.3 Research Findings for Existing Literature……………………..17
2.1.4 Problem Definition……………………………………………..19
iv
2.1.5 Survey of Tools and Technologies Used…………………….....20
2.2 Standards………………………………………………………………....27
2.3 Software Requirement Specification……………………………………..28
2.3.1 Introduction………………………………………………….....28
[Link] Purpose……………………………………………….28
[Link] Intended Audience and Reading Suggestions……......28
[Link] Project Scope …………………………………….......30
2.3.2 Overall Description…………………………………………….30
[Link] Product Perspectives………………………………................30
[Link] Product Features……………………………………...32
2.3.3 External Interface Requirements……………………………….35
[Link] User Interfaces………………………………………..35
[Link] Hardware Interfaces………………………………….36
[Link] Software Interfaces……………………………….…..36
2.3.4 Other Non-functional Requirements……………………….…..38
[Link] Performance Requirements……………………….….38
[Link] Safety Requirements……………………………..…..38
[Link] Security Requirements……………………………….39
2.4 Cost Analysis……………………………………………………….…….39
2.5 Risk Analysis……………………………………………………….…….40
3. Methodology Adopted
3.1 Investigative Techniques………………………………………………....41
3.2 Proposed Solution………………………………………………………..43
3.3 Work Breakdown Structure………………………………………….…...46
3.4 Tools and Technology…………………………………………………....50
4. Design Specifications
4.1 System Architecture……………………………………………………...51
4.2 Design Level Diagrams……………………………………………….….52
4.2.1 Component Diagram……………………………………….…..52
4.2.2 Data Design………………………………………………….…55
v
4.3.1 Use Case Diagram……………………………………………...63
4.3.2 Use Case Template……………………………………………..65
4.3.3 Activity Diagram……………………………………………….69
4.3.4 Sequence Diagram……………………………………………...72
4.3.5 Class Diagram………………………………………………….74
4.3.6 State Chart Diagram…………………………………………....76
vi
7. Project Metrics
7.1 Challenges Faced......................................................................................116
7.2 Relevant Subjects.....................................................................................117
7.3 Interdisciplinary Knowledge Sharing......................................................118
7.4 Peer Assessment Matrix..........................................................................118
7.5 Role Playing and Work Schedule............................................................119
7.6 Student Outcomes Description and Performance Indicators……..........121
7.7 Brief Analytical Assessment..................................................................124
APPENDIX A: References……………………………………………..………..126
APPENDIX B: Plagiarism Report……………………………………………...128
APPENDIX C: Technical Course Writing Proof…………………....………….129
vii
LIST OF TABLES
8 Release Details 33
9 Features 34
viii
22 Work Accomplished 111-112
ix
LIST OF FIGURES
6 Gantt Chart 53
11 ER Diagram 63
16 Sequence Diagram 74
17 Class Diagram 77
21 code 1 83
x
22 code 2 83
23 code 3 84
24 code 4 84
25 code 5 85
26 code 6 85
27 code 7 86
28 code 8 86
29 code 9 86
30 code 10 86
31 Welcome Screen 89
32 Login Screen 90
33 Sign Up Screen 90
34 About Menu 91
35 About Menu 91
37 Feedback 92
xi
LIST OF ABBREVIATIONS
DL Deep Learning
xii
INTRODUCTION
The language barrier or communication barrier has been overcome by making learning
a common language or using language to language digital translator. But for the people
who are not able to hear and speak, such technologies and ideas make no difference.
Therefore “Sign Language Identifier” is an interface that will get rid of this
communication divide.
Sign language is any means of communication that uses body movements, especially
hands and arms when spoken communication is impossible or not desirable. Therefore
sign language is the bridge gap wherever vocal communication is not possible. Only
5% of the population is hearing impaired, so most of the people don’t learn sign
language and the medium of its education is also less. Therefore the communication
gap between the two communities still persists. Due to this many deaf and dumb people
were left uneducated and jobless. The pen and paper approach is a bad idea for
communication as it is very time consuming and also fails while communicating with
uneducated people. The schools are made for the deaf and dumb people where they are
tought sign language and they learn many skills and do activities using sign language.
But this defines the boundary for them and they are left segregated from the rest of the
world. This not only affects them but also their country, as it loses many great thinkers
and intellectuals which could have helped in uplifting the nation just because of the
communication barrier.
Medical field has worked on helping the deaf and dumb people. Cochlear implant is the
cure fort the deaf and dumb children under the age of six but it also has strict selection
process for the surgery and is also expensive.
1
Then there came the person interpreters who translate the communication for hearing
impaired people and became their speakers. But interpreters aren't usually available,
and also could be expensive. So Deaf people do not have many options for
communicating with a hearing person, and all of the alternatives do have major flaws.
So a framework needs to be built that provides a helping-hand for deaf and dumb people
to communicate using sign language. This would be a user-friendly environment for
the user by providing text output for a sign gesture input and speech output for ISL
input. There are some frameworks that are built for the same purpose, but they have
many flaws which make them imperfect to use.
Basic sign language system is based on the 5 parameters and they are hand and head
recognition, hand and head orientation, hand movement, shape of hand and location of
hand and head. Capture of hand movements is via webcam. Complex video
backgrounds in another gap in sign language recognition where it poses a major
challenge to extract signers hand shapes accurately. This SLR system focuses on putting
a simple constant background with the signer's shirt matching the background. In
cluttered video backgrounds tracking hands has become simpler but tracking each
finger movements remains quite a challenging task. So by using different technologies
of machine and deep learning and computer vision in “Sign Language Identifier” will
eradicate the above complexities.
This system will be user friendly and all the complexities in recording will be solved
using different technologies. Therefore, users will not have to be careful or take certain
measures to use the system. The system is based on Indian Sign Language, it is basically
made for the indian users or for those who use indian sign language. Previously no
system was designed for ISL, in order to use it at work or while communication people
usually have to learn other sign language. It has been observed that people who learn
sign language during adolescence or teenage are not able to grap it better than those
who learnt it when they were born. The World Bank strongly stated that children whose
first language is not used at school experience a low level of learning and are much less
likely to contribute to the country's economy and intellectual development. The main
purpose of education is to contribute skilled manpower and it's hard to get that kind of
man if they learn in their second language. Above all language is the instrument used
2
by people to think, analyse and relate to other people and environment. Language is
therefore the basis of self awareness, for sensing of meaning and for intellectual
development. Therefore switching from one sign language to another will degrade their
performance at work, their creativity and they will find it difficult to express themselves
even if they are using a sign language translator.
Therefore, The Sign Language Interpreter will provide the people using indian sign
language, a user friendly interface that will help them to have communication with
people who have no knowledge about the sign language in a much more natural way
and this will boost their confidence and productivity. Following are benefits that users
will get.
1. Since deaf people are usually deprived of normal communication with other people,
they have to rely on an interpreter or some visual communication. Now the interpreter
can not be available always, so this project can help eliminate the dependency on the
interpreter.
2. The system can be extended to incorporate the knowledge of facial expressions and
body language too so that there is a complete understanding of the context and tone of
the input speech.
3. A progressive web based application will increase the reach to more people.
4. Integrating hand gesture recognition system using computer vision for establishing
2-way communication system.
5. It will be a scalable project which can be extended to capture the whole vocabulary
of ISL through manual and non manual signs.
3
1.1.1 Technical terminology
I. Machine Learning: To get better insight about how “sign language
identification” problems can be looked at and hence, solved.
II. Deep Learning: To apply the most suitable deep learning algorithm (CNN) for
sign language identification, and also knowledge of optimizers, loss functions.
III. Natural Language Processing: For matching gif or image corresponding to
speech from a bag of words.
IV. Image Processing: To capture the image from the camera and further removal
of noise using noise reduction filters and further enhancing the input for better
output and ease of prediction for the trained model.
V. Software Engineering: Software Development Lifecycle, Preparation of SRS,
Working as per Scrum model, Shaping functional and non-functional
requirements, Understanding and communication of ideas through UML
Diagrams.
VI. CNN: Convolution Neural Network is a class of deep neural networks, applied
to analyzing visual imagery.
VII. OpenCV: OpenCV (Open Source Computer Vision Library) is an open source
computer vision and machine learning software library. It provides a common
infrastructure for computer vision applications and to accelerate the use of
machine perception.
VIII. Tkinter: Tkinter is the standard GUI library for Python. Python when combined
with Tkinter provides a fast and easy way to create GUI applications. Tkinter
provides a powerful object-oriented interface to the Tk GUI toolkit.
IX. Flask: Flask is a web application framework written in Python.
4
1.1.2 Problem statement
To design “Sign Language Identifier”, a web app for deaf and dumb people is to be
deigned keeping in view that sign language is used by the hearing impaired people for
their communication. These kinds of systems create a gap between normal people and
hearing impaired people. Sign Language Identifier is a helping aid for speech impaired
to communicate with the rest of the world using sign language. This application will
convert the sign language gestures to text/speech and the speech/text back to sign
language.
1.1.3 Goal
1. To survey and analyze the current status of usage and effectiveness of Indian Sign
Language (ISL) solutions.
2. To propose and implement an interactive solution that will convert sign language to
text and speech and vice versa.
A. To collect ISL dataset and create ISL dictionary for commonly used words and
to preprocess them.
B. To apply Deep Learning Model for training on collected dataset.
3. To test and demonstrate the usability and effectiveness of the proposed solution for
targeted user groups (Deaf and dumb persons).
1.1.4 Solution
Solution to the problem of communication of the deaf and dumb people is to build a
computer that can be programmed in such a way that it can translate sign language to
text format; and speech to ISL format, the difference between the normal people and
the deaf community can be minimized. Indian Sign Language (ISL) Interpretation
system is a good way to help the Indian hearing impaired people to interact with normal
people with the help of a computer. Compared to other sign languages, ISL
interpretation is an area which gained less attention by the researchers since American
Sign Language (ASL) has most of the signs, single handed and thus, complexity is less.
Also, ASL already has a standard database that is available for use. When compared
with ASL, Indian Sign Language relies on both hands and thus, an ISL recognition
system is more complex.
5
1.2 Need Analysis
Sign Language is the most natural and expressive way for hearing impaired people.
People who are not deaf, never try to learn sign language for interacting with the deaf
people. India does not have a single well-reputed junior or senior college for them. The
deaf and mute people have to drop out of education or take admission for long distance
courses or they have to take admission in colleges meant for those without any physical
challenges. In regular colleges, they face a series of hurdles. To overcome low esteem
born out of their disability while studying in colleges, where most of the students have
no physical challenges, is pretty difficult. And most of them do not seek admission or
prefer to drop out after joining. Besides, none of these colleges avail interpreters or
lecturers who can teach physically challenged students in sign language. The students,
therefore, have to get help from students or close friends/relatives to study. They aren’t
able to become self dependent in most of the cases and rely on mercy of others.
Every day thousands of local businesses around the globe face problems with providing
their services to them. People who are talented and capable of bringing a good change
in society are hindered by their disability, which is not just their loss but even nation’s
loss.
Not even here, everywhere they feel the disability they are born with. This leads to
isolation of the deaf people. But if the computer can be programmed in such a way that
it can translate sign language to text format; and speech to ISL format, the difference
between the normal people and the deaf community can be minimized. Indian Sign
Language (ISL) Interpretation system is a good way to help the Indian hearing impaired
people to interact with normal people with the help of a computer. Compared to other
sign languages, ISL interpretation is an area which gained less attention by the
researchers since American Sign Language (ASL) has most of the signs, single handed
and thus, complexity is less. Also, ASL already has a standard database that is available
for use. When compared with ASL, Indian Sign Language relies on both hands and
thus, an ISL recognition system is more complex.
6
1.3 Research Gaps
Sign language plays a vital role in the life of any auditory impaired individual but it
does not get the same status as any other language. Hand gestures are most widely used
as a medium of sign language based communication framework among various forms
of gestures. Recognizing gestures under dynamic parameters such as lighting
conditions, multiple hands as background, left handed person, right handed person, size
of finger etc. are thought provoking fields in automate Indian sign language recognition
. A review is based on datasets of hand gestures, alphabets, digits, words and sentences
with various real time conditions used by different researchers at national and
international level is summarized in Table 1.
7
head ● Indian Sign Language Computer Science, Dublin,
coordination contains gestures July, Technical Report No.
involving both hands. TCD-CS-93-11, 1993
Therefore no system was Mandeep Kaur Ahuja,
made to identify it. Amardeep Singh, “Hand
● Sign language involve the Gesture Recognition Using
words or sentences made PCA”, International Journal
using both hands, hand and of Computer Science
head coordination, these Engineering and Technology
words and sentences could (IJCSET ), Volume 5, Issue
not be identified 7, pp. 267-27, July 2015
8
Rigid system of ● In order to record the Alina Kuznetsova, Laura
maintaining the video, specific distance Leal-Taix´
specific distance should be maintain from e, Bodo Rosenhahn
from the webcam the camera which give rise Institute fuer
while recording to the complication in Informationsverarbeitung,
recording the video Leibniz University
● It is also not possible for Hannover
the user to calculate the Appelstr. 9A, Hannover,
distance very time before 30167, Germany
recording the video
● This will give rise to [Link] [Link]
inappropriate results and Kumar Segment, Track,
also efficiency of the Extract, Recognize and
system will decrease Convert Sign Language
Videos to Voice/Text
9
1.4 Problem Definition and Scope
To bridge the gap between deaf and dumb and the rest , the main purpose of this project
is to design a web app that can convert sign language to english text; and speech to ISL.
Sign Language is an incredible advancement that has grown over the years. It is
primarily used by deaf and dumb to communicate with rest of the world. Unfortunately,
not everyone can understand and interpret sign language which leads to a
communication gap between deaf and dumb and the rest. The main purpose of the
project is to eliminate this communication gap so that sign language can be understood
by common people without the help of any interpreter. This project aims at building a
web based application that can convert sign language to text;and speech to ISL. Anyone
willing to communicate through sign language will be able to login through the app.
Hand gestures of the signer will be captured by the camera of the device which will be
further converted to text and speech recorded by the audio will be converted to ISL.
Signer will be at a fixed distance from the camera. The dataset for training the model
will contain gestures from Indian Sign Language only.
5. User has a steady pace while making signs so that the camera could detect
the signs properly.
10
6. The signs are made in a controlled environment keeping a fixed background.
1.6 Standards
1. ISL(Indian Sign Language) standard: Used to create Indian Sign Language
dictionary.
2. Web 2.0 : Internet technology to create online applications that behave
dynamically.
3. Camera measurement standards :
3.1 Resolution Measurements - ISO 12233
3.2 Noise Measurements – ISO 15739
3.3 Sensitivity Measurements – ISO 12232
3.4 OECF (tone reproduction) Measurements – ISO 14524
3.5 Chromatic Displacement Measurements – ISO 19084
3.6 Color Characterization Test Procedures – ISO 17321
3.7 Image Stabilization Measurements – ISO 20954
3.8 Low Light Measurements – ISO 19093
3.9 Shooting Time Lag Measurements – ISO 15781
3.10 Shading Measurements – ISO 17957
11
1.7 Approved Objectives
1. To create a web app to enable Indian Sign Language to be converted into
English text; and speech to Indian Sign Language.
2. To perform tasks for conversion of Indian Sign Language to English
2.1 Capture Video
2.2 Generate text
2.3 Display text
3. To perform the task for conversation of Speech to ISL
3.1 Record the voice
3.2 Generate the ISL gif
3.3 Display the ISL gif
4. To design and implement the web application.
5. To test and validate the web app as per functionality and user experience.
12
1.8 Methodology Used
ISL to text
For the sign language recognition system, a vision based approach will be used. The
system will contain four modules/derived algorithms: preprocessing and segmentation,
feature extraction, sign recognition and sign to text. Indian Sign Language is the area
of concern. The signs will be captured using a webcam where the signer must be directly
in front of the webcam. Then the output will be in the form of text. Following is
methodology for sign language to text.
13
1.8.1 Preparing the dataset
A dataset of 26 English language alphabets in Indian Sign Language will be created.
The images in the dataset are to be segmented. The dataset will then be divided for
training and testing. A dictionary for ISL will be created so that idea extends to
sentences as well.
a. Segmentation: Here the image is converted into small segments in order to get
more accurate image attributes. Here the segmentation of hands is carried out to
separate objects and the background. To make segmentation more robust, so
that the feature extraction and sign recognition of the image is accurate, the
BGR image is converted to HSV, then fed into inRange() function which
performs segmentation.
14
1.8.5 Template Matching and Sign Recognition
The image so generated is matched with labels with help of predict_class() method of
CNN model. Say the sign for ‘A’ is performed, then the class is ‘A’ which is the folder's
name in the dataset.
Speech to ISL
Firstly a dataset of alphabets and gifs is created. Post that, the coding is done for the
conversion of speech to ISL. Firstly, the microphone is activated and input voice is
thresholded. Then, the voice recognition takes place using Google API. The spoken
word is processed to text and then text is searched with labelled folders of the dataset.
If word exists, then corresponding GIF is displayed, else the plot of alphabets of which
letter is composed of is displayed on the screen.
The deaf and dumb person will simply perform the actions according to sign language
and that will be converted to text which could be easily understandable by the other
person. Hence this app will assist them and translate the sign language to text; and
speech to ISL as fast as the person speaks and not make them feel isolated from society.
15
1.10 Novelty of Work
The novelty of the work lies in the fact that the sign language used is Indian Sign
Language (ISL). Though a lot of work is done in American Sign Language (ASL), ISL
has a bit less research when compared with ASL and the work done in ISL is restricted
to 26 alphabets of English or maximum to maximum, Classification and Feature
Extraction takes place. Only a few out of which employ Image Processing tasks to make
it efficient. Hardly, any work is done by combining all techniques of Image Processing,
Machine Learning, Neural Network and Natural Language Processing (NLP).
Generally, only one or combination of two techniques is used/ solution is hardware
based/ glove based, which makes it less easily accessible, lesser ease of avail, so it’s
not able to reach more people, who are in need. Also, lesser systems are practical and
if they are practical, they have less accuracy, so the “Sign Lnaguage Interpreter”, works
on making it more feasible and practical keeping in view the needs of deaf and dumb.
In SLI various Image Processing techniques are used like applying different filters to
extract frames out of videos, followed by image pre processing operations like Image
Segmentation, Morphological Processing (Dilation and Erosion), then some of
Machine Learning and Deep Learning tasks like Template Matching, Feature
Extraction, then NLP tasks and at last SLI gets a text, which brings in novelty of the
project.
16
REQUIREMENT ANALYSIS
Later when the AI era came, many apps, system for the same were made but none could
succeed to provide real time communication. It either failed in having a proper dataset
or had a lot of constraint while using the system. Nowadays a lot of efforts are made to
make such devices that overcome all the constraints and make a user friendly interface
for the hearing impired people.
17
with great solutions but it has not been implemented yet. The following are the existing
systems and solutions developed for sign language identifiers.
MotionSavvy: It is a tablet app that understands sign language. It is the software that
converts the american sign language to text or voice. It also has a voice recognition
system which helps hearing people to convert their voice to American Sign
Language(ASL). It currently identifies 100 words. There are many signs in american
language itself. The MotionSavvy case embeds the Leap, and the MotionSavvy
software leverages the Leap’s 3D motion recognition, which detects when a person is
using ASL.
Re-voice Glove: It is the glove based approach and is a data glove. It is an electronic
device that converts the gestures made by the gloves into text or voice. It has a screen
and speaker attached to the back of the glove, which makes it portable. It has many
sensors attached to it that detects its motion. It could be attached to the smart devices
and then make the gesture and type the word it represents. This way the glove learns to
identify the signs. Therefore, it is not specific to a particular sign language and it also
bridges the communication gap between the people that uses different sign language.
AI interpreter: It translates the user’s sign in real time. It is based on the chinese sign
language. Its dataset contains 900 chinese phrases and can detect up to 1000 words. It
has one drawback that it requires white background. The Tecent YouTu Lab has
decided to install it in every railway station, bus stands and airport to help the hearing
impaired people get their tickets or get queries without any problem.
AR App: Augmented reality app is the mobile app used to convert American Sign
Language to text and voice. It also has a voice recognition system that converts the
voice to sign language. It uses a mobile camera for recording the gestures and mobile
speaker for voice recognition. It is basically a mobile solution as an instant sign
language interpretation app that leverage computer vision models over the cloud, and
present users the visual translations in augmented reality to empower both the sign
language users and nonusers for more collaborations.
18
Google Translator: It is an AI app that converts the sign language to text/speech. It
works by placing a smartphone in front of the user while the app translates gestures or
sign language into text and speech. The app, called GnoSys, uses neural networks and
computer vision to recognise the video of sign language speakers, and then smart
algorithms translate it into speech. The translation software in the market are either
slow or expensive, or rely on old technology which does not allow scaling to other
markets outside the country of origin. But this app is a compellingly fast, easy,
comfortable and economical solution. It can translate as quickly as the person speaks,
translate any sign language and can be plugged into many products, such as video chat
applications, AI assistants, etc. The pocket interpreter for the deaf relies on superior
new technology: AI and neural networks. All the translation happens in the cloud. It
just requires a camera on the device facing the signing person, and a connection to the
internet.
The work in this field began using the VPL data glove. It is the most commercially
available glove. It was developed by Zimmerman[11] during the 1970s. It is based upon
patented optical fiber sensors along the back of the fingers. Star-ner and Pentland
developed a glove-environment system capable of recognizing 40 signs from the
American Sign Language (ASL) with a rate of 5Hz. Then after many years in 1996,
Christopher Lee and Yangsheng Xu developed a glove-based gesture recognition
system that was able to recognize 14 of the letters from the hand alphabet, learn new
gestures and able to update the model of each gesture in the system in online mode,
with a rate of 10Hz. Over the years advanced glove devices have been designed such
as the Sayre Glove, Dexterous Hand Master and Power Glove[1]. Then in 1999, another
research was done by Hyeon-Kyu Lee and Jin H. Kim. Hyeon-Kyu Lee et al presented
work on real-time hand-gesture recognition using HMM (Hidden Markov Model).
Kjeldsen and Kendersi devised a technique for doing skintone segmentation in HSV
space, based on the premise that skin tone in images occupies a connected volume in
HSV space. They further developed a system that used a back-propagation neural
network to recognize gestures from the segmented hand images[1]. Next in
19
2001,[7]Mohamed proposed a vision based recognizer to automatically classify Arabic
sign language. A set of statistical moments for feature extraction and support vector
machines was used for classification which provided an average recognition rate of
87%. Then in 2012, Reyadh [8] worked on an alphabet sign recognition system with a
recognition rate with naked hand of 50%, red hand of 75%, black Hand of 65% and
white hand of 80%. The sign alphabet images were mapped to histograms to uniquely
represent the sign images. KNN algorithm which measures the distance between these
sign histograms was used for classification. The process was simple and effective only
on images and produced a very low recognition rate on video data. A lot of research
was done in 2015. The method proposed [10] involves extracting the hand gestures
from original color images. The segmented hand positions shape modulated using
Chan-Vese (CV) active contour model and obtained 92.1% recognition rate. Kishore
PVV, proposed 4-Camera model. The segmented hand gestures with extracted shapes
created a feature matrix described by elliptical Fourier descriptors which were classified
with back propagation algorithms trained using artificial neural networks. The normal
recognition rate in the proposed 4 Camera model for sign language recognition is about
92.23%. Then, Etsuko Ueda and Yoshio Matsumoto presented a novel technique-a
hand-pose estimation that can be used for vision-based human interfaces. In this
method, the hand regions are extracted from multiple images obtained by a multi
viewpoint camera system, and constructing the “voxel Model”[2]. Also then, Chan Wah
Ng, Surendra Ranganath presented a hand gesture recognition system. They used image
furrier descriptors as their prime feature and classified them with the help of RBF
network. Their system’s overall performance was 90.9%. Claudia Nolker and Helge
Ritter presented a hand gesture recognition model based on recognition of finger [Link]
their approach they found full identification of all finger joint angles and based on that
a 3D model of hand was prepared using neural networks. Basic sign language system
is based on the 5 parameters and they are hand and head recognition, hand and head
orientation, hand movement, shape of hand and location of hand and head. Among the
5 parameters above, the above system only met one parameter that also partially as they
were not able to recognize the gestures formed using both hands. Hence they didn't get
much success. Next research was also done in Hand Gesture Recognition Using PCA
in [6]: In this paper author presented a scheme using a database driven hand gesture
recognition based upon skin color model approach and thresholding approach along
with an effective template matching which can be effectively used for human robotics
20
applications and similar other applications.. Initially, the hand region was segmented
by applying a skin color model in YCbCr color space. In the next stage thresholding is
applied to separate foreground and background. Finally, template based matching
technique was developed using Principal Component Analysis (PCA) for recognition.
Next, after 1 year in 2016, the dynamic time wrapping based level building (LB-DTW)
algorithm was proposed [9] to solve sign sequence segmentation and sign recognition.
This LB-DTW introduces two problems in recognition. One was under the bad
relationship the recognition rate was very low and HMM was incorporated to improve
recognition rate by calculating the similarity between sign model and testing sequence.
On the other hand, the grammar constraint and sign length constraint were employed to
improve recognition rate whereas in experiments with a KINECT data set of chinese
sign language containing 100 sentences composed of 5 signs each, the proposed method
showed superior recognition performance and lower computation compared to other
existing techniques.
21
2.1.5 Survey of Tools and Technologies Used
Table 3 shows Tools and Technologies used in various works carried out by researchers
(as discussed in literature survey)
Table 3: Tools and their working
Tools Working
Glove based Gesture Recognition It has an electronic glove worn in one hand
and all the gestures or movement made by the
hand are captured based on the technology
used to make that glove. It can be connected
to a computer using a wire or wireless
module.
Technology Working
Optical Fibre Sensors Fiber optic sensors work based on the principle
that light from a laser or any superluminescent
source is transmitted via an optical fiber,
22
experiences changes in its parameters either in the
optical fiber or fiber Bragg gratings and reaches a
detector which measures these changes
23
Back Propagation Algorithm is used to adjunct the
weights in order to improve the efficiency of the
model.
24
Table 5: Tools and Technologies used in various works carried out by researchers
25
Surinder network
Ranganath
(2015)
26
learning of Hand Master model of each g/document/
gestures for and Power gesture 509165
human/robot Glove
interfaces
27
and Convert =[Link].
Sign 1887&rep=r
Language ep1&type=p
Videos to df#page=45
Voice/Text
28
Markov grammar
Model constraints and
sign length
constraints are
applied, the
efficiency of the
recognition
system
increases.
2.1.6 Summary
Concluding the work of all the researchers, the methods of making sign language
identifiers can be broadly divided into a glove based approach and a vision based
approach irrespective of the sign language being used. SLI is built based on the later
method. Although accuracy achieved is quite high using the glove based approach but
the inconvenience, cost, more prone to damage, uncertainty, tricky and risky to use
dominate its pros. Further dissecting the vision based work, a technology is needed to
identify and reliably distinguish the signs. Mainly researchers have used ANN, KNN,
HMM, SVM etc which all are the machine learning methods to classify the objects. SLI
also has a machine learning method for recognition of the signs but a more sophisticated
and rigid in terms of identifying and classifying the image, that is CNN (convolution
neural network). All the methods as mentioned above used by the researchers are good
enough to identify the image but require more preprocessing, additional techniques and
more micro details in order to make it work with sufficient accuracy. Whereas CNN is
the evolved version of the ANN specifically for the image related classification that
reduces the work, boasts accuracy and makes it flexible for any addition of the property
without distorting the entire neural network. Now talking in terms of the dataset
required for the testing and training. Researchers have used colored images, dictionaries
according to the technique they have used. In SLI for testing and training, the dataset
of morphologically processed images is used. The benefit of doing so is the speed,
precision and increased accuracy which gives the result close to glove based approach.
29
In addition to the above similarities in one way or another, SLI has additional features
of converting the english text to sign language for the two way communication. Adding
this feature has made it a complete project for the deaf and dumb people to have
communication reliably and without any difficulty, which is the main motive behind
making this Sign Language Identifier. For more convenience we have it in the form of
a website to make it portable as nowadays everyone has a mobile phone and its interface
is user friendly.
2.2 Standards
1. ISL (Indian Sign Language) standard: Used to create Indian Sign Language
Dictionary.
2. Web 2.0 : Internet technology to create online applications that behave
dynamically.
3. Camera measurement standards :
3.1 Resolution Measurements - ISO 12233
3.2 Noise Measurements – ISO 15739
3.3 Sensitivity Measurements – ISO 12232
3.4 OECF (tone reproduction) Measurements – ISO 14524
3.5 Chromatic Displacement Measurements – ISO 19084
3.6 Color Characterization Test Procedures – ISO 17321
3.7 Image Stabilization Measurements – ISO 20954
3.8 Low Light Measurements – ISO 19093
3.9 Shooting Time Lag Measurements – ISO 15781
3.10 Shading Measurements – ISO 17957
4. Video quality: atleast 360 pixels.
30
2.3 Software Requirement Specification
Clear requirements help to get a better vision for the project and helps developers to
make code in a certain way. SRS for SLI are given below.
2.3.1 Introduction
This document lays out a project plan for the development of a “Sign Language
Interpreter” system.
The intended readers of this document are current and future developers working on
“Sign Language Interpreter” and the sponsors of the project. The plan will include, but
is not restricted to, a summary of the system functionality, the scope of the project from
the perspective of the “Sign Language Identifier” team, scheduling and delivery
estimates, project risks and how those risks will be mitigated, the process by which the
project, and metrics and measurements will be developed will be recorded throughout
the project.
[Link] Purpose
Sign Language is the most natural and expressive way for hearing impaired people.
Unfortunately, most of the common people are not familiar with sign language. This
leads to a large communication gap between deaf and dumb and the rest. The SLI
project aims at eliminating this communication gap. The main purpose of this project
is to build a web based app that can convert ISL into text; and speech to ISL.
31
Reading Suggestions
About SLI
SLI (Sign Language Identifier) is a web app designed to convert the Indian Sign
Language to text and speech to ISL. The main objective of the app is to reduce the
communication gap between deaf and dumb and the rest of the world. You simply need
to perform signs in front of the camera, and SLI will provide you output w.r.t.
corresponding sign.
2. The signer must have a steady pace while making signs so that the camera could
detect the signs properly.
4. The user must keep the camera at a fixed angle so that it can record the signs
properly and accurately.
32
[Link] Project Scope
To bridge the gap between deaf and dumb and the rest , the main purpose of this project
is to design a web app that can convert sign language into text; and speech to ISL.
Anyone willing to communicate through sign language will be able to login through
the app. Hand gestures of the signer will be captured by the camera of the device which
will be further converted to text . The quality of the webcam will be more than 360
pixels. Signer will be at a fixed distance from the camera. The dataset for training the
model will contain gestures from Indian Sign Language only.
First the user will enter their username and password, then sql server will check for its
validity from the “user” database. Authenticated users will be allowed to access the web
app otherwise the user has to make a user id before using the product. The web app will
33
ask for the permission of the webcam. Figure 3 shows the Layered architecture of the
product. The user will press the start button of the webcam for the video input and stop
button once they're done. This recorded video is then converted into an image sequence,
which is preprocessed and through template matching from “the sign to text” database,
corresponding text is saved. On pressing the “convert sign language to “text” button,
the saved text will appear in the text box. In order to convert it to speech to ISL, a
person has to press on the live voice button and start saying the text what he/she wants
to convert to ISL.
34
[Link].2 User characteristics
The web app is a software, designed for the deaf and dumb people to help them in
having conversion with the world without any barrier. It converts the sign language to
text; and speech to ISL. No prior expertise is required to use this software except for
knowledge of indian sign language.
[Link].1 Objectives
The web app, sign language identifiers to fill the gap in communication between the
hearing impaired and other people. Table 7 shows visions, goals, initiatives and
customers for SLI.
Visions “Sign Language Identifier” to be used all over India and by the
people who use indian sign language
35
[Link].2 Release
Table 8 provides an insight to release details.
Table 8: Release Details
Dependencies ● Keras
● Tensorflow
● CV2
● Numpy
● OS
● Sklearn
● Speech_recognition
● [Link]
● Flask
● Pillow
● HTML and CSS
● Itertools
● String
● sklearn
36
[Link].3 Features
Table 9 provides an insight to the features of SLI, it’s description, purpose, user value
and assumption.
Table 9: Features
Purpose To make the interface user friendly and help the user get
output with convenience so that it can have the real time
conversation with hearing person
User problem Users were not able to have real time conversation due to
difficulty in recording the video with accuracy,
maintaining certain distance the webcam and having
particular color background
User Value Now user can record the video without worrying about the
distance and color of the background
Assumption User must have webcam and the recorded frame should
only contain the user
37
[Link].4 User Interface
1. The user uses various buttons of the application in the system to operate it.
a. START BUTTON: This would start the application and hence the user
(speech-impaired person) would give input to the application through gestures.
b. STOP BUTTON: This would make the application stop and hence the user
can close the application.
c. Live Voicing: This would start recording the voice of the speaker to convert
it to ISL.
2. Input for the application would be provided by any camera and speaker attached to
the system to recognize gestures.
Output
1. Output would be provided for various gestures in a separate window in the form of
text in case of ISL to text.
38
● Strategic use of color and texture
● Feedback mechanism
● Purposeful Layout
● Efficient
Functional Requirements
● Authentication
a. New users will have to create an id using attributes: username, email
address and password.
b. The user will use username and password to sign in to the web
application.
● Capturing Video
a. Web app will start capturing video of a person doing sign language on
pressing the webcam button.
b. Web app will stop the webcam and store the captured video on the
repressing of the same button.
c. Web app will alert the user if distance between user and webcam is either
too much or too less.
● Text generation
a. Web app will convert the sign language video to text.
b. Web app will give an alert message if text related to sign language is not
found.
39
● Display Generated Text
a. Web app will display the text generated corresponding to the video on
the screen.
b. Web app will not display anything but give an alert message if text is
not generated.
c. The generated text will be displayed in natural english language.
● Record Speech
a. Web app will record the voice of the person pressing the Live voicing
button.
b. Web app will close the recording interface on speaking “good bye”.
● Display GIF
A. Web app opens a separate window displaying the output.
B. Web app will show the gif of the person doing ISL matching the
recorded speech.
40
2.3.4 Other Non-Functional Requirements
To judge the operation of the system (SLI), rather than specific behavior, there are
certain non functional requirements given below.
The user login and sign up just requires email id, username and password, hence no
personal details are asked for, ensuring no pose to security threat. Passwords could be
stored securely, by firstly modifying password with a unique way and then applying
hashing method to make it more secure. Hashing and salting passwords could be used.
For more sensitive information, asymmetric encryption could be employed. But since,
it is planned to keep it simple in the initial stage, so only, limited attempts are permitted
for wrong username or password. More features as proposed could be added later on.
41
SLI will use a firewall authentication system for user login and WAF (web application
firewall) to protect the database.
Service Budget
Note: These prices are for Google Cloud. The above cost analysis has been done with
reference to [Link]
Egress from Cloud SQL: See Network Egress Pricing in the above link.
42
2.5 Risk Analysis
There is already interest among researchers in various fields in applying different
methods to sign language applications. In particular, some applications / methods use
gesture recognition touch on sign language recognition as a domain. These methods
work on algorithms to detect fingers, hands, and human gestures. However, by framing
sign language recognition as an application area, they risk misrepresenting sign
language recognition as a gesture recognition problem, ignoring the complexity of sign
language as well as the broader context within which such systems must function. In
this work, the sign language will be translated into text; and speech to ISL without
changing the context with the help of natural language processing techniques.
The only drawback with this approach is the risk of losing information. The gestures of
different letters are very similar and differ only in the location of the thumb and
contours of hand. If the silhouette of the hand was taken, it would not be possible to
determine which letter was being presented. Contours are similar to silhouettes but
focus on extracting edges on an image. Some methods derive the edges from the
silhouette, making them equivalent; However, edge detection techniques can also be
applied to images.
43
METHODOLOGY ADOPTED
44
a software based approach. MotionSavvy (A tablet app
Hardware based solutions are that understands sign
expensive and not affordable by language)
common people whereas
software based approaches are
cheap and affordable.
45
3.2 Proposed Solution
ISL to text
For the sign language recognition system, a vision based approach will be used. The
system will contain four modules/derived algorithms: preprocessing and segmentation,
feature extraction, sign recognition and sign to text. Indian Sign Language is the area
of concern. The signs will be captured using a webcam where the signer must be directly
in front of the webcam. Then the output will be in the form of text. Following is
methodology for sign language to text.
The main steps for converting sign language to text are given below.
c. Segmentation: Here the image is converted into small segments in order to get
more accurate image attributes. Here the segmentation of hands is carried out to
separate objects and the background. To make segmentation more robust, so
that the feature extraction and sign recognition of the image is accurate, the
BGR image is converted to HSV, then fed into inRange() function which
performs segmentation.
46
d. Morphological filtering: The image attributes extracted from the
segmentations consist of noise, therefore morphological filtering removes the
noise and gives a smooth contour. The output of these filterings are image
attributes which are useful for feature extraction and sign recognition. Dilation
and Erosion are used for morphological filtering.
47
Speech to ISL
Firstly a dataset of alphabets and gifs is created. Post that, the coding is done for the
conversion of speech to ISL. Firstly, the microphone is activated and input voice is
thresholded. Then, the voice recognition takes place using Google API. The spoken
word is processed to text and then text is searched with labelled folders of the dataset.
If word exists, then corresponding GIF is displayed, else the plot of alphabets of which
letter is composed of is displayed on the screen.
Figure 4 shows a block diagram for ISL to text conversion as well as speech to ISL.
48
3.3 Work Breakdown Structure
“Sign Language Identifier” has various modules with major modules being Data
Acquisition, Image preprocessing, Feature Extraction, Deep Learning Model, Template
Matching and Text Output. They further have submodules as shown in Figure 5 given
below. Along with that, for speech to ISL, the modules are shown in the same figure
i.e. Figure 5.
49
Following Gantt Chart (Figure 6) shows tentative project timeline and %age of tasks
completed along with start date and end date which is shown in detail in Gantt Chart
(Table 12) along with Task Description and Status (in terms of progress).
50
Gantt Chart Table
Table 12 below shows the gantt chart table in detail.
Table 12: Gantt Chart
51
Training Dataset 3d 06-01-20 06-13-20 100% Completed
52
3.4 Tools and Technology
Following is the tentative list of tools and technology to be used while making the Sign
Language Identifier (SLI). Some of the tools could be modified based on outcomes. As
of now, given below is the list.
● Numpy
○ Working with arrays
● SpeechRecognition
○ Speech Recognition into Python application
● [Link]
○ Plotting graphs
● OpenCV
○ CV (Computer Vision) library for images and videos
● OS
○ For manipulating files
● Pillow
○ Image manipulation
● String
○ For punctuations
● Itertools
○ For count
● PyAudio
○ Audio I/O library
● Anaconda Navigator and PyCharm
○ Python offline coding
● Google Colab
○ Cloud coding
● Neural Networks (Keras, Tensorflow, CNN)
○ For classification and training
53
DESIGN SPECIFICATIONS
54
“Sign Language Identifier” has various modules with major modules being Data
Acquisition, Image preprocessing, Feature Extraction, Deep Learning Model, Template
Matching and Text Output. They further have submodules as shown in Figure 10 given
below. Along with that, for speech to ISL, the modules are shown in the same figure
i.e. Figure 8.
55
4.2 Design Level Diagrams
Design Level Diagrams include Component Design, Data Design, Interface Design
given in section 4.2.1, 4.2.2, 4.2.3 respectively.
Figures 9,10 serve the following purposes for the SLI project
1. It visualizes the components of the SLI system.
2. It describes the organization and relationships of the components.
3. It represents the dependencies of the components on each other.
Figure 9 shows the visualization of components for the conversion of ISL to text
and figure 10 shows the same for conversion of speech to ISL.
56
Figure 9 : Component Design for ISL to English text
57
4.2.2 Data Design
In data design, there will be three databases for storing user information and image to
text information.
58
There are basically four relationships established between different tables. They are as
follows.
1. Provides: The relationship between the user table and saved image table is
provided as the user will capture the image and it will be stored corresponding
to the respective user. The relation between these tables is weak as the saved
image table is a weak entity. The saved image table will use its foreign key and
the primary key of the user table to uniquely identify its tuple.
2. Preprocessing: The relation between saved image table and image sequence
table is preprocessing as the image from the saved image table will be
preprocessed and converted into image sequence of segmented image and it is
also a weak relation as both the tables are weak entities. The tables will use their
foreign key and primary key of the user table to uniquely identify its tuple.
3. Conversion: The relation between image sequence and image to text table is
conversion as the image taken from the image sequence will serve as primary
key and will give the corresponding text, which is basically the english
translation of the image containing sign language words.
4. Record speech: This is the relation between user table and speech recognition
table. The speech input by the user is converted to ISL using this relation. The
speech is converted into text using google api and then using this table we
convert it to ISL gif or ISL images.
59
Figure 11: ER Diagram
60
4.3 Analysis Diagrams
Analysis Model is used to describe the model, information and structure of the system
and they are converted to architecture, interface and component level design in the
'design modeling'. It is a technical representation of the system. It acts as a link between
system description and design model. In Analysis Modelling, information, behavior and
functions of the system are defined and translated into the architecture, component and
interface level design in the design modeling.
The model can be analysed through various diagrams which are shown below.
As shown in Figure 12, there are two users in the SLI project. One is the deaf and dumb
person and the other user is the system/admin. The user (specially abled person) has to
login to the app before using it. Then the webcam will be switched on and the person
has to perform some actions according to the Indian Sign Language that will be
converted to Text Output through a series of steps which are performed in sequence as
shown below. At last, there is optional feedback.
61
Figure 12: Use Case Diagram of ISL to text
62
4.3.2 Use Case Template
Table 13 shows Use Case Template for ISL to text and Table 14 is Use Case
Template for speech to ISL.
Description SLI 1.1 is a web app to convert ISL to Text. This app will help
the deaf and dumb people to communicate with the rest of the
world. The app captures the gestures made by the deaf and dumb
people through webcam and then using different technologies
like image acquisition, image preprocessing, feature extraction,
template matching and applying DL model. The video captured
will have several frames and the corresponding sign will be
converted to text on screen.
Goal The goal of this system is to fill the gap between deaf society
and the rest of the world. Using this app, the deaf and dumb
people will be able to explain their thoughts to the rest of the
world and will make them feel less isolated from the society and
brings in sense of equality.
63
3. The user will perform the sign language actions in front
of the camera.
4. The system will convert the speech to ISL.
5. Feedback page will be displayed.
Post Conditions Customers can provide feedback after using the web app post
login only.
Actors ● Users
● System
Includes ● Login
● Start Webcam
● Gesture capture
64
● Preparing image sequence
● Image acquisition
● Image preprocessing
● Feature extraction
● Template matching
● Text Generation
● Text Output
Extends ● Feedback
Description SLI 1.2 is a system to convert speech to ISL. This app will help
the deaf and dumb people to communicate with the rest of the
world. The system captures the voice and then displays the gif
or sequence of plots of alphabets corresponding to the sign so
performed.
Goal The goal of this system is to fill the gap between deaf society
and the rest of the world. Using this system, the deaf and dumb
people will be able to explain their thoughts to the rest of the
world and will make them feel less isolated from the society and
brings in sense of equality.
65
Pre Conditions Microphone must be in working condition.
Actors ● Users
● System
Extends ● Feedback
66
Modification 10th April, 2020
History
67
4.3.3 Activity Diagram
Activity diagram is basically a flowchart to represent the flow from one activity to
another activity, just the difference being that parallel activities going on can be shown.
The activity can be described as an operation of the system. Figures 24 and 25 represent
the activity diagram for the “Sign Language Identifier” project . It serves the following
purposes:
1. It draws the activity flow of the system.
2. It describes the sequence from one activity to another.
Figure 14 shows the activity diagram for conversion of ISL to text and figure 15 shows
the activity diagram for conversion of speech to ISL.
68
Figure 14: Activity Diagram of ISL to text
69
Figure 15: Activity Diagram of speech to ISL
70
4.3.4 Sequence Diagram
Figure 16 depicts interaction between the objects in a sequential order. There are 6
objects named User, System, Webcam, Image Preprocessor, Database and Template
Matching. The diagram below shows the order in which the interaction between the
objects takes place.
One more feature has been added to the project, i.e. Converting Speech to Indian Sign
language (ISL). This feature will help to communicate two-way, which means, normal
people can also talk to the deaf and dumb. They just have to speak (a dictionary has
been created with some words) and whatever they speak will be converted to ISL ,
which will be of more help for deaf people. Following are the steps followed for
converting speech to ISL
71
If said word is present in the dataset
1: Live Voice Recording
2: Recognition of Speech using the Google Speech API
3: Text Preprocessing
4: Dictionary based Machine Translation
5: A gif is displayed performing the ISL.
72
4.3.5 Class Diagram
Following Class Diagram (Figure 17) shows various classes, attributes and functions
and Table 15 also shows the same.
73
Figure 17: Class Diagram
74
4.3.6 State Chart Diagram
A state chart diagram is used to represent the condition of the system at finite instances
of time. The following diagram i.e. Figure 18 is serving the purpose for the SLI project.
There are nine states and occurrence of a particular event causes transition from one
state to another.
75
Figure 19 shows the State Chart Diagram for conversion of speech to ISL.
76
IMPLEMENTATION AND EXPERIMENTAL RESULTS
This section deals with discussion of implementation and experimentation with regards
to the project. It also mentions all the test plans including the features to be tested, the
test cases and discusses the inference drawn from the results. Furthermore, the
procedural workflow and algorithm used are mentioned in it.
5.2.1 Data
● Data Sources: Data sources include collection of dataset partly by creation of
dataset on own and partly by using dataset from various github repositories.
77
● Dataset Preprocessing: Since, the algorithm requires the dataset to have
segmented images, so a script was developed which can automatically perform
this daunting task.
● Overall dataset information: Dataset includes training_set which has folders of
A, B, C, …., Z which each folder contains 1200 images and testing_set also has
similar subfolder arrangement, just the difference being the no. of images which
are 251.
● For captured data during service avail by user: The image of only the hand
portion was captured and segmentation was performed which was matched with
templates (in form of images) by the trained model (CNN).
78
5.3 Working of the Project
5.3.1 Procedural Workflow
Sign language Identifier is extremely intuitive to use and operate. The first-time user
has to first sign up and then login to use. There are two modules :- sign language to
text and speech to sign language. The user has the choice to invoke the necessary
module as and when required.
On choosing sign language to text, the webcam automatically gets started. The webcam
detects the sign made by the signer and correspondingly converts the same to text. The
session can be terminated by clicking the “Close Session” button on the interface.
On choosing speech to sign language, the microphone gets activated. The live voice is
recorded through the microphone and is recognized using google speech API, resulting
in production of the text. If the user says “goodbye”, then the session is terminated else
the relevant ISL gif is displayed.
79
5.3.2 Algorithmic Approaches Used
Figure 21 : code 1
The values of l_h, l_s, l_v, u_h, u_s, u_v in Figure 37 are respective positions of the
window “Slider”. This window is created with the help of code shown in Figure 22.
Figure 22 : code 2
80
[Link].2 Building the CNN model
Figure 23 shows the code for building the CNN.
Figure 23 : code 3
81
[Link].3 Recognizing image and predicting the result
Figure 25 and Figure 26 show the code for predicting the result.
Figure 25 : code 5
Figure 26 : code 6
82
[Link] Speech to ISL
● Firstly, live voice is recorded i.e. audio is obtained from the microphone.
● Then speech is recognized using Google Speech Recognition and
corresponding text is obtained.
● After this, with the help of dictionary based machine translation, relevant ISL
gif or plot of composed alphabets is displayed.
The code for the same is shown in the Figure 27-30 below.
Figure 27 : code 7
Figure 28 : code 8
83
Figure 29 : code 9
Figure 30 : code 10
84
5.3.3 Project Deployment
Our project, Sign Language Identifier (SLI) basically consists of two modules -(i) ISL
to text and (ii) Speech to ISL.
It has four main menus : HOME, LOGIN, SIGNUP, ABOUT. The already registered
users can directly login by filling in the username and password. The new user first
needs to sign up by filling in : email id, username and password. There is an ABOUT
menu which displays quick information about our project.
After successful login, the user has a choice between the two modules: sign language
to text and speech to sign language. First module (sign language to text) is implemented
using CNN (Convolutional Neural Network) and various image processing techniques.
The second module (speech to sign language) is implemented using Google audio api
for speech recognition, and dictionary based machine translation. After this the final
step in the project was to deploy the model. Deployment of the model was done using
Flask. Flask is a web application framework written in Python. The deployed model
worked perfectly when running on the flask server.
Process of deployment:
The existing functions were embedded into html and css webpage and to have python
as a backend, flask provides a facility of routing to the particular webpage.
1. Model Building: Deep Learning Model (CNN) pipeline was built to classify
Signs performed into A-Z in English.
2. Webpage template: Here, the website was designed where users can avail
services of ISL to text; and speech to ISL. Website design uses following
technology stack:
a. Front-end: HTML, CSS
b. Backend: Python
3. Predict class and send results: Next, use the saved model to predict the class of
the signs performed and send the results back to the webpage.
85
Below given is sample deployment code (not complete code)
@[Link]('/')
@[Link]('/[Link]')
def index():
return render_template("[Link]")
@[Link]("/[Link]", methods=['POST','GET'])
def login():
return render_template("[Link]")
@[Link]("/[Link]")
def signup():
return render_template("[Link]")
@[Link]("/[Link]")
def about():
return render_template("[Link]")
86
5.3.4 System Screenshots
SLI has 4 main menus: HOME, LOGIN, SIGNUP, ABOUT.
87
The second menu is Login (Figure 32) for already registered users. It asks for Username
and Password. Click on the LOGIN button to complete the task.
The third menu is SignUp (Figure 33) for new users. It asks for Email Id, Username
and Password. Click on the SignUp button to complete the task.
88
Figure 34 displays the ABOUT menu.
On successful login, the following screen (as shown in figure 35) is displayed from
where you can choose between services available which are: Sign Language to Text
and other one is Speech to Sign Language.
89
If the user chooses Sign Language to Text option, then the following screen appears
(Figure 36) and then the user performs sign language and corresponding text output is
displayed.
After the user is done with the task, he/she can click on LOGOUT which ends the
session and brings the user to the homepage but if the user clicks on the “Close
Session!!” button (as shown in Figure 36), then SLI asks for Feedback (Figure 37)
which is optional. The user can simply log out then if he/she isn’t interested or can write
feedback and click on “Submit!!” button.
90
If the user avails the facility of speech to ISL, then the following screen (Figure 38)
appears where the user needs to click on the “Activate Microphone” button.
91
The user speaks “FLOWER IS BEAUTIFUL” and the corresponding GIF is loaded as
given in Figure 40.
92
5.4 Testing Process
Testing is the most important phase where normal and corner cases are tested
and that’s where the developer gets an idea, how the product is performing in
real time.
93
II. Different frequency voices: Voice to be heard can be of different
frequencies, it should have no impact on the conversion.
III. Standard English Speech: Speech to sign conversion on standard english
speech was tested.
IV. Interrogative speech: Interrogative speeches were tested.
V. Greeting speech:Greeting Speeches were tested.
VI. Imperative Speech: Imperative Speeches were tested
Speech to Text
1. Computer recorded speech: Checking the result with the computer voice.
2. Real time speech: Live testing the result with real time speech.
94
5.4.5 Test Cases
There are two modules in SLI:
1. ISL to text
2. Speech to ISL
The test cases for ISL to text are given in table below (Table 16)
1 ISL Alphabet
2 Blank
The test cases for speech to ISL are given in table below (Table 17)
2 Declarative Sentence
3 Interrogative Sentence
4 Imperative Sentence
5 Proper Nouns
95
5.4.6 Test Results
On performing the test cases for ISL to text, following results were obtained and are
given in table below (Table 18)
1.2 Alphabet B
2 Blank Pass
5 Similar ISL 5.1 ‘I’ and ‘L’ 5.1 ‘I’ and ‘L’
96
Alphabets Pass
5.1 ‘I’ and ‘L’
5.2 ‘D’ and ‘P’ 5.2 ‘D’ and ‘P’
Fail
97
The test cases for speech to ISL are given in table below (Table 19)
98
4 Imperative Sentence Pass
99
Table 20 shows the detailed result of all the ISL alphabets
A Yes
B Yes
C Yes
D No
100
E Yes
F Yes
G Yes
H Yes
I Yes
J Yes
101
K Yes
L Yes
M Yes
N Yes
O Yes
P Yes
102
Q Yes
R Yes
S Yes
T Yes
U Yes
V Yes
103
W Yes
X Yes
Y Yes
Z Yes
104
identified it as ‘P’. Another parameter could be Training and Testing Accuracy and also
Training and Testing loss, which is discussed in detail in graphs given below namely
Figure 41 and Figure 42. Figure 41 summarizes history for model accuracy and it can
be noted that as the number of epochs increase, the model accuracy initially shot up and
then later on after some epochs it remained nearly constant giving final accuracy of
0.9895 for training data and 0.9765 for test data. Figure 42 summarizes history for loss
and it can be seen that the loss was huge in the very beginning of epoch and as the
number of epochs increased, loss decreased to 0.0115 for training data.
105
5.6 Inferences Drawn
Following are the inferences drawn from testing performed on Sign Language
Identifier:
1. Sign Language Identifier is suitable for all the alphabets (A-Z) performed
according to Indian Sign Language.
2. Sign Language Identifier gives results with very good accuracy of 97.25 for
static images but in case of dynamic images, the output comes with some
fluctuations and the user has to do some movements for getting the result.
3. Sign Language Identifier is majorly suitable with white plain background while
converting Indian Sign Language to Text.
4. Sign Language Identifier successfully converts all the words which are there in
our system to gif while performing speech to text. Apart from the words which
are not in the system, it successfully shows the letter by letter image of the word
spoken by the user.
5. Sign Language Identifier is suitable if there is no background noise while
performing speech to ISL.
106
5.7 Validation Of Objectives
Table 21 shows the status (successful/unsuccessful) of the objectives for the ISL
identifier.
107
CONCLUSIONS AND FUTURE SCOPE
Objectives Discussion
108
Machine (SVM), Hidden Markov
Models (HMM) etc.
3. To test and demonstrate the usability The SLI prototype is being trained on the
and effectiveness of the proposed dataset collected. After training
solution for targeted user groups (Deaf efficiency of the model will be
and dumb persons). evaluated.
109
6.2 Conclusions
In the Prototype of Sign Language Identifier the working model of the “Sign Language
Identifier'' is made, which works on the images captured by the camera. For ISL to text
conversion, the user performs the sign in the rectangular region defined. Then
corresponding text appears on the screen. Internally, firstly from video, frames are
extracted and then based on HSV color model and inRange() function of opencv, the
thresholding takes place. Different thresholding techniques like Simple Thresholding,
Adaptive Thresholding, Otsu Thresholding were used but inRange() function gave
better results, then the model was trained by CNN model.
For Speech to ISL conversion, simply speech is converted to text and then the label is
checked.
There are many research gaps in glove based recognition systems like the VPL
dataglove or other glove systems that use wired networks makes it importable and there
is a lot of noise and delay in the system, that's why the vision based recognition system
is made. HMM model is used by many researchers and it has high efficiency but it is
only used by those researchers who aimed at recognizing very few words or alphabets.
Aim is to make the real time conversation using “Sign Language Identifier” between
hearing impaired and other people, that's why CNN model for image recognition is used
as its efficiency is high in recognizing large numbers of images. For the same reason it
will be modified to an interface that will recognize the gestures in the video and for
easy access a web app will be made.
110
6.3 Environmental (/Economic/ Social) Benefits
This app is of no harm to the environment and has lots of social and economic benefits.
Social benefits include that it will be helping that section of the society which feels
completely isolated because no one is able to understand what they want to say/express.
The deaf and dumb people always feel deprived of society. They are not given equal
opportunity in society. Hence, by using this app they just need to perform Indian sign
language in front of a webcam and hence the sign language will be converted to text;
and speech will be converted to ISL as fast as they perform the actions. Along with that,
it has a lot of economic benefits too. As the app will be made by using the Indian Sign
Language, it will be of great benefit to the economy as it will reduce the use of the apps
and datasets made by other countries.
6.4 Reflections
The whole journey of building “Sign Language Identifier” has been a valuable
experience, starting with the discovery of possible opportunities to think of the idea to
the phase where the same idea was actually deployed. The team gained insight into the
field of software development and now in the future, members shall feel more confident
in the process of project development. Furthermore, it was learnt how to analyze the
existing frameworks and perform literature surveys and utilize that analysis to identify
the problem statement, research gaps and come up with the solution ideas. It was a
learning of how to incorporate and take care of the user requirements. It was the time
when the importance of documentation was realized and what are techniques involved
in being organized about it. One of the takeaways was how to manage the resources in
an efficient manner and most importantly to use common sense and build a viable and
efficient model, but best takeaway was development of analytical skills while working
in the team and discussing each point of the assigned task in hand in detail. The whole
project helped us in exploring the skills as a computer engineer and improved
confidence levels, ability to work under pressure and helped in learning project
management techniques. It aided the members to be familiarized with the working and
delivering of projects and how to build an entire product from just an idea.
111
6.5 Future Work Plan
The system currently converts sign language to text; and speech to ISL, not sign
language to speech, which can be included later on. Other feature extraction algorithms
can be added in conducting experiments for more accurate results. Same goes for more
classifiers like Support Vector Machine (SVM), Principal Component Analysis (PCA)
and many more. Their combination could be used for improvement in recognition rate.
The Natural Language Processing (NLP) tasks could be focussed on, to facilitate proper
context. For the web app, the user login and sign up just requires an email id, username
and password, hence no personal details are asked for, ensuring no pose to security
threat. Currently, limited attempts are permitted for wrong username or password.
Passwords could be stored securely, by firstly modifying password with a unique way
and then applying hashing method to make it more secure. Hashing and salting
passwords could be used. For more sensitive information, asymmetric encryption could
be employed. As of now, dataset is also limited, more dataset can be added for better
results, which can be prepared by taking the pictures of different humans performing
signs, different backgrounds. As of now, it is limited to the machine on which it runs,
future work includes a chrome extension for SLI and a progressive web application. A
digital avatar could be made for better interaction and real time experience.
112
PROJECT METRICS
113
7.2 Relevant Subjects
The Capstone Project requires the knowledge of multiple subjects. Some of them are
direct learnings from courses (taught in Institute), and some are the skills to be learnt
by the developers themselves. Table 23 provides information about the relevant
subjects used to successfully complete the project.
UCS615 Image Processing To capture the image from the camera and
further removal of noise using noise
reduction filters and further enhancing the
input for better output and ease of prediction
for the trained model.
114
7.3 Interdisciplinary Knowledge Sharing
From doing this project, the members gained knowledge with regards to computer
engineering aspects like image processing, machine learning, deep learning, natural
language processing, web development and also about some aspects of biology and
computer vision. Learning was also with regards to target users- the deaf and the dumb
people and also got some knowledge about the Indian Sign language. Major revelation
was about their everyday problems, how they face such challenges and what new ways
have been brought into their lives through modernization. Note was made on the various
devices available to them and how much they cost and how much efficient they are in
changing their lives and helping them.
Evaluation of
S1 S2 S3 S4
Evaluation By S1 5 5 5 5
S2 5 5 5 5
S3 5 5 5 5
S4 5 5 5 5
115
Table 25: Student Information
S4 101703545 Simran
116
● Text preprocessing
● Template Matching and Sign
Recognition using appropriate
algorithm (coding)
● Conversion of Speech to Text
● Creation of GUI
117
7.6 Student Outcomes Description and Performance Indicators (A-K
Mapping)
Table 27: AK Mapping of various concepts
SO Description Outcome
A3 Applying engineering techniques for Used google speech API for speech
solving computing problems. recognition and various algorithms
detect and preprocess images
captured from webcam.
B2 Use appropriate methods, tools and Proper research on ISL was done in
techniques for data collection. order to prepare an accurate dataset.
B3 Analyze and interpret results with respect Made the model portable and
to assumptions, constraints and theory. affordable.
118
multidisciplinary teams. multitasking and punctuality.
G2 Deliver well-organized and effective oral Helped the panel to understand and
presentation. give suggestions.
119
enhance self-learning. products, analyzing their drawbacks
and removing them from our model
was the goal.
120
7.7 Brief Analytical Assessment
Q1. What sources of information did your team explore to arrive at the list of
possible Project Problems?
Ans: The different sources used were first going through research papers and taking
out keywords. Then, we discussed our approach with our mentor. We studied different
types of sign language such as American Sign language etc to understand the language
in a detailed manner. We explored already available solutions and their drawbacks.
Then after seeing the pros and cons of different modules, we finalized the project.
Q4. How did your team share responsibility and communicate the information
of schedule with others in team to coordinate design and manufacturing
dependencies?
Ans: Our team had good communication throughout the project. We had one official
whatsapp group with our mentor and one was of all the team members. Our team leader
used to assign tasks to all group members and we were asked to finish it well in time.
121
Whatever problems came across were discussed and if we were not able to solve it on
our own then we used to ask our mentor. Also, there were regular meetings with ma’am
and among ourselves.
Q5. What resources did you use to learn new materials not taught in class for the
course of the project?
Ans: We used various resources such as the internet, books in the library. We went
through a lot of research papers, understanding of various products already made in this
field, and understanding their drawbacks. We also consulted our mentor whenever in
doubt. This project wouldn’t have been possible without the guidance of our mentor.
Q6. Does the project make you appreciate the need to solve problems in real life
using engineering and could the project development make you proficient
with software development tools and environments?
Ans: This one-year project taught us a lot of things. Not only the working of the project
is important but also it taught us how the documentation should go side by side which
keeps us in track of what part is left, what is done. Also, it taught us how real-life
problems are much different than the theory one. We learn better when we are doing
things and seeing it rather than reading from the book. We learnt how to use various
new languages, software and tools.
122
APPENDIX A: REFERENCES
[1] Christopher Lee and Yangsheng Xu, “Online, interactive learning of gestures for
human robot interfaces” Carnegie Mellon University, the Robotics Institute, Pittsburgh,
Pennsylvania, USA, 1996
[2] Etsuko Ueda, Yoshio Matsumoto, Masakazu Imai, Tsukasa Ogasawara, ”Hand Pose
Estimation for Vision Based Human Interface”, IEEE Transactions on Industrial
Electronics,Vol.50,No.4,pp.676- 684,2003.
[3] G. AnanthRao P.V.V .Kishore Ain Shams Engineering Journal Volume 9, Issue 4,
December 2018, Pages 1929-1939 Selfie video based continuous Indian sign language
recognition system
[5] Mahesh Kumar N B Assistant Professor (Senior Grade), Bannari Amman Institute
of Technology, Sathyamangalam, Erode, India. Conversion of Sign Language into
Text
[6] Mandeep Kaur Ahuja, Amardeep Singh, “Hand Gesture Recognition Using PCA”,
International Journal of Computer Science Engineering and Technology (IJCSET ),
Volume 5, Issue 7, pp. 267-27, July 2015.
[8] Naoum, Reyadh, Hussein H. Owaied, and Shaimaa Joudeh. Development of a new
Arabic sign language recognition using k-nearest neighbor algorithm.“ (2012).
[9] Pattern Recognition Letters Volume 78, 15 July 2016, Pages 28-35 Continuous sign
language recognition using level building based on fast hidden Markov model
[10] [Link] [Link] Kumar Segment, Track, Extract, Recognize and Convert
Sign Language Videos to Voice/Text
123
[11] Richard Watson, “Gesture recognition techniques”, Technical report, Trinity
College, Department of Computer Science, Dublin, July, Technical Report No. TCD-
CS-93-11, 1993
124
APPENDIX B: PLAGIARISM REPORT
125
APPENDIX C: TECHNICAL WRITING COURSE PROOF
126
Name: Sanjana Vashisth
Roll No. : 101703482
127
Name: Sargunpreet Kaur
Roll No.:101703488
128
Name: Simran
Roll No.: 101703545
129
The specific requirements for using the SLI app effectively include: 1. Users must log in with a user ID before accessing the app . 2. A webcam and speaker must be connected to the device. The webcam's resolution should be at least 360 pixels . 3. Users must maintain a fixed distance from the webcam, keeping a stable pace when performing signs and ensuring a plain background for accurate detection . 4. The app's interface requires a responsive, easy-to-use layout, ensuring security and efficient operation . 5. For sign language to be converted to text, users input gestures which the app captures and processes into text displayed on the screen. The conversion should happen with a time delay of less than 2 seconds . 6. Speech to ISL conversion involves pressing the "Live Voicing" button, speaking the desired text, and if "goodbye" is spoken, the recording session terminates .
The conversion of Indian Sign Language (ISL) gestures to text using the SLI web app primarily involves a vision-based approach with several technical methods. The process begins with image acquisition, where gestures are captured through a webcam. The signer must be in a defined rectangular region to ensure the system captures the video effectively, from which frames are extracted . Image preprocessing follows, involving segmentation and morphological filtering to isolate the hands and reduce noise, thereby enhancing accuracy . The segmented images are converted from BGR to HSV color models, and the inRange() function is used for thresholding . Feature extraction and sign recognition are key stages that involve using a trained dataset of Indian Sign Language alphabets. The dataset is divided for training and testing, and a CNN model is employed for classification to recognize gestures from the video frames . These steps ensure that ISL gestures are accurately converted into English text, which is then displayed on the screen via the web app .
Glove-based sign language systems have significant limitations when compared to vision-based systems. Glove systems, such as the VPL data glove and others, are inherently dependent on wearable technology that can be expensive and requires maintenance and care, making them less accessible to most users . They often offer limited recognition capabilities due to reliance on the specific sensors embedded in the gloves, which may not cover all possible gestures in a dynamic language like ISL . In contrast, vision-based systems use cameras to capture a broader range of gestures and can be more easily updated with advancements in software technology. They also eliminate the dependence on physical hardware, reducing the barriers to entry and making the solutions more scalable and practical for everyday use .
The system described in Source 4, known as the Sign Language Identifier (SLI), aims to bridge the communication gap by providing a web application that can convert Indian Sign Language gestures into text and translate speech into ISL. The app records gestures through a camera, processes them to produce corresponding text, and offers a reverse conversion from speech to ISL, allowing hearing individuals to communicate back with sign language users. This system empowers both parties to engage in seamless communication, thereby narrowing the understanding gap .
The SLI app ensures user accessibility by being a web-based application that can convert Indian Sign Language (ISL) to text and vice versa, catering to both signers and non-signers. For signers, it uses a webcam to capture gestures, which are then converted to text, facilitating communication with non-signers who don't understand ISL. For non-signers, the app converts speech to ISL, displaying corresponding gifs, making communication smoother for both parties . Additionally, the app has a user-friendly interface, requiring only basic interactions such as starting and stopping recordings, and does not require prior expertise in using technology, making it accessible for the target audience . The app is also portable and can be accessed via any device with a webcam and speaker, leveraging common technologies such as HTML and CSS for widespread accessibility ."
Mobile and cloud-based solutions for sign language interpretation are advantageous because they leverage AI and computer vision technologies to provide real-time translation of sign language to text and speech, enabling portability and ease of use through mobile devices. These solutions are economical and scalable, allowing integration with various applications like video chat and AI assistants, and they eliminate the dependency on physical hardware like gloves, which are expensive and cumbersome to use. Additionally, they offer user-friendly interfaces accessible on mobiles, thereby increasing accessibility for wider user groups, including those using Indian Sign Language .
Developing efficient Indian Sign Language (ISL) interpretation systems presents several challenges compared to American Sign Language (ASL). ISL is more complex due to its reliance on both hands for most of the signs, whereas ASL primarily uses single-handed gestures, making it simpler . Additionally, ISL lacks a comprehensive standardized database, unlike ASL, which has a well-established dataset, adding complexity to building effective ISL systems . This complexity is compounded by the need for extensive research and database creation for ISL, because researchers have traditionally focused more on ASL . Furthermore, ISL systems must handle dynamic parameters such as lighting conditions and varied gesture sizes, which require sophisticated image processing and machine learning techniques .
Neural networks and computer vision enhance sign language interpretation systems by enabling accurate and efficient gesture recognition. Convolutional Neural Networks (CNNs) are particularly effective for analyzing visual imagery, which allows for real-time conversion of sign language gestures into text or speech, and vice versa. This system facilitates seamless communication for hearing-impaired individuals, bridging the gap between them and those who are not familiar with sign language . Mobile solutions utilizing neural networks can translate gestures quickly and economically, leveraging cloud computing to process video inputs and provide instant translations, thus making the technology portable and accessible . These systems use computer vision techniques to improve gesture recognition accuracy, integrating noise reduction, and enhancement processes to better support model training and output prediction ."}
The absence of a comprehensive Indian Sign Language (ISL) dictionary significantly impacts the development of sign language interpretation technology. This limitation results in incomplete datasets for real-time conversation, as many datasets contain only alphabets or specific words, hindering the ability to frame proper grammatical sentences in ISL . Furthermore, ISL is inherently more complex than American Sign Language (ASL), as it relies on both hands, making recognition systems more challenging to develop . These challenges restrict the practical deployment and scalability of ISL interpretation technologies, as real-time and comprehensive communication capabilities are crucial but hard to achieve without a full dictionary .
Key components of the use case diagrams for the SLI project include actors (users such as deaf and dumb individuals and the system/admin), and use cases (like gesture capture, image preprocessing, feature extraction, and text output) for the conversion from ISL to text and from speech to ISL . These components contribute to system functionality by outlining the necessary steps for users to interact with the system, such as logging in, starting the webcam, performing sign language actions, and converting these gestures into text or speech into ISL, thus facilitating communication between deaf individuals and others . The diagrams are instrumental in defining expected behavior, helping to ensure that user interaction with the application is smooth and intuitive .