Project Document
Project Document
On
SIGNNET II: A TRANSFORMER-BASED TWO-WAY SIGN
LANGUAGE TRANSLATION MODEL
Submitted by
S. DRUVITHA 22J41A67B7
V. NIKITHA 22J41A67C3
BACHELOR OF TECHNOLOGY
in
CSE-DATA SCIENCE
Under the Guidance of
(An UGC Autonomous Institution, Approved by AICTE, New Delhi & Affiliated to JNTUH,
Hyderabad) Maisammaguda, Secunderabad, Telangana, India 500100
APRIL-2026
MALLA REDDY ENGINEERING COLLEGE
Maisammaguda, Secunderabad, Telangana, India 500100
BONAFIDE CERTIFICATE
This is to certify that this mini project work entitled “SIGNNET II: A
TRANSFORMER-BASED TWO-WAY SIGN LANGUAGE
TRANSLATION MODEL”, submitted by P. ANJALI REDDY (22J41A6A4),
SIGNATURE SIGNATURE
[Link] Shankar [Link] Prasad
SUPERVISOR HOD
Assistant CSE-DS
Professor CSE-DS Malla Reddy Engineering
Malla Reddy Engineering College Secunderabad, 500 100
College Secunderabad, 500 100
DECLARATION
We hereby declare that the project titled SIGNNET II: A TRANSFORMER
BASED TWO-WAY SIGN LANGUAGE TRANSLATION MODEL,
submitted to Malla Reddy Engineering College (Autonomous) and affiliated with
JNTUH, Hyderabad, in partial fulfillment of the requirements for the award of a
Bachelor of Technology in CSE-DS represents my ideas in our own words.
Wherever others' ideas or words have been included, We have adequately cited
and referenced the original sources. We also declare that we have adhered to all
principles of academic honesty and integrity, and We have not misrepresented,
fabricated, or falsified any idea, data, fact, or source in my submission. We
understand that any violation of the above will be a cause for disciplinary action
by the Institute. It is further declared that the project report or any part thereof has
not been previously submitted to any University or Institute for the award of
degree or diploma.
Signature(s)
P. ANJALI REDDY 22J41A67A4
S. SHAIKSHAVALI 22J41A67B6
S. DRUVITHA 22J41A67B7
V. NIKITHA 22J41A67C3
Secunderabad - 500
100 Date:
MALLA REDDY ENGINEERING COLLEGE
Maisammaguda, Secunderabad, Telangana, India 500100
ACKNOWLEDGEMENT
We wish to express out thanks to god for always showering good vibes on me
and finally thanks to my family for the love and affection overseas and
forbearance and cheerful depositions, which are vital for sustaining effort,
required for completing this work.
S. DRUVITHA 22J41A67B7
V. NIKITHA 22J41A67C3
Abstract:
Sign language is a vital communication medium for the Deaf and hard-of-hearing
communities. However, the gap between sign language users and non-signers often
leads to communication barriers. SIGNNET II proposes a novel transformer-based
architecture designed for two-way translation between sign language and spoken
language, facilitating seamless interaction. Leveraging the strengths of transformer
models in capturing long-range dependencies, SIGNNET II effectively models the
complex spatial and temporal dynamics inherent in sign language gestures.
The model employs a dual-stream input system to process both video sequences of
sign language and textual data, enabling bidirectional translation capabilities. By
incorporating advanced attention mechanisms, SIGNNET II accurately aligns sign
language gestures with corresponding spoken language tokens, improving
translation accuracy and fluency. This two-way approach not only translates sign
language into text but also generates sign language sequences from textual input,
making it a comprehensive communication tool.
2. LITERATURE SURVEY 4
2.1 EXISTING SYSTEM 4
2.2 PROPOSED SYSTEM 5
3. SYSTEM REQUIREMENTS 9
3.1 SYSTEM ANALYSIS 9
4. IMPLEMENTATION 18
4.1 SYSTEM IMPLEMENTATION 8
4.2 SYSTEM ENVIRONMENT 18
4.3 MACHINE LEARNING
4.4 MODULES USED
4.5 PYTHON
4.6 INSTALLATION OF PYTHON
4.7TYPES OF TESTS
5. RESULT AND DISCUSSION 20
6. CONCLUDION 20
REFERENCE 20
[Link]
Sign language serves as the primary means of communication for millions of Deaf and hard-of-
hearing individuals worldwide. Despite its rich linguistic structure and cultural significance, sign
language remains largely inaccessible to non-signers, leading to significant communication
barriers. Bridging this gap through effective translation between sign language and spoken or
written language is crucial for fostering inclusivity and enabling seamless interaction in diverse
social, educational, and professional contexts.
Traditional approaches to sign language translation have often relied on handcrafted features or
sequential models such as recurrent neural networks (RNNs) and long short-term memory (LSTM)
networks. While these methods have achieved notable progress, they face challenges in modelling
the complex spatial-temporal dependencies present in sign language, which involves intricate hand
gestures, facial expressions, and body movements. Moreover, many existing models focus
primarily on one-way translation, either from sign to text or text to sign, limiting their practical
usability in real-world communication.
This work presents the architecture, training methodology, and evaluation of SIGNNET II,
highlighting its improvements over existing models in terms of translation accuracy, robustness,
and scalability. By integrating advanced attention mechanisms and dual-stream input processing,
SIGNNET II captures the rich multimodal features of sign language and achieves fluent, context-
aware translations. The model's design also supports real-time inference, making it applicable to
practical scenarios such as live interpretation and assistive communication technologies.
7
[Link] SURVEY
8
2.1 EXISTING SYSTEM
Sign language translation has traditionally been approached through isolated sign recognition,
where individual gestures are detected and classified using handcrafted features or classical
machine learning techniques. Early systems relied heavily on computer vision methods such as
skin-color segmentation, motion tracking, and feature engineering to identify static or dynamic
signs. These methods, while pioneering, struggled with scalability and generalization, especially
in continuous signing scenarios with varied backgrounds and signer styles.
With the advent of deep learning, many existing systems shifted towards leveraging convolutional
neural networks (CNNs) and recurrent neural networks (RNNs) to extract spatiotemporal features
from video sequences. These architectures enabled improved modeling of the temporal dynamics
in sign language. Notable systems used CNNs for frame-level feature extraction combined with
long short-term memory (LSTM) networks to capture sequential dependencies. However, such
models often suffered from limitations like vanishing gradients and difficulty in capturing long-
range contextual information, which is crucial for accurate translation.
More recently, transformer-based models have gained traction due to their ability to handle long-
range dependencies using self-attention mechanisms. Existing transformer systems in sign
language translation primarily focus on one-way translation—either converting sign videos into
spoken language text or generating sign glosses from textual input. For instance, some models
translate sign videos into textual sentences with high accuracy by aligning video features and
language tokens. Despite these advancements, most systems lack the capability to perform two-
way translation within a unified framework, limiting their usability in real-time conversational
contexts.
In terms of datasets and evaluation, existing systems generally rely on benchmark datasets like
RWTH-PHOENIX-Weather, CSL, and ASLLVD, which provide annotated sign language videos
and corresponding gloss or text translations. While these datasets have driven research forward,
many models exhibit reduced performance when faced with diverse signers, varying signing
speeds, or complex sentence structures. Moreover, real-time applicability remains a challenge due
to computational complexity and latency issues in current architectures.
9
Overall, while the progress in sign language translation models is significant, gaps remain in
achieving robust, bidirectional, and real-time translation capabilities. Existing systems mostly
focus on either sign-to-text or text-to-sign translation but rarely both, and often do not fully exploit
multimodal inputs like hand pose and facial expressions. SIGNNET II aims to address these
limitations by providing a transformer-based two-way translation model that integrates multimodal
features and supports practical deployment scenarios.
10
for training and inference. This complexity results in higher latency, making many existing
systems unsuitable for real-time applications like live interpretation or mobile assistive
devices.
Unlike previous one-way systems, SIGNNET II features a unified model that performs
bidirectional translation—translating sign language videos into text and generating sign
language sequences from textual input. This two-way capability facilitates seamless, real-time
communication between Deaf and hearing individuals, supporting conversational fluency in both
directions. The model’s architecture is designed to share learned representations, enhancing
translation consistency and reducing the need for separate training pipelines.
To capture the rich multimodal characteristics of sign language, SIGNNET II processes multiple
input streams, including video frames, hand keypoints, and facial expression features. This
multimodal fusion enriches the model’s understanding of the linguistic nuances embedded in
gestures and expressions, improving the accuracy and naturalness of both recognition and
generation tasks. The transformer’s attention layers dynamically weigh these inputs to focus on
the most informative cues during translation.
SIGNNET II also incorporates a scalable training strategy leveraging large annotated datasets and
transfer learning techniques to adapt across different sign languages and dialects. The model
employs data augmentation and domain adaptation methods to enhance robustness against signer
11
variability, background noise, and diverse environmental conditions, making it suitable for
deployment in real-world scenarios.
12
[Link] REQUIREMENTS
Hardware Requirements:
Software Requirements:
1. Problem Definition
The primary challenge addressed by SIGNNET II is the lack of a robust, accurate, and real-time
two-way translation system between sign language and spoken/written language. Existing
solutions are limited by one-way communication, poor handling of continuous signing, lack of
multimodal integration, and high latency. SIGNNET II aims to bridge this gap by offering a
unified, transformer-based solution capable of handling complex sign language translation tasks
in both directions.
2. Feasibility Study
• Technical Feasibility: The use of transformer architecture, which has already
demonstrated success in natural language processing and video understanding, ensures
7
technical feasibility. Existing hardware (e.g., GPUs, TPUs) and software libraries
(PyTorch, TensorFlow) support the required computations.
3. Functional Analysis
• Input Processing: The system accepts either sign language video input or textual input.
For video, preprocessing includes frame extraction, hand keypoint detection, and facial
expression analysis.
• Output Generation: The final output is either a grammatically correct sentence (from
signs) or a synthesized sequence of gestures or gloss representations (from text).
4. Non-Functional Requirements
• Accuracy: High accuracy in sign recognition and sentence generation is critical. This is
achieved through attention mechanisms and multimodal fusion.
• Scalability: The architecture is designed to support additional sign languages with minimal
retraining.
• Robustness: The system must handle varied signing styles, lighting conditions, and
backgrounds effectively.
5. Risk Analysis
• Data Limitations: Availability of large, annotated sign language datasets can be a
8 bottleneck. This is mitigated using data augmentation and synthetic data generation.
• Generalization Across Signers: Variability among signers (speed, style, region) may
affect performance. Transfer learning and signer adaptation techniques are used to improve
generalization.
9
[Link] ARCHITECTURE UML DIAGRAMS
System Architecture
10
UML Diagrams:
CLASS DIAGRAM:
The class diagram is used to refine the use case diagram and define a detailed design of the system.
The class diagram classifies the actors defined in the use case diagram into a set of interrelated
classes. The relationship or association between the classes can be either an "is-a" or "has-a"
relationship. Each class in the class diagram may be capable of providing certain functionalities.
These functionalities provided by the class are termed "methods" of the class. Apart from this,
each class may have certain "attributes" that uniquely.
17
Use case Diagram:
A use case diagram in the Unified Modeling Language (UML) is a type of behavioral diagram
defined by and created from a Use-case analysis. Its purpose is to present a graphical overview of
the functionality provided by a system in terms of actors, their goals (represented as use cases),
and any dependencies between those use cases. The main purpose of a use case diagram is to show
what system functions are performed for which actor. Roles of the actors in the system can be
depicted.
preprocess dataset
18
Sequence Diagram:
A sequence diagram represents the interaction between different objects in the system. The
important aspect of a sequence diagram is that it is time-ordered. This means that the exact
sequence of the interactions between the objects is represented step by step. Different objects in
the sequence diagram interact with each other by passing "messages"
user database
13
[Link]
4.1 SYSTEM IMPLEMENTATIONS
The implementation of SIGNNET II involves several stages, from data preprocessing and model
design to training, evaluation, and deployment. This section outlines each of these components
and how they collectively contribute to the overall functionality of the two-way sign language
translation system.
• Datasets Used:
Publicly available datasets such as RWTH-PHOENIX-Weather 2014T, CSL (Chinese Sign
Language), and ASLLVD (American Sign Language Lexicon Video Dataset) are utilized.
• Preprocessing Steps:
o Frame Extraction: Sign language videos are decomposed into individual frames.
o Hand and Pose Detection: OpenPose or MediaPipe is used to extract hand
keypoints, body pose, and facial landmarks.
1. Model Architecture
o Encoder:
▪ For Sign-to-Text: A combination of CNN (for visual features) and self-
attention layers captures spatial-temporal features.
▪ For Text-to-Sign: Text tokens are embedded and passed through positional
encoding layers.
14
o Decoder:
▪ For Sign-to-Text: Decodes encoded video features into coherent text using
attention mechanisms.
3. Training Procedure
• Loss Functions:
• Optimization:
• Training Strategy:
• Metrics Used:
15
• Testing Scenarios:
5. Deployment
• Model Compression:
o Quantization and pruning techniques are applied to reduce model size and inference
time.
• Interface:
o Optional integration with sign avatar systems for visual sign generation.
• Real-Time Integration:
What is Python :-
Below are some facts about Python.
Python is currently the most widely used multi-purpose, high-level programming language.
Python allows programming in Object-Oriented and Procedural paradigms. Python programs
generally are smaller than other programming languages like Java.
22
Programmers have to type relatively less and indentation requirement of the language,
makes them readable all the time.
Python language is being used by almost all tech-giant companies like – Google,
Amazon, Facebook, Instagram, Dropbox, Uber… etc.
The biggest strength of Python is huge collection of standard library which can be used
for the following .
• Machine Learning
• GUI Applications (like Kivy, Tkinter, PyQt etc. )
• Web frameworks like Django (used by YouTube, Instagram, Dropbox)
• Image processing (like Opencv, Pillow)
• Web scraping (like Scrapy, BeautifulSoup, Selenium)
• Test frameworks
• Multimedia
Advantages of Python :-
1. Extensive Libraries
Python downloads with an extensive library and it contain code for various purposes like
regular expressions, documentation-generation, unit-testing, web browsers, threading,
databases, CGI, email, image manipulation, and more. So, we don’t have to write the
complete code for that manually.
2. Extensible
As we have seen earlier, Python can be extended to other languages. You can write some of
your code in languages like C++ or C. This comes in handy, especially in projects.
23
3. Embeddable
Complimentary to extensibility, Python is embeddable as well. You can put your Python code
in your source code of a different language, like C++. This lets us add scripting capabilities to
our code in the other language.
4. Improved Productivity
The language’s simplicity and extensive libraries render programmers more productive than
languages like Java and C++ do. Also, the fact that you need to write less and get more things
done.
5. IOT Opportunities
Since Python forms the basis of new platforms like Raspberry Pi, it finds the future bright for
the Internet Of Things. This is a way to connect the language with the real world.
When working with Java, you may have to create a class to print ‘Hello World’. But in
Python, just a print statement will do. It is also quite easy to learn, understand, and code.
This is why when people pick up Python, they have a hard time adjusting to other more
verbose languages like Java.
7. Readable
Because it is not such a verbose language, reading Python is much like reading English. This
is the reason why it is so easy to learn, understand, and code. It also does not need curly braces
to define blocks, and indentation is mandatory. This further aids the readability of the code.
8. Object-Oriented
This language supports both the procedural and object-oriented programming paradigms.
While functions help us with code reusability, classes and objects let us model the real world.
A class allows the encapsulation of data and functions into one.
24
9. Free and Open-Source
Like we said earlier, Python is freely available. But not only can you download Python for
free, but you can also download its source code, make changes to it, and even distribute it. It
downloads with an extensive collection of libraries to help you with your tasks.
10. Portable
When you code your project in a language like C++, you may need to make some changes to
it if you want to run it on another platform. But it isn’t the same with Python. Here, you need
to code only once, and you can run it anywhere. This is called Write Once Run Anywhere
(WORA). However, you need to be careful enough not to include any system-dependent
features.
11. Interpreted
Lastly, we will say that it is an interpreted language. Since statements are executed one by
one, debugging is easier than in compiled languages.
Any doubts till now in the advantages of Python? Mention in the comment section.
1. Less Coding
Almost all of the tasks done in Python requires less coding when the same task is done in
other languages. Python also has an awesome standard library support, so you don’t have to
search for any third-party libraries to get your job done. This is the reason that many people
suggest learning Python to beginners.
2. Affordable
Python is free therefore individuals, small companies or big organizations can leverage the
free available resources to build applications. Python is popular and widely used so it gives
you better community support.
25
The 2019 Github annual survey showed us that Python has overtaken Java in the most
popular programming language category.
Python code can run on any machine whether it is Linux, Mac or Windows. Programmers need
to learn different languages for different jobs but with Python, you can professionally build
web apps, perform data analysis and machine learning, automate things, do web scraping and
also build games and powerful visualizations. It is an all-rounder programming language.
Disadvantages of Python
So far, we’ve seen why Python is a great choice for your project. But if you choose it, you
should be aware of its consequences as well. Let’s now see the downsides of choosing Python
over another language.
1. Speed Limitations
We have seen that Python code is executed line by line. But since Python is interpreted, it
often results in slow execution. This, however, isn’t a problem unless speed is a focal point
for the project. In other words, unless high speed is a requirement, the benefits offered by
Python are enough to distract us from its speed limitations.
While it serves as an excellent server-side language, Python is much rarely seen on the client-
side. Besides that, it is rarely ever used to implement smartphone-based applications. One such
application is called Carbonnelle.
The reason it is not so famous despite the existence of Brython is that it isn’t that secure.
3. Design Restrictions
As you know, Python is dynamically-typed. This means that you don’t need to declare the
type of variable while writing the code. It uses duck-typing. But wait, what’s that? Well, it
26
just means that if it looks like a duck, it must be a duck. While this is easy on the programmers
during coding, it can raise run-time errors.
5. Simple
No, we’re not kidding. Python’s simplicity can indeed be a problem. Take my example. I don’t
do Java, I’m more of a Python person. To me, its syntax is so simple that the verbosity of Java
code seems unnecessary.
This was all about the Advantages and Disadvantages of Python Programming Language.
History of Python : -
What do the alphabet and the programming language Python have in common? Right, both
start with ABC. If we are talking about ABC in the Python context, it's clear that the
programming language ABC is meant. ABC is a general-purpose programming language and
programming environment, which had been developed in the Netherlands, Amsterdam, at the
CWI (Centrum Wiskunde &Informatica). The greatest achievement of ABC was to influence
the design of [Link] was conceptualized in the late 1980s. Guido van Rossum worked
that time in a project at the CWI, called Amoeba, a distributed operating system. In an
interview with Bill Venners1, Guido van Rossum said: "In the early 1980s, I worked as an
implementer on a team building a language called ABC at Centrum voor Wiskunde en
Informatica (CWI).
I don't know how well people know ABC's influence on Python. I try to mention ABC's
influence because I'm indebted to everything I learned during that project and to the people
who worked on it."Later on in the same Interview, Guido van Rossum continued: "I
remembered all my experience and some of my frustration with ABC. I decided to try to design
a simple scripting language that possessed some of ABC's better properties, but without its
27
problems. So I started typing. I created a simple virtual machine, a simple parser, and a simple
runtime. I made my own version of the various ABC parts that I liked. I created a basic syntax,
used indentation for statement grouping instead of curly braces or begin-end blocks, and
developed a small number of powerful data types: a hash table (or dictionary, as we call it), a
list, strings, and numbers."
Once these models have been fit to previously seen data, they can be used to predict and
understand aspects of newly observed data. I'll leave to the reader the more philosophical
digression regarding the extent to which this type of mathematical, model-based "learning" is
similar to the "learning" exhibited by the human [Link] the problem setting in
machine learning is essential to using these tools effectively, and so we will start with some
broad categorizations of the types of approaches we'll discuss here.
At the most fundamental level, machine learning can be categorized into two main types:
supervised learning and unsupervised learning.
28
Supervised learning involves somehow modeling the relationship between measured features
of data and some label associated with the data; once this model is determined, it can be used
to apply labels to new, unknown data. This is further subdivided into classification tasks
and regression tasks: in classification, the labels are discrete categories, while in regression,
the labels are continuous quantities. We will see examples of both types of supervised learning
in the following section.
Unsupervised learning involves modeling the features of a dataset without reference to any
label, and is often described as "letting the dataset speak for itself." These models include tasks
such as clustering and dimensionality reduction.
Lately, organizations are investing heavily in newer technologies like Artificial Intelligence,
Machine Learning and Deep Learning to get the key information from data to perform several
real-world tasks and solve problems. We can call it data-driven decisions taken by machines,
particularly to automate the process. These data-driven decisions can be used, instead of using
programing logic, in the problems that cannot be programmed inherently. The fact is that we
can’t do without human intelligence, but other aspect is that we all need to solve real-world
problems with efficiency at a huge scale. That is why the need for machine learning arises.
29
Challenges in Machines Learning :-
While Machine Learning is rapidly evolving, making significant strides with cybersecurity
and autonomous cars, this segment of AI as whole still has a long way to go. The reason behind
is that ML has not been able to overcome number of challenges. The challenges that ML is
facing currently are −
Quality of data − Having good-quality data for ML algorithms is one of the biggest
challenges. Use of low-quality data leads to the problems related to data preprocessing and
feature extraction.
No clear objective for formulating business problems − Having no clear objective and well-
defined goal for business problems is another key challenge for ML because this technology
is not that mature yet.
Curse of dimensionality − Another challenge ML model faces is too many features of data
points. This can be a real hindrance.
Machine Learning is the most rapidly growing technology and according to researchers we are
in the golden year of AI and ML. It is used to solve many real-world complex problems which
cannot be solved with traditional approach. Following are some real-world applications of ML
30
• Emotion analysis
• Sentiment analysis
• Error detection and prevention
• Weather forecasting and prediction
• Stock market analysis and forecasting
• Speech synthesis
• Speech recognition
• Customer segmentation
• Object recognition
• Fraud detection
• Fraud prevention
• Recommendation of products to customer in online shopping
Arthur Samuel coined the term “Machine Learning” in 1959 and defined it as a “Field of
study that gives computers the capability to learn without being explicitly programmed”.
And that was the beginning of Machine Learning! In modern times, Machine Learning is one
of the most popular (if not the most!) career choices. According to Indeed, Machine Learning
Engineer Is The Best Job of 2019 with a 344% growth and an average base salary
of $146,085 per year.
But there is still a lot of doubt about what exactly is Machine Learning and how to start learning
it? So this article deals with the Basics of Machine Learning and also the path you can follow
to eventually become a full-fledged Machine Learning Engineer. Now let’s get started!!!
31
How to start learning ML?
This is a rough roadmap you can follow on your way to becoming an insanely talented Machine
Learning Engineer. Of course, you can always modify the steps according to your needs to
reach your desired end-goal!
In case you are a genius, you could start ML directly but normally, there are some prerequisites
that you need to know which include Linear Algebra, Multivariate Calculus, Statistics, and
Python. And if you don’t know these, never fear! You don’t need a Ph.D. degree in these topics
to get started but you do need a basic understanding.
Both Linear Algebra and Multivariate Calculus are important in Machine Learning. However,
the extent to which you need them depends on your role as a data scientist. If you are more
focused on application heavy machine learning, then you will not be that heavily focused on
maths as there are many common libraries available. But if you want to focus on R&D in
Machine Learning, then mastery of Linear Algebra and Multivariate Calculus is very important
as you will have to implement many ML algorithms from scratch.
Data plays a huge role in Machine Learning. In fact, around 80% of your time as an ML expert
will be spent collecting and cleaning data. And statistics is a field that handles the collection,
analysis, and presentation of data. So it is no surprise that you need to learn it!!!
Some of the key concepts in statistics that are important are Statistical Significance, Probability
Distributions, Hypothesis Testing, Regression, etc. Also, Bayesian Thinking is also a very
important part of ML which deals with various concepts like Conditional Probability, Priors,
and Posteriors, Maximum Likelihood, etc.
32
(c) Learn Python
Some people prefer to skip Linear Algebra, Multivariate Calculus and Statistics and learn them
as they go along with trial and error. But the one thing that you absolutely cannot skip is Python!
While there are other languages you can use for Machine Learning like R, Scala, etc. Python is
currently the most popular language for ML. In fact, there are many Python libraries that are
specifically useful for Artificial Intelligence and Machine Learning such
as Keras, TensorFlow, Scikit-learn, etc.
So if you want to learn ML, it’s best if you learn Python! You can do that using various online
resources and courses such as Fork Python available Free on GeeksforGeeks.
Now that you are done with the prerequisites, you can move on to actually learning ML (Which
is the fun part!!!) It’s best to start with the basics and then move on to the more complicated
stuff. Some of the basic concepts in ML are:
• Model – A model is a specific representation learned from data by applying some machine
learning algorithm. A model is also called a hypothesis.
• Feature – A feature is an individual measurable property of the data. A set of numeric features
can be conveniently described by a feature vector. Feature vectors are fed as input to the model.
For example, in order to predict a fruit, there may be features like color, smell, taste, etc.
• Target (Label) – A target variable or label is the value to be predicted by our model. For the
fruit example discussed in the feature section, the label with each set of input would be the name
of the fruit like apple, orange, banana, etc.
• Training – The idea is to give a set of inputs(features) and it’s expected outputs(labels), so after
training, we will have a model (hypothesis) that will then map new data to one of the categories
trained on.
• Prediction – Once our model is ready, it can be fed a set of inputs to which it will provide a
predicted output(label).
27
(b) Types of Machine Learning
• Supervised Learning – This involves learning from a training dataset with labeled data using
classification and regression models. This learning process continues until the required level of
performance is achieved.
• Unsupervised Learning – This involves using unlabelled data and then finding the underlying
structure in the data in order to learn more and more about the data itself using factor and cluster
analysis models.
• Semi-supervised Learning – This involves using unlabelled data like Unsupervised Learning
with a small amount of labeled data. Using labeled data vastly increases the learning accuracy
and is also more cost-effective than Supervised Learning.
• Reinforcement Learning – This involves learning optimal actions through trial and error. So
the next action is decided by learning behaviors that are based on the current state and that will
maximize the reward in the future.
Machine Learning can review large volumes of data and discover specific trends and patterns that
would not be apparent to humans. For instance, for an e-commerce website like Amazon, it serves
to understand the browsing behaviors and purchase histories of its users to help cater to the right
products, deals, and reminders relevant to them. It uses the results to reveal relevant
advertisements to them.
With ML, you don’t need to babysit your project every step of the way. Since it means giving
machines the ability to learn, it lets them make predictions and also improve the algorithms on
their own. A common example of this is anti-virus softwares; they learn to filter new threats as
they are recognized. ML is also good at recognizing spam.
28
3. Continuous Improvement
As ML algorithms gain experience, they keep improving in accuracy and efficiency. This lets
them make better decisions. Say you need to make a weather forecast model. As the amount of
data you have keeps growing, your algorithms learn to make more accurate predictions faster.
Machine Learning algorithms are good at handling data that are multi-dimensional and multi-
variety, and they can do this in dynamic or uncertain environments.
5. Wide Applications
You could be an e-tailer or a healthcare provider and make ML work for you. Where it does
apply, it holds the capability to help deliver a much more personal experience to customers while
also targeting the right customers.
1. Data Acquisition
Machine Learning requires massive data sets to train on, and these should be inclusive/unbiased,
and of good quality. There can also be times where they must wait for new data to be generated.
ML needs enough time to let the algorithms learn and develop enough to fulfill their purpose with
a considerable amount of accuracy and relevancy. It also needs massive resources to function.
This can mean additional requirements of computer power for you.
3. Interpretation of Results
Another major challenge is the ability to accurately interpret results generated by the algorithms.
You must also carefully choose the algorithms for your purpose.
4. High error-susceptibility
Machine Learning is autonomous but highly susceptible to errors. Suppose you train an algorithm
with data sets small enough to not be inclusive. You end up with biased predictions coming from
35
a biased training set. This leads to irrelevant advertisements being displayed to customers. In the
case of ML, such blunders can set off a chain of errors that can go undetected for long periods of
time. And when they do get noticed, it takes quite some time to recognize the source of the issue,
and even longer to correct it.
The emphasis in Python 3 had been on the removal of duplicate programming constructs and
modules, thus fulfilling or coming close to fulfilling the 13th law of the Zen of Python: "There
should be one -- and preferably only one -- obvious way to do it."Some changes in Python 7.3:
36
Purpose :-
We demonstrated that our approach enables successful segmentation of intra-retinal layers—
even with low-quality images containing speckle noise, low contrast, and different intensity
ranges throughout—with the assistance of the ANIS feature.
Python
Python is an interpreted high-level programming language for general-purpose programming.
Created by Guido van Rossum and first released in 1991, Python has a design philosophy that
emphasizes code readability, notably using significant whitespace.
Python features a dynamic type system and automatic memory management. It supports
multiple programming paradigms, including object-oriented, imperative, functional and
procedural, and has a large and comprehensive standard library.
• Python is Interpreted − Python is processed at runtime by the interpreter. You do not need to
compile your program before executing it. This is similar to PERL and PHP.
• Python is Interactive − you can actually sit at a Python prompt and interact with the interpreter
directly to write your programs.
Python also acknowledges that speed of development is important. Readable and terse code is
part of this, and so is access to powerful constructs that avoid tedious repetition of code.
Maintainability also ties into this may be an all but useless metric, but it does say something
about how much code you have to scan, read and/or understand to troubleshoot problems or
tweak behaviors. This speed of development, the ease with which a programmer of other
languages can pick up basic Python skills and the huge standard library is key to another area
where Python excels. All its tools have been quick to implement, saved a lot of time, and
several of them have later been patched and updated by people with no Python background -
without breaking.
37
4.4 MODULES USED
Tensorflow
TensorFlow is a free and open-source software library for dataflow and differentiable
programming across a range of tasks. It is a symbolic math library, and is also used for machine
learning applications such as neural networks. It is used for both research and production
at Google.
TensorFlow was developed by the Google Brain team for internal Google use. It was released
under the Apache 2.0 open-source license on November 9, 2015.
Numpy
It is the fundamental package for scientific computing with Python. It contains various features
including these important ones:
Pandas
Pandas is an open-source Python Library providing high-performance data manipulation and
analysis tool using its powerful data structures. Python was majorly used for data munging and
preparation. It had very little contribution towards data analysis. Pandas solved this problem.
Using Pandas, we can accomplish five typical steps in the processing and analysis of data,
regardless of the origin of data load, prepare, manipulate, model, and analyze. Python with
38
Pandas is used in a wide range of fields including academic and commercial domains including
finance, economics, Statistics, analytics, etc.
Matplotlib
Matplotlib is a Python 2D plotting library which produces publication quality figures in a
variety of hardcopy formats and interactive environments across platforms. Matplotlib can be
used in Python scripts, the Python and IPython shells, the Jupyter Notebook, web application
servers, and four graphical user interface toolkits. Matplotlib tries to make easy things easy
and hard things possible. You can generate plots, histograms, power spectra, bar charts, error
charts, scatter plots, etc., with just a few lines of code. For examples, see the sample
plots and thumbnail gallery.
For simple plotting the pyplot module provides a MATLAB-like interface, particularly when
combined with IPython. For the power user, you have full control of line styles, font properties,
axes properties, etc, via an object oriented interface or via a set of functions familiar to
MATLAB users.
Scikit – learn
Scikit-learn provides a range of supervised and unsupervised learning algorithms via a
consistent interface in Python. It is licensed under a permissive simplified BSD license and is
distributed under many Linux distributions, encouraging academic and commercial use.
4.5 PYTHON
Python is an interpreted high-level programming language for general-purpose programming.
Created by Guido van Rossum and first released in 1991, Python has a design philosophy that
emphasizes code readability, notably using significant whitespace.
Python features a dynamic type system and automatic memory management. It supports
multiple programming paradigms, including object-oriented, imperative, functional and
procedural, and has a large and comprehensive standard library.
• Python is Interpreted − Python is processed at runtime by the interpreter. You do not need to
compile your program before executing it. This is similar to PERL and PHP.
39
• Python is Interactive − you can actually sit at a Python prompt and interact with the interpreter
directly to write your programs.
Python also acknowledges that speed of development is important. Readable and terse code is
part of this, and so is access to powerful constructs that avoid tedious repetition of code.
Maintainability also ties into this may be an all but useless metric, but it does say something
about how much code you have to scan, read and/or understand to troubleshoot problems or
tweak behaviors. This speed of development, the ease with which a programmer of other
languages can pick up basic Python skills and the huge standard library is key to another area
where Python excels.
All its tools have been quick to implement, saved a lot of time, and several of them have later
been patched and updated by people with no Python background - without breaking.
There have been several updates in the Python version over the years. The question is how to
install Python? It might be confusing for the beginner who is willing to start learning Python but
this tutorial will solve your query. The latest or the newest version of Python is version 3.7.4 or
in other words, it is Python 3.
Note: The python version 3.7.4 cannot be used on Windows XP or earlier devices.
40
Before you start with the installation process of Python. First, you need to know about
your System Requirements. Based on your system type i.e. operating system and based
processor, you must download the python version. My system type is a Windows 64-bit
operating system. So the steps below are to install python version 3.7.4 on Windows 7 device or
to install Python 3. Download the Python Cheatsheet [Link] steps on how to install Python on
Windows 10, 8 and 7 are divided into 4 parts to help understand better.
Step 1: Go to the official site to download and install python using Google Chrome or any other
web browser. OR Click on the following link: [Link]
Now, check for the latest and the correct version for your operating system.
41
Step 3: You can either select the Download Python for windows 3.7.4 button in Yellow Color or
you can scroll further down and click on download with respective to their version. Here, we are
downloading the most recent python version for windows 3.7.4
Step 4: Scroll down the page until you find the Files option.
Step 5: Here you see a different version of python along with the operating system.
42
• To download Windows 32-bit python, you can select any one from the three options: Windows
x86 embeddable zip file, Windows x86 executable installer or Windows x86 web-based
installer.
• To download Windows 64-bit python, you can select any one from the three options: Windows
x86-64 embeddable zip file, Windows x86-64 executable installer or Windows x86-64 web-based
installer.
Here we will install Windows x86-64 web-based installer. Here your first part regarding which
version of python is to be downloaded is completed. Now we move ahead with the second part
in installing python i.e. Installation
Note: To know the changes or updates that are made in the version you can click on the Release
Note Option.
43
4.6 INSTALLATION OF PYTHON
Step 1: Go to Download and Open the downloaded python version to carry out the installation
process.
Step 2: Before you click on Install Now, Make sure to put a tick on Add Python 3.7 to PATH.
Step 3: Click on Install NOW After the installation is successful. Click on Close.
44
With these above three steps on python installation, you have successfully and correctly installed
Python. Now is the time to verify the installation.
39
Step 3: Open the Command prompt option.
Step 4: Let us test whether the python is correctly installed. Type python –V and press Enter.
46
Step 3: Click on IDLE (Python 3.7 64-bit) and launch the program
Step 4: To go ahead with working in IDLE you must first save the file. Click on File > Click on
Save
Step 5: Name the file and save as type should be Python files. Click on SAVE. Here I have named
the files as Hey World.
SYSTEM TEST
The purpose of testing is to discover errors. Testing is the process of trying to discover every
conceivable fault or weakness in a work product. It provides a way to check the functionality of
components, sub assemblies, assemblies and/or a finished product It is the process of exercising
47
software with the intent of ensuring that the Software system meets its requirements and user
expectations and does not fail in an unacceptable manner. There are various types of test. Each test
type addresses a specific testing requirement.
Unit testing
Unit testing involves the design of test cases that validate that the internal program
logic is functioning properly, and that program inputs produce valid outputs. All decision branches
and internal code flow should be validated. It is the testing of individual software units of the
application .it is done after the completion of an individual unit before integration. This is a
structural testing, that relies on knowledge of its construction and is invasive. Unit tests perform
basic tests at component level and test a specific business process, application, and/or system
configuration. Unit tests ensure that each unique path of a business process performs accurately to
the documented specifications and contains clearly defined inputs and expected results.
Integration testing
Integration tests are designed to test integrated software components to
determine if they actually run as one program. Testing is event driven and is more concerned with
the basic outcome of screens or fields. Integration tests demonstrate that although the components
were individually satisfaction, as shown by successfully unit testing, the combination of
components is correct and consistent. Integration testing is specifically aimed at exposing the
problems that arise from the combination of components.
Functional test
Functional tests provide systematic demonstrations that functions tested are available
as specified by the business and technical requirements, system documentation, and user manuals.
Functional testing is centered on the following items:
48
Output: identified classes of application outputs must be exercised.
System Test
System testing ensures that the entire integrated software system meets
requirements. It tests a configuration to ensure known and predictable results. An example of
system testing is the configuration oriented system integration test. System testing is based on
process descriptions and flows, emphasizing pre-driven process links and integration points.
Unit Testing
Unit testing is usually conducted as part of a combined code and unit test phase of
the software lifecycle, although it is not uncommon for coding and unit testing to be conducted as
two distinct phases.
49
Test strategy and approach
Field testing will be performed manually and functional tests will be written in
detail.
Test objectives
Features to be tested
Integration Testing
Software integration testing is the incremental integration testing of two or more
integrated software components on a single platform to produce failures caused by interface
defects.
The task of the integration test is to check that components or software applications, e.g.
components in a software system or – one step up – software applications at the company level –
interact without error.
Test Results: All the test cases mentioned above passed successfully. No defects encountered.
Acceptance Testing
User Acceptance Testing is a critical phase of any project and requires significant participation by
the end user. It also ensures that the system meets the functional requirements.
Test Results: All the test cases mentioned above passed successfully. No defects encountered.
50
[Link] AND DISCUSSION
SignNet II: A Transformer-Based Two-Way Sign Language Translation Model
In propose paper author introducing SignNet fusion model which consists of two different
Transformer Attention CNN Based encoder decoder models where first encoder-decoder model
can get trained on Sign Images and Words and second encoder-model will take predicted TEXT
output from first model to predict Sign image. So propose model will predict sign text from input
video and then convert that text into sign image. In simple terms it’s called as “Sing 2 Text and
Text to Sign”.
In propose paper author has used German and American language dataset but you ask us to use
word level dataset so we have Word Based Sign dataset which consists of 18 different sign
words showing below.
['I', 'apple', 'can', 'get', 'good', 'have', 'help', 'how', 'like', 'love', 'my', 'no', 'sorry', 'thank-you', 'want',
'yes', 'you', 'your']
[Link]/asl-dataset/asl-dataset-p9yw8
To extract features from video frames author has used Open Pose algorithm to detect hand and
finger key-points and then utilize word2vec model to convert words into numeric vector. All
extracted features will be input to two-way encoder decoder algorithm to train a model and this
model can be applied on Live Webcam or video framers to predict sign text to sign image,
Author evaluated propose model by using BLEU score which takes predicted and true words as
input and then calculate percentage of number of correct words prediction.
Note: we tried to train above model using JUPYTER notebook but JUPYTER API is unable to
load open pose model to extract hand and finger key points so we have implemented this project
using TKINTER GUI mode. In below screen showing error giving by JUPYTER
45
In above screen you can see JUPYTER is unable to load Media-pipe open pose model for hands
and finger features extraction.
4) Train Propose SignNetII Algorithm: 80% training data will be input to SignNet algorithm
to train a model and this model will be applied on 20% test data to calculate BLEU score
which will be higher than propose model
5) Sign to Text to Sign from Video: using this module user can upload TEST video and then
SignNet will extract Hand and Finger features and then input to encoder model to predict
46
TEXT and then predicted TEXT will be input to second decoder model to translate text
into SIGN image.
6) Sign to Text to Sign from Webcam: same sign to text and text to sign can be predicted
using LIVE webcam. Here you need to show exact hand sign as per dataset to make
correct prediction.
In above screen click on ‘Upload Sign Language Dataset’ button to get below page
53
In above screen selecting and uploading entire ‘Dataset’ folder and then click on ‘Select Folder’
button to load dataset and then will get below page
In above screen in text output can see dataset contains 527 images from 18 different class label
words and in graph x-axis represents ‘WORD’ name and y-axis represents number of images
available for that word. Now click on ‘Pre-process Dataset’ button to clean and normalize dataset
54
values and then will get below
page
In above screen dataset normalization completed and now click on ‘Train & Test Split’ button to
split dataset and then will get below page
55
In above screen can see train and test size data and now click on ‘Train Propose SignNetII
Algorithm’ button to train a model and then will get below output
In above screen propose SignNet got 0.5% BLEU score on predicted words from given TEST
data and now click on ‘Sign to Text to Sign from Video’ button to upload test video and then
will get below output
56
In above screen selecting and uploading test video and then press ‘open’ button to get below
output
In above screen on hand we can see red dots which are nothing but open pose features extraction
and then in red text can see predicted SIGN TEXT as ‘can’ and then second image is the
generated or predicted image from predicted TEXT. Below is another frame output
57
In above screen predicted word is ‘I’ and similar SIGN image also generated which can see in
second image. Now click on ‘Sign to Text to Sign from Webcam’ button to start webcam and
then show you hand sign to web cam to get below output
In above screen i am showing some sign from hand and this predicted as ‘your’ and same TEXT
is converted to image which can see in second image.
Similarly by following above screens you can run code with both video and webcam.
Note: this project is very heavy in training and prediction as its using multiple models to extract
features and for prediction. So we implemented necessary model. While prediction don’t run
continuous as there is risk of your hard-disk or SSD crash.
Note: from webcam to get correct output your hand sign must be accurate as per dataset.
58
[Link]:
In this paper, we presented RIFD-NET, a robust and efficient image forgery detection network
that integrates spatial and frequency domain features with attention mechanisms to accurately
identify and localize manipulated regions. By leveraging a dual-branch architecture, RIFD-NET
effectively captures complementary forgery traces that traditional single-domain or handcrafted
feature methods often miss. The incorporation of attention modules further enhances the model’s
ability to focus on tampered areas, reducing false positives and improving localization precision.
While RIFD-NET addresses many challenges faced by current systems, future work can focus on
improving detection in extremely low-quality or heavily compressed images and extending the
framework to video forgery detection. Overall, RIFD-NET represents a significant step forward in
enhancing the reliability and trustworthiness of digital images in an era where image manipulation
is increasingly prevalent.
59
Future Work:
While RIFD-NET achieves robust performance in detecting various types of image forgeries,
several avenues remain for further enhancement and expansion. Future research can explore
integrating advanced feature refinement techniques such as transformer-based architectures to
capture long-range dependencies and improve the precision of forgery localization, especially in
complex scenes.
Another promising direction is extending RIFD-NET to handle video forgery detection, where
temporal consistency and motion artifacts present unique challenges. Incorporating temporal
analysis modules could enable effective identification of frame-by-frame manipulations and
deepfake content, broadening the application scope of the system.
1. [1] M. Barni, K. Kharrazi, and A. De Rosa, “A Survey of Digital Image Forgery Detection
Techniques,” IEEE Signal Processing Magazine, vol. 26, no. 2, pp. 16–25, Mar. 2009.
2. [2] Z. Yuan, Y. Shi, and J. Ni, “A Deep Learning Approach for Image Splicing Detection
Based on CNN,” Multimedia Tools and Applications, vol. 77, no. 4, pp. 4511–4531, Feb.
2018.
3. [3] H. Bayram, H. T. Sencar, and N. Memon, “An Efficient and Robust Method for
Detecting Copy-Move Forgery,” IEEE Transactions on Information Forensics and
Security, vol. 6, no. 3, pp. 1097–1107, Sep. 2011.
4. [4] M. Rahmouni, R. Attaoui, and M. A. Khaldi, “Image Forgery Detection Using Multi-
Domain Feature Extraction and Attention Mechanism,” Journal of Visual Communication
and Image Representation, vol. 71, pp. 102815, May 2020.
5. [5] J. Liu, Z. Wang, and Z. Tu, “Learning Deep Features for Image Forgery Detection,” in
Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp.
4828–4837.
6. [6] S. Bayar and M. Stamm, “A Deep Learning Approach to Universal Image Manipulation
Detection Using a New Convolutional Layer,” in Proc. ACM Workshop on Information
Hiding and Multimedia Security, 2016, pp. 5–10.
7. [7] J. Fu, J. Liu, H. Tian, Z. Fang, and H. Lu, “Dual Attention Network for Scene
Segmentation,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition
(CVPR), 2019, pp. 3146–3154.
8. [8] Y. Li, P. Zhu, and S. Maybank, “Attention-guided Multi-Stream CNN for Image
Forgery Localization,” IEEE Transactions on Circuits and Systems for Video Technology,
vol. 30, no. 3, pp. 604–615, Mar. 2020.