(10) Publication No: IN202641060000 A1 (43) Publication Date: 22-05-2026
(22) Date of Filing: 12-05-2026 Journal No: 21/2026
(19) Intellectual Property Office, India
(12) INDIAN PATENT APPLICATION
(54) Title: Pocketable sign language translator ai-based gesture and voice converter for persons with
hearing and speech impaired
(51) International Classification: (71) Name of Applicant(s):
G10L 15/22, G06N 3/08, G06N 3/04, G10L 15/26 Francis Xavier Engineering College,
G10L 15/18
(72) Name of Inventor(s):
(21) Application No: 202641060000 1. Maharaja M
(31) Priority Document No: -- 2. Esha Gobika G
3. Gopika J
(32) Priority Date: --
4. Dharanya Rani G
(86) International Application No: -- 5. Maajid Hassan S
Filing Date: -- 6. Dr. R. Ravi
(87) International Publication No: --
(61) Patent of Addition to Application No: --
Filing Date: --
(62) Divisional to Application No: --
Filing Date: --
(57) Abstract:
This invention establishes a system called the Pocketable Sign Language Translator: AI-Based Gesture & Voice Converter for
Persons with Hearing and Speech Impaired; this system is designed to improve communication accessibility for individuals with
hearing and speech impairments. The technologies used in creating this system include Artificial Intelligence (AI), Computer
Vision, Deep Learning, Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and Augmented Reality (AR). The system
will provide seamless two-way communication by using a pocket-mounted camera to capture hand gestures while a microphone
captures spoken language from normal persons. A Convolutional Neural Network (CNN) and Transformer-based AI model will
process and analyze gesture patterns to ensure accurate sign language recognition. The recognized signs are then converted into
natural voice output through an integrated Text-to-Speech (TTS) engine. Conversely, spoken input from a normal person is
converted into live text captions using Automatic Speech Recognition (ASR) models. These captions are instantly displayed on a
Wrist Band Display, providing a compact and accessible reading interface for deaf users. The Wrist Band Display may also
support synchronized visual feedback and Augmented Reality (AR)- based caption overlays for improved accessibility. This
proposed system offers a compact, portable, economical, and user-friendly design to provide dependable real-time
communication support for hearing- and speech-impaired individuals, thereby promoting independence, accessibility, and
inclusive participation in healthcare, education, workplaces, and public environments.
FORM 2
THE PATENTS ACT, 1970
[39 of 1970]
&
THE PATENTS RULES, 2003
COMPLETE SPECIFICATION
[See section 10 and rule 13]
“POCKETABLE SIGN LANGUAGE TRANSLATOR AI-BASED GESTURE AND VOICE
CONVERTER FOR PERSONS WITH HEARING AND SPEECH IMPAIRED”
Applicant Name Nationality Address
103/G2, Bypass Road,
Vannarpettai,
FRANCIS XAVIER ENGINEERING COLLEGE Indian
Tirunelveli - 627003,
Tamil Nadu, India
PREAMBLE OF THE DESCRIPTION
The following specification particularly describes the invention and the manner in which it is
to be performed.
Signature Not Verified
1 Digitally Signed.
Name: [Link]
Date: 12-May-2026 07:57:44
Reason: Patent Efiling
Location: DELHI
POCKETABLE SIGN LANGUAGE TRANSLATOR AI-BASED GESTURE AND
VOICE CONVERTER FOR PERSONS WITH HEARING AND SPEECH
IMPAIRED
FIELD OF INVENTION
[01] The present invention relates to the field of assistive communication technologies and
artificial intelligence systems. More specifically, it is a Pocketable Sign Language
Translator: AI-Based Gesture & Voice Converter for Persons with Hearing and Speech
Impaired that integrates computer vision, speech recognition, and real-time captioning
technologies for two-way communication support.
BACKGROUND OF THE INVENTION
[02] People with hearing and speech impairments face communication challenges that
affect their participation in social, educational, healthcare, and professional environments.
Traditional communication methods may not always provide fast or convenient interaction
in real-time situations.
[03] Existing communication systems often require bulky wearable devices, manual
interpretation, or external assistance. Many hearing- and speech-impaired individuals
depend on others for communication, reducing independence and accessibility in daily life.
[04] Advancements in Artificial Intelligence (AI), Deep Learning, and Computer Vision
technologies have enabled the development of intelligent communication systems capable
of recognizing hand gestures and converting them into meaningful speech or text.
However, many existing solutions still lack portability, real-time responsiveness,
contextual understanding, and user-friendly interfaces suitable for daily use.
2
[05] In addition, communication barriers during emergencies, healthcare interactions,
and public activities may increase social isolation and difficulty in accessing immediate
support for hearing- and speech-impaired individuals.
[06] Therefore, there is a need for a portable and intelligent communication system that
can accurately recognize sign language gestures, convert them into speech, and
simultaneously provide live captioning support for spoken conversations. The present
invention addresses these challenges by introducing a Pocketable Sign Language
Translator that integrates AI-driven gesture recognition, speech synthesis, voice-to-text
conversion, and smart display technologies for seamless communication assistance.
SUMMARY OF THE INVENTION
[07] The objective of the present invention is to develop a Pocketable Sign Language
Translator: AI-Based Gesture & Voice Converter for Persons with Hearing and Speech
Impaired that enables efficient two-way communication in real time. The system uses a
pocket-mounted camera to continuously capture dynamic hand gestures and process them
using advanced computer vision and deep learning algorithms.
[08] A Convolutional Neural Network (CNN) is utilized to extract spatial gesture
features, while a Transformer-based model analyses sequential gesture patterns for accurate
and context-aware recognition. The recognized signs are converted into natural voice
output through an integrated Text-to-Speech (TTS) engine. Conversely, spoken input from
normal persons is captured using a microphone and converted into live text captions
through Automatic Speech Recognition (ASR) models. These captions are displayed on a
3
dedicated Wrist Band Display and may also be extended using Augmented Reality (AR)
overlays to improve accessibility and communication effectiveness in dynamic
environments.
BRIEF DESCRIPTION OF FIGURES
[09] Fig. 1 Working Diagram of Integrated Elderly Care Platform Using Wearable
Sensors and Vision-Based Gesture Alerts
[10] Fig. 2 vitals displayed on the device and mobile application
[11] Fig. 3 Fall detected and displayed on the device and mobile application
[12] Fig. 4 To send message “SOS” using 0 fingers for hand gestures
[13] Fig. 5 To send message “Medical Help” using 1 fingers for hand gestures
[14] Fig. 6 To send message “Emergency Help” using 2 fingers for hand gestures
[15] Fig. 7 To send message “Need Food” using 3 fingers for hand gestures
[16] Fig. 8 To send message “Need Water” using 4 fingers for hand gestures
[17] Fig. 9 To send message “Washroom” using 5 fingers for hand gestures
DETAILED DESCRIPTION OF FIGURES
[18] Fig. 1 shows the prototype model of the Pocketable Sign Language Translator system
with a pocket-mounted AI camera, microphone, and Wrist Band Display. The system uses
CNN and Transformer- based models for sign language recognition. Recognized gestures
are converted into speech using TTS, while spoken input is converted into text using ASR.
The compact wearable design enables real-time communication assistance.
4
[19] Fig. 2 shows the bilingual communication capability of the system supporting Sign
→ Voice and Voice → Text interaction. Hand gestures captured by the AI camera are
converted into speech output using gesture recognition and TTS technology. Spoken input
from normal users is converted into live text captions using ASR. The translated text is
displayed on the Wrist Band Mobile Display for communication assistance.
[20] Fig. 3 Shows the real-time sign language recognition interface of the system. The
AI camera captures hand gestures and detects hand landmarks using computer vision
techniques. The recognized gesture is translated into the phrase “Thank you” and
converted into speech output. The interface also displays live video feed and gesture
tracking points.
[21] Fig. 4 shows the Gesture Dictionary management interface used for storing
customized hand gesture mappings. Each gesture is represented using a binary finger
pattern linked to a specific phrase or message. In this figure, the gesture pattern “01110” is
assigned to the phrase “I Need Water.” The module improves flexibility and personalized
communication support.
[22] Fig. 5 demonstrates the real-time Speech-to-Text communication interface of the
system. Spoken input from the normal user is captured through the microphone and
processed using ASR technology. In this figure, the sentence “Honesty is the best policy”
is converted into live text captions. The generated text is displayed on the wearable
interface for communication assistance.
5
[23] Fig. 6 demonstrates shows the Sign-to-Voice communication process between a
deaf person and a normal person. The AI camera captures sign language gestures and
processes them using deep learning-based gesture recognition algorithms. The recognized
signs are converted into speech output through the integrated TTS module. The system
enables real-time communication and accessibility support.
[24] Fig. 7 shows the wearable Wrist Band Display used during real- time
communication. The display provides translated messages and communication feedback
for the deaf user. The wearable design improves portability and enables continuous
interaction in daily environments.
[25] Fig. 8 shows the Voice → Text communication interface of the system. Spoken
input from the normal person is converted into live subtitles using ASR technology. The
generated text captions are displayed on the Wrist Band Mobile Display along with live
video feed. The interface assists hearing-impaired users during communication.
[26] Fig. 9 shows the hardware components of the Pocketable Sign Language Translator
system. The figure includes the pocket- mounted AI camera, hand gesture detection area,
and Wrist Band Mobile Display. The AI camera processes gestures using computer vision
techniques while the display provides translated communication output. The figure
demonstrates the compact wearable communication setup.
6
We Claim:
1. AI-based pocket-mounted sign language translation system for real-time Sign → Voice
communication.
2. Automatic Speech Recognition (ASR)-based Voice → Text conversion for live caption
generation and bilingual communication support.
3. CNN and LSTM-based gesture recognition model for accurate and contextual hand sign
detection.
4. Wristband mobile display system for showing live captions, captured video, and
communication output in real time.
5. Compact, portable, and wearable-free communication assistance system with Augmented
Reality (AR)-based caption support for improved accessibility and independent communication.
7
POCKETABLE SIGN LANGUAGE TRANSLATOR AI-BASED GESTURE AND VOICE
CONVERTER FOR PERSONS WITH HEARING AND SPEECH IMPAIRED
ABSTRACT
This invention establishes a system called the Pocketable Sign Language Translator: AI-
Based Gesture & Voice Converter for Persons with Hearing and Speech Impaired; this
system is designed to improve communication accessibility for individuals with hearing
and speech impairments. The technologies used in creating this system include Artificial
Intelligence (AI), Computer Vision, Deep Learning, Automatic Speech Recognition
(ASR), Text-to-Speech (TTS), and Augmented Reality (AR). The system will provide
seamless two-way communication by using a pocket-mounted camera to capture hand
gestures while a microphone captures spoken language from normal persons. A
Convolutional Neural Network (CNN) and Transformer-based AI model will process and
analyze gesture patterns to ensure accurate sign language recognition. The recognized
signs are then converted into natural voice output through an integrated Text-to-Speech
(TTS) engine. Conversely, spoken input from a normal person is converted into live text
captions using Automatic Speech Recognition (ASR) models. These captions are instantly
displayed on a Wrist Band Display, providing a compact and accessible reading interface
for deaf users. The Wrist Band Display may also support synchronized visual feedback
and Augmented Reality (AR)- based caption overlays for improved accessibility. This
proposed system offers a compact, portable, economical, and user-friendly design to provide
dependable real-time communication support for hearing-and speech-impaired individuals,
thereby promoting independence, accessibility, and inclusive participation in healthcare,
education, workplaces, and public environments.