0% found this document useful (0 votes)
33 views93 pages

Object Recognition for the Visually Impaired

The document is a project report titled 'Object Recognition for Visually Impaired,' submitted by students for a Bachelor of Technology in Information Technology at Anil Neerukonda Institute of Technology and Sciences. It outlines the development of an Android application that integrates voice commands and real-time object detection to assist visually impaired individuals, enhancing their independence and accessibility. The report includes sections on introduction, literature survey, system analysis, design, and testing, highlighting the project's objectives and significance in assistive technology.

Uploaded by

Sai
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
33 views93 pages

Object Recognition for the Visually Impaired

The document is a project report titled 'Object Recognition for Visually Impaired,' submitted by students for a Bachelor of Technology in Information Technology at Anil Neerukonda Institute of Technology and Sciences. It outlines the development of an Android application that integrates voice commands and real-time object detection to assist visually impaired individuals, enhancing their independence and accessibility. The report includes sections on introduction, literature survey, system analysis, design, and testing, highlighting the project's objectives and significance in assistive technology.

Uploaded by

Sai
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

OBJECT RECOGNITION FOR VISUALLY

IMPAIRED
A Project report submitted in partial fulfilment of the requirements for
the award of the degree of

BACHELOR OF TECHNOLOGY
IN
INFORMATION TECHNOLOGY

G Santhoshi Rupa A21126511021


S Chandra Shekhar Raju A21126511053
P Goutham Ganesh
A21126511051
M Vinay A21126511039

Under the Guidance of


Dr. B. Kiran Kumar
Associate Professor

DEPARTMENT OF INFORMATION TECHNOLOGY


ANIL NEERUKONDA INSTITUTE OF TECHNOLOGY AND
SCIENCES (AUTONOMOUS)
(Permanent Affiliation by Andhra University & Approved by AICTE
Accredited by NBA (ECE, EEE, CSE, IT, Mech. Civil & Chemical) & NAAC
A+) Sangivalasa, Bheemili Mandal, Visakhapatnam dist.(A.P)

2024-25

I
ANIL NEERUKONDA INSTITUTE OF TECHNOLOGY AND
SCIENCES (AUTONOMOUS)
(Permanent Affiliation by Andhra University & Approved by AICTE
Accredited by NBA (ECE, EEE, CSE, IT, Mech. Civil & Chemical) & NAAC
A+) Sangivalasa, Bheemili Mandal, Visakhapatnam dist.(A.P)

DEPARTMENT OF INFORMATION TECHNOLOGY

CERTIFICATE

This is to certify that the project report entitled “Object Recognition For
Visually Impaired” submitted by G Santhoshi Rupa (A21126511021), S
Chandra Shekhar Raju (A21126511053), M Vinay (A21126511059), P
Goutham Ganesh (A21126511051) , in partial fulfilment of the requirements
for the award of the degree of Bachelor of Technology in Information
Technology of Anil Neerukonda Institute of technology and sciences,
Visakhapatnam is a record of bonafide work carried out under my guidance and
supervision.

Project Guide Head of the Department


Dr. B. Kiran Kumar Dr. M. Rekha Sundari
Associate Professor Professor
Department of IT Department of IT
ANITS ANITS

II
III
ACKNOWLEDGEMENT

We would like to express our deep gratitude to our project guide Dr. B. Kiran
Kumar , Professor, Department of Information Technology, ANITS, for his guidance
with unsurpassed knowledge and immense encouragement. We are grateful to Prof.
[Link] Sundari, Head of the Department, Information Technology, for providing us
with the required facilities for the completion of the project work.
We are very much thankful to the Principal and Management, ANITS,
Sangivalasa, for their encouragement and cooperation to carry out this work.
We express our thanks to all teaching faculty of Department of IT, whose
suggestions during reviews helped us in accomplishment of our project. We would like to
thank all non-teaching staff of the Department of IT, ANITS for providing great
assistance in accomplishment of our project.
We would like to thank our parents, friends, and classmates for their
encouragement throughout our project period. At last but not the least, we thank everyone
for supporting us directly or indirectly in completing this project successfully.

PROJECT STUDENTS

G Santhoshi Rupa (A21126511021)

S Chandra Shekhar Raju (A21126511053)

M Vinay (A21126511039)

P Goutham Ganesh (A21126511051)

IV
DECLARATION

We hereby declare that the project work entitled “ OBJECT RECOGNITION FOR
VISUALLY IMPAIRED ” submitted to the Anil Neerukonda Institute of Technology and
Sciences is a record of an original work done by G Santhoshi Rupa (A21126511021), S
Chandra Shekhar Raju (A21126511053), M Vinay (A21126511039) , P Goutham
Ganesh (A21126511051) under the esteemed guidance of Dr. B. Kiran Kumar, Professor
of Information Technology, Anil Neerukonda Institute of Technology and Sciences, and this
project work is submitted in the partial fulfillment of the requirements for the award of
degree Bachelor of Technology in Information Technology. This entire project is done
with the best of our knowledge and have not been submitted for the award of any other
degree in any other universities.

G Santhoshi Rupa (A21126511021)

S Chandra Shekhar Raju (A21126511053)

M Vinay (A21126511039)

P Goutham Ganesh (A21126511051)

V
VI
INDEX

Chapter Name Page No.

Abstract i
List of Figures ii
List of Tables iii
List of Abbrevations iv

1. INTRODUCTION 11
1.1 Problem Statement and Motivation 12
1.2 Research Objectives 12
1.3 Project Scope and Direction 12
1.4 Impact, significance and contribution 13
1.5 Background Information 13
1.5.1 Project Field 13
1.5.2 Historical development prior to the project 13

2. LITERATURE SURVEY 14
3. SYSTEM ANALYSIS 18
3.1 Existing System 19
3.2 Disadvantages 19
3.3 Proposed System 19
3.4 Advantages 20
4. REQUIRMENT ANALYSIS 21
4.1 Function and non-functional Requirements 22
4.2 Hardware Requirements 23
4.3 Software Requirements 23
4.4 Architecture 24
5. SYSTEM DESIGN 25
5.1 Introduction of Input 26
5.2 DATA SET 27
5.3 UML Diagram 29
5.4 Data Flow Diagram 34

VII
6. ANDROID ENVIRONMENT 35
6.1 Software Installation 36
6.2 Software Development Life Cycle 38
7. CODING 40
7.1 [Link] 41
7.2 Object detection helper class 41
7.3 splash screen 42
8. SYSTEM STUDY AND TESTING 43
8.1 Feasibility Study 44
8.2 Types of Test & Test Cases 47
9. RESULTS 48
10. CONCLUSION 52
11. REFERENCES 54
12. PAPER PUBLICATIONS WITH PLAGIARISM REPORT 57

VIII
IX
ABSTRACT

This innovative application seamlessly integrates voice-activated commands, real-time


object detection, and communication features to deliver an intuitive and accessible user
experience. Users interact through clear voice prompts, enabling accurate interpretation
and hands-free operation for tasks such as making calls, sending messages, and managing
contacts. The system leverages advanced speech recognition and visual recognition
algorithms to assist users—especially those with visual impairments or mobility challenges
—in identifying objects within their surroundings, fostering independence and smart
interaction. By combining voice control with object detection, the application extends
beyond conventional voice assistants[3], offering a novel and inclusive interface. Future
enhancements will aim to improve recognition accuracy, reduce response time, and ensure
compatibility across multiple devices, highlighting the potential of integrating artificial
intelligence with assistive technology to promote greater accessibility and user
engagement. The system also features customizable voice profiles for personalized
responses and supports multilingual input for diverse users. Its user-friendly interface
ensures ease of use without requiring technical expertise. Continuous learning capabilities
allow the application to adapt to individual user preferences over time. Integration with
navigation tools can further aid users in unfamiliar environments. This application sets a
foundation for the next generation of intelligent assistive technologies.

Keywords: Voice Commands, Object Detection, Accessibility, Assistive Technology,


Voice Recognition, Hands-free Communication, Inclusive User Interface, Smart
Interaction, Visual Assistance.

X
XI
LIST OF FIGURES

[Link] Figure Number and Name Page No

1 4.4 Architecture 24

2 5.3.1 Use Case Diagram 30

3 5.3.2 Class Diagram 31

4 5.3.3 Sequence Diagram 31

5 5.3.4 Collaboration Diagram 32

6 5.3.5 Activity Diagram 32

7 5.3.6 Component Diagram 33

8 5.3.7 Deployment Diagram 33

9 5.3.8 ER Diagram 34

10 5.3.9 Data Flow Diagram 34

XII
LIST OF TABLES

[Link] Table Number and Name Page No


1. 8.3 Testing Cases 47

XIII
LIST OF ABBREVATIONS

[Link] Abbreviation Abbreviation Name

1. AI Artificial Intelligence

2. API Application Programming Interface

3. TTS Text-to-Speech

4. SDK Software Development Kit

5. IDE Integrated Development Environment

6. APK Android Package Kit

7. XML Extensible Markup Language

8. YOLO You Only Look Once

9. UX User Experience

10. UI User Interface

11. SSD Single Shot Multibox Detector

12. GUI Graphical User Interface

XIV
CHAPTER-1
INTRODUCTION

1
1. Introduction

This project focuses on developing an Android-based application to assist visually


impaired[2] individuals using real-time object detection. By integrating voice commands [1],
visual recognition, and communication features, the app enhances accessibility and
independence. It leverages machine learning and deep learning models for accurate and
efficient object identification. The solution aims to provide an intuitive user interface with
hands-free interaction and audio feedback. Ultimately, it bridges the gap between
technology and inclusivity, empowering users in their daily lives.

1.1 Problem Statement and Motivation


Existing applications lack a comprehensive solution for users requiring accessible and
intuitive interaction. The absence of seamless integration between voice commands, object
detection, and communication features hinders user experience. Current systems lack the
ability to add contacts through voice input and provide visual assistance. This application
addresses these gaps, aiming to enhance accessibility and inclusivity.

This application is motivated by the vision of enhancing user interactions through cutting-
edge technology. By combining voice-activated commands, object detection, and seamless
communication features, it aspires to provide an inclusive and intuitive user experience.
The motivation lies in creating a versatile tool that prioritizes accessibility and caters to
diverse user needs, showcasing the potential for innovation in user interface design.

1.2 Research Objectives

The project aims to develop an innovative application that seamlessly integrates voice-
activated commands, object detection, and communication features. The objective is to
enhance user experience by leveraging voice clarity for accurate interpretation,
incorporating visual recognition[2] through object detection, and enabling hands-free calls
and [Link] project seeks to create a versatile, inclusive, and intuitive user interface,
addressing a variety of user needs and showcasing the potential for advanced technological
solutions.

1.3 Project Scope and Direction

The scope of this project encompasses the development of an application that


revolutionizes user interactions by integrating voice-activated commands, object detection,
and communication features. It includes accurate voice interpretation, visual recognition
through object detection, hands-free calls, messaging, and seamless contact management
via voice

2
3
input. The application aims to provide a versatile and visually assisted user experience,
demonstrating potential advancements in inclusive and intuitive interface design.
Additionally, it enhances accessibility for visually impaired [2] users, ensuring ease of
navigation and interaction. Furthermore, it integrates real-time feedback mechanisms to
refine user experience and adaptability. This project sets the foundation for future
advancements in assistive technology, bridging the gap between digital interfaces and
human interaction.
1.4 Impact, Significance and Contribution

This study's findings advance the field of assistive technology by:


1. Improves accessibility for impaired individuals through voice and object recognition.
2. Integrates advanced AI and machine learning for seamless real-time assistance.
3. Enhances user independence and daily navigation with smart assistive
technology. 4 Provides a foundation for future innovations in assistive and IoT-
based solutions.

1.5 Background Information

1.5.1 Project Field

This study is at the intersection of deep learning, computer vision, and assistive
technology[2]. Object recognition plays a crucial role in enhancing accessibility for visually
impaired individuals by enabling them to interact with their surroundings through audio
feedback. Leveraging deep learning models, object detection and classification can provide
real-time assistance, allowing users to navigate their environment more effectively and
independently.

1.5.2 Historical Development. Prior to the project

The advancement of deep learning-based object recognition [13] has significantly improved
assistive technologies:
1. Viola-Jones Algorithm (2001) introduced real-time object detection using Haar
cascades, laying the foundation for modern detection techniques.
2. YOLO (You Only Look Once) (2016) revolutionized real-time object detection
by providing fast and accurate predictions, making it suitable for real-world
applications.
3. Efficient Net (2019) introduced scalable deep learning[12] models that balance
performance and efficiency, improving accuracy while reducing computational
costs.
4. Recent advancements in AI-powered assistive technology have integrated speech
synthesis and edge computing, allowing visually impaired[2] individuals to receive

4
CHAPTER-2
LITERATURE
SURVEY

5
6
2. Literature Survey

2.1 Related Work

This paper presents the development of home appliances based on voice command using
Android. This system has been designed to assist and provide support to elderly and
disabled people at home. Google application has been used as voice recognition and
processes the voice input from the smartphone. In this paper, the voice input has been
captured by the Android device and sent to the Arduino Uno. The Bluetooth module in
Arduino Uno receives the signal[10] and processes the input to control the light and fan. The
proposed system is intended to control electrical appliances with a relatively user-friendly
interface and ease of installation. We have demonstrated up to 20 meters of range to
control the home appliances[1] via Bluetooth. The system effectively reduces the need for
physical interaction with switches, promoting convenience and accessibility. Future
enhancements may include integrating Wi-Fi modules for extended range and IoT-based
remote control capabilities. The components used in this project are low-cost and easily
available in the market, making it a cost-effective solution. The system is flexible and can
be modified to control additional appliances as per the user’s requirement. Implementation
of safety features such as overload protection can further enhance its reliability. Voice-
controlled automation not only improves the lifestyle of users but also contributes to
energy efficiency by ensuring better control over devices. This project showcases the
potential of combining Android technology with embedded systems to create smart and
accessible home automation solutions.

Mobile devices or portable smart gadgets are largely used in our day-to-day life now.
These devices can be very useful in helping people who are visually impaired and assist
them to make their life easier. This paper presents an Android application that recognizes
objects using real-time object and text detection by scanning them. The application can run
on the device without any remote server. Our solution uses voice feedback to tell the user
about the detected object. The app doesn't need any photograph to detect the object. To
enable robust recognition, we first segment the object from the background using
TensorFlow Machine Learning API based on iterative diagram cuts. We then formulate the
recognition problem as an instance retrieval task and the user gets to know the object
through a text-to-speech method. The system is designed to help a visually impaired [2]
person know about the environment around them. This app can be run on any low-end
smartphone devices. Additionally, the lightweight design ensures minimal battery
consumption, enhancing usability for long durations. The system architecture supports
offline functionality, ensuring reliable performance in areas without internet connectivity.
Future developments may include multi-language support and gesture-based controls for
more intuitive interaction.

7
This paper presents a Voice Assistant, an Android-based personal assistant application for
mobile phones, allowing voice control for the Serbian language. The native interface is
provided for a large vocabulary continuous speech recognition system based on the open-
source Kaldi speech recognition toolkit. Several acoustic models were trained using a
database of about 70,000 utterances, by adding different amounts of noise. The results are
provided for a test database[11] of about 4,500 utterances and a test vocabulary of more than
14,000 words. The system demonstrates high accuracy in recognizing voice commands,
even in noisy environments. It supports various functionalities such as making calls,
sending messages, and performing web searches, all through voice interaction. The use of
Kaldi allows flexibility and customization, enabling developers to further optimize the
system for specific applications. The application provides a user-friendly interface that
caters to both tech-savvy and non-technical users. Future improvements could involve real-
time learning capabilities and support for additional regional dialects, enhancing
accessibility and user experience.

Portable devices are nowadays largely used among people. These devices have a lot of
potential in aiding visually impaired people on a daily basis. This paper presents an
android application for smartphone, especially made to assist these persons. The
application uses MEMS sensors from a smartphone and also the information received from
a few external sensorial modules. The hardware modules form together an assistive
portable system. Communication between smartphone and external modules is made via
Bluetooth and Wi- Fi. Because the users are visually impaired, this application interface is
designed to meet the necessary requirements. Moreover, the communication between
smartphone and its user is made through a text-to- speech software module. The assistive
activities covered by these modules, range from making a phone call, to indoor and
outdoor guidance. According to tests, the assistant system that uses this android application
proves to be efficient, portable, small, cost-effective and does not require many hours of
training.

In this modern century, everyone benefits from smart assistants. where we can converse
with intelligent helpers and ask robots questions. Nearly everyone in the globe is drawn to
this technology in a variety of ways, including smartphones, laptops, desktops, etc. It is a
smart assistant with speech recognition intelligence[6] that accepts user input in the form of
text or voice, processes it, and then provides the output in a variety of ways, such as by
speaking the user's search results. Google Assistant, Siri, and Alexa are a few examples of
intelligent assistants. They are unable to perform voice recognition, and there are still
problems with human interaction. Those Google Assistant problems, in order to
communicate with users, they must have access to WiFi and the internet also they are
using very limited language so users can get many problems in using them. Google
Assistant is a feature of Android smartphones. However, Internet connections are required
for this application to function.

8
9
However, our proposed system can also function without an Internet connection. So that to
retifice the problem the to implement multiple languages to that so that we are using
raspberry pi for that in the raspberry pi the data of multiple languages is loaded and stored
so that the user can access any language through the voice assistant [3] so that the user gets
information about recent apps, daily news, geolocation, and Wikipedia. Through voice
assistance, the user also gets benefits in using automation technology.

10
CHAPTER-3
SYSTEM
ANALYSIS

11
12
3. SYSTEM ANALYSIS
The System Analysis identifies key limitations in current assistive technologies, such as
poor integration and limited accessibility. The proposed system addresses these gaps with
a unified approach combining voice commands, object detection, and communication
features. It ensures offline functionality and user-friendly design tailored for the visually
impaired. This approach enhances inclusivity, usability, and independence for users.
3.1 Existing System
Current systems lack a unified solution for comprehensive user interaction. They often
miss seamless integration of voice commands[1], object detection, and communication
features, resulting in fragmented user experiences. The absence of intuitive voice-driven
additions to contacts and visual assistance through object detection leaves users without a
versatile and inclusive interface, prompting the need for a more advanced solution. Most
existing applications rely heavily on internet connectivity, limiting their usability in offline
scenarios. Additionally, they often fail to support regional languages, which restricts
accessibility for a diverse user base. The user interfaces are generally not optimized for
visually impaired individuals, making navigation difficult. Furthermore, security and data
privacy concerns are often overlooked, putting user information at risk.
3.2 Disadvantages
1. Lack of unified solution: Current systems lack comprehensive user interaction.
2. Fragmented experiences: Missing integration of voice commands, object detection,
communication.
3. Need for advancement: Absence of intuitive features prompts the search for improvement.

3.3 Proposed System


The proposed system introduces an innovative Android-based application designed
specifically for the visually impaired, seamlessly integrating voice-activated commands,
object detection, and communication features to create a comprehensive and accessible
user experience. Users can interact naturally with the application using voice input,
supported by advanced speech recognition[6] and clear voice clarity, allowing for easy
navigation and control without the need for visual cues.
The core component of the application is real-time object detection powered by efficient
deep learning models, which allows the system to recognize and announce objects in the
environment through voice output. This feature significantly improves spatial awareness
and environmental understanding[7] for visually impaired users, enhancing safety and
independence.
Additionally, the system supports hands-free communication, enabling users to make calls,

13
send messages, and manage contacts entirely through voice commands. Contact
management is simplified with the ability to add, delete, or call contacts using natural
language[8], improving usability for those unable to rely on touch-based navigation. To
ensure wide accessibility, the application is designed to function smoothly even on low-
end smartphones and offers offline capabilities for essential functions like object detection
and calling saved contacts, reducing dependency on continuous internet connectivity.
The interface incorporates inclusive design principles, featuring large, high-contrast icons,
screen reader compatibility, and support for regional languages. Integration with text-to-
speech (TTS) and speech-to-text (STT) services allows seamless interaction and auditory
feedback for both visually and physically challenged individuals. Security and data privacy
are also prioritized, with local data processing where possible and user permissions
managed transparently. Overall, the proposed system bridges the gap between assistive
needs and modern smartphone capabilities [4], promoting independence, inclusivity, and
ease of use in daily life.
3.4 Advantages
1. Unified Need application: Integrating voice commands, object detection, and
communication features seamlessly.
2. Effortless user interaction: Precise voice commands and clear voice clarity aid interaction.
3. Visual enhancement: Object detection enables hands-free calls, messaging,
and intuitive contact management.
4. Offline Functionality: Core features such as object detection and calling saved
contacts work without internet access, enhancing reliability.
5. Accessibility Optimized Interface: Large icons, high-contrast themes, and screen
reader compatibility improve usability for visually impaired users.
6. Secure and Private: Emphasis on local data processing and secure permission
handling ensures user data privacy.

14
15
CHAPTER-4
REQUIREMENT
ANALYSIS

16
4. REQUIREMENT ANALYSIS
Requirement’s analysis is very critical process that enables the success of a system or
software project to be assessed. Requirements are generally split into two types: Functional
and non-functional requirements.
4.1 Functional and non-functional requirements
Functional Requirements: These are the requirements that the end user specifically
demands as basic facilities that the system should offer. All these functionalities need to be
necessarily incorporated into the system as a part of the contract. These are represented or
stated in the form of input to be given to the system, the operation performed and the
output expected. They are basically the requirements stated by the user which one can see
directly in the final product, unlike the non-functional requirements.
Examples of functional requirements:
1. Authentication of user whenever he/she logs into the system
2. System shutdown in case of a cyber-attack
3. Voice command[5] execution for making phone calls or sending messages
4. Object detection and recognition with voice feedback
5. Real-time text recognition and conversion to speech for visually impaired users
Non-functional requirements
These are basically the quality constraints that the system must satisfy according to the
project contract. The priority or extent to which these factors are implemented varies from
one project to other. They are also called non-behavioral requirements.
They basically deal with issues like:
1. Portability
2. Security
3. Maintainability
4. Reliability
5. Scalability
6. Performance
7. Reusability
8. Flexibility
Examples of non-functional requirements :

1. Emails should be sent with a latency of no greater than 12 hours from such an activity.

17
18
2. The processing of each request should be done within 10 seconds
3. The site should load in 3 seconds whenever of simultaneous users are > 10000
4. The system must be accessible on Android devices with OS version 6.0 and above.
5. The application should function offline with at least 90% of the core features available.
6. User data must be encrypted using AES-256 encryption standards.
7. The system should be able to handle at least 1,000 voice commands per hour
without degradation performance.

4.2 Hardware Requirements


1. Processor - I3/Intel Processor
2. RAM - 8 GB
3. Hard Disk - 1TB

4.3 Software Requirements


1. Operating System - Windows 10
2. JDK - java
3. Plugin - Kotlin
4. SDK - Android
5. IDE - Android studio

19
4.4. Architecture
The system architecture is designed to enable seamless interaction through voice
commands[5]. The user initiates input via the microphone, which is processed by the speech
recognizer. The converted text is analyzed, and corresponding functions like object
detection, contact addition, or SMS sending are triggered. The architecture supports
modular integration of features such as object detection using a dedicated model and
communication via content providers. This layered design ensures scalability, easy
maintenance, and adaptability across devices.

20
21
CHAPTER-5
SYSTEM
DESIGN

22
5. SYSTEM DESIGN
System design is the blueprint of the software system, translating requirements into
detailed architecture. It defines the structure, components, interfaces, and data flow for
efficient functioning. The goal is to ensure scalability, performance, security, and
maintainability of the application. Both high-level design (HLD) and low-level design
(LLD) are considered to meet functional and non-functional requirements. A well-
structured design serves as the foundation for smooth implementation and deployment of
the system.

5.1 Introduction of Input Design

INPUT DESIGN
The input design is the link between the information system and the user. It comprises the
developing specification and procedures for data preparation and those steps are necessary
to put transaction data in to a usable form for processing can be achieved by inspecting the
computer to read data from a written or printed document or it can occur by having people
keying the data directly into the system. The design of input focuses on controlling the
amount of input required, controlling the errors, avoiding delay, avoiding extra steps and
keeping the process simple. The input is designed in such a way so that it provides security
and ease of use with retaining the privacy. Input Design considered the following things:
1. What data should be given as input?
2. How the data should be arranged or coded?
3. The dialog to guide the operating personnel in providing input.
4. Methods for preparing input validations and steps to follow when error occur.

OBJECTIVES

1. Input Design is the process of converting a user-oriented description of the input into
a computer-based system. This design is important to avoid errors in the data input
process and show the correct direction to the management for getting correct
information from the computerized system.

2. It is achieved by creating user-friendly screens for the data entry to handle large
volume of data. The goal of designing input is to make data entry easier and to be free
from errors. The data entry screen is designed in such a way that all the data
manipulates can

23
24
be performed. It also provides record viewing facilities.

3. When the data is entered it will check for its validity. Data can be entered with the
help of screens. Appropriate messages are provided as when needed so that the user
will not be in maize of instant. Thus the objective of input design is to create an input
layout that is easy to follow.

OUTPUT DESIGN
A quality output is one, which meets the requirements of the end user and presents the
information clearly. In any system results of processing are communicated to the users and
to other system through outputs. In output design it is determined how the information is to
be displaced for immediate need and also the hard copy output. It is the most important
and direct source information to the user. Efficient and intelligent output design improves
the system’s relationship to help user decision-making.
1. Designing computer output should proceed in an organized, well thought out
manner; the right output must be developed while ensuring that each output
element is designed so that People will find the system can use easily and
effectively. When analysis design computer output, they should Identify the
specific output that is needed to meet the requirements.
2. Select methods for presenting information
3. Create document, report, or other formats that contain information produced by
the system.
The output form of an information system should accomplish one or more of the
following objectives.
1. Convey information about past activities, current status or projections of the Future.
2. Signal important events[10], opportunities, problems, or warnings
3. Trigger an action.
4. Confirm an action.
5.2 DATA SET
The dataset includes a wide variety of indoor and outdoor images to ensure balanced and
diverse object detection scenario. Indoor images capture household items, furniture, doors,
and appliances to aid navigation inside homes. Outdoor scenes feature objects like
vehicles, street signs, people, and pathways for real-world environment interaction. All
images are labeled with bounding boxes and object classes to support training robust
detection models. The dataset is optimized for Android deployment to assist visually
impaired users in real- time object recognition. A mix of lighting conditions and
backgrounds ensures the model performs reliably across different environments.

25
1. OUT DOOR

2. IN DOOR

26
27
5.3 UML DIAGRAM

1. UML stands for Unified Modelling Language. UML is a standardized general-


purpose modelling language in the field of object-oriented software engineering.
The standard is managed, and was created by, the Object Management Group.
2. The goal is for UML to become a common language for creating models of object-
oriented computer software. In its current form UML is comprised of two major
components: A Meta-model and a notation. In the future, some form of method or
process may also be added to; or associated with, UML.
3. The Unified modelling Language is a standard language for specifying,
Visualization, Constructing and documenting the artifacts of software system, as
well as for business modelling and other non-software systems.
4. The UML represents a collection of best engineering practices that have proven
successful in the modelling of large and complex systems.
5. The UML is a very important part of developing objects-oriented software and the
software development process. The UML uses mostly graphical notations to express
the design of software projects.

GOALS

The Primary goals in the design of the UML are as follows:

1. Provide users a ready-to-use, expressive visual modelling Language so that they


can develop and exchange meaningful models.
2. Provide extendibility and specialization mechanisms to extend the core concepts.
3. Be independent of particular programming languages and development process.
4. Provide a formal basis for understanding the modelling language.
5. Encourage the growth of OO tools market.
6. Support higher level development concepts such as collaborations, frameworks,
patterns and components.
7. Integrate best practices.

28
5.3.1 USE CASE DIAGRAM

A use case diagram in the Unified Modeling Language (UML) is a type of behavioral
diagram defined by and created from a Use-case analysis. Its purpose is to present a graphical
overview of the functionality provided by a system in terms of actors, their goals (represented
as use cases), and any dependencies between those use cases. The main purpose of a use case
diagram is to show what system functions are performed for which actor. Roles of the actors
in the system can be depicted.

29
30
5.3.2 CLASS DIAGRAM

In software engineering, a class diagram in the Unified Modelling Language (UML) is a type
of static structure diagram that describes the structure of a system by showing the system's
classes, their attributes, operations (or methods), and the relationships among the classes. It
explains which class contains information.

5.3.3 SEQUENCE DIAGRAM

A sequence diagram in Unified Modelling Language (UML) is a kind of interaction diagram


that shows how processes operate with one another and in what order. It is a construct of a
Message Sequence Chart. Sequence diagrams are sometimes called event diagrams, event
scenarios, and timing diagrams.

31
5.3.4 COLLABORATION DIAGRAM

In collaboration diagram the method call sequence is indicated by some numbering technique
as shown below. The number indicates how the methods are called one after another. We
have taken the same order management system to describe the collaboration diagram. The
method calls are similar to that of a sequence diagram. But the difference is that the sequence
diagram does not describe the object organization whereas the collaboration diagram shows
the object organization.

5.3.5 ACTIVITY DIAGRAM

Activity diagrams are graphical representations of workflows of stepwise activities and


actions with support for choice, iteration and concurrency. In the Unified Modelling
Language, activity diagrams can be used to describe the business and operational step-by-step
workflows of components in a system. An activity diagram shows the overall flow of control.

32
33
5.3.6 COMPONENT DIAGRAM

5.3.7 DEPLOYMENT DIAGRAM

34
5.3.8 ER DIAGRAM

5.4 DATA FLOW DIAGRAM

35
36
CHAPTER-6
ANDROID
ENVIRONMENT

37
6. ANDROID ENVIRONMENT

The Android environment provides a powerful open-source platform for mobile


application development. It includes the Android SDK, Android Studio, Java/Kotlin
support, and a rich set of APIs. Android offers a user-friendly interface and flexible tools
for building feature- rich apps. It supports rapid development and deployment across a
wide range of devices.

6.1 Software Installation


Software Installation of JDK
Kit
This Java Development Kit (JDK) allows you to code and run Java programs. It's possible
that you install multiple JDK versions on the same PC. But it’s recommended that you
install only latest version.
How to install Java for Windows
Step 1 . Go to link Click on JDK Download for Java

Step 2 : Next ,
1. Accept License Agreement
2. Download Java 8 JDK for your version 32 bit or JDK 8 download for windows 10
64 bit.

Step 3 : when you click on the Installation link the popup will be open. Click on I reviewed
and accept the Oracle Technology Network License Agreement for Oracle Java SE and you
will be redirected to the login page. If you don't have an oracle account you can easily sign
up by adding basics details of yours

38
39
Step 4 : once the Java JDK 8 download is complete, run the exe for install JDK. Click Next

Step 5 : Select the PATH to install Java in Windows and click next./

Step 6 : Once you install Java in windows, click

Close If you see a screen like below, Java is

installed.

Android Studio IDE and SDK Installation

Installing Android software is probably the most challenging part of this


project After Installation:

STEPS FOR EXECUTING THE PROJECTS

Step 1: Open Android Studio

Step2: Choose a virtual device or Physical device from the menu

Step3: Click on the project Run

Step4: View the application performance on virtual or Physical device.

40
6.2 Software Development Life Cycle
The meaning of Agile is swift or versatile. “Agile process model" refers to a software
development approach based on iterative development. Agile methods break tasks into
smaller iterations, or parts do not directly involve long term planning. The project scope
and requirements are laid down at the beginning of the development process. Plans
regarding the number of iterations, the duration and the scope of each iteration are clearly
defined in advance. Each iteration is considered as a short time "frame" in the Agile
process model, which typically lasts from one to four weeks. The division of the entire
project into smaller parts helps to minimize the project risk and to reduce the overall
project delivery time requirements. Each iteration involves a team working through a full
software development life cycle including planning, requirements analysis, design, coding,
and testing before a working product is demonstrated to the client.

41
42
Principles of Agile Model

1. To establish close contact with the customer during development and to gain a clear
understanding of various requirements, each Agile project usually includes a customer
representative on the team. At the end of each iteration stakeholders and the customer
representative review, the progress made and re-evaluate the requirements.
2. Requirement change requests from the customer are encouraged and efficiently
incorporated.
3. It emphasizes on having efficient team members and enhancing communications
among them is given more importance. It is realized that enhanced communication
among the development team members can be achieved through face-to-face
communication rather than through the exchange of formal documents.
4. It is recommended that the development team size should be kept small (5 to 9 people)
to help the team members meaningfully engage in face-to-face communication and
have colla realize collaborative work environment.
5. Agile development process usually deploys Pair Programming. In Pair programming,
two programmers work together at one work-station. One does code while the other
reviews the code as it is typed in. The two programmers switch their roles every hour
or so.
Advantages
1. Working through Pair programming produce well written compact programs which
has fewer errors as compared to programmers working alone .
2. It reduces total development time of the whole project. Customer representatives get
the idea of updated software products after each iteration. So, it is easy for him to
change any requirement if needed.
Disadvantages
1. Due to lack of formal documents, it creates confusion and important decisions taken
during different phases can be misinterpreted at any time by different team members.
2. Due to the absence of proper documentation, when the project completes and the
developers are assigned to another project, maintenance of the developed project can
become a problem.

43
CHAPTER-7
CODING

44
45
7. CODING
7.1 [Link]

7.2 Object Detection Helper Class

46
7.3 Splash Screen

47
48
CHAPTER-8
SYSTEM STUDY AND TESTING

49
8. SYSTEM STUDY AND TESTING
8.1 Feasibility Study
The feasibility of the project is analyzed in this phase and business proposal is put forth
with very general plan for the project and some cost estimates. During system analysis the
feasibility study of the proposed system is to be carried out. This is to ensure that the
proposed system is not a burden to the company. For feasibility analysis, some
understanding of the major requirements for the system is essential.
Three key considerations involved in the feasibility analysis are :
1. ECONOMICAL FEASIBILITY
2. TECHNICAL FEASIBILITY
3. SOCIAL FEASIBILITY

 ECONOMICAL FEASIBILITY
This study is carried out to check the economic impact that the system will have on the
organization The amount of fund that the company can pour into the research and
development of the system is limited. The expenditures must be justified. Thus, the
developed system as well within the budget and this was achieved because most of the
technologies used are freely available. Only the customized products had to be purchased.

 TECHNICAL FEASIBILITY
This study is carried out to check the technical feasibility, that is, the technical
requirements of the system. Any system developed must not have a high demand on the
available technical resources. This will lead to high demands on the available technical
resources. This will lead high demands being placed on the client. The developed system
must have a modest requirement, as only minimal or null changes are required for
implementing this system.

 SOCIAL FEASIBILITY
The aspect of study is to check the level of acceptance of the system by the user. This
includes the process of training[9] the user to use the system efficiently. The user must not
feel threatened by the system, instead must accept it as a necessity. The level of acceptance
by the users solely depends on the methods that are employed to educate the user about the
system and to make him familiar with it. His level of confidence must be raised so that he
is also able to make some constructive criticism, which is welcomed, as he is the final user
of the system.

 SYSTEM TESTING
The purpose of testing is to discover errors. Testing is the process of trying to discover every

50
51
conceivable fault or weakness in a work product. It provides a way to check the
functionality of components, sub-assemblies, assemblies and/or a finished product It is the
process of exercising software with the intent of ensuring that the software system meets
its requirements and user expectations and does not fail in an unacceptable manner. There
are various types of tests. Each test type addresses a specific testing requirement.

8.2 TYPES OF TESTS &

CASE UNIT TESTING

Unit testing involves the design of test cases that validate that the internal program logic is
functioning properly, and that program inputs produce valid outputs. All decision branches
and internal code flow should be validated. It is the testing of individual software units of
the application .it is done after the completion of an individual unit before integration. This
is a structural testing, that relies on knowledge of its construction and is invasive. Unit tests
perform basic tests at component level and test a specific business process, application,
and/or system configuration. Unit tests ensure that each unique path of a business process
performs accurately to the documented specifications and contains clearly defined inputs
and expected results.

INTEGRATION TESTING

Functional tests provide systematic demonstrations that functions tested are available as
specified by the business and technical requirements, system documentation, and user
manuals.
Functional testing is centred on the following items:

Valid Input : identified classes of valid input must be accepted.


Invalid Input : identified classes of invalid input must be
rejected. Functions : identified functions must be exercised.
Output : identified classes of application outputs must be exercised.

Systems/Procedures: interfacing systems or procedures must be invoked. Organization and


preparation of functional tests is focused on requirements, key functions, or special test
cases. In addition, systematic coverage pertaining to identify Business process flows; data
fields, predefined processes, and successive processes must be considered for testing.
Before functional testing is complete, additional tests are identified and the effective value
of current tests is determined.

52
SYSTEM TEST

System testing ensures that the entire integrated software system meets requirements. It
tests a configuration to ensure known and predictable results. An example of system
testing is the configuration-oriented system integration test. System testing is based on
process descriptions and flows, emphasizing pre-driven process links and integration
points.

WHITE BOX TESTING

White Box Testing is a testing in which in which the software tester has knowledge of the
inner workings, structure and language of the software, or at least its purpose. It is
purpose. It is used to test areas that cannot be reached from a black box level.

BLACK BOX TESTING

Black Box Testing is testing the software without any knowledge of the inner workings,
structure or language of the module being tested. Black box tests, as most other kinds of
tests, must be written from a definitive source document, such as specification or
requirements document, such as specification or requirements document. It is a testing in
which the software under test is treated, as a black box. you cannot “see” into it. The test
provides inputs and responds to outputs without considering how the software works.

UNIT TESTING

Unit testing is usually conducted as part of a combined code and unit test phase of the
software lifecycle, although it is not uncommon for coding and unit testing to be
conducted as two distinct phases.

Test strategy and approach

Field testing will be performed manually and functional tests will be written in detail.
Test objectives
 All field entries must work properly.
 Pages must be activated from the identified link.
 The entry screen, messages and responses must not be delayed.
Features to be tested
 Verify that the entries are of the correct format
 No duplicate entries should be allowed
 All links should take the user to the correct page.

53
INTEGRATION TESTINGSS

Software integration testing is the incremental integration testing of two or more integrated
software components on a single platform to produce failures caused by interface defects.
The task of the integration test is to check that components or software applications, e.g.,
components in a software system or – one step up – software applications at the company
level – interact without error.
Test Results: All the test cases mentioned above passed successfully. No defects
encountered.

ACCEPTANCE TESTING

User Acceptance Testing is a critical phase of any project and requires significant
participation by the end user. It also ensures that the system meets the functional
requirements.
Test Results: All the test cases mentioned above passed successfully. No defects
encountered.

8.3 TESTING CASES

54
55
56
CHAPTER-9
RESULTS

57
58
9. RESULTS

The object detection system accurately identified electronic devices such as a TV and a cell
phone with high confidence scores. The results demonstrate the model’s effectiveness in
recognizing and labeling real-world objects.

9.1 INPUT THROUGH MIC

fig-1: Splash Screen

59
9.2 OBJECT DETECTION

fig-2: Object Detection 1

fig-3: Object Detection 2

60
61
9.3 ADDING CONTACT

fig-4: Contact saving

62
9.4 SENDING MESSAGES

fig-5: Sending Messages

63
64
CHAPTER-9
CONCLUSION

65
9. CONCLUSION
In conclusion, this application signifies a transformative leap in user interaction, uniting
voice commands[5], object detection, and communication. Its seamless integration offers
unparalleled accessibility, with clear voice interpretation and visual recognition. The
inclusion of hands-free calls, messaging, and intuitive contact management illustrates a
commitment to user-centric design. This innovative fusion pioneers a versatile, visually
assisted interface, promising inclusivity and enhanced user experiences across diverse needs.
It empowers visually impaired and physically challenged users to interact with their
environment more independently.
The offline functionality ensures consistent performance even in areas with poor internet
connectivity. Lightweight design makes it suitable for low-end smartphones [4], expanding
its reach. Future enhancements could include multi-language support and AI-driven
personalization. Overall, the system lays a strong foundation for the next generation of
assistive technologies.

Future Enhancement
 Future enhancements could explore expanding the application's language
recognition capabilities, ensuring a more diverse user base. Integration with
emerging technologies, like augmented reality, could further enrich the visual
assistance aspect.

 Continuous updates may introduce additional functionalities, such as advanced voice


commands and broader compatibility, solidifying the application's position as a
cutting- edge and adaptable solution for diverse user needs.

66
67
CHAPTER-10
REFERENCES

68
10. REFERENCES
[1] Norhafizah bt Aripin; M. B. Othman ;Voice control of home appliances using Android;
27-28 August 2014.
[2 Md. Amanat Khan Shishir; Shahariar Rashid Fahim; Fairuz Maesha Habib; Tanjila
Farah; Eye Assistant : Using mobile application to help the visually impaired; 03-05 May
2019.
[3] Branislav Popović; Edvin Pakoci; Nikša Jakovljević; Goran Kočiš; Darko Pekar; Voice
assistant application for the Serbian language; 24-26 November 2015.
[4] Laviniu Ţepelea; Ioan Gavriluţ Electronics and Telecommunications Department,
University of Oradea, Oradea, Romania ; Alexandru Gacsádi; Smartphone application to
assist visually impaired people; 01-02 June 2017.
[5] Rajakumar P; K. Suresh; Boobalan M; M. Gokul; [Link] Kumar; Archana R; IoT
Based Voice Assistant using Raspberry Pi and Natural Language Processing; 08-09
December 2022.
[6] B. Popović, E. Pakoci, S. Ostrogonac, D. Pekar, “Large Vocabulary Continuous Speech
Recognition for Serbian Using the Kaldi Toolkit, in Proc. 10th Digital Speech and Image
Processing, DOGS, Novi Sad, 2014, pp. 31-34.
[7] A. Stolcke, J. Zheng, W. Wang, and V. Abrash, “SRILM at Sixteen: Update and
Outlook,” in Proc. IEEE Workshop on Automatic Speech Recognition and Understanding
ASRU, Waikoloa, 2011.
[8] R. Kneser and H. Ney, “Improved Backing-Off for M-Gram Language Modeling,” in
Proc. 20th Int. Conf. on Acoustics, Speech and Signal Processing ICASSP, Detroit, 1995,
pp. 181-184.
[9] D. Povey, D. Kanevsky, B. Kingsbury, B. Ramabhadran, G. Saon, and K.
Visweswariah, “Boosted MMI for Model and Feature-Space Discriminative Training,” in
Proc. 33rd Int. Conf. on Acoustics, Speech and Signal Processing ICASSP, Las Vegas,
2008, pp. 4057- 4060.
[10] D. Povey and P. C. Woodland, “Minimum Phone Error and I- Smoothing for
Improved Discriminative Training,” in Proc. 27th Int. Conf. on Acoustics, Speech and
Signal processing ICASSP, Orlando, 2002, pp. I105-108.
[11] S. Suzić, B. Popović, V. Delić, and D. Pekar, “Serbian Mobile Speech Database
Collection and Evaluation”, in Proc. Int. Conf. on Electronics, Telecommunications,
Automation and Informatics ETAI, ETAI 1-1, Ohrid, 2015.
[12] S. S. K. S. Swain et al. (2020), "Object Detection and Recognition in Real-Time
using Deep Learning Techniques," International Journal of Computer Applications.
[13] Y. LeCun, Y. Bengio, and G. Haffner (1998), "Gradient-Based Learning Applied to
Document Recognition," Proceedings of the ΙΕΕΕ.

69
70
[14] K. Simonyan and A. Zisserman (2014), "Very Deep Convolutional Networks for
Large- Scale Image Recognition," arXiv preprint arXiv:1409.1556.
[15] Μ. Α. Η. Β. Μ. Nasir et al. (2017), "Smart Walking Stick for Blind People,"
International Journal of Engineering and Technology.
[16] IoT Enabled Automated Object Recognition for the Visually Impaired Author
links open overlay panel Md. Atikur Rahman, Muhammad Sheikh Sadi
[17] CICERONE- A Real Time Object Detection for Visually Impaired People
Therese Yamuna Mahesh, S S Parvathy, Shibin Thomas, Shilpa Rachel Thomas and Thomas
Sebastian
[18] Blind Person Assistant: Object Detection Authors: Aniket Birambole, Pooja Bhagat,
Bhavesh Mhatre, Prof. Aarti Abhyankar
[19] OBJECT DETECTION SYSTEM FOR THE BLIND WITH VOICEGUIDANCE
Rajeshvaree Ravindra Karmarkar Walchand College of Engineering, Sangli Vikas
Honmane Walchand College of Engineering, Sangli
[20] Real-Time Object Detection for Visually Challenged People Sunit Vaidya; Naisha
Shah; Niti Shah; Radha Shankarmani
[21] Object Recognition System for the Visually Impaired: A Deep Learning Approach
using Arabic Annotation by Nada Alzahrani andHeyam H. Al-Baity
[22] Third Eye: Object Recognition and Speech Generation for Visually Impaired Author
links open overlay panel Koppala Guravaiah,Yarlagadda Sai Bhavadeesh ,Peddi Shwejan ,
Allu Harsha Vardhan , S Lavanya.
[23] Object Detection for Blind People Using Yolov3 Authors: Prof. Pranjali Deshmukh,
Ajinkya Khedkar, Shubham Kulkarni, Shriram Morkhandikar
[24] Implementation Of Object Detection and Identification For Visually Impaired
Mohammad Hussain Zaidi; Ahmad Yasir; Ajay Shanker Singh; Anandhan. K
[25] Using Object Detection Technology to Identify Defects in Clothing for Blind People
by Daniel Rocha, Leandro Pinto ,José Machado ,Filomena Soares and Vítor Carvalho

71
CHAPTER-11
PAPER PUBLICATION
WITH
PLAGIARISM REPORT

72
73
Object Recognition for Visually Impaired
B. Kiran Kumar1, G Santhoshi Rupa2, S Chandra Shekhar Raju3, M Vinay4, P Goutham Ganesh5
1
Associate Professor, Anil Neerukonda Institute of Technology and Sciences, Visakhapatnam, India
2,3,4,5
Student, Anil Neerukonda Institute of Technology and Sciences, Visakhapatnam, India

Abstract—In an era where technology plays a pivotal role


in enhancing accessibility, this innovative application leverages project addresses these challenges by offering a unified
voice-activated commands, object detection, and communication system that enhances accessibility, independence, and user
functionalities to create an inclusive and user-friendly interface. conve- nience. The primary objective is to develop an
Designed to assist users in various tasks, the application en- intelligent and adaptive interface that facilitates accurate voice
ables seamless interaction through advanced speech recognition interpretation, real-time object recognition, and effortless
technology, ensuring precise interpretation of voice commands.
By prioritizing voice clarity and natural language processing, the communication. By enabling users to make calls, send
system minimizes errors in command execution, making it messages, and manage con- tacts through voice input, the
particularly useful for individuals with visual impairments or application eliminates the need for manual interactions,
those who require hands-free operation. At the core of this making it particularly beneficial for visually impaired
system is a sophisticated object detection module that utilizes individuals and those requiring hands-free operations. The
computer vision algorithms to recognize and identify objects in
real time. This visual recognition feature empowers users by scope of this project encompasses advanced [7.],[8.] speech
providing descriptive feedback on their surroundings, thereby recognition, computer vision-based object detection, and
enhancing spatial awareness and navigation. The integration of seamless contact management. Through the integration of
object detection with voice commands fosters a comprehensive these technologies, the application introduces a novel
interactive experience, bridging the gap between visual and audi- approach to assistive solutions, enhancing user experience and
tory inputs. Beyond object recognition, the application extends
its capabilities to communication by enabling users to make demon- strating the potential of AI-driven interaction. This
phone calls and send messages using only voice prompts. This fusion of voice commands and visual assistance represents a
hands-free functionality significantly improves accessibility and significant leap towards inclusivity and intuitive user interface
convenience, particularly for individuals with physical design, ensuring that users can interact effortlessly with their
limitations. The system is further enhanced by its ability to
digital environment.
synchronize with the user’s contact list, allowing seamless
retrieval of saved contacts as well as the addition of new ones II. RELATED WORK
through voice input. By incorporating contact management via
voice-driven interactions, the application eliminates the need for The integration of voice recognition, object detection, and
manual typing, making communication more efficient and assistive technology has been widely explored in various
accessible. The fusion of voice command- driven interactions domains, particularly in enhancing accessibility and user in-
with real-time object detection introduces a novel approach to
teraction. Several research studies have contributed to this
assistive technology. Unlike conventional assistive applications
that focus solely on either voice commands or object recognition, field, each addressing different aspects of voice-assisted ap-
this system integrates both aspects into a unified framework, plications.[1] Norhafizah bt Aripin and M. B. Othman (2014)
thereby expanding its utility across diverse user scenarios. Such presented a voice-controlled home automation system using
an approach not only demonstrates the potential of AI-driven Android smartphones. Their system leveraged Google’s voice
interfaces but also sets a new benchmark for designing accessible
and intuitive digital solutions.
recognition to process voice commands, which were transmit-
ted to an Arduino Uno via Bluetooth for controlling household
I. INTRODUCTION appliances such as lights and fans. The study demonstrated the
In today’s rapidly evolving technological landscape, acces- system’s efficiency in enhancing accessibility, particularly for
sibility and seamless interaction remain at the forefront of elderly and disabled users, with an operational range of up to
user experience design. This project introduces an innovative 20 meters.[2]. Md. Amanat Khan Shishir et al. (2019)
application that integrates voice-activated commands, object introduced ”Eye Assistant,” an Android application designed
detection, and communication features to create an inclusive to support visually impaired users by enabling real-time object
and intuitive user interface. Motivated by the vision of and text recognition. The system utilized TensorFlow-based
enhanc- ing digital interactions, the application aims to bridge machine learning for background segmentation and object
the gap between voice-driven control and visual assistance, retrieval, providing auditory feedback through a text-to-
catering to users who require a more accessible and efficient speech engine. Notably, this solution functioned without a
interaction model. Existing solutions lack a comprehensive remote server, mak- ing it compatible with low-end
approach that seamlessly combines voice recognition, real- smartphones, thereby ensuring accessibility for a broader user
time object detec- tion, and hands-free communication. Users base.[6.],[11.] A number of speech and language resources for
often struggle with limited functionalities, such as the inability the Serbian language have been created during the last decade.
to add contacts via voice commands or receive real-time [3.] Branislav Popovic´ et al. (2015) developed a Serbian-
object descriptions. This language voice assistant for An- droid mobile devices,
leveraging the Kaldi speech recognition toolkit. Their system
incorporated multiple acoustic models trained with over
70,000 utterances, enhancing its ability to recognize the proposed system delivers a versatile, accessible, and user-
continuous [9.],[10.] speech across diverse environments. The friendly interface suitable for diverse needs. 3.4 Advantages
assistant demonstrated notable accuracy, even in noisy of the Proposed System Comprehensive Integration: A unified
conditions, offering an extensive vocabulary of over 14,000 platform that seamlessly combines voice commands, object
words.[4]. Laviniu T¸ epelea and Ioan Gavrilu¸t (2017) detection, and communication functionalities. Enhanced User
proposed a smartphone-based assistive system for visually Interaction: Precise voice recognition ensures clear, accurate
impaired individuals, integrating MEMS sensors and external command execution. Improved Visual Assistance: Real-time
hardware modules to facilitate indoor and outdoor navigation. object detection facilitates hands-free communication and
Com- munication between the smartphone and external intu- itive contact management, making the system more
modules was established via Bluetooth and Wi-Fi, ensuring accessible for visually impaired users. The proposed
seamless data transmission. The system also incorporated text- system pioneers an inclusive and intuitive user interface,
to-speech technology, enabling users to make calls and bridging the gaps in current assistive technologies and
interact with the environment efficiently. This cost-effective expanding its potential applications across various domains.
and portable solution required minimal training, making it
highly practical for real-world applications.[5]. Rajakumar P IV. PROPOSED SYSTEM
et al. (2022) explored an IoT-based voice assistant utilizing The proposed system integrates voice commands, object
Raspberry Pi and Natural Language Processing (NLP). Unlike detection, and communication features into a single mobile
conventional AI assistants such as Google Assistant, Siri, and application. Key features include:
Alexa, their system sup- ported offline voice recognition and
• Voice-Based Interaction: Users can perform actions
multilingual capabilities. By preloading language models onto
such as making calls, sending messages, and managing
Raspberry Pi, the as- sistant could function without an active
contacts via voice commands.
internet connection, overcoming major limitations of existing
• Object Detection: Advanced machine learning models
cloud- dependent voice assistants. The system provided users
identify and describe objects in real-time, aiding visually
with real-time updates on news, geolocation, and Wikipedia
impaired users.
searches, along with automation features for smart
• Hands-Free Communication: The application allows
environments.
These studies highlight the importance of integrating speech users to add new contacts, make calls, and send messages
recognition and object detection for enhanced accessibility. without manual input.
However, existing solutions lack a comprehensive framework • Offline Mode: Core functionalities remain operational

that combines voice commands, object recognition, and com- without internet connectivity, enhancing accessibility in
munication features in a single application. rural areas.
Advantages of the Proposed System Comprehensive In-
III. EXISTING SYSTEM tegration: A unified platform that seamlessly combines
Current assistive technologies lack a unified approach that voice commands, object detection, and communication
seamlessly integrates voice commands, object detection, and functionalities. Enhanced User Interaction: Precise voice
communication features into a single solution. Many ex- recognition ensures clear, accurate command execution.
isting systems provide isolated functionalities, leading to a Improved Visual Assistance: Real-time object detection
fragmented user experience. The absence of intuitive voice- facilitates hands-free communication and intuitive
driven contact management and real-time object detection for contact management, making the system more accessible
visual assistance restricts accessibility and usability. These for visually impaired users. The proposed system
limitations highlight the need for a more versatile and in- pioneers an inclusive and intuitive user interface,
clusive interface to enhance user interaction. 3.2 Limitations bridging the gaps in current assistive technologies and
of Existing Systems Lack of Integration: Existing solutions expanding its potential applications across various
do not provide a comprehensive platform combining voice domains.
control, object recognition, and communication. Fragmented
User Experience: The absence of seamless interoperability V. METHODOLOGY
between these features reduces usability and effectiveness. A. Data Collection The dataset comprises real-world im-
Limited Accessibility: The lack of intuitive, voice-driven ages for object detection and voice command datasets for
functionalities necessitates a more adaptive and user-friendly speech recognition. To enhance data quality, preprocessing
system. 3.3 Proposed System To address these challenges, we techniques such as noise reduction, feature extraction, and
propose an intelligent and inclusive application that integrates dataset augmentation are applied, ensuring improved accuracy
voice-activated commands, real-time object detection, and en- and robustness of the models. B. Feature Extraction Speech
hanced communication capabilities. This system allows users Recognition: Utilizes Google’s Speech-to-Text API for
to interact effortlessly using precise voice recognition, ensur- precise voice command processing. Object Detection:
ing a seamless and intuitive experience. The incorporation Implements a pre-trained YOLO model to identify and
of object detection technology enhances visual recognition, classify objects in real-time. Communication Module:
enabling hands-free operations such as calling, messaging, Seamlessly integrates with phone contacts, enabling hands-
and contact management. By merging these functionalities, free calling and mes- saging functionalities. C. Model
Training Speech-to-Text
Fig. 2. Sample-Code 1

Fig. 1. A. System Architecture

Model: Trained using Recurrent Neural Networks (RNN),


[12.] Deep learning with phoneme recognition to enhance
command interpretation. Object Detection Model: Fine-tuned
leveraging TensorFlow and OpenCV, ensuring high-accuracy
real-time object recognition.
D. System Implementation The application is developed
using Kotlin in Android Studio, incorporating TensorFlow Fig. 3. Sample-Code 2
Lite for lightweight on-device processing. This ensures
efficient mobile deployment while maintaining real-time
tated the evaluation of critical areas inaccessible to black-box
responsive- ness and accuracy.
testing, ensuring structural integrity and functional correctness.
VI. TESTING AND EVALUATION C. Black Box Testing
The system was evaluated across multiple Android devices Black Box Testing assessed the system without knowledge
based on key performance metrics: of internal implementation, focusing on input-output valida-
A. System Testing tion. Test cases were designed based on functional require-
ments and specifications, ensuring compliance with expected
System testing was conducted to validate the entire inte- behaviors.
grated software system against specified requirements. The
testing process ensured the system’s consistency, reliability, D. Unit Testing
and expected functional outcomes. Unit testing was conducted alongside code development to
ensure individual software components functioned correctly.
B. White Box Testing This phase verified code logic, handled edge cases, and
White Box Testing involved an in-depth assessment of the detected anomalies before integration.
internal structure, logic, and codebase. This approach facili-
E. Test Strategy and Approach
• Field testing was conducted manually to validate real-
world performance.
• Functional tests were designed to verify system
compo- nents and user interactions.
F. Test Objectives
The testing process aimed to ensure:
• Proper validation of all input fields.
• Correct activation of navigational links.
• Responsive user interface, ensuring minimal delays.

G. Integration Testing
Integration testing was performed incrementally to assess
in- teractions between software components. This phase Fig. 5. Object Detection - 2
validated seamless communication between modules, ensuring
stability and error-free operations at both system and
enterprise levels. • Augmented Reality Integration: Enhancing real-world
object recognition.
H. Acceptance Testing • Adaptive Learning: Implementing AI-driven personal-
User Acceptance Testing (UAT) was conducted with end- ization for improved accuracy and[14.] Large-Scale Image
users to verify the system met functional and accessibility Recognition.
requirements. By incorporating these advancements, the system aims to
Test Results: The system was successfully validated with provide an inclusive and accessible user experience for a
no defects encountered, confirming its readiness for deploy- diverse range of users via various integrations[15.] like in
ment. smart stick.

VII. OUTPUT REFERENCES

[1]. Norhafizah bt Aripin; M. B. Othman


;Voice control of home appliances using Android; 27-28
August 2014.
[2]. Md. Amanat Khan Shishir; Shahariar Rashid Fahim;
Fairuz Maesha Habib; Tanjila Farah; Eye Assistant : Using
mobile application to help the visually impaired; 03- 05 May
2019.
[3]. Branislav Popovic´; Edvin Pakoci; Niksˇa Jakovljevic´;
Goran Kocˇisˇ; Darko Pekar; Voice assistant appli- cation for
the Serbian language; 24-26 November 2015.
[4]. Laviniu T¸ epelea; Ioan Gavrilu¸t Electronics and
Telecommuni- cations Department, University of Oradea,
Oradea, Romania ; Alexandru Gacsa´di; Smartphone
application to assist visually impaired people; 01-02 June
2017. [5]. Rajakumar P; K. Suresh; Boobalan M; M. Gokul;
[Link] Kumar; Archana R; IoT Based Voice Assistant using
Fig. 4. Object Detection - 1
Raspberry Pi and Natural Language Processing; 08-09
December 2022. [6]. B. Popovic´,E. Pakoci, S. Ostrogonac,
D. Pekar, “Large Vocabulary Contin- uous Speech
VIII. CONCLUSION AND FUTURE WORK Recognition for Serbian Using the Kaldi Toolkit, in Proc.
This study presents a novel assistive application that in- 10th Digital Speech and Image Processing, DOGS, Novi
tegrates voice commands, object detection, and hands-free Sad, 2014, pp. 31-34.
communication for visually and hearing-impaired users. The [7]. A. Stolcke, J. Zheng, W. Wang, and V. Abrash, “SRILM
proposed system demonstrates improved accessibility and us- at Sixteen: Update and Outlook,” in Proc. IEEE Workshop
ability, outperforming traditional assistive technologies. on Automatic Speech Recognition and Understanding ASRU,
Future enhancements include: Waikoloa, 2011.
[8]. R. Kneser and H. Ney, “Improved Backing-Off for M-
• Multilingual Support:[13.] Expanding language
Gram Language Modeling,” in Proc. 20th Int. Conf. on
recognition capabilities with help of gradients.
Acoustics, Speech and Signal Processing ICASSP, Detroit,
1995, pp. 181-184.
[9]. D. Povey, D. Kanevsky, B. Kingsbury, B.
Ramabhadran,
G. Saon, and K. Visweswariah, “Boosted MMI
for Model and Feature-Space Discriminative
Training,” in Proc. 33rd Int. Conf. on Acoustics,
Speech and Signal Processing ICASSP.
[10]. D. Povey and P. C. Woodland, “Minimum Phone Error and I-
Smoothing for Improved Discriminative Training,” in Proc. 27 th Int.
Conf. on Acoustics, Speech and Signal Processing ICASSP, Orlando,
2002, pp. I105-108.
[11]. S. Suzić, B. Popović, V. Delić, and D. Pekar, “Serbian Mobile
Speech Database Collection and Evaluation”, in Proc. Int. Conf. on
Electronics, Telecommunications, Automation and Informatics ETAI,
ETAI 1-1, Ohrid, 2015
[12]. S. S. K. S. Swain et al. (2020), "Object Detection and Recognition
in Real-Time using Deep Learning Techniques," International Journal of
Computer Applications.
[13]. Y. LeCun, Y. Bengio, and G. Haffner (1998), "Gradient-Based
Learning Applied to Document Recognition," Proceedings of the ΙΕΕΕ.
[14]. K. Simonyan and A. Zisserman (2014), "Very Deep Convolutional
Networks for Large-Scale Image Recognition," arXiv preprint
arXiv:1409.1556.
[15]. Μ. Α. Η. Β. Μ. Nasir et al. (2017), "Smart Walking Stick for Blind
People," International Journal of Engineering and Technology.

You might also like