Augmented Reality Language Learning Project
Augmented Reality Language Learning Project
INSTITUTE OF ENGINEERING
ADVANCED COLLEGE OF ENGINEERING AND MANAGEMENT
DEPARTMENT OF ELECTRONICS AND COMPUTER ENGINEERING
KALANKI, KATHMANDU
[CT654]
A Minor Project Report On
“AUGMENTED REALITY LANGUAGE LEARNING”
Submitted By:
Aditya Joshi (ACE078BCT005)
Anup Joshi (ACE078BCT011)
Asim Basnet (ACE078BCT017)
Bikrant Budhathoki (ACE078BCT023)
Project Supervisor:
Er. Laxmi Prasad Bhatt
Er. Ukesh Thapa
A Major Project Final report submitted to the department of Electronics and Computer
Engineering in the partial fulfillment of the requirements for degree of Bachelor of
Engineering in Computer Engineering
Kathmandu, Nepal
7 March, 2025
AUGMENTED REALITY LANGUAGE LEARNING
Submitted By:
Aditya Joshi (ACE078BCT005)
Anup Joshi (ACE078BCT011)
Asim Basnet (ACE078BCT017)
Bikrant Budhathoki (ACE078BCT023)
Supervised by:
Er. Laxmi Prasad Bhatt
Er. Ukesh Thapa
Submitted to:
"DEPARTMENT OF ELECTRONICS AND COMPUTER ENGINEERING"
ADVANCED COLLEGE OF ENGINEERING AND MANAGEMENT
Balkhu, Kathmandu
7 March, 2025
II
LETTER OF APPROVAL
The undersigned certify that they have read and recommended to the Institute of Engineering
for acceptance, a project report entitled “ AUGMENTED REALITY LANGUAGE
LEARNING ” submitted by:
In the partial fulfillment of the requirements for the degree of Bachelor’s Degree in Computer
Engineering.
……………………….. ………………………..
Project Supervisor Project Supervisor
Academic Project Coordinator Er. Ukesh Thapa
Er. Laxmi Prasad Bhatt Department of Electronics and Computer
Department of Electronics and Computer Engineering
Engineering
………………………..
Academic Project Coordinator
Er. Laxmi Prasad Bhatt
Department of Electronics and Computer Engineering
7 March, 2025
III
COPYRIGHT
The author has agreed that the library, Advanced College of Engineering and Management,
may make this report freely available for inspection. Moreover, the author has agreed that
permission for extensive copying of this project report for scholarly purposes may be granted
by the supervisors who supervised the project work recorded herein or, in their absence, by
the Head of the Department wherein the project was done. It is understood that recognition
will be given to the report's author and the Department of Electronics and Computer
Engineering, Advanced College of Engineering and Management for any use of the material
of this project report. Copying publication or the other use of this project for financial gain
without the approval of the Department an author’s written permission is prohibited.
Request for permission to copy or to make any other use of the material in this report in
whole or in should be addressed to:
Head of Department
Department of Electronics and Computer Engineering
Advanced College of Engineering and Management
Balkhu, Kathmandu
Nepal
IV
ACKNOWLEDGMENT
We would like to express our profound gratitude and deep regards to our respected supervisor
Er. Laxmi Prasad Bhatt, for his insightful advice, motivating suggestions, invaluable
guidance, help, and support in the successful completion of this project and also for his
constant encouragement and advice throughout our Bachelor's program.
We express our deep gratitude to Er. Prem Chandra Roy, Head of the Department of
Electronics and Computer Engineering, Er. Dhiraj Pyakurel Deputy Head, Department Of
Electronics and Computer Engineering, Er. Laxmi Prasad Bhatt, Academic Project
Coordinator, Department Of Electronics and Computer Engineering, for their support,
co-operation, and coordination.
The in-time facilities provided by the department throughout the Bachelor's program are also
equally acknowledgeable.
We would also like to express our sincere gratitude to Er. Ukesh Thapa for his valuable
advice and guidance throughout the project.
We would like to convey our thanks to the teaching and non-teaching staff of the Department
of Electronics & Communication and Computer Engineering, ACEM for their invaluable
help and support throughout Bachelor’s Degree. We are also grateful to all our classmates for
their help, encouragement, and invaluable suggestions.
Finally, yet more importantly, We would like to express our deep appreciation to our
grandparents, parents, and siblings for their perpetual support and encouragement
throughout the Bachelor’s degree period.
Project Members:
Aditya Joshi (ACE078BCT005)
Anup Joshi (ACE078BCT011)
Asim Basnet (ACE078BCT017)
Bikrant Budhathoki (ACE078BCT023)
V
ABSTRACT
VI
TABLE OF CONTENTS
VII
6.2 Result Analysis 19
6.3 User Interface 21
CHAPTER 7: CONCLUSION, LIMITATIONS AND FUTURE ENHANCEMENT 22
7.1 Conclusion 22
7.2 Limitations 22
7.3 Future Enhancement 22
REFERENCES 23
APPENDIX A: ADDITIONAL TOOLS 24
VIII
LIST OF TABLE
X
LIST OF FIGURES
X
LIST OF ABBREVIATIONS
AI – Artificial Intelligence
AR – Augmented Reality
VR – Virtual Reality
XI
CHAPTER 1: INTRODUCTION
Learning a new language is a handy task from the beginning, however, in our fast-paced,
globalized world, it has become even more important. The unique charm of language is lost
in traditional education through rote-learning of vocabulary and grammar patterns. Such
approaches might get mundane and boring and the learning becomes more of a chore than an
exciting journey. The solution is to incorporate AR (Augmented Reality) in language
learning. This has the potential to effectively change learning to an engaging, interactive and
immersive experience through AR. Now, combine this with AR and object detection and
maybe real-time translation and make every aspect of daily life a learning experience. To
illustrate, point the camera of your smartphone at something like a product name or a nature
object and it would display what it was called in multiple languages. This kind of task
transforms everyday interactions into captivating, informative challenges and puts a virtual
language instructor right inside the learner's pocket. The project intends to implement this
through an android app that will provide AR applications to improve language learning. The
app will allow a user to scan objects with their camera in an AR space during its first phase of
the rollout. Fast recognition of objects translated into several languages. The idea itself is
simple, but the impact it has on how humans acquire a new language and bridge the gap
between nations is enormous.
1.1 Background
In our world full of connections, having the knowledge of more than one language is among
the must-have skills. Learning a language can make way for many things, be it to travel or to
work, or just to get to know new people. Traditional ways to learn a language can be
incredibly dull. They tend to be non-interactive and require an infinite amount of
memorization which can sometimes make it difficult for one to stay engaged with them. AR
has the ability to provide real-life experiences and to make learning much more immersive.
Point your camera at an object and learn its name in several languages. It's easy, fun and
much more enjoyable than swiping flashcards or grinding away at yet another quiz app. An
android that has been built from inception to product laced with AR, image processing, object
detection and AR based real-time translation of text and voice. Scan an object using the
application, it will be recognized with AR, and then it will show its name in different
languages. It will be like carrying around a localized version of your very own language tutor
1
in your pocket. We want language learning to become something that everyone enjoys doing
and we believe that the best way to achieve that is by turning the very objects you see around
you into an opportunity to learn. So, it is just about combining tech choice with creative
process for making it easier, quicker and interesting.
1.2 Motivation
The absence of context and interaction is what makes language learning so boring and
difficult. Visualizing objects in real time, with the translations overlaid directly over them,
could create a more interactive and natural learning experience. The following are the reasons
for this project: Learning tool that is a true use case for the real-world and not limited to
books and flashcards. There are no long-term, everyday applications of AR object detection
for language. We were also motivated to experiment with the latest in ARCore and AI
technology in real-world scenarios.
People who are learning a language face several big challenges that make it a less effective
and enjoyable process. One of the biggest problems is that there are no contextual learning
tools that help connect words to objects. Most applications today are only focused either on
language translation or AR experience and not both at the same time and in a proper way.
Many of these solutions are also designed with complicated interfaces or require technical
knowledge, making them difficult for the average user to get used to. Learners often struggle
to find tools that are easy to use and can provide interactivity, immersiveness and blend
manual immersive learning with digital immersive learning. Due to this, learners often give
up their new language learning process and many of them only speak a few handful of
languages which creates a language barrier.
2
1.4 Project Objectives
General Objectives:
● To develop an Android-based AR app that detects an object and translates its name
into other languages in real time.
● To provide a brief description of the detected object to enhance learning and
engagement.
Specific Objective:
● To provide real-time object detection using the Tiny YOLOv2 CNN model.
This project focuses on developing an Android app that allows users to point their camera at
objects, identify them using AR, and display the object's name in different languages. Using
AR technology, the language-learning app blends interactive and conventional learning
methods. The application's usage of augmented technologies will enable users to scan
real-world objects and receive translations quickly, aiding with context learning and memory.
For all kinds of learners in the modern world, our effort will contribute to making language
acquisition more efficient and easy. Additionally, the usage of AR in educational tools is
highlighted, which paves the way for future advancements in interactive application
development and language pedagogy.
3
CHAPTER 2: LITERATURE REVIEW
Gómez et al. in 2023 state that AR (AR) is bridging the real and the virtual world to create
engaging learning environments. AR tools, such as marker-based, markerless, and
geolocation-based applications, are utilized at all levels of education, improving motivation,
engagement, and comprehension of abstract concepts through the use of AR books, games,
and note-taking systems. AR provides benefits like higher retention, emotional connection,
and autonomous learning, while also highlighting challenges such as technological
complexity, lack of training, and potential distractions. It holds that, though AR has great
potential for transforming education, its effective implementation demands overcoming these
barriers via proper training and resources [1].
Gómez et al. in 2023 explore The Effect of AR (AR) on Language Acquisition, examining
and evaluating its benefits and problems. A comparative analysis of several studies published
in WOS and SCOPUS reveals that, in general, AR-based language learning projects improve
student performance and attitudes, although limitations were noted. As suggested by the
investigations, it is not the learning process itself that augments the effect of AR; rather, it is
the motivation and positive attitudes toward learning a language that are enhanced more
significantly by AR than by conventional methods. Improvements brought about by
well-accomplished studies include grammar and reading comprehension, reduced anxiety
about language acquisition, and other areas. However, limitations, such as the small scope of
research in this field, should be taken into consideration [2].
Mekni et al. in 2024 outline several future avenues for the development of AR systems,
focusing on advancements in Head Mounted Displays (HMDs), wearables, and addressing
system limitations. Present-day HMDs are bulky and suffer from limited fields of view,
resolution, and contrast, meaning they require significant input improvements to become
more user-friendly. Wearable devices, such as data gloves and suits, will also need to become
lighter, smaller, and more efficient. Key research challenges for AR systems include time
delays between inputs and percepts, hardware/software failures, and registration inaccuracies
all of which need to be addressed for future progress [3].
4
In this study, Alsing (2018) discusses the development of a Post-it note recognition system
that adopts the concept of CNNs and can run directly on mobile devices. These systems are
combined with multiple object detection frameworks to produce real-time models capable of
detecting Post-it notes. The transfer learning achieved by training a CNN on a small dataset
results in high performance, yielding impressive mAP scores. The models are sufficiently
feasible for detection on low-resource devices with rapid reliability. Future improvements
would ideally include increasing the training dataset to enhance performance in diverse
environments and considering Capsule Networks to further improve the understanding of
part-whole relationships and viewpoint variations, thereby boosting accuracy and recall on a
broader scale [4].
In this paper Bell et al. (2016) present the Inside-Outside Net (ION), an object detector that
exploits information both inside and outside the region of interest. ION uses external
contextual features through spatial RNNs and applies skip pooling within the region of
interest to learn multi-scale features and hierarchical representation. Extensive
experimentation showed that ION boosts object detection performance immensely, with an
mAP of 76.4% on the PASCAL VOC 2012 dataset compared to the previous 73.9%. For the
MS COCO dataset, ION improved mAP from 19.7% to 33.1% with a focus on detecting
small objects. This was further validated in the 2015 MS COCO Detection Challenge where
it was awarded Best Student Entry and third overall, confirming the significance of
contextual information and multi-scale representations for improved object detection [5].
Redmon et al. (2017) introduced YOLO9000, a real-time object detection system capable of
detecting over 9000 object categories. The authors proposed YOLOv2-an upgrade upon
YOLO and was benchmarked against state-of-the-art methods on standard detection tasks
contrary to PASCAL VOC and COCO. YOLOv2 allows multi-scale training, hence a wider
flexibility in speed-accuracy trade-off. It attains 76.8% mAP at 67 FPS on the VOC 2007 and
78.6% mAP at 40 FPS, thus showing a faster frame-rate than others-Faster RCNN with
ResNet and SSD case-in. YOLO9000 was suggested as a new model for the joint training of
the detection and classification tasks. As a result, the model combines the COCO detection
dataset and ImageNet classification dataset to make predictions on unseen classes [6].
5
2.1 Review Summary
Table 2.1: Literature Review Summary
2015 S. Bell et al. Inside Outside Net: arXiv preprint Proposes ION, an object
Detecting Objects in detector using context
Context and RNNs.
6
CHAPTER 3: REQUIREMENT ANALYSIS
This chapter provides a detailed overview of the project requirements, including hardware
and software specifications, as well as functional and non-functional requirements.
Additionally, it includes a feasibility study for the development of the AR language
translation and learning app.
3.1.1 Unity
Unity is a real-time 3D development engine used for games, simulations, and AR/VR
applications. It provides a user-friendly interface, a large asset store, and supports multiple
programming languages (mainly C#). Unity allows developers to build immersive augmented
reality (AR) experiences using frameworks like AR Foundation, Vuforia, and ARCore.
7
3.1.2 AR Foundation
● Motion Tracking – Tracks the phone’s position and orientation in real-time without
requiring external sensors.
● Image & Object Tracking – Detects and tracks 2D images or 3D objects in real-time,
allowing for interactive AR experiences.
3.2 Functional Requirements
8
3.3 Non-Functional Requirements
9
CHAPTER 4: METHODOLOGY
The proposed project aims to develop an application that can identify an object in real world
through a device (android device) camera and will be able to translate the object’s name in
different languages.
10
4.1.2 Data Preprocessing
Data preprocessing ensures that the raw data is ready for model training. For real-time object
detection, images and video frames are processed and converted into the appropriate format,
so they can be fed to the Tiny YOLOv2 model. This step ensures the data is clean,
standardized, and ready for machine learning. Images and video frames are resized and
normalized according to the specifications required by the Tiny YOLOv2 model.
11
4.1.6 Deployment
Deployment involves integrating the trained models into the Android app, using Barracuda
for efficient performance. The fine-tuned Tiny YOLOv2 model handles real-time object
detection, identifying objects like a person, bottle, cat, cow, dog, chair, motorbike, etc., from
live video feeds. The app utilizes the Lingva Translate API for online translations, providing
support for multiple languages. ARCore enables real-time object name overlays on the
detected objects, ensuring a seamless user experience in augmented reality environments.
4.2 YOLO
As our project is based on real-time object detection we are using YOLO architecture. YOLO
is a real time object detection algorithm which is a single stage object detector that uses
Convolution Neural Network to predict the bounding boxes and class probabilities of objects
in the input image. The YOLO algorithm divides the input image into a grid of cells. Each
cell predicts an object class as well as the corresponding bounding box coordinates if it
contains an object. The overall processing of the image is fast and efficient since YOLO
processes the entire image in a single run. It has been introduced into various versions, and
we are using Tiny YOLOv2 for this project.
The YOLOv2 object detector uses a single stage object detection network. YOLO v2 is faster
than two-stage deep learning object detectors, such as regions with convolutional neural
networks (Faster R-CNNs). The YOLO v2 model runs a deep learning CNN on an input
image to produce network predictions.
Tiny YOLOv2 is a lightweight and fast version of YOLOv2, designed for real-time object
detection on low-power devices. It reduces computational complexity while maintaining
reasonable accuracy, making it suitable for embedded systems, mobile devices, and real-time
applications.
12
The benefits of Tiny YOLOv2 include speedy inference, making it well-suited for real-time
applications, reduced computational cost, allowing it to run efficiently on low-power devices,
and a lightweight model, which requires less storage and memory. However, Tiny YOLOv2
also has some shortcomings, such as lower accuracy compared to full YOLO versions, the
absence of a passthrough layer, which limits its ability to detect smaller objects effectively,
and less robustness, making it less reliable in detecting overlapping or complex objects.
The AR Language Learning App is being built following the Agile methodology due to its
flexibility, enabling adaptation to changes that occur on complicated components like object
detection , translations, and AR overlays. Incremental delivery by iterative development
allows for functional modules to become available gradually; this thus reduces risks while
ensuring continuous feedback on the user's experience. Agile allows cooperation between
developers, designers, and testers towards supplementing features such as the Lingva
Translate API. Agile is a user-centered, effective, and scalable methodology for developing
such innovative applications.
13
CHAPTER 5: SYSTEM DESIGN AND ARCHITECTURE
14
5.2 Use Case Diagram:
15
5.3 Flowchart:
16
5.4 Data Flow Diagram (DFD):
17
CHAPTER 6: RESULTS AND ANALYSIS
Based on the our model, we have obtained different the following output:
18
6.2 Result Analysis
We obtained various metrics which are presented below:
19
6.2.3 Precision confidence curve
20
6.3 User Interface
We have focused on creating a user-friendly interface for both learners and educators,
incorporating features such as real-time object detection and language translation, while
keeping the design simple and intuitive within the given time constraints.
21
CHAPTER 7: CONCLUSION, LIMITATIONS AND FUTURE ENHANCEMENT
7.1 Conclusion
This project will be able to detect objects in real time and use the Lingva API for the
language translation. The AR app developed successfully integrates object detection and
language translation capabilities, enhancing user interaction through augmented reality. Key
achievements include:
7.2 Limitations
● Only compatible with android devices.
● Cannot detect moving objects.
● Cannot detect objects in a poor lighting environment.
● Can only handle 4-5 object detection models for now.
● Simple user interface.
22
REFERENCES
[3] Mekni, M., & Lemieux, A. (2014). “AR: Applications, challenges and future trends,”
Applied Computational Science, 20, 205-214.
[4] Alsing, T. (2018). Real-time post-it note detection using convolutional neural networks
for mobile applications. Proceedings of the International Conference on Mobile Vision
Systems, 85-92.
[5] S. Bell, C. L. Zitnick, K. Bala, and R. Girshick. Inside Outside net: Detecting objects in
context with skip pooling and recurrent neural networks. arXiv preprint arXiv:1512.04143,
2015.
[6] Redmon, J., & Farhadi, A. (2017). YOLO9000: Better, Faster, Stronger. arXiv preprint
arXiv:1612.08242.
23
APPENDIX A: ADDITIONAL TOOLS
CNN
CNNs are an advanced form of Artificial Neural Networks (ANNs), designed specifically to
process grid-like data structures, such as images. By utilizing convolutional layers, CNNs
automatically extract important features like edges, textures, and patterns from raw image
data, allowing the network to learn and recognize complex visual patterns. This makes CNNs
particularly effective for tasks such as image classification, object detection, and facial
recognition, where data patterns and spatial relationships play a crucial role in understanding
visual content.
PASCAL VOC
PASCAL VOC is a format for annotation in object detection for widely known computer
vision tasks. It consists of annotations for images, object classes, and bounding boxes. The
VOC format finds its application in object detection, segmentation, and classification. It is
primarily used to evaluate and benchmark object detection models. This dataset involves
pixel-wise segmentation, bounding box annotations, and object class labels for different
objects across various environments.
Barracuda
Barracuda is the machine learning inference engine for Unity which supports running
machine learning models such as ONNX models in Unity applications. It allows you to
perform efficient and high-performance inference for real-time object detection in Unity
environments. This project used Barracuda to load and execute the model for object detection
tasks trained with Tiny YOLOv2 in the Unity platform.
24