0% found this document useful (0 votes)
6 views8 pages

Complete File Raman

The document provides an overview of various machine learning algorithms and their applications in computer vision, highlighting recent advancements such as Vision Transformers and YOLOv8. It discusses the methodology, findings, and contributions of specific research articles, emphasizing the importance of model efficiency and self-supervised learning. Additionally, it explores real-world applications, particularly in healthcare, showcasing how machine learning can enhance patient care and address complex visual tasks.

Uploaded by

vanshgangwar1211
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views8 pages

Complete File Raman

The document provides an overview of various machine learning algorithms and their applications in computer vision, highlighting recent advancements such as Vision Transformers and YOLOv8. It discusses the methodology, findings, and contributions of specific research articles, emphasizing the importance of model efficiency and self-supervised learning. Additionally, it explores real-world applications, particularly in healthcare, showcasing how machine learning can enhance patient care and address complex visual tasks.

Uploaded by

vanshgangwar1211
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Full Name & Indexing Metrics (IF / SJR etc) Homepage Link

Publisher

IEEE Transactions on Scopus: Yes Web of Impact Factor: 6.5 [Link]


Software Engineering Science (SCIE): Yes SJR: 1.447 h-index: .org/tsse
Publisher: IEEE 193
Computer Society

Journal of Systems Scopus: Yes Web of Impact Factor: 4.1 [Link]


and Software Science (SCIE): Yes SJR: 0.975 h-index: [Link]/journal/journ
Publisher: Elsevier 135 al-of-systems-and-
software

Software: Practice Scopus: Yes Web of Impact Factor: 2.0 [Link]


and Experience Science (SCIE): Yes SJR: 0.675 h-index: [Link]/journal/1097
Publisher: John Wiley 79 024x
& Sons Ltd.

CM Transactions on Scopus: Yes Web of Impact Factor: 2.7 [Link]


Computer Systems Science (SCIE): Yes SJR: 0.583 h-index: nal/tocs
Publisher: 76
Association for
Computing Machinery
(ACM)

International Journal Web of Science Impact Factor: 1.5 [Link]


of Computers and (ESCI): Yes SJR: 0.413 h-index: [Link]/journals/tjca2
Applications 26 0
Publisher: Taylor &
Francis Ltd.

Part B – Research Article Analysis


Topic: Unleashing the Potential of Machine Learning – An Exploration of State-of-the-
Art Algorithms and Real-World Applications in Computer Vision

Title: Vision Transformers for Image Recognition: A Survey


Authors: Ashish Vaswani et al. (Extended by several authors, 2023)
Journal: IEEE Access
Year: 2023

Research Objectives:
This paper studies the rise of Vision Transformers (ViTs) as a new deep-learning model
for image recognition. The main goal is to explain how ViTs differ from CNNs
(Convolutional Neural Networks) and why they perform better for many vision tasks.
Methodology:
The authors review and compare more than 80 recent transformer-based models. They
analyze their architecture, training methods, datasets used (ImageNet, COCO, CIFAR),
and performance.

Major Findings:
Vision Transformers often achieve higher accuracy and flexibility than traditional CNNs,
especially when trained with large datasets. However, they require more computation
and memory.

Contribution to the Field:


The paper helps researchers understand how transformers are changing modern
computer vision. It also guides new work on making ViTs faster and lighter for real-world
devices.

Article 2

Title: YOLOv8: Real-Time Object Detection with Improved Accuracy and Speed
Authors: Glenn Jocher et al.
Journal: Journal of Machine Learning Research (preprint on arXiv 2024)
Year: 2024

Research Objectives:
This work introduces an improved version of the YOLO (You Only Look Once) family
for real-time object detection. The goal is to create a model that detects multiple objects
quickly and accurately on normal hardware.

Methodology:
The authors designed a new architecture with lightweight convolution layers, anchor-
free detection, and better loss functions. They tested the model on the COCO dataset
and compared it with YOLOv5, YOLOv7, and Faster-RCNN.

Major Findings:
YOLOv8 achieved faster inference speed (over 60 FPS on a GPU) and higher mAP
(Mean Average Precision > 53%) than previous versions. It also performs well for small-
object detection, which was a weakness in older models.

Contribution to the Field:


This paper shows how modern ML models balance accuracy and speed. YOLOv8 is
now widely used in real-world systems like traffic monitoring, factory automation, and
drone vision.

Article 3
Title: Self-Supervised Learning for Medical Image Analysis: A Comprehensive Review
Authors: J. Chen, S. Gao, and Y. Zhou
Journal: Information Sciences
Year: 2025

Research Objectives:
The study explores how self-supervised learning (SSL) can help train medical image
models when labeled data is limited. The aim is to make machine learning useful in real
hospital and diagnostic scenarios.

Methodology:
The paper reviews various SSL techniques such as contrastive learning, masked auto-
encoders, and generative pre-training. It compares their performance on MRI, CT, and
X-ray datasets.

Major Findings:
SSL models achieve accuracy close to supervised models but require far fewer labeled
samples. They are effective in detecting tumors and lung diseases.

Contribution to the Field:


This review proves that SSL is a powerful tool for medical image analysis, reducing
cost and time for dataset labeling. It opens a new path for applying computer-vision
models in healthcare

Part C: Research Trends

1. Vision Transformers and Model Efficiency

One major research trend is the rise of Vision Transformers (ViTs). These models
perform very well in image recognition and object detection tasks compared to older
CNN-based models. Researchers are now focusing on making these models smaller,
faster, and more energy efficient. This is important so they can run on devices like
mobile phones, drones, and real-time systems. Earlier, most studies only focused on
improving accuracy, but now there is a balance between accuracy, speed, and power
usage. This shift helps bring computer vision models closer to real-world use.

2. Self-Supervised and Multimodal Learning

Another growing trend is self-supervised learning (SSL) and multimodal learning.


Self-supervised learning helps models learn from data without needing human-labeled
examples, saving time and cost. Multimodal learning allows a single model to work with
different types of data, such as text, images, and audio. For example, models like CLIP
can understand both pictures and written text together. Research in this area is moving
toward creating one general model that can perform many tasks efficiently. At the same
time, scientists are paying more attention to making AI models fair, explainable, and
privacy-friendly.

3. Generative Models and Real-World Applications

The third important trend is the use of generative models and their role in real-time
applications. Generative models, like diffusion models, can create new and realistic
images from text descriptions. Real-time models, such as YOLO, are used for detecting
objects in videos, cameras, and self-driving cars. These technologies are now widely
used in fields like healthcare, robotics, and security. However, as they become more
powerful, researchers are also focusing on solving ethical issues such as fake content,
copyright, and responsible use of AI-generated media.

Overall Trends

Computer vision research is changing quickly. The focus is moving from supervised
learning (using labeled data) to self-supervised learning (using unlabeled data).
Researchers now care more about model speed, energy saving, and ethical use.
Generative and multimodal models are becoming more common and helpful in real-life
work.
GEN740 RESEARCH AND PUBLICATION ETHICS

SCHOOL OF COMPUTER APPLICATION

LOVELY PROFESSIONAL UNIVERSITY

GEN740

Submitted By Submitted To:


Raman Goyal [Link] Morya
(42500144 )
Unleashing the Potential of Machine Learning: An
Exploration of State-of-the-Art Algorithms and Real-
World Applications in Computer Vision
Raman Goyal Sonia Morya
Department of Computer Application Department of Food and Technology
Lovely Professional University Lovely Professional University
Jalandhar, India Jalandhar, India
ramangoyal87021@[Link] sonia.25123@[Link]

Abstract—This research paper explores the transformative


recognition, image segmentation, face detection, and scene
impact of machine learning techniques in the field of computer understanding, among others [2].
vision. With the rapid evolution of machine learning algorithms,
computer vision has witnessed remarkable advancements and The paper will delve into the fundamental concepts and
diverse applications. methodologies of machine learning, offering insights into the
underlying algorithms and techniques employed. Additionally,
study will examine popular machine learning algorithms used in
This study dives into the intersection of machine learning and computer vision, and outline their strengths and limitations.
computer vision, presenting a comprehensive analysis of state-of-
the-art algorithms and their real-world use cases. Notable machine
learning algorithms are shown in the context of their applications in
This research paper will also explore real-world applications
computer vision. of machine learning in computer vision. It will delve into a real-
life use case that involves object recognition and classification,
The paper highlights one real-world example in particular: the as well as pose estimation. By highlighting these applications,
use of machine learning techniques in long-term care for elderly the paper aims to showcase the impact of machine learning in
patients using object recognition and pose estimation. Experimental addressing complex visual tasks, with applications to long-term
results and shed light on algorithm performance and inherent hospital care of elderly patients.
limitations. The findings provide valuable insights into the untapped
potential of machine learning in computer vision. II. MACHINE LEARNING ALGORITHMS
Keywords—machine learning, algorithms, computer vision This section will delve into four commonly used methods
used in machine learning that have proven to be highly effective
in a variety of applications. By exploring these methods, we gain
I. INTRODUCTION insight into the diverse techniques employed in machine learning
In recent years, the use of machine learning has transformed and their practical implementation.
numerous industries and paved the way for groundbreaking
innovations. Machine learning enables computer systems to learn A. Linear Regression
from data and make predictions or decisions without being Linear regression is one of the most commonly understood
programmed to do so [1]. It provides valuable insights that have techniques used for machine learning. A very simple linear
positioned them as powerful tools for data-driven decision- regression model consists of one input variable (x), and one
making and automation. output variable (y) [4]. By plotting each data point on a set of
axes, a “line of best fit” may be drawn through the data (see Fig.
The objective of this research paper is to provide a 1). Subsequently, one may predict the value of the output variable
comprehensive exploration of machine learning and its (y) for any given input (x).
applications, with a particular focus on computer vision.
Computer vision, a field concerned with enabling computers to
One particular application of this approach that is more
commonly used is least squares regression. Minimizing the sum
of the squared residuals (distance from the line of best fit) allows
the computer to determine the best internal parameters for the
computer to classify future data.

In practice, linear regression usually contains a combination


of input variables, which allow for greater complexity in the
fitting process for the line of best fit. There are also numerous
more optimization methods to improve internal parameters,
however the basis of linear regression remains the same.
Fig. 1: Simplified linear regression model [5]

extract information and understand visual data, has witnessed


remarkable progress due to the integration of machine learning
techniques. The combination of advanced algorithms and large-
scale data availability has fueled breakthroughs in object
B. QDA abstract than the other three methods outlined above. It involves
Quadratic discriminant analysis follows a very similar providing data to the computer with no explicit labels or
premise to linear regression, with the main difference being the information about the expected output, which is characterized by
relationship of the input to the output variable. While linear the lack of a training set [10]. The computer then performs
regression assumes a linear relationship between input and output complex processing tasks to uncover hidden patterns about the
variables, QDA assumes a quadratic relationship [6]. More data, and perform its own categorization.
specifically, it assumes a Gaussian distribution of the data and
formulates a curve to fit the data as best as possible.

Fig. 4: Unsupervised Learning Visual Aid [11]


One notable drawback of unsupervised learning is that it is
much more computationally intensive compared to supervised
learning techniques, and usually requires more time to uncover
the patterns hidden in the data.
Fig. 2: QDA model [7]
III. EXPERIMENTAL RESULTS AND ANALYSIS
As a result, the “line of best fit” starts assuming a “curve of
best fit”. One advantage of this approach over linear regression The following section discusses a real-life application of
is its precision. For data that do not possess any sort of natural machine learning in computer vision through the development of
association, or contain a lot of outliers, a curve of best fit would an application to assist hospital caregivers to turn immobile
be more appropriate to classify it [6]. One downside of it is its patients. The application was developed in Python using a
computational complexity compared to linear regression. Raspberry Pi and a mounted camera setup. The use of
TensorFlow, OpenCV and Google MediaPipe Pose (all open-
source Python libraries) were used for the machine learning
C. kNN procedures and integration with computer vision.
The k-Nearest Neighbors algorithm is a technique that uses
proximity to determine the classification of a new data point [8].
Fig. shows an example of how a person’s posture could be
Selecting the k nearest neighbors (where k is some constant) and
tracked using a set of coordinates on the person’s notable
seeing where the larger proportion of neighbors are classified
extremities. Once all of the desired points are shown on the
allows the computer to find the appropriate classification for that
screen, MediaPipe Pose will constantly track the person’s
new data point.
position and any changes in it.

The main use for tracking the coordinates above was to


facilitate proper turning of immobile patients by hospital
caregivers. Their reason for doing so was because negligence in
conducting this process may lead to stress ulcers, sores, or even
death for these patients. As a result, computer-assisted
procedures could lower the chance of this happening by a
significant margin. Fig. 6 shows an example of the proper
position and steps a caregiver should follow when turning a
patient.

The researchers first provided the machine learning model with


Fig. 3: Simplified kNN model [9] training data, which consisted of one of the researchers posing in
various ways, and telling the computer whether or not the pose was
More specifically, the distance to each neighbor from the acceptable or not for that specific phase of the turning process. By
unclassified data point is stored in a list, which is subsequently repeating this process multiple times, as well as using the machine
sorted in ascending order. Some variations of the algorithm learning algorithms outlined above to process the data, the
may also assign additional weight to each category based on the computer was able to learn and create a classifier for newposes
relative proximity of each data point to another.

D. Unsupervised Learning
Training a “custom model” is also a very widely used
technique within machine learning. However, it is notably more
If any of the three required steps were done incorrectly, the 7. “9.2.8 - Quadratic Discriminant Analysis (QDA) | STAT 508,”
red light in the top-right corner would flash (see Fig. 7) and the PennState: Statistics Online Courses.
[Link] IBM, “What is the
user would not be able to proceed until the patient was in the k-nearest neighbors algorithm? | IBM,” [Link]. [Link]
correct position for the next part. Otherwise, the light would flash 8. JavaTpoint, “K-Nearest Neighbor(KNN) Algorithm for Machine
green. Learning - Javatpoint,” [Link], 2021.
[Link]
machine- learning
Results of testing the application showed that it had a high 9. D. Johnson, “Unsupervised Machine Learning: What is,
degree of accuracy when it came to assessing the correct poses, Algorithms, Example,” [Link], Sep. 21, 2019.
which only improved with more rounds of training and testing [Link]
the existing data. However, the researchers also noted that some 10. “Unsupervised Machine learning - Javatpoint,”
[Link]. [Link]
parts of the turning process were physically impossible to be machine-learning
analyzed by the camera; so, there is still room for human error. 11. “Pose landmarks detection task guide | MediaPipe,” Google Developers.
[Link]
Nevertheless, it is crucial to note that the use of the ark
application is not redundant, as it streamlines and adds precision
to most of the remaining turning process, ensuring that any
mistakes caught will be highly minimized.

Data Collection for the Fall Prevention Project uses the Intel
RealSense D435 camera [3], a depth camera frequently used in
robotics and virtual reality applications. MediaPipe, a computer
vision framework used in Python [4], is combined with the depth
camera to obtain motion data throughout a patient’s fall. Markers
are placed on the user’s shoulders, wrists, elbows, hips, knees,
and ankles in order to track their coordinate motion.
IV. CONCLUSION
In conclusion, the use of machine learning has accelerated
growth in various fields, with computer vision standing out as a
prime example of its transformative capabilities. By leveraging
machine learning algorithms, computer vision has revolutionized
how we interact with visual data.

One field that has particularly benefited from machine


learning in computer vision is the medical field. The integration
of machine learning algorithms in medical applications has
resulted in significant advancements, improving both patient
care and hospital processes.

As we move forward, it is crucial to harness the full potential


of machine learning while ensuring responsible practices and
ethical considerations. By doing so, we can unlock even greater
possibilities for innovation, discovery, and societal impact,
ultimately improving the way we interact with visual data and
transforming numerous fields for the better.

REFERENCES
1. S. Brown, “Machine learning, explained,” MIT Sloan, Apr. 21,
2021. [Link]
learning- explained
2. IBM, “Computer Vision,” IBM, 2019.
[Link]
vision
3. “The Complete Guide to Machine Learning Steps,”
[Link]. [Link]
learning- tutorial/machine-learning-steps
4. Jason Brownlee, “Linear Regression for Machine Learning,”
Machine Learning Mastery, Mar. 24, 2016.
[Link]
learning/
5. “Linear Regression in Machine learning - Javatpoint,”
[Link]. [Link]
in- machine-learning
6. “Linear & Quadratic Discriminant Analysis · UC Business Analytics
R Programming Guide,” [Link]. [Link]
[Link]/discriminant_analysis.

You might also like