0% found this document useful (0 votes)
6 views1 page

Classroom Activity Video Analysis Methodology

The project methodology involves uploading student images and names to a database, creating a list of activities with competency weightages, and using computer vision for face recognition to generate focused videos. These videos will be analyzed by a multimodal vision language model to provide competency scores and remarks for each student. If the model fails, simpler algorithmic approaches will be explored, and five cameras will be installed for comprehensive classroom coverage.

Uploaded by

J.V. Raghunath
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views1 page

Classroom Activity Video Analysis Methodology

The project methodology involves uploading student images and names to a database, creating a list of activities with competency weightages, and using computer vision for face recognition to generate focused videos. These videos will be analyzed by a multimodal vision language model to provide competency scores and remarks for each student. If the model fails, simpler algorithmic approaches will be explored, and five cameras will be installed for comprehensive classroom coverage.

Uploaded by

J.V. Raghunath
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Project Methodology

We are planning to get two sample activity videos and its database to analyse them. Here is
the methodology that we devised for this project,

1)​ Upload the face images and name of each student from a particular class to the
database
2)​ Have a set of activity list for those videos
a)​ We assume same activity for each kid of the class at a given time
b)​ We create possible list of competencies and its weightages along with its
description for the given activity
(For example, the activity of paper folding has precise measurements (60%),
sturdiness of the model (30%), time taken (10%) as its individual competencies
and weightages)
3)​ Use computer vision to generate the names by face recognition and export a video that
focuses on them primarily
4)​ We feed this video and competency description(s) into a “multimodal VLM” (vision
language model), that uses computer vision and Natural Language Processing, and
provides competency score(s) and remarks based on the performance, as a PDF report,
for each of the students
5)​ If VLM step doesn’t work, we have to break it down into further simple algorithmic cases
of the competencies

Camera - We are planning to fix five cameras in the ceiling, one in each corner of the classroom
and one more in the center of the ceiling to get the full coverage of activities. We will focus on
each student (or) a small group for about 5 minutes to understand their activity and then we
move to neighbouring students. The camera should focus only on this particular group which is
done through the camera’s pan-tilt-zoom (PTZ) controls.

You might also like