Project Methodology
We are planning to get two sample activity videos and its database to analyse them. Here is
the methodology that we devised for this project,
1) Upload the face images and name of each student from a particular class to the
database
2) Have a set of activity list for those videos
a) We assume same activity for each kid of the class at a given time
b) We create possible list of competencies and its weightages along with its
description for the given activity
(For example, the activity of paper folding has precise measurements (60%),
sturdiness of the model (30%), time taken (10%) as its individual competencies
and weightages)
3) Use computer vision to generate the names by face recognition and export a video that
focuses on them primarily
4) We feed this video and competency description(s) into a “multimodal VLM” (vision
language model), that uses computer vision and Natural Language Processing, and
provides competency score(s) and remarks based on the performance, as a PDF report,
for each of the students
5) If VLM step doesn’t work, we have to break it down into further simple algorithmic cases
of the competencies
Camera - We are planning to fix five cameras in the ceiling, one in each corner of the classroom
and one more in the center of the ceiling to get the full coverage of activities. We will focus on
each student (or) a small group for about 5 minutes to understand their activity and then we
move to neighbouring students. The camera should focus only on this particular group which is
done through the camera’s pan-tilt-zoom (PTZ) controls.