A Comparative Study
A Comparative Study
net/publication/359964811
CITATIONS READS
0 128
3 authors, including:
All content following this page was uploaded by R. A. H. M. Rupasingha on 19 April 2022.
Sabaragamuwa University of Sri Lanka Sabaragamuwa University of Sri Lanka Information Systems
Belihuloya, Sri Lanka Belihuloya, Sri Lanka Sabaragamuwa University of Sri Lanka
rangikaishani75@[Link] hmrupasingha@[Link] Belihuloya, Sri Lanka
btgsk2000@[Link]
Abstract — Depending on the different activities of the behaviors that is unique to each student. In addition to the
students, their target goals may change. There it is more results, it is important to know what students can obtain based
important to identify the relation in their educational activities, on their various extracurricular activities and behavior
other extracurricular activities, and behaviors. Therefore, patterns and it is very important to understand the relationship
classification algorithms are used to identify such patterns. It between students’ performance and their other extracurricular
can be used to make predictions in education and it makes a activities, behavior in the university.
huge contribution, especially in a competitive world. We
examine what is the most appropriate classification algorithm Predicting and analyzing students’ performance is an
for predicting students’ academic performance and those important process of universities and academic institutions. It
decisions based on information obtained from the graduates of helps students to work with the understanding of their own
the Sabaragamuwa University of Sri Lanka. The data set is capacity and instructors to do their job better efficiently and
preprocessed by cleaning and removing unwanted attributes effectively with the understanding of students' capacity.
and removing missing values etc. The preprocessed data is Furthermore, the students can predict and then take necessary
ranked using Waikato Environment for Knowledge Analysis action to improve themselves. And also, it shows how the
(WEKA) data mining tool and to improve the results hyper- students manage their behaviors, daily routine schedules,
parameter tuning was used. The classification algorithms such extracurricular activities with their academic activities with
as Random Forest, Naïve Bayes, Multi-Layer Perceptron time management. Furthermore, if we can identify that, we
(MLP), and Decision Tree (J48) were used with the 80%
can test the talents and results of new students, and also, we
training data and 20% testing dataset for creating a model.
Random Forest algorithm presents best results considering that
can let them know how hard they have to work to reach their
well-known statistical accuracy metrics such as accuracy, academic goals. After identifying the students’ performance,
precision, recall, f-measure, Mean Absolute Error (MAE), and we can pay special attention to weak students and can point
Root Mean Square Error (RMSE) values. Based on the Random out the factors they need to improve.
Forest algorithm results, we can enhance the level of education We address this using ML algorithms. It is an advanced
and skills of students and create plans in the present to reach method of information discovery and data mining. It is
students’ future goals.
important to understand the user and handle a wide range of
Keywords—academic performance, prediction, classification,
data sets[2]. There are various techniques for machine
machine learning learning. But classification is the most used technology.
Classification is a task that should be valued in machine
I. INTRODUCTION learning and also future planning and knowledge
discovery[3]. Education data mining (EDM) is an important
Student performance is a major aspect of all educational
field focused on developing ways for extracting knowledge
institutions. As a result, students’ achievement evaluation is
from data generated in the education field. EDM goals are
required to support in order to reach the higher grade of the
recognizing the students’ behaviors, promoting education
institutions. This can be evaluated by machine learning (ML)
support systems. EDM can be used by higher education
algorithms using educational data on current students’
institutions to make predictions and forecast students’
performance. Higher educational institutions are one of the
performance[4].
most important factors of the development of any country[1].
Therefore, it should be considered very seriously. Nowadays There are various ML algorithms are available to classify
education system is not focused on classroom teaching and it the data such as decision tree, Support Vector Machine,
tends to online learning and teaching, web-based education Random Forest, Naïve Bayes, K-nearest neighbor, and MLP
systems, seminars, and workshops. etc. These ML algorithms can be applied to educational data
then institutions and students can focus on what to teach and
Today Grade Point Average (GPA) system is used in the
how to teach to achieve the objectives and goals. The main
university system of Sri Lanka to measure students’
objective of this research is to classify the collected student
performance. Their final results calculate on the quizzes,
information and create a model for student’ performance
classroom assessments, and final examination results. In
prediction using classification algorithms. In this approach,
addition to their academic activities, university students
we have done a comparative study with different four ML
engage in various types of extracurricular activities. Such as
algorithms namely Random Forest, decision tree, Naïve
athletics, other sports, traveling, hobbies, special events in the
Bayes, and MLP. Among them Random Forest show better
university so on. These factors affect the students’
results overcoming the results of other ML algorithms.
performance with their final result. There is also a pattern of
386
Authorized licensed use limited to: University of Otago. Downloaded on April 19,2022 at 10:02:07 UTC from IEEE Xplore. Restrictions apply.
tested the data set with 32 attributes and after 8 most A. Data collection
significant attributes. After that process, OneR, REP Tree,
Simple Logistic, JRip, Naïve Bayes, and Decision Tree (J48)
have more than 70%. But Random Tree algorithm results have
less than 70%. Result of the implementation, they found not
only performance at school time but also study time affect the
students’ final grade.
When comparing the existing approaches, most of the
approaches used only a few attributes, and some approaches
used a few data for their model. But less number of attributes
and fewer data is not enough for getting an accurate result.
And also existing works apply one algorithm or compare with
a few classification algorithms. In our approach, we increased
the number of attributes and compare with four classification
algorithms to find the best algorithm with the highest accuracy
and less error rate.
III. PROPOSED APPROACH
The proposed approach is explaining in the Fig. 2. It
contains the main steps such as data collection, data
preprocessing, and classification process.
387
Authorized licensed use limited to: University of Otago. Downloaded on April 19,2022 at 10:02:07 UTC from IEEE Xplore. Restrictions apply.
B. Data Preprocessing
The data preprocessing process is used to remove
inconsistent, noisy, and incomplete data in the raw dataset.
The raw data is included low-quality values and it may be
affected to the data mining process. Therefore, data
preprocessing plays a major role in the data mining process.
Data cleaning, data transformation, and data reduction are
applied for preprocessing the data. In the beginning, data set
contained 21 attributes, but the data was reduced to 17
attributes including one dependent variable and other
independent variables. Some attributes were removed
because they were identified as not important to build a
prediction model. The department, district, types of hobbies,
and extracurricular activities are removed. And also, hyper-
parameter tuning was used in the WEKA data mining tool to
rank attributes. Monthly family income and degree type are
shown in low ranks. Therefore, the top 17 attributes are
selected from the ranked list as shown in Table II and Fig. 4
presents value details of the ranked attributes.
Fig. 4: Attributes in the sample
TABLE II. THE VALUE DETAILS OF RANKED ATTRIBUTES
C. Classification
Attribute Values After the data preprocessing, the WEKA 3.8.5 data mining
tool is used for the classification and evaluation process. The
Hobbies time in the 0-2h, 2h-4h, 4h-6h, 6h>
semester (hours)
input data set is stored in CSV file format. The data set is
Ways of gathering Gathering knowledge through Internet,
categorized as training data set and test data set. It is split 80%,
extra knowledge Joining Knowledge sharing sessions, 20% respectively. The model is developed targeting the
Academic Discussion with friends, Self- students’ final results based on the final class as high class or
study, Referring Books low class obtained by the students. After partitioning the data
Doing past papers Frequently, Normal, Infrequently for training and testing, the training data was applied into the
Random Forest, Naïve Bayes, decision tree, and MLP
Browsing social media Frequently, Normal, Infrequently classification algorithms to evaluate the result and find the
best classification model.
Chatting with friends Frequently, Normal, Infrequently
Each model is tested by selected testing data set and the
General English A, B, C, S, F Random Forest classification algorithm shows better results
Results in Advanced based on the results of the accuracy, precision, recall, f-
Level exam measure, MAE, and RMSE values. It was selected as the best
Extracurricular activity 0-2h, 2h-4h, 4h-6h, 6h>
time in the semester
classification algorithm for predicting academic performance
(hours) of students. Random Forest is powerful and popular
Hobbies time in exam 0-2h, 2h-4h, 4h-6h, 6h> classification algorithm that consists of several individual
period (hours) decision trees works together. Every tree in the Random
Making short notes Frequently, Normal, Infrequently Forest spreads a class prediction. It has more chances to build
the best prediction model. Furthermore, Random Forest can
Study time in exam 0-2h, 2h-4h, 4h-6h, 6h> handle the missing values as well as can maintain the accuracy
period (hours) of a huge data set.
Study time in the 0-2h, 2h-4h, 4h-6h, 6h>
semester (hours) IV. EXPERIMENTS AND EVALUATIONS
Extracurricular time in 0-2h, 2h-4h, 4h-6h, 6h>
exam period (hours) The experimental platform used Microsoft Windows 10 on
Following Extra Yes (During the university period), Yes a PC with Processor Intel® Core (TM) i3-5005U CPU @
courses (During after A/L period ,During the 2.00GHz, RAM 4.0GB. The WEKA 3.8.5 data mining tool is
university period), Yes (During after A/L used as the training and testing environment. The collected
period), No 200 data using a Google form from the graduates in the
Satisfaction of the Strongly Agree, Agree, Neutral, Disagree, Sabaragamuwa University of Sri Lanka was used as a data set
followed extra Courses Strongly Disagree, Have not followed
for the evaluation process.
Attending lectures Frequently, Normal, Infrequently
In this study Decision Tree (J48) and Random Forest in
Gender Female, Male Trees classifier, Naïve Bayes in Bayes classifier, and MLP in
Functions classifier is used for the comparison and
Class (Target variable) High Class (First class and Second-class evaluations. The accuracy results of these algorithms are
upper division), Low Class (Second class shown in Fig. 5.
lower division and normal pass)
388
Authorized licensed use limited to: University of Otago. Downloaded on April 19,2022 at 10:02:07 UTC from IEEE Xplore. Restrictions apply.
100%
Based on the MAE and RMSE, the lowest error increases
90% the accuracy of the algorithm. These two criteria represent in
80% (4) and (5). Here n is sample size, is actual value, and is
70%
predicted value. As shown in the Table IV, the RMSE value
is lowest in the Random Forest, and MLP show the lowest
60% error in MAE. But MAE of the Random Forest is also a low
50% value comparing with other algorithms.
1
&
!"# = $| − | (4)
40%
30%
20% '(
1
&
10%
!*# = + $( − ), (5)
0%
'(
Decision Tree Naïve Bayes Random Multilayer
(J48) Forest Perceptron
= (1)
Classifier TN FP FN TP
+
Decision Tree (J48) 127 33 3 37
= (2)
Naïve Bayes 108 31 22 39
+
Random Forest 130 1 0 69
2 × ×
− = (3)
MLP 130 2 0 68
+
Following Table VI shows the accuracy of the Random
Forest for different training data and testing data set. Based on
Table III shows precision, recall, and f-measure values of the result, Random Forest shows the best accuracy of 97.5%
four algorithms. Compared with all used classification in 80%, 20% percentages for training and testing data
algorithms, the highest precision, recall, and f-measure respectively.
values are shown in the Random Forest algorithm. TABLE VI. ACCURACY OF RANDOM FOREST ALGORITHM
389
Authorized licensed use limited to: University of Otago. Downloaded on April 19,2022 at 10:02:07 UTC from IEEE Xplore. Restrictions apply.
According to the above evaluations, Random Forest is the [7] H. H. Patel and P. Prajapati, “Study and Analysis of Decision Tree
best algorithm for predicting the academic performance of the Based Classification Algorithms,” Int. J. Comput. Sci. Eng., vol.
6, no. 10, pp. 74–78, 2018.
students. It presents not only the highest values in accuracy, [8] M. A. P. Praveen Kumar and M. Sundararajan, “Predicting
precision, recall, and f-measure and also the lowest error value Educational Performances of Under Graduates Students Using
in RMSE. The lowest error rate shown in MLP but the error Rule Based Decision Tree J48 Classification,” Journalof Crit.
rate in the random forest is also very low comparing with other Rev., vol. 7, no. 14, p. 2020, 2020.
algorithms. [9] A. Khalaf, “Selection of Best Decision Tree Algorithm for
Prediction and Classification of Students’ Action,” Am. Int. J. Res.
As compared to existing approaches, the accuracy od the Sci. Technol. Eng. Math., vol. 16, no. October, p. 26, 2016.
models is measured based on the higher attributes and the [10] A. K. Hamoud, A. S. Hashim, and W. A. Awadh, “Predicting
Student Performance in Higher Education Institutions Using
extracurricular activities in addition to the students’ academic Decision Tree Analysis,” Int. J. Interact. Multimed. Artif. Intell.,
activities. This research shows that Random Forest is shown vol. 5, no. 2, p. 26, 2018.
high accuracy compared to other approaches. [11] E. Osmanbegović and M. Suljić, “Data Mining Approach for
Predicting Student Performance,” Econ. Rev. – J. Econ. Bus., vol.
V. CONCLUSION 10, no. 1, 2012.
[12] D. C. R. Durgesh Ugale, Jeet Pawar, Sachin Yadav, “Students
Performance Prediction using Data Mining Techniques,” Int. Res.
The purpose of this study is to predict the academic J. Eng. Technol., vol. 7, no. 5, p. 1081, 2020.
performance of the students based on their other activities [13] T. Jeevalatha, N. A. [Link], and D. S. Kumar, “Performance
such as extra-curricular activities, behaviors in the university, Analysis of Undergraduate Students Placement Selection using
daily routine, and hobbies. Information obtained online from Decision Tree Algorithms,” Int. J. Comput. Appl., vol. 108, no. 15,
pp. 27–31, 2014.
the graduates of the Sabaragamuwa University of Sri Lanka. [14] H. Bydžovská, “A comparative analysis of techniques for
The prediction results help to enhance the academic level of predicting student performance,” Proc. 9th Int. Conf. Educ. Data
the students as well as the quality of the institution. Mining, EDM 2016, vol. 2, pp. 306–311, 2016.
The preprocessed data are applied and decision tree (J48) [15] G. Kaur and W. Singh, “Prediction Of Student Performance Using
and Random Forest in trees classifier, Naïve Bayes in Bayes Weka Tool,” Res. Cell An Int. J. Eng. Sci., vol. 17, no. January,
pp. 2229–6913, 2016.
classifier, and MLP in Functions classifier are used as [16] R. Asif, A. Merceron, S. A. Ali, and N. G. Haider, “Analyzing
machine learning algorithms to compare the evaluation undergraduate students’ performance using educational data
results. All the algorithms are tested with supplied test set mining,” Comput. Educ., vol. 113, pp. 177–194, 2017.
option to calculate the accuracy of the algorithms. Out of the [17] Y. K. Salal, S. M. Abdullaev, and M. Kumar, “Educational data
selected four algorithms, random forest shows better results mining: Student performance prediction in academic,” Int. J. Eng.
Adv. Technol., vol. 8, no. 4C, pp. 54–59, 2019.
and it is the best for forecasting the students’ performance.
Random Forest algorithm presents 97.5% accuracy and also
highest precision, recall, and f-measure values with the lower
error rate in MAE and RMSE values. As a result of this study,
the higher education institutions will be able to identify the
performance of the students and identify the weak students
and take steps that need to be taken to help them to reach the
good final result with the highest final class. This can be used
not only in higher education institutions but also in schools,
tuition classes, etc.
We are intent to enlarge the data set and attributes for
applying to ensemble learning and clustering algorithms.
Also, the future plan is to continue this research by obtaining
data from the students in whole university system. It will be
an in-depth study of how it affects to predict students’
performance.
REFERENCES
[1] A. Namoun and A. Alshanqiti, “Predicting student performance
using data mining and learning analytics techniques: A systematic
literature review,” Appl. Sci., vol. 11, no. 1, pp. 1–28, 2021.
[2] M. Durairaj and C. Vijitha, “Educational Data mining for
Prediction of Student Performance Using Clustering Algorithms,”
Int. J. Comput. Sci. Inf. Technol., vol. 5, no. 4, pp. 5987–5991,
2014.
[3] K. David Kolo, S. A. Adepoju, and J. Kolo Alhassan, “A Decision
Tree Approach for Predicting Students Academic Performance,”
Int. J. Educ. Manag. Eng., vol. 5, no. 5, pp. 12–19, 2015.
[4] A. M. Shahiri, W. Husain, and N. A. Rashid, “A Review on
Predicting Student’s Performance Using Data Mining
Techniques,” Procedia Comput. Sci., vol. 72, pp. 414–422, 2015.
[5] A. Kalpesh, G. Aditya, D. Amiraj, J. Rohit, and H. Vipul,
“Predicting Students’ Performance Using Id3 and C4.5
Classification Algorithms,” Int. J. Data Min. Knowl. Manag.
Process, vol. 3, no. 5, pp. 39–52, 2013.
[6] S. Astha, K. Vivek, K. Rajwant, and H. D., “Predicting student
performance using data mining techniques,” Int. J. Pure Appl.
Math., vol. 119, no. 12, pp. 221–227, 2018.
390
Authorized licensed use limited to: University of Otago. Downloaded on April 19,2022 at 10:02:07 UTC from IEEE Xplore. Restrictions apply.
View publication stats