Predicting Student Dropout with ML
Predicting Student Dropout with ML
Intelligence applications worldwide. Few of the Data mining unexpected surge and for IT professionals having Machine
applications are: Learning skills.
Market Segmentation Predictive analytics is a section of data mining
Fraud detection sighted at obtaining predictions about future results based on
Market based analysis factual data and techniques such as statistical modelling
Trend Analysis algorithms and machine learning. The science of predictive
Prediction Systems analytics can produce insights with a vital degree of precision
Customer churn and accuracy. With the assistance of sophisticated data
mining tools and algorithm models, any industry can now use
Insurance and health car
past and current data to reliably predict trends and
performances which are seconds, days, or years into the
II. PROBLEM DEFINITION INCLUDING THE SIGNIFICANCE AND
future.
OBJECTIVE
Few examples of the various industries which use
There is a distinct rise in the demand for engineers in the IT the data mining techniques to predict the future trends and
industry too in our country. Though there is a dire need, most analyze them are:
of the students are not aware of this. Therefore, though the Aerospace: Used to forecast the result of particular
industry is rising in demand, the students are not able to meet maintenance operations on aircraft reliability, use of fuel
their requirements and the eligibility criteria. Due to various and availability.
factors, many may undergo severe issues that may force them Automotive: To consolidate records of segment
to discontinue their studies. It might be financial, personal, or sturdiness and predict the failure rate in the upcoming
any unforeseen demanding situations. Thus, the plans of vehicle manufacturing. The driver’s behaviour
administration and management of the colleges waste huge is studied to improve the driver assistance technologies
chunks of money on discontinued students' resources. A drop and, ultimately, autonomous vehicles.
out prediction system can help the institutions to balance and Energy: Predict long-term price and supply-demand
use their finances and resources. ratios and to determine the impact of weather disasters,
failure of equipment, rules and additional variables on
III. OBJECTIVES: service costs.
Help universities cope up with the phenomenon of Financial services: To develop credit-risk models.
adapting to rising rate of dropouts Predict the trends of the financial market. Predict the
Prevention and intervention services to upgrade the influence of new strategies, laws and regulations on
students’ retention rate. business markets.
Manufacturing: Foretell the location and machine failure
IV. SCOPE OF THE PROJECT: rates. To improve the efficiency of raw material
deliveries based on predicted future demands.
To study the how efficient the campus education
methods are which will be extremely helpful for both Law enforcement: Use the trends of crime data to
students and Institution. determine neighbourhood that may require further
security at definite times of the year.
To Build a predictive model which can be used to
forecast the probability of a randomly chosen student, Retail: Develop an online client in real-time to decide
who will be tested if he/she will graduate or not. whether providing extra product knowledge or incentives
will enhance the probability of a finished transaction.
To Identity the factors which affect the probability of a
student to dropout of a technical education. B. Existing System:
An educational institute contains student records that are
V. TECHNIQUES AND METHODS: affluence of information but are too huge for one person to
Data mining techniques understand in their entirety. Finding essential features from
Feature selections this data is an important task in educational research. Finding
Classification algorithms in ML the academic and financial status of each student in an
Prediction based on attributes and performance institute is a wearisome task. Hence, the modification of the
system includes time consumption, less efficiency and less
VI. LITERATURE SURVEY user satisfaction. Furthermore, this system is a manual
process that supplements the limitation.
A. Introduction to the Problem Domain Terminology
C. Design of the Proposed System:
Today Artificial Intelligence (AI) is beyond the technologies
of blockchain and quantum computing. It has also found its The proposed system uses two data mining techniques for
place in smaller projects and is made easier as even the predictive analysis.
common man also to work with. Machine Learning models Decision Tree Classifier
are created to re-train the existing models for better Gaussian Naïve Bayes Algorithm
performance and resulted in an efficient way to solve a
problem. High-Performance Computing (HPC), which are
now easily in reach to everyone which resulted in an
VII. RELATED WORKS: or fails within the current semester before attempting the
1) Data processing Approach for Predicting Student and ultimate examination. The author uses SVM, Naive
Institution's Placement Percentage by Professor Ashok Bayes, Random Forest Classifier and Gradient Boosting
M, Professor Apoorva A, 2016 International Conference to compute the result. Boosting is an ensemble learning
on Computational Systems and knowledge Systems for algorithm which mixes various learning algorithm to
Sustainable Solutions during this paper author has used get better predictive performance.
the info mining technique for the prediction of the 5) Student Placement Analyzer: A Recommendation
student’s placement. For the prediction of student's System Using Machine Learning, Apoorva Rao R,
placement author has divided the information into the 2 Deeksha K C, Vishal Prajwal R, Vrushak K, Nandini,
segments, first segment is that the training segment JARIIE-ISSN(O)-2395-4396 Now-a-days
which is historic data of passed out students. Another institutions face many challenges regarding student
segment consists of current data of scholars, supported placements. For educational institutions it's much
the historic data author has designed the algorithm for difficult task to stay record of each single student and
calculating the position chances. Author has used the predict the location of student manually. To beat these
varied data processing algorithms like decision tree, challenges, concept of machine learning and various
Naive Bayes, neural network and therefore the prosed algorithms are explored to predict the results of class
algorithm were applied, and decision are made with the students. For this purpose, training data set is historical
assistance of confusion matrix. data of past students and this is often wont to train the
2) Student Placement Analyzer: A Recommendation model. This software package predicts placement status
System Using Machine Learning, by Senthil Kumar in 5 categories viz dream company, core company, mass
Thangavel, Divya Bharathi P, Abijith Sankar, recruiter, not eligible and not inquisitive
International Conference on Advanced Computing and about placements. This method is additionally helpful to
Communication Systems (ICACCS -2017) ), Jan. 06 - weaker students. Institutions can provide extra care
07, 2017, Coimbatore, INDIA during this paper author is towards weaker students in order that they'll improve
concern about the challenges face by any institute their performance. By use Naïve Bayes algorithm all the
regarding the position. the location prediction is information are going to be monitor and appropriate
extremely complex when the quantity of the entities decision are going to be provided.
increases in any institute. With the assistance of machine
learning this complex problem of prediction are often VIII. TECHNOLOGIES AND METHODS PYTHON
easily solved. during this paper all the tutorial record of A. History of Python:
student is taken into consideration. Various classification
Guido van Rossum developed Python in the early nineties at
and data making algorithms are used like Naïve Bayes,
the National Research Institute for Mathematics and
Decision Tree, SVM and Regressions. After the
Computer Science, Netherlands. Python is a derivation of
prediction of the scholars will be placed in of the given
many other languages, namely ABC, Modula-3, C, C++,
category that's Core Company, dream company or
Algol68, Smalltalk, and Unix shell and few other scripting
support services.
languages. Python has been copyrighted immediately after its
3) A Placement Prediction System Using K-Nearest
establishment. Python source code is open to all with the
Neighbors Classifier, by Animesh Giri, M Vignesh V
GNU General Public License (GPL). A dedicated technical
Bhagavath, Bysani Pruthvi, Naini Dubey, Second
team works for the development of Python while the major
International Conference on Cognitive Computing and
decisions regarding the technology are still taken by Mr
data Processing (CCIP), 2016 the position prediction
Guido van Rossum.
system predicts the probability of scholars getting placed
in various companies by applying K-Nearest Neighbors B. Input as CSV File:
classification. The result obtained is additionally Data Science often requires the action of reading and taking
compared with the results obtained from other machine inputs from CSV(Comma Separated Values) files as a
learning models like Logistic Regression and SVM. the prerequisite to several tasks. Usually, the inputs from which
tutorial history of student together with their skill sets come in from varied sources are converted into CSV files. To
like programming skills, communication skills, make it easier, Python offers a library called the Pandas
analytical skills and team work is considered which is library which renders different features to read the CSV files
tested by companies during recruitment process. Data of according to the developer's need. The CSV file is a text file,
past two batches are taken for this system. similar to a Microsoft Excel sheet, the only difference is that
4) Class Result Prediction using Machine Learning, by the values in the columns are separated by a comma in CSV
Pushpa S K, professor, Manjunath T N, Professor and files.
Head, Mrunal T V, Amartya Singh, C Suhas, For example, suppose the data is present in the file
International Conference on Smart Technology for Smart named [Link]. This file can be created using windows
Nation, 20172017 during this paper, the results of a notepad by just copying and pasting the data and saving it
category is predicted using machine learning. with an extension ".csv"
Performance of scholars in past semester together import pandas as pd
with ample internal examinations of this semester is dt_data= pd.read_csv(‘path/[Link]’)
taken into account to predict whether the scholar passes print(dt_data)
C. Operations using NumPy: problems which none of the other supervised learning
NumPy (Numerical Python) is a python library that provides algorithms can do. The purpose of using a Decision Tree as a
high level functions to operate mathematically on arrays. It classifier is to build a training model to predict the target
consists of multidimensional array objects and a compilation variable's class/value by acquiring simple decision tree rules
of methods for processing array. deduced from earlier data(training data).
Using NumPy, the following operations can be To predict a class label of a record, the values of the
performed − root attribute are compared with the attribute of the record.
Mathematical and logical operations. Based on this comparison, we understand the branch
corresponding to that value and move to the next node.
Fourier transforms and routines, which are used for shape
manipulation.
XI. CONCLUSION
Operations related to linear algebra. NumPy also has
preprocessed functions for linear algebra and random Looking at the statistics of the last few years, the number of
number generation. student dropout from the educational institute is rapidly
increasing. It affects the educational institutions besides the
D. Key Features of Pandas: future of the student hence posing as a major threat to all
Fast and efficient DataFrame object with efficient educational institutions. The reasons for dropping out may
indexing methods vary. It depends on various factors that lead to predicting the
Different formats of the inputs can be stored into the student who may drop out. In this project, we concentrate on
memory using the library tools the reason behind the student's drop out. So, data collection
Alignment of data and integrated handling of data that is plays a vital role in this paper. The collected data is evaluated
missing. by several techniques under data preprocessing methods. The
Data sets can be reshaped and pivoted. data which is collected by various resources contains many
Label-based slicing determinants like Academics, Demographical factors,
large data sets can be efficiently indexed and subsetting Psychological factors, Health issues etc. which play
can be used. important roles in a student's decision to drop out. This
Insertion and deletion of Columns from a data structure. prediction system will help to predict the student who will be
Group by data used for aggregation and transformations. choosing to drop out of the registered course. Identifying
them at an early stage will prevent the student to drop out and
High-performance merging and joining of data.
the institution can monitor and give valuable counselling to
Time Series functionality.
them in persuading to change the mind of the student from
dropping out. It will also help the student to know the right
IX. NAÏVE BAYES ALGORITHM path to achieve their dreams.
Naïve Bayes classifier is an algorithm that is based on the
Bayes theorem. It has a substantial independence assumption. REFERENCES
It is also called "independent feature model". It presupposes
[1] K. Bonneau and D. Management, “Brief 3: What is a
the presence or absence of a particular feature of a class is
Dropout?” pp. 14–17, 2007.
irrelevant to the presence or absence of any other feature in a
[2] A. Pradeep, S. Das, and J. J. Kizhekkethottam,“Students
supposed class. Naïve Bayes classifier can be trained in a
dropout factor prediction using EDM techniques,” Proc.
supervised learning model too. It applies a method of highest
IEEE Int. Conf. Soft- Computing Netw. Secur. ICSNS
similarity. It has worked in a complicated real-world
2015, 2015. A. Cheah et al., “Analyzing Students
problem. It needs only a small quantity of training data. It
Records to Identify Patterns of Students’ Performance,”
determines parameters for classification. The variance of the
pp. 544–547.
variable needs to be calculated for each class and not the
[3] K. Chai, H. T. Hn, and H. L. Cheiu, ‘Naive-Bayes
entire matrix. Naïve bayes is mainly used when the inputs are
Classification Algorithm’, Bayesian Online Classif. Text
high. It gives output in a better sophisticated form. Every
Classif. Filter, pp. 97–104, 2002.
attribute's probability is displayed in the Prediction table.
[4] S. S. Panwar, ‘of Computer © I a E M E Data Reduction
Machine learning and data mining methods are based on
Techniques To Analyze Nsl-Kdd Dataset’, pp. 21–31,
naïve bayes classification.
𝑃(𝐻|𝑋) = 𝑃(𝑋|𝐻) 𝑃(𝐻) 2014.
Bayes theorem: [5] V. Hegde and S. G. Kini, ‘Multivariate and Multi-
𝑃(𝑋)
Where P(H|X ) is the posterior probability of H conditioned Behavioral Student Dropout Prediction Using Naïve
on X Bayesian Algorithm’.
P(X|H) is the posterior probability of X conditioned on H [6] A. Cheah et al., “Analyzing Students Records to Identify
P(H)is the prior probability of H Patterns of Students’ Performance,” pp. 544–547.
P(X) is the prior probability of X [7] V. Hegde, “Dimensionality Reduction Technique for
Developing Undergraduate Student Dropout Model
X. DECISION TREE CLASSIFIER using Principal Component Analysis through R
Package,” pp. 1–6, 2016.
The Decision Tree Classifier algorithm comes under
[8] M. Nasiri, B. Minaei, and F. Vafaei, “Predicting GPA
Supervised Learning. The decision tree algorithm can also be
and academic dismissal in LMS using educational data
used for solving regression and classification