90.weather Prediction Using Machine Learning Algorithms
90.weather Prediction Using Machine Learning Algorithms
LEARNING ALGORITHMS
Kochi, Amrita Vishwa Vidyapeetham, Kochi, Amrita Vishwa Vidyapeetham, Kochi, Amrita Vishwa Vidyapeetham,
India India India
Kochi, India Kochi, India Kochi, India
aiswaryashaji.98@[Link] amrithavihar123@[Link] rajiprithviraj@[Link]
Abstract— Weather forecasts have grown increasingly computational systems and provide quick and accurate
significant in recent years since they can save us time, money, forecasts that we can utilise in our everyday lives.
property, or even our lives. Despite the fact that India has a large (1) One of the paper's major advances is that machine
number of weather stations, they are mainly located in inhabited learning may be used to anticipate weather conditions over
regions such as cities, suburbs, or towns. This makes weather short periods of time using less resource-intensive
forecasting in isolated regions more imprecise, which can be
equipment.
inconvenient for individuals such as farmers who rely largely on
weather reports in their daily work. In this paper, we are predicting (2) A thorough assessment of the suggested technique, as well
the weather by analyzing features like temperature, apparent as a comparison of numerous machine learning models for
temperature, humidity, wind speed, wind bearing, visibility, cloud forecasting future weather conditions.
cover with Random Forest, Decision Tree, MLP classifier, Linear
regression, and Gaussian naive Bayes are examples of machine II. LITERATURE REVIEW
learning methods. Based on the results obtained a comparative Several researchers compared weather predictions using
study is done concerning the accuracy. different methods and a few of them are mentioned below.
Keywords— Machine Learning; Prediction; Weather In,[19] the author introduced a methodology of using the R
Forecast; Classification. tool, and a comparison of Decision Tree and Random Forest
was conducted. The algorithm is executed using Rattle, an R
I. INTRODUCTION
GUI for analysis of data mining algorithms, that will predict
Pattern extraction from data sets was done manually in the whether it rains the next day or not. The redistribution error
past. The collecting, manipulation, and storage of data sets rate is compared; and found that a random forest has less error
have expanded tremendously in the modern computer era. value than a decision tree since the decision tree is an
Pattern recognition becomes extremely sophisticated as a ensemble of trees. [8] The author presented the approaches
result of this. The computerized application of specialized for improving the accuracy of random forest classifiers. To
algorithms to detect specific patterns from massive data sets improve the accuracy, a weighted hybrid decision tree model
is known as data mining. Machine learning is a type of data is used. Disjoint partitions of the training dataset and ranking
mining in which a model is created by learning concepts with of training bootstrap samples are two more methods used.
a computer. This model learns by training and testing for the Both of these are leading effective learning and classification
supplied large data sets, and it predicts future data instances using the Random Forest classifier.[18] The objective of their
using the learning principle. Modeling refers to the practice project was to build a desktop application that predicts
of creating a classification model to anticipate an outcome weather automatically and extracts data that is needed to
Classification is the data mining process of creating a model capture global data. The author collected a dataset on the
based on one or more attributes of categorical variables to actual weather of Nashville city from [Link].
predict the value of a target categorical variable. Based on the After pre-processing, some features with empty or invalid
training set and class labels, this can classify the given data. data are eliminated. After splitting the dataset into training
Since weather systems can extend a significant distance in all and testing data. The first phase of results brings the accuracy
directions over time, the weather of one locality can have a of prediction by adding more features. Since the predicted
massive effect on the weather of others. In this paper, we results are continuous numerical values, the author used
offer a method for predicting weather conditions by Random Forest Regression (RFR) which is considered a
combining historical weather data from nearby cities with superior regressor. Several machine learning strategies with
data from a single city. These data are combined and used to the RFR process are also considered in building the model.
build simple and basic machine learning models that can The model performance is measured using Mean Squared
correctly foretell weather conditions for the next several days. Error. And concluded that the system operates at a high level
These simple models may be run on low-cost, constrained of efficiency without malfunctioning. As it is an application,
the author points out that it will not consume much RAM and
Authorized licensed use limited to: SASTRA. Downloaded on September 25,2025 at 07:15:40 UTC from IEEE Xplore. Restrictions apply.
forests, decision trees, bagged trees, boosted trees, and According to Wikipedia, a Random Forest is a classifier that
boosted stumps in a large-scale empirical comparison. They averages multi-criteria trees from multiple subsets of a
also look at how applying Platt Scaling and Isotonic dataset to improve the dataset's projected accuracy. The
Regression to calibrate the model’s influences their random forest, rather than relying on a single decision tree,
performance. One of the most important aspects of their work incorporates inputs from each tree and forecasts the eventual
was the requirement for a variety of performance indicators output based on the majority votes of projections.
to evaluate the learning techniques. They concluded that
learning approaches like as boosting, random forests,
B. Decision Tree Algorithm
bagging, and SVMs perform remarkably well. [2] In this
research, the researchers identified three data mining A sort of predictive modeling known as decision tree
techniques: Nave Bayes, back-propagated neural networks, analysis can be used to a wide range of scenarios. An
and the C4.5 decision tree algorithms. They used these algorithmic technique can also be used to generate decision
algorithms to estimate the survivability rate of the SEER trees, which can segment data in a variety of ways based on
breast cancer data set, and they determined that these three particular parameters. Decision trees are the most powerful
categorization systems were the most effective at forecasting algorithm in the realm of supervised algorithms.
cancer survival [Link] data mining strategies are
reviewed for accuracy, with the goal of having high accuracy C. Gaussian Naïve Bayes
in addition to high precision and recall metrics. The C4.5
Continuous data is supported by the Gaussian Naive Bayes
algorithm is more accurate. [20] The authors of this paper
version, which follows the Gaussian normal distribution. The
presented a novel use of NN techniques in severe numerical
Naive Bayes classification methods, which are supervised
modeling of the environment. They created an HGCM, a
complicated hybrid environmental numerical model that machine learning classification algorithms, are based on the
combines deterministic modeling and machine learning Bayes theorem. It's a simple categorization method with a lot
of power. They're advantageous when the inputs'
techniques in a synergetic way. This method uses neural
dimensionality is high. The Naive Bayes Classifier can also
networks as a statistical or machine learning tool to create
handle complex classification issues.
highly accurate and quick simulations of the most time-
consuming deterministic model components. Other
complicated numerical models utilized outside of the realm D. MLP Classifier
of environmental modeling applications, such as advanced Backpropagation is used to train a multi-layer perceptron
models in computational physics, chemistry, biology, and so (MLP) technique in the MLP Classifier. A multilayer
on, can benefit from the established hybrid modeling perceptron can have more than one linear layer (combinations
paradigm and related NN emulation technology. [7] The of neurons). The input layer receives our data, and the output
paper analyses the performance of datasets using several layer receives our output. By increasing the number of hidden
classification techniques, with accuracy and execution time layers, we may build the model as complex as we like.
as an evaluation criteria. The performance of classification
techniques is observed to vary with diverse datasets. Dataset,
Number of instances and attributes, and Type of attributes are E. Linear regression
all factors that influence the classifier's performance. Other In machine learning, linear regression lets you uncover
data sets used in the comparison yielded excellent patterns and relationships in data so you can make an
performance with J48 and NaiveBayesUpdatable. [14] informed choice or prediction. In machine learning, linear
Anomaly Detection System (ADS) monitors a system's regression models a linear connection between data features.
behavior and marks significant departures from expected A linear relationship across continuous variables is
behavior as anomalies. Anomaly detection is used to detect modeled by linear regression. We examine two variables, one
computer network assaults, malicious actions in computer of which is a predictor and the other of which is a response.
systems, and Web-based system misuses. The paper
discusses anomalies and many supervised and unsupervised
anomaly detection strategies, as well as individual K-means IV. EXPERIMENTAL TECHNOLOGY
and Id3 Decision Tree usage, comparative research, and the
proposed system's combined approach. To summarise, the
training instance is first partitioned into k different clusters
using the k-Means technique. The ID3 decision tree in each A. Dataset
cluster learns the cluster's sub-classifies and divides the For our proposed system weather data is collected from
decision space into classification sections. [Link] and processed using python. We are considering
the attributes and their summary for the prediction. We
focused on eight factors in the dataset and they are
III. ALGORITHMS TAKEN FOR COMPARISON temperature, apparent temperature, humidity, wind speed,
wind bearing, visibility, cloud cover, and pressure. There is a
total of 27 summaries and they are: Partly cloudy, Mostly
A. Random Forest Algorithm cloudy, Overcast, Foggy, Breezy and mostly cloudy, breezy
The supervised learning method is used by Random Forest, and partly cloudy, Humid and mostly cloudy, Humid and
a well-known machine learning algorithm. It can be used for partly cloudy, Clear, Breezy and overcast, Light rain, Breezy
both classification and regression issues in machine learning. and foggy, Dry and partly cloudy, Windy and Foggy, windy,
Authorized licensed use limited to: SASTRA. Downloaded on September 25,2025 at 07:15:40 UTC from IEEE Xplore. Restrictions apply.
Drizzle, Dry, Windy and partly cloudy, Breezy, Humid and regression and MLP classifiers) with the test dataset and final
overcast, Windy and overcast, Dry and mostly cloudy, Windy predicted value. It shows only the selected labels which are
and mostly cloudy, Rain, Dangerously windy and partly filtered from the first phase which gives us the accuracy of
cloudy, Breezy and dry, Windy and dry. each analysis part. Comparing the four analysis part, Part 3
and 4 gives a better accuracy with 51% that is shown under
Random forest, Gaussian Naïve, and MLP Classifiers.
B. Pre-processing
The pre-processing of data is the first phase in the
process's commencement. After acquiring the dataset, the D. Accuracy Table
first step to do is pre-processing. Because the dataset will be
TABLE 1 ACCURACY TABLE
comprised of data gathered from multiple sources, that can be
incomplete, inconsistent, or inaccurate. Thus, pre-processing Analysis part Algorithms Accuracy
has a great role. Once the dataset is ready it must be put in
1 Random Forest + Linear 0.2354
CSV file format. As we use python, it has many libraries for Regression
pre-processing. Two libraries used here are NumPy and 2 Decision Tree + Linear 0.2163
Pandas. NumPy is a Python library that allows you to perform Regression
scientific calculations. Using this we can also add large 3 Random Forest + MLP 0.5138
multidimensional arrays and matrices to our codes. Pandas is Classifier
4 Gaussian Naïve + MLP 0.5103
a data manipulation library written in Python that is open- Classifier
source. It's a powerful platform for importing and managing
datasets. During data pre-processing, it's vital to find and
handle missing values correctly. We can remove a feature
with a null value for a specific row or a column with more
V. CONCLUSION AND FUTURE ENHANCEMENT
than 75% of the entries missing. Since we only require
numbers, another option utilized is encoding category data in
the equation, which can generate some complications. As a
result, we'll turn it into numerical values. The next stage in We performed hybrid comparative research in which we used
data pre-processing is to split the dataset. The data should be machine learning approaches to offer weather forecasts in this
split up into two: training and testing. publication. Intelligent models can be created using machine
learning technologies that are far simpler than traditional
physical models. They need limited resources and maybe run
C. Proposed Model on nearly any computer, including mobile devices. The
This is an analysis paper using the weather dataset. We Random Forest and Gaussian Nave with MLP Classifier
have a hybrid model, that has four analysis parts consisting models predict weather features more correctly than the other
of two phases of machine learning algorithms each. 70% of hybrid models described here, according to our evaluation
the original dataset is split into training data and 30% for test results. To forecast the weather in a specific place, we also
data. analyze past data from surrounding locations. We show that
i. The first analysis part consists of the Random focusing primarily on the location where weather forecasting
Forest algorithm and Linear Regression. is done is ineffective. The accuracy of a dataset can be
ii. The second analysis part consists of a Decision improved by focusing on only two or three features.
Tree and Linear Regression.
iii. The third analysis part consists of Random
Forest and MLP Classifier. VI. GRAPHICAL REPRESENTATION
iv. The fourth analysis part consists of Gaussian
Naïve and MLP Classifiers.
Initially, the feature importance score of each feature is
calculated using the machine learning algorithms coming in
the first phase of every analysis part (i.e., Random forest,
Decision tree, and Gaussian Naïve respectively). The
strategies that determine a score for all of the input features
for a particular model are referred to as feature importance.
The 'importance' of each characteristic is simply represented
by the score. We are giving an input threshold frequency
concerning the importance score of each feature in the first
phase. The only features which are greater than the given
input threshold frequency will be taken. These selected
features will display the actual weather and final predicted
weather which is received after training with the respective Fig. 1. Graphical Representation of accuracy obtained
algorithms.
In the second phase of every analysis part, a confusion matrix
is created using the machine learning algorithms (Linear
Authorized licensed use limited to: SASTRA. Downloaded on September 25,2025 at 07:15:40 UTC from IEEE Xplore. Restrictions apply.
VII. REFERENCES [14] Rao, K. Hanumantha, et al. "Implementation of
anomaly detection technique using machine
[1] Aher, Sunita B., and L. M. R. J. Lobo. "Comparative learning algorithms." International journal of
study of classification algorithms." International computer science and telecommunications 2.3
Journal of Information Technology 5.2 (2012): 239- (2011): 25-31.
243. [15] Rich Caruana and Alexandru Niculescu-Mizil.
[2] Bellaachia, Abdelghani, and Erhan Guven. 2006. An empirical comparison of supervised
"Predicting breast cancer survivability using data learning algorithms. In Proceedings of the 23rd
mining techniques." Age 58.13 (2006): 10-110. international conference on Machine learning.
[3] Devasena, C. Lakshmi. "Comparative analysis of Association for Computing Machinery, New York,
random forest, REP tree and J48 classifiers for credit NY, USA, 161–168.
risk prediction." International Journal of Computer [16] Sharma, Aman Kumar, and Suruchi Sahni. "A
Applications (2014): 0975-8887. comparative study of classification algorithms for
[4] Grover, Aditya; Kapoor, Ashish; Horvitz, Eric spam email data analysis." International Journal on
(2015). [ACM Press the 21th ACM SIGKDD Computer Science and Engineering 3.5 (2011):
International Conference - Sydney, NSW, Australia 1890-1895.
(2015.08.10-2015.08.13)] Proceedings of the 21th [17] Singh, Nitin; Chaturvedi, Saurabh; Akhter, Shamim
ACM SIGKDD International Conference on (2019). [IEEE 2019 International Conference on
Knowledge Discovery and Data Mining - KDD '15 Signal Processing and Communication (ICSC) -
- A Deep Hybrid Model for Weather Forecasting. , NOIDA, India (2019.3.7-2019.3.9)] 2019
(), 379–386. International Conference on Signal Processing and
[5] Holmstrom, Mark, Dylan Liu, and Christopher Vo. Communication (ICSC) - Weather Forecasting
"Machine learning applied to weather Using Machine Learning Algorithm.
forecasting." Meteorol. Appl (2016): 1-5. [18] Singh, Shashank & Faraz, Ahmed & Nagrami, &
[6] Kalmegh, Sushilkumar. "Analysis of weka data Pillai, Aditya. (2020). WEATHER PREDICTION
mining algorithm reptree, simple cart and BY USING MACHINE LEARNING.
randomtree for classification of indian [19] T R, Prajwala. (2015). A Comparative Study on
news." International Journal of Innovative Science, Decision Tree and Random Forest Using R Tool.
Engineering & Technology 2.2 (2015): 438-446. IJARCCE. 196-199.
[7] Kharche, Deepali, K. Rajeswari, and Deepa Abin. 10.17148/IJARCCE.2015.4142.
"Comparison of different datasets using various [20] Vladimir M. Krasnopolsky; Michael S. Fox-
classification techniques with weka." International Rabinovitz (2006). Complex hybrid models
Journal of Computer Science and Mobile combining deterministic and machine learning
Computing 3.4 (2014): 389-393. components for numerical climate modeling and
[8] Kulkarni, Vrushali Y. and Pradeep K. Sinha. weather prediction. 19(2), 122–134.
“Effective Learning and Classification using
Random Forest Algorithm.” (2014).
[9] Latinne, P., Debeir, O., & Decaestecker, C.
(2001). Limiting the Number of Trees in Random
Forests. Lecture Notes in Computer Science, 178–
187
[10] Mishra, Ajay Kumar, and Bikram Kesari Ratha.
"Study of random tree and random forest data
mining algorithms for microarray data
analysis." International Journal on Advanced
Electrical and Computer Engineering 3.4 (2016): 5-
7.
[11] Nalluri, Sravani; Ramasubbareddy, Somula;
Kannayaram, G (2019). Weather Prediction Using
Clustering Strategies in Machine Learning. Journal
of Computational and Theoretical Nanoscience,
16(5), 1977–1981
[12] Pourdarab, Sanaz, Ahmad Nadali, and Hamid
Eslami Nosratabadi. "A hybrid method for credit
risk assessment of bank customers." International
Journal of Trade, Economics and Finance 2.2
(2011): 125-131.
[13] Radhika, Y., and M. Shashi. "Atmospheric
temperature prediction using support vector
machines." International journal of computer
theory and engineering 1.1 (2009): 55.
Authorized licensed use limited to: SASTRA. Downloaded on September 25,2025 at 07:15:40 UTC from IEEE Xplore. Restrictions apply.