0% found this document useful (0 votes)
12 views10 pages

Springer Format123

This document discusses the use of machine learning techniques for predicting crop yields based on environmental factors and historical data. Various algorithms, particularly ensemble methods like Bagging Regressor, have shown high accuracy in forecasting yields, with significant predictors identified as rainfall, temperature, and soil fertility. The research emphasizes the potential of data-driven decision-making in enhancing agricultural productivity and sustainability.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
12 views10 pages

Springer Format123

This document discusses the use of machine learning techniques for predicting crop yields based on environmental factors and historical data. Various algorithms, particularly ensemble methods like Bagging Regressor, have shown high accuracy in forecasting yields, with significant predictors identified as rainfall, temperature, and soil fertility. The research emphasizes the potential of data-driven decision-making in enhancing agricultural productivity and sustainability.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Crop Prediction Based on Characteristics of the

Agricultural Environment Using Various Feature


Selection Techniques and Classifiers

S. P. RAJA, BARBARA SAWICKA, ZORAN STAMENKOVIC AND

G. MARIAMMAL
1
School of Computer Science and Engineering, Vellore Institute of Tech-
nology, Vellore, Tamil Nadu 632014, India
2
Department of Plant Production Technology and Commodities Science,
University of Life Sciences in Lublin, 20-950 Lublin, Poland

3 IHP—Leibniz-Institut für innovative Mikroelektronik, 15236 Frankfurt (Oder),


Germany

4 Department of Computer Science and Engineering, Kalasalingam Academy of


Research and Education, Tamil Nadu 626126, India

Abstract: Accurate prediction of crop yield is critical for agricultural plan-


ning and food security. This system explores the application of machine learn-
ing techniques to forecast crop yields based on historical data and environ-
mental factors. We utilize various algorithms, including linear regression, de-
cision trees, random forests, and neural networks, to develop predictive mod-
els, but Bagging Regressor seems to be the best predictive model which pro-
vides better and accurate results. The dataset comprises weather conditions,
soil properties, and crop management practices collected over multiple grow-
ing seasons. Our analysis shows that machine learning models, particularly
ensemble methods and deep learning approaches, outperform traditional sta-
tistical methods in predicting crop yields. Feature importance analysis reveals
that rainfall, temperature, and soil fertility are the most significant predictors.
The models demonstrate high accuracy and generalizability across different
crops and regions, suggesting their potential for large-scale agricultural de-
ployment. This research highlights the promise of Machine learning in en-
hancing agricultural productivity and sustainabilitythrough data-driven deci-
sion-making.

Keywords: Agriculture, Bagging Regressor, Crop Prediction

Crop prediction in agriculture is a complicated process [1] and multiple mod-


els have been proposed and tested to this end. The problem calls for the use
of assorted datasets, given that crop cultivation depends on biotic and abiotic
factors [2]. Biotic factors include those elements of the environment that oc-
cur as a result of the impact of living organisms (microorganisms, plants, an-
imals, parasites, predators, pests), directly or indirectly, on other living organ-
isms. This group also includes anthropogenic factors (fertilization, plant pro-
tection, irrigation, air pollution, water pollution and soils, etc.). These factors
may contribute to the occurrence of many changes in The associate editor
coordinating the review of this manuscript and approving it for publication
was Yongming Li . the yield of crops, cause internal defects, shape defects
and changes in the chemical composition of the plant yield. The shaping of
the environment as well as the growth and quality of plants is influenced by
abiotic and biotic factors. Abiotic factors can be divided into physical, chem-
ical, and other. The recognized physical factors include: mechanical vibra-
tions (vibration, noise), radiation (e.g., ionizing, electromagnetic, ultraviolet,
infrared); climatic conditions (atmospheric pressure, temperature, humidity,
2 S. P. RAJA AND BARBARA SAWICKA
air movements, sunlight); soil type, topography, soil rockiness, atmosphere,
and water chemistry, especially salinity. The chemical factors include: prior-
ity environmental poisons, such as sulfur dioxide and derivatives, PAHs; ni-
trogen oxides and derivatives, fluorine, and its compounds, lead and its com-
pounds, cadmium and its compounds, nitrogen fertilizers, pesticides, carbon
monoxide. The others are: mercury, arsenic, dioxins and furans, asbestos, and
aflatoxins [3]. Abiotic factors also include bedrock, relief, climate, and water
conditions - all of which affect its properties. Soil-forming factors have a di-
versified effect on the formation of soils and their agricultural value [4]. Pre-
dicting crops yields is neither simple nor easy. The methodology for predict-
ing the area under cultivation is, according to Myers et al. [5] and Muriithi
[6], a set of statistical and mathematical techniques useful in an evolving and
improving optimization process. It also has important uses in design, devel-
opment, and formulation new as well as improving existing products. Presen-
tation or performance of statistical analysis requires the possession of numer-
ical data. Based on them, conclusions are drawn as to various phenomena and
further, on this basis, binding economic decisions can be made. According to
Muriithi [6], the better you describe certain phenomena in terms of numbers,
the more you can say about them, and with increasing data accuracy you can
also obtain more accurate information and make more accurate decisions. The
biggest problem in the temperate climate zone is assessment of agroclimatic
factors in terms of shaping the yield of winter plant species, mainly cereals.
The key factor influencing wintering yield, which provides access to days
with a temperature over of 5◦ C, their number and frequency, and the number
of days in the wintering period with temperatures above 0◦C and 5◦C. A num-
ber of these can be estimated on the basis of public statistics and yield regres-
sion statistics in years. Developed models for checking the situation that as-
sess whether they want to be a probation of state policy in the field of inter-
vention in the cereal market. Efficient forecasting of productivity requires
forecasting of agrometeorological factors. Aspects related to the variability of
these factors may pose a particular problem [7]. Many researchers have dealt
with this issue with varying degrees of success [8]–[10]. Grabowska et al. [9]
predicted narrow-leaf lupine yields for 2050-2060 using weather models and
three climate change scenarios for Central Europe: E-GISS model, HadCM3
and GFDL. The fit of the models was assessed by means of the determination
coefficient R2, corrected coefficient of determination R2adj, standard error
of estimation and the coefficient of determination R2pred calculated using the
Cross Validation procedure. The selected equation was used to forecast lupine
yield under the conditions of doubling the CO2 content in the atmosphere.
These authors stated that the influence of meteorological factors on the yield
of narrow-leaved lupine varied depending on the location of the station. The
temperature (maximum, average, minimum) at the beginning of the growing
season, as well as rainfall during the flowering - technical maturity period,
most often had a significant influence on the yield. It has been shown that the
predicted climate changes will have a positive effect on the lupine yield. The
simulated profitability was higher than that observed in 1990-2008, and
HadCM3 was the most favorable scenario. Dąbrowska-Zielińska et al. [8] as-
sessed the usefulness of plant biophysical parameters, calculated from the
ranges of reflected electromagnetic radiation recorded by the new generation
satellites Sentinel-2 and Proba-V, for forecasting crop yields in Poland. In
2016-2018, ground measurements were carried out in arable fields in the area
included in the global crop monitoring network GEO Joint Experiment of
CROP YIELD PREDICTION

3
Crop Assessment and Monitoring JECAM. Classification of crops was per-
formed using optical and radar images Sentinel-1 and RadarSat-2. The PRO-
totypical model of Biomass and Evapotranspiration PRO was used to simulate
the growth of winter wheat cultivation, to forecast its biomass size. Got high
accuracy of 94% of the size of biomass modeled with real biomass. Li et al.
[10] found that accurate, high-resolution yield maps are needed to identify
spatial patterns of yield variability, to identify key factors influencing yield
variability, and to provide detailed management information in precision
farming. Varietal differences may significantly affect the forecasting of po-
tato tuber yields with the use of remote sensing technologies. These authors
argue that improving potato crop forecasting with remote sensing of un-
manned aerial vehicles (UAVs) by incorporating varietal information into
machine learning methods has the best chance at present. There are different
challenges in this research area. Currently, crop prediction [11] models gen-
erate actual results that are satisfactory, though they could perform better.
This paper attempts to propose an enhanced crop prediction model that ad-
dresses these issues. The prediction process [12] depends on the two funda-
mental techniques of feature selection [FS] and classification. Prior to the ap-
plication of FS techniques, sampling techniques are applied to balance an im-
balanced dataset.

1.1 RELATED WORK

A. BASED ON SOIL CONDITIONS

The approach attempts to supplant conventional laboratory approaches in or-


der to eliminate drawbacks such as manual involvement, time consumption,
human error, and uncertain predictions. The signal processing method im-
proved the quality of the original image through the use of filters and by com-
puting the features in the enhanced images. The proposed algorithm uses
color quantization and texture-based feature extraction by applying the Gabor
filter and Laws’ mask. Matching is achieved by applying statistical measure-
ments like the mean, standard deviation, skewness, and kurtosis. You et al.
[15] posited an adaptable and precise technique to anticipate yields by em-
ploying openly accessible remote sensing data. The methodology enhances
existing procedures in three different ways. To begin with, a remote detecting
network is applied to propose a working methodology. Next, a novel dimen-
sionality reduction procedure is presented that uses a convolutional neural
network (CNN) alongside long-term memory. Finally, a Gaussian process is
used to investigate and examine the spatio-transient structure of the data and
enhance its accuracy. Anantha et al. [16] implemented a recommendation sys-
tem using an associate ensemble model with majority voting. The random
tree, Chi-square Automatic Interaction Detection (CHAID), kNN, and Naive
Bayes (NB) are used as learners to help determine the most appropriate crop,
taking into consideration soil parameters, with the results showing high accu-
racy and potency. The classified image generated by these techniques consists
of ground truth-applied mathematics information Further, it incorporates such
data as the parameters of the square measure in terms of the weather and crop
yield, as well as state and district-wise crop produce. All of the above are
employed to predict specific crop yields in a given set of circumstances. Rale
4 S. P. RAJA AND BARBARA SAWICKA
et al. [17] developed a forecasting model which uses the default settings along
with RF regression for crop yield production.

B. BASED ON ENVIRONMENTAL CONDITIONS

Jones et al. [18] modified the Decision Support System for Agrotechnology
Transfer (DSSAT) crop model, using a decision support system algorithm.
However, it is increasingly difficult to sustain DSSAT crop models, given the
different sets of code in operation for different crops. The new design uses a
multi-modular approach, comprising a cropping template as well as soil and
weather modules. Further, there is a module that monitors light and water in
the crops, soil, and environment. Fernando et al. [19] studied data on annual
coconut production from 1971 to 2001 in a particular region and assessed its
economic impact. The research revealed that the loss sustained by the econ-
omy in crop shortage terms was around US $50 million. Ji et al. [20] advanced
an estimation technique to predict rice yields. The study attempted to deter-
mine the effectiveness of artificial neural networks (ANN) in predicting rice
yield in mountainous regions. It assessed the efficacy of the ANN, relative to
biological parametric variations, and compared the efficiency of multiple bi-
linear regression models with the ANN model. Boryan et al. [21] proposed a
decision tree-based technique to depict openly accessible state-level crop
cover groups, in accordance with guidelines laid down by the Cropland Data
Layer (CDL) and National Agricultural Statistics Service (NASS), and utiliz-
ing ground truth collected during the June Agricultural Survey. The proposed
work outlines the NASS CDL program. It presents information dealing with
handling strategies, order and approval, precision evaluation, and CDL item
particulars, and product cost estimation procedure. Hansen and Loveland [22]
proposed the use of Landsat to acquire satellite imagery that facilitates remote
sensing of the environment. Current strategies for monitoring land cover
changes across massive swathes of land commonly utilize Landsat infor-
mation. Bolton and Friedl [23] created a precise model to forecast maize and
soybean yield in the Central United States. Part of their examination included
testing the capacity of the MODIS (Moderate Resolution Imaging Spectrora-
diometer) to catch between-yearly fluctuations in yields. Their outcomes
demonstrate that the MODIS two-band Enhanced Vegetation Index outper-
forms the generally utilized Normalized Difference Vegetation Index in re-
spect to anticipating maize yields. Taking into consideration data using veg-
etation phenology obtained from the MODIS has fundamentally enhanced the
model execution internally as well as crosswise, over the years. Dempewolf
et al. [24] designed and developed a practical wheat yield prediction model
for the Punjab Province of Pakistan. Shannon and Motha [25] examined the
agricultural lands of North America, Central America, and the Caribbean fol-
lowing various weather and climate-related natural disasters. The latest cli-
mate and weather data is needed to help farmers manage agricultural risks.
The study discusses climatic uncertainties in agriculture such as drought,
flood, typhoons, extreme heat, and freezing temperatures. A decision support
system is used to prepare adequately for hazard management prior to the oc-
currence of a disaster. The Agro Climate Research Centre and Agro Meteor-
ological Department play a critical role in agriculture-based risk management
activities. Manjula and Djodiltachoumy [26] analyzed crop yield prediction
data supported by association rules for the chosen region, that is, the district
of Madras in an Asian nation. Eswari and Vinitha [27] employed the Bayesian
CROP YIELD PREDICTION

5
network classification supervised learning model in their proposed approach
Environmental characteristics such as temperature and rainfall are analyzed
alongside crop information to classify crops like rice, coconut, areca nut,
black pepper, and dry ginger. Bayesian network classification is employed to
explore the dataset.

C. SURVEY OF MACHINE LEARNING TECHNIQUES FOR CROP


PREDICTION

Shivnath and Santanu [28] devised a machine learning approach to examine soil
fertility and plant nutrient management. The backpropagation network (BPN) used
is trained with inputs on crop growth characteristics, nutrient reserves in the soil,
and external applications for crop production. The ML system follows the 3 steps
of sampling (different soils with similar properties and completely different param-
eters), backpropagation, and weight change. Paul et al. [29] designed a system that
uses data processing techniques to foretell the class of the soil datasets analyzed in
terms of crop yields. The process of prognosticating crop yields is formalized as a
sorting rule, using NB and kNN clustering. Pudumalar et al. [30] devised an exact-
ness agriculture approach, which is a smart farming technique that uses information
on soil property, soil type, and crop yields to help farmers determine the most ap-
propriate cultivable crops based on soil parameters. A new ensemble model using
the random tree, CHAID, kNN, and NB is proposed to recommend crops for a
specific land area. Bodake et al. [31] developed a soil-based fertilizer guidance
system that facilitates topical soil examination to help farmers cultivate the right
crop. The tool is intended to be made available in the local language so farmers
experience no difficulty in comprehension. Heupel et al. [32] proposed an unsuper-
vised fuzzy classification approach that suggests crop types with produce harvested
in early spring. The classification results are expected to improve with time. Liu et
al. [33] investigated the probability of implementing multi-temporal Sentinel-2 sat-
ellite images to discern heavy metal-induced stress (i.e., Cd stress) in rice crops in
four study areas in Zhuzhou City of Hunan Province in China. Priya et al. [34]
advocated a crop yield prediction approach using the RF rule. Real-time infor-
mation from Tamil Nadu state in India was used to develop the models, which were
tested on several samples. The predictions generated help farmers forecast crop
yields prior to cultivation. Archana et al. [35] proposed an ontology-based recom-
mendation system for crop quality and fertilizer use, successfully bridging the gap
between the farming community and technological applications. The system pre-
dicts a relevant crop, taking into consideration the geographical area and soil type,
and offers guidance on appropriate fertilizer use. The recommendation system uses
the RF rule and k-means clump rule. Brogi et al. [36] proposed a superintended
classification methodology to classify the Essential Commodities Act (ECA) and
map regions on the basis of similar soil characteristics. Ali Al-Naji et al. [39] pro-
posed method which is referred as a non-contact vision system based on a standard
video camera to solve the irrigation-based problems in the agriculture. The authors
have used the feedforward back propagation neural network to analyze/irrigate the
soil which is captured at various times, distances and illumination levels.

1.2 PREPROCESSING
6 S. P. RAJA AND BARBARA SAWICKA
Sampling techniques are applied during preprocessing to balance the dataset
and maximize the prediction performance [37]. The sampling techniques used
include the ROSE, SMOTE and MWMOTE. ROSE is used for binary classi-
fication in the presence of rare classes and SMOTE for better classifier per-
formance in the ROC space, while MWMOTE handles imbalanced dataset
issues in crop prediction.

1.3 REGRESSION TECHNIQUES

A. DECISION TREE

A decision tree is a flowchart-like tree structure generally used in supervised


machine learning for classification and prediction. A DT can be turned into a
set of rules where each path, heading from the root node to every leaf node,
is a rule. In a decision tree, every internal node represents a test/condition or
an attribute, every branch is a result of the test, and every leaf node has a class
ascribed to it that is reachable if the attribute fulfils the condition of the branch
leading to it. A famous example of a decision tree is the C4.5 by Ross Quin-
lan. There are, broadly, two types of decision trees, categorical and continu-
ous variable, based on target attribute types. A decision tree starts with a root
node that is compared with other attributes/features in the dataset for a perfect
split. A perfect split implies that the number of outputs of one class are on
one side of the tree and those of the other class on the other. In this way, every
node gets split until it reaches a perfect split, the outcome of which becomes
the leaf node of a tree. The real challenge in constructing a DT is attribute
selection. That is, given the large number of attributes available, it is difficult
to select the ones to be used as root nodes or internal nodes. To this end, there
are two techniques that can be applied, Information Gain and Gini Index:

Information Gain (T, X) = Entropy (T) – Entropy (T, X)

, where T refers to the current state and X to the selected attribute;

Gini index = 1 − 6(p)∧ 2 = 1 − [(p+) ∧ 2 + (p−) ∧ 2]

, (2) where p+ represents the probability of Yes/Good and p- the probability


of No/Bad.

B. K-NEAREST NEIGHBOR (KNN)

One of the most commonly used supervised and nonparametric machine


learning techniques is the kNN, used in classification and regression prob-
lems. Supervised algorithms are the ones with labelled data. Labeled data re-
fers to input data that is already tagged with the correct output in supervised
learning. Supervised learning algorithms take the data and attempt to make
models that predict the output data, given relevant inputs. There is, however,
no actual learning in the k-NN algorithm, which follows the ‘‘lazy learning’’
principle, where all the work happens at the time a prediction is required. The
algorithm depends on distances between points, which can be ascertained us-
ing one of a few methods. A key aspect for consideration is that the distance
is always required to be either zero or positive. This is done by squaring the
CROP YIELD PREDICTION

7
distance or raising it to a certain power or taking the absolute values. Methods
to find distances include the following: Manhattan Distance This distance is
easier to calculate than the others.

|x2 − x1| + |y2 − y1|

(3) (i) Euclidean Distance This is the distance between two points, used in
regular geometry.
((x2 − x1) 2 + (y2 − y1) 2 ) 1 /2

(4) (ii) Hamming Distance This method finds distances by depending on com-
mon neighbors.

|x1 − y1 | + |x2 − y2|

If x1 and y1 are of the same type, their difference is 0, else it is 1. (iii) Min-
kowski Distance Similar to the Euclidean distance, an ‘‘n’’ value is needed
here,

((x2 − x1) p + (y2 − y1) p ) 1 /p

where xi and yi are the x and y coordinates of a point on an xy plane. The k-


nearest neighbors (KNN) method is a supervised machine learning algorithm
used to solve classification and regression problems. kNN works on the prem-
ise of similar entities existing in close proximity. Related data fields would
therefore occur nearby. This helps us in mapping similarities between datasets
and a given query. Before we implement the kNN method, all the labeled data
must be pre-processed. Firstly, all the data must be normalized. Next, feature
selection must be employed to delete the irrelevant features as kNN doesn’t
work well when too many features are present. Missing values cannot be tol-
erated, thus in the case of missing values that particular row must be deleted.
Hereon, we can move towards the implementation of the kNN algorithm.
First, data is loaded into the model. kNN being a supervised learning method
requires data to be loaded in labeled form. Next, K is declared according to
the desired number of neighbours. Then, for every element in the dataset, the
‘‘distance’’ or ‘‘relation’’ with the query input is calculated by the machine
learning algorithm. The distance between the element and the query input is
then added to an ordered collection and is subsequently sorted in increasing
order of the distances. Lastly, the first K items of the collection are selected
and the output, depending upon the model being a regression or a classifica-
tion problem, is returned by taking a mean, in case of the former, and taking
the mode in case of the latter. Choosing k is an important factor as it heavily
influences the result of our ML model. If the value of K is too low, the model
suffers from instability and the results become increasingly inaccurate. Con-
versely, an extremely high value of k will start furnishing an increased num-
ber of errors in the model. Therefore, the value of k must be balanced between
the two extremums. In the case of a model where a vote is required to get the
output, K should be taken as an odd number to ensure a deciding game. The
chief advantages of the KNN algorithm are, firstly, it is simple and relatively
easy to implement. Next, the algorithm can serve multiple purposes, right
from classification and regression to searching problems. Furthermore, the
8 S. P. RAJA AND BARBARA SAWICKA
algorithm can be improved by adding additional training data. The main dis-
advantage of KNN is that the speed of the algorithm goes on decreasing as
our dataset becomes larger and larger as the cost of computation keeps in-
creasing. Therefore, it is not suitable in cases where immediate results are
required. Secondly, we must accurately determine the value of k to get appro-
priate results. This process can be difficult sometimes. Also, the KNN algo-
rithm requires a large amount of memory to store large sets of data. The KNN
method can be used for predicting crop yield using a set of known parameters,
namely, rainfall, temperature, humidity, and soil moisture. The value of crop
yield is calculated by using the values of the nearest neighbors. KNN has
yielded suitable accuracy in predicting crop yield. This model can be further
enhanced by adding additional features and more data from all the seasons.
KNN has also been applied for predictive analysis of paddy production and
has shown better and faster results compared to the SVM algorithm.

C. BAGGING

Bagging, also known as bootstrap aggregation, is used with decision trees,


where it significantly raises the stability of models in terms of advancing ac-
curacy and diminishing variance so as to eliminate overfitting. In ensemble
machine learning, bagging takes numerous weak models specializing in dis-
tinct sections of the feature space and aggregates them to pick the best pre-
diction. An ensemble set of learners is developed and built utilizing the learn-
ing algorithm, but with each learner instructed on a different set of data. Such
a process is termed bootstrap aggregating or bagging. Initially, numerous sub-
sets of the data are created, with each being a subset of the initial data. So, a
subset containing n’’ values have an original dataset comprising n different
instances. n’’ of these are grabbed at random with replacements from the ini-
tial data. An n’’, randomly grabbed and put in the bag, is chosen from the
entire data collection. It is picked again at random and bagged, implying a
certain degree of repetition, resulting in some recurrence. Such recurrence,
however, creates no problems and is to be anticipated. So then, m groups or
bags are created altogether, with each holding n’’ different data instances,
randomly picked, with replacements. Thus, n is the number of training in-
stances in the original data, n’’ the number of instances in individual bags,
and m the number of bags. We nearly always want n’’ to be less than n, usu-
ally by about 60%. Therefore, each of these bags has, as a rule of thumb, about
60% as many training instances as the original data. Each of the data groups
is used to practice a different model. There are m different models, each one
practiced on different data, producing an ensemble of different learning algo-
rithms, along with an ensemble of models to be queried identically. Each
model is queried with the equivalent x and all of their output collected. The y
output of specific models is taken with their mean to generate the y for the
ensemble. For example, assuming that there are L bootstrap samples, each of
size b, then

z 1 1 ,z 1 2 , . . . ,z 1 B , z 2 1 ,z 2 2 , . . . ,z 2 B , . . . , z L 1 ,z L 2 , . . . ,z L

B Whereas, assuming z l b ≡ b-th observation of the l-th bootstrap sample, L


number of nearly independent weak learners can be fitted in w1(.), w2(.), . . .
,wL(.).
CROP YIELD PREDICTION

D. RANDOM FOREST (RF)

The random forest (RF) is one of the most successful supervised machine
learning algorithms. The RF algorithm embodies the essence of ensemble
learning in that it links multiple classifiers to resolve a complicated problem,
thereby enhancing the performance of the model. In this method, the ‘‘forest’’
that is built is a set of decision trees. Characteristics in the RF are randomly
picked in each decision split. The correlation between trees is diminished by
randomly picking features that promote prediction and result in greater effi-
ciency. Random Forest is an ML classification algorithm that works by divid-
ing the dataset into subsets or decision trees and aggregating the outputs of
all the trees to produce the final output. Random Forest comes under the Bag-
ging category of ensemble learning techniques. The Row and Feature samples
from the main dataset are randomly selected and fed into the decision trees in
the Random Forest Technique. The analyst chooses the number of decision
trees for the model. Each decision tree works on the data and predicts a result
based upon its calculation. Random Forest doesn’t take the result from any
one of the decision trees but combines the outputs from all the decision trees.
Random Forest takes the majority of the result (in case the result is in a Bool-
ean form) or the mean/median of the result (in case the result is in numerical
form). Thus, a higher number of decision trees gives a more accurate result
and circumvents the problem of overfitting. The Random Forest technique
provides several advantages. Firstly, Random Forest is simple and relatively
easy to understand and is therefore extremely popular. It is also capable of
performing both classification and regression tasks. It is also suitable for han-
dling large sets of data with high dimensionality and most importantly, it
makes the model much more precise and resolves the overfitting issue. Ran-
dom Forest cannot be used in case of extrapolation of data as it could produce
inaccurate results. Although Random Forest can be used for both regression
and classification, it is better suited for classification tasks. Also, it does not
produce proper results when dealing with sparse data. Random Forest also
needs more time for implementation and requires larger data and greater re-
sources. In the presence of correlated predictors, Random Forest is known to
produce inexact results. Given that each bagged tree is identically dissemi-
nated, the expectation of an average of B trees is the equivalent of the expec-
tation of each. Since this accounts for the bias of bagged trees being the same
as that of individual trees, a change may only be affected through variance
reduction. This contrasts with advancing, where the trees are grown adap-
tively to exclude bias, and hence are not identically distributed. An average
of B identically distributed random variables has a variance of σ 2 . If the
variables are completely identically distributed, but with positive pairwise
correlation ρ, the variance of the average is given as

ρσ2 + 1 − ρ B σ 2

(5) It is observed that as B increases, the value of the second term shifts neg-
ligibly while that of the first term remains unchanged. Consequently, the size
of the correlation of the bagged trees limits the benefits of averaging. The RF
focuses on bagging variance minimization by cutting the correlation between
10 S. P. RAJA AND BARBARA SAWICKA
the trees without increasing the variance excessively. The tree-growing pro-
cess makes this procedure possible through picking input variables at random.

F. YIELD PREDICTION REPRESENTATION USING GRADIUM

Crop yield prediction using Gradio is an innovative approach that leverages


the power of artificial intelligence to forecast agricultural productivity. By
integrating Gradio's interactive interface with advanced machine learning al-
gorithms, farmers and researchers can input various factors such as weather
conditions, soil quality, and crop management practices to generate accurate
predictions. This enables them to make informed decisions regarding crop
selection, resource allocation, and yield optimization. With Gradio's user-
friendly interface, crop yield prediction becomes accessible and beneficial for
a wide range of users in the agricultural sector, ultimately contributing to
more sustainable and efficient farming practices.

1.4. CONCLUSION

Predicting crops for cultivation in agriculture is a difficult task. This paper has used a
range of feature selection and classification techniques to predict yield size of plant
cultivations. The results depict that an ensemble technique offers better prediction ac-
curacy than the existing classification technique. Forecasting the area of cereals, po-
tatoes and other energy crops can be used to plan the structure of their sowing, both
on the farm and country scale. The use of modern forecasting techniques can bring
measurable financial [Link]. A third level heading in 9-point
font size at the end of the paper is used for general acknowledgments, for example:
This study was funded by X (grant number Y).

References
1. R. Jahan, ‘‘Applying naive Bayes classification technique for classification of
improved agricultural land soils,’’ Int. J. Res. Appl. Sci. Eng. Technol.,
vol. 6, no. 5, pp. 189–193, May 2018.

2. B. B. Sawicka and B. Krochmal-Marczak, ‘‘Biotic components influencing the


yield and quality of potato tubers,’’ Herbalism, vol. 1, no. 3, pp. 125–136, 2017.

3. B. Sawicka, A. H. Noaema, and A. Gáowacka, ‘‘The predicting the size of the


potato acreage as a raw material for bioethanol production,’’ in Alternative En-
ergy Sources, B. Zdunek, M. Olszáwka, Eds. Lublin, Poland: Wydawnictwo
Naukowe TYGIEL, 2016, pp. 158–172.

4. B. Sawicka, A. H. Noaema, T. S. Hameed, and B. Krochmal-Marczak, ‘‘Biotic


and abiotic factors influencing on the environment and growth of plants,’’ (in
Polish), in Proc. Bioróżnorodność Środowiska Znaczenie, Problemy,
Wyzwania. Materiały Konferencyjne, Puławy, May 2017. [Online]. Available:
[Link]

You might also like