Sugarcane Yield Using SML
Sugarcane Yield Using SML
Received: 31 January 2019; Revised: 02 March 2019; Accepted: 20 March 2019; Published: 08 August 2019
Abstract—Agriculture is the most important sector in the agriculture sector for yield prediction/crop forecasting is
Indian economy and contributes 18% of Gross Domestic limited to empirical methods using ground-based
Product (GDP). India is the second largest producer of observations and productions reports gathered by various
sugarcane crop and produces about 20% of the world's organizations from different sources: meteorological data,
sugarcane. In this paper, a novel approach to sugarcane agro-meteorological(yield), soil (water holding capacity),
yield forecasting in Karnataka(India) region using Long- and remotely sensed agricultural statistics. Based on
Term-Time-Series (LTTS), Weather-and-soil attributes, meteorological and agronomic data, several indices are
Normalized Vegetation Index(NDVI) and Supervised derived which are deemed to be relevant variables in
machine learning(SML) algorithms have been proposed. determining crop yield. For instance, crop water
Sugarcane Cultivation Life Cycle (SCLC) in satisfaction, surplus and excess moisture, average soil
Karnataka(India) region is about 12 months, with moisture. As Crop production rate depends on the
plantation beginning at three different seasons. Our geography of a region(e.g. hill area, river ground, depth
approach divides yield forecasting into three stages, region, etc), weather condition (e.g. temperature, cloud,
i)soil-and-weather attributes are predicted for the duration rainfall, humidity etc), soil type (e.g. sandy, salty, clay,
of SCLC, ii)NDVI is predicted using Support Vector peaty, saline soil etc), soil composition (e.g. PH value,
Machine Regression (SVR) algorithm by considering nitrogen, phosphate, potassium, organic carbon, calcium,
soil-and-weather attributes as input, iii)sugarcane crop is magnesium, sulphur, manganese, copper, iron etc) and
predicted using SVR by considering NDVI as input. Our harvesting methods, various combinations of subsets of
approach has been verified using historical dataset and these influencing parameters have been used by different
results have shown that our approach has successfully prediction models for crop yield prediction using ground-
modeled soil and weather attributes prediction as 24 steps based observations[3-7]. Prediction models are broadly
LTTS with accuracy of 85.24% for Soil Temperature classified into two types: i)traditional statistics model (e.g.
given by Lasso algorithm, 85.372% accuracy for multiple linear regression model), which formulates a
Temperature given by Naive-Bayes algorithm, accuracy single predictive function holding entire sample space. i.e.
for Soil Moisture is 77.46% given by Naive-Bayes, it generates a global model over the entire sample space.
NDVI prediction with accuracy of 89.97% given by ii)machine learning technique, which is emerging
SVR-RBF, crop prediction with accuracy of 83.49% technology for knowledge mining that relates input and
given by SVR-RBF. output variables which is hard to obtain statistically. In
traditional statistical methods, the structure of the data
Index Terms—Agriculture, NDVI, Machine Learning, model needs to be assumed priory, whereas machine
Support Vector Regression, Crop Prediction. learning techniques need not assume this structure. This
is a useful characteristic for machine learning techniques
to model complex, non-linear behavior in crop yield
I. INTRODUCTION prediction. Recent researches in the field of yield
prediction have focused on Remote Sensing(RS) data and
Long Term Time Series (LTTS) forecasting has been a
Machine Learning(ML) techniques, as RS data is cost
useful tool for governments, planning commissions and effective compared to ground observation data. RS data,
decision makers in various applications such as solar such as Special Vegetation Indices (SVIs) derived using
energy, wind power energy, economic forecasting, and in
multi hyper-spectral calibrated data. The one among is
the agriculture sector. Historically, LTTS has been
the Normalized Difference Vegetation Index(NDVI), a
applied at the regional and national level for planning, type of SVIs, used with machine learning parametric
import and export decision making and policy algorithms like Linear Regressions, Naïve Bayes,
decisions[1,2]. Traditionally, application of LTTS in the
This work is open access and licensed under the Creative Commons CC BY 4.0 License. Volume 11 (2019), Issue 8
Sugarcane Crop Yield Forecasting Model Using Supervised Machine Learning
Gaussian Process and Simple Neural Networks, as well as using Lasso Regression with an accuracy of about
non-parametric algorithms like Support Vector 80%[15]. FASAL combines conventional methods of
Machine(SVM), Decision Trees(DT) and Nearest forecasting with remotely sensed data to make multiple
Neighbors to predict yield successfully. Parametric ML in-season yield predictions. Applications of predictive
algorithms are derived from traditional statistical methods. empirical models using remotely sensed data in crop
The latest development in yield prediction being an yield prediction have been popular, and successful in
application of Deep Learning(DL) techniques using RS predicting yield efficiently and quantitatively. Previous
dataset[8,9]. NDVI is calculated as normalization of studies have established that NDVI data which is derived
reflectance values from Near Infrared(NIR) and red from satellite images, normally used for monitoring
brands ranging from -1.0 to +1.0. Higher NDVI is an vegetation health and changes in growth patterns, could
indication of greener surface and lower NDVI indicates be used for in-season yield prediction in a larger region.
less green surface[10]. Crop yield has been predicted not Many researchers have also established a relation
only at the beginning of the sowing phase but also at between various weather parameters, such as surface
various intervals and phases during crop cultivation temperature, precipitation, soil moisture and cloud cover
season depending on purpose and organizations. with NDVI values. Past research studies in soil
As defined by the Food and Agriculture of the United parameters, weather, NDVI and yield predictions have
Nations, crop forecasting is the art of predicting crop shown a strong relationship between these four types of
yields and production before the harvest actually takes parameters. Although previous studies have successfully
place, typically a couple of months in advance[11]. applied various machine learning algorithms to crop yield
Defining time horizon for crop yield forecasting in terms prediction considering NDVI, weather parameters and
of time series forecasting methodologies is an important soil parameters as attributes, inherent dependencies of
aspect. In common research practice, forecasting horizons these attributes requires careful selection of these
are categorized as short-term, medium-term and long- attributes to address curse of dimension problem in
term. Short-term forecasting horizons are closer to the machine learning.
end of observation time period, and long-term forecasting
horizons are far from the end observation time period [12,
13]. In other words, single step forecasting is short-term, II. RELATED WORKS
two steps forecasting is a medium-term and more than
two steps forecasting is long-term forecasting, where a Crop yield forecasting is a unified, bio-socio-system
step is time unit of observations, like a month in monthly comprised of a complex interaction among soil, air, water,
and crops grown in it, where a comprehensive model is
observed time series. In many applications of time series,
required. Crop yield forecasting models could be
the boundary for categorization of time series forecasting
may not be defined clearly. Total yield being a one-time categorized based on attribute measurement methods,
outcome in season, time horizons for yield prediction such as ground-based observed data, remotely sensed
data, and a combination of both ground-based and remote
needs to be defined with different parameters to be
sensed. Another way to categorize yield forecasting
predicted for meaningful results. Yield prediction could
be useful in achieving maximum yield rate of the crop models could be as classical empirical models and
using the limited land resource as part of agricultural machine learning models[1-7]. Researchers have been
using periodical, cost-effective and comprehensive
planning in an agro-based country [10]. Antecedent
remote sensed data, which provides information about
determination of problems associated with crop yield
indicators can help to increase the yield rate of crops. earth surface for yield prediction. Two approaches have
Crop selector could be applicable to minimizing losses been used to obtain a quantitative relationship between
remotely sensed data and crop yield. Studies in the first
when unfavorable conditions may occur and this selector
group of approaches incorporate remotely sensed data
could be used to maximize crop yield rate when the
potential exists for favorable growing conditions. into agro-meteorological models such as SAFY and Aqua
Maximizing production rate of the crop is an interesting Crop [16]. These approaches predict yield accurately and
model crop development, but requires large amount and
research field to agro-meteorologists which play a
complicated field inputs like water balance models,
significant role in the national economy[14]. In India,
farmers' conditions are worsening day by day, even fertilizers, etc. derived from remotely sensed data. The
during favorable condition as well as unfavorable second group of classical approaches to yield forecasting
is based on empirical relationships like regression.
condition. During favorable condition, with bumper yield,
Advanced ML techniques[17], Adaptive Neuro Fuzzy-
farmers are getting less price because of surplus yield
than demand. During unfavorable condition, because of Inference Systems [18], Multi-Layer Perceptron (MLP),
loss of crop. Crop yield prediction could be useful in Artificial Neural Network (ANN) [19], Bayes Net, XY-
Fused Networks, Supervised Kohonen Network, Counter
selecting crop to minimize loss to farmers'.
Propagation ANN, have been applied in yield prediction
Forecasting Agricultural output using Space, Agro-
metrological and Land based observations (FASAL)[15], for yields like rice, wheat using various indices like
forecasts multiple in-season yields, namely pre-season, NDVI, SVI and LAI [20, 21].
India is the second largest sugarcane production
early-season and mid-season. Recent developments in
country in the world. Sugarcane is cultivated across all
FASAL shows that it is able to predict statewide yield
states in India and across different seasons. India has
most of its sugarcane cultivation located on the sub- the various combination of attributes using correlation
tropical belt, Uttaranchal, Bihar, Uttar Pradesh, Punjab, matrix and feature selection techniques.
and Haryana are important sugarcane growing states in This paper is arranged as follows. Section III describes
the Indian sub-tropical region. Sugarcane is also grown in sugarcane crop cultivation in India, Section IV outlines
a few minor regions and pockets of Madhya Pradesh, dataset used for yield prediction and dataset pre-
Rajasthan, Assam, and West Bengal, but these states processing. In section V, modeling time series dataset as
have a low throughput when compared to the sub-tropical a supervised ML problem and implementation of long
and tropical belts. The growth of sugarcane is massive term time series yield prediction model are explained, in
and expansive in the tropical belt and states such as section VI, we have discussed experimental results and
Maharashtra, Tamil Nadu, Andhra Pradesh, Karnataka, analyzed the impact of the curse of dimensionality on
and Gujarat, as sugarcane is a tropical crop, thus all the sugarcane yield prediction outcome using SVR and other
agro-climatic conditions for the cultivation of sugarcane popular supervised ML algorithms, and section VII
are met at these states. Growth Cycle of sugarcane yield concludes this paper and lists the scope for future
plays a very important role in crop yield prediction. So, in enhancements.
this section, we discuss the conditions favorable for
sugarcane growth along with the crop cycle. Depending
upon the sowing time and variety of the crop, it takes III. SUGARCANE CROP CULTIVATION IN INDIA
about 12 to 18 months for sugarcane crop to mature.
Fig.1 shows the Gantt chart for sugar cultivation in
Generally, the months of January to March are considered
India.
for sowing, and harvesting is done from December to
March. Once harvesting is completed, a ratoon crop is
cultivated from the re-growth. A Ratoon Crop is the new
crop which is cultivated using the stubble left behind
from the previous harvest. In India, it is a common
practice to take one ratoon after a normally planted crop.
In a few countries, 2-6 ratoon crops are allowed[22].
Sugarcane demands high water and high nutrient
consumption for a long duration. The range of climatic
conditions where sugarcane will be grown is wide,
sugarcane is grown ranging from sub-tropical to tropical
conditions. Temperatures below 20oC and above 50oC
are not suitable for sugarcane growth. For optimum
productivity, the requirement of 750-1200mm of rainfall
needs to be satisfied during the growth period. Well-
drained alluvial to medium black cotton soils with neutral
pH (6.0-7.0) and optimum depth (>60 cm) are good for
sugarcane growth. Optimum productivity is also being Fig.1. Gantt chart for sugar cultivation
obtained in sandy to sandy-loam soils with near neutral
pH under assured irrigated conditions of North India[22- As sugarcane is cultivated by planting in either
27]. Sugarcane planting could be done in three seasons January-February, July-August or October-November,
namely, Spring, Winter, and Adsali. Spring planting is with maturity duration of 12-18 months, it is challenging
also called as "Suru", is done during January to February, task for predicting crop yield, as time series problem
Winter planting, also called as "Pre-Seasonal" planting, is across various states of India. In Karnataka and
done during October to November, and Adsali planting is Maharashtra states, sugarcane variety with 12 months
done during July to August[22]. maturity is cultivated, whereas for different sub-regions
In this paper, we are proposing novel crop yield planting time varies. SCLC has four phases,
forecasting model using long term time series and support a)Germination and Establishment, b)Tillering, c)Grand
vector regression, for sugarcane as a primary crop in a Growth, d)Ripening and Maturity. Germination and
multi-crop system with multiple, unknown inter-season Establishment phase lasts for 15 days, whereas Tillering
secondary crops, using remote sensed NDVI data and phase is for about 4 months, Grand Growth phase is about
ground-based observed, highly co-related weather and 4.5 months, Ripening and Maturity phase lasts for 3
soil data. We are also comparing the accuracy metrics of months. As plantation start time varies in the Karnataka
SVR with other popular supervised ML algorithms like region, start and end of each phase vary in different
Gaussian Process Regression(GPR)[23] and Linear regions[22].
Regression(LR). Parametric ML algorithms, GPR and LR
have been chosen for comparison purpose because of
their derivation from classic empirical models. We are IV. DATASET DESCRIPTION
also analyzing curse of dimension problem of ML
algorithms by using dimension reduction techniques like Dataset is a critical component of any ML algorithm
Lasso Regression ML algorithm[24] as well as selecting and it needs to be understood and pre-processed before
applying ML algorithms in any domain. Dataset used in and soil data recorded at every hour at latitude and
this research comprises Weather and Soil Dataset(WSD), longitude, and yield per hector with number of hectors
NDVI dataset and Sugarcane crop yield dataset. WSD is cultivated, gathered at every season of primary crop
downloaded from sugarcane and secondary crops like maize, rice for entire
[Link] district, where season for sugarcane is entire year and for
.246N74.737E [28], for the village Shirdhan, located at secondary crops, three seasons in a year.
latitude and longitude of 16.2458oN, 74.737oE of Fig.2 shows the average yearly pattern of each
Belagavi district, Karnataka(India). WSD dataset attributes at Shirdhan. Temperature and Soil Temperature
attributes are listed in Table 1. follows a similar pattern. Soil Moisture, Relative
Humidity, and Dew Temperature follow similar trends,
Table 1. List of attributes and units of measurement Sunshine Duration and Precipitation follows the opposite
Attribute Name Measurement Units trend. Closely looking at NDVI trend, which is lowest
during January and highest during September and
Temperature (2m above October, and starts reducing during November and
Celsius December. So we can safely assume that crop sowing
ground) (T)
Dew Point Temperature (2m
Celsius starts during January and harvesting starts during
above ground) (DPT) November.
Soil Temperature (0-10cm
Celsius
below ground) (ST) A. Dataset Distribution And Outliers
Soil Moisture (0-10cm below
ground) (SM) m3.m−3 Using various graphs, we can understand the
Precipitation (P) mm distribution of each attributes in dataset independently.
Relative Humidity (2m Box and Whisker plot is shown in Fig.3, indicates
%
above ground) (RH)
weather attributes are skewed or have outliers. Attributes
Sunshine Duration (SD) W/m2
like Precipitation, NDVI, and Evapotranspiration have
Evapotranspiration (E) mm
outliers. Outliers in weather dataset could not be
NDVI range(-1 ,1)
neglected, as they provide very important information
about the nature of overall weather condition. Histogram
graph, as shown in Fig.4, groups each attributes in the
number of bins and provides the number of observations
in each bin, and shape of bins lets us understand the kind
distribution each attribute is following.
that attribute Temperature has a positive relationship with accordance with ML approaches. As shown in Fig.7, the
Soil Temperature and Dew Point Temperature, negative first module known as Dataset Pre-processing
relation with Soil Moisture, Relative Humidity, Module(DPM) re-samples, scales and normalizes each
Evapotranspiration, and Precipitation. After careful attribute, select independent and important attributes and
analyses of the correlation matrix, we could select either divides into training and testing dataset. The second
Temperature or Soil Temperature, either Dew Point module i.e. Training and Testing Module(TTM) trains
Temperature or Relative Humidity, either Sunshine SVR algorithm with RBF kernel and other Supervised
Duration or Soil Moisture. Correlation matrix also ML algorithms for comparison purpose, and verifies
conveys that, many features have a strong relationship trained algorithms using hold-out data and compares
between them. So feature selection or dimensionality various regression metrics to evaluate efficiency and
reduction should be applied before learning prediction accuracy of a trained module. The third module is known
function. Criteria for Selection of attributes in high as Prediction Module(PRM) forecasts weather and soil
dimensional dataset, feature selection or dimension attributes, NDVI, and finally sugarcane crop yield. Each
reduction techniques have been used to reduce feature to module has three sub-modules: Weather and Soil
boost algorithm performance, increase the accuracy of Attribute Module(WASAM), which deals with weather
estimators. Univariate feature selection method selects and soil attributes pre-processing, training, testing, and
the best feature based on statistical tests like best scoring prediction. NDVI Module(NDVIM) for NDVI values,
feature, best percentile feature, false positive rate, false and Sugarcane Crop Yield Forecasting Module (SCYFM)
discovery rate, family-wise error, hyper-parameter search for sugar- cane yield. All three modules are implemented
estimator. Recursive feature elimination method, which using Sci-Kit Learn package version 0.1.91 and Python
eliminates features by comparing a small set of features 3.7.
recursively, L1 based feature selection methods based on
linear regression and Linear SVM, Tree-Based Feature
Selection compute feature importance. In this paper, ML-
based feature selection methods are used to calculate the
importance of each feature.
output attributes, and ML algorithm learns function of total samples. In NDVI module, weighted SVR
which relates output(s) to input attributes. Time series algorithm is trained using previous 8 years of WS and
dataset consists of observations indexed by time period, NDVI datasets, where WS dataset is recorded every hour
i.e. frequency, but observation at frequency 't' is and down-scaled to every 15 days, as input attributes and
independent and identically distributed. However, NDVI dataset is observed every 15 days as output to
observation at frequency 't' depends on observation at forecast NDVI as a time-independent attribute. In
previous frequency t-1 in the long term and has a practice, NDVI is derived from greenery and solar
regular pattern. So observation at time 't+1' will be radiance and is not time-dependent as weather and soil.
dependent on observation at a time 't', by considering Weights of samples are increased in accordance with the
this, we have converted time series dataset into chronological order of dataset, with an assumption of
supervised machine learning problem. recent past WS attribute values have more influence as
NDVI Forecasting for the duration of SCLC, 12 compared to past values. According to the correlation
months is considered as supervised ML problem with matrix, it is observed that WS attributes are completely
NDVI values as output and WS attributes as input. WS independent, and in order to reduce the influence of
dataset has been down-sampled to the frequency of NDVI, dependent attributes, feature selection algorithms like
but each attribute is measured on a different scale. In Lasso, Decision Trees are used to select the best features.
NDVIM submodule, attributes are scaled between 0 and 1 NDVI module has been trained and tested with various
using normalization according to equation (1) combinations of features. In CYPM module, SVR
algorithm is trained using average NDVI values for the
xi − xmin past 8 years from 28 districts of Karnataka state, where
NX i = (1) sugarcane is cultivated, as a dataset with NDVI at every
xmax − xmin
15 days over one year period as input attributes and
yearly sugarcane crop yield as output. If we consider one
NXi is the normalized value of observation Xi, xi is the location then dataset will have only 8 samples, so all
value of ith observation for attribute x, xmax is the districts of Karnataka is considered for training SVR. In
maximum value observed for attribute x and x min is the this module, training dataset has been divided into
minimum value observed for attribute x. SCYFM is various sizes between 20% and 90% in the step of 5%
considered as a ML problem with sugarcane crop yield as each, and comparison of training time and r2 score values
output and NDVI values as input. Crop yield is recorded have been analyzed to understand the minimum number
at every season, where the season is 12 months long and of training samples required for better accuracy. Other
NDVI values are derived after every 15 days. We have ML algorithms like GPR, DT, and Lasso are trained and
four years of yield data per district, which makes it tested for comparison purpose.
difficult for the ML algorithm with one district dataset.
So we have considered all sugarcane growing districts in C. Prediction Module
Karnataka state for training and testing purpose, whereas In Prediction Module, each attribute of WS dataset is
yield is predicted at Shirdhan. Mean, Median, and
predicted for the next season in real time. Predicted WS
Standard Deviation of NDVI values for each phase of
attributes are used as input to the prediction of NDVI
sugarcane cultivation is considered as one input attribute. values and predicted NDVI values are used as input to
Phase wise NDVI aggregation is done according to sugarcane crop yield prediction.
planting timing in each district. In Bijapur and Belagavi
districts, sugarcane planting is done in January, So NDVI
for first 15 days of January is considered as phase 1
VI. EVALUATION RESULTS AND ANALYSIS
NDVI, whereas in Coastal districts, planting is done in
November, so NDVI for first 15 days of November is The implemented model has been evaluated using
considered as Phase 1 NDVI. Phase 2 NDVI values are accuracy matrix by running experiments with various test
calculated as aggregate of NDVI values of the next four and train sizes for SVR algorithm and comparing with
months after the first 15 days of January and November Lasso, Naive-Bayes and Decision Tree algorithms.
respectively for each region. Phase duration for different WASAM model has been evaluated using hold-out sizes
regions in Karnataka is shown in Fig.1. ranging from 5% to 50%. According to Fig.8, Naive-
Bayes algorithm is performing better as comparing to the
B. Training And Testing Module
other three algorithms. Accuracies of Soil Temperature,
In WASAM submodule, each attribute listed in Table.1 Soil Moisture, and Temperature prediction are more than
is observed at every hour, down-sampled to every 15 80% whereas, for Precipitation, accuracy is low and is
days, and are forecasted individually for the duration of about 35%. According to Fig.9, the dataset used has 180
one year by modeling 24 steps LTTS as supervised ML samples. Various algorithms are used such as SVR, Lasso,
regression. WASAM is trained iteratively as SVR Naïve-Bayes, and Decision Tree Regressor. These
algorithm with RBF kernel, C value as 100 and gamma algorithms have been used to predict features such as Soil
value as 0.1, using previous 20 years observations over a Temperature, Temperature, Soil Moisture, and
prior 6, 12, 18, and 24 observations as input set and next Precipitation. These algorithms have been run for several
24 observations as multiple outputs. Learned Model has times for different dataset samples with the base sample
been tested using held-out dataset samples of 75% to 5% size taken as 100 samples of data and then incremented
Fig.10. Box and Whisker and Line Plot for NDVI prediction
Fig.11. Box and Whisker and Line Plot for Crop prediction.
In Fig.10, SVR-RBF yields 89.97% accuracy for 112
samples of data in NDVI prediction. When we look at Fig.11 represents the crop prediction where the dataset
the entire range of accuracies and perform analysis, has 380 samples in total, and 180 samples of data have
minimum accuracy obtained for SVR-RBF is 82.02% been taken as base for training the dataset, the testing
with 100 samples considered, and maximum accuracy dataset consists of multiple slices which have been
obtained is 89.97% for 112 samples. When the median is increased periodically, sample size for every run varies
considered for the entire range, the accuracy obtained is
which results in various accuracy scores from which the accuracy of 85.24% for Soil Temperature given by Lasso,
maximum score would be considered. SVR-RBF yields 85.372% accuracy for Temperature given by Naive Bayes,
83.49% accuracy for 122 samples of data in crop accuracy for Soil Moisture is 77.46% given by Naive
prediction. After looking at the entire range of accuracies Bayes, accuracy for Precipitation is 28.69% given by
and performing analysis, minimum accuracy obtained for Naive Bayes for weather and soil attribute prediction,
SVR-RBF is 74.53% with 146 samples considered, and 89.97% for NDVI forecasting and 83.49% for final yield
maximum accuracy obtained is 83.49% for 122 samples. prediction. In the future, we can consider predicting
When the median is considered for the entire range, the sugar-cane yield considering variable growth periods of
accuracy obtained is 80.41%. About 40% of the obtained 12 months and 18 months. Since SVR has emerged as a
accuracies lie greater than the median, proving the better algorithm, we can also use various Kernel
algorithm implementation to be constant. After looking at functions to reduce noise in the dataset, to consider a
the entire range of accuracies and performing analysis, better seasonal variation. We can also think of using
minimum accuracy obtained for GPR is 74.29% with 148 ensemble learning methods and compare accuracy with
samples considered, and maximum accuracy obtained is SVR.
81.71% for 122 samples. When the median is considered
for the entire range, the accuracy obtained is 78.39%. REFERENCES
About 48% of the obtained accuracies lie below the
[1] Antti Sorjamaa, Jin Hao, Nima Reyhani, Yongnan Ji,
median and 52% of the accuracies lie greater than the Amaury Lendasse, Methodology for long-term prediction
median. After looking at the entire range of accuracies of time series, Neurocomputing, Volume 70, Issues 16–18,
and performing analysis, minimum accuracy obtained for 2007, Pages 2861-2869, ISSN 0925-2312,
Kernel Ridge-RBF is 17.03% with 110 samples [Link]
considered, and maximum accuracy obtained is 21.66% [2] H. Aghighi, M. Azadbakht, D. Ashourloo, H. S. Shahrabi,
for 150 samples. When the median is considered for the and S. Radiom, "Machine Learning Regression
entire range, the accuracy obtained is 19.68%. About Techniques for the Silage Maize Yield Prediction Using
43% of the obtained accuracies lie below the median and Time-Series Images of Landsat 8 OLI," in IEEE Journal
of Selected Topics in Applied Earth Observations and
57% of the accuracies lie greater than the median. This
Remote Sensing, vol. 11, no. 12, pp. 4563-4577, Dec.
algorithm has given the least accurate prediction for the 2018.
crop. After looking at the entire range of accuracies and [3] W.G.N.N. Jayawardhana, V.M.I. Chathurange, Extraction
performing analysis, minimum accuracy obtained for of Agricultural Phenological Parameters of Sri Lanka
Lasso Regression is 21.77% with 148 samples considered, Using MODIS, NDVI Time Series Data, Procedia Food
and maximum accuracy obtained is 27.03% for 120 Science, Volume 6, 2016, Pages 235-241, ISSN 2211-
samples. When the median is considered for the entire 601X, [Link]
range, the accuracy obtained is 24.41%. About 48% of [4] Y.R. Lai, M.J. Pringle, P.M. Kopittke, N.W. Menzies, T.G.
the obtained accuracies lie below the median and 52% of Orton, Y.P. Dang, An empirical model for prediction of
wheat yield, using time-integrated Landsat NDVI,
the accuracies lie greater than the median. This algorithm
International Journal of Applied Earth Observation and
has given a very low accurate prediction for the crop. The Geoinformation, Volume 72, 2018, Pages 99-108, ISSN
line graph also depicts similar results for the applied 0303-2434, [Link]
algorithms for crop prediction. The fall and rise of [5] Saeed, Umer & Dempewolf, Jan & Becker-Reshef, Inbal
accuracies for crop prediction can be observed in Fig.11. & Khan, Ahmad & Ahmad, Ashfaq & Aftab Wajid, Syed.
These accuracies have also been plotted for exactly the (2017). Forecasting wheat yield from weather data and
same parameters as the Box and Whisker plot. The line MODIS NDVI using Random Forests for Punjab province,
graph depicts the trend in the accuracies for changing Pakistan. International Journal of Remote Sensing. 38.
dataset sample sizes. Thus for the crop prediction, we can 4831-4854. 10.1080/01431161.2017.1323282.
[6] Manasah S. Mkhabela, Milton S. Mkhabela, Nkosazana N.
conclude that SVR-RBF is the best performing algorithm
Mashinini, Early maize yield forecasting in the four agro-
when compared to the other three algorithms and leads ecological regions of Swaziland using NDVI data derived
the pack with 83.46% of accuracy. from NOAA's-AVHRR, Agricultural and Forest
Meteorology, Volume 129, Issues 1–2, 2005, Pages 1-9,
ISSN 0168-1923,
VII. CONCLUSION AND FUTURE SCOPE [Link]
[7] Prasad, Anup & Singh, R & Tare, V & Kafatos, Menas.
Earlier researchers have worked on predicting (2007). Use of vegetation index and meteorological
sugarcane yield for small regions where weather parameters for the prediction of crop yield in India.
conditions and sowing start times were the same. In this International Journal of Remote Sensing. 28. 5207-5235.
research, we have successfully modeled sugarcane yield 10.1080/01431160601105843.
prediction considering different sowing start period under [8] Ahmed, Nesreen & Atiya, Amir & Gayar, Neamat & El-
different conditions in India region with an overall Shishiny, Hisham. (2010). An Empirical Comparison of
accuracy of 83.49%. We have also developed a model Machine Learning Models for Time Series Forecasting.
Econometric Reviews. 29. 594-621.
which can predict crop yield in real time by predicting 10.1080/07474938.2010.481556.
weather conditions using time series forecasting and [9] X.E. Pantazi, D. Moshou, T. Alexandridis, R.L. Whetton,
predicting NDVI using predicted weather parameters and A.M. Mouazen, Wheat yield prediction using machine
using NDVI predicting real-time crop. We have achieved learning and advanced sensing techniques, Computers and
Electronics in Agriculture, Volume 121, 2016, Pages 57- [20] P. Bose, N. K. Kasabov, L. Bruzzone, and R. N. Hartono,
65, ISSN 0168-1699, "Spiking Neural Networks for Crop Yield Estimation
[Link] Based on Spatiotemporal Analysis of Image Time Series,"
[10] J. Huang, H. Wang, Q. Dai and D. Han, "Analysis of in IEEE Transactions on Geoscience and Remote Sensing,
NDVI Data for Crop Identification and Yield Estimation," vol. 54, no. 11, pp. 6563-6573, Nov. 2016.
in IEEE Journal of Selected Topics in Applied Earth [21] Fang, Hongliang & Liang, Shunlin & Hoogenboom,
Observations and Remote Sensing, vol. 7, no. 11, pp. Gerrit. (2011). Integration of MODIS LAI and vegetation
4374-4384, Nov. 2014. index products with the CSM-CERES-Maize model for
[11] Anup K. Prasad, Lim Chai, Ramesh P. Singh, Menas corn yield estimation. International Journal of Remote
Kafatos, Crop yield estimation model for Iowa using Sensing - INT J REMOTE SENS. 32. 1039-1065.
remote sensing and surface parameters, International 10.1080/01431160903505310.
Journal of Applied Earth Observation and Geoinformation, [22] [Link]/advertisements/20170809_advt_36.pdf
Volume 8, Issue 1, 2006, Pages 26-33, ISSN 0303-2434, [23] Singh, Arti & Ganapathysubramanian, Baskar & Singh,
[Link] Asheesh & Sarkar, Soumik. (2015). Machine Learning for
[12] David M. Johnson, An assessment of pre- and within- High-Throughput Stress Phenotyping in Plants. Trends in
season remotely sensed variables for forecasting corn and Plant Science. 21. 10.1016/[Link].2015.10.015.
soybean yields in the United States, Remote Sensing of [24] R. Tibshirani, "Regression shrinkage and selection via the
Environment, Volume 141, 2014, Pages 116-128, ISSN lasso: a retrospective," J. R. Statist. Soc. B (2011), p. 10,
0034-4257, [Link] 2011.
[13] Yaping Cai, Kaiyu Guan, Jian Peng, Shaowen Wang, [25] [Link]
Christopher Seifert, Brian Wardlow, Zhan Li, A high- areas-of-india/
performance and in-season classification system of field- [26] [Link]
level crop types using time-series Landsat data and a cultivation-in-india-conditions- production-and-
machine learning approach, Remote Sensing of distribution/20945
Environment, Volume 210, 2018, Pages 35-47, ISSN [27] [Link]
0034-4257, [Link] l
[14] Yaoliang Chen, Dengsheng Lu, Lifeng Luo, Yadu Pokhrel, [28] [Link]
Kalyanmoy Deb, Jingfeng Huang, Youhua Ran, Detecting 246N74.737E
irrigation extent, frequency, and timing in a heterogeneous [29] [Link]
arid agricultural region using MODIS time series, Landsat
imagery, and ancillary data, Remote Sensing of
Environment, Volume 204, 2018, Pages 197-211, ISSN
0034-4257, [Link] Authors’ Profiles
[15] Parihar, J & Oza, Markand. (2006). FASAL: An
integrated approach for crop assessment and production
Ramesh Medar, Assistant Professor,
forecasting. Proceedings of SPIE - The International
Department of Computer Science and
Society for Optical Engineering. 6411.
Engineering, KLS Gogte Institute of
10.1117/12.713157.
Technology, Belagavi, Karnataka, India.
[16] Steduto, Pasquale & Hsiao, Theodore & Raes, Dirk &
Pursuing Ph.D. in the domain Machine
Fereres, E. (2009). AquaCrop—The FAO Crop Model to
learning, Data mining. Completed [Link].
Simulate Yield Response to Water: I. Concepts and
in the year 2011 from KLS Gogte Institute
Underlying Principles. Agronomy Journal - AGRON J.
of Technology. Total teaching experience of 12.5 years, 5 years
101. 10.2134/agronj2008.0139s.
of research experience. Published a few papers in national,
[17] X.E. Pantazi, D. Moshou, T. Alexandridis, R.L. Whetton,
international journals. Presented papers in national, international
A.M. Mouazen, Wheat yield prediction using machine
conferences. Member of LMISTE, CSTA.
learning and advanced sensing techniques, Computers and
Electronics in Agriculture, Volume 121, 2016, Pages 57-
65, ISSN 0168-1699,
Dr. Vijay S Rajpurohit, working as
[Link]
Professor in the Department of Computer
[18] D Jones, E.M Barnes, Fuzzy composite programming to
Science and Engg at Gogte Institute of
combine remote sensing and crop models for decision
Technology, Belagavi, Karnataka, India.
support in precision crop management, Agricultural
Completed B.E. in Computer Science and
Research Division, University of Nebraska, Agricultural
Engg. from Karnataka University Dharwad,
Systems, Volume 65, Issue 3, 2000, Pages 137-158, ISSN
[Link]. at N.I.T.K Surathkal and Ph.D.
0308-521X, [Link]
from Manipal University, Manipal in 2009. His research areas
521X(00)00026-3.
include Image Processing, Cloud Computing, and Data
[19] Jiang, Dong & Yang, X.H. & Clinton, Nicholas & Wang,
Analytics. He has published a good number of papers in
Naijiang. (2004). An artificial network model for
Journals, International and National conferences. Dr. V. S.
estimating crop yields using remotely sensed information.
Rajpurohit is the reviewer for a few international journals and
International Journal of Remote Sensing. 25. 1723-1732.
conferences. He is the associate editor for two international
10.1080/0143116031000150068.
journals and Senior Member of the International Association of
CS and IT. He is also the life member of SSI, ISC and ISTE
associations.