CROP YIELD PREDICTION USING MACHINE LEARNING
ALGORITHMS
Dr. K. Sujatha1a, Karri Sri Chaithanya1b
Gedda Jayadhar Pavan Kalyan2, Moturu Vijaya3, Kothapalli Ashok4
1a) Professor, Dept. of Computer Science, Dadi Institute of Engineering & Technology, Anakapalli, India.
1b, 2, 3, 4) Student, Dept. of Computer Science, Dadi Institute of Engineering & Technology, Anakapalli, India.
hodcse@[Link]
chaithanya20043@[Link], jayadhargedda@[Link]
vijayamoturu270@[Link], kothapalliashok70@[Link]
Abstract: Crop yield prediction plays a critical role in optimizing agricultural practices and ensuring food
security. Accurate forecasting of crop production is essential for efficient resource allocation, planning, and
decision-making in agriculture. This project focuses on predicting crop yields using various machine learning
algorithms, with an emphasis on the Bagging Regressor model due to its ability to handle complex, high-vari-
ance agricultural data. The study incorporates key factors influencing crop yield, including soil characteristics,
climatic conditions, crop types, and historical yield data. Data preprocessing techniques such as cleaning, fea-
ture selection, and normalization are applied to enhance model performance. Multiple machine learning al-
gorithms, including Decision Trees, Random Forests, and Gradient Boosting, are evaluated and compared to
identify the most accurate and reliable model. The Bagging Regressor, by aggregating predictions from multiple
models trained on different subsets of the data, significantly improves prediction accuracy and reduces overfit-
ting. The model's performance is validated using various evaluation metrics such as Mean Absolute Error
(MAE), R-squared, and Root Mean Squared Error (RMSE). The results demonstrate that the Bagging Regressor
provides the most accurate crop yield predictions, making it a valuable tool for precision agriculture. The final
model is integrated into a user-friendly system for real-time yield predictions, helping farmers and stakeholders
make informed decisions. This project contributes to the advancement of data-driven agriculture, offering in -
sights for sustainable farming practices and enhanced food security.
Keywords: Agriculture, Training and Testing, Crop yield prediction
1. INTRODUCTION
Agriculture plays a vital role in ensuring food security and supporting economies worldwide. Accurately
predicting crop yield is a key challenge for farmers, agricultural planners, and policymakers, as it helps optimize
resources and improve decision-making. Crop yield prediction involves estimating the quantity of crop produc-
tion per unit area of land, which is essential for planning agricultural operations and managing food supply
chains.
The use of machine learning (ML) techniques in agriculture has opened new possibilities for analyzing complex
and diverse datasets to make precise predictions. This project focuses on utilizing various machine learning
algorithms to predict crop yield based on factors such as soil characteristics, climatic conditions, crop types, and
historical data. By applying ML models, the project aims to estimate the expected yield per quintal of land accu -
rately, thereby helping stakeholders make informed decisions. The study explores multiple machine learning
algorithms, including Linear Regression, Decision Trees, Random Forests, and Bagging Regressor, to evaluate
their accuracy and efficiency in yield prediction.
This initiative seeks to provide a practical solution to challenges in the agricultural sector by enhancing produc-
tivity, reducing uncertainties, and promoting resource-efficient farming practices. The insights derived from this
project will contribute to sustainable agriculture and help address global food security concerns.
2. LITERATURE SURVEY
The application of machine learning in crop yield prediction has gained significant attention due to its abil -
ity to analyze complex datasets and provide accurate forecasts. A variety of approaches have been employed to
leverage environmental, climatic, and soil-related data for predicting crop yields. Researchers have explored the
use of statistical and machine learning models to understand the intricate relationships between the influencing
factors and yield outcomes.
Early studies utilized traditional statistical methods such as Linear Regression, which provided a straightforward
approach to modeling relationships between variables. While effective in simple scenarios, these models often
struggled to capture the non-linear dependencies commonly observed in agricultural data. This limitation led to
the adoption of more advanced machine learning algorithms capable of addressing the complexity of crop yield
prediction.
Decision Trees and ensemble methods, such as Random Forests, have been widely adopted due to their ability to
handle large datasets and capture non-linear interactions between variables. These models are particularly effec -
tive in accounting for the variability introduced by factors such as soil characteristics, crop types, and climatic
conditions. Comparative studies have consistently highlighted their superior accuracy over traditional methods,
making them a preferred choice for many researchers.
Among these techniques, Bagging Regression has demonstrated exceptional performance in yield prediction
tasks. By averaging predictions from multiple models built on random subsets of the data, Bagging Regression
effectively reduces variance and enhances robustness. Its ability to handle noise and avoid overfitting makes it
particularly suitable for agricultural datasets, where variability is high.
Overall, the literature underscores the potential of machine learning in transforming agricultural practices. By
integrating diverse data sources and employing advanced algorithms, researchers are paving the way for data-
driven, sustainable farming systems that improve yield forecasting and resource management.
3. METHODOLOGY
The methodology for crop yield prediction involves a series of systematic steps to ensure accurate and
reliable outcomes. The process begins with the collection of historical data from diverse sources, including agri-
cultural surveys, weather records, soil profiles, and past crop yields. This data is then cleaned to address missing
values, remove outliers, and ensure consistency across variables.
Key features influencing crop yield, such as soil properties, climatic factors, and crop types, are identified using
techniques. The dataset is then divided into training and testing subsets to facilitate model development and
evaluation. Machine learning algorithms such as Bagging Regression, Random Forests, or Decision trees are
employed, and the models are trained on the training dataset to learn patterns and relationships. Once trained,
the models are tested on the testing dataset to assess their performance using metrics like Mean Absolute Error
(MAE) and R-squared.
The best-performing model is then used for making predictions on new data. Finally, the model is integrated into
a practical system or interface, allowing for user-friendly predictions. Continuous updates and retraining are
implemented to adapt the model to new data, ensuring its relevance and accuracy over time.
3.1 Machine Learning Algorithm
The Bagging Regressor algorithm plays a pivotal role in this project due to its ability to enhance prediction
accuracy and robustness. Bagging works by creating multiple subsets of the dataset through random sampling
with replacement and training individual models on these subsets. The final prediction is obtained by averaging
the outputs of these models, which significantly reduces variance and minimizes overfitting.
In the context of crop yield prediction, where datasets often exhibit high variability due to environmental and
agricultural factors, the Bagging Regressor's ensemble approach is particularly beneficial. It effectively handles
non-linear relationships between input features such as soil properties, weather conditions, and crop types,
which might not be captured well by simpler algorithms. Moreover, its ability to combine multiple weak learn-
ers into a strong predictive model ensures consistent performance, even with noisy or incomplete data.
Compared to other algorithms, the Bagging Regressor offers superior accuracy by leveraging diversity in the
base models, thereby mitigating the risks of relying on a single model's prediction. This makes it an ideal choice
for this project, where precise crop yield estimates are critical for resource planning and decision-making in
agriculture. Its scalability and adaptability further enhance its suitability for real-world applications, ensuring
reliable predictions across varying datasets and conditions.
3.2 Architecture
3.3 Sample Dataset
4. RESULTS AND DISCUSSIONS
The results of the crop yield prediction models indicate that the Bagging Regressor significantly outperforms
other machine learning algorithms in terms of accuracy and reliability. After training and testing the models on a
diverse dataset that includes soil properties, climatic conditions, and historical yield data, the Bagging Regressor
achieved a high R-squared value, demonstrating its ability to explain a large proportion of the variance in crop
yields. Additionally, the model showed a low Mean Absolute Error (MAE) and Root Mean Squared Error
(RMSE), indicating precise and consistent predictions. In comparison to other algorithms like Random Forests,
Decision Trees, and Gradient Boosting, the Bagging Regressor exhibited superior generalization, reducing the
risk of overfitting and providing more accurate yield predictions across different crop types and regions. These
results highlight the effectiveness of the Bagging Regressor for handling complex, high-variance agricultural
data and emphasize its potential as a valuable tool for precision farming.
4.1 Accuracy
Table 4.2.1. Accuracy
4.2 Error
Table 4.2.2. Comparing MSE, RMSE, MAE
4.3 Screenshots
Table 4.3.1. Models Accuracy
Table 4.3.2. Comparing Accuracy and Error
5. CONCLUSION
This project demonstrates the effectiveness of machine learning, particularly the Bagging Regressor, in predict -
ing crop yields accurately. By utilizing data on soil, weather, and crop history, the model achieved superior ac-
curacy compared to other algorithms, reducing variance and overfitting. The results validate the potential of
machine learning for enhancing decision-making in agriculture, offering a reliable tool for improving crop man-
agement and resource allocation. With continuous updates and retraining, this approach can contribute to sus -
tainable farming practices and help address food security challenges.
Fig.1. Displayed webpage of the project
Fig.2. Output represents the amount of crop yield produced in quintals per hectare
7. References
1. R. Jahan, ‘‘Applying naive Bayes classification technique for classification of improved agricultural land soils,’’
Int. J. Res. Appl. Sci. Eng. Technol., vol. 6, no. 5, pp. 189–193, May 2018.
2. B. B. Sawicka and B. Krochmal-Marczak, ‘‘Biotic components influencing the yield and quality of potato tubers,’’
Herbalism, vol. 1, no. 3, pp. 125–136, 2017.
3. B. Sawicka, A. H. Noaema, and A. Gáowacka, ‘‘The predicting the size of the potato acreage as a raw material for
bioethanol production,’’ in Alternative Energy Sources, B. Zdunek, M. Olszáwka, Eds. Lublin, Poland: Wydawnictwo
Naukowe TYGIEL, 2016, pp. 158–172.
4. B. Sawicka, A. H. Noaema, T. S. Hameed, and B. Krochmal-Marczak, ‘‘Biotic and abiotic factors influencing on the
environment and growth of plants,’’ (in Polish), in Proc. Bioróżnorodność Środowiska Znaczenie, Problemy,
Wyzwania. Materiały Konferencyjne, Puławy, May 2017. [Online]. Available: [Link]
5. R. H. Myers, D. C. Montgomery, G. G. Vining, C. M. Borror, and S. M. Kowalski, ‘‘Response surface methodology: A
retrospective and literature survey,’’ J. Qual. Technol., vol. 36, no. 1, pp. 53–77, Jan. 2004.
6. D. K. Muriithi, ‘‘Application of response surface methodology for optimization of potato tuber yield,’’ Amer. J. Theor.
Appl. Statist., vol. 4, no. 4, pp. 300–304, 2015, doi: 10.11648/[Link].20150404.20.
7. M. Marenych, O. Verevska, A. Kalinichenko, and M. Dacko, ‘‘Assessment of the impact of weather conditions on the
yield of winter wheat in Ukraine in terms of regional,’’ Assoc. Agricult. Agribusiness Econ. Ann. Sci., vol. 16, no. 2,
pp. 183–188, 2014.
8. J. R. Olędzki, ‘‘The report on the state of remotesensing in Poland in 2011–2014,’’ (in Polish), Remote Sens. Environ.,
vol. 53, no. 2, pp. 113–174, 2015.
9. K. Grabowska, A. Dymerska, K. Poáarska, and J. Grabowski, ‘‘Predicting of blue lupine yields based on the selected
climate change scenarios,’’ Acta Agroph., vol. 23, no. 3, pp. 363–380, 2016.
10. D. Li, Y. Miao, S. K. Gupta, C. J. Rosen, F. Yuan, C. Wang, L. Wang, and Y. Huang,‘‘Improving po tato yield predic-
tion by combining cultivar information and UAV remote sensing data using machine learning,’’ Remote Sens., vol.
13, no. 16, p. 3322, Aug. 2021, doi: 10.3390/rs13163322.
11. N. Chanamarn, K. Tamee, and P. Sittidech, ‘‘Stacking technique for academic achievement prediction,’’ in Proc. Int.
Workshop Smart Info-Media Syst., 2016, pp. 14–17.
12. W. Paja, K. Pancerz, and P. Grochowalski, ‘‘Generational feature elimination and some other ranking feature selection
methods,’’ in Advances in Feature Selection for Data and Pattern Recognition, vol. 138. Cham, Switzerland: Springer,
2018, pp. 97–112.
13. D. C. Duro, S. E. Franklin, and M. G. Dubé, ‘‘A comparison of pixelbased and object-based image analysis with selected
machine learning algorithms for the classification of agricultural landscapes using SPOT-5 HRGimagery,’’ Remote Sens.
Environ., vol. 118, pp. 259–272, Mar. 2012.
14. S. K. Honawad, S. S. Chinchali, K. Pawar, and P. Deshpande, ‘‘Soil classification and suitable crop prediction,’’ in Proc.
Nat. Conf. Comput. Biol., Commun., Data Anal. 2017, pp. 25–29.
15. J. You, X. Li, M. Low, D. Lobell, and S. Ermon, ‘‘Deep Gaussian process for crop yield prediction based on remote sens-
ing data,’’ in Proc. AAAI Conf. Artif. Intell., 2017, vol. 31, no. 1, pp. 4559–4565.
16. D. A. Reddy, B. Dadore, and A. Watekar, ‘‘Crop recommendation system to maximize crop yield in ramtek region using
machine learning,’’ Int. J. Sci. Res. Sci. Technol., vol. 6, no. 1, pp. 485–489, Feb. 2019.
17. N. Rale, R. Solanki, D. Bein, J. Andro-Vasko, and W. Bein, ‘‘Prediction of CropCultivation,’’ in Proc. 19th Annu. Comput.
Commun. Workshop Conf. (CCWC), Las Vegas, NV, USA, 2019, pp. 227–232.
18. J. Jones, G. Hoogenboom, C. Porter, K. Boote, W. Batchelor, L. Hunt, P. Wilkens, U. Singh, A. Gijsman, and J. Ritchie,
‘‘The DSSAT cropping system model,’’ Eur. J. Agronomy, vol. 18, nos. 3–4, pp. 235–265, 2003.
19. M. T. N. Fernando, L. Zubair, T. S. G. Peiris, C. S. Ranasinghe, and J. Ratnasiri, ‘‘Economic value of climate variability
impact on coconut production in Sri Lanka,’’ in Proc. AIACC Working Papers, vol. 45, 2007, pp. 1–7.
20. B. Ji, Y. Sun, S. Yang, and J. Wan, ‘‘Artificial neural networks for rice yield prediction in mountainous regions,’’ J. Agri-
cult. Sci., vol. 145, no. 3, pp. 249–261, Jun. 2007.
21. C. Boryan, Z. Yang, R. Mueller, and M. Craig, ‘‘Monitoring U.S. agriculture: The U.S. department of agriculture, national
agricultural statistics service, cropland data layer program,’’ Geocarto Int., vol. 26, no. 5, pp. 341–358, 2011.
22. M. C. Hansen and T. R. Loveland, ‘‘A review of large area monitoring of land cover change using Landsat data,’’ Remote
Sens. Environ., vol. 122, pp. 66–74, Jul. 2012.
23. D. K. Bolton and M. A. Friedl, ‘‘Forecasting crop yield using remotely sensed vegetation indices and crop phenology
metrics,’’ Agricult. Forest Meteorol., vol. 173, pp. 74–84, May 2013.
24. J. Dempewolf, B. Adusei, I. Becker-Reshef, M. Hansen, P. Potapov, A. Khan, and B. Barker, ‘‘Wheat yield forecasting for
Punjab province from vegetation index time series and historic crop statistics,’’ Remote Sens., vol. 6, no. 10, pp. 9653–
9675, Oct. 2014.
25. H. D. Shannon and P. M. Raymond, ‘‘Managing weather and climate risk to agriculture in North America,’’ Central
Amer. Caribbean, vol. 10, pp. 50–56, Dec. 2015.