0% found this document useful (0 votes)
33 views18 pages

Federated Learning in Crop Yield Prediction

This document provides a comprehensive review of federated learning (FL) techniques and their applications in crop yield prediction, highlighting the advantages of decentralized data processing while maintaining privacy. It examines mathematical foundations, compares machine learning models, and discusses real-world implementations along with challenges such as data heterogeneity and communication overhead. The review aims to address existing research gaps and propose future directions for integrating FL with emerging agricultural technologies.

Uploaded by

hw459277
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
33 views18 pages

Federated Learning in Crop Yield Prediction

This document provides a comprehensive review of federated learning (FL) techniques and their applications in crop yield prediction, highlighting the advantages of decentralized data processing while maintaining privacy. It examines mathematical foundations, compares machine learning models, and discusses real-world implementations along with challenges such as data heterogeneity and communication overhead. The review aims to address existing research gaps and propose future directions for integrating FL with emerging agricultural technologies.

Uploaded by

hw459277
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

MethodsX 14 (2025) 103408

Contents lists available at ScienceDirect

MethodsX
journal homepage: [Link]/locate/methodsx

Federated learning for crop yield prediction: A comprehensive


review of techniques and applications
Vani Hiremani a, Raghavendra M. Devadas b,∗, Preethi b, R. Sapna b, T. Sowmya c,
Praveen Gujjar d, N. Shobha Rani e, K.R. Bhavya f
a
Symbiosis Institute of Technology, Symbiosis International (Deemed) University, Pune, India
b
Department of Information Technology, Manipal Institute of Technology Bengaluru, Manipal Academy of Higher Education, Manipal, India
c
Department of Computer Science and Engineering, Manipal Institute of Technology Bengaluru, Manipal Academy of Higher Education, Manipal,
India
d
Faculty of Management Studies, JAIN (Deemed-to-be University), Bengaluru, India
e
MURTI Research Center, Smart Agriculture Labs, Department of Artificial Intelligence and Data Science, GITAM School of Technology, GITAM
(Deemed to be) University, Bengaluru, India
f
Department of Computer Science and Engineering, GITAM School of Technology, GITAM (Deemed to be) University, Bengaluru, India

r e v i e w h i g h l i g h t s

• To summarize and classify different approaches of federated learning applicable to crop yield prediction along with their techniques, merits, and
demerits in a systematic manner.
• To examine the mathematical foundations of federated learning—including objective functions and optimization techniques—with a focus on
their application in agricultural contexts.
• Comparing the performance of various machine learning models employed in federated learning frameworks to predict crop yields for determining
which models work best and most efficiently.
• To present practical examples of implementations of federated learning in agriculture, while simultaneously identifying the challenges such as
data heterogeneity besides transmission overhead and propose future research directions.

a r t i c l e i n f o a b s t r a c t

Name of the reviewed methodology: The demand for food all over the world requires the implementation of advanced technologies
Federated learning to improve agricultural productivity. Federated Learning (FL) as a decentralized approach to ma-
chine learning facilitates collaborative model training on different data sources while maintaining
Keywords:
privacy—making it highly applicable technology for sensitive agricultural data. This paper offers
Federated learning
Crop yield prediction
a systematic overview of the recent knowledge on the application of FL towards the prediction
Machine learning of crop yield and other agricultural uses. We discussed the mathematical basis of FL, the variety
Precision agriculture of machine learning models used, the types of used agricultural data, and the major performance
Data privacy metrics. The paper presents real-world applications and lists the current limitations, including
Rice cultivation communication overhead, data heterogeneity, and interpretability issues. Lastly, we introduce
open research directions to inform the development of FL in precision agriculture.


Corresponding author.
E-mail address: [Link]@[Link] (R.M. Devadas).

[Link]
Received 26 April 2025; Accepted 29 May 2025
Available online 30 May 2025
2215-0161/© 2025 Published by Elsevier B.V. This is an open access article under the CC BY license
([Link]
V. Hiremani, R.M. Devadas, Preethi et al. MethodsX 14 (2025) 103408

Specifications table

Subject area: Computer Science


More specific subject area: Federated learning
Name of the reviewed methodology: Federated learning
Keywords: Federated Learning, Crop Yield Prediction, Machine Learning, Precision Agriculture, Data Privacy, Rice
Cultivation
Resource availability: NA
Review question: • What are the mathematical foundations and the latest mathematical contributions in Federated learning?
• What are the model type/frameworks, data sources, results contribute and limitations of past studies on crop
yield using Federated learning?
• How Federated learning has contributed region wise in terms of features, the dataset size.
• What other studies have contributed to the research by integrating Federated learning with latest
technologies?

Background

Due to increased food demand globally because of population growth and climate change, the agricultural sector is more chal-
lenging than ever. Crop yield prediction plays a very significant role in ensuring food security and efficient resource management in
agriculture [1]. Traditional crop yield prediction relies on centralized data processing, which may cause problems concerning data
privacy, accessibility, and the capability to take advantage of localized data [2]. Fig. 1 shows the yield of the wheat crop across
selected countries measured in tons. Federated Learning (FL) has arisen as a promising approach to address these challenges by
facilitating collaborative model training across various decentralized data sources, all while safeguarding the privacy of sensitive
information [3,4].
Federated learning is a decentralized paradigm of machine learning. Here, several clients can train a shared model together without
sharing their local data. In simpler words, as an alternative to uploading raw data to a central server, each of the clients computes
updates for its local dataset and shares only the model updates with the server [5].
Adopting FL improves the privacy of data and reduces issues of latency in those rural agricultural areas characterized by poor con-
nectivity [6]. Its applications based on IoT technologies in precision agriculture improve applicability since such technologies permit
real-time analysis and decision-making on any data [7]. The global federated learning market was estimated to be approximately
USD 119.4 million in 2022. Additionally, it is projected to expand at a compound annual growth rate (CAGR) of 12.7% from 2023
to 2030, as shown in Fig. 2.
Continuous novelty in machine learning practices and algorithms improves the efficiency of federated learning by a considerable
amount, thereby making it more attractive for various uses. As AI technology keeps advancing, federated learning must adapt to its

Fig. 1. Wheat yield (Source: Food and Agriculture Organization of the United Nations; Bayliss-Smith & Wanmali (1984); Brassley (2000); Broadberry
et al. (2015)).

2
V. Hiremani, R.M. Devadas, Preethi et al. MethodsX 14 (2025) 103408

Fig. 2. FL market (Source: APAC market CAGR).

Fig. 3. FL model lifecycle and system actors [8].

Table 1
Types of federated learning.

Type Description Use Cases

Centralized Federated Learning (CFL) A server collects and aggregates model updates from clients [9] Mobile applications, healthcare data
Decentralized Federated Learning (DFL) No central server; clients communicate directly with each other [10]. Peer-to-peer applications, IoT devices
Vertical Federated Learning (VFL) Different parties hold different features of the same dataset [11] Collaborative research, multi-party analytics
Horizontal Federated Learning (HFL) Different parties hold the same features but on different datasets [12] Federated learning in agriculture

superior techniques, making it improve on the efficiency of training the model across distributed devices. Therefore, this continuous
process will make federated learning attractive and enable it to support different industries, like health and finance, while providing
very strong data privacy throughout the decentralized networks.
The Federated Learning (FL) process involves identifying a problem, collecting and storing client data, prototyping in simulations
(if needed), federated model training, and evaluation. Once a suitable model is selected, it undergoes deployment through quality
assurance, A/B testing, and a staged rollout as shown in Fig. 3.
Table 1 categorizes the different types of federated learning, highlighting their characteristics and applicable use cases in various
domains.
Table 2 outlines the primary challenges faced in federated learning implementations, detailing their implications for model per-
formance and data management.
The schematic diagram (Fig. 4) gives a relative visualization of federated and centralized learning frameworks for farm-level
crop yield estimation. The centralized learning depiction, on the left, illustrates the central server receiving individual farm data,

3
V. Hiremani, R.M. Devadas, Preethi et al. MethodsX 14 (2025) 103408

Table 2
Common challenges in federated learning.

Challenge Description

Data Heterogeneity Differences in data distribution across clients can lead to biased models [13]
Communication Bottlenecks Limited bandwidth may slow down the aggregation of model updates [14]
Security Risks Vulnerabilities in model updates can expose sensitive information [15]
Scalability As the number of clients increases, maintaining efficiency becomes difficult.

Fig. 4. Federated and centralized learning.

which is then trained and redistributed back to farms after being trained in a global model. This arrangement, though simple, creates
concerns about data privacy, communication costs, and scalability. To its right, a federated learning system is illustrated, where
every farm trains a local model on local data separately. Model updates (e.g., gradients or weights) are sent from these farms to the
central server, which takes these updates and aggregates them into an enhanced global model, which is then sent back to farms.
The integration of icons for security acknowledges federated learning’s privacy-preserving nature such that raw data never transmits
outside local devices. Technically, this system minimizes bandwidth consumption, facilitates heterogeneity in data, and increases
trust and openness—features well-suited for agricultural applications that have distributed and sensitive data.
Federated learning has emerged as a central framework in recent studies, facilitating collaborative machine learning while over-
coming privacy, regulatory, and data silo issues. The following is an elaboration of its role and importance as established in peer-
reviewed research papers:

Role of federated learning

• Privacy-Preserving Collaboration: FL allows for several different entities (e.g., hospitals, devices, or organizations) to jointly train
machine learning models without sharing raw data, maintaining privacy in the data and following directives such as GDPR and
HIPAA [16,17].
• Distributed and Decentralized Learning: FL facilitates distributed learning from decentralized sources of data such that knowledge
aggregation from various environments is done while leaving data localized [18].
• Application in Multiple Domains: Examples that prove FL’s applicability across various domains include healthcare (e.g., privacy-
enhancing medical AI), IoT (e.g., network intrusion detection), smartphones (e.g., prediction-based texting), and finance, etc.
[19].

Relevance in current research

• Scalability and Real-World Deployments: FL has been scaled up to millions of devices, supporting production applications across
organizations such as Google and Apple, and has been shown to provide meaningful privacy assurances through methods like
differential privacy [20].
• Overcoming Data Silos: FL enables sensitive or proprietary-data-containing institutions (e.g., hospitals, banks) to cooperate and
enhance model accuracy while respecting sharing restrictions, thereby unlocking siloed data value [21].
• Algorithmic Innovations: Research has developed sophisticated aggregation algorithms (e.g., FedAvg, FedProx) and
communication-efficient protocols for overcoming federated settings’ special circumstances, including communication cost and
client data heterogeneity [22].

Federated learning has exhibited great promise in numerous real-world agricultural use cases, including

4
V. Hiremani, R.M. Devadas, Preethi et al. MethodsX 14 (2025) 103408

• Multi-crop Yield Prediction: FL has been utilized to forecast yields for key crops like maize, wheat, rice, and soybean from
geographically dispersed farms. This allows for model training collectively without violating farm-specific data privacy.
• Pest and Disease Detection: FL models trained using distributed image data (e.g., from smartphone or drone images) detect crop
diseases such as rust in wheat or potato late blight, allowing intervention at an early stage while maintaining data locality.
• Precision irrigation and water management: Farms can utilize sensor-inclusive farm networks and FL to collectively model optimal
irrigation schedules to conserve water and optimize schedules without centralizing raw sensor data.
• Soil Health Monitoring: FL models can be trained on spatially distributed data for soil properties (such as nutrient status, pH, and
texture) to create targeted fertilization policies without compromising data confidentiality.
• Livestock and Farm Resource Management: FL facilitates aggregation of productivity data and health monitoring data from live-
stock farms (e.g., dairy production, detection of diseases) to assist decision making while protecting sensitive operational infor-
mation.
• Climate-Smart Agriculture: FL promotes inter-regional collaboration to simulate the impact of climates on crops, enabling farmers
to gain from knowledge generated from varying climates while ensuring local data remains secured.
This review is motivated by an increased interest in federating learning techniques in predicting the yields of crops, but most
especially staple crops like rice, which are fundamental crop staples for food security globally. Recent studies indicate the capability
of FL to boost up the accuracy of prediction even with privacy-preserving data handling, where traditional methods cannot [23]. This
paper provides an exhaustive review of federated learning applications in crop yield prediction and their present state of research.

Research gaps and contributions

Although interest in Federated Learning has been on the rise in agriculture, no in-depth and specialized review focusing on FL
methodologies in the domain of crop yield prediction currently exists. Reviews of the available literature tend to give an overall
summary of machine learning in agriculture without considering the special problems of FL—data privacy, model transfer to hetero-
geneous sources, and communication overhead during the training process.
Moreover, most FL-based work targets specific types of crops, e.g., soybeans, while generalizability to broader agricultural frame-
works remains limited. Comparative evaluation of FL methodologies versus baseline (centralized/traditional) methodologies has also
been largely neglected. Discussion of the potential convergence of FL with future agricultural technologies, e.g., IoT networks and
remote sensing setups, which can immensely boost the accuracy of the models as well as enable real-time decision-making capacities,
has been rare.
This review addresses these gaps by:
• Focusing exclusively on FL for crop yield prediction
• Analyzing and comparing various FL models and performance metrics
• Highlighting challenges like privacy, convergence, and system heterogeneity
• Exploring integration of FL with IoT and remote sensing for agricultural applications

Motivation

The rising demand for efficient and sustainable agricultural practices has pushed the implementation of data-driven technology
to improve yield prediction. Traditional centralized machine learning methods are hindered by issues of data privacy, ownership,
and the costs of transmission, particularly when handling region-specific and sensitive agricultural information. Federated Learning
provides a compelling solution by allowing collaborative model training on decentralized data sets while maintaining privacy. This
shift in paradigm provides a chance to enable farmers and agricultural institutions to become empowered with predictive information
while maintaining the security of the data.

Objectives

The overall goal of this review is to critically evaluate the application of Federated Learning to agricultural crop yield forecasting
with a focus on its mathematical basis, optimization approaches, and incorporation in machine learning models. To this end, the
review seeks to
(i) Highlight major mathematical background and terminology employed in FL to enhance interpretability to a wide range of
audiences
(ii) Compare the performance of different ML models used in FL frameworks for the task of crop yield forecasting,
(iii) Identify translational applications to real-world agricultural sectors outside of specific crop types
(iv) Outline current challenges and knowledge gaps to inform the next generation of FL development in precision agriculture.
Below are the review questions framed to carry out this work
• What are the mathematical foundations and the latest mathematical contributions in Federated learning?
• What are the model type/frameworks, data sources, results contribute and limitations of past studies on crop yield using Federated
learning?
• How Federated learning has contributed region wise in terms of features, the dataset size.
• What other studies have contributed to the research by integrating Federated learning with latest technologies?

5
V. Hiremani, R.M. Devadas, Preethi et al. MethodsX 14 (2025) 103408

Table 3
List of notations used.

Symbol Description

(x) Global model parameters


(xt ) Model parameters at iteration t
(xtk ) Local model parameters of client k at round t
(F(x)) Global objective function
(Fi (x)) Local objective function at client i
(ni ) Number of data samples at client i
(n) Total number of data samples across all clients
(α) Learning rate
(hi ) Local loss function
(∇hi ) Gradient of the local loss function
(J) Number of clients participating in a round
(mj ) Number of data samples on client j
(m) Total number of data samples across selected clients
(l(⋅)) Loss function (e.g., cross-entropy, MSE)
(f (⋅)) Prediction function (e.g., neural network, SVM)
(xk,i ) i-th input sample at client k
(yk,i ) Target label corresponding to x_{k,i}
(Δxit ) Difference between client i’s model and global model at round t
(μ) Proximal term coefficient in FedProx
(θ) SVM model parameters
(𝜙(⋅)) Kernel or feature mapping function in SVM
(ξi ) Slack variable in SVM optimization
(C) Regularization parameter in SVM

Method details

Question 1: What are the mathematical foundations and the latest mathematical contributions in Federated learning?
Table 3 represents a list of notations used as part of mathematical background. The fundamental strategy in federated learning
involves the collaborative training of a global model using diverse and distinct decentralized datasets. Thus, it minimizes local
objectives with gradient descent to generate models at clients and, using the federation of updates through a weighted average
periodically, generates global models that accumulate models through aggregating client-updates weighted average of. In effect, global
objectives ensure all proportionate influence. To reduce communication overhead, clients send just model updates (differences) rather
than raw data. The local objective function ensures that all clients optimize their models based on their respective data characteristics.
The goal of FL is to minimize a global objective function(as shown in Eq. (1)) across all participating devices [24].
∑D
dni
min F(𝐱) = F (𝐱) (1)
𝐱
i=1
d i

Where, Fi (x) represents the local objective for client i and ni and n are the data size for client i and the total data size, respectively.
This formulation ensures that clients contribute proportionally to the global model based on their data volume.
Each client updates its local model by applying T steps of gradient descent [25] as shown in Eq. (2):
( )
ut+1 = ut − α∇hi ut (2)
Here, u being the model and 𝛼 for learning rate. Gradient descent helps adjust model weights in the direction that minimizes
the loss. In one of the early study (FedMDFG), fairness is addressed by introducing a multi-objective optimization framework with
a fairness-driven objective. Their method refines gradient descent by ensuring a fair descent direction (via cosine similarity) and
optimizing the step size for improved fairness and convergence in federated learning as proven in Fig. 5.
Next equation represents the global model update in federated learning, combining the weighted contributions of local updates
from all clients [27]. This update rule ensures that models with larger data volumes contribute more to the global model as shown
in Eq. (3).
∑𝐽
𝑚𝑗 ( 𝑗 )
U𝑡+1 = U𝑡 + U𝑡+1 − U𝑡 (3)
𝑗=1
𝑚

Here,

Ut+1 : Global model parameters at round t+1


𝜃 t : Global model parameters at round t
J: Total number of participating clients
mj : Number of data samples for client j
𝐽

m= 𝑚𝑗 The overall count of data samples from all clients
𝑗=1

6
V. Hiremani, R.M. Devadas, Preethi et al. MethodsX 14 (2025) 103408

Fig. 5. Comparison of various federated learning algorithms on the CIFAR-100 and MNIST datasets [26].

Fig. 6. Comparison of FedAvg and FedSGD [28].

𝑈𝑡𝑗+1 Updated local model parameters for client j after round t+1

In one of the past works authors show that Federated Averaging (FedAvg) demonstrates superior performance with fewer local
epochs (e.g., E=1E = 1 vs. E=5E = 5) and exhibits lower accuracy variance across evaluation rounds compared to Federated Stochastic
Gradient Descent (FedSGD) as shown in Fig. 6.
The loss function in federated learning is a local objective. It represents the loss to be optimized by the different clients for their
respective dataset [24]. Aggregation of those local objectives leads to minimizing a global objective function. Thus, knowledge from

7
V. Hiremani, R.M. Devadas, Preethi et al. MethodsX 14 (2025) 103408

Fig. 7. Robust FL performance [29].

all clients will remain reflected in the global model while data privacy is held in place as shown in Eq. (4).

1 ∑ ( (
dk
) )
Lk (u) = l f u; xk,i , yk,i (4)
𝑑k i=1 k

Here,

Lk (u): Local objective function for client k.


w: Model parameters.
dk : Number of data samples for client k.
xk,i : Input data sample i from client k.
yk,i : Corresponding label for xk,i .
𝓁 k : Loss function for client k.
f(w;xk,i ): Model’s predicted output for input xk,i .

This ensures models adapt to local patterns while maintaining alignment with the global objective. In one study, the local objective
function incorporated a reweighting mechanism to address noisy-labeled data effectively. By assigning appropriate local weights to
the loss function, edge devices collaboratively mitigated the impact of noise, resulting in improved prediction accuracy (shown in
Fig. 7) and robustness in both full-data and part-data network setups.
The Communication Update (Model Differences) captures the difference of the locally trained model as compared to the global
model, which is then allocated across the server. It serves as a difference critical in efficiently aggregating the update and ensuring
that all such improvements are reflected through this global model while communicating reasonably as shown in Eq. (5).

Δuit = uit − ut (5)

This delta update helps the server efficiently reconstruct model changes without transferring the full model. In federated learning,
loss function selection is critical as it affects the performance of the model, especially when there is statistical heterogeneity and noisy
data. The FedProx algorithm has stability by incorporating a proximal term that, together with corrected loss functions, improves the
robustness of local models against noise and bias(shown in Fig. 8). This synergy would effectively mitigate the effects of noisy data
on convergence leading to better accuracy and stability in federated learning scenarios.
This section will explore various machine learning models applicable in federated learning, particularly for agricultural tasks such
as crop harvest prediction besides disease detection. In the context of federated learning for crop yield forecast, several machine
learning models have been commissioned to enhance accuracy and robustness against data heterogeneity. Linear Regression is often
utilized for its simplicity and interpretability [31], represented by the Eq. (6):

y = β0 + β1 x1 + β2 x2 + … + β𝑛 xn + 𝜖 (6)

where y is the predicted yield, xi are the input features, and 𝜖 represents the error term. The study [32] utilizes AI-based Federated
Learning and multi-regression analysis to estimate crop loss, highlighting the benefits of diversifying crops for smallholder farmers
on limited land. FL facilitates an irrigation advisory system using Logistic Regression, employing client-server architecture within
the Flower framework to aggregate model parameters from multiple clients. This system predicts irrigation needs based on diverse
agricultural data inputs like temperature and soil moisture [33]. Decision Trees present a non-linear methodology by splitting data
into branches constructed on feature values, allowing for complex decision-making processes. Random Forests, an ensemble method
that merges multiple decision trees, have shown significant improvements in predictive performance. For instance, a comparison
study highlighted that Random Forests achieved an R2 score of 0.946 when predicting rice crop yields in Karnataka [34].

8
V. Hiremani, R.M. Devadas, Preethi et al. MethodsX 14 (2025) 103408

Fig. 8. Comparing convergence rates of FedProx [30].

Table 4
Inclusion and exclusion criteria.

Inclusion Criteria 1. “Federated learning” AND “Crop Yield”


2. “Federated learning” AND “Agricultural yield
3. “Federated learning” AND “Agricultural produce”
4. “Federated learning” AND “Harvest”
5. “Federated learning” AND “Crop production”
6. “Federated learning” AND “Crop harvest”
7. "Federated Learning" AND "Explainable AI" OR "XAI"

Exclusion Criteria 1. “Federated learning” AND “Crop disease”


2. “Federated learning” AND “Crop disease prediction”
3. “Federated learning” AND “Crop disease detection”

Support Vector Machines (SVM) are also notable for their effectiveness in classification tasks within federated learning frameworks
[35]. The optimization problem for SVM can be expressed as as shown in Eq. (7):

1 || ||2 ∑n
min ||𝜽|| + D mi (7)
𝜽 2 || ||
i=1

where D is the penalty parameter and mi represents the slack variables. This formulation maximizes the margin while allowing for
some misclassifications. In FL, each client computes gradients locally and sends parameter updates or dual variables to the server for
secure aggregation.
This approach has been successfully applied to classify healthy versus diseased crops, demonstrating robust performance in hetero-
geneous data environments [36]. Deep learning models, especially Convolutional Neural Networks and Recurrent Neural Networks,
have become widely used in the context of complex pattern extraction from big data. For instance, CNN architectures are successfully
applied in image classification problems concerning crop disease identification and greatly surpass traditional methods. In terms of
performance metrics, several studies have compared these models within federated learning contexts. One study reported that using
federated averaging with the ResNet-16 results in an accuracy of 92.5% in comparison to the RMSE values where it is lower than
traditional centralised models [3]. Another analysis showed that an LSTM and Bi-LSTM model resulted in an accuracy of 93.7%, with
precision and recall values around 92 and 91, respectively [37]. In summary, the integration of these machine learning models within
federated learning frameworks is promising for improving crop yield prediction while addressing challenges due to heterogeneity
and privacy.

Literature search methodology

To ensure a comprehensive and systematic review, we conducted a structured literature search using multiple academic databases,
including IEEE Xplore, SpringerLink, ScienceDirect, ACM Digital Library, Google Scholar, and arXiv. The search covered publications
from January 2017 to March 2025, encompassing peer-reviewed journal articles, conference papers, and relevant preprints. Table 4
displays the search terms used as part of the inclusion and exclusion criteria

9
V. Hiremani, R.M. Devadas, Preethi et al. MethodsX 14 (2025) 103408

Filtering and inclusion criteria


• Only English-language publications were considered.
• Studies were included if they addressed federated learning in agricultural or environmental contexts.
• For evaluation and benchmarking, we selected studies reporting at least one performance metric (e.g., RMSE, accuracy, R2 ).
• Duplicates and purely theoretical papers without empirical evidence were excluded.
Evaluation criteria and interpretation
In evaluating the performance of machine learning models within federated learning frameworks for crop yield prediction, several
metrics are considered to assess both accuracy and operational efficiency. These include:
• Root Mean Squared Error (RMSE): Measures the average magnitude of prediction errors. Lower RMSE values indicate better
predictive performance.
• Accuracy: Primarily used in classification tasks to determine the proportion of correct predictions.
• Precision and Recall: Important for evaluating classification models, particularly in detecting specific crop conditions or events.
• R-squared (R2 ): Represents the proportion of the variance in the dependent variable that is predictable from the independent
variables.
• Convergence Speed: Reflects how quickly the model reaches a stable performance during training, important in communication-
constrained FL settings.
• Communication Overhead: Quantifies the amount of data exchanged between clients and the central server, a key efficiency metric
in federated learning.
These metrics provide a multi-faceted view of model performance across diverse data types and agricultural contexts.
Question 2: What are the model type/frameworks, data sources, results contributed and limitations of past studies on crop yield
using Federated learning?
For reasons of dataset heterogeneity across the studied papers in terms of crops, regions, and testing regimes, direct numeric
comparison of measures like accuracy or RMSE might not always make sense. Therefore, Table 5 emphasizes the presentation of
relative performance (improvement compared to the baseline in the same paper) and contextual information (testing setups and
model employed as well as the type of data) to enable fair and informative study-to-study comparison.
Table 4 displays studies carried out in crop yield.
Table 5 depicts a holistic overview of the crop yield prediction studies including the models/frameworks employed, the source of
the data, the results obtained, and their limitations. Covering the period from 2019 to 2024, the table showcases both federated and
centralized AI strategies in various agricultural environments.
Analysis
1. Model Types and Frameworks
• A range of deep learning methodologies are utilized, ranging from CNNs to RNNs, LSTMs, DRQN, and ensemble models.
• Centralized and decentralized frameworks of federated learning play a central role in the latest research, mirroring a trend

towards privacy-enhancing, distributed AI.


• Other sophisticated approaches including sICA, AdaBoost, and hybrid ensemble classifiers are also employed, describing

methodological diversity in the approach to yield prediction.


2. Data Sources
• Data types vary from public sets of data (e.g., soil, weather, and crop information) to remote sensing images of high resolution,

hyperspectral and LiDAR information, and ICRISAT datasets.


• Several studies focus on multi-modal or region-dependent data, which calls for localized and contextual AI models in agricul-

ture.
3. Achievements Realized
• Models exhibit excellent forecasting capability:

➢ RMSE values ranging from 0.496 t/ha (Peng et al.)


➢ Up to 97% accuracy in federated learning (Mukherjee & Buyya)
➢ R2 values ranged from 0.82 to 0.93 for maize (Aviles Toledo et al.)
➢ Scores of F1, precision, and recall over 90% in ensemble systems (Tripathi & Biswas)
• These results validate the capability of AI models to make accurate forecasts of crop yield, especially when fueled by substantial

and quality-focused datasets.


4. Limitations Identified
• Data heterogeneity and communication expenses are ongoing issues, particularly in federated learning.
• Some models suffer from limited generalizability because of regional specificity or data bias.
• Others find difficulty with the multi-source data integration complexity or the rigidity of some algorithms (e.g., the limited

state description of Elavarasan et al. using the Q-learning algorithm).


CNNs and RNNs are used for different applications in agricultural analytics. CNNs are specifically useful in spatial data tasks like
satellite imaging interpretation or farm image acquisition by drones to evaluate plant health or identify disease. RNNs—such as LSTM
types—are suitable for temporal sequence modeling and are useful in forecasting trends in yield from time-series data like rainfall,
temperature, and phenological data. These can be trained locally in federated learning architectures with respect to applicable data

10
V. Hiremani, R.M. Devadas, Preethi et al. MethodsX 14 (2025) 103408

Table 5
Studies on crop yield.

Author, Year Model Type/ Framework Data Sources Results Achieved Limitations

Khaki, S., et al., merges Convolutional Neural Publicly available data. RMSE (Corn):9% The model’s performance was
(2019) [38] Networks (CNNs) besides (weather data, soil, and RMSE (soybean):8% for sensitive to various factors, including
Recurrent Neural Networks management data) yields weather, soil, and management
(RNNs) practices, indicating that external
conditions could impact its accuracy
Aviles Toledo, Long Short-Term Memory (LSTM) High-resolution R2 (maize): 0.82 to 0.93 The findings are not generalizable to
et al. (2024) hyperspectral imagery, other crops or agricultural contexts,
[39] LiDAR point clouds, and data integration and processing
are complex. It does not address
potential biases in the data collection
process.
Elavarasan, D., Deep Recurrent Q-Network multi-modal data sources Accuracy: 93.7%, The Q-learning method has
et al. (2020) (DRQN), Recurrent Neural limitations in describing states, which
[40] Network (RNN), can hinder effective crop yield
prediction.
Pham, H.T., Higher-order spatial independent satellite-based vegetation Prediction limit its applicability to other crops or
et al. (2022) component analysis (sICA) condition indices (VCI) accuracy:18.5% to 45%, regions, the optimal number of
[41] and thermal condition average error:5%, subregions is not predetermined and
indices (TCI), Historical can vary, potentially affecting model
VCI and TCI time series performance.
data, administrative
maps
Peng, Dailiang, Deep learning model Weather Forecast Data, Root Mean Square Error limited to China’s main
et al. (2024) Remote Sensing Data (RMSE):0.496 t/ha wheat-producing areas
[42]
Tripathi, D., & Precise Ensemble Expert System International Crops Accuracy: 90.11%, f-1 Complexity of Factors, Data
Biswas, S.K. for Crop Yield Prediction Research Institute for the score:90.20%, Dependency, limited to the specific
(2024) [43] (PEESCYP), Multiple Imputation Semi-Arid Tropics precision:90.39%, recall: datasets used for training
by Chained Equations (MICE), (ICRISAT), 90.10%.
T, M., federated learning framework scattered and siloed data soybean yield prediction data heterogeneity and
Makkithaya, K., related to weather, soil, communication costs between client
& G, N.V. (2022) and crop management devices, performance may vary
[3] grounded on the availability and
accuracy of the data
Mukherjee, A., & centralized and decentralized Publicly available data. Centralized federated Potential limitations could include
Buyya, R. federated learning frameworks learning frameworks challenges in data heterogeneity
(2024) [9] prediction accuracy: among clients, which may affect
≥97% and decentralized model performance.
federated learning Communication overhead in
frameworks prediction decentralized frameworks.
accuracy: ≥97%
j, J., Koh, J.G., & AdaBoost, Linear Discriminant Climatic conditions in Prediction for various Potential limitations could include
Lee, S.K. (2023) Analysis Classifier, Stochastic meteorological data ensemble model the generalizability of the results
[44] Gradient Boosting Classifier, across different climates or regions
Blending Models, Quadratic
Discriminant Analysis Classifier,
Random Forest Regressor,
Bagging Classifier, Extra Trees
Classifier
Li, A., Markovic, federated learning framework National Agricultural Improved performance data heterogeneity, communication
M., Edwards, P., Statistics Service of the by 15.5% to 20%, Local costs, limit the generalizability of the
& Leontidis, G. United States model sizes were results
(2023) [45] reduced by up to 84%,

types (e.g., images or time-series) and contribute to a global model, thus allowing spatial and temporal insight in a privacy-preserving
fashion.
Question 3: How Federated learning has contributed region wise in terms of features, the dataset size.
Table 6 indicates region wise study in federated learning.
Table 6 provides a region-wise overview of some AI-based agricultural research, including the geographical focus, the nature of
the data sets, and the aspects considered in each. The dates range from 2019 to 2024 and represent an international diversity with
contributions from the United States, India, Vietnam, China, Korea, Brazil, and Argentina.
Analysis

1. Regional Coverage
• The United States appears most frequently, particularly the Midwestern region, underscoring its role as a hub for precision

agriculture research.

11
V. Hiremani, R.M. Devadas, Preethi et al. MethodsX 14 (2025) 103408

Table 6
Region -wise study.

Paper title/Study Region Dataset size Features included

Khaki, S., et al. (2019) Midwestern United 300GB Genetic Improvements, Management Practices
[37] States
Aviles Toledo, et al. Midwestern United high-resolution hyperspectral imagery, LiDAR Genetic Data, Temporal Features
(2024) [38] States point clouds, and environmental data
collected over a two-year period
Elavarasan, D., P. M. India 200,000 records Agricultural Practices,
DURAIRAJ VINCENT
Member, I., & Durairaj,
P.M. (2020) [39]
Pham, H.T., et al. (2022) Vietnam Spatio temporal resolution of 4 km and a Normalized Difference Vegetation Index
[40] 7-day composite (NDVI), Land Surface Temperature (LST)
Peng, D., et al. (2024) China 40 days prior to harvest and 25 days of future Remote Sensing Data, Temporal Data
[41] weather forecast data
Tripathi, D., & Biswas, India 15-year period (2003–2017) of all Assam’s Socioeconomic Factors, Historical Yield
S.K. (2024) [42] districts Records
T, M., Makkithaya, K., & India 2TB WRF-HRRR Computed Dataset, Geospatial
G, N.V. (2022) [3] Data
Mukherjee, A., & Buyya, United States 5-year data Federated Learning Frameworks, Application
R. (2024) [9] Context
j, J., Koh, J.G., & Lee, Korea Raw data Federated Learning Frameworks, Machine
S.K. (2023) [43] Learning Frameworks
Li, A., et al. (2023) [44] United States, Brazil, The size of the dataset may vary significantly Federated Learning Frameworks, Localized
Argentina. dependent on the number of participating Data Processing
farms plus the amount of data each farm
contributes.


India is also a significant focus, with multiple studies analyzing large-scale agricultural practices, historical data, and environ-
mental factors.
• Other countries represented include Vietnam, China, Korea, and South American nations, indicating a growing global interest

in applying advanced AI and data-driven methods to regional agricultural challenges.


2. Dataset Scale and Type
• Dataset sizes are wide-ranging—raw data or multi-year histories (e.g., 15 years of Tripathi & Biswas) to enormous amounts

such as 2TB (T. Makkithaya et al.) and 300GB (Khaki et al.).


• Numerous analyses point to the trends of utilizing high-resolution multimodal sources (e.g., Aviles Toledo et al. employ-

ing hyperspectral and LiDAR imagery) toward increasingly dense, multidimensional data sets to support finer granularity of
analysis.
3. Specifications Encompassed
• Temporal information, remote sensing indicators (such as NDVI and LST) and environmental/management-related factors are

typically incorporated, exemplifying the multifactor and time-dependent nature of agricultural predictive tasks.
• Several investigations combine the socioeconomic and local aspects, recognizing the contextual nature of yield estimation and

optimization.
• Federated learning frameworks make an appearance in latter entries along with localized processing of the data, reflecting a

strategic shift towards privacy-conscious, distributed AI solutions designed for a variety of different geographic regions.

Question 4: What other studies have contributed to the research by integrating Federated learning with latest technologies?
Table 7 details study conducted on integration of technologies with federated learning
The Table 7 presents a comparative summary of recent scholarship (2019-2025) on the incorporation of emerging technologies
used in machine learning and data analysis, with a focus on the application to the domains of IoT, remote sensing, and federated
learning (FL). The analysis indicates the integrating technologies utilized, as well as the primary benefits derived from their applica-
tion.
Analysis

1. Technological Integration Trends


• A significant focus on Deep Learning (DL) emerges in most of the studies and usually in association with IoT, cloud computing,

remote sensing, and UAVs, demonstrating its core importance to enhance model precision.
• Federated Learning emerges as a central theme in later work (2023–2025) and is often paired with Explainable AI (XAI), at-

tention mechanisms, and decentralized designs. This trend reflects increased interest in addressing concerns over data privacy,
interpretability, and efficiency in a distributed setting.
2. Advantages Achieved
• Improved Accuracy is the most frequently occurring result, being seen in almost all the entries, confirming the usefulness of

AI-based solutions for predictive modeling.

12
V. Hiremani, R.M. Devadas, Preethi et al. MethodsX 14 (2025) 103408

Table 7
Studies with integration of technologies.

Author, Year Integration Technology Benefits realized

Khaki, S., et al. (2019) [37] Deep Learning, IoT, Remote Sensing Improved Accuracy
Aviles Toledo, C., Crawford, Multi-Modal Remote Sensing, Deep Learning Enhanced Forecast Accuracy, Interpretability
M.M., & Tuinstra, M. (2024) [38] Architectures, Attention Mechanisms
Elavarasan, D., P. M. DURAIRAJ Deep Learning Models, Cloud Computing Increased Accuracy, Resource Optimization
VINCENT Member, I., & Durairaj,
P.M. (2020) [39]
Pham, H.T., Awange, J.L., Kuhn, Satellite Imagery, Machine Learning Models, Enhanced Accuracy, Resource Optimization
M., Nguyen, B., & Bui, L.K. Phenological Metrics
(2022) [40]
Peng, D., et al. (2024) [41] Deep Learning Models, Multi-Source Data Fusion, UAV Improved Accuracy, Timely Decision-Making, Adaptation to
Remote Sensing Climate Variability
Tripathi, D., & Biswas, S.K. Ensemble Learning Techniques, Data Analytics Improved Accuracy, Adaptability
(2024) [42]
T, M., Makkithaya, K., & G, N.V. Collaborative machine learning across decentralized Data Privacy and Security, Improved Prediction Accuracy,
(2022) [3] data sources Reduced Latency, Resource Efficiency
Mukherjee, A., & Buyya, R. Centralized and Decentralized Federated Learning Data Privacy
(2024) [9]
j, J., Koh, J.G., & Lee, S.K. (2023) Federated Learning, Ensemble Models Improved Accuracy, Scalability
[43].
Li, A., Markovic, M., Edwards, P., Model Pruning, Decentralized Architecture, Reduced Communication Costs, Energy Efficiency, Localized
& Leontidis, G. (2023) [44] Communication Protocols Adaptation
Lopez-Ramos et al., (2024) [45] FL and Explainable Artificial Intelligence Allows for privacy-preserving model training from distributed
data while enhancing interpretability and trust through
explanation methods
Rahmati, Mohammad (2025) FL and Explainable Artificial Intelligence Enhanced Predictive Accuracy, Data Privacy Protection,
[46] Interpretability of Models, Optimized for Low-Resource
Environments
Briola et al., (2024) [47] FL and Explainable Artificial Intelligence Enhanced User Privacy, High Accuracy and Performance,
Explainability Maintained
Lawrence Anebi Enyejo et al., FL, computational geometry and Explainable Artificial Enhances model transparency and interpretability, fostering
(2024) [48] Intelligence trust and accountability in decentralized systems
Daole, M et al., (2024) [49] FL, explainable classifiers, Fuzzy Rule-based Classifiers Offers advantages in classification performance while
preserving data privacy for participants in heterogeneous
settings
Uppu Nirosha, G. Vennila (2025) FL, attention-based graph neural networks, recurrent High precision, high correlation
[50] neural networks


Optimization of resources, privacy of data, and interpretability become key advantages in contemporary deployments, partic-
ularly in federated and decentralized settings.
• Certain studies bring to the forefront specific results including lower communication expenses, low-resource optimization, and

timely decision making, demonstrating the operational benefits of these integrations.


3. Emerging Focus Areas:
• The union of FL with Explainable AI is taking center stage (Lopez-Ramos et al., Rahmati, Briola et al.), providing a performance-

transparency balance, especially in heterogeneous and privacy-sensitive settings.


• The incorporation of latest architectures such as attention-based graph neural networks and fuzzy rule-based classifiers signals

a trend towards increasingly adaptive and interpretable AI systems.

Models utilizing federated learning frameworks achieve consistently higher accuracy (≥97%) compared to traditional machine
learning methods, indicating their robustness in diverse datasets and contexts as shown in Fig. 9.
It can be seen from Fig. 9, the Federated (Centralized) and Federated (Decentralized) models reflect the greatest precision, both
with almost 100%. DRQN and Ensemble models are next with slightly reduced precision at a lower level than the first two. CNN-RNN
has slightly lower precision than Ensemble. LSTM has the lowest precision of the compared varieties, slightly above 80%. The graph
indicates federated learning models, both decentralized and centralized, to be superior to conventional and hybrid deep models in
this specific setting.
It is evident from Fig. 10 that accuracy trends vary significantly by region, with federated learning and ensemble models yielding
the highest performance in the United States, highlighting the role of localized data.
Technologies like deep learning and IoT are most frequently integrated, showing their wide applicability and scalability, while
federated learning is emerging for privacy-sensitive contexts as evidenced in Fig. 11.
Table 8 outlines a comparison of our study with state-of-the art studies.
Comparative Table 8 provides a systematic analysis of two notable review articles compared to the proposed review. Although
study 1 provides domain-related information specific to agriculture, the discussion is concept-oriented and restrictive about crop
diversity and model performance. Study 2, though broader in the domain scope, has only broad observations and lacking in domain-
related depth, especially in mathematical foundations and optimization strategies. On the contrary, the proposed review stands

13
V. Hiremani, R.M. Devadas, Preethi et al. MethodsX 14 (2025) 103408

Fig. 9. Accuracy by model type.

Fig. 10. Average accuracy by region.

out with its holistic and agriculture-oriented treatment of FL, including exhaustive mathematical expressions, incorporation of new
technologies like IoT and remote sensing, relative performance evaluation across a wide variety of ML models, and specific discussion
on data privacy, communication overhead, and heterogeneity. It further depicts distinct research gaps and targeted directions, setting
it apart as the field’s unique contribution.

Observations

The deep learning models, such as CNN, RNN, and LSTM, have shown consistent improvements in the accuracy and RMSE of
crop yield prediction across datasets. For instance, Mukherjee & Buyya (2024) [9], reported a prediction accuracy ≥97% using
federated learning frameworks. Technologies such as IoT, remote sensing, ensemble models, and federated learning combine to
enhance accuracy, scalability, and data privacy. For example, Khaki et al. (2019) [37] achieved higher accuracy with IoT and remote
sensing, while Li et al. (2023) [44] reduced communication costs by up to 84% using model pruning. Most models are limited
geographically, for instance, Peng et al. (2024) [41] focused on wheat producing areas of China, while Tripathi & Biswas (2024) [42]
proposed methods for estimating crop yield for crops like rice and wheat in India.

14
V. Hiremani, R.M. Devadas, Preethi et al. MethodsX 14 (2025) 103408

Fig. 11. Integration technology usage frequency.

Table 8
Comparative review with existing studies.

Aspect Study 1 [51] Study 2 [52] Proposed review

Scope and Domain Coverage Federated Learning in Agriculture FL across Healthcare, Agriculture, Federated Learning for Crop Yield
Education Prediction
Focus on Agriculture Yes (Agriculture-specific) Partially (One of three domains) Yes (Extensively agriculture-focused)
Mathematical Foundations & Discussed at conceptual level General overview, less Detailed equations and optimization
Optimization mathematical depth formulations included
Integration with Emerging Tech Covered (IoT, edge computing, General discussion, not Extensively discussed with examples
(IoT, Remote Sensing) sensor networks) agriculture-specific
Crop Diversity & Generalizability Focused primarily on limited Not emphasized Covers multiple crops beyond rice (e.g.,
crops like rice and wheat maize, soybean)
Model Performance Comparison Limited comparison across Not focused on performance Comparative analysis across models and
models metrics datasets
Discussion on Data Privacy & Covered briefly Discussed cross-domain FL issues Discussed in detail with solutions
Communication
Frameworks & Model Types Overview of FL frameworks and Broad architectural classifications Diverse ML models: SVM, LSTM, CNN, etc.
Reviewed types (HFL, VFL, CFL)
Challenge Identification General challenges mentioned Highlights high-level concerns Specific to FL in agriculture: data, comm.,
heterogeneity
Future Directions Outlined broadly Focused on cross-domain trends Research gaps + targeted future research
areas

As Aviles Toledo et al. (2024) [38] would explain, multi-modal and high-resolution data sources, such as hyperspectral imagery
and LiDAR, improve the prediction accuracy but at the cost of increased complexity. Most of the models fail to generalize across
regions and crop types due to heterogeneity in data, variation in external conditions, and biases in data collection. For example,
Pham et al. (2022) [40] noted the limited applicability of the model to other crops or regions. Federated learning frameworks are
much effective in decentralized and sensitive data, as found in works by Mukherjee & Buyya (2024) [9] and Li et al. (2023) [44]
though with the challenge of high communication overhead. Although federated learning has some benefits, it has several challenges
for real-world applications. Model inversion and gradient leakage represent potential sources of security risk that could compromise
sensitive information [17]. Convergence may be unstable for non-IID distributions of data found in agricultural datasets as well.
Edge devices may not have enough computational or energy resources for full model training, requiring adaptive participation or
light-weight models.

Research gaps

The existing studies do not have a clear framework that can integrate multi-source data across different regions and crops. Most of
the deep learning-based models, like CNN and RNN, act as black boxes and hence require development in the field of explainable AI
toward agricultural predictions. High-resolution data, along with multi-modal frameworks, adds to the complexity and hence limits
scalability because of higher computational demands. While most related works in federated learning report significant variability
in client data and its implications for model performance, the solution remains largely under-explored. While federated learning

15
V. Hiremani, R.M. Devadas, Preethi et al. MethodsX 14 (2025) 103408

provides data privacy, it is still a challenge to get the best performance without compromising on privacy. Models are limited in general
application by potential biases in the way data is collected-for example, reliance on weather stations or soil samples. Despite impressive
progress made, a few critical challenges in using Federated Learning in agricultural contexts, specifically crop yield prediction, persist.
One significant problem lies in data heterogeneity—the nonIID nature associated with agricultural data by area and crop type usually
distorts model convergence and generalizability. A corresponding bottleneck also exists due to communication overhead, especially
in far-away farming regions where internet access is poor, impacting FL deployment scalability.
An additional key limitation includes a lack of interpretability for FL models. Although robust, most deep learning-based FL models
are black boxes, preventing trust and uptake among agricultural users. Moreover, data labeling quality varies by farm, potentially
compromising local training reliability. The computational load placed upon edge devices also inhibits small-scale farmer involvement
where infrastructure may not be up-to-date.

Future directions

Researchers should prioritize developing models that are generalizable across diverse agricultural contexts and crop types. Lever-
aging techniques like transfer learning and domain adaptation could enable predictive frameworks to adapt effectively to varying
conditions. Additionally, scalability remains a critical area for improvement. Federated learning frameworks should focus on opti-
mizing communication protocols and employing model pruning to reduce communication overhead and latency, making them more
practical for widespread deployment. Emerging technologies offer significant potential for advancing this field. Attention mechanisms,
generative models, and graph-based learning could be integrated to enhance feature representation and decision-making capabilities.
Climate variability is another pressing challenge, necessitating the development of robust predictive systems that can incorporate
extreme weather conditions and long-term climatic shifts through multi-source data fusion and temporal modeling. Improving the
quality of data input is equal in importance. The issues of bias in data collection and balancing out with high-quality datasets would
result in the reliability and utility of the prediction models. They should also investigate light, energy-efficient models designed
specifically for resource-constrained environments so that they are pragmatically usable without reducing their accuracy. Few more
areas where future studies can be performed by creating hybrid FL-XAI frameworks that blend explainability with privacy.
Creating lightweight FL models adapted for deployment on low-resource edge devices.
Developing domain-specific FL benchmarks for agricultural use to provide standardized evaluations.
Examining asynchronous and adaptive aggregation techniques for handling data and resource variability.
Investigating privacy-preserving data augmentation methods to enhance learning from sparse data or noisy data. Finally, there
is a need to bridge the gap between academic advancements and practical implementations. Researchers should collaborate with
stakeholders in agriculture, such as farmers, policymakers, and technology providers, to design systems that meet real-world needs
while incorporating cutting-edge methodologies.

Conclusion

Federated learning in crop yield prediction is one of the hopeful advancements in agricultural technology toward overcoming the
problems posed in terms of data privacy and heterogeneity, requiring a decentralized approach to data processing. The present review
analyzes recent studies on models of various types, including Convolutional Neural Networks (CNNs), Long Short-Term Memory
(LSTM) networks, and ensemble methods, which are implemented to predict crop yields in diverse datasets, ranging from weather
data, soil conditions, and remote sensing. Notably, improvements made by federated learning frameworks surpass 97% prediction
accuracies reported by some studies. Despite these advancements, the field remains in its infancy, with many studies highlighting
limitations such as sensitivity to external factors, challenges in generalizability across different crops and regions, and complexities in
data integration. Future research should be directed toward enhancing model robustness against environmental variability, improving
data integration techniques, and exploring the scalability of federated learning frameworks. In addition, research gaps involving the
representation and interpretability of complex models and crop diversity should be targeted to increase broader acceptance. In
conclusion, though federated learning represents a transformative approach to crop yield prediction, further efforts are required in
refining these models and applying them in various agricultural settings. Further collaboration between researchers and practitioners
will be important in tapping into these technologies fully to meet global food production needs under the changing climatic conditions.

Ethics statements

No data was collected from any social media platforms.

Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to
influence the work reported in this paper.

16
V. Hiremani, R.M. Devadas, Preethi et al. MethodsX 14 (2025) 103408

CRediT authorship contribution statement

Vani Hiremani: Conceptualization. Raghavendra M. Devadas: Methodology. Preethi: Formal analysis. R. Sapna: Data curation.
T. Sowmya: Visualization. Praveen Gujjar: Writing – review & editing. N. Shobha Rani: Writing – review & editing. K.R. Bhavya:
Supervision.

Acknowledgments

This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.

References

[1] O. Elijah, et al., An overview of Internet of Things (IoT) and data analytics in agriculture: benefits and challenges, IEEE Internet Things J. 5 (5) (2018) 3758–3773,
doi:10.1109/jiot.2018.2844296.
[2] M. Jahanbakht, et al., Internet of underwater things and big marine data analytics–a comprehensive survey, IEEE Commun. Surv. Tutor. 23 (2) (2021) 904–956,
doi:10.1109/comst.2021.3053118.
[3] T. Manoj, et al., A federated learning-based crop yield prediction for agricultural production risk management, in: Proceedings of the IEEE Delhi Section
Conference (DELCON), 2022.
[4] G.M. Pakadang, Y.T. Muryanto, Application of federated learning for smart agriculture system, J. Leg. Subj. (43) (2024) 36–47.
[5] A. Yousefpour, et al., All one needs to know about fog computing and related edge computing paradigms: a complete survey, J. Syst. Archit. 98 (2019) 289–330,
doi:10.1016/[Link].2019.02.009.
[6] O. Elijah, T.A. Rahman, I. Orikumhi, C.Y. Leow, M.N. Hindia, An overview of Internet of Things (IoT) and data analytics in agriculture: benefits and challenges,
IEEE Internet Things J. 5 (5) (2018) 3758–3773, doi:10.1109/JIOT.2018.2844296.
[7] M. Weiss, et al., Remote sensing for agricultural applications: a meta-review, Remote Sens. Environ. 236 (2019) 111402, doi:10.1016/[Link].2019.111402.
[8] Kairouz, P., et al. Advances and open problems in federated learning. 2021.
[9] Mukherjee, A., and R. Buyya. "Federated learning architectures: a performance evaluation with crop yield prediction application." ArXiv, 2024,
[Link] Accessed 19 Dec. 2024.
[10] D.C. Nguyen, et al., Federated learning for internet of things: a comprehensive survey, IEEE Commun. Surv. Tutor. 23 (3) (2021) 1622–1658,
doi:10.1109/comst.2021.3075439.
[11] R. Fu, et al., FEAST: a communication-efficient federated feature selection framework for relational Data, in: Proceedings of the ACM on Management of Data,
1, 2023, pp. 1–28, doi:10.1145/3588961.
[12] J. Wen, et al., A survey on federated learning: challenges and applications, Int. J. Mach. Learn. Cybern. 14 (2) (2022) 513–535, doi:10.1007/s13042-022-01647-y.
[13] V. Ramesh, P. Kumaresan, Advancements in machine learning and deep learning techniques for crop yield prediction: a comprehensive review, Nat. Environ.
Pollut. Technol. 23 (4) (2024) 2071–2086.
[14] I. Dubey, D. Motwani, Fusion of deep learning techniques with time series analysis for crop yield prediction from satellite remote sensing data, ShodhKosh: J.
Vis. Perform. Arts 5 (1) (2024).
[15] B.S. Devi, Optimization of crop yield prediction through linear modelling and deep learning techniques used in precision agriculture, Commun. Appl. Nonlinear
Anal. 32 (3) (2024) 577–591.
[16] B. Yurdem, M. Kuzlu, M.K. Gullu, F.O. Catak, M. Tabassum, Federated learning: overview, strategies, applications, tools and future directions, Heliyon 10 (19)
(2024) e38137, doi:10.1016/[Link].2024.e38137.
[17] J. Curl, X. Xie, Societal impacts and opportunities of federated learning, Chin. J. Sociol. 11 (1) (2025) 90–100, doi:10.1177/2057150x251314299.
[18] M.R. Berkani, A. Chouchane, Y. Himeur, A. Ouamane, S. Miniaoui, S. Atalla, W. Mansoor, H. Al-Ahmad, Advances in federated learning: applications and
challenges in smart building environments and beyond, Computers 14 (4) (2025) 124, doi:10.3390/computers14040124.
[19] A. Karunamurthy, K. Vijayan, P.R. Kshirsagar, K.T. Tan, An optimal federated learning-based intrusion detection for IoT environment, Sci. Rep. 15 (1) (2025),
doi:10.1038/s41598-025-93501-8.
[20] K. Daly, H. Eichner, P. Kairouz, H.B. McMahan, D. Ramage, Z. Xu, Federated learning in practice: reflections and projections, in: Proceedings of the IEEE 6th Inter-
national Conference on Trust, Privacy and Security in Intelligent Systems, and Applications (TPS-ISA), 2024, pp. 148–156, doi:10.1109/tps-isa62245.2024.00026.
[21] P. Kairouz, H.B. McMahan, B. Avent, A. Bellet, M. Bennis, A.N. Bhagoji, K. Bonawitz, Z. Charles, G. Cormode, R. Cummings, R.G.L. D’Oliveira, H. Eichner,
S.E. Rouayheb, D. Evans, J. Gardner, Z. Garrett, A. Gascón, B. Ghazi, P.B. Gibbons, M. Gruteser, Z. Harchaoui, C. He, L. He, Z. Huo, B. Hutchinson, J. Hsu,
M. Jaggi, T. Javidi, G. Joshi, M. Khodak, J. Konecný, A. Korolova, F. Koushanfar, S. Koyejo, T. Lepoint, Y. Liu, P. Mittal, M. Mohri, R. Nock, A. Özgür, R. Pagh,
Q. Yang, D. Ramage, R. Raskar, M. Raykova, D. Song, W. Song, S. U. Stich, Z. Sun, A. T. Suresh, F. Tramèr, P. Vepakomma, J. Wang, L. Xiong, Z. Xu, F. X. Yu,
H. Yu, S. Zhao, Advances and open problems in federated learning, 2021, doi:10.1561/9781680837896.
[22] M. Moshawrab, M. Adda, A. Bouzouane, H. Ibrahim, A. Raad, Reviewing federated learning aggregation algorithms; strategies, contributions, limitations and
future perspectives, Electronics 12 (10) (2023) 2287, doi:10.3390/electronics12102287.
[23] A.B. Arrieta, et al., Explainable artificial intelligence (XAI): concepts, taxonomies, opportunities and challenges toward responsible AI, Inf. Fusion 58 (2019)
82–115, doi:10.1016/[Link]ffus.2019.12.012.
[24] W.Y.B. Lim, et al., Federated learning in mobile edge networks: a comprehensive survey, IEEE Commun. Surv. Tutor. 22 (3) (2020) 2031–2063,
doi:10.1109/comst.2020.2986024.
[25] L. Bottou, Large-scale machine learning with stochastic gradient descent, in: Proceedings of the COMPSTAT’2010, 2010, pp. 177–186.
[26] Z. Pan, S. Wang, C. Li, H. Wang, X. Tang, J. Zhao, FedMDFG: federated learning with multi-gradient descent and fair guidance, in: Proceedings of the AAAI
Conference on Artificial Intelligence, 37, 2023, pp. 9364–9371, doi:10.1609/aaai.v37i8.26122.
[27] H.B. McMahan, et al., Communication-efficient learning of deep networks from decentralized data, in: Proceedings of the International Conference on Artificial
Intelligence and Statistics, 2016.
[28] L. Chen, F. Ang, Y. Chen, W. Wang, Robust federated learning with noisy labeled data through loss function correction, IEEE Trans. Netw. Sci. Eng. 10 (3) (2023)
1501–1511, doi:10.1109/TNSE.2022.3227287.
[29] P. Zheng, et al., Federated learning in heterogeneous networks with unreliable communication, IEEE Trans. Wirel. Commun. 23 (2024) 3823–3838.
[30] J. Friedman, et al., Additive logistic regression: a statistical view of boosting (With discussion and a rejoinder by the authors), Ann. Stat. 28 (2) (2000),
doi:10.1214/aos/1016218223.
[31] M. Penmetsa, R.N.V. Jagan Mohan, Multi-crop analysis using multi-regression via AI-based federated learning, in: Algorithms in Advanced Artificial Intelligence,
CRC Press, 2024, pp. 487–491.
[32] D. Gardas, R. Karthi, Crop irrigation advisory system using federated logistic regression, in: Proceedings of the International Conference on Computational
Intelligence in Data Science, Cham, Springer Nature Switzerland, 2024.
[33] S. Senthil, K. Raja, Comparison of machine learning algorithms for the prediction of rice crop yield in Karnataka, in: Proceedings of the Fourth International
Conference on Smart Technologies in Computing, Electrical and Electronics (ICSTCEE), 2023, pp. 1–8.
[34] Q. Yang, et al., Federated Learning: Privacy and Incentive, Springer Nature, 2020.
[35] C. Shorten, T.M. Khoshgoftaar, A survey on image data augmentation for deep learning, J. Big Data 6 (1) (2019), doi:10.1186/s40537-019-0197-0.

17
V. Hiremani, R.M. Devadas, Preethi et al. MethodsX 14 (2025) 103408

[36] V. Kiran Kumar, et al., Optimizing LSTM and Bi-LSTM models for crop yield prediction and comparison of their performance with traditional machine learning
techniques, Appl. Intell. 53 (23) (2023) 28291–28309.
[37] S. Khaki, et al., A CNN-RNN framework for crop yield prediction, Front. Plant Sci. 10 (2020).
[38] A. Toledo, Claudia, et al., Integrating multi-modal remote sensing, deep learning, and attention mechanisms for yield prediction in plant breeding experiments,
Front. Plant Sci. 15 (2024).
[39] D. Elavarasan, P.M. Vincent, Crop yield prediction using deep reinforcement learning model for sustainable agrarian applications, IEEE Access 8 (2020)
86886–86901.
[40] H.T. Pham, et al., Enhancing crop yield prediction utilizing machine learning on satellite-based vegetation health indices, Sensors 22 (3) (2022) 719.
[41] D. Peng, et al., A deep–learning network for wheat yield prediction combining weather forecasts and remote sensing data, Remote Sens. 16 (19) (2024) 3613.
[42] D. Tripathi, S.K. Biswas, Design of a precise ensemble expert system for crop yield prediction using machine learning analytics, J. Forecast. 43 (8) (2024)
3161–3176.
[43] O. Khin, J.G. Koh, S.K. Lee, Harvest forecasting improvement using federated learning and ensemble model, Korean Inst. Smart Media 12 (10) (2023) 9–18.
[44] A. Li, et al., Model pruning enables localized and efficient federated learning for yield forecasting and data sharing, Expert Syst. Appl. 242 (2024) 122847.
[45] Lopez-Ramos, L.M., et al. Interplay between federated learning and explainable artificial intelligence: a scoping review. 2024, 10.48550/arxiv.2411.05874.
[46] Rahmati, M. Federated learning and explainable AI for personalized healthcare in resource-limited settings. 2025, 10.21203/[Link]-5691431/v1.
[47] E. Briola, C.C. Nikolaidis, V. Perifanis, N. Pavlidis, P. Efraimidis, A federated explainable AI model for breast cancer classification, in: Proceedings of the European
Interdisciplinary Cybersecurity Conference, 2024, pp. 194–201, doi:10.1145/3655693.3660255.
[48] L.A. Enyejo, M.B. Adewoye, U.N. Ugochukwu, Interpreting federated learning (FL) models on edge devices by enhancing model Explainability with computational
geometry and advanced database architectures, Int. J. Sci. Res. Comput. Sci. Eng. Inf. Technol. 10 (6) (2024) 332–354, doi:10.32628/cseit24106185.
[49] M. Daole, P. Ducange, F. Marcelloni, A. Renda, Trustworthy AI in heterogeneous settings: federated learning of explainable classifiers, in: Proceedings of the
IEEE International Conference on Fuzzy Systems (FUZZ-IEEE), 2024, pp. 1–9, doi:10.1109/fuzz-ieee60900.2024.10612109.
[50] U. Nirosha, G. Vennila, Enhancing crop yield prediction for agriculture productivity using federated learning integrating with graph and recurrent neural networks
model, Expert Syst. Appl. (2025) 128312 ISSN 0957-4174, doi:10.1016/[Link].2025.128312.
[51] K.R. Žalik, M. Žalik, A review of federated learning in agriculture, Sensors 23 (23) (2023) 9566, doi:10.3390/s23239566.
[52] A. Chauhan, A. Jot, I. Kaur, R. Mohana, A review on federated learning in healthcare, agriculture, and education, and its future prospects, in: Proceed-
ings of the International Conference on Artificial Intelligence and Emerging Technology (Global AI Summit), 2024, pp. 501–506, doi:10.1109/globalaisum-
mit62156.2024.10947884.

18

You might also like