Statistical Analysis of Yield in Major Crops Under Varying
Environmental and Agronomical Conditions Using RStudio
Ayaz Haider Khan
2022-ag-6654
PBG-514
Department of Plant Breeding and Genetics, University of
Agriculture Faisalabad
1
Table of contents
1. Introduction ........................................................................................................................................... 4
Objectives: ................................................................................................................................................ 4
Methods: ................................................................................................................................................... 4
2. Summary of variables ........................................................................................................................... 5
3. Analysis of variance ............................................................................................................................ 10
Assumptions:........................................................................................................................................... 10
Fully fixed model ........................................................................................................................................ 10
Objective: ................................................................................................................................................ 10
Findings: ................................................................................................................................................. 11
Implications: ........................................................................................................................................... 11
Fully random model .................................................................................................................................... 11
Objective: ................................................................................................................................................ 11
Findings: ................................................................................................................................................. 11
Implications: ........................................................................................................................................... 12
Mixed model ............................................................................................................................................... 12
Objective: ................................................................................................................................................ 12
Findings: ................................................................................................................................................. 12
Interpretation:.......................................................................................................................................... 12
4. Association between Categorical Variables ........................................................................................ 13
Crop vs Soil Type ....................................................................................................................................... 13
Result: ..................................................................................................................................................... 13
Interpretation:.......................................................................................................................................... 13
Implications: ........................................................................................................................................... 13
Fertilizer Use vs Irrigation Use ................................................................................................................... 14
Result: ..................................................................................................................................................... 14
Interpretation:.......................................................................................................................................... 14
Implications: ........................................................................................................................................... 14
Region vs Weather Condition ..................................................................................................................... 14
Result: ..................................................................................................................................................... 14
Interpretation:.......................................................................................................................................... 15
Implications: ........................................................................................................................................... 15
Crop Distribution (Goodness-of-Fit Test)................................................................................................... 15
2
Result: ..................................................................................................................................................... 15
Interpretation:.......................................................................................................................................... 15
Implications: ........................................................................................................................................... 15
5. Calculation of Heritability and Extraction of Variance Components ................................................. 16
Objective ................................................................................................................................................. 16
Results ..................................................................................................................................................... 16
Variance Proportions ("Heritability") ..................................................................................................... 16
Interpretation ........................................................................................................................................... 16
6. Conclusion .......................................................................................................................................... 17
3
1. Introduction
The data has been acquired from Kaggle under the name “Agriculture Crop Yield Dataset”
which shows the relative yield of crops under different variables of environment, soil, crop etc.
The data considered has a total of 400 observations including both qualitative and quantitative
variables and values.
Objectives:
1. The objective of the analysis is to find the most effecting factors contributing to the total
yield of the crops.
2. Test association between different variables for relative impacts.
3. Quantify variance of Region and Soil Type attributable to yield.
Methods:
The dataset was uploaded and analyzed in RStudio using the R Commander interface. Summary
statistics for each variable were generated using R Commander and exported (Table 1 and 2).
Various statistical analyses were performed using base R functions to address the hypotheses and
research questions.
1. Descriptive Statistics and Data Visualization
Summary statistics were calculated for all variables in the dataset, including measures such as
mean, median, quartile 1, quartile 3, maximum and minimum values. Visualizations, including
histograms and boxplots, were used to assess the distribution and identify potential outliers. The
summary statistics and corresponding plots are presented in Fig 1-10.
2. ANOVA
Analysis of Variance conducted using base functions available in RStudio under library “lme4”
to find to most effecting factors for Yield (tons/hectare). After the analysis, results and
implications were discussed.
3. X2 Analysis
Association between categorical variables was measured using base functions available in
RStudio under library “rcompanion”. After the analysis, results and implications were discussed.
4
4. Heritability and Variance Analysis
Heritability and Variance analysis conducted using base functions available in RStudio under
library “lme4” to find to most effecting factors having variances attributable to Yield
(tons/hectare). After the analysis, results and implications were discussed.
2. Summary of variables
• Regions: The dataset includes four regions: East, North, South, and West, with respective
counts of 88, 125, 85, and 102. The North region has the highest count at 125, while the
South region has the lowest count at 85. The distribution of counts suggests a relatively
balanced but slightly higher concentration of data from the North region.
Fig. 1: Bar plot showing the frequency of crop yield entries across different Regions.
• Soil: Soil types in the dataset include Chalky, Clay, Loam, Peaty, Sandy, and Silt, with
respective counts of 63, 65, 71, 63, 61, and 77. The most frequent soil type is Silt with 77
counts, while Sandy has the fewest, with 61 counts. The distribution shows some
variation in soil type representation, with a noticeable concentration of Loam and Silt.
Fig. 2: Bar plot showing the frequency of crop yield entries across different Soil Types.
5
• Crops: The dataset contains data on six types of crops: Barley, Cotton, Maize, Rice,
Soybean, and Wheat, with respective counts of 62, 69, 62, 74, 65, and 68. Rice is the
most common crop, with 74 counts, while Barley and Maize are tied for the least
frequent, both with 62 counts. The distribution appears relatively balanced with some
slight variation in crop representation.
Fig. 3: Bar plot showing the frequency of crop yield entries across different Crops.
• Rainfall (mm): Rainfall in the dataset ranges from 104.20 mm to 997.40 mm, with a
mean value of 547.90 mm and a median of 545.40 mm. The closeness of the mean and
median suggests a fairly symmetric distribution, though some outliers cause a bit
skewness. The first quartile (Q1) is 334.00 mm, and the third quartile (Q3) is 783.10 mm.
Fig. 4: Histogram showing the distribution of Rainfall (mm) across crop yield records.
• Temperature (°C): Temperature values range from 15.03°C to 39.99°C, with a mean of
27.45°C and a median of 27.79°C. The mean being slightly lower than the median
6
suggests a minor left skew in the distribution. The first quartile (Q1) is 21.33°C and the
third quartile (Q3) is 33.49°C. Outliers may be present at the lower or upper ends.
Fig. 5: Histogram showing the distribution of Temperature (°C) across crop yield records.
• Fertilizer: The fertilizer variable has two categories: TRUE and FALSE, with counts of
197 for TRUE and 203 for FALSE. The dataset is slightly more weighted toward the
FALSE category, though the distribution is nearly balanced.
Fig. 6: Bar plot showing the frequency of crop yield entries across Fertilizer usage.
• Irrigation: The irrigation variable has two categories: TRUE and FALSE, with counts of
210 for TRUE and 190 for FALSE. TRUE (irrigation) has a slightly higher count,
suggesting a greater number of instances where irrigation is used.
7
Fig. 7: Bar plot showing the frequency of crop yield entries across Irrigation usage.
• Weather: The weather variable includes three categories: Cloudy, Rainy, and Sunny, with
respective counts of 139, 112, and 149. The most common weather condition is Sunny,
with 149 counts, followed by Cloudy with 139 counts. Rainy days have the lowest count
at 112. The distribution is somewhat uniform, with Sunny being slightly more prevalent.
Fig. 8: Bar plot showing the frequency of crop yield entries across different Weather
Conditions.
• Days To Harvest (days): The number of days until harvest varies from 60.00 days to
149.00 days, with a mean of 105.10 days and a median of 105.00 days, indicating a
nearly symmetrical distribution. The first quartile (Q1) is 80.00 days and the third
quartile (Q3) is 131.00 days.
8
Fig. 9: Histogram showing the distribution of Days To Harvest across crop yield records.
• Yield: Yield ranges from 0.62 to 8.88, with a mean of 4.69 and a median of 4.67,
suggesting a roughly symmetric distribution with minimal skewness. The first quartile
(Q1) is 3.52 and the third quartile (Q3) is 5.98.
Fig. 10: Histogram showing the Yield (tons/hectare) in data records.
Table 1: Summary statistics of quantitative variables
Variable Minimum 1st Quartile Median Mean 3rd Quartile Maximum
Rainfall 104.20 334.00 545.40 547.90 783.10 997.40
Temperature 15.03 21.33 27.79 27.45 33.49 39.99
Days_Harvest 60.00 80.00 105.00 105.10 131.00 149.00
Yield 0.62 3.52 4.67 4.69 5.98 8.88
Table 2: Summary statistics of qualitative variables
Irrigation FALSE TRUE - - - -
9
Count 190 210 - - - -
Fertilizer FALSE TRUE - - - -
Count 203 197 - - - -
Weather Cloudy Rainy Sunny - - -
Count 139 112 149 - - -
Region East North South West - -
Count 88 125 85 102 - -
Soil Chalky Clay Loam Peaty Sandy Silt
Count 63 65 71 63 61 77
Crop Barley Cotton Maize Rice Soybean Wheat
Count 62 69 62 74 65 68
3. Analysis of variance
Assumptions:
1- Independence: Observations should be independent of each other.
2- Homoscedasticity: The variance of residuals should be constant across groups and levels
of predictors.
3- Normality: The residuals (not the raw variables) should follow a normal distribution.
Independence is likely satisfied because the data appears to be from distinct observations, and
random effects account for any potential grouping. Homoscedasticity is likely satisfied because
the yield distribution is symmetric and predictor group sizes are reasonably balanced. Normality
is likely satisfied because the yield variable is approximately normally distributed (mean ≈
median) with minimal skew.
Fully fixed model
Objective:
Understand the influence of Region, Soil, Crop, Fertilizer, Irrigation and Weather when
considered completely fixed against a response variable of Yield. The formula used will be as
below:
𝑌𝑖𝑒𝑙𝑑 ~ 𝑟𝑒𝑔𝑖𝑜𝑛 + 𝑠𝑜𝑖𝑙 + 𝑐𝑟𝑜𝑝 + 𝑓𝑒𝑟𝑡𝑖𝑙𝑖𝑧𝑒𝑟 + 𝑖𝑟𝑟𝑖𝑔𝑎𝑡𝑖𝑜𝑛 + 𝑤𝑒𝑎𝑡ℎ𝑒𝑟 + 𝑅
10
Findings:
Analysis revealed that Region has a highly significant effect on Yield (tons/hectare), with a p-
value less than 0.01. Fertilizer and Irrigation were found to have even stronger effects, both
showing highly significant impacts with p-values less than 0.001. In contrast, the remaining
factors—Soil, Crop, and Weather—did not contribute significantly to the variability in Yield
(tons/hectare), as their p-values exceeded 0.05, indicating negligible or no statistically significant
influence.
Implications:
The findings suggest that changes in Region, Fertilizer, and Irrigation will significantly influence
the final crop Yield (tons/hectare). The p-values indicate strong evidence that these factors are
driving variations in Yield (tons/hectare). This means that agricultural practices and
environmental conditions related to these factors should be closely considered to optimize Yield
(tons/hectare). On the other hand, the factors of Soil, Crop, and Weather do not appear to
contribute meaningfully to the variability in Yield (tons/hectare) in this model.
Fully random model
Objective:
Understand the influence of Region, Soil, Crop, Fertilizer, Irrigation and Weather when
considered completely random against a response variable of Yield. The formula used will be as
below:
𝑌𝑖𝑒𝑙𝑑 ~ (1|𝑅𝑒𝑔𝑖𝑜𝑛) + (1|𝑆𝑜𝑖𝑙) + (1|𝐶𝑟𝑜𝑝) + (1|𝐹𝑒𝑟𝑡𝑖𝑙𝑖𝑧𝑒𝑟) + (1|𝐼𝑟𝑟𝑖𝑔𝑎𝑡𝑖𝑜𝑛)
+ (1|𝑊𝑒𝑎𝑡ℎ𝑒𝑟) + 𝑅
Findings:
Crop and Region had zero variance, indicating no random effect on yield (tons/hectare). In
contrast, Irrigation and Fertilizer showed substantial random effects with variances of 1.01 and
0.92, respectively, while Weather had a minor random effect (variance = 0.0067). The residual
variance was 1.89, suggesting notable unexplained variability. For fixed effects, the intercept
(baseline yield) was estimated at 4.66 tons/hectare with a standard error of 0.98 and a significant
t-value of 4.71 (p < 0.001), indicating the yield is significantly above zero.
11
Implications:
The analysis shows that Crop and Region have no significant effect on Yield (tons/hectare),
while Weather, Irrigation, and Fertilizer contribute random effects, indicating they explain some
yield variability. The fixed intercept provides a baseline yield estimate. However, a singularity fit
warning suggests possible multicollinearity or overparameterization, meaning some parameters
may be redundant or not estimable. This raises concerns about model stability and limits
confidence in the interpretation and generalizability of the random effects.
Mixed model
Objective:
Understand the influence of Region, Soil, Crop, Fertilizer, Irrigation and Weather when Region
and Soil are considered as random and others as fixed against a response variable of Yield. The
formula used will be as below:
𝑌𝑖𝑒𝑙𝑑 ~ (1|𝑅𝑒𝑔𝑖𝑜𝑛) + (1|𝑆𝑜𝑖𝑙) + 𝐶𝑟𝑜𝑝 + 𝐹𝑒𝑟𝑡𝑖𝑙𝑖𝑧𝑒𝑟 + 𝐼𝑟𝑟𝑖𝑔𝑎𝑡𝑖𝑜𝑛 + 𝑊𝑒𝑎𝑡ℎ𝑒𝑟 + 𝑅
Findings:
The model results indicate that Soil Type contributes some variability to Yield (tons/hectare)
(variance = 0.02391), while Region shows no significant impact (variance = 0.00000). Most of
the unexplained variability in Yield (tons/hectare) is captured by the residual variance (1.88638).
Among the fixed effects, the Intercept is highly significant at 3.31 tons per hectare, representing
the baseline Yield (tons/hectare). Fertilizer and Irrigation both have strong, statistically
significant positive effects on Yield, with estimates of 1.37 and 1.44 respectively. In contrast,
none of the crops (Cotton, Maize, Rice, Soybean, Wheat) differ significantly in Yield
(tons/hectare) compared to the reference crop. Additionally, sunny weather appears to negatively
affect Yield (tons/hectare), as indicated by a negative t-value (-1.744).
Interpretation:
None of the crops play a significant role in the variability of Yield (tons/hectare), likely
contributing to a singularity fit error due to a lack of meaningful differences among crop types.
Fertilizer and Irrigation both have highly significant and positive effects, increasing Yield by
12
approximately 1.37 and 1.44 tons per hectare, respectively. In contrast, sunny climate slightly
reduces Yield (tons/hectare), as suggested by its negative effect. Regarding random effects, Soil
Type contributes a small amount of variability to Yield, while Region shows no significant
influence.
4. Association between Categorical Variables
Chi-square tests the fitness of data and is only used for qualitative variables. Within my dataset
qualitative variables are Region, Soil, Crop, Fertilizer, Irrigation, Weather. So, for these ones chi-
square can be applied to while others have been accounted for in ANOVA.
Crop vs Soil Type
Result:
Statistic Value
X2 19.26
Degrees of Freedom (df) 25
p-value 0.78
Interpretation:
The results of the chi-square test suggest that there is no statistically significant association
between crop type and soil type in the data analysed. With a chi-square value of 19.265 and 25
degrees of freedom, the calculated p-value is 0.7842, which is much higher than the conventional
threshold of 0.05. This high p-value indicates that any observed differences between crop types
and soil types are likely due to random chance rather than a true underlying relationship.
Implications:
Given the lack of statistical significance, there is no strong evidence to suggest that the type of
soil influences the type of crop that can be grown, based on the current dataset. The result
implies that other factors, outside of soil type, may be more important when considering crop
suitability.
13
Fertilizer Use vs Irrigation Use
Result:
Statistic Value
X2 0.15
p-value 0.70
Cramér’s V 0.02
Interpretation:
The analysis suggests that there is no significant association between the use of fertilizer and the
use of irrigation, as evidenced by the high p-value. A high p-value indicates that the observed
relationship could likely be due to random chance rather than a meaningful statistical
relationship. Additionally, the Cramér's V value of 0.019 further supports this conclusion by
revealing an extremely weak effect size, meaning any potential relationship between these two
variables is practically negligible.
Implications:
The lack of a significant relationship between fertilizer use and irrigation practices implies that
the decision to use one does not necessarily depend on or influence the use of the other. In
practical terms, this suggests that agricultural strategies focused on fertilizer and irrigation may
be independent decisions for farmers or land managers, and interventions aimed at improving
one practice may not automatically impact the other. This can inform resource allocation and
policy decisions, as efforts to promote more efficient irrigation or fertilizer usage might need to
be addressed separately, without assuming an interdependent relationship between the two
practices.
Region vs Weather Condition
Result:
Statistic Value
X2 11.19
Degrees of Freedom (df) 6
p-value 0.08
14
Interpretation:
The analysis indicates that the p-value is marginally above the 0.05 threshold, suggesting that the
result is not statistically significant. However, the p-value is closer to 0.05 compared to other
tests, implying the possibility of a weak trend between the variables. Despite this, the evidence is
not strong enough to confidently conclude a meaningful or substantial association between the
variables. Additionally, since no expected counts fall below 5, there is no need to conduct
Fisher’s test.
Implications:
The marginally higher p-value suggests that while there may be a weak trend between the
variables, any potential relationship is not statistically significant enough to inform major
decisions or interventions. This implies that, although further investigation might be worth
considering, the current data does not provide robust support for a strong link between the
variables. In practice, this could mean that any observed trend may be more coincidental than
causal, and resources or actions based on this potential relationship should be carefully weighed.
Crop Distribution (Goodness-of-Fit Test)
Result:
Statistic Value
X2 1.61
Degrees of Freedom (df) 5
p-value 0.90
Interpretation:
The analysis reveals a very high p-value, which leads us to fail to reject the null hypothesis of
uniform distribution. This indicates that there is no significant deviation from a uniform
distribution among the crop types in the dataset. In other words, the distribution of crop types is
roughly even across the observed data.
Implications:
Since the crop types are evenly distributed, this suggests that there is no clear preference or
imbalance in the types of crops represented in the dataset. In practical terms, this could imply
15
that factors influencing crop selection are fairly balanced, and there may be no strong trend or
bias towards any specific crop type. This could inform agricultural strategies or research, as it
suggests that crop distribution in the dataset does not need to be adjusted for imbalances, and any
interventions or policies may need to account for the diversity of crop types equally.
5. Calculation of Heritability and Extraction of Variance
Components
Objective
The objective is to quantify the variance in crop yield (Yield_tons_per_hectare) attributable to
three factors: soil type (random effect), region (random effect), and residual variance
(unexplained variation). Additionally, we aim to estimate the "heritability," which is the
proportion of total variance explained by each factor, providing insight into the relative
importance of soil type, region, and other unaccounted influences on crop yield.
Results
The VarCorr() output shows:
Group Variance (vcov) Std. Dev. (sdcor)
Soil_Type 0.0239 0.1546
Region 0.0000 0.0000
Residual 1.8864 1.3735
Variance Proportions ("Heritability")
The analysis shows that soil type accounts for 1.25% of the variance in crop yield, while region
contributes 0% to the variance. The remaining 98.75% of the variance is attributed to residual
(unexplained) factors. This suggests that soil type has a small influence on crop yield, while
region has no measurable effect, and most of the variation in yield remains unexplained by these
factors.
Interpretation
Soil type explains a small but non-zero portion (1.25%) of the variance in crop yield, with a
variance component of 0.0239. This suggests that while differences in soil type have a weak
16
influence on yield, other factors are likely to have a greater impact. Therefore, while improving
soil properties could have some benefit, it may not be the most effective focus if resources are
limited, as other yield drivers might provide more significant improvements.
Region shows no detectable effect on crop yield in this dataset, as it explains 0% of the variance
(variance component = 0). This implies that crop yields do not vary systematically across
regions. The regions in this study may be too similar in terms of climate and management
practices, or regional effects might be captured by other variables such as weather patterns or
irrigation practices, which could better explain yield differences.
The residual variance accounts for 98.75% of the variation in crop yield, indicating that most of
the yield variability remains unexplained by the model. This unexplained variation could be due
to missing key predictors, such as pest pressure, farmer expertise, or microclimates, as well as
potential measurement errors in yield data or random environmental fluctuations. Identifying and
incorporating these factors could improve the model's ability to explain yield variation.
6. Conclusion
In conclusion, the analysis of agricultural crop yield data reveals the following key findings:
• Fixed Effects: Region, Fertilizer, and Irrigation significantly influence crop yield,
meaning changes in these factors will directly affect yield.
• Random Effects: Weather, Irrigation, and Fertilizer show random effects on yield, with
Irrigation and Fertilizer exhibiting higher variances.
• Mixed Effects: Soil Type contributes to yield variability, while Fertilizer and Irrigation
have a strong positive effect. Weather shows a slight negative effect on yield.
From the chi-square analysis, no significant relationships were found between crop type and soil
type, or between fertilizer and irrigation use. Additionally, no strong association was found
between region and weather conditions, although a weak trend was observed. Crop distribution
appears to be uniform across the dataset.
17
Variance analysis indicates that soil type accounts for just 1.25% of the yield variance, while the
majority (98.75%) remains unexplained. This suggests that other factors, such as pest pressure or
microclimates, need to be explored to gain a fuller understanding of yield variability.
18