Mapping crops within the growing season
Project-I (AI67001) report submitted to
Indian Institute of Technology Kharagpur
in partial fulfilment for the award of the degree of
Dual Degree
in
Centre of Excellence in Artificial Intelligence
by
S Rohit Chandra Jagadeesh
(18MT3AI32)
Under the supervision of
Prof Adway Mitra
Centre of Excellence in Artificial Intelligence
Indian Institute of Technology Kharagpur
Autumn Semester, 2022-23
November 28, 2022
DECLARATION
I certify that
(a) The work contained in this report has been done by me under the guidance of my
supervisor.
(b) The work has not been submitted to any other Institute for any degree or diploma.
(c) I have conformed to the norms and guidelines given in the Ethical Code of Conduct
of the Institute.
(d) Whenever I have used materials (data, theoretical analysis, figures, and text) from
other sources, I have given due credit to them by citing them in the text of the
thesis and giving their details in the references. Further, I have taken permission
from the copyright owners of the sources, whenever necessary.
Date: November 28, 2022 (S Rohit Chandra Jagadeesh)
Place: Kharagpur (18MT3AI32)
i
CENTRE OF EXCELLENCE IN ARTIFICIAL
INTELLIGENCE
INDIAN INSTITUTE OF TECHNOLOGY KHARAGPUR
KHARAGPUR - 721302, INDIA
CERTIFICATE
This is to certify that the project report entitled “Mapping crops within the grow-
ing season” submitted by S Rohit Chandra Jagadeesh (Roll No. 18MT3AI32) to
Indian Institute of Technology Kharagpur towards partial fulfilment of requirements for
the award of degree of Dual Degree in Centre of Excellence in Artificial Intelligence is a
record of bona fide work carried out by him under my supervision and guidance during
Autumn Semester, 2022-23.
Prof Adway Mitra
Date: November 28, 2022 Centre of Excellence in Artificial Intelligence
Place: Kharagpur Indian Institute of Technology Kharagpur
Kharagpur - 721302, India
ii
Acknowledgements
I want to thank my project supervisor, Professor Dr. Adway Mitra, whose expertise
was invaluable in formulating the research questions that helped push my limits and
work on various domains.
I was fortunate to be admitted by the Indian Institute of Technology Kharagpur, where
I met many excellent classmates and had an unforgettable time with them. Secondly, I
would like to thank my parents, who care about me all the time.
Date: November 28, 2022 (S Rohit Chandra Jagadeesh)
Place: Kharagpur (18MT3AI32)
iii
Abstract
Name of the student: S Rohit Chandra Jagadeesh Roll No: 18MT3AI32
Degree: Dual Degree
Department: Centre of Excellence in Artificial Intelligence
Thesis title: Mapping crops within the growing season
Thesis supervisor: Prof Adway Mitra
Month and year of thesis submission: November 28, 2022
The goal of this study was to map crops across the Continental US (CONUS) before the
harvest, and to estimate the earliest date of classification by which crops can be mapped
with sufficient accuracy (90% of full-season accuracy).
The first step in the crop classification was to perform Multivariate Spatio-Temporal
Clustering (MSTC) of annual MODIS-derived NDVI trajectories to create phenologically
similar regions, or phenoregions.
The second step was to assign crop labels to phenoregions based on spatial concordance
between phenoregions and crop classes from CDL using Mapcurves.
iv
Contents
Declaration i
Certificate ii
Acknowledgements iii
Abstract iv
1.5
1 Introduction 1
2 Study area and datasets 3
2.1 Study area and training data . . . . . . . . . . . . . . . . . . . . . . . . . . 3
2.2 Remote sensing data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
3 Methods 4
3.1 Development of Phenoregions . . . . . . . . . . . . . . . . . . . . . . . . . 4
3.2 Spatio-temporal variability in crop phenology . . . . . . . . . . . . . . . . 4
3.3 Climatic ecoregions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
3.4 Crop classification model . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
3.4.1 Cluster-then-label model training . . . . . . . . . . . . . . . . . . . 5
3.5 Evaluation metrics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
3.5.1 Accuracy assessment . . . . . . . . . . . . . . . . . . . . . . . . . . 5
3.5.2 Shannon diversity of crop types . . . . . . . . . . . . . . . . . . . . 6
4 Results 7
4.1 Model training to address spatial variability . . . . . . . . . . . . . . . . . 7
4.2 Mapping crop types across the continental United States . . . . . . . . . . 7
4.3 Within-season mapping of crops . . . . . . . . . . . . . . . . . . . . . . . . 8
v
Contents vi
5 Discussion 9
6 Limitations, challenges and future steps 10
7 Conclusions 11
7.1 Data Availability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
7.2 Appemdix A. Supplementary data . . . . . . . . . . . . . . . . . . . . . . . 12
Chapter 1
Introduction
Accurate and timely monitoring of crops over national scales is critical for crop production
forecasts, water management, assessment and management of disaster and disturbance
impacts and characterizing land use for Earth system modeling
Near real-time national scale crop mapping is challenging because
1) crop phenology changes quickly over relatively short time scales, thus requiring remote
sensing data with a high temporal frequency,
2) crop-specific land cover maps, required for model development, need to be available
over large spatial scales,
3) crop phenology varies across space due to differences in environmental growing condi-
tions,
4) interannual variations in crop phenology caused by variations in climate make a classifier
trained on a single year perform poorly in another year, and
5) a spatial crop classifier needs to be efficient to be nationally scalable.
Unsupervised methods like k-means clustering, the ISODATA algorithm and Gaussian
mixture models have been used in the past to cluster features derived from a time series
of remotely sensed vegetation indices (Gumma et al., 2016; Skakun et al., 2017; Xiong
et al., 2017; Wang et al., 2019). Crop type labels were then assigned to these clusters
using spectral matching techniques or using spatially aggregated crop statistics at the
administrative level.
Supervised methods like decision tree algorithms (Pittman et al., 2010), support vector
machines (Waldner et al., 2015a), random forests (Shao and Lunetta, 2012), neural net-
works (Shao et al., 2010) and, more recently, deep learning approaches (Kussul et al.,
2017; Zhong et al., 2019) have also been successfully applied for crop classification at
small scales.
1
Chapter 1. Introduction 2
The choice of classification algorithm requires considering the type and volume of data,
target accuracy, ease of use, speed and scalability, usually posing trade-offs and compro-
mises
One of the challenges in large area crop mapping is the variation in the timing of crop
phenological development across climate zones, since it is influenced by climate, soil, to-
pography, etc., as well as farm technology, management practices, fertilization, irrigation,
etc.
Clustering algorithms have been successfully used (Hargrove and Hoffman, 2004) to create
ecoregions: regions on a map within which exist similar combinations of ecologically
relevant conditions like temperature, precipitation, soil and topographic properties
The objectives for this study were as follows:
• To create a national, crop-specific land cover map (with all of the crop types, as included
in the CDL) for the CONUS using time series of MODIS-derived Normalized Difference
Vegetation Index (NDVI) as inputs to a generalized cluster-then-label crop classifier.
• To create crop maps in near real-time during the current growing season and to study
the rate of increase in mapping accuracy as the season progresses for 8 major crop types
grown in the US: corn, soybeans, winter wheat, fallow/idle cropland, other hay/non alfalfa,
alfalfa, sorghum and rice
Chapter 2
Study area and datasets
2.1 Study area and training data
Crop type and acreage information collected in surveys from farmers during the current
growing season are used to train the CDL classifier.
We downloaded the CDL for 2008–2018 from the USDA NASS Data Portal (USDA,
2019b). Our crop classification model was trained over 2008–2014 and applied to the
period 2000–2018. The period of 2015–2018 was used as test years for the classifier.
2.2 Remote sensing data
Time series of smoothed and gap-filled NDVI generated from Collection 5 data streams
from Terra (MOD13Q1) and Aqua (MYD13Q1) satellite instruments for the CONUS
were downloaded from the Oak Ridge National Laboratory (ORNL) Distributed Active
Archive Center for Biogeochemical Dynamics (DAAC) for the period 2000-01-01 through
2018-12-31 (Spruce et al., 2016). The MODIS NDVI data set, at a spatial resolution of 231
m and an 8-day temporal frequency, was generated using the NASA Stennis Time Series
Product Tool (TSPT) (Spruce et al., 2011) to remove clouds and otherwise clean and
filter the time series temporally. The smoothed, gap-filled data set is nearly complete,
with few missing values, and is ideal for many phenological analyses and applications.
Files are available in netCDF format, one per year, for the period 2000–2018, as a time
series of 8- day maximum-value composited MODIS NDVI in Lambert Azimuthal Equal
Area projection.
3
Chapter 3
Methods
3.1 Development of Phenoregions
Phenoregions are regions having similar annual profiles of NDVI “greenness” phenology
through space and time. Hargrove and Hoffman (2004) developed Multivariate Spatio-
Temporal Clustering (MSTC) based on a non-hierarchical k-means algorithm (Hartigan,
1975) for classification of phenoregions (White et al., 2005), classification of remote sensing
data (Hoffman et al., 2010), analysis of dynamic climate regimes in Global Circulation
Models (GCMs) (Hoffman et al., 2008), and detection of disturbance from phenological
time series (Mills et al., 2011).
MSTC does not explicitly use geographic location during classification and does not im-
pose spatial contiguity. Thus, a phenoregion may be comprised of many spatially disjoint
agricultural fields, so long as they have similar phenological profiles.
3.2 Spatio-temporal variability in crop phenology
This spatio-temporal variability in crop phenology adds complexity in phenology-based
identification of crop types. We address temporal variability by training the classifier on
multiple years, and we address spatial variability using ecoregions, thus developing a more
robust and accurate general crop classification model.
4
Chapter 3. Methods 5
3.3 Climatic ecoregions
Clustering algorithms have been widely used for classification of ecoregions (Hargrove and
Hoffman, 2004; Williams et al., 2008; Kumar et al., 2011). We used the same MSTC algo-
rithm (Hargrove and Hoffman, 2004) to divide the CONUS into 500 synoptic ecoregions,
representing regions with similar crop growing conditions
3.4 Crop classification model
3.4.1 Cluster-then-label model training
Crop pixels from the CDL and the spatially concordant pixels from phenoregions present
within each ecoregion were randomly divided into training (70%) and validation (30%)
sets for each year
Mapcurves, a quantitative method that calculates the spatial concordance between two or
more categorical maps and provides an assignment of labels between the maps (Hargrove et
al.,2006), was used to compare phenoregions with the CDL and assign crop type labels to
entire phenoregions. Mapcurves calculates a pairwise Goodness-of-fit statistic (GOF) over
all categories in the two maps being compared. The GOF statistic between a phenoregion,
P, with a crop type, C, was defined as follows:
where AP and AC represent the area under P and C, respectively, and AP AC represents
the area that is common to P and C.
3.5 Evaluation metrics
3.5.1 Accuracy assessment
While all crop types contained in the CDL were analyzed and mapped in our study, we
focus our accuracy assessment here on the 8 dominant crop types (by area across CONUS):
corn, soybeans, winter wheat, fallow/idle cropland, other hay/non alfalfa, alfalfa, sorghum
and rice
Crop types other hay/non alfalfa and fallow/idle cropland are referred to as other hay
and fallow, respectively in all tables and figures. Three metrics were used to evaluate the
Chapter 3. Methods 6
accuracy of classification: Producer’s Accuracy, User’s Accuracy and Overall Accuracy,
as defined in Eqs. (2), (3) and (4).
Producer’s Accuracy is the accuracy of the map from the map producer’s point of view,
quantifying the probability that a feature class on the ground is correctly classified by the
map.
User’s Accuracy is the accuracy from the user’s perspective, and quantifies the reliability
of the map, i.e., the probability that a feature on the map will actually be present on the
ground
Overall Accuracy quantifies the fraction of the reference CDL pixels that are correctly
mapped by our crop classification method.
User’s Accuracy is the most relevant for a farmer or resource manager; thus, we focus
our discussion on User’s Accuracy, and include the Producer’s Accuracy statistics in the
Supplementary Material.
3.5.2 Shannon diversity of crop types
Omission and commission errors due to confusion between crop types are larger in regions
with diverse crop types, and when the cultivated field sizes are smaller than the resolution
of MODIS products. We calculate the Shannon Diversity Index (H) to quantify the
diversity of crop types within an area:
where pi is the proportion of map grid cells belonging to a crop type i in the mapped area.
When only one kind of crop exists in an area, H has a value of zero, and crop classification
is easy. H increases when there are more crop types present and their probabilities are
more uniform within the area.
Chapter 4
Results
4.1 Model training to address spatial variability
Training for smaller regions reduced the spatial variability and thus allowed more special-
ized models for those regions.
4.2 Mapping crop types across the continental United
States
The cluster-then-label model was also applied to the test data set with four never-seen-
before years 2015–2018. Overall Accuracy for the years 2015–2017, over all 102 crop types,
is slightly lower (compared to 2008–2014) at 58% and is 53% for 2018.
Region A from the Corn Belt, where corn and soybeans are the dominant crops, shows
broad agreement between the cluster-then-label-based map and CDL.
Region B in winter wheat-producing areas in Kansas demonstrates broad agreement be-
tween the two maps.
Region C from Central Valley, California, exhibits immense diversity in crop types grown
across small-sized fields and thus represents a difficult-to-classify region; yet the cluster-
then-label model is able to classify the crop types in this region with reasonable accuracy.
The cluster-then-label model performs well in terms of Overall Accuracy (Fig. 4(a)) in
major crop growing regions with large field sizes and lower diversity, but has comparatively
lower accuracy in regions with high crop diversity and smaller field sizes (Fig. 4(b)).
7
Chapter 4. Results 8
4.3 Within-season mapping of crops
While the accuracy of classification is low during early winter months, a large improvement
is observed during July when corn and soybeans reach maturity.
Chapter 5
Discussion
The goal of this study was to map crops across the CONUS during the active current
growing season as they grow, a critical step in near real-time crop health monitoring.
Accuracy in crop type classification was improved when cluster-thenlabel models were
trained within each ecoregion, compared to being trained for individual states or the
entire CONUS.
Pixel-wise Producer’s Accuracy and User’s Accuracy for major crops like corn, soybeans
and winter wheat were greater than those of lesscommonly grown crops during both
training and testing periods.
While the cluster-then-label method exploits the salient differences in phenological devel-
opment of the crops, errors in the crop type classification sometimes remain, due to the
inherent similarity in NDVI profiles among crop types.
9
Chapter 6
Limitations, challenges and future
steps
The ability to create a gap-filled remote sensing product that spans the whole CONUS
is critical for near real-time crop health monitoring and commodity yield predictions. It
requires remote sensing products that are corrected for missing values due to clouds/snow
cover. NDVI values were used as an integrative proxy to capture crop land surface phenol-
ogy. Past studies have included additional spectral bands spanning optical, Near Infrared,
Short Wave Infrared (SWIR) and Synthetic Aperture Radar (SAR), as well as indices that
are derived from them, like Enhanced Vegetation Index (EVI), Green Chlorophyll Vege-
tation Index (GCVI), Land Surface Water index (LSWI), Normalized Difference Tillage
Index (NDTI), among others. The addition of these bands and indices has been shown to
improve classification accuracy, and future studies could include such additional metrics.
The Mapcurves algorithm and cluster-then-label model assigned a single crop label (having
the best Goodness-of-fit) to the entire phenoregion based on a single majority “winner
takes all” strategy; however, other overlapping crop types might also be significant.
Given the amount of remote sensing and CDL data available, more sophisticated machine
learning, deep learning or Bayesian algorithms could also be tested.
The CDL, which served as our training data, itself is a classification product and likely
contains errors that will be propagated forward.
As crop mapping models increase in prognostic power, this limited availability of public,
error-free training data may become the greatest limitation to future progress in remote
sensing-based national crop classification.
10
Chapter 7
Conclusions
The goal of this study was to produce national-scale crop maps at 8-day intervals during
the growing season. We first developed a cluster-then-label approach to create end-of-
growing season crop maps for CONUS. This was done using a generalized classification
approach consisting of two steps:
1) creating phenoregions based on Multivariate Spatio-Temporal Clustering of annual time
series of 8-day NDVI collected for every 231 m pixel on the ground across CONUS for the
years 2000–2018, and
2) assigning crop labels to phenoregions based on the degree of spatial concordance be-
tween crop growing areas and entire, individual phenoregions
Spatial and temporal variability in phenology increases the challenges of national crop
mapping and were addressed by training the cluster-then-label models within each ecore-
gion and on multiple years (2008–2014), respectively.
We then used this approach to generate crop maps for CONUS well before harvest, and
to estimate the earliest time during the growing season by which crops could be mapped
with sufficient accuracy.
Running updated projections of final crop yields for each planted crop during the growing
season, estimated from historical productivity data per hectare within each ecoregion,
may be a feasible next step.
7.1 Data Availability
The data products from this study are publicly available at doi:[Link]
The data collection includes annual crop type maps for the period 2000–2018. It also in-
cludes the earliest dates of classification for eight dominant crop types (corn, soybeans,
11
Chapter 7. Conclusions 12
winter wheat, fallow/idle cropland, alfalfa, other hay/nonalfalfa, sorghum, and rice) dur-
ing 2015.
7.2 Appemdix A. Supplementary data
Supplementary data to this article can be found online at https:// [Link]/10.1016/[Link].2020.112048.