0% found this document useful (0 votes)
7 views44 pages

Multimodal AI for Wheat Breeding Insights

This study introduces a multimodal large language model (WBLM) for wheat breeding that integrates UAV remote sensing technology and cross-domain data to enhance breeding efficiency and accuracy. By employing supervised fine-tuning, retrieval-augmented generation, and reinforcement learning, the WBLM demonstrates superior performance in wheat yield prediction and provides decision support for various breeding tasks. The research aims to address challenges in wheat breeding through innovative data fusion and intelligent solutions to support sustainable agricultural development.

Uploaded by

distaste19
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views44 pages

Multimodal AI for Wheat Breeding Insights

This study introduces a multimodal large language model (WBLM) for wheat breeding that integrates UAV remote sensing technology and cross-domain data to enhance breeding efficiency and accuracy. By employing supervised fine-tuning, retrieval-augmented generation, and reinforcement learning, the WBLM demonstrates superior performance in wheat yield prediction and provides decision support for various breeding tasks. The research aims to address challenges in wheat breeding through innovative data fusion and intelligent solutions to support sustainable agricultural development.

Uploaded by

distaste19
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Multimodal large language model for wheat breeding:

a new exploration of smart breeding


Guofeng Yanga, Yu Lib, Yong Hea, Zhenjiang Zhoua, Lingzhen Yec, Hui

Fanga, Yiqi Luod, Xuping Fenga*

Affiliations
a
College of Biosystems Engineering and Food Science, Zhejiang University,
Hangzhou 310058, Zhejiang, China
b
Zhejiang Society of Agricultural Machinery, Hangzhou 310003, Zhejiang, China
c
College of Agriculture and Biotechnology, Zhejiang University, Hangzhou
310058, Zhejiang, China
d
Soil and Crop Sciences Section, School of Integrative Plant Science, Cornell
University, Ithaca 14853, New York, USA
*
Correspondence: fengxp@[Link]

Abstract
UAV remote sensing technology has become a key technology in crop breeding,
which can achieve high-throughput and non-destructive collection of crop phenotyping
data. However, the multidisciplinary nature of breeding has brought technical barriers
and efficiency challenges to knowledge mining. Therefore, it is important to develop a
smart breeding goal tool to mine cross-domain multimodal data. Based on different pre-
trained open-source multimodal large language models (MLLMs) (e.g., Qwen-VL,
InternVL, Deepseek-VL), this study used supervised fine-tuning (SFT), retrieval-
augmented generation (RAG), and reinforcement learning from human feedback
(RLHF) technologies to inject cross-domain knowledge into MLLMs, thereby
constructing multiple multimodal large language models for wheat breeding (WBLMs).
The above WBLMs were evaluated using the newly created evaluation benchmark in
this study. The results showed that the WBLM constructed using SFT, RAG and RLHF
technologies and InternVL2-8B has leading performance. Then, subsequent
experiments were conducted using the WBLM. Ablation experiments indicated that the
combination of SFT, RAG, and RLHF technologies can improve the overall generation
performance, enhance the generated quality, balance the timeliness and adaptability of
the generated answer, and reduce hallucinations and biases. The WBLM performed best
in wheat yield prediction using cross-domain data (remote sensing, phenotyping,
weather, germplasm) simultaneously, with R2 and RMSE of 0.821 and 489.254 kg/ha,
respectively. Furthermore, the WBLM can generate professional decision support
answers for phenotyping estimation, environmental stress assessment, target
germplasm screening, cultivation technique recommendation, and seed price query
tasks. This study aims to provide intelligent and integrated solutions for wheat breeding
goals, to help breeding work be carried out efficiently, to accelerate the breeding process
of excellent varieties, and to provide scientific basis and technical support for achieving
sustainable agricultural development and ensuring food security.

Keywords
Data fusion; Remote sensing; Multimodal large language model; Cross-domain
knowledge; Smart breeding

1. Introduction

As one of the basic food crops for mankind, wheat breeding is facing
unprecedented challenges and opportunities due to the global food crisis and the
promotion of sustainable agricultural development (Cheng et al., 2024; Senapati et al.,
2022). Although traditional breeding methods have achieved remarkable results, their
efficiency and accuracy have gradually revealed their limitations in the face of complex
and changing climatic conditions, increasingly serious threats from pests and diseases,
and constantly upgrading consumer demands (Hu et al., 2022; Xiong et al., 2021). Due
to the diversity of wheat varieties, breeding information has long lacked a unified tool,
and data knowledge has shown an "isolated" distribution, which has created barriers to
the learning of wheat breeding knowledge (Bhat et al., 2023). Simultaneously, since
wheat breeding involves the intersection of multiple disciplines such as biology,
genetics, weather, and soil science, professionals have to cross literature and data from
many fields when engaged in breeding work, and even need to write code to access data,
which greatly limits their work efficiency (Kaur et al., 2021). Smart breeding, as an
innovative mode that integrates modern information technology, biotechnology, and
agricultural science, is gradually becoming a key path to solving this problem (H. Li et
al., 2024; Xu et al., 2022).
With the rapid development of unmanned aerial vehicle (UAV) technology, its
application in the agricultural field has become more and more extensive, especially
UAVs have shown great potential in intelligent breeding (Das et al., 2021; Fei et al.,
2023; Jiang et al., 2021). With its advantages of high efficiency, precision and flexibility,
it is becoming an indispensable tool in smart breeding and can provide important
monitoring data for smart breeding. Image recognition technology can accurately
extract wheat growth indicators from RS images, such as chlorophyll content (Feng et
al., 2023; Zhang et al., 2024) and leaf area index (Chen et al., 2022; Du et al., 2023);
text analysis technology can extract valuable germplasm descriptions (L. Liu et al.,
2024; Sansaloni et al., 2020), breeding optimization strategies (J. Xu et al., 2023; Yao
et al., 2022) and market trends (Garg et al., 2022; Padhy et al., 2024) from research
manuscripts, reports and databases. The complementarity of image and text information
not only enhances the credibility and interpretability of the data, but also promotes the
integration of interdisciplinary knowledge, opening up new avenues for wheat breeding
research. At present, there is an urgent need to integrate cross-domain data (RS data,
phenotyping data, environment data, germplasm data, cultivation data, and price data),
use artificial intelligence algorithms and big data technology to build a cross-scale and
cross-domain comprehensive analysis tool for wheat breeding, accelerate the screening
and optimization of wheat target varieties, and improve the efficiency and accuracy of
wheat breeding.
In recent years, the large language model (LLM) has attracted much attention due
to their powerful text generation and understanding capabilities (X. Huang et al., 2024;
Lin et al., 2024). After fine-tuning, the LLM has shown strong interactive capabilities
and the potential to improve productivity (Min et al., 2024). However, since the LLM
can only process plain text and cannot process images, voice, and video, their
application scope is limited (Wang et al., 2023). In the context of cross-domain, multi-
modal data fusion applications, several multimodal large language models (MLLM)
have been developed to enhance the ability of LLM to perceive and understand visual
signals (AI et al., 2024, p. 0; Bai et al., 2023; Z. Chen et al., 2024; DeepSeek-AI et al.,
2024; GLM et al., 2024; F. Li et al., 2024). Although a lot of study has been done to
explore the limitations and effectiveness of MLLM, the current open-source general
MLLM still has the problem of insufficient accuracy in professional applications (D.
Huang et al., 2024; C. Li et al., 2024), such as wheat breeding selection. When faced
with wheat breeding query, general MLLM often evades the question or gives irrelevant
answers due to lack of breeding knowledge. This hinders the further exploration and
application of MLLM in the field of wheat breeding.
However, the research and application of MLLM in crop breeding is still in its early
stages and faces many challenges and opportunities (Zhu et al., 2024). The MLLM
needs to integrate more complex data from different sources, including but not limited
to RS data, phenotyping data, and environment data during the crop growth period. Data
from these sources vary significantly in format, dimensionality, and sparsity. How to
efficiently and accurately integrate these cross-domain heterogeneous data is a
difficulty in current research (Kuska et al., 2024). MLLM needs to have strong learning
and generalization capabilities to adapt to the needs of smart breeding and predict the
performance of crops under different environmental conditions, which places extremely
high demands on the model's algorithm design and computing resources. Additionally,
professional knowledge and experience in the field of crop breeding are crucial for the
development and application of the model (J. Li et al., 2024). How to integrate
professional knowledge into model design to make the model more in line with actual
breeding needs, while discovering new breeding strategies and optimization solutions
through the model, is another issue that needs in-depth exploration. MLLM is expected
to reveal the deep mechanisms of crop trait formation by integrating cross-modal
information and accelerate the screening and breeding process of new varieties.
This study aims to innovatively construct a MLLM for wheat breeding (WBLM)
through cross-domain data fusion and cutting-edge technology application, and explore
its potential in wheat breeding goals. The purposes of this study are to: (i) evaluate the
contribution of the integrated application of domain knowledge technologies
(supervised fine-tuning, retrieval augmented generation, and reinforcement learning
from human feedback) to achieving wheat breeding goals, and analyze the performance
of cross-domain data fusion in wheat yield prediction; (ii) explore the response of
WBLM in coping with multidimensional breeding goals, and generate personalized
decision support from the aspects of phenotyping estimation, environmental stress
assessment, target germplasm screening, cultivation technique recommendation, seed
price query; (iii) release the study dataset to promote research and application
innovation in this field.

2. Study area and data

2.1 Experimental sites and design

The experiment was conducted for two years (2021-2022 and 2022-2023), with
305 and 351 varieties, respectively. The varieties in the second year included the
varieties in the first year and 46 varieties were added (Table S1). The experimental sites
were set up in Changxing and Yuhang Agricultural Experimental Bases of Zhejiang
University Agricultural Experiment Station in Zhejiang Province, China (Fig.1a, Table
S2). Specific planting information is shown in Table S3. The experimental fields were
not sprayed with pesticides and herbicides, and weeds were not controlled manually or
mechanically. Maintaining the natural growth environment is to develop competitive
wheat varieties that have a greater competitive advantage than wheat weeds. The water
required for wheat growth comes from natural rainfall. Both places have a subtropical
monsoon climate.
Fig. 1. (a) Experimental sites. (b) Multi-source data acquisition.

2.2. Data acquisition

2.2.1. UAV data


For the experimental field, we used multiple sensors and multiple UAVs to acquire
hyperspectral (HS), lidar, multispectral (MS), and RGB (red, green, blue) data on the
same day and compiled them into a RS dataset (Fig. 1b). The Matrice 300 RTK (DJI,
China) was selected as the flight platform to carry the Zenmuse L1 (DJI, China) lidar
camera and the FireflEYE S185 (Cubert, Germany) hyperspectral camera to obtain lidar
data and HS data respectively. FireflEYE S185 covers the visible and near-infrared
bands from 450 to 950 nm, has 125 spectral channels, and can achieve synchronous
frame imaging, spectral resolution: 8 nm @ 532 nm, spectral sampling interval 4 nm.
Zenmuse L1 has a range of 190m @ 10%, 100 klx, point cloud data rate multi-echo:
maximum 480,000 points/second, elevation accuracy: 5cm @ 50m, plane accuracy:
10cm @ 50m. The Phantom 4 multispectral (DJI, China) and the Mavic 3E (DJI, China)
were used to obtain MS data and RGB image respectively. The MS includes five bands,
blue: 450 nm ± 16 nm; green: 560 nm ± 16 nm; red: 650 nm ± 16 nm; red edge: 730
nm ± 16 nm; near infrared: 840 nm ± 26 nm. The Mavic 3E is equipped with a wide-
angle camera 4/3 CMOS, with 20 million effective pixels and a maximum photo size
of 5280×3956. The telephoto camera has an equivalent focal length of 162 mm, 12
million pixels, and 56x hybrid zoom.
UAVs were used to collect data during the wheat growth period many times, all
under low wind speed and clear and cloudless conditions at noon using different UAVs
and sensors. During automatic data collection, the UAV's flight altitude was 15 m, and
other flight parameters such as overlap rate were set according to the parameters
recommended by the equipment manufacturer. During manual data collection, the
Mavic 3E UAV flight altitude was maintained at 3 m and a 7x zoom was set to obtain
images and video data of the experimental field. In particular, since the UAV used
supports image-free control technology and the experimental field area is relatively
small, no ground control points were set and no geometric correction and registration
of the data was performed.

2.2.2. Phenotypic data


Phenotypic data collection involves chlorophyll content, leaf area index (LAI),
canopy height (CH) and yield (Fig. 1b). Field phenotyping data of wheat were collected
six times during 2021-2022 (on January 19, March 8, April 2, April 15, April 29, and
May 18, 2022) and nine times during 2022-2023 (on December 27, 2022; February 16,
March 4, March 19, April 7, April 22, May 1, May 13, May 19, 2023). UAV RS and
phenotyping data were collected on the same day.
The chlorophyll content of the wheat canopy was measured using a SPAD-502
PLUS chlorophyll meter (Konica Minolta, Japan). Five wheat plants with
representative growth in each wheat planting plot of the field experiment were
randomly selected. For each wheat plant, three leaves without pests and diseases,
physiological spots, and mechanical damage were selected, and the tip, middle, and
base of each leaf were measured. All the measured values were averaged as the SPAD
value of the canopy leaves of the wheat germplasm in the plot. The LAI of the wheat
canopy in each plot was measured using a LAI-2000C plant canopy analyzer (LI-COR,
USA). Five sampling points with uniform growth were selected at the four corners and
the center of each plot, and the LAI values were measured and recorded using a LAI-
2200C according to the standard method. Each sampling point was measured three
times and the average value was taken as the LAI of the sampling point. We recorded
the average of the LAI of the five sampling points after scatter correction as the LAI of
the plot. In each plot, five CH samples were randomly selected for measurement using
a tape measure, and the average value was taken as the CH of the plot. The manually
measured CH was used to verify the CH derived from the UAV light detection and
ranging (lidar) point cloud. Wheat was harvested when mature in each plot. After
threshing, the grain from each yield plot was weighed and yield was expressed as kg
ha-1 and standardized to a moisture content of 12.5% (State Administration for Market
Regulation and Standardization Administration of the P.R.C, 2023).

2.2.3. Environmental data


The weather data continuously recorded by meteorological equipment at each
agricultural experimental base and the meteorological observation station of the China
Meteorological Administration were used to provide us with weather data for two-year
growing season (Fig. 1b). The data mainly include daily average temperature, dew point
temperature, precipitation, net solar radiation intensity, wind speed and other data. The
average annual temperature is 15.6 ℃, the average relative humidity is 76%, the annual
rainfall is 1309 mm, and the average of 1810 hours of sunlight per year at Changxing
Agricultural Experiment Station. The average annual temperature is 16.2 ℃, the
average relative humidity is 68%, the annual rainfall is 1400 mm, and the average of
1970 hours of sunlight per year at Yuhang Agricultural Experiment Station.
The test soil was collected before sowing at different locations (0-20 cm) in the
experimental field. The collected soil was placed in turnover boxes and then sent for
testing in time to determine the physical and chemical properties of the soil. Soil pH =
6.2, total N = 1.32 g kg-1, available potassium = 94.9 mg kg-1, available phosphorus =
1.9 mg kg−1 and soil organic C = 12.4 g kg−1 at Changxing Agricultural Experiment
Station. Soil pH = 6.14, total N = 1.68 g kg−1, available potassium = 187 mg kg−1,
available phosphorus = 63.6 mg kg−1 and soil organic C = 16 g kg−1 at Yuhang
Agricultural Experiment Station.

2.3. Data processing


2.3.1. Data preprocessing
The acquired UAV data were preprocessed to generate digital orthophoto map
(DOM) for subsequent processing according to each plot. Hyperspectral data were
converted into reflectance images after radiometric correction using the supporting
Cubert Utils Touch software (Cubert, Germany) and radiometric correction plates. Then,
Agisoft Metashape (Agisoft, Russia) was used to align photos, create dense point clouds,
generate grids, and generate textures to get DOM for subsequent processing. The MS
data was imported into Pix4Dmapper (Pix4D, Switzerland) software for initialization
processing, point cloud and texture, DSM and DOM generation operations. The
irradiance value captured by the light intensity sensor during flight was used to
compensate the MS bands for illumination, eliminating the interference of ambient light
on data collection. The bands are then radiometrically calibrated using the radiometric
calibration plate image acquired simultaneously during data collection. Furthermore,
the radiometrically calibrated images were stitched to generate DOM. Finally, the DOM
was normalized to obtain reflectance images for subsequent generation of vegetation
indexes (VI). The lidar data were processed using DJI Terra (DJI, China) to generate
three-dimensional point clouds. The RGB images were imported into Pix4Dmapper to
stitch RGB DOM. In addition, one frame of the video obtained by manual flight is
extracted every five seconds as an image for storage. These images were divided
according to different germplasms (different plots). Then, the wheat heads (WH) in the
images were manually annotated to produce a WH dataset.

2.3.2. Spectral data processing


The processed MS and HS bands were used for VIs calculation to analyze the
canopy spectral features of different plots. VIs commonly used for phenotyping
estimation and grain yield prediction were selected, including NDVI, SAVI, kNDVI,
NIRv and PSRI (Table 1).
The fractional vegetation cover (FVC) of the plots, i.e., the percentage of the
vegetation area to the plot area (Yang et al., 2022), was calculated as a valuable
indicator of crop density and structural information. The vegetation area of the plot was
extracted from the MS image by excluding the background soil using a Transformer-
based segmentation algorithm, similar to related study (Cui et al., 2023). The number
of vegetation pixels in each plot was then divided by the total number of pixels in that
plot to calculate FVC (Maimaitijiang et al., 2020).

2.3.3. Lidar data processing


The CH was extracted from LiDAR point clouds and used as canopy structure
features. The bare ground digital elevation model (DEM) was created using point cloud
data acquired before the emergence of seedlings in the field. Then, the digital surface
model (DSM) representing all objects (vegetation) on the ground was constructed based
on the point cloud data acquired at different times. Thus, the CH is obtained by pixel-
level subtraction of the generated DSM and DEM (Maimaitijiang et al., 2020). Finally,
the obtained lidar data were used to extract CH in different plots in the wheat field.
In addition, the measurement tool of Terra software (DJI, China) was used to mark
the reference surface for canopy volume (CV) measurement of each plot based on the
lidar point cloud and the boundaries of different plots. The sum of the excavated volume
and the filled volume above and below the reference surface is taken as the spatial
volume of the reference surface. It should be noted that since there are two types of
reference planes, the lowest point (the plane on which it is located) and the average
plane, the average value of the volumes obtained from the lowest point and the average
plane is taken as the spatial volume of the plot (Table 1).

2.3.4. RGB image processing


To obtain the lodging level of different plots in the wheat field, we refer to a study
to construct a wheat lodging area segmentation model (Zhang et al., 2023). Then we
use the ratio of the lodging area extracted from the plot to the plot area to determine the
plant lodging (PL) level of the plot. In our experiment, the lodging levels were divided
into: no lodging (0), slight lodging (0-50%), severe lodging (50-100%), and special (no
crop or only a few plants) (Saskatchewan Seed Growers’ Association, 2024).
We implemented the detection and counting of WH in the field based on the two-
stage method FR-Transformer proposed in our previous study (Zhu et al., 2022) and
RGB images. Then, we used this method to obtain the number of WH in the
preprocessed WH images of each plot, and averaged the number of WH in each plot.
Finally, the ground coverage area of the WH image was calculated using the UAV flight
altitude and camera field of view (Avola et al., 2021), and the number of WH per unit
area of the plot was further calculated.
Due to the lack of herbicide spraying and manual weeding, weeds in wheat fields
grow randomly and in large numbers. To study the environmental (weed) stress of wheat
growth, we focused on the severity of weeds in each plot and the area 10 to 20 cm wide
outside the plot in the wheat field. We referred to a study (Anderegg et al., 2023) to
extract the number of pixels in the plot and specific area and the number of weed pixels
in them, and then used the ratio of the number of weed pixels to the number of pixels
in the plot and specific area to classify each plot according to the weed level (WL). The
weed level can be roughly divided into: no weeds (0-10%), slight weeds (10-40%),
moderate weeds (40-70%), and severe weeds (70-100%).

Table 1. Definition of features extracted from different sensors


Sensor Features Formulation References

MS (Spectral Normalized Difference NDVI =


(NIR – R) (Tucker,
(NIR + R)

information) Vegetation Index 1979)

Soil-Adjusted Vegetation SAVI =


(1 + L)×(NIR – R)
, L = 0.5 (Huete, 1988;
(NIR + R + L)

Index Richardson

and Everitt,

1992)

NIR−R 2
kernel Normalized kNDVI = tanh (( ) ), where σ is a tunable (Camps-Valls

Difference Vegetation length-scale parameter intended to capture et al., 2021)

Index nonlinear sensitivity of NDVI to vegetation

density. If σ = 0.5(NIR + R), which simplifies

to kNDVI = tanh ((NDVI)2).

Near-Infrared Reflectance NIRv = NIR × NDVI (Badgley et

of Vegetation al., 2017)


Plant Senescence PSRI =
(R – G)
, R, G, and NIR denote the (Cao et al.,
NIR

Reflectance Index 2019)


reflectance of the red, green, and near-infrared

bands from the 5-band Phantom 4 multispectral,

respectively.

HS (Spectral Normalized Difference NDVI =


(NIR – R) (Tucker,
(NIR + R)

information) Vegetation Index 1979)

Soil-Adjusted Vegetation SAVI =


(1 + L)×(NIR – R)
, L = 0.5 (Huete, 1988;
(NIR + R + L)

Index Richardson

and Everitt,

1992)

NIR−R 2
kernel Normalized kNDVI = tanh (( ) ), where σ is a tunable (Camps-Valls

Difference Vegetation length-scale parameter intended to capture et al., 2021)

Index nonlinear sensitivity of NDVI to vegetation

density. If σ = 0.5(NIR + R), which simplifies

to kNDVI = tanh ((NDVI)2).

Near-Infrared Reflectance NIRv = NIR × NDVI (Badgley et

of Vegetation al., 2017)

Plant Senescence PSRI =


(R680 – R500)
, R680, R500 and R750 denote the (Cao et al.,
R750

Reflectance Index 2019)


reflectance at 680 nm, 500 nm, and 750 nm

respectively.

LiDAR Canopy Height (CH) CH = DSM - DEM (Maimaitijian

(Structural g et al., 2020)

information)

LiDAR Canopy Volume (CV) CV = Excavated volume + Filled volume, DJI [Link]

(Structural Terra software (DJI) [Link]/d

information) ji-terra

RGB image Plant Lodging (PL) PL =


Number of lodging pixels in the plot
× 100% (Zhang et al.,
Total number of the plot pixels

(Structural 2023)

information)
RGB video Wheat Head (WH) WH = The number of wheat head detection (Zhu et al.,

(Reproduction 2022)

information)

RGB image Weed Level (WL) WL = (Anderegg et

(Environment Number of weed pixels in the plot and specific area


× 100% al., 2023)
Total number of the plot and specific area pixels

al

information)

MS Fractional Vegetation Cov FVC =


Number of crop pixels in the plot
× 100% (Cui et al.,
Total number of the plot pixels

(Environment erage (FVC) 2023)

al

information)

* R, G and NIR are the pixel values of the red, green and near-infrared bands, respectively.

2.4. Construction of Cross-domain knowledge base

The cross-domain knowledge base (in Chinese and English) consists of multi-
source datasets in the field (Fig. 2a) and external domain knowledge base (Fig. 2b). The
multi-source datasets in the field include UAV RS data, phenotyping data, weather data.
The external domain knowledge base includes:

(1) wheat germplasm data (22k), mainly including variety name, place of origin,
nutritional quality, resistance and agricultural traits. In terms of nutritional quality, key
parameters such as crude protein, lysine, and sedimentation value are involved. In terms
of resistance, it covers important characteristics such as disease resistance (stripe rust,
leaf rust, powdery mildew, etc.), drought resistance, and cold resistance. In addition,
agronomic traits such as maturity, plant height, thousand grain weight, and grain
hardness are also recorded.

(2) wheat cultivation technique data (10k), covers the entire process from pre-
planting preparation to post-harvest storage, mainly including the selection of suitable
varieties, soil preparation, sowing, fertilizer management, irrigation and pest and
disease control techniques.

(3) wheat plant protection technique data (10k), mainly including disease
prevention and control, pest prevention and control, weed management, as well as new
techniques such as precision plant protection based on artificial intelligence and RS,
agricultural UAV flight control and green prevention and control.

(4) wheat seed price data (20k), mainly including observation point, variety name,
price, specification, planting area and time, etc. For observation points, the economic
level and planting demand of different regions lead to price differences; for varieties,
different varieties have different characteristics and application scenarios, and high-
quality varieties are usually more expensive; different specifications of packaging and
planting areas affect the final selling price; prices change over time due to changes in
market supply and demand, especially natural disasters or policy adjustments.

The wheat germplasm data comes from the Chinese Crop Germplasm Information
Network ([Link] The data on wheat cultivation technique and wheat
plant protection technique come from search engines (Google, Bing, Baidu), and the
search terms include "wheat cultivation technique, wheat plant protection technique,
小麦栽培技术, 小麦植保技术". The wheat seed price data comes from the National Seed

Market Monitoring Information Release Platform ([Link] All data from


this study have been made publicly available
([Link]

3. Methods
Fig. 2. A workflow for the construction and application of a multimodal large
language model for wheat breeding (WBLM). (a) Multi-source dataset construction.
(b) External domain knowledge base construction. (c) The WBLM with domain
knowledge is constructed using supervised fine-tuning, retrieval augmented
generation, and reinforcement learning from human feedback. (d) The user's question
(image and text) is sent to WBLM. (e) The WBLM answers the question.

3.1. Construction of WBLM


We construct a WBLM with domain knowledge using the supervised fine-tuning
(SFT), retrieval augmented generation (RAG) and reinforcement learning from human
feedback (RLHF) technologies based on different MLLMs and cross-domain
knowledge base. WBLM consists of two parts. The part 1 builds a RAG system based
on the LlamaIndex large model application framework ([Link]
and the Milvus vector database ([Link] (Fig. 2c, yellow line). The part 2
uses STF and RLHF methods to obtain a WBLM that follows breeding preferences (Fig.
2c, red line). We regard the input of part 1 as a query (Fig. 2d), and the original query
and the output of part 1 as a new query. The part 2 receives the new query, processes it,
and then outputs the result to the Q/A system (Fig. 2e).
To make WBLM better adapt to breeding tasks, we use RAG technology to
conveniently and efficiently supplement the knowledge not covered by STF and RLHF
methods. First, we convert the data of the files in the external domain knowledge base,
use BGE-M3 to generate embedding vectors (J. Chen et al., 2024), and then import the
data into the Milvus vector database. Then, LlamaIndex receives user input and initiates
a retrieval request to the Milvus database based on the input, using a hybrid retrieval
(BGE-M3) and re-ranking (BGE-Reranker-V2-M3) pipeline. LlamaIndex merges the
retrieval results with the input to form a new prompt. After that, LlamaIndex inputs the
new prompt to InternLM2.5-7B-Chat to achieve reasoning and answer generation based
on the retrieved knowledge (Cai et al., 2024).
The training process using the STF and RLHF methods is divided into three stages.
In the first stage, we collect demonstration data and train a supervised policy. Based on
the cross-domain knowledge base, we artificially construct a question-answering
dataset that we hope the model generates aligned answers to SFT the pre-trained MLLM.
For questions we want the model to answer well, we collect the answers we want the
model to output to increase the probability of the model generating the expected
answers. We choose currently popular and competitive open-source MLLMs for model
construction and fine-tuning, including Qwen-VL (Qwen-VL-Chat), InternVL
(InternVL2-8B and InternVL2-2B), Yi-VL (Yi-VL-6B), Deepseek-VL (DeepSeek-VL-
7B-chat and DeepSeek-VL-1.3B-chat) and GLM-4 (GLM-4V-9B), all of which support
English and Chinese. It cannot be denied that we mainly provide a method to build an
MLLM suitable for wheat breeding, rather than having to choose a specific open-source
MLLM, as models with better performance are constantly emerging. We used Scalable
lightWeight Infrastructure for Fine-Tuning (SWIFT) to fine-tune the MLLM (The
ModelScope Team, 2024). We refer to the SWIFT recommended training script, in
which the LoRA (Low-Rank Adaptation) fine-tuning method is selected, the number of
epochs is set to 5, and no further hyperparameter adjustment is performed. All fine-
tuning and inference were performed on four V100 NVIDIA GPUs with 32GB.
Specifically, dataset D consists of the model input prompt and the answer that the
breeder expects the model to output, denoted by x and y respectively. The model is
denoted by 𝜋𝜋𝜃𝜃 , and the length of y is T. When prompt x is given, the probability that
the model generates answer y can be expressed as formula (1). Where πθ �yt �x, y1:t−1�

denotes the probability of the model outputting the t-th token given the input before the
t-th token.
T

πθ (𝑦𝑦|𝑥𝑥) = �� πθ �yt �x, y1:t−1 �� (1)


t=1

The SFT stage uses the following loss to train the model. For simplicity, we call
the model trained in this stage the SFT model.
T

LSFT = -𝐸𝐸(𝑥𝑥,𝑦𝑦)∼𝐷𝐷 [logπθ (𝑦𝑦|𝑥𝑥)] = -𝐸𝐸(𝑥𝑥,𝑦𝑦)∼𝐷𝐷 �� log πθ �yt �x, y1:t−1 �� (2)
t=1

The second stage collects comparative data consisting of questions and different
answers. Labels indicate their preferred answer given the input. A reward model (RM)
is then trained to predict the breeder's preferred answer. Specifically, the model takes

prompt and answer as input and outputs a scalar value. We use 𝑟𝑟𝜙𝜙 to denote the RM,

and 𝑟𝑟𝜙𝜙 (x,y) denotes the scalar output of the RM given prompt x and answer y. The

dataset D used to train the RM consists of the model input prompt, the answer that the
breeder wants the model to output, and the answer that the breeder does not want the
model to output, denotes by x, 𝑦𝑦𝑤𝑤 , and𝑦𝑦𝑡𝑡 respectively. We are given data (x, 𝑦𝑦𝑤𝑤 , 𝑦𝑦𝑡𝑡 ),
and use the maximum likelihood estimation loss to train the RM. σ denotes the
sigmoid function. During training, K answers are selected each time and combined in
pairs.

LR (𝑟𝑟𝜙𝜙 ) = -E(𝑥𝑥,𝑦𝑦𝑤𝑤 ,𝑦𝑦𝑡𝑡)∼𝐷𝐷 �logσ(𝑟𝑟𝜙𝜙 (𝑥𝑥, 𝑦𝑦𝑤𝑤 ) − 𝑟𝑟𝜙𝜙 (𝑥𝑥, 𝑦𝑦𝑡𝑡 ))� (3)

In the third stage, proximal policy optimization (PPO) is used to optimize the policy
for the RM. We use the output of the RM as a scalar reward and use the PPO algorithm
to fine-tune the SFT model to optimize this reward (Schulman et al., 2017). Among
them, the second and third stages can be carried out iteratively. We collect more
comparison data on the current best policy, use it to train a new RM, and then train a
new policy. In practice, most of our comparison data comes from supervised policies,
and some of it comes from PPO policies. After training the RM, we have obtained an
approximate reward function. Consider the model as a policy in the reinforcement
learning (RL) problem, and the input and output of the model can be regarded as the
state and action of RL respectively. We can use RL algorithms to train the model and
train the model to output the answer with the highest reward. However, the RL training
process is very unstable, so the SFT model is used as a reference model, denoted by
𝜋𝜋ref . The KL divergence regularization term with the reference model is added to the
objective function (4) to limit the update range of the model. Among them, the
hyperparameter β controls the degree of deviation from the reference model.

J𝑟𝑟𝜙𝜙 (𝜋𝜋𝜃𝜃 ) = 𝐸𝐸𝑥𝑥∼𝐷𝐷, 𝑦𝑦∼𝜋𝜋𝜃𝜃 �𝑟𝑟𝜙𝜙 (𝑥𝑥, 𝑦𝑦)� − 𝛽𝛽𝐷𝐷𝐾𝐾𝐾𝐾 �𝜋𝜋𝜃𝜃 (𝑦𝑦|𝑥𝑥)�𝜋𝜋ref (𝑦𝑦|𝑥𝑥)� (4)

According to the objective function (4), the KL divergence term is expanded, and
the final comprehensive reward is:

r(𝑥𝑥,𝑦𝑦) = 𝑟𝑟𝜙𝜙 (𝑥𝑥, 𝑦𝑦) − β(log𝜋𝜋𝜃𝜃 (𝑦𝑦|x) − log𝜋𝜋ref (𝑦𝑦|x)) (5)

The objective function is:

J𝑟𝑟𝜙𝜙 (𝜋𝜋𝜃𝜃 ) = 𝐸𝐸𝑥𝑥∼𝐷𝐷, 𝑦𝑦∼𝜋𝜋𝜃𝜃 [𝑟𝑟(𝑥𝑥, 𝑦𝑦)] (6)

We created three different datasets: (1) The SFT dataset (75k) is created using
multi-source datasets from the field, which contains questions and answers for training
the SFT model. 80% of this dataset is used for training, and 20% is used for accuracy
testing of wheat breeding model evaluation benchmark (phenotyping estimation task
and environmental stress assessment task). (2) The RM dataset, with breeder rankings
of model outputs, is used to train the RM. (3) The PPO dataset, without any breeder
labels, which are used as inputs for RLHF fine-tuning. External domain knowledge base
is mainly used by RAG methods. The dataset format refers to SWIFT (Table S4).
3.2. WBLM evaluation benchmark
In order to comprehensively evaluate the performance of MLLM in scientific
breeding work, it is necessary to build a standardized machine and human evaluation
benchmark. This benchmark aims to integrate challenges from various wheat breeding
tasks. This benchmark aims to integrate challenges from various wheat breeding tasks.
We have specifically designed many professional Chinese and English questions on
wheat breeding and corresponding standard answers, covering five tasks: phenotyping
estimation, environmental stress assessment, target germplasm screening, cultivation
technique recommendation, and seed price query. The phenotyping estimation includes
seven subtasks (Yield, SPAD, LAI, CH, CV, WH and PL), the environmental stress
assessment includes two subtasks (WL and FVC), the target germplasm screening
includes five subtasks: high quality (HQ), disease resistance (DS), drought resistance
(DR), maturity period (MP), adapt to mechanized (AM), the cultivation technique
recommendation includes two subtasks: cultivation technique (CT) and plant protection
technique (PPT), and the seed price query includes one subtask seed price (SP) query.
Through objective evaluation indicators and manual scoring and sorting, the evaluation
team conducted a detailed evaluation of the answers of multiple MLLMs including
WBLM, covering three aspects: accuracy, stability, and reasoning.
(1) Accuracy
We evaluate different subtasks under five tasks, respectively. For the phenotyping
estimation task and the environmental stress assessment task, the RMSE and R2
evaluation indicators are selected for regression tasks (estimating the values of FVC,
SPAD, LAI, CH, CV, WH, and Yield), and the accuracy evaluation indicator is selected
for the classification task (estimating the categories of WL and PL). For each subtask
under the target germplasm screening task and cultivation technique recommendation
task, each MLLM is tested multiple times (1k), and each test is manually judged
whether the answer is correct. Then the proportion of the number of correct answers to
the total number of tests is used as the evaluation indicator. For the seed price query
task, we counted whether each result after multiple (1k) tests was consistent with the
data in the knowledge base (price ± 10%), and took the proportion of the total number
of consistent data to the total number of tests as the evaluation indicator.
(2) Stability
We evaluate the consistency and robustness of the answers. The evaluation
indicator of stability is the proportion of times that satisfy the stability requirements in
the total number of tests (1k), where the stability requirement is that the deviation of
the numerical value in the answer is within ±10%, and whether the text satisfies the
requirement is manually judged. Consistency: Since cross-domain data are related, the
answers given after inputting different data from cross-domain data into the model
should remain basically consistent. Robustness: The model's answer or performance
should not change significantly when faced with small changes in input data.
Robustness requires that the model should have a certain degree of fault tolerance.
(3) Reasoning
We evaluate the logical deduction, inductive reasoning, and explanation of the
answers. The evaluation indicator of reasoning is the proportion of the total reasoning
score of a single model after multiple tests (1k) to all scores. After one test, x MLLM
models output x answers. After manually analyzing x answers, we select a score from 1
to x points for each answer and the score can only be selected once until all answers are
scored. We sum up the answer scores of different models after multiple tests as the total
reasoning score of the corresponding model. Then calculate the proportion of the total
reasoning score of different models to all scores ((1 + x) ∗ x / 2). Logical deduction:
The model can deduce new conclusions or answers through logical reasoning based on
known facts, rules and conditions. This requires the model to understand the logical
relationships in the question and accurately apply these relationships to deduce answers.
Inductive reasoning: The model can summarize general rules or patterns from a series
of specific examples or observations and give reasonable answers based on them.
Inductive reasoning requires the model to have the ability to abstract and generalize
information. Explanation: The reasoning process is explained or visualized to a certain
extent to help user understand the source and basis of the answer.

4. Results and Discussion


4.1. Evaluation of WBLM

The newly constructed wheat breeding model evaluation benchmark was used to
evaluate closed-source MLLMs (Qwen 2.5, ERNIE Bot 4 Turbo, GPT-4o Plus and
Gemini Pro 1.5, abbreviated as Qwen, ERNIE Bot, ChatGPT, Gemini) and WBLMs
built based on different open-source MLLMs (Qwen-VL, InternVL, Yi-VL, Deepseek-
VL and GLM-4) (Fig. 3-5, Table S5).
Accuracy evaluation: On various tasks of the evaluation benchmark, the
comprehensive performance of WBLM based on InternVL2-8B is better than that of
WBLM built on other open-source MLLMs, and is significantly better than closed-
source ChatGPT, Gemini, and Qwen. Among them, for LAI evaluation, the R2 based
on InternVL2-8B is 0.795 and the RMSE is 1.302. The R2 and RMSE of other open-
source MLLMs range from [0.674, 0.772] and [1.542, 2.486], and the R2 and RMSE of
closed-source MLLMs range from [0.118, 0.142] and [3.61, 3.985], respectively.
ChatGPT shows leading performance among open-source models, although still far
behind the WBLM based on InternVL2-8B. In addition, we find that consistent with
existing study (Z. Chen et al., 2024), the performance of MLLMs with small parameters
in the same series is lower than that of MLLMs with larger parameters, such as
InternVL2-8B and InternVL2-2B. We attribute this modest decline to the smaller size
of the MLLMs.
Stability evaluation: Compared with closed-source MLLMs, WBLMs based on
open-source models outperform in stability tests. Specifically, the WBLM based on
InternVL2-8B achieved the best performance in consistency, with a stability score of
0.811, while the stability score of the lowest-performing Qwen 2.5 was only 0.053. The
robustness scores of closed-sources MLLM ([0.829, 0.895]) are higher than those of
WBLM based on open-source MLLM ([0.735, 0.802]), because closed-source MLLM
cannot accurately answer relevant breeding questions in most tasks, and the answers
are still wrong and similar when faced with slight changes in input data. However,
among multiple WBLMs based on open-source MLLM, WBLM based on InternVL2-
8B has a higher robustness score.
Reasoning evaluation: The WBLM based on InternVL2-8B outperforms other
MLLMs in logical deduction and explanation evaluation, with the highest scores of
0.158 and 0.15, respectively. The WBLM based on GLM-4V-9B-chat outperforms other
MLLMs in inductive reasoning evaluation, with the highest score of 0.16. Compared
with accuracy and stability evaluation, reasoning evaluation tests the model's breeding
decision support and problem-solving capabilities more. Excellent reasoning ability can
enhance users’ trust in MLLM and improve user satisfaction (J. Li et al., 2024).
InternVL2-8B based WBLM shows leading performance in most benchmark tests
(Fig. 3-5, Table S5). However, the closed-source commercial MLLMs used in this study
encountered significant difficulties in the wheat breeding evaluation benchmark test
and performed poorly. The main reason is that closed-source commercial MLLM lacks
domain knowledge related to wheat breeding, while open-source MLLM obtains it
using a combination of domain knowledge technologies (Ming and Li, 2024). Although
closed-source MLLMs represented by ChatGPT achieved better performance on the
popular MLLM evaluation benchmarks (Fu et al., 2024; Y. Liu et al., 2024; Yue et al.,
2024), we did not build WBLM based on closed-source MLLMs because these MLLMs
do not support further implementation of the process of this study. If closed-source
MLLMs support richer fine-tuning and downstream task customization, they may gain
greater commercial value and more users. Although all open-source MLLMs can handle
multimodal tasks, different network architectures have different abilities to solve
semantic alignment between modalities when integrating visual and textual data, thus
affecting the accuracy of the final results. In addition, different training processes, tricks,
and private training data are also important factors affecting model performance (D.
Huang et al., 2024). In the future, researchers can build breeding models based on
MLLM with better performance to provide stronger breeding assistance efficiency and
results.
Fig. 3. Comparison of different MLLMs on the evaluation benchmark (accuracy).

Fig. 4. Comparison of different MLLMs on the evaluation benchmark (stability),


which includes consistency and robustness.

Fig. 5. Comparison of different MLLMs on the evaluation benchmark (reasoning).


This figure shows the total reasoning score as a proportion of all scores after multiple
tests of a single MLLM.

4.2. Impact of domain knowledge on breeding goals

4.2.1. Contribution of the combination of domain knowledge and technology to


breeding goals
We evaluated WBLM on the wheat breeding model evaluation benchmark and
observed strong performance of WBLM based on InternVL2-8B. In this section, we
conduct ablation experiments to test the performance of different domain knowledge
technologies combinations in different tasks (Fig. 6 and Table S6).
For the phenotyping estimation task, Yield, SPAD, LAI, CH, and CV show poor
performance under the MLLM-only method due to the same lack of domain knowledge.
Compared with the MLLM-only method, the MLLM and SFT method, the MLLM and
RAG method, and the MLLM and RLHF method all greatly improve the prediction
performance of these subtasks. Additionally, for the WH and PL subtasks, MLLM has
certain detection and classification capabilities. With the addition of SFT, RAG, and
RLHF technologies, the performance is further improved to achieve relatively high R2
and low RMSE. For the environmental stress assessment task, the results showed that
accurate classification of WL was simpler than specific values estimation of FVC.
When the method combining MLLM, SFT, RAG, and RLHF was used, the best
performance for WL prediction was 0.929 (accuracy), and the best performance for
FVC prediction was 0.841 (R2) and 0.052 (RMSE).
When only MLLM is used, the prediction results for the target germplasm
screening task and the seed price query task are poor. The main reason is that the MLLM
lacks domain knowledge and price dynamic data, so it cannot answer accurately.
Furthermore, the results show that when the MLLM is combined with one or more
domain knowledge technologies (SFT, RAG, RLHF), better prediction performance can
be achieved. For the cultivation technique recommendation task, because the MLLM
has been trained with many images and texts from various data sources, and some wheat
cultivation technique knowledge is also covered, thus it can answer some questions well.
It is undeniable that when one or more domain knowledge technologies (SFT, RAG,
RLHF) are combined, the cultivation technique recommendation is more accurate and
the data is traceable.
The method combining MLLM, SFT, RAG and RLHF performs better in different
subtasks of the five tasks. It is worth mentioning that by analyzing the prediction results
of the above tasks, it can be clearly found that hallucination exists when only the MLLM
method is used, which directly reflects the necessity of combining MLLM with domain
knowledge technology. With the increase of domain knowledge technology, the R2 of
the methods in the phenotyping estimation and environmental stress assessment tasks
gradually increased, and the RMSE decreased. Simultaneously, the consistency of the
prediction results in the target germplasm screening, cultivation technique
recommendation, and seed price query tasks gradually increased. This shows that all
methods based on domain knowledge technology can enhance the model's
understanding of specific fields to some extent, improve data interpretation capabilities,
and thus optimize prediction accuracy and reliability.
However, when the task involves many rare or long-tail scenarios (e.g.,
phenotyping estimation tasks), the method combining MLLM, SFT and RLHF
outperforms the method combining MLLM, RLHF and RAG and the method
combining MLLM, SFT and RAG. When the task involves obtaining knowledge from
many structured and domain-related data (e.g., the seed price query task), the method
combining MLLM, RLHF and RAG and the method combining MLLM, SFT and RAG
outperform the method combining MLLM, SFT and RLHF. This may be due to the
unique advantages of different domain knowledge technologies in dealing with specific
tasks. This result is consistent with previous studies (Balaguer et al., 2024; Giuffre et
al., 2024; J. Li et al., 2024), which showed that SFT directly optimizes model output by
annotating data, RAG integrates external multimodal knowledge to improve answer
quality, and RLHF iteratively optimizes generated content to improve task orientation
and adaptability. Therefore, the method combining MLLM, SFT, RAG, and RLHF
combines the advantages of different technologies to achieve better performance. It
should be noted that this method may be affected by the update frequency and accuracy
of the knowledge base, resulting in a decrease in the quality of answers (Siriwardhana
et al., 2023). In addition, RLHF requires a large amount of interaction data to optimize
the MLLM, and sample efficiency issues may occur in the process (Yuan et al., 2024).
Future study can integrate more advanced domain knowledge technologies (e.g.,
knowledge graphs) to extract better domain knowledge to achieve more accurate
generation prediction capabilities for different breeding tasks.
Fig. 6. Prediction performance of different tasks for methods with different
combinations of domain knowledge technologies.

4.2.2. Performance of cross-domain data combination in wheat yield prediction


When only RS data is used, the R2 of wheat yield prediction is 0.707 and the RMSE
is 599.532 kg/ha (Fig. 7 and Table S7). The RS data comes from processed data from
different UAV sensors. As shown in many previous studies, different information such
as VI (spectral), FVC (environment), CH (canopy structure) have become the most
commonly used RS indicators in crop yield prediction due to their stable and superior
performance (Roth et al., 2022; Skobalski et al., 2024; T. Xu et al., 2023). However, a
wide range of environmental conditions such as growing soil, water availability and
atmospheric conditions, as well as different germplasm genes of wheat, also have a
significant impact on yield, which may reduce model performance.
The combination of RS and phenotyping data can significantly improve the
prediction accuracy compared with using RS data alone. The canopy structure features
(LAI) and physiology and biochemistry features (SPAD) obtained from phenotyping
data contain independent information about canopy growth and structure, as well as
information that indirectly reflects the relative content of current chlorophyll in plant
leaves. Furthermore, phenotyping data have the advantage of higher accuracy to a
certain extent, and can directly capture more subtle changes and differences, which is
particularly important for accurately assessing wheat growth conditions and predicting
yields. Therefore, the combination of RS and phenotyping data can improve the
prediction accuracy. Many previous studies have verified the potential of coupled RS
and phenotyping information in crop yield prediction (Duan et al., 2017; Jin et al., 2024).
The combination of RS and weather data slightly improved the prediction accuracy
compared with RS data alone, with an increase of 0.014 in R2 and a decrease of 9.84
kg/ha in RMSE. The information from weather data tends to add supplementary
information to RS information in wheat yield prediction, but to a lesser extent than the
information contained in phenotyping data. It shows that RS information and weather
information at different growth stages are prone to information overlap, further
affecting the prediction ability. When the relevant information based on germplasm data
(high yield, HQ, DS, DS, MP and AM) was added to the information from RS data, the
prediction accuracy of all methods was slightly improved. It shows that using macro-
growth conditions provided by RS data and micro-genetic traits provided by germplasm
data can provide a more comprehensive assessment of crop growth potential and yield.
The yield prediction performance from the combination of RS, phenotyping, and
weather data; RS, phenotyping, and germplasm data; and RS, weather, and germplasm
data is better than the combination of the two data types, with R2 varying from 0.734 to
0.808 and RMSE ranging from 494.967 kg/ha to 572.755 kg/ha. These results indicate
that RS, phenotyping, weather, and germplasm information provide unique and
complementary information that is helpful for wheat yield prediction. These results at
least partially further explain the benefits of using cross-domain data in wheat yield
prediction. For this benefit, we further utilized all the data obtained (RS, phenotyping,
weather, germplasm), and the results showed that the yield prediction performance was
the best, with an R2 of 0.821 and an RMSE of 489.254 kg/ha. Related studies have also
shown that using one or more data from RS, environment and germplasm can improve
the performance of model yield prediction (Feng et al., 2024; Maimaitijiang et al., 2020;
Tian et al., 2022). These studies use data fusion as one of the data-level solutions to
improve the accuracy of yield prediction. It is undeniable that as the data scale increases,
the model can usually achieve improved performance, but this increases the complexity
of the model, requiring more powerful computing resources and more complex
algorithms (J. Li et al., 2024). And data quality and missing issues directly affect the
prediction accuracy of the model (Yang et al., 2023). In addition, obtaining high-quality
yield data requires a lot of field experiments and data processing, which increases the
cost of data acquisition (Ruan et al., 2022). Although data fusion can improve the
generalization ability of the model, the model may still experience performance
degradation when facing a completely new planting environment or germplasm (Gu et
al., 2024).
The predicted values of wheat yield obtained by the methods (trained using data
from different data sources) were compared with the corresponding observed values
using the scatter plot (Fig. 7). As expected from the comparable R2 and RMSE values,
the values obtained by the method trained only with RS data are widely scattered on
both sides of the regression line, while the values obtained by the method trained with
all multi-source data are closely around the regression line. It is worth noting that the
data from the China Rural Statistical Yearbook show that the yield per unit area of
wheat in Zhejiang Province in 2022 is 4230.2 kg/ha (National Bureau of Statistics of
China, 2022), where the points to the right of the purple dashed line in Fig. 7 represent
germplasms that exceed the yield per unit area. The germplasms with yields exceeding
4230.2 kg ha−1 are from China, Japan, Argentina, Italy, Turkey, and Australia. These
germplasms showed higher FVC, CH, CV, and WH compared to other germplasms. It
should be emphasized that these germplasms achieved higher yields in an environment
with stronger weed competition, thus indicating that these germplasms have stronger
competitiveness and adaptability. These wheat germplasms are more valuable in
production because their growth shows a certain degree of freedom from weed control
activities, reducing production costs and improving yield sustainability.
Fig. 7. Cross-validation scatter plot of measured and predicted wheat yield. The black
solid line indicates a 1:1 relationship. The blue points to the right of the purple dashed
line represent germplasms with wheat yields exceeding 4230.2 kg ha−1.

4.3. Breeding goals response to WBLM

We compare WBLM with the easily accessible MLLM: Qwen 2.5, ERNIE Bot 4
Turbo, GPT-4o Plus, and Gemini Pro 1.5 in different scenarios (the five tasks of this
study) to explore wheat breeding goals response capabilities of MLLM. Since our
WBLM and all other MLLMs support both Chinese and English, we chose Chinese and
English for testing. The results show that WBLM, Qwen and ChatGPT can correctly
and automatically adjust to the corresponding language to answer according to the
user's query language, while the other two MLLMs need to explicitly specify the answer
language. We aim to demonstrate the practicality and versatility of WBLM and other
MLLMs in real breeding applications, providing analysis and insights from the
perspective of actual user experience whenever possible.
We can use WBLM and commercial MLLMs to estimate the yield of different
germplasms (Fig. 8). Specifically, we provide different phenotyping images and related
growth information. Except for WBLM, other MLLMs cannot answer the questions
accurately, and WBLM's answers are simple and direct. Some MLLMs cannot answer
correctly, but ideas and methods are provided to further answer the question. The results
show that WBLM has unique and superior capabilities in yield phenotyping estimation
than other MLLMs.

Fig. 8. Examples of phenotyping (yield) estimation on different MLLMs.


We can use WBLM and commercial MLLMs to screen target germplasm. We test
five target germplasm characteristics: HQ, DS, DR, MP, and AM, either individually or
in combination. Regarding the screening of AM (plant height) (Fig. 9a), Qwen and
ERNIE Bot were unable to answer correctly, while other MLLMs were able to answer
correctly. For the HQ screening (Fig. 9b), Qwen and ERNIE Bot cannot correctly
answer the specific variety name and preservation organization, while other MLLMs
can answer correctly. For the MP and AM screening (Fig. 9c), we found that except for
Gemini, other MLLMs can correctly output specific wheat variety names that satisfy
the requirements. For DS, DR and HQ screening (Fig. 9d), each MLLM can correctly
output specific wheat varieties. Analyzing the above results, except for WBLM, which
has a clear data source, the data sources of other MLLMs are unknown. In addition, the
answers of different MLLMs vary greatly, so users need to verify and refer to them with
caution. In addition, if MLLM cannot answer the question correctly by self-judgment,
some solutions or methods will be provided.

Fig. 9. Examples of target germplasm screening on different MLLMs.


We can use WBLM and commercial MLLMs to assess the environmental stress
(FVC and WL) of different germplasms (Fig. 10). Only WBLM can accurately output
the value of FVC and the level of WL. ChatGPT can answer questions in the
environmental stress assessment, but the answers are inaccurate. ERNIE Bot and
Gemini can only partially answer and the answers are inaccurate. Qwen cannot answer
either. Furthermore, we found that when there was no word limit on the answer, WBLM
had a deeper understanding of the knowledge in the field of wheat breeding and the
answers were more focused and specific.

Fig. 10. Examples of environmental stress assessment on different MLLMs.


We can use WBLM and commercial MLLMs to recommend efficient, innovative
and sustainable PT and PP techniques (Fig. 11). Focus on their ability to accurately
answer questions under different techniques, different growth stages, and different
growth environments. The scope of the recommendation: From soil preparation
technique before wheat planting, to wheat sowing time, sowing amount and depth; from
wheat fertilization management (base fertilizer), to when to apply topdressing and what
fertilizer to use; from how to effectively control weeds in field management, to wheat
disease and pest control in specific growth periods; from introducing wheat soil testing
and formulation technique, to introducing wheat “one spray and three preventions”
technique. The results show that the performance of WBLM is on par with Qwen and
ChatGPT, and slightly better than the answers of other MLLMs. This reflects WBLM's
superior understanding and reasoning of PT and PP technique, resulting in better
recommendation results than the answers of other MLLMs.
Fig. 11. Examples of cultivation technique recommendation on different MLLMs.
We can use WBLM and commercial MLLMs to query the seed varieties,
specifications and prices at a specified observation point and time (Fig. 12). It can be
seen that except for WBLM, none of them can provide accurate answers. Other MLLMs
choose to give factors that affect wheat seed prices and suggest methods to obtain
relevant data. Through this experiment, we found that for highly dynamic information
such as price data, MLLMs need to rely on external data sources to ensure the real-time
and accuracy of the data. Although MLLMs can provide trend analysis, principle
explanations, or forecasting methodologies based on historical data, they are unable to
capture new market conditions, stock prices, commodity prices, or other information
that requires real-time updates. Real-time and price data are usually collected and
published by specialized financial data providers, market analysis companies or
government statistics departments. These data often need to be obtained through
subscription services, API interfaces or public market reports.

Fig. 12. Examples of seed price query on different MLLMs.

5. Conclusion

In this study, we innovatively applied MLLM to wheat breeding and explored its
potential for application in wheat breeding objectives. Based on different pre-trained
open-source MLLMs, we used SFT, RAG, and RLHF technologies to inject cross-
domain knowledge into MLLMs, thereby constructing multiple WBLMs. Among them,
the WBLM constructed using SFT, RAG and RLHF technologies and InternVL2-8B
has leading performance on the evaluation benchmark newly created in this study. The
WBLM has competitive advantages over leading proprietary MLLMs, especially in
wheat breeding tasks. Then, subsequent experiments were conducted using the WBLM.
Ablation experiments indicated that the combination of SFT, RAG, and RLHF
technologies can improve the overall generation performance, enhance the generated
quality, balance the timeliness and adaptability of the generated answer, and reduce
hallucinations and biases. The WBLM performed best in wheat yield prediction using
cross-domain data (RS, phenotyping, weather, germplasm) simultaneously, with R2 and
RMSE of 0.821 and 489.254 kg/ha, respectively. Furthermore, the performance of
WBLM and other MLLMs for wheat breeding goals was compared from a qualitative
perspective of answer generation, and it was found that WBLM could generate more
professional decision support answers. This study aims to provide an intelligent and
integrated solution for wheat breeding goals, to help breeding work be carried out
efficiently, to accelerate the breeding process of excellent varieties, and to provide
scientific basis and technical support. In future, we will continue to explore deeper data
fusion technology, optimize the generalization ability of the model, and enhance the
interpretability of the model to ensure the transparency and reliability of breeding
decisions. We will also expand the scope of application of the model to cover more crop
varieties and environmental conditions to meet the breeding needs of different regions.

CRediT authorship contribution statement

Guofeng Yang: Methodology, Data curation, Formal analysis, Validation,


Visualization, Writing – original draft. Yu Li: Writing – review & editing, Project
administration. Yong He: Conceptualization, Project administration. Zhenjiang Zhou:
Formal analysis, Writing – review & editing. Lingzhen Ye: Data curation, Project
administration. Hui Fang: Data curation, Writing – review & editing. Yiqi Luo:
Writing – review & editing. Xuping Feng: Conceptualization, Supervision, Funding
acquisition, Writing - review & editing.
Declaration of Competing Interest

The authors declare that they have no known competing financial interests or
personal relationships that could have appeared to influence the work reported in this
paper.

Acknowledgments

Thanks to the Changxing Agricultural Experiment Station of Zhejiang University


and the Yuhang Agricultural Experiment Station of Zhejiang University for providing
the experimental base. Thanks to Professor Xianchun Xia and Dr. Jindong Liu (Chinese
Academy of Agricultural Sciences) for providing wheat germplasm resources. Thanks
to the members of the Digital Phenotyping Research Group of Zhejiang University for
their help in data collection. This research was supported by Zhejiang Provincial Key
R&D Program of China (Grant No. 2022C02013).

Appendix A. Supplementary data

Supplementary material

Data Availability

The dataset for the current study is available in the Zenodo repository
at [Link]

References

AI, 01, Young, A., Chen, B., Li, C., Huang, C., Zhang, Ge, Zhang, Guanwei, Li, H., Zhu, J., Chen,
J., Chang, J., Yu, K., Liu, P., Liu, Q., Yue, S., Yang, Senbin, Yang, Shiming, Yu, T., Xie, W.,
Huang, W., Hu, X., Ren, X., Niu, X., Nie, P., Xu, Y., Liu, Y., Wang, Y., Cai, Y., Gu, Z., Liu,
Z., Dai, Z., 2024. Yi: Open Foundation Models by [Link].
[Link]
Anderegg, J., Tschurr, F., Kirchgessner, N., Treier, S., Schmucki, M., Streit, B., Walter, A., 2023.
On-farm evaluation of UAV-based aerial imagery for season-long weed monitoring under
contrasting management and pedoclimatic conditions in wheat. Comput. Electron. Agric.
204. [Link]
Avola, D., Cinque, L., Fagioli, A., Foresti, G.L., Pannone, D., Piciarelli, C., 2021. Automatic
estimation of optimal UAV flight parameters for real-time wide areas monitoring. Multimed.
TOOLS Appl. 80, 25009–25031. [Link]
Badgley, G., Field, C.B., Berry, J.A., 2017. Canopy near-infrared reflectance and terrestrial
photosynthesis. Sci. Adv. 3. [Link]
Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., Zhou, J., 2023. Qwen-VL:
A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and
Beyond. [Link]
Balaguer, A., Benara, V., Cunha, R.L. de F., Filho, R. de M.E., Hendry, T., Holstein, D., Marsman,
J., Mecklenburg, N., Malvar, S., Nunes, L.O., Padilha, R., Sharp, M., Silva, B., Sharma, S.,
Aski, V., Chandra, R., 2024. RAG vs Fine-tuning: Pipelines, Tradeoffs, and a Case Study
on Agriculture. [Link]
Bhat, J.A., Feng, X., Mir, Z.A.A., Raina, A., Siddique, K.H.M., 2023. Recent advances in artificial
intelligence, mechanistic models, and speed breeding offer exciting opportunities for
precise and accelerated genomics-assisted breeding. Physiol. Plant. 175.
[Link]
Cai, Z., Cao, M., Chen, H., Chen, Kai, Chen, Keyu, Chen, Xin, Chen, Xun, Chen, Zehui, Chen, Zhi,
Chu, P., Dong, X., Duan, H., Fan, Q., Fei, Z., Gao, Y., Ge, J., Gu, C., Gu, Y., Gui, T., Guo,
A., Guo, Q., He, C., Hu, Y., Huang, T., Jiang, T., Jiao, P., Jin, Z., Lei, Z., Li, Jiaxing, Li,
Jingwen, Li, L., Li, S., Li, W., Li, Y., Liu, H., Liu, J., Hong, J., Liu, Kaiwen, Liu, Kuikun,
Liu, X., Lv, C., Lv, H., Lv, K., Ma, L., Ma, R., Ma, Z., Ning, W., Ouyang, L., Qiu, J., Qu,
Y., Shang, F., Shao, Y., Song, D., Song, Z., Sui, Z., Sun, P., Sun, Y., Tang, H., Wang, B.,
Wang, G., Wang, Jiaqi, Wang, Jiayu, Wang, R., Wang, Y., Wang, Z., Wei, X., Weng, Q., Wu,
F., Xiong, Y., Xu, C., Xu, R., Yan, H., Yan, Y., Yang, X., Ye, H., Ying, H., Yu, Jia, Yu, Jing,
Zang, Y., Zhang, C., Zhang, L., Zhang, Pan, Zhang, Peng, Zhang, R., Zhang, Shuo, Zhang,
Songyang, Zhang, Wenjian, Zhang, Wenwei, Zhang, Xingcheng, Zhang, Xinyue, Zhao, H.,
Zhao, Q., Zhao, X., Zhou, F., Zhou, Z., Zhuo, J., Zou, Y., Qiu, X., Qiao, Y., Lin, D., 2024.
InternLM2 Technical Report. [Link]
Camps-Valls, G., Campos-Taberner, M., Moreno-Martinez, A., Walther, S., Duveiller, G., Cescatti,
A., Mahecha, M.D., Munoz-Mari, J., Javier Garcia-Haro, F., Guanter, L., Jung, M., Gamon,
J.A., Reichstein, M., Running, S.W., 2021. A unified vegetation index for quantifying the
terrestrial biosphere. Sci. Adv. 7. [Link]
Cao, Z., Yao, X., Liu, H., Liu, B., Cheng, T., Tian, Y., Cao, W., Zhu, Y., 2019. Comparison of the
abilities of vegetation indices and photosynthetic parameters to detect heat stress in wheat.
Agric. For. Meteorol. 265, 121–136. [Link]
Chen, J., Xiao, S., Zhang, P., Luo, K., Lian, D., Liu, Z., 2024. BGE M3-Embedding: Multi-Lingual,
Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge
Distillation. [Link]
Chen, Z., Jia, K., Wei, X., Liu, Y., Zhan, Y., Xia, M., Yao, Y., Zhang, X., 2022. Improving leaf area
index estimation accuracy of wheat by involving leaf chlorophyll content information.
Comput. Electron. Agric. 196. [Link]
Chen, Z., Wang, Weiyun, Tian, H., Ye, S., Gao, Z., Cui, E., Tong, W., Hu, K., Luo, J., Ma, Z., Ma,
J., Wang, J., Dong, X., Yan, H., Guo, H., He, C., Shi, B., Jin, Z., Xu, C., Wang, B., Wei, X.,
Li, W., Zhang, W., Zhang, B., Cai, P., Wen, L., Yan, X., Dou, M., Lu, L., Zhu, X., Lu, T.,
Lin, D., Qiao, Y., Dai, J., Wang, Wenhai, 2024. How Far Are We to GPT-4V? Closing the
Gap to Commercial Multimodal Models with Open-Source Suites.
[Link]
Cheng, S., Feng, C., Wingen, L.U., Cheng, H., Riche, A.B., Jiang, M., Leverington-Waite, M.,
Huang, Z., Collier, S., Orford, S., Wang, Xiaoming, Awal, R., Barker, G., O’Hara, T., Lister,
C., Siluveru, A., Quiroz-Chavez, J., Ramirez-Gonzalez, R.H., Bryant, R., Berry, S., Bansal,
U., Bariana, H.S., Bennett, M.J., Bicego, B., Bilham, L., Brown, J.K.M., Burridge, A., Burt,
C., Buurman, M., Castle, M., Chartrain, L., Chen, B., Denbel, W., Elkot, A.F., Fenwick, P.,
Feuerhelm, D., Foulkes, J., Gaju, O., Gauley, A., Gaurav, K., Hafeez, A.N., Han, R., Horler,
R., Hou, J., Iqbal, M.S., Kerton, M., Kondic-Spica, A., Kowalski, A., Lage, J., Li, X., Liu,
H., Liu, S., Lovegrove, A., Ma, L., Mumford, C., Parmar, S., Philp, C., Playford, D.,
Przewieslik-Allen, A.M., Sarfraz, Z., Schafer, D., Shewry, P.R., Shi, Y., Slafer, G., Song,
Baoxing, Song, Bo, Steele, D., Steuernagel, B., Tailby, P., Tyrrell, S., Waheed, A., Wamalwa,
M.N., Wang, Xingwei, Wei, Y., Winfield, M., Wu, S., Wu, Y., Wulff, B.B.H., Xian, W., Xu,
Yawen, Xu, Yunfeng, Yuan, Q., Zhang, X., Edwards, K.J., Dixon, L., Nicholson, P., Chayut,
N., Hawkesford, M.J., Uauy, C., Sanders, D., Huang, S., Griffiths, S., 2024. Harnessing
landrace diversity empowers wheat breeding. Nature. [Link]
07682-9
Cui, S., Chen, W., Gu, W., Yang, L., Shi, X., 2023. SiamC Transformer: Siamese coupling swin
transformer Multi-Scale semantic segmentation network for vegetation extraction under
shadow conditions. Comput. Electron. Agric. 213.
[Link]
Das, S., Christopher, J., Apan, A., Choudhury, M.R., Chapman, S., Menzies, N.W., Dang, Y.P., 2021.
Evaluation of water status of wheat genotypes to aid prediction of yield on sodic soils using
UAV-thermal imaging and machine learning. Agric. For. Meteorol. 307.
[Link]
DeepSeek-AI, Liu, A., Feng, B., Wang, Bin, Wang, Bingxuan, Liu, B., Zhao, C., Dengr, C., Ruan,
C., Dai, D., Guo, D., Yang, D., Chen, D., Ji, D., Li, E., Lin, F., Luo, F., Hao, G., Chen, G.,
Li, G., Zhang, H., Xu, H., Yang, H., Zhang, Haowei, Ding, H., Xin, H., Gao, H., Li, H., Qu,
H., Cai, J.L., Liang, J., Guo, J., Ni, J., Li, J., Chen, J., Yuan, J., Qiu, J., Song, J., Dong, K.,
Gao, K., Guan, K., Wang, L., Zhang, Lecong, Xu, L., Xia, L., Zhao, L., Zhang, Liyue, Li,
Meng, Wang, M., Zhang, Mingchuan, Zhang, Minghua, Tang, M., Li, Mingming, Tian, N.,
Huang, P., Wang, P., Zhang, P., Zhu, Q., Chen, Q., Du, Q., Chen, R.J., Jin, R.L., Ge, R., Pan,
R., Xu, R., Chen, R., Li, S.S., Lu, S., Zhou, Shangyan, Chen, S., Wu, S., Ye, S., Ma, S.,
Wang, S., Zhou, Shuang, Yu, S., Zhou, Shunfeng, Zheng, S., Wang, T., Pei, T., Yuan, T.,
Sun, T., Xiao, W.L., Zeng, W., An, W., Liu, W., Liang, W., Gao, W., Zhang, W., Li, X.Q.,
Jin, X., Wang, Xianzu, Bi, X., Liu, Xiaodong, Wang, Xiaohan, Shen, X., Chen, Xiaokang,
Chen, Xiaosha, Nie, X., Sun, X., Wang, Xiaoxiang, Liu, Xin, Xie, X., Yu, X., Song, X.,
Zhou, X., Yang, X., Lu, X., Su, X., Wu, Y., Li, Y.K., Wei, Y.X., Zhu, Y.X., Xu, Y., Huang,
Y., Li, Yao, Zhao, Yao, Sun, Y., Li, Yaohui, Wang, Yaohui, Zheng, Y., Zhang, Y., Xiong, Y.,
Zhao, Yilong, He, Y., Tang, Y., Piao, Y., Dong, Y., Tan, Y., Liu, Yiyuan, Wang, Yongji, Guo,
Y., Zhu, Y., Wang, Yuduan, Zou, Y., Zha, Y., Ma, Y., Yan, Y., You, Y., Liu, Yuxuan, Ren,
Z.Z., Ren, Z., Sha, Z., Fu, Z., Huang, Z., Zhang, Zhen, Xie, Zhenda, Hao, Z., Shao, Z., Wen,
Z., Xu, Z., Zhang, Zhongyu, Li, Zhuoshu, Wang, Z., Gu, Z., Li, Zilin, Xie, Ziwei, 2024.
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.
[Link]
Du, R., Chen, J., Xiang, Y., Zhang, Z., Yang, N., Yang, X., Tang, Z., Wang, H., Wang, X., Shi, H.,
Li, W., 2023. Incremental learning for crop growth parameters estimation and nitrogen
diagnosis from hyperspectral data. Comput. Electron. Agric. 215.
[Link]
Duan, T., Chapman, S.C., Guo, Y., Zheng, B., 2017. Dynamic monitoring of NDVI in wheat
agronomy and breeding trials using an unmanned aerial vehicle. FIELD CROPS Res.
[Link]
Fei, S., Hassan, M.A., Xiao, Y., Su, X., Chen, Z., Cheng, Q., Duan, F., Chen, R., Ma, Y., 2023. UAV-
based multi-sensor data fusion and machine learning algorithm for yield prediction in wheat.
Precis. Agric. 24, 187–212. [Link]
Feng, A., Zhou, J., Vories, E., Sudduth, K., 2024. Prediction of cotton yield based on soil texture,
weather conditions and UAV imagery using deep learning. Precis. Agric. 25, 303–326.
[Link]
Feng, Z., Guan, H., Yang, T., He, L., Duan, J., Song, L., Wang, C., Feng, W., 2023. Estimating the
canopy chlorophyll content of winter wheat under nitrogen deficiency and powdery mildew
stress using machine learning. Comput. Electron. Agric. 211.
[Link]
Fu, C., Chen, P., Shen, Y., Qin, Y., Zhang, M., Lin, X., Yang, J., Zheng, X., Li, K., Sun, X., Wu, Y.,
Ji, R., 2024. MME: A Comprehensive Evaluation Benchmark for Multimodal Large
Language Models. [Link]
Garg, M., Kaur, S., Sharma, A., Kumari, A., Tiwari, V., Sharma, S., Kapoor, P., Sheoran, B., Goyal,
A., Krishania, M., 2022. Rising Demand for Healthy Foods-Anthocyanin Biofortified
Colored Wheat Is a New Research Trend. Front. Nutr. 9.
[Link]
Giuffre, M., Kresevic, S., Pugliese, N., You, K., Shung, D.L., 2024. Optimizing large language
models in digestive disease: strategies and challenges to improve clinical outcomes. LIVER
Int. [Link]
GLM, T., Zeng, A., Xu, B., Wang, B., Zhang, C., Yin, D., Zhang, D., Rojas, D., Feng, G., Zhao, H.,
Lai, H., Yu, H., Wang, H., Sun, Jiadai, Zhang, Jiajie, Cheng, J., Gui, J., Tang, J., Zhang,
Jing, Sun, Jingyu, Li, J., Zhao, L., Wu, L., Zhong, L., Liu, M., Huang, M., Zhang, P., Zheng,
Q., Lu, R., Duan, S., Zhang, S., Cao, S., Yang, S., Tam, W.L., Zhao, W., Liu, Xiao, Xia, X.,
Zhang, Xiaohan, Gu, X., Lv, X., Liu, Xinghan, Liu, Xinyi, Yang, X., Song, X., Zhang,
Xunkai, An, Y., Xu, Y., Niu, Y., Yang, Y., Li, Y., Bai, Y., Dong, Y., Qi, Z., Wang, Zhaoyu,
Yang, Z., Du, Z., Hou, Z., Wang, Zihan, 2024. ChatGLM: A Family of Large Language
Models from GLM-130B to GLM-4 All Tools. [Link]
Gu, Y., Wang, Y., Wu, Y., Warner, T., Guo, T., Ai, H., Zheng, H., Cheng, T., Zhu, Y., Cao, W., Yao,
X., 2024. Novel 3D photosynthetic traits derived from the fusion of UAV LiDAR point
cloud and multispectral imagery in wheat. REMOTE Sens. Environ. 311.
[Link]
Hu, N., Du, C., Zhang, W., Liu, Y., Zhang, Y., Zhao, Z., Wang, Z., 2022. Did Wheat Breeding
Simultaneously Improve Grain Yield and Quality of Wheat Cultivars Releasing over the
Past 20 Years in China? Agron.-BASEL 12. [Link]
Huang, D., Yan, C., Li, Q., Peng, X., 2024. From Large Language Models to Large Multimodal
Models: A Literature Review. Appl. Sci.-BASEL 14. [Link]
Huang, X., Ruan, W., Huang, W., Jin, G., Dong, Y., Wu, C., Bensalem, S., Mu, R., Qi, Y., Zhao, X.,
Cai, K., Zhang, Y., Wu, S., Xu, P., Wu, D., Freitas, A., Mustafa, M.A., 2024. A survey of
safety and trustworthiness of large language models through the lens of verification and
validation. Artif. IN℡LIGENCE Rev. [Link]
Huete, A.R., 1988. A soil-adjusted vegetation index (SAVI). Remote Sens. Environ. 25, 295–309.
[Link]
Jiang, Z., Tu, H., Bai, B., Yang, C., Zhao, B., Guo, Z., Liu, Q., Zhao, H., Yang, W., Xiong, L., Zhang,
J., 2021. Combining UAV-RGB high-throughput field phenotyping and genome-wide
association study to reveal genetic variation of rice germplasms in dynamic response to
drought stress. NEW Phytol. 232, 440–455. [Link]
Jin, Z., Guo, S., Li, S., Yu, F., Xu, T., 2024. Research on the rice fertiliser decision-making method
based on UAV remote sensing data assimilation. Comput. Electron. Agric. 216.
[Link]
Kaur, B., Sandhu, K.S., Kamal, R., Kaur, K., Singh, J., Roeder, M.S., Muqaddasi, Q.H., 2021. Omics
for the Improvement of Abiotic, Biotic, and Agronomic Traits in Major Cereal Crops:
Applications, Challenges, and Prospects. PLANTS-BASEL 10.
[Link]
Kuska, M.T., Wahabzada, M., Paulus, S., 2024. AI for crop production - Where can large language
models (LLMs) provide substantial value? Comput. Electron. Agric. 221.
[Link]
Li, C., Gan, Z., Yang, Z., Yang, J., Li, L., Wang, L., Gao, J., 2024. Multimodal Foundation Models:
From Specialists to General-Purpose Assistants. Found. TRENDS Comput. Graph. Vis. 16,
1–214. [Link]
Li, F., Zhang, R., Zhang, H., Zhang, Y., Li, B., Li, W., Ma, Z., Li, C., 2024. LLaVA-NeXT-Interleave:
Tackling Multi-image, Video, and 3D in Large Multimodal Models.
[Link]
Li, H., Li, X., Zhang, P., Feng, Y., Mi, J., Gao, S., Sheng, L., Ali, M., Yang, Z., Li, L., Fang, W.,
Wang, W., Qian, Q., Gu, F., Zhou, W., 2024. Smart Breeding Platform: A web-based tool
for high-throughput population genetics, phenomics, and genomic selection. Mol. PLANT
17, 677–681. [Link]
Li, J., Xu, M., Xiang, L., Chen, D., Zhuang, W., Yin, X., Li, Z., 2024. Foundation models in smart
agriculture: Basics, opportunities, and challenges. Comput. Electron. Agric. 222.
[Link]
Lin, Z., Guan, S., Zhang, W., Zhang, Huiyan, Li, Y., Zhang, Huaping, 2024. Towards trustworthy
LLMs: a review on debiasing and dehallucinating in large language models. Artif.
IN℡LIGENCE Rev. [Link]
Liu, L., Zhan, J., Yan, J., 2024. Engineering the future cereal crops with big biological data:
towardan intelligence-driven breeding by design. J. Genet. Genomics Yi Chuan Xue Bao.
[Link]
Liu, Y., Li, Z., Yang, B., Li, C., Yin, X., Liu, C., Jin, L., Bai, X., 2024. On the Hidden Mystery of
OCR in Large Multimodal Models. [Link]
Maimaitijiang, M., Sagan, V., Sidike, P., Hartling, S., Esposito, F., Fritschi, F.B., 2020. Soybean
yield prediction from UAV using multimodal data fusion and deep learning. REMOTE Sens.
Environ. [Link]
Min, B., Ross, H., Sulem, E., Ben Veyseh, A.P., Nguyen, T.H., Sainz, O., Agirre, E., Heintz, I., Roth,
D., 2024. Recent Advances in Natural Language Processing via Large Pre-trained
Language Models: A Survey. ACM Comput. Surv. 56. [Link]
Ming, Y., Li, Y., 2024. How Does Fine-Tuning Impact Out-of-Distribution Detection for Vision-
Language Models? Int. J. Comput. Vis. [Link]
National Bureau of Statistics of China, 2022. China Rural Statistical Yearbook: 2022. China
Statistics Press, Beijing.
Padhy, A.K., Kaur, P., Singh, S., Kashyap, L., Sharma, A., 2024. Colored wheat and derived products:
key to global nutritional security. Crit. Rev. FOOD Sci. Nutr. 64, 1894–1910.
[Link]
Richardson, A.J., Everitt, J.H., 1992. Using spectral vegetation indices to estimate rangeland
productivity. Geocarto Int. 7, 63–69. [Link]
Roth, L., Barendregt, C., Betrix, C.-A., Hund, A., Walter, A., 2022. High-throughput field
phenotyping of soybean: Spotting an ideotype. REMOTE Sens. Environ.
[Link]
Ruan, G., Li, X., Yuan, F., Cammarano, D., Ata-UI-Karim, S., Liu, X., Tian, Y., Zhu, Y., Cao, W.,
Cao, Q., 2022. Improving wheat yield prediction integrating proximal sensing and weather
data with machine learning. Comput. Electron. Agric. 195.
[Link]
Sansaloni, C., Franco, J., Santos, B., Percival-Alwyn, L., Singh, S., Petroli, C., Campos, J., Dreher,
K., Payne, T., Marshall, D., Kilian, B., Milne, I., Raubach, S., Shaw, P., Stephen, G., Carling,
J., Saint Pierre, C., Burgueno, J., Crosa, J., Li, H., Guzman, C., Kehel, Z., Amri, A., Kilian,
A., Wenzl, P., Uauy, C., Banziger, M., Caccamo, M., Pixley, K., 2020. Diversity analysis of
80,000 wheat accessions reveals consequences and opportunities of selection footprints.
Nat. Commun. 11. [Link]
Saskatchewan Seed Growers’ Association, 2024. SaskSeed Guides (2024).
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., Klimov, O., 2017. Proximal Policy Optimization
Algorithms. [Link]
Senapati, N., Semenov, M.A., Halford, N.G., Hawkesford, M.J., Asseng, S., Cooper, M., Ewert, F.,
van Ittersum, M.K., Martre, P., Olesen, J.E., Reynolds, M., Roetter, R.P., Webber, H., 2022.
Global wheat production could benefit from closing the genetic yield gap. Nat. FOOD 3,
532–541. [Link]
Siriwardhana, S., Weerasekera, R., Wen, E., Kaluarachchi, T., Rana, R., Nanayakkara, S., 2023.
Improving the Domain Adaptation of Retrieval Augmented Generation (RAG) Models for
Open Domain Question Answering. Trans. Assoc. Comput. Linguist.
[Link]
Skobalski, J., Sagan, V., Alifu, H., Al Akkad, O., Lopes, F.A., Grignola, F., 2024. Bridging the gap
between crop breeding and GeoAI: Soybean yield prediction from multispectral UAV
images with transfer learning. ISPRS J. Photogramm. REMOTE Sens.
[Link]
State Administration for Market Regulation, Standardization Administration of the P.R.C, 2023.
Wheat.
The ModelScope Team, 2024. SWIFT:Scalable lightWeight Infrastructure for Fine-Tuning [WWW
Document]. URL [Link] (accessed 8.2.24).
Tian, Z., Zhang, Y., Liu, K., Li, Z., Li, M., Zhang, H., Wu, J., 2022. UAV Remote Sensing Prediction
Method of Winter Wheat Yield Based on the Fused Features of Crop and Soil. REMOTE
Sens. 14. [Link]
Tucker, C.J., 1979. Red and photographic infrared linear combinations for monitoring vegetation.
Remote Sens. Environ. 8, 127–150. [Link]
Wang, X., Chen, G., Qian, G., Gao, P., Wei, X.-Y., Wang, Y., Tian, Y., Gao, W., 2023. Large-scale
Multi-modal Pre-trained Models: A Comprehensive Survey. Mach. Intell. Res. 20, 447–482.
[Link]
Xiong, W., Reynolds, M.P., Crossa, J., Schulthess, U., Sonder, K., Montes, C., Addimando, N.,
Singh, R.P., Ammar, K., Gerard, B., Payne, T., 2021. Increased ranking change in wheat
breeding under climate change. Nat. PLANTS. [Link]
w
Xu, J., Zhu, C., Su, M., Li, S., Chao, H., Chen, M., 2023. CropGF: a comprehensive visual platform
for crop gene family mining and analysis. DATABASE- J. Biol. DATABASES CURATION
2023. [Link]
Xu, T., Wang, F., Shi, Z., Xie, L., Yao, X., 2023. Dynamic estimation of rice aboveground biomass
based on spectral and spatial information extracted from hyperspectral remote sensing
images at different combinations of growth stages. ISPRS J. Photogramm. REMOTE Sens.
[Link]
Xu, Y., Zhang, X., Li, H., Zheng, H., Zhang, J., Olsen, M.S., Varshney, R.K., Prasanna, B.M., Qian,
Q., 2022. Smart breeding driven by big data, artificial intelligence, and integrated genomic-
enviromic prediction. Mol. PLANT 15, 1664–1695.
[Link]
Yang, G., He, Y., Zhou, Z., Huang, L., Li, X., Yu, Z., Yang, Y., Li, Y., Ye, L., Feng, X., 2022. Field
Monitoring of Fractional Vegetation Cover Based on UAV Low-altitude Remote Sensing
and Machine Learning, in: 2022 10th International Conference on Agro-Geoinformatics
(Agro-Geoinformatics). pp. 1–6. [Link]
Geoinformatics55649.2022.9859063
Yang, G., Li, X., Liu, P., Yao, X., Zhu, Y., Cao, W., Cheng, T., 2023. Automated in-season mapping
of winter wheat in China with training data generation and model transfer. ISPRS J.
Photogramm. REMOTE Sens. [Link]
Yao, E., Blake, V.C., Cooper, L., Wight, C.P., Michel, S., Cagirici, H.B., Lazo, G.R., Birkett, C.L.,
Waring, D.J., Jannink, J.-L., Holmes, I., Waters, A.J., Eickholt, D.P., Sen, T.Z., 2022.
GrainGenes: a data-rich repository for small grains genetics and genomics. DATABASE-
J. Biol. DATABASES CURATION 2022. [Link]
Yuan, Y., Hao, J., Ma, Y., Dong, Z., Liang, H., Liu, J., Feng, Z., Zhao, K., Zheng, Y., 2024. Uni-
RLHF: Universal Platform and Benchmark Suite for Reinforcement Learning with Diverse
Human Feedback. Presented at the The Twelfth International Conference on Learning
Representations.
Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., Sun, Y.,
Wei, C., Yu, B., Yuan, R., Sun, R., Yin, M., Zheng, B., Yang, Z., Liu, Y., Huang, W., Sun,
H., Su, Y., Chen, W., 2024. MMMU: A Massive Multi-discipline Multimodal
Understanding and Reasoning Benchmark for Expert AGI.
[Link]
Zhang, B., Gu, L., Dai, M., Bao, X., Sun, Q., Zhang, M., Qu, X., Li, Z., Zhen, W., Gu, X., 2024.
Estimation of grain filling rate of winter wheat using leaf chlorophyll and LAI extracted
from UAV images. FIELD CROPS Res. 306. [Link]
Zhang, G., Yan, H., Zhang, D., Zhang, H., Cheng, T., Hu, G., Shen, S., Xu, H., 2023. Enhancing
model performance in detecting lodging areas in wheat fields using UAV RGB Imagery:
Considering spatial and temporal variations. Comput. Electron. Agric.
[Link]
Zhu, J., Yang, G., Feng, X., Li, X., Fang, H., Zhang, J., Bai, X., Tao, M., He, Y., 2022. Detecting
Wheat Heads from UAV Low-Altitude Remote Sensing Images Using Deep Learning
Based on Transformer. REMOTE Sens. 14. [Link]
Zhu, W., Han, R., Shang, X., Zhou, T., Liang, C., Qin, X., Chen, H., Feng, Z., Zhang, H., Fan, X.,
Li, W., Li, L., 2024. The CropGPT project: Call for a global, coordinated effort in precision
design breeding driven by AI using biological big data. Mol. PLANT 17, 215–218.
[Link]

Common questions

Powered by AI

Integrating RS data with phenotyping data significantly enhances the accuracy of wheat yield predictions compared to using RS data alone. While RS data provides valuable spectral, environmental, and structural indicators such as Vegetation Index (VI), Fractional Vegetation Cover (FVC), and Canopy Height (CH), the addition of phenotyping data introduces crucial physiological and biochemical traits like Leaf Area Index (LAI), which reflect the actual biological performance of crops. This combined approach leverages the strengths of each data type, leading to improved prediction precision and reduced errors, thus optimizing yield forecasting and supporting effective breeding programs .

The combination of MLLM with supervised fine-tuning (SFT), retrieval augmented generation (RAG), and reinforcement learning from human feedback (RLHF) is particularly effective in addressing rare or long-tail scenarios in phenotyping estimation tasks because each method contributes a unique advantage. SFT provides specificity by using annotated data to refine model outputs. RAG offers enhanced quality through external knowledge integration, which helps with infrequent data occurrences. RLHF promotes adaptability by continuously refining task-specific responses through expert-driven feedback loops. These complementary strengths effectively manage the complexities and uncertainties inherent in long-tail phenotyping scenarios .

To embed professional knowledge into MLLM, the document suggests several mechanisms including supervised fine-tuning, retrieval augmented generation, and reinforcement learning from human feedback. These techniques can help align the model outputs with actual breeding needs by integrating domain-specific annotations, adding external structured knowledge, and iteratively refining model responses through expert feedback. Additionally, collaborating with agricultural experts to annotate data and define breeding algorithms can enable models to discover novel breeding strategies and optimize decision-making processes. This alignment with domain expertise is crucial for advancing realistic applications in wheat breeding .

Cross-modal information integration plays a critical role in enhancing the understanding of crop trait formation and the breeding process by fusing diverse data types such as remote sensing, phenotypic, and environmental data. This integration allows MLLM to develop a more comprehensive understanding of crop traits by considering various influencing factors. It improves the model's ability to predict crop performance under different environmental conditions, thus aiding in efficient screening and breeding of new crop varieties by revealing the complex mechanisms behind trait formation. This holistic approach accelerates and optimizes decision-making within smart breeding frameworks .

Releasing datasets created during MLLM research for wheat breeding to the broader scientific community offers several potential benefits. It promotes collaboration and innovation, as researchers can build upon existing data to develop new models and approaches. Shared datasets enhance transparency and reproducibility of research findings, which is vital for scientific validation. They also allow comparative analyses across different modeling techniques, contributing to the discovery of best practices in crop breeding. Moreover, access to comprehensive datasets may reduce redundant data collection efforts and foster interdisciplinary advancements by integrating knowledge from various fields, ultimately accelerating progress in smart agriculture .

Each domain knowledge technology offers distinct benefits when used with MLLM in crop breeding tasks. Supervised fine-tuning (SFT) directly optimizes model outputs by specifically annotating data, making it advantageous for tasks requiring precise data interpretation. Retrieval augmented generation (RAG) integrates external multimodal knowledge to enhance answer quality, which is particularly useful for structured query tasks like seed price predictions. Reinforcement learning from human feedback (RLHF) helps iteratively optimize content generated by the model, enhancing task orientation and adaptability. These technologies, when combined, leverage their strengths to yield better prediction and classification performances in challenging tasks such as germplasm screening and phenotyping under various environmental and data-related constraints .

The deployment of UAVs improves the evaluation process of crop yield and stress by providing high-resolution imagery and multispectral data that capture comprehensive views of crop fields. This enables precise measurement of vegetation indexes such as canopy structure and chlorophyll content. UAVs also facilitate frequent data collection across different growth stages and environmental conditions, offering dynamic insights into biomass distribution and stress factors impacting crop health. Such data supports more accurate phenotyping and stress assessments, leading to better-informed breeding decisions and enhanced crop yield predictions due to the UAVs' ability to cover large areas efficiently while reducing labor and time costs .

The main challenges in integrating cross-domain data for developing MLLM in crop breeding include dealing with data that vary significantly in format, dimensionality, and sparsity from different sources such as RS data, phenotyping data, and environmental data during the crop growth period. Efficient and accurate integration of this heterogeneous data is difficult. Additionally, there is a need for models to have strong learning and generalization capabilities to predict crop performance under varying environmental conditions, which demands high expertise in model algorithm design and computing resources. Furthermore, incorporating professional knowledge and experience in crop breeding to make the models align with real breeding needs and discovering new breeding strategies are ongoing challenges .

Domain knowledge technology reduces the likelihood of hallucination in MLLM predictions by grounding the model's processes in well-defined, expert-validated data and providing structured frameworks for response generation. Supervised fine-tuning utilizes curated datasets specific to breeding tasks, ensuring that output is directly relevant. Retrieval augmented generation allows models to access and incorporate reliable external databases, thereby anchoring responses in factual knowledge. Reinforcement learning from human feedback iteratively corrects errors and aligns model behavior with domain expectations. These technologies collectively create a more robust framework that minimizes the chance of spurious or irrelevant outputs by reinforcing the model's understanding and narrowing the scope for speculative responses .

Combining domain knowledge technologies such as supervised fine-tuning (SFT), retrieval augmented generation (RAG), and reinforcement learning from human feedback (RLHF) with MLLM significantly improves its application to wheat breeding tasks. For example, without these technologies, MLLM shows poor performance in phenotyping estimation tasks due to a lack of domain knowledge. However, the integration of SFT, RAG, and RLHF enhances the prediction performance of subtasks like Yield, SPAD, LAI, CH, and CV by improving R2 and reducing RMSE. For environmental stress assessment, combining these technologies yields the best performance in tasks such as WL prediction with high accuracy and FVC prediction with better R2 scores. They also improve prediction performance in tasks like seed price query and target germplasm screening, where MLLM alone performs poorly due to insufficient domain knowledge and data dynamics .

You might also like