Module 5
Module 5
edu)
Home > Course Materials > Crop Genetics > Linkage
Linkage
By Thomas Lübberstedt, Arden Campbell, Deborah Muenchrath, Laura Merrick, Shui-Zhang Fei (ISU)
Except otherwise noted, this work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
Introduction
Genes located on the same chromosome are
genetically linked. Genetic linkage analysis can be
used to determine the order of genes on
chromosomes. Closely linked genes are not
segregating independently, like genes located on
different chromosomes. This has different
implications, e.g., in relation to trait correlations.
Moreover, linked genes can be used as genetic
markers, which have become an important tool in
plant breeding.
Objectives
• Develop an understanding of the genetic
basis of linkage.
• Gain awareness on how to detect the
occurrence of linkage.
• Review the principles of genetic map
Fig. 1 Genes located on the same chromosome are
construction.
genetically linked.
• Become familiar with the concept of linkage
disequilibrium.
Crossover and Recombination
Genetic Organization
Fig. 2 Genetic mapping involves specifying which chromosome a gene is located on, along with the position on that
chromosome. Illustration by Iowa State University.
Genes are physically organized on chromosomes. Each gene is located at a particular “address” (particular
position on a speci�c chromosome, which can be identi�ed by genetic mapping). Inheritance of genes located
on different chromosomes follows the rules of independent assortment. Since plant species have multiple
chromosomes, independent assortment is true for the majority of genes. In contrast, linked genes located on
the same chromosome are more likely to cosegregate, i.e., being jointly transmitted to offspring more often
than expected by independent assortment. The biological process that separates linked genes is the crossing-
over (C.O., or crossover), which occurs during meiosis, and leads to genetic recombination.
Crossing-Over
During meiosis of diploid organisms, the chromatids of homologous chromosomes pair and form bivalents.
During Meiosis I, homologous chromatids pair to physically exchange chromosome segments. The
chromosomal site, where this reciprocal exchange of homologous chromosome segments takes place, is called
a chiasma. Thus, crossing-over involves not completely understood mechanisms for identi�cation of
homologous sites of chromatids, breakage and rejoining of chromosomes.
Genetic Distance
Crossing-over events occur more or less random during meiosis. In most plant species, one to few crossing-
over events occur per meiosis and chromosome. Thus, the closer the genes are physically linked on the same
chromosome, the less likely they will get separated, and consequently, the less likely genetically recombinant
gametes will be produced. This is the underlying principle of genetic maps: the genetic distance between genes
re�ects the probability of a crossing-over between linked genes.
Recombination
Linkage Phase
Linkage phase is the physical arrangement of linked genes in a chromosome. A double heterozygote with a
genotype of AaBb could be in one of the two linkage phases. Conventionally, when linked dominant alleles are
located on the same homologous chromosome and the linked recessive alleles are on the other homologous
chromosome, for example, AB/ab, it is said the genes are linked in coupling phase. When a dominant allele at
one locus is on the same homologous chromosome as a recessive allele of the other linked gene, for example
Ab/aB, it is said that the genes are linked in repulsion phase (Fig. 4).
This knowledge is crucial, as linkage detection and distance estimation is based on the observed parental and
non-parental gametes.
In case of close linkage, non-parental gametes and respective offspring are underrepresented.
An example is the Australian sheep blow�y, Lucilia cuprina. Normal blow�ies have a green thorax and surround
themselves in a brown cocoon during their pupal stage. However, recessive genes (here marked a and b) can
cause the �y to develop a purple thorax and spin a black puparium.
Fig. 5 Australian sheep blow�y, Lucilia cuprina. Photo by �r0002, licensed under CC BY-NC via Wikimedia Commons.
Using Testcrosses
For detection of linkage, appropriate testcrosses need to be conducted. The linkage phase is known, if two
homozygous parental genotypes (AABB and aabb) are crossed to produce the respective F1 (AaBb).
Fig. 7
The non-parental recombinant gametes have the genotype Ab and aB, whereas the parental gametes have the
genotype AB and ab.
Usually the phenotype cannot be observed in (haploid) gametes, but only in diploid plants. Thus, to determine
whether two loci are linked, offspring need to be produced. This can be achieved by self pollination of the AaBb
– F1, by production of doubled haploid offspring, or by a testcross.
In this particular example, a backcross (BC) of the F1 to the aabb parent would be the best option.
Testcross Gametes
For detection of linkage, a Chi-Square test can be employed. The Chi-Square test compares observed with
expected frequencies. In this case, the null hypothesis to determine expected frequencies is the assumption of
independent assortment. Under this assumption, equal frequencies of all four gametes are expected. In case of
linkage, BC1 individuals carrying non-parental gametes are underrepresented, leading to a statistically
signi�cant Chi-Square value. This means that the null hypothesis of independent assortment would be rejected
and linkage assumed.
Table 1 An example of the detection of linkage in Drosophila melanogaster using a Chi-Square test. d: difference between
the observed number and expected number. The signi�cantly higher Chi-Square values reject the null hypothesis and
strongly indicate the presence of linkage.
To better understand the use of Chi-Square in determining linkage, two numerical examples based on the cross
schemes described in Figs. 8 and 9 are provided here. In both examples, a sample size of 2,000 BC1 individuals
has been used.
The Chi-Square test sums up over all squared differences between observed and expected values, divided by
expected values.
In example A, observed and expected values are equal, thus the Chi-Square value = 0.
In example B, the squared differences between observed and expected values is in all cases 90,000, to be
divided by the expected 500 = 180. As there are four genotypic classes, the Chi-Square value is 720, which is
signi�cantly larger than the tabulated value of 3.81 (p = 5%).
In conclusion, example A is in agreement with independent assortment, whereas in example B, linkage has been
detected.
Fig. 10 Expected vs. obsevered distributions of phenotypes
In conclusion, example A is in agreement with independent assortment, whereas in example B, linkage has been
detected.
Genetic Distance
In case of complete linkage of two genes, no recombinants would be expected. The recombinant frequency
would be 0%, which represents the lower limit of recombinant frequencies.
In case of random segregation, the expected numbers of recombinant and nonrecombinant alleles are equal.
Thus, the upper limit of recombinant frequencies in case of unlinked or loosely linked genes is 50%.
Even for gene pairs located at the different ends of the same chromosome, recombination frequency can reach
50%. The procedure to determine recombination frequencies between any pair of genes is called two-point
analysis.
Study Question 1
You have a F1 plant heterozygous at two loci that are 12 map units apart on the same chromosome. The
F1 received linked recessive alleles from one parent and linked dominant alleles from the other parent.
Study Question 2
You have a F1 plant heterozygous at two loci that are 12 map units apart on the same chromosome. The F1
received linked recessive alleles from one parent and linked dominant alleles from the other parent.
Since linkage cannot be detected in the F1, you self-pollinate the F1 and evaluate the F2. What
would be the F2 and testcross percentages if the F1 percentages in case of repulsion phase of the
recessive alleles ?
Check
Three-Point Analysis
Purpose
Whereas two-point testcrosses establish linkage between pairs of genes, three-point testcrosses facilitate
establishment of the order of genes on chromosomes, as prerequisite to establish genetic maps. If a third locus
with alleles C and c (C is dominant over c) is added to the case mentioned in Genetic Distance, where A and B
are linked in coupling phase and the dominant allele C is in coupling with A and B, then eight
different testcross progeny would result from a backcross with the recessive parent.
2 a c b 173
3 A c b 52 } Recombinants, single
crossover AC
4 a C B 46
5 A C b 22 } Recombinants, single
crossover CB
6 a c B 22
7 A c B 4 } Recombinants, double
crossover AC, CB
8 a C b 2
Total = 500
Fig. 12 Example for a testcross (backcross) as �rst step toward a three-point analysis. Adapted from Russell, 2010.
Frequency Chart
• 20.8% for AC (AC recombinants are in classes 3, 4, 7, and 8; thus, the recombination rate between A and
C is (52+46+4+2/500) * 100% = 20.8%)
Once linkage between pairs of three (or more) genes has been established, the next question is how they are
arranged in linear order on chromosomes, which could be ABC, ACB, or CAB.
1 A C B 179
} Parentals, no crossover
2 a c b 173
3 A c b 52 Recombinants, single
}
4 a C B 46 crossover AC
5 A C b 22 Recombinants, single
}
6 a c B 22 crossover CB
7 A c B 4 Recombinants, double
}
8 a C b 2 crossover AC, CB
Total = 500
Gene Order
The most likely gene order minimizes the sum of pairwise recombination frequencies within a three-gene
interval, which would be:
Thus, the most likely gene order is ACB. In other words, the interval between A and B can be subdivided into the
intervals between AC and CB.
2 a c b 173
3 A c b 52 } Recombinants, single
crossover AC
4 a C B 46
5 A C b 22 } Recombinants, single
crossover CB
6 a c B 22
7 A c B 4 } Recombinants, double
crossover AC, CB
8 a C b 2
Total = 500
Expressed yet another way: incorrectly ordered genes would increase the total map length, because part of the
recombination events would be counted twice. If ACB is the true order, then the genetic length of, e.g., ABC
would be in�ated, because recombinants for the segment BC would be counted two times: for the interval BC, in
addition to the same interval within the segment A(C)B. Algorithms of mapping programs use this principle
(minimizing the genetic distance) for three-point-analyses.
Double Crossovers
Genotype of gamete
Class Number Origins
from heterozygous parent
1 A C B 179 Parentals,
} no
2 a c b 173 crossover
Recombinants,
3 A c b 52
single
}
crossover
4 a C B 46
AC
Recombinants,
5 A C b 22
single
}
crossover
6 a c B 22
CB
Genotype of gamete
Class Number Origins
from heterozygous parent
Recombinants,
7 A c B 4
double
}
crossover
8 a C b 2
AC, CB
Total = 500
Coefficient of Coincidence and Interference
Crossover events in adjacent chromosome regions might affect each other, a phenomenon called interference.
Most typically, a crossover event in one region tends to suppress a crossover in the adjacent regions. The
extent of interference is expressed by the coe�cient of coincidence, which is equal to the observed frequency
of double crossovers / expected frequency of double crossovers.
1 A C B 179
} Parentals, no crossover
2 a c b 173
3 A c b 52 Recombinants, single
}
4 a C B 46 crossover AC
5 A C b 22 Recombinants, single
}
6 a c B 22 crossover CB
7 A c B 4 Recombinants, double
}
8 a C b 2 crossover AC, CB
Total = 500
The expected frequency of double crossovers is the product of two single crossovers in adjacent regions
assuming there is no interference.
In this example, this expected frequency is 0.21(recombination frequency for AC) * 0.10 (recombination
frequency for CB) = 0.021.
The observed frequency of double crossover events is 6/500 in the example, resulting in 0.012. Thus, the
coe�cient of coincidence in this example is 0.012/0.021 = 0.58.
Interference is de�ned as 1 − coe�cient of coincidence, which would be 0.42 in this example. A value of zero
for interference would mean that a crossover in one region does not affect crossovers in the adjacent region.
Interference of 1 means, that crossovers in one region suppress crossovers in the adjacent region. Negative
values are possible and have been reported in some instances, which means that crossovers in one region
stimulate crossovers in the adjacent region.
Map Functions
Measurement Units
The purpose of genetic maps is to report the length of chromosome intervals, chromosomes, and whole
genomes. Since recombination frequencies converge to a value of 50% as reported above, indicating absence
of linkage, recombination frequencies are not additive and, thus, not useful to describe the distance between
genes that are located far apart. When recombination frequency reaches 50%, it would be impossible to tell
whether the genes are located far apart on the same chromosome or on different chromosomes.
Instead, estimates of the number of crossover events are used as additive measure of genetic map distances.
The unit for measuring genetic distances is Morgan (M), or usually centiMorgan (cM). In contrast to
recombination frequencies, map units expressed in cM are additive. One Morgan re�ects the observation of one
crossover event per single meiosis. One cM is a distance between genes that produces 1% recombinants in the
offspring. Typical lengths of genetic maps in maize, for example, vary between 1,600 to 2,000 cM, which means
that on average, 1.6 – 2 crossovers occur per chromosome and single meiosis in maize (maize has 10
homologous chromosome pairs).
Frequency Conversion
As mentioned in the previous page, direct observation of crossover events is cumbersome. For that reason
most genetic maps published to date are based on the conversion of recombination frequencies into crossover
frequencies. The main obstacle on translating recombination frequencies into crossover frequencies is the
variable and unknown degree of interference in different genome regions. While it has been possible in the
earlier example to determine the degree of interference, and thus frequency, of double crossover events, in the
genetic interval between A and B by adding C, the degree of interference between AC and CB is unknown. This
could be addressed by observing segregation of further genes within these two regions (if available), but this
issue could ultimately only be addressed by complete genome sequencing of all offspring in a mapping
population, which at this point is still too costly.
Visual Relationship
Instead, map functions have been developed, that translate recombination frequencies into crossover
frequencies, and thus cM (see Fig. 14 below). Figure 14 clearly shows, that there is an approximately linear
relationship between recombination rates (y-axis) and crossover rates (x-axis). However, with increasing map
distances, recombination rates converge to 50%. In other words, gene pairs with crossover rates of 80 cM or
200 cM, respectively, would be nearly indistinguishable based on recombination rates, which would result in
recombination rates between 40 and 50%.
Fig. 15 Relationship between crossover rates (x-axis) and recombination rates (y-axis). Adapted from Russell, 2010.
The various available map functions make different assumptions on the extent of interference. For example, the
Haldane mapping function assumes absence of interference. In contrast, the Kosambi function assumes
presence of interference.
Other Types of Maps
Genetic maps can be generated in other ways than using testcrosses. Examples include somatic cell
hybridization and tetrad analysis. In plants, interspecies addition lines such as oat-maize addition lines created
by distant hybridization have been developed as tool for mapping of genes. If two genes appear on the same
addition segment, they are genetically linked. Besides genetic maps, cytological and physical maps can be
established.
Fig. 17 Cytogenetic map of maize chromosomes. Illustration by Neuffer et al., 1997. Used with permission.
Factors Influencing Linkage Mapping
Linkage mapping based on testcrosses can be affected by selection or incomplete penetrance, among others.
Selection in the most extreme case would be due to lethality of gametes (gametic selection) or zygotes (zygotic
selection). If a backcross is used for linkage detection, as described above, lethality of male gametes carrying
for example the a allele would lead to only two classes of BC progeny, if AaBb is crossed as pollinator to aabb.
In that case, only AaBb (parental) and Aabb (recombinant) genotypes would be obtained. Zygotic selection
affects the viability of particular genotypes. If the aa genotype in the example above is lethal, then the aa
offspring derived from self pollination of an AaBb genotype would be missing. Incomplete penetrance means
that a genotype which is supposed to express, for example, red �owers, has to a certain extent white �owers. In
other words, there is no 100% match between genotype and phenotype, but due to environmental factors, the
phenotype might differ. As for selection, incomplete penetrance alters the frequency of expected genotypes in
testcrosses, which is the basis for detection of linkage.
Consequences and Applications of Linkage
The main application of linkage is in genetic mapping of genes using molecular markers. Once genes have been
mapped and closely-linked markers identi�ed, those markers can be used for marker-aided selection
procedures. Technological progress in DNA methods has been and still is rapid, so that thousands of markers
can be produced at low cost in any species of interest. Moreover, novel genomic selection strategies
addressing complex inherited traits are being developed.
Linkage can in some cases be confused with pleiotropy. If a favorable character (e.g. resistance) is always
inherited together with an unfavorable trait (e.g., lodging), a negative pleiotropic effect might be assumed,
which might alternatively be caused by two closely linked genes. Whereas close linkage can be resolved to �nd
favorable genotypes for both traits, this is not true for pleiotropy. Linkage reduces the possible genetic variation
in small populations. With increasing numbers of generations, or population sizes, genetic variation can be
increased. Similarly, inbreeding reduces the opportunity for effective recombination.
Linkage Disequilibrium
Genotype Distribution
Although allele frequencies at individual loci are expected to be stable in case of random mating, genotype
frequencies at two or more loci jointly do not achieve this equilibrium after one generation of random mating.
To illustrate this point, consider two populations, one consisting of entirely AABB genotypes and the other
consisting entirely of aabb genotypes. Assumed they are mixed equally and allowed to randomly mate. The �rst
generation would consist of the three genotypes AABB, AaBb, and aabb in the proportions 1/4 : 1/2 : 1/4.
However, for two loci, each with two alleles, nine genotypes are possible. (For n alleles at each locus and k loci,
there are:
possible genotypes). Continued random mating would produce the missing genotypes, but they would not
appear at the equilibrium frequencies immediately.
Equilibrium
Consider the following table based on two alleles at each of two loci:
Alleles A a B b
Allele Frequencies PA Pa PB Pb
Gametic Types AB Ab aB ab
In linkage equilibrium, the expected gamete frequencies can be calculated from the marginal allele frequencies.
For example, in equilibrium, the frequency of gamete AB (PAB) would be expected to be equal to the product of
the frequencies of the A allele (PA) and the B allele (PB).
This is valid under the following conditions: PA + Pa = 1; PB + Pb = 1; and PAB + PAb + PaB + Pab = 1.
If, for example, the allele frequencies of PA = Pa and PB = Pb are 0.5, then the frequencies of all gametes are
0.25.
A measure for Disequilibrium, D = PAB - PA*PB. D = 0 in case of equilibrium. If D differs from 0, it re�ects
presence of Disequilibrium. In other words, the frequency of a gamete differs from its expected frequency
based on marginal probabilities of the respective individual alleles.
LD and Mapping
Linkage disequilibrium is the non-random association of alleles at different loci. LD is extensively used in
mapping human disease genes using natural populations (Association mapping).
In plants, gene mapping has been conducted mainly by using mapping families because of the ease with which
mapping families are created, but LD mapping using natural populations is increasing rapidly because such
populations are large in size and have much greater allelic diversity.
Fig. 18 Linkage disequilibrium and equilibrium. When LD is present, all individuals possessing red alleles in locus A have
green alleles in locus B. When the two loci are in equilibrium, individuals having red alleles in locus A could have any
alleles in locus B. Adapted from Rafalski, 2002.
LD Statistic D’
|D’| =
|D’| =
It can be shown that after t generations of random mating, the remaining disequilibrium is given by:
where, D0 is the disequilibrium in generation 0 and c is the recombination fraction, with c = 0.5 for
independently segregating loci, which is identical to a recombination frequency of 50% (the range of c is from 0
to 0.5, whereas the range of r is from 0% to 50%). The dissipation of disequilibrium relative to generation 0 is
given in Fig. 18.
Recombination and LD
Generally, deviations from independence at multiple loci are referred to as linkage disequilibrium, even if genetic
linkage is not the cause (in other words, alleles are not physically linked). Unless two loci are known to reside
on the same chromosome, the term Gametic Disequilibrium should be used to describe disequilibrium among
loci. Whereas recombination and crossover frequencies, as mentioned initially, are used to describe the
distance between genes from a chromosomal perspective, linkage disequilibrium is mostly used to describe a
property of populations. However, both terms are closely related.
Genetic Markers
Genetic Markers
Overview
Genetic variation results from differences in DNA sequences and, within a population, occurs when there is
more than one allele present at a given locus. Such populations are referred to as populations that are
polymorphic or segregating at that locus. The opposite situation is when all members of the population are
homozygous for the same allele, in which case the population is said to be �xed or monomorphic for that allele.
A genetic marker is a DNA sequence that exhibits polymorphism among individuals and can thus be used to
identify a particular locus (although not necessarily a gene) on a particular chromosome; the marker itself may
be part of a gene or may have no known function. Markers are inherited in a Mendelian fashion and facilitate
the study of inheritance of a trait or sometimes a linked gene. Markers are used to identify, map, and isolate
genes, select desired genotypes, and detect genetic variation or determine genetic relationships among
individuals. Markers are regions of genomes that are heritable, often easy to document, and useful for detecting
genetic variation.
Three Types
Three Types
Genetic markers generally do not represent target genes of interest to a breeding program, but instead are
useful as 'signs' or 'tags', particularly when they are closely linked to genes that control a trait of interest. A
genetic map constructed with genetic markers is similar to a road map. Linkage groups in a genetic map
represent roads whereas individual markers on each linkage group represent signs or landmarks that help plant
breeders to navigate through the plant genome and �nd the genes of interest.
• Morphological markers
• Biochemical markers
• Molecular markers
Morphological Markers
These types of markers (also called visible or classical markers) are phenotypic traits with only a few distinct
morphs or variants (e.g., �ower color or seed shape), usually due to one or perhaps two gene loci so they are
not strongly affected by the environment. Inheritance patterns of visible and morphological characters have
been used to map genes to particular chromosome segments and to identify linkage groups. Such markers are
limited in number compared to the abundance of DNA markers, however, and may be in�uenced by
developmental stage of the plant.
Biological Markers
Isozymes (sometimes called allozymes) are allelic variants of a single enzyme that share the same function,
but may differ in level of activity due to differences in amino acid sequence. Isozymes are proteins for which
variation can be detected by differential separation using electrophoresis, a technique for separating
macromolecules (DNA, RNA, protein) on a gel by means of an electric �eld and speci�c chemical staining.
Isozymes have codominant expression, meaning that both homozygotes can be distinguished from
the heterozygote and neither allele is recessive. In contrast to codominant markers, dominant markers are
either present or absent.
In comparison to visible polymorphisms they reveal more of the underlying genetic variation.
However isozymes are gene products, so they reveal only a small subset of the actual variation in DNA
sequences between individuals and do not reveal variation in the non-coding regions of the genome. In general,
such markers are limited in number and have limited use in genetic mapping studies.
Fig. 20 Electrophoresis is a laboratory technique used to separate DNA, RNA or protein molecules based on their size and
electrical charge. Illustration adapted from NIH-NHGRI, 2011.
Molecular Markers
Molecular or DNA markers reveal sites of variation in DNA. Variability in DNA facilitates �ner scale mapping and
detection. Mapping is the process of making a representative diagram cataloging genes and other features of a
chromosome and showing their relative positions. Many of these molecular markers avoid the limitations
associated with visible and biochemical markers. They facilitate evaluation of genome-wide coverage and are
not affected by environmental factors or developmental stages. They allow high resolution of genetic diversity
to be detected. Molecular markers have added substantial amounts of information to our genetic maps.
Any DNA sequence can be genetically mapped, like genes leading to plant phenotypes as long as there is a
polymorphism available for the sequence to be mapped, i.e., two or more different alleles. This can basically be
a single nucleotide polymorphism (SNP), a single nucleotide variant at a particular position within the target
sequence, or an insertion / deletion (INDEL) polymorphism. Any target sequence can be ampli�ed by the
Polymerase chain reaction (PCR), and subsequently be visualized to generate "molecular phenotypes"
comparable to visual phenotypes, that can be observed by using appropriate equipment.
Various molecular methods have been developed to visualize SNPs or INDEL polymorphisms at low cost and
high throughput, which will be presented in detail in the Molecular Genetics and Biotechnology course. The
main use of those SNPs and INDEL polymorphisms is as molecular markers. By genetic mapping as described
above, linkage between genes affecting agronomic traits or morphological characters, and DNA-based SNP or
INDEL markers can be established. It can be more effective in the context of plant breeding, to select indirectly
for such DNA markers, than directly for target genes. This is due to lower costs for DNA analyses, the ability to
run multiple such assays (for multiple target genes) in parallel, the ability to select early and to discard
undesirable genotypes or to perform selection before �owering, codominant inheritance of markers, among
others.
For example, both ginkgo trees (Ginkgo biloba) and asparagus (Asparagus o�cinalis) are dioecious species.
Male plants are preferred for ginkgo tree because fruits produced from female trees have an unpleasant smell
whereas male asparagus plants are preferred because of their higher yield potential. Unfortunately, sex
expression will take years to occur for both species. If a DNA marker that either directly affects sex expression
or is linked to genes that affect sex expression can be identi�ed, selection of male plants can be conducted in
early seedling stage rather than waiting for many years. Occurrence of environment conditions favoring
selection for disease, insect resistant plants or drought tolerant plants such as the prevalence of the particular
disease or insect or drought is not always reliable. Selection using DNA markers can overcome these
limitations as they are not affected by the environment.
Polymorphism
Polymorphism involves one of two or more variants of a particular DNA sequence. The most common type of
polymorphism involves variation at a single base pair, also called single nucleotide polymorphism (SNP) (Fig.
22). Polymorphisms can also be much larger in size and involve long stretches of DNA. Tandem repeat is a
sequence of two or more DNA base pairs that is repeated in such a way that the repeats are generally
associated with non-coding DNA. In contrast, SNPs can sometimes be identi�ed that occur within coding
sequences (that is within genes), as well as in non-coding DNA.
Fig. 22 Example for a SNP (yellow highlighted) in a population of six diploid genotypes. Individuals 1, 4, 5, 6 are
T/C heterozygotes, individual 2 a C/C homozygote, and individual 3 a T/T homozygote. Adapted from NIH-NHGRI, 2011.
Types of Biochemical/Molecular Markers
Protein-based DNA-based
Isozymes RFLP RAPD AFLP SSR SNP
No. of loci 30-50 100s ~Unlimited ~Unlimited 10s 10s
Degree of
Low-medium Meduim-high Medium-high Medium-high High High
polymorphism
Nature of gene action Codominant Codominant Dominant Dominant Codominant Codominant
Reproducibility High High Low-medium Medium-high High High
Amount of DNA per
Not applicable mg ng ng ng ng
sample
Method* Biochemical DNA-DNA hybridization PCR PCR PCR PCR
Ease of array? Easy Di�cult Easy Moderate Easy-moderate Easy
Can be automated? Di�cult Di�cult Yes Yes Yes Yes
Equipment cost Inexpensive Expensive Moderate Expensive Expensive Expensive
Development cost Inexpensive Expensive Moderate Expensive Very Expensive
Assay cost Inexpensive Expensive Moderate Expensive expensive Expensive
* ‘PCR’ means Polymerase Chain Reaction ampli�cation of genomic DNA fragments, a method that uses short,
single-stranded DNA sequences, known as primers, to hybridize with the sample DNA Table 2. Comparison
among widely used molecular markers. Adapted from Nageswara-Rao and Soneji, 2008.
SSR and SNP Markers
SSR markers remain useful to plant breeders due to their abundance and convenience with which they are
assessed, but they serve most likely as linked markers. SNP markers, however can either be linked to or directly
reside in a gene of interest and are hugely abundant. For these reasons, they are increasingly becoming the
marker of choice.
Uses of Molecular Markers
Molecular markers are useful for both applied and basic genetic research. Here are some examples:
Genetic mapping
Molecular markers provide a means to map genes to more speci�c chromosome segments than is possible
using visible markers.
Genetic relationships within families, genera, species, or cultivars can be determined from molecular markers.
Much information about the evolution of crops has been learned using molecular markers. The markers also
enable breeders to monitor the genetic diversity among breeding lines to broaden the genetic base and reduce
the risk of widespread genetic vulnerability to detrimental conditions.
Molecular Marker 'Fingerprints'
Individuals possessing more markers in common than could occur by random chance are closely related. Such
molecular �ngerprints have been used successfully in court to prove misappropriation of proprietary breeding
lines.
Isolate genes
Molecular markers are used to map candidate genes in a much �ner scale and can eventually isolate candidate
genes by positional cloning. Isolated genes can be used to study gene regulation or to directly improve
agronomic performance by genetic transformation. Although molecular markers have many applications and
provide useful tools to plant breeders, lines must still be evaluated under normal production conditions before
their release.
Reflection
The Module Re�ection appears as the last "task" in each module. The purpose of the Re�ection is to enhance
your learning and information retention. The questions are designed to help you re�ect on the module and
obtain instructor feedback on your learning. Submit your answers to the following questions to your instructor.
1. In your own words, write a short summary (< 150 words) for this module.
2. What is the most valuable concept that you learned from the module? Why is this concept valuable to
you?
3. What concepts in the module are still unclear/the least clear to you?
References
Falconer, D.S., and Trudy F.C. Mackay. Introduction to Quantitative Genetics (4th edition). San Francisco, CA:
Benjamin Cummings.
Fehr, Walter R. 1987: Principles of Cultivar Development. Macmillan Publishing Company: New York.
Nageswara-Rao, M., J.R. Soneji, C. Chen, S. Huang, and F.G. Gmitter. 2008. Characterization of zygotic and
nucellar seedlings from sour orange-like citrus rootstock candidates using RAPD and EST-SSR markers. Tree
Genet Genomes. 4:113-124. DOI: 10.1007/s11295-007-0092-2
Neuffer M.G., Coe, E.H. and Wessler, S.R. (1997) Mutants of Maize. Plainview, NY: Cold Spring Harbor
Laboratory Press
NIH-NHGRI (National Institutes of Health. National Human Genome Research Institute). 2011. Talking Glossary
of Genetic Terms. [available online September 23, 2011, [Link]
Neuffer M.G., Coe, E.H. and Wessler, S.R. (1997) Mutants of Maize. Plainview, NY: Cold Spring Harbor
Laboratory Press
Pierce, Benjamin A. 2008: Genetics: A Conceptual Approach. W.H. Freeman and Company: 160-199, 335-339.
New York, NY.
Rafalski, A. 2002 Applications of single nucleotide polymorphisms in crop genetics. Current Opinion in Plant
Biology 2002, 5:94–100
Russell, Peter J. 2012: iGenetics: A Molecular Approach. Pearson Education, Inc., San Francisco, Calif.
Acknowledgements
This module was developed as part of the Bill & Melinda Gates Foundation Contract No. 24576 for Plant
Breeding E-Learning in Africa.
Crop Genetics Linkage Author: Thomas Lübberstedt, Arden Campbell, Deborah Muenchrath, Laura Merrick, and
Shui-Zhang Fei (ISU)
Multimedia Developers: Gretchen Anderson, Todd Hartnell, and Andy Rohrback (ISU)
How to cite this module: Lübberstedt, T., A. Campbell, D. Muenchrath, L. Merrick, and S. Fei. 2016. Linkage. In Crop
Genetics, interactive e-learning courseware. Plant Breeding E-Learning in Africa. Retrieved from
[Link]
Source URL: [Link]