LECTURE TWO
Learning objectives
i) Definition of terms: population, species, gene pool, fixed locus, gene frequency, genetic
structure.
ii) Relate Mendel’s law of segregation with allele and gene frequencies.
iii) Computation of allele frequencies.
iv) Computation of genetic frequencies.
HARDY-WEINBERG LAW
Important definitions and overview
Population genetics is essentially the study of allele and genotype frequencies within populations
of organisms
Population: a localized group of interbreeding individuals
Species: a group of populations whose members are capable of interbreeding
Gene pool: In a population genetics sense, a population or species consists of a gene pool. A
gene pool "consists of all alleles at all gene loci in all individuals in a population."
Fixed locus (fixed allele): A locus for which only a single allele exists for an entire gene pool is
considered to be fixed, i.e., a fixed locus
Gene frequency (allele frequency): alleles at not-fixed loci possess a frequency that is
somewhere between 0.0 and 1.0. We describe this frequency as allele frequency or, less
correctly but more commonly, as gene frequency
Genotype frequency: Any one individual may be homozygous for only one allele of the one or
more present in the population (at a given locus) or a given individual may be heterozygous at
that locus. Thus, three alleles (or many more) can exist in a population (with associated allele
frequencies) but only up to two alleles at a time can exist within a given individual
Genetic structure: Genetic structure is a population genetics term used to refer to a population’s
allele and genotype frequencies. Evolution may be defined as change over time of a population's
genetic structure "Evolution is a generation-to-generation change in a population's frequencies
alleles and genotypes. A change in a population's genetic structure."
HARDY–WEINBERG EQUILIBRIUM
Short introduction:
An interesting question was raised following a lecture by Reginald Punnett on February 28,
1908. Statistician Udny Yule (1871–1951) asked about the distribution of brachydactyly, an
inherited condition where the fingers or toes are very short.
Most types of brachydactyly result from a dominant allele, and Yule wondered why, if the trait
were dominant, it was not more common in our species. Yule hypothesized that there should be
three cases of brachydactyly to every one person with the normal phenotype, which was clearly
not the case (Stern 1965). At first glance, it might seem to make sense that a dominant trait will
become more common in the population over time. The reason why it does not, and a
fundamental basis of microevolutionary theory, can be understood by looking at what has
become known as
Hardy–Weinberg equilibrium or the Hardy–Weinberg law. Following Yule’s question, Punnett
turned for assistance to his colleague and cricket-playing friend, mathematician Godfrey Hardy
(1877–1947). Hardy was apparently interested in pure mathematics, and as Punnett noted,
‘‘Knowing that Hardy had not the slightest interest in genetics I put my problem to him as a
mathematical one’’ (Stern 1965:220). Hardy quickly solved the problem mathematically, and
published his treatment of Mendelian proportions in populations in 1908. Hardy was too
humble, not realizing that his mathematical insight had solved a basic dilemma in genetics.
Recognizing the significance of Hardy’s insights for the development of genetics, Punnett began
referring to the principle as Hardy’s law (Stern 1965). Later, it was found that a German
physician, Wilhelm Weinberg (1862–1937), had arrived independently at the same conclusion
and had presented it in a lecture in 1908. In order to recognize both men’s independent
accomplishments, geneticist Curt Stern suggested that the basic principle be henceforth known
as the Hardy–Weinberg law (Stern 1943).
Before getting into the definition, derivation, and application of Hardy–Weinberg equilibrium,
some basic ideas for measuring genetic variation in a population need to be clarified, specifically
the ideas of genotype frequencies and allele frequencies.
The Hardy-Weinberg theorem characterizes the distributions of genotype
frequencies in populations that are not evolving, and is thus the fundamental null
model for population genetics.
RECALL
Basic Mendelian Genetics
Under the now-discredited theory of blending inheritance, the hereditary
material was conceived as a fluid that combines the traits from two individuals
into phenotypically intermediate offspring.
Given observed patterns of resemblance between parents and offspring, blending
inheritance may seem intuitively reasonable, as it did to many of Charles
Darwin’s contemporaries.
This mode of inheritance, however, posed problems for Darwin’s theory of
natural selection (1859), which depends on the existence of heritable trait
variation in populations of organisms.
Blending inheritance would quickly erode such variation, since all traits would
be combined from one generation to the next until all individuals shared the
same blended phenotype.
In his famous experiments on pea plants, Gregor Mendel rejected this hereditary
mechanism in favor of particulate inheritance by demonstrating that alternative
versions of genes (alleles) account for variations in inherited characters, though
he didn’t actually know about genes as such.
Although Mendel published his results in 1866, his work remained obscure until
its rediscovery in 1900 (reviewed in Monaghan & Corcos 1984), which helped
give rise to the modern field of genetics.
Mendel’s Law of Segregation, states that a diploid individual carries two
individual copies of each autosomal gene (i.e., one copy on each member of a
pair of homologous chromosomes).
Each gamete produced by a diploid individual receives only one copy of each
gene, which is chosen at random from the two copies found in that individual.
Under Mendel’s Law of Segregation, each of the two copies in an individual has
an equal chance of being included in a gamete, such that we expect 50% of an
individual’s gametes to contain one copy, and 50% to contain the other copy
(Figure 1).
Figure 1: Mendel's Law of Segregation
An individual’s genotype is the combination of alleles found in that individual at
a given genetic locus.
If there are two alleles in a population at locus A (A and a), then the possible
genotypes in that population are AA, Aa, and aa.
Individuals with genotypes AA and aa are homozygotes (i.e., they have two
copies of the same allele).
Individuals with genotype Aa are heterozygotes (i.e., they have two different
alleles at the A locus).
If the heterozygote is phenotypically identical to one of the homozygotes, the
allele found in that homozygote is said to be dominant, and the allele found in
the other homozygote is recessive.
Even after many geneticists had accepted Mendel’s laws, confusion arose
regarding the maintenance of genetic variation in natural populations.
Some opponents of the Mendelian view contended that dominant traits should
increase and recessive traits should decrease in frequency, which is not what is
observed in real populations.
Hardy (1908; Figure 2) refuted such arguments in a paper that, along with an
independently published paper by Weinberg (1908; Figure 3) laid the foundation
for the field of population genetics (Crow 1999; Edwards 2008).
GENOTYPIC AND ALLELIC FREQUENCIES
Are Used to Describe the Gene Pool of a Population
A. Computing Genotype Frequencies
Let’s use the standard model of a hypothetical locus with two alleles, A and a. When we have
two alleles, we will have three genotypes: AA, Aa, and aa. To make things easy, let us suppose
that the two alleles are codominant, so that we can tell the difference between people having
these three genotypes. Imagine that we visit a population and test the genotype of 150 people,
and obtain the following numbers:
AA = 54
Aa = 72
aa = 24
The first thing we would want to do is to figure out the genotype frequencies, which are the
proportions of each genotype. To do this, we simply divide each number by the total number of
individuals. Thus, because 54 of 150 people have the AA genotype, the genotype frequency is
54/150 = 0.36. We can do this for all three genotypes:
The sum of the genotype frequencies adds up to 1 (0.36 + 0.48 + 0.16 = 1.0).
Although we use genotype frequencies in population genetics, note that we could also express
these proportions as percentages by multiplying each proportion by 100: AA = 36%, Aa = 48%,
and aa = 16%. Note that these add up to 100%. At this point, it is useful to consider another
characteristic of genotype frequencies. Suppose that you were to return to this population and
choose a single person at random without knowing that individual’s genotype. What is the
probability that this person would have the genotype AA? The answer is 0.36, which is the
genotype frequency, because this number also represents the number of times that a specific
event occurs (54 people have genotype AA) divided by the total number of events (there are 150
people).
B. Computing Allele Frequencies
What are the number of A alleles and the number of a alleles in the hypothetical population
described above?
Given the genotype numbers AA = 54, Aa = 72, and aa = 24, what are the relative frequencies of
the A and a alleles? Start by counting the number of A alleles, remembering to count the A allele
twice for the AA genotype and once for the Aa genotype. There are 54(2) + 72(1) = 180 A
alleles. We repeat this procedure for the number of a alleles, getting 72(1) + 24(2) = 120 a
alleles. Thus, this hypothetical population has 180 A alleles and 120 a alleles for a total of 300
alleles (which works out since there are 150 people, each with two alleles). Therefore, the
relative frequency of the A allele is 180/300 = 0.6, and the relative frequency of the a allele is
120/300 = 0.4.
A simple way to keep all of these calculations straight is to make a table that lists the number of
A alleles and the number of a alleles for each of the three genotypes,
EXAMPLE 2.1 How to Compute Allele Frequencies. In this example, we use the hypothetical
example from the text of a single locus with two alleles, A and a. The data are
Number of people with genotype AA = 54
Number of people with genotype Aa = 72
Number of people with genotype aa = 24
In the following table, we list the number of A and a alleles and the total number of all alleles for
each genotype, and then sum each column:
There are a total of 180 A alleles and 120 a alleles in this population, for a total of 300 alleles
(twice the number of people sampled, because each person has two alleles (diploid number)).
From the genotype frequency, one can calculate the allele frequencies.
As a check, the allele frequencies should add up to 1.0, which they do (0.6 + 0.4 = 1.0).
Because we use allele frequencies extensively in population genetics equations, we use symbols
to refer to the different allele frequencies. Although you can feel free to create any symbol you
wish, the conventional format is to use the symbol p to refer to one allele and the symbol q to
refer to the other allele. In Example 2.1, p is the relative frequency of the A allele and q is the
relative frequency of the a allele.
Note that in the case of two alleles, there are only two allele frequencies, and these numbers
must add up to 1.0. In Example 2.1, this is seen because 0.6 + 0.4 = 1.0. This property allows us
a convenient check on our math; if the two allele frequencies do not add up to 1.0 (or very close,
in the case of irrational numbers and roundoff), then an error was made in computing the allele
frequency.
Mathematically, we can express the sum of the allele frequencies using the
Following simple formula:
p+q=1
Where p and q are allele frequencies for a genetic locus with two alleles.
Note that this relationship allows us to predict one allele frequency if we know the other.
When an allele is rare, the corresponding homozygote genotype is even rarer; the genotype
frequency is the square of the allele frequency.
WHAT IS HARDY–WEINBERG EQUILIBRIUM?
At the most basic level, Hardy–Weinberg is a mathematical expression describing the expected
genotype frequencies in a new generation.
The law is actually a mathematical model that evaluates the effect of reproduction on the
genotypic and allelic frequencies of a population. It makes several simplifying assumptions
about the population and provides two key predictions if these assumptions are met. For an
autosomal locus with two alleles, the Hardy–Weinberg law can be stated as follows:
Assumptions If a population is large, randomly mating, and not affected by mutation, migration,
or natural selection, then:
Prediction 1 the allelic frequencies of a population do not change; and
Prediction 2 the genotypic frequencies stabilize (will not change) after one generation in the
proportions p2 (the frequency of AA), 2pq (the frequency of Aa), and q2 (the frequency of aa),
where p equals the frequency of allele A and q equals the frequency of allele a.
The Hardy–Weinberg law indicates that, when the assumptions are met, reproduction alone does
not alter allelic or genotypic frequencies and the allelic frequencies determine the frequencies of
genotypes.
The statement that genotypic frequencies stabilize after one generation means that they may
change in the first generation after random mating, because one generation of random mating is
required to produce Hardy–Weinberg proportions of the genotypes. Afterward, the genotypic
frequencies, like allelic frequencies, do not change as long as the population continues to meet
the assumptions of the Hardy–Weinberg law. When genotypes are in the expected proportions of
p2, 2pq, and q2, the population is said to be in Hardy–Weinberg equilibrium.
Imagine that we have a locus with two alleles, A and a, where p is the frequency of the A allele
and q is the frequency of the a allele.
If we know that p = 0.7 and q = 0.3 in the parental gene pool, we can quickly figure out the
expected distribution of genotypes in the next generation as:
Genotypic Frequencies at Hardy–Weinberg Equilibrium
How do the conditions of the Hardy–Weinberg law lead to genotypic proportions of p2, 2pq, and
q2? Mendel’s principle of segregation says that each individual organism possesses two alleles at
a locus and that each of the two alleles has an equal probability of passing into a gamete. Thus,
the frequencies of alleles in gametes will be the same as the frequencies of alleles in the parents.
Suppose we have a Mendelian population in which the frequencies of alleles A and a are p and
q, respectively. These frequencies will also be those in the gametes. If mating is random (one of
the assumptions of the Hardy–Weinberg law), the gametes will come together in random
combinations.
The multiplication rule of probability can be used to determine the probability of various
gametes pairing. For example, the probability of a sperm containing allele A is p and the
probability of an egg containing allele A is p.
Applying the multiplication rule, we find that the probability that these two gametes will
combine to produce an AA homozygote is p × p = p2. Similarly, the probability of a sperm
containing allele a combining with an egg containing allele a to produce an aa homozygote is q
× q = q2. An Aa heterozygote can be produced in one of two ways: (1) a sperm containing allele
A may combine with an egg containing allele a (p × q) or (2) an egg containing allele A may
combine with a sperm containing allele a (p × q). Thus, the probability of alleles A and a
combining to produce an Aa heterozygote is 2pq. In summary, whenever the frequencies of
alleles in a population are p and q, the frequencies of the genotypes in the next generation will
be p2, 2pq, and q2.