© GTS367 University of Pretoria 2026
GTS 367 – Population and Evolutionary Genetics
Tutorial 6: Natural Selection in finite populations
The theory of near neutrality divides mutations into three broad categories based on if 2Ns > 1, - 1 < 2Ns < 1 and 2Ns < -1. For
these three areas u(s,N) is respectively approximately equal to 2s, 1/2N and 0 for new mutations. First, we will compare the relative
strength of selection and drift. For a neutral allele, if three hundred populations started with an allele frequency of 0.5, we would
expect 150 of those populations to become fixed for A1 and 150 for A2 since drift would be the only factor changing allele
frequencies. On the other hand, if the mutation had a large enough selective coefficient, we would expect the favoured allele to fix
in all 300 populations. The theory of near neutrality claims that as s gets smaller (or N gets smaller) alleles will start to behave as if
they are neutral.
Use all your decimal points when working. Report only three (unless you have very small values, then report in scientific notation
with four digits, e.g. 5.678x10-10).
1. You will test this broad prediction by running 1000 simulated populations in the program PopG for 2000 generations. Each
simulation must be run with a diploid population size of 500. Consider an allele A1, with fA1(0) = 0.001. Set s equal to the values
listed in the table below with fitness (or viabilities) of the three genotypes calculated as vA1A1 = 1 + 2s, vA1A2 = 1 + s and vA2A2 = 1.
[13.5]
a) Record the number of times A1 was fixed and lost for each of the 1000 simulations. (4)
s -0.1 -0.01 -0.001 -0.0005 0.0005 0.001 0.01 0.1
A1 fixed
A1 lost
b) In the next 300 simulated populations, we set s = 0.01, calculating viabilities as before, this time for 6000 generations, and
varying the population size. Again, you must record the number of times A1 fixed and was lost. For this exercise, keep the
allele frequency as before. (3.5)
N 10 50 100 500 1000 5000 10000
A1 fixed
A1 lost
c) Given that the strength of drift is proportional to 1/2N, how would you explain the outcome of the simulations? Summarize
the outcomes and then compare them to the relationship of 1/2N to s. (5)
d) What determines if selection or drift is more powerful? (1)
2. Use the values obtained in question 1a for the number of A1 fixing in each simulation, for a diploid population of size 1000.[15]
s -0.1 -0.01 -0.001 -0.0005 0.0005 0.001 0.01 0.1
A1 fixed 0
2Ns -200
u(s,N)a 0
u(s,N)b 4.2402
x 10-175
For a-c, use the table above as template for your own work.
a) Calculate 2Ns for each s to divide the table into three parts where the approximate u(s,N) is 0, 1/2N and 2s. (3.5)
b) Give these approximate values of u(s,N) in the row marked u(s,N)a. (3.5)
c) Calculate u(s,N) using equation 8.1 in the row marked u(s,N)b. Do these values differ much from the approximations?
(3.5)
If your calculator does not want to work out the values, work the numerator and denominator out separately and then
divide them.
1 − 𝑒𝑒 −2𝑠𝑠
1 − 𝑒𝑒 −4𝑁𝑁𝑁𝑁
d) To compare observed and expected numbers it is normally best to do a statistical test, however, when any expected value
is less than 5, the χ2 test becomes unreliable. This requirement excludes all but the two final columns. Therefore, calculate
χ2 value for the last two columns. (3.5)
1
© GTS367 University of Pretoria 2026
(𝐸𝐸 − 𝑂𝑂)2
𝜒𝜒 2 = �
𝐸𝐸
Hint: in this case, you need to sum your calculations of A1 fixed and A2 fixed, as these are the two classes. You need to calculate this
sum for both of the last columns.
e) Compare your critical value of χ2 for 1 df that is 5.024 (corrected for two tests, to give an overall significance value of p<
0.05). (1)
3. Researchers sequenced 420 nucleotides from a human and a mouse that code for 140 amino acids that are part of a larger
protein. These two taxa split 96 million years ago. By comparing the two sequences they count 45 synonymous and 5
nonsynonymous differences. Assume all of these are substitutions. They calculated that 80% of all mutations would be non-
synonymous and the remainder synonymous. The potential number of synonymous (S) and nonsynonymous (N) sites are then
obtained as S = 420x0.2 = 84 and N = 420x0.8 = 336. [6]
a) Calculate the proportion of synonymous (pS) and nonsynonymous differences (pN). (1)
b) For species separated by so many years, observed mutations will often affect the same codon, leaving uncertainty as to
which happened first. Depending on the sequence of mutations, the same mutation could be counted as nonsynonymous
or synonymous. We will ignore this complication here. Estimate the number of synonymous (dS) and nonsynonymous (dN)
substitutions per site using the Jukes-Cantor correction as:
3 4𝑝𝑝 3 4𝑝𝑝
𝑑𝑑𝑆𝑆 = − 𝑙𝑙𝑙𝑙 �1 − 𝑆𝑆�3� 𝑎𝑎𝑎𝑎𝑎𝑎 𝑑𝑑𝑁𝑁 = − 𝑙𝑙𝑙𝑙 �1 − 𝑁𝑁�3�
4 4
Note that this is now differences per site, rather than per gene. (1)
c) What is the rate of synonymous and nonsynonymous substitutions? Use the equations:
𝑑𝑑𝑆𝑆 𝑑𝑑𝑁𝑁
𝑟𝑟𝑆𝑆 = 𝑎𝑎𝑎𝑎𝑎𝑎 𝑟𝑟𝑁𝑁 =
2𝑇𝑇 2𝑇𝑇
which are the same as equation 8.3 except that we have already divided by the number of nucleotides (L - 84 and 336,
respectively) when we calculated pS and pN and doing so again would be incorrect. (1)
d) Explain why this correction makes a large difference to synonymous, but not to nonsynonymous differences per
nucleotide. (2)
e) If we assume that the nonsynonymous mutations can be divided into two classes: a) having no fitness effect and b) being
lethal, what percentage of nonsynonymous mutations can be considered lethal? You have to assume that the
synonymous rate = neutral mutation rate, and then use rN = (1-α)µ (1)
4. In a study containing a set of Adh sequences in Drosophila melanogaster, 100 synonymous substitutions and 10
nonsynonymous substitutions were observed. The authors concluded that negative selection must be at play, and most
nonsynonymous mutations are removed from the population. However, when the authors examined fixed differences between D.
melanogaster and either D. simulans or D. yakuba they found that 15 of the fixed differences between these species were
nonsynonymous, and 45 were synonymous. [10]
a) What is the null hypothesis in this case? (1)
b) What neutrality test should be applied in this case and why? (2)
c) Using the table below, estimate the ratios of the neutrality test (2)
Nonsynonymous (N) Synonymous (S)
Within species (S)
Between species (F)
d) Calculate if these results are significant using a 𝜒𝜒 2 test, with one degree of freedom for p < 0.01 (already corrected,
critical value is 6.635). (5)
Hint: remember how we calculated contingency tables for the GWAS exercise in Tutorial 4