Chapter 5
Genome
Sequences and
Evolution
Photo of intracellular
bacterium courtesy of
Gregory P. Henderson and
Courtesy of Keith Grant J. Jensen, California
Weller/USDA. Institute of Technology.
University College of Medicine.
Courtesy of Eishi Noguchi, Drexel
© Photodisc.
5.1 Introduction
of organism increases with its complexity.
Figure 5.1: The minimum gene number required for any type
Courtesy of Carolyn B. Marks Courtesy of Rocky Mountain
and David H. Hall, Albert Laboratories, NIAID, NIH
Einstein College of Medicine,
Bronx, NY.
5.2 Prokaryotic Gene Numbers Range
Over an Order of Magnitude
• The minimum
number of
genes for a
parasitic
prokaryote is
about 500; for a
free-living
nonparasitic
prokaryote, it is
about 1,500.
Table 5.1: Genome sizes and
gene numbers are known
from complete sequences
for several organisms.
Figure 5.2: The number of genes in bacterial and
archaeal genomes is proportional to genome size.
5.2 Prokaryotic Gene Numbers Range
Over an Order of Magnitude
• pathogenicity islands – DNA segments that are
present in pathogenic bacterial genomes but absent
in their nonpathogenic relatives.
• horizontal transfer – The transfer of DNA from one
cell to another by a process other than cell division,
such as bacterial conjugation.
5.3 Total Gene Number Is Known for
Several Eukaryotes
• There are 6000 genes in yeast; 21,700 in a
nematode worm; 17,000 in a fly; 25,000 in the small
plant Arabidopsis; and probably 20,000 to 25,000 in
mammals.
Figure 5.3: The number of genes
in a eukaryote varies from 6000–
32,000 but does not correlate
with genome size or the
organism complexity.
Figure 5.4: The S. cerevisiae genome of 13.5 Mb has 6000 genes, almost all
uninterrupted, while the Sc. pombe genome of 12.5 Mb has 5000 genes.
5.3 Total Gene Number Is Known for
Several Eukaryotes
• monocistronic mRNA – mRNA that encodes one
polypeptide.
• polycistronic mRNA – mRNA that includes coding
regions representing more than one gene.
5.3 Total Gene Number Is Known for
Several Eukaryotes
Figure 5.5: Functions of Drosophila genes based on
comparative genomics of twelve species.
Adapted from Drosophila 12 Genomes Consortium,
“Evolution of genes and genomes on the Drosophila
phylogeny,” Nature 450 (2007): 203–218.
5.4 How Many Different Types of
Genes Are There?
• The sum of the number of unique genes and the
number of gene families is an estimate of the
number of types of genes.
Figure 5.6: Many genes are
duplicated, and as a result
the number of different gene
families is much smaller than
the total number of genes.
5.4 How Many Different Types of
Genes Are There?
• orthologous genes
(orthologs) –
Related genes in
different species.
• The minimum size of
the proteome can be
estimated from the
number of types of
genes.
Figure 5.7: Fruit fly genome can be divided
into genes present in all eukaryotes, genes
present in all multicellular eukaryotes, genes
specific to flies.
5.5 The Human Genome Has Fewer
Genes Than Originally Expected
• Only 1% of the human
genome consists of
exons.
• The exons comprise about
5% of each gene, so
genes (exons plus introns)
comprise about 25% of
the genome.
• The human genome has
about 20,000 genes.
Figure 5.9: Genes occupy 25% of the human
genome, but protein-coding sequences are
only a small part of this fraction.
5.5 The Human Genome Has Fewer
Genes Than Originally Expected
Figure 5.10: The average human gene is 27 kb long and has 9 exons,
usually comprising 2 longer exons at each end and 7 internal exons.
5.5 The Human Genome Has Fewer
Genes Than Originally Expected
• Roughly 60% of human genes are alternatively
spliced.
• Up to 80% of the alternative splices change protein
sequence, so the human proteome has 50,000 to
60,000 members.
5.6 How Are Genes and Other
Sequences Distributed in the Genome?
• Repeated sequences (present in more than one copy)
account for more than 50% of the human genome.
• The great bulk of repeated sequences consist of copies
of nonfunctional transposons.
• There are many duplications of large chromosome
regions.
Figure 5.12: The largest
component of the human
genome consists of transposons.
5.7 The Y Chromosome Has Several
Male-Specific Genes
• The Y chromosome has about 60 genes that are
expressed specifically in the testis.
• The male-specific genes are present in multiple copies in
repeated chromosomal segments.
• Gene conversion between multiple copies allows the
active genes to be maintained during evolution.
Figure 5.13: The Y
chromosome consists of
X-transposed regions, X-
degenerate regions, and
amplicons.
5.8 How Many Genes Are Essential?
• Not all genes are essential. In
yeast and flies, deletions of
less than 50% of the genes
have detectable effects.
• When two or more genes are
redundant, a mutation in any
one of them might not have
detectable effects.
Figure 5.14: Essential yeast
genes are found in all classes.
5.8 How Many Genes Are Essential?
• We do not fully understand the persistence of genes that
are apparently dispensable in the genome.
Figure 5.15: A systematic analysis of loss of function for 86% of worm genes shows
that only 10% have detectable effects on the phenotype.
5.8 How Many Genes Are Essential?
• synthetic lethal – Two mutations that are viable by
themselves but cause lethality when combined.
• synthetic genetic array analysis (SGA) – An
automated technique in budding yeast whereby a
mutant is crossed to an array of approximately 5000
deletion mutants to determine if the mutations
interact to cause a synthetic lethal phenotype.
5.8 How Many Genes Are Essential?
Figure 5.16: The chart shows how many lethal
interacting genes there are for each test gene.
5.9 About 10,000 Genes Are Expressed at
Widely Differing Levels in a Eukaryotic Cell
• In any particular cell, most genes are expressed at
a low level.
• scarce (complex) mRNA – mRNA that consists of
a large number of individual mRNA species, each
present in very few copies per cell.
– This accounts for most of the sequence complexity
in RNA.
5.9 About 10,000 Genes Are Expressed at
Widely Differing Levels in a Eukaryotic Cell
• Only a small number of genes, whose products
are specialized for the cell type, are highly
expressed.
– abundance – The average number of mRNA
molecules per cell.
– abundant mRNA – Consists of a small number of
individual species, each present in a large number of
copies per cell.
5.9 About 10,000 Genes Are Expressed at
Widely Differing Levels in a Eukaryotic Cell
Figure 5.17: Hybridization between Figure 5.18: The abundances of yeast
excess mRNA and cDNA identifies mRNAs vary from less than 1 per cell to
components in chick oviduct cells, more than 100 per cell.
characterized by the Rot1/2 of reaction.
5.9 About 10,000 Genes Are Expressed at
Widely Differing Levels in a Eukaryotic Cell
• mRNAs expressed at low levels overlap extensively
when different cell types are compared.
– housekeeping gene – A gene that is (theoretically)
expressed in all cells because it provides basic
functions needed for sustenance of all cell types.
5.9 About 10,000 Genes Are Expressed at
Widely Differing Levels in a Eukaryotic Cell
• The abundantly expressed mRNAs are usually
specific for the cell type.
– luxury gene – A gene encoding a specialized function
(usually) synthesized in large amounts in particular
cell types.
• About 10,000 expressed genes might be common to
most cell types of a multicellular eukaryote.
5.10 Expressed Gene Number Can Be
Measured En Masse
• DNA microarray technology allows a snapshot to
be taken of the expression of the entire genome in
a yeast cell.
• About 75% (approximately 4,500 genes) of the
yeast genome is expressed under normal growth
conditions.
5.10 Expressed Gene Number Can Be
Measured En Masse
• DNA microarray technology allows for detailed comparisons
of related animal cells to determine (for example) the
differences in expression between a normal cell and a
cancer cell.
Figure 5.19: “Heat map” of 59
invasive breast tumors from
women who breastfed for at
least 6 months (red lines) or
who never breastfed (blue
lines).
Image courtesy of Rachel E. Ellsworth, Clinical Breast
Care Project, Windber Research Institute.
5.11 DNA Sequences Evolve by Mutation
and a Sorting Mechanism
• The probability of a mutation is influenced by the
likelihood that the particular error will occur and the
likelihood that it will be repaired.
• synonymous mutation – A change in DNA
sequence in a coding region that does not alter the
amino acid that is encoded.
• nonsynonymous mutation – A change in DNA
sequence in a coding region that alters the amino
acid that is encoded.
5.11 DNA Sequences Evolve by Mutation
and a Sorting Mechanism
• In small populations, the frequency of a mutation will
change randomly and new mutations are likely to be
eliminated by chance.
• fixation – The process by which a new allele
replaces the allele that was previously predominant
in a population.
5.11 DNA Sequences Evolve by Mutation
and a Sorting Mechanism
• The frequency of a neutral mutation largely depends on
genetic drift, the strength of which depends on the size
of the population.
• The frequency of a mutation that affects phenotype will
be influenced by negative or positive selection.
Figure 5.21A: The fixation or loss of alleles by Figure 5.21B: The fixation or loss of alleles by
random genetic drift in populations of 10. random genetic drift in populations of 100.
Data courtesy of Kent E. Holsinger,
University of Connecticut
[[Link]
5.12 Selection Can Be Detected by
Measuring Sequence Variation
• The ratio of nonsynonymous to synonymous
substitutions in the evolutionary history of a gene is
a measure of positive or negative selection.
• Low heterozygosity of a gene might indicate recent
selective events.
• genetic hitchhiking – The change in frequency of
a genetic variant due to its linkage to a selected
variant at another locus.
5.12 Selection Can Be Detected by
Measuring Sequence Variation
• Comparing the rates of substitution among related
species can indicate whether selection on the gene has
occurred.
• linkage disequilibrium – A nonrandom association
between alleles at two different loci, often as a result of
linkage.
Figure 5.22: A higher
number of nonsynonymous
substitutions in lysozyme
sequences in the cow/deer
lineage as compared to the
pig lineage.
Adapted from N. H. Barton, et al. Evolution. Cold Spring Harbor
Laboratory Press, 2007. Original figure appeared in J. H.
Gillespie, The Causes of Molecular Evolution. Oxford University
Press, 1991.
5.12 Selection Can Be Detected by
Measuring Sequence Variation
Figure 5.23: The fraction of recombinants between an allele of G6PD and
alleles at nearby loci on a human chromosome remains low.
Adapted from E. T. Wang, et al., Proc. Natl.
Acad. Sci. USA 103 (2006): 135-140.
5.12 Selection Can Be Detected by
Measuring Sequence Variation
• Most functional genetic variation in the human
species affects gene regulation and not variation in
proteins.
5.13 A Constant Rate of Sequence
Divergence Is a Molecular Clock
• The sequences of orthologous genes in different
species vary at nonsynonymous sites (where
mutations have caused amino acid substitutions)
and synonymous sites (where mutation has not
affected the amino acid sequence).
• Synonymous substitutions accumulate about 10
times faster than nonsynonymous substitutions.
5.13 A Constant Rate of Sequence
Divergence Is a Molecular Clock
• The evolutionary divergence between two DNA
sequences is measured by the corrected
percentage of positions at which the
corresponding nucleotides differ.
• Substitutions can accumulate at a more or less
constant rate after genes separate, so that the
divergence between any pair of globin sequences
is proportional to the time since they shared
common ancestry.
5.13 A Constant Rate of Sequence
Divergence Is a Molecular Clock
Figure 5.24: Divergence
of DNA sequences
depends on evolutionary
separation. Each point on
the graph represents a
pairwise comparison.
5.13 A Constant Rate of
Sequence Divergence Is
a Molecular Clock
Figure 5.25: All globin genes have evolved by a
series of duplications, transpositions, and
mutations from a single ancestral gene.
5.13 A Constant Rate of Sequence
Divergence Is a Molecular Clock
Figure 5.26: Nonsynonymous
site divergences between pairs
of β-globin genes allow the
history of the human cluster to
be reconstructed.
5.13 A Constant Rate of Sequence
Divergence Is a Molecular Clock
• codon bias – A higher usage of one codon in genes
to encode amino acids for which there are several
synonymous codons.
5.14 The Rate of Neutral Substitution
Can Be Measured from Divergence of
Repeated Sequences
• The rate of substitution per year at neutral sites is
greater in the mouse genome than in the human
genome, probably because of a higher mutation rate.
Figure 5.27: An ancestral
consensus sequence for a
family is calculated by
taking the most common
base at each position.
5.15 How Did Interrupted Genes Evolve?
• An interesting evolutionary question is whether
genes originated with introns or were originally
uninterrupted.
• “introns late” model – The hypothesis that the
earliest genes did not contain introns, and that
introns were subsequently added to some genes.
5.15 How Did Interrupted Genes Evolve?
• Interrupted genes that correspond either to proteins
or to independently functioning noncoding RNAs
probably originated in an interrupted form (“introns
early” hypothesis).
• exon shuffling – The hypothesis that genes have
evolved by the recombination of various exons
encoding functional protein domains.
Figure 5.28: An exon surrounded by flanking sequences that is
translocated into an intron can be spliced into the RNA product.
5.15 How Did Interrupted Genes Evolve?
• The interruption allowed base order to better
satisfy the potential for stem–loop extrusion from
duplex DNA, perhaps to facilitate recombination
repair of errors.
• A special class of introns is mobile and can insert
themselves into genes.
5.16 Why Are Some Genomes So Large?
• There is no clear correlation
between genome size and
genetic complexity.
• C-value – The total amount
of DNA in the genome (per
haploid set of
chromosomes)
• C-value paradox – The
lack of relationship
between the DNA content
(C-value) of an organism Figure 5.29: DNA content of the haploid
genome increases with morphological
and its coding potential. complexity of lower eukaryotes.
5.16 Why Are Some Genomes So Large?
• There is an increase in
the minimum genome size
associated with organisms
of increasing complexity.
• There are wide variations
in the genome sizes of
organisms within many
taxonomic groups.
Figure 5.30: The minimum genome size
found in each taxonomic group increases
from prokaryotes to mammals.
5.17 Morphological Complexity Evolves by
Adding New Gene Functions
• In general, comparisons of eukaryotes to
prokaryotes, multicellular to unicellular eukaryotes,
and vertebrate to invertebrate animals show a
positive correlation between gene number and
morphological complexity as additional genes are
needed with generally increased complexity.
• Most of the genes that are unique to vertebrates are
concerned with the immune or nervous systems.
5.17 Morphological Complexity Evolves by
Adding New Gene Functions
Figure 5.31: Human genes can be classified according to how widely
their homologs are distributed in other species.
5.17 Morphological Complexity Evolves by
Adding New Gene Functions
Figure 5.32: Common eukaryotic proteins are
involved with essential cellular functions.
Figure 5.33: Increasing complexity in eukaryotes is accompanied by accumulation
of new proteins for transmembrane and extracellular functions.
5.18 Gene Duplication Contributes to
Genome Evolution
• Duplicated genes can diverge to generate different
genes, or one copy might become an inactive
pseudogene.
5.18 Gene Duplication Contributes to
Genome Evolution
Figure 5.34: After a globin gene
has been duplicated,
differences can accumulate
between the copies.
5.19 Globin Clusters Arise by
Duplication and Divergence
• All globin genes are descended
by duplication and mutation from
an ancestral gene that had three
exons.
• The ancestral gene gave rise to
myoglobin, leghemoglobin, and
α- and β-globins.
Figure 5.35: Each of the α-like
• The α- and β-globin genes and β-like globin gene families is
separated in the period of organized into a single cluster,
early vertebrate evolution, which includes functional genes
after which duplications and pseudogenes (ψ).
generated the individual
clusters of separate α- and
β-like genes.
5.19 Globin Clusters Arise by
Duplication and Divergence
• nonallelic genes – Two (or more) copies of the
same gene that are present at different locations in
the genome (contrasted with alleles, which are
copies of the same gene derived from different
parents and present at the same location on the
homologous chromosomes).
• When a gene has been inactivated by mutation, it
can accumulate further mutations and become a
pseudogene (ψ), which is homologous to the
functional gene(s) but has no functional role (or at
least has lost its original function).
5.19 Globin Clusters Arise by
Duplication and Divergence
Figure 5.36: Different hemoglobin genes are expressed during
embryonic, fetal, and adult periods of human development.
5.19 Globin Clusters Arise by
Duplication and Divergence
Figure 5.37: Clusters of β-globin genes and pseudogenes are found in vertebrates.
5.20 Pseudogenes Have Lost
Their Original Functions
• Processed pseudogenes result from reverse
transcription and integration of mRNA transcripts.
• Nonprocessed pseudogenes result from incomplete
duplication or second-copy mutation of functional genes.
• Some pseudogenes
might gain functions
different from those
of their parent genes,
such as regulation of
gene expression, and
take on different names.
Figure 5.38: Many changes have occurred in a
β-globin gene since it became a pseudogene.
5.20 Pseudogenes Have Lost
Their Original Functions
Table 5.6: Most human RP pseudogenes are of recent origin; many are
shared with the chimpanzee but absent from rodents.
Adapted from S. Balasubramanian, et al.,
Genome Biol. 20 (2009): R2.
5.21 Genome Duplication Has Played a
Role in Plant and Vertebrate Evolution
• Genome duplication occurs when polyploidization
increases the chromosome number by a multiple of
two.
• autopolyploidy – Polyploidization resulting from
mitotic or meiotic errors within a species.
• allopolyploidy – Polyploidization resulting from
hybridization between two different but
reproductively compatible species.
5.21 Genome Duplication Has Played a
Role in Plant and Vertebrate Evolution
• Genome duplication events can be obscured by the
evolution and/or loss of duplicates as well as by
chromosome rearrangements.
• Genome duplication has been detected in the
evolutionary history of many flowering plants and of
vertebrate animals.
• 2R hypothesis – The hypothesis that the early
vertebrate genome underwent two rounds of
duplication.
5.21 Genome Duplication Has Played a
Role in Plant and Vertebrate Evolution
Figure 5.39: At left, a constant rate of gene duplication and loss shows an
exponentially decreasing age distribution of duplicated gene pairs.
Adapted from G. Blanc and K. H. Wolfe,
Plant Cell 16 (2004): 1667-1678.
5.22 What Is the Role of Transposable
Elements in Genome Evolution?
• Transposable elements tend to increase in copy
number when introduced to a genome but are kept
in check by negative selection and transposition
regulation mechanisms.
5.23 There Can Be Biases in Mutation,
Gene Conversion, and Codon Usage
• Mutational bias can account for a high AT content in
organismal genomes.
• Gene conversion bias, which tends to increase GC
content, can act in partial opposition to the
mutational bias.
• Codon bias might be a result of adaptive
mechanisms that favor particular sequences, and of
gene conversion bias.