MZO-004
SYSTEMATICS, BIODIVERSITY
Indira Gandhi AND EVOLUTION
National Open University
School of Sciences
Block
2
EVOLUTIONARY BIOLOGY-II
UNIT 3
The Three Domain System 49
UNIT 4
Molecular Phylogeny 62
UNIT 5
Molecular Clocks 91
UNIT 6
Divergence of Bacterial and Archaeal Genome 104
UNIT 7
Bacterial Genome Evolution 142
COURSE NAME: SYSTEMATICS, BIODIVERSITY AND EVOLUTION COURSE CODE: MZO-004
Course Design Committee
Prof. Sujatha Varma Prof. Bharati Chauhan Prof. Amrita Nigam (Retd.)
Former Director, School of Department of Zoology School of Sciences, IGNOU
Sciences, IGNOU, Maidan Garhi University of Rajasthan Maidan Garhi, New Delhi-110068
New Delhi-110068 Jaipur-302004
Prof. Neera Kapoor
Prof. Neeta Sehgal Prof. Devinder Kaur Kochar School of Sciences, IGNOU
Department of Zoology Department of Zoology, Punjab Maidan Garhi, New Delhi-110068
University of Delhi, Delhi-110007 Agricultural University
Ludhiana-141004 Dr. Siya Ram
Prof. Sarita Kumar School of Sciences, IGNOU
Department of Zoology, Acharya Prof. Varsha Wankhade Maidan Garhi, New Delhi-110068
Narender Dev College, University Department of Zoology
of Delhi, Delhi-110007 Savitribai Phule University Dr. Ravi Rajwanshi
Pune-411007 School of Sciences, IGNOU
Prof. S. Dinakaran Prof. Sarita Sachdeva Maidan Garhi, New Delhi-110068
Department of Zoology Department of Biotechnology
Madurai Kamaraj University Manav Rachna International
Tamil Nadu-625021 University, Sector-43, Faridabad
Haryana-121004
Block Preparation Team
Dr. Anjali S. Nawani Dr. Neelam Gandhi School of Sciences
Department of Zoology, Department of Zoology,
Dr. Ravi Rajwanshi
Sri Venkateswara College, Hansraj College, University
School of Sciences, IGNOU
South Campus,University of Delhi, of Delhi, Delhi, 110007
Maidan Garhi, New Delhi-
Delhi 110021 (Unit-3) (Units- 6 and 7)
110068 (Units-3, 4, 5, 6 and 7)
Dr. Jyoti Taneja,
Department of Zoology,
Daulat Ram College, University of Delhi,
Delhi 110007 (Units-4 and 5)
Programme Coordinators : Prof. Neera Kapoor and Dr. Siya Ram
Course Coordinators : Dr. Ravi Rajwanshi
Course Editor : Dr. Tapati Das,
Associate Professor,
Department of Ecology and Environmental Sciences,
Assam University, Silchar 788011, Assam
Production : Mr. Tilak Raj
AR, MPDD, IGNOU
Acknowledgement:
• Dr. Ravi Rajwanshi and Mr. Ajit Kumar - Suggestions for figures and cover design.
• Mr. Manoj Kumar, Assistant for word processing and CRC preparation.
August, 2024
Indira Gandhi National Open University, 2024
ISBN: ..................................
All rights reserved. No part of this work may be reproduced in any form, by mimeograph or any other
means, without permission in writing from Indira Gandhi National Open University.
Further information on Indira Gandhi National Open University courses may be obtained from the
University’s office at Maidan Garhi, New Delhi-110 068 or IGNOU website [Link].
Printed and published on behalf of Indira Gandhi National Open University, New Delhi by the Registrar,
MPDD, IGNOU.
Printed at:
Unit 3 The Three Domain System
UNIT 3
THE THREE DOMAIN SYSTEM
Structure
3.1 Introduction 3.4 From Five Domain to
Three Domain
Objectives The Role of Horizontal
3.2 Linking Classification and gene Transfer
Phylogeny 3.5 Summary
3.3 Methods of Inferring 3.6 Terminal Questions
Evolutionary Relationship
3.7 Answers
3.1 INTRODUCTION
In the previous units of the block you have learnt about the idea of evolution
and the origin of life. In the present Unit, you are going to learn about how
biologists distinguish and cateogorize the millions of species on Earth. You will
learn how understanding has been developed regarding the evolutionary
relationships among different species. Additionally, you will discover various
methods for classifying a species based on similarities between its
characteristics and those of its close relatives.
In this unit, we will survey the diversity among different species and describe
hypotheses regarding how it evolved. As we do so, our emphasis will shift
from the process of evolution (the evolutionary mechanisms described in later
units) to its pattern (observations of evolution’s products over time).
To set the stage for surveying life’s diversity, this unit explains tracing of
phylogeny (the evolutionary history of a species or group of species) by the
biologists. As we will see, biologists reconstruct and interpret phylogenies
using systematic (a discipline that focus on classifying the organisms and
determining their evolutionary relationships).
You will also study about the concept of the universal tree of life in this Unit,
which was developed using systematics and phylogenies. You will also
discover how there are three basic domains of life forms into which all species
can be classified.
This is an introductory unit of the block 2. You will learn details of all the topics
covered here in the next few units of the block.
49
Block 2 Evolutionary Biology-II
Objectives
After studying this unit, you would be able to:
explain the difference between the previously described method of
classification and the concept of phylogentic trees;
elaborate the concept of Universal tree of life; and
comprehend the concept of three major domains of life forms.
3.2 LINKING CLASSIFICATION AND
PHYLOGENY
Organisms share many characteristics because of common ancestry. As a
result, we can learn a great deal about a species if we know its evolutionary
history. For example, an organism is likely to share many of its genes,
metabolic pathways, and structural proteins with its close relatives. We will
consider practical applications of such information in the later units, but first we
will look at how we can interpret and use diagrams that represent evolutionary
history.
You have already learnt that by 18th century Carolus Linnaeus instituited the
binomial nomenclature for naming individual species. In addition to naming
species, Linnaeus also grouped them into a hierarchy of increasingly inclusive
categories. The first grouping is built into the binomial: Species that appear to
be closely related are grouped into the same genus. For example, the leopard
(Panthera pardus) belongs to a genus that also includes the African lion
(Panthera leo), the tiger (Panthera tigris), and the jaguar (Panthera onca).
Beyond genera, taxonomists employ progressively more comprehensive
categories of classification. The taxonomic system named after Linnaeus, the
Linnaean system, places related genera in the same family, families into
orders, orders into classes, classes into phyla (singular, phylum), phyla into
kingdoms, and, more recently, kingdoms into domains (Fig. 3.1). The resulting
biological classification of a particular organism is somewhat like a postal
address identifying a person in a particular apartment, in a building with many
apartments, on a street with many apartment buildings, in a city with many
streets, and so on.
However, characters that are useful for classifying one group of organisms
may not be appropriate for other organisms. For this reason, the larger
categories often are not comparable between lineages; that is, an order of
snails does not exhibit the same degree of morphological or genetic diversity
as an order of mammals. Furthermore, as we’ll see, the placement of species
into orders, classes, and so on, does not necessarily reflect evolutionary
history.
50
Unit 3 The Three Domain System
Fig. 3.1: Linnaean classification. At each level, or “rank,” species are placed in
groups within more inclusive groups.
The evolutionary history of a group of organisms can be represented in a
branching diagram called a phylogenetic tree. As in Figure 3.2, the branching
pattern often matches with the classification done by the taxonomists for the
groups of organisms nested within more inclusive groups.
There are two main difficulties in aligning Linnaean classification with
phylogeny that have led some systematists to propose that classification be
based entirely on evolutionary relationships.
1. Over the course of evolution, a species which has lost a key feature
shared by its close relatives has been misclassified in Linnaean system.
The DNA or other new evidence indicates that an organism should be
reclassified in order to accurately reflect its evolutionary history.
51
Block 2 Evolutionary Biology-II
2. Another issue is that while the Linnaean system may distinguish groups,
such as amphibians, mammals, reptiles, and other classes of
vertebrates, it tells us nothing about these groups’ evolutionary
relationships to one another.
Regardless of how groups are named, a phylogenetic tree represents a
hypothesis about evolutionary relationships. These relationships often are
depicted as a series of dichotomies, or two-way branch points. Each branch
point represents the divergence of two evolutionary lineages from a common
ancestor. In Fig. 3.2, for example, branch point 3 represents the common
ancestor of taxa A, B, and C. The position of branch point 4 to the right of 3
indicates that taxa B and C diverged after their shared lineage split from the
lineage leading to taxon A. Note also that tree branches can be rotated around
a branch point without changing their evolutionary relationships.
In Fig. 3.2, taxa B and C are sister taxa, groups of organisms that share an
immediate common ancestor (branch point 4) and hence are each other’s
closest relatives. In addition, this tree, like most of the phylogenetic trees in
this book, is rooted, which means that a branch point within the tree (often
drawn farthest to the left) represents the most recent common ancestor of all
taxa in the tree. The term basal taxon refers to a lineage that diverges early in
the history of a group and hence, like taxon G in Figure 3.2, lies on a branch
that originates near the common ancestor of the group. Finally, the lineage
leading to taxa D–F includes a polytomy, a branch point from which more than
two descendant groups emerge. A polytomy signifies that evolutionary
relationships among the taxa are not yet clear.
Fig. 3.2: Reading a phylogenetic tree.
52
Unit 3 The Three Domain System
Let’s summarize three key points about phylogenetic trees.
Key Point Explanation Example
Phylogenetic trees are Although closely related Even though crocodiles
intended to show organisms often resemble phylogenetically are more
patterns of descent, not one another due to their closely related to birds than
phenotypic similarity. common ancestry, they to lizards, they look more
may not if their lineages like lizards because
have evolved at different morphology has changed
rates or faced very dramatically in the bird
different environmental lineage.
conditions.
The sequence of Generally, unless given The tree in Figure 3.3 does
branching in a tree does specific information about not indicate that the wolf
not necessarily indicate what the branch lengths evolved more recently than
the actual (absolute) in a phylogenetic tree the European otter; rather,
ages of the particular mean—for example, that the tree shows only that the
species. they are proportional to most recent com- mon
time—you should ancestor of the wolf and
interpret the diagram otter (branch point 1) lived
solely in terms of patterns before the most recent
of descent. No common ancestor of the
assumptions should be wolf and coyote (2). To
made about when par- indicate when wolves and
ticular species evolved or otters evolved, the tree
how much change would need to include
occurred in each lineage. additional divergences in
each evolutionary lineage,
as well as the dates when
those splits occurred.
We should not assume Figure 3.3 does not indicate
that a taxon on a that wolves evolved from
phylogenetic tree coyotes or vice versa. We
evolved from the taxon can infer only that the
next to it. lineage leading to wolves
and the lineage leading to
coyotes both evolved from
the common ancestor 2.
That ancestor, which is now
extinct, was neither a wolf
nor a coyote. However, its
descendants include the
two extant (living) species
shown here, wolves and
coyotes.
53
Block 2 Evolutionary Biology-II
Fig. 3.3: The link between Classification and Phylogeny.
SAQ 1
State True or False:
a) The Linnaean system of classification can not tell us about evolutionary
history of a species.
b) The sequence of branching in a tree does not necessarily indicate when
particular species evolved or how much change occurred in each
lineage.
c) We should not assume that a taxon on a phylogenetic tree evolved from
the taxon next to it.
3.2 METHODS OF INFERRING EVOLUTIONARY
DATA
There are many ways in which researchers construct trees like those we’ve
considered in the above section. The detail of these methods are given in next
54 few units of this block. Here we will review the methods briefly.
Unit 3 The Three Domain System
To infer phylogeny, systematists must gather as much information as possible
about the morphology, genes, and biochemistry of the relevant organisms. It is
important to focus on features that result from common ancestry, because only
such features reflect evolutionary relationships.
Analyses of molecular and morphological homologies: The phenotypic
and genetic similarities due to shared ancestry are called homologies. For
example, the similarity in the number and arrangement of bones in the
forelimbs of mammals is due to their descent from a common ancestor with
the same bone structure; this is an example of a morphological homology. In
the same way, genes or other DNA sequences are homologous if they are
descended from sequences carried by a common ancestor.
Homologies are used to create phylogenetic trees: A widely used set of
methods of inferring phylogeny from these homologous characters is known
as cladistics. Using this methodology, biologists attempt to place species into
groups called clades, each of which includes an ancestral species and all of its
descendants. You will study these methods in detail in next few units.
Maximum Parsimony and Maximum Likelihood: As the database of DNA
sequences that enables us to study more species, the difficulty of building the
phylogenetic tree that best describes their evolutionary history also grows.
What if you are analyzing data for 50 species? There are 3x1076 different ways
to arrange 50 species into a tree and which tree in this huge forest reflects the
true phylogeny? Systematists can never be sure of finding the most accurate
tree in such a large data set, but they can narrow the possibilities by applying
the principles of maximum parsimony and maximum likelihood.
Phylogenetic trees as hypothesis: This is a good place to reiterate that any
phylogenetic tree represents a hypothesis about how the organisms in the tree
are related to one another. The best hypothesis is the one that best fits all the
available data. A phylogenetic hypothesis may be modified when new
evidence compels systematists to revise their trees. Indeed, while many older
phylogenetic hypotheses have been supported by new morphological and
molecular data, others have been changed or rejected.
3.4 FROM FIVE DOMAIN TO THREE DOMAIN
CLASSIFICATION
In recent decades, systematists have gained insight into the very deepest
branches of the tree of life by analyzing DNA sequence data.
Initially taxonomists classified all known species into two kingdoms: plants and
animals. Classification schemes with more than two kingdoms gained broad
acceptance in the late 1960s, when many biologists recognized five kingdoms:
Monera (prokaryotes), Protista (a diverse kingdom consisting mostly of
unicellular organisms), Plantae, Fungi, and Animalia. This system highlighted
the two fundamentally different types of cells, prokaryotic and eukaryotic, and
set the prokaryotes apart from all eukaryotes by placing them in their own
kingdom, Monera.
55
Block 2 Evolutionary Biology-II
However, phylogenies based on genetic data soon began to reveal a problem
with this system: Some prokaryotes differ as much from each other as they do
from eukaryotes. Such difficulties have led biologists to adopt a three-domain
system. A domain is a larger, more inclusive category than a kingdom. Under
this system, there are the three domains - Bacteria, Archaea, and Eukarya.
The validity of these domains is supported by many studies, including a recent
study that analyzed nearly 100 completely sequenced genomes.
The domain Bacteria contains most of the currently known prokaryotes, while
the domain Archaea consists of a diverse group of prokaryotic organisms that
inhabit a wide variety of environments. The domain Eukarya consists of all the
organisms that have cells containing true nuclei. This domain includes many
groups of single-celled organisms as well as multicellular organisms such as
plants, fungi, and animals.
The three-domain system highlights the fact that much of the history of life has
been about single-celled organisms. The two prokaryotic domains consist
entirely of single-celled organisms, and even in Eukarya, only the branches
labeled in blue type (land plants, fungi, and animals) are dominated by
multicellular organisms. Of the five kingdoms previously recognized by
taxonomists, most biologists continue to recognize Plantae, Fungi, and
Animalia, but
not Monera and Protista. The kingdom Monera is obsolete because it would
have members in two different domains. The kingdom Protista has also
crumbled because it includes members that are more closely related to plants,
fungi, or animals than to other protists.
3.4.1 The Important Role of Horizontal Gene Transfer
In the phylogeny shown in Figure 3.4, the first major split in the history of life
occurred when bacteria diverged from other organisms. If this tree is correct,
eukaryotes and archaea are more closely related to each other than either is
to bacteria.
The phylogenetic tree represented in figure 3.4 is based on sequence data for
rRNA and other genes. For simplicity, only some of the major branches in
each domain are shown. Lineages within Eukarya that are dominated by
multicellular organisms (land plants, fungi, and animals) are in blue type, while
the two lineages denoted by an asterisk are based on DNA from cellular
organelles. All other lineages consist solely or mainly of single-celled
organisms.
56
Unit 3 The Three Domain System
Fig. 3.4: The three domains of life.
As mentioned above this reconstruction of the tree of life is based in part on
sequence comparisons of rRNA genes, which code for the RNA components
of ribosomes. However, some other genes reveal a different set of
relationships. For example, researchers have found that many of the genes
that influence metabolism in yeast (a unicellular eukaryote) are more similar to
genes in the domain Bacteria than they are to genes in the domain Archaea -
a finding that suggests that the eukaryotes may share a more recent common
ancestor with bacteria than with archaea.
What causes trees based on data from different genes to yield such different
results? Comparisons of complete genomes from the three domains show that
there have been substantial movements of genes between organisms in the
different domains. These took place through horizontal gene transfer, a
process in which genes are transferred from one genome to another through
57
Block 2 Evolutionary Biology-II
mechanisms such as exchange of transposable elements and plasmids, viral
infection and perhaps
perhaps fusions of organisms (as when a host and its
endosymbiont become a single organism). Recent research reinforces the
view that horizontal gene transfer is important. For example, a 2008 analysis
indicated that, on average, 80% of the genes in 181 prokaryotic genomes had
moved between species at some point during the course of evolution. As
phylogenetic trees are based on the assumption that genes are passed
vertically from one generation to the next, the occurrence of such horizontal
transfer events helps
helps to explain why trees built using different genes can give
inconsistent results.
Horizontal gene transfer can also occur between eukaryotes. For example,
over 200 cases of the horizontal transfer of transposons have been reported in
eukaryotes, including humans and other primates, plants, birds, and the gecko
shown in Figure 3.5. Recent genetic evidence indicates that this gecko is one
of 17 reptile species that acquired the transposon SPIN as a result of
horizontal gene transfer. The transposon may have been transferred from
one species to another by the feeding activities of blood-sucking insects.
Fig. 3.5: A recipient of transferred genes: the Mediterranean house gecko
(Hemidactylus turcicus).
Nuclear genes have also been transferred horizontally from one eukaryote to
another. The Scientific Skills exercise describes one such example, giving
you the opportunity to interpret data on the transfer of a pigment gene to an
aphid from another species.
58
Unit 3 The Three Domain System
Fig. 3.6: A tangled web of life. Horizontal gene transfer may have been so
common in the early history of life that the base of a “tree of life” might
be more accurately portrayed as a tangled web.
Overall, horizontal gene transfer has played a key role throughout the
evolutionary history of life and it continues to occur today. Some biologists
have argued that horizontal gene transfer was so common that the early
history of life should be represented not as a dichotomously branching tree like
that in Figure 3.4, but rather as a tangled network of connected branches (Fig.
3.6). Although scientists continue to debate whether early steps in the history
of life are best represented as a tree or a tangled web, in recent decades there
have been many exciting discoveries about evolutionary events that occurred
over time. You will explore such discoveries in the rest of this unit’s, beginning
with Earth’s earliest inhabitants, the prokaryotes.
This is our best current understanding of some of the high plots in our history,
in which, metaphorically speaking, the human species developed as one twig
in a gigantic tree, the great Tree of Life (Fig. 3.7). A species is likely to
"branch" throughout time, giving rise to two new species with distinct
evolutionary modifications of some of its traits. This process may also
continue for the descendants of the original species, who may continue to
evolve. Over the span of many millions of years, numerous branching and
modification events have led to the evolution of millions of distinct types of
species from a single ancestral organism at the very base, or root, of the tree.
59
Block 2 Evolutionary Biology-II
Fig. 3.7: The great Tree of Life. Represents three main domain Archaea, Bacteria
and Eukarya. Human beings is onle one twig of this tree of life.
SAQ 2
a) What is horizontal gene transfer? How is it different from vertical gene
transfer?
b) What could be different methods of horizontal gene tranfer?
c) Are the trees based on data from different genes to yield different
results? Why?
3.5 SUMMARY
• Linnaeus’s binomial classification system gives organisms two-part
names: a genus plus a specific epithet.
• In the Linnaean system, species are grouped in increasingly broad taxa:
related genera are placed in the same family, families in orders, orders
in classes, classes in phyla, phyla in kingdoms, and (more recently)
kingdoms in domains.
• Systematists depict evolutionary relationships as branching phylogenetic
trees. Many systematists propose that classification be based entirely on
evolutionary relationships.
• Unless branch lengths are proportional to time or genetic change, a
phylogenetic tree indicates only patterns of descent. Much information
can be learned about a species from its evolutionary history; hence,
phylogenies are useful in a wide range of applications.
• Organisms with similar morphologies or DNA sequences are likely to be
more closely related than organisms with very different structures and
60 genetic sequences.
Unit 3 The Three Domain System
• A clade is a grouping that includes an ancestral species and all of its
descendants.
• Among phylogenies, the most parsimonious tree is the one that requires
the fewest evolutionary changes. The most likely tree is the one based
on the most likely pattern of changes. Well-supported phylogenetic
hypotheses are consistent with a wide range of data.
• Past classification systems have given way to the current view of the
tree of life, which consists of three great domains: Bacteria, Archaea,
and Eukarya.
• Phylogenies based in part on rRNA genes suggest that eukaryotes are
most closely related to archaea, while data from some other genes
suggest a closer relationship to bacteria.
• Genetic analyses indicate that extensive horizontal gene transfer has
occurred throughout the evolutionary history of life.
3.5 TERMINAL QUESTIONS
1. Why classical taxonomy can not explain evolutionary relationship among
different species? Explain giving examples.
2. Write a brief note on how horizontal gene transfer giving different
examples of prokaryotes and eukaryotes.
3.6 ANSWERS
Self-Assessment Questions
1. a) True; b) True; c) True.
2. a) The Horizontal Gene Transfer is a process in which genes are
transferred from one genome to another. It is different from vertical
gene transfer where genes are transferred from one generation to
other generation within the same species/genome.
b) Different mechnisms of Horizontal gene transfer exist such as
exchange of transposable elements and plasmids, viral infection
and perhaps fusions of organisms (as when a host and its
endosymbiont become a single organism).
c) Yes. The trees based on data from different genes to yield different
phylogenetic trees. The phylogenetic trees are based on the
assumption that genes are passed vertically from one generation
to the next, the occurrence of horizontal transfer events explain the
different results.
Terminal Questions
1. Refer to Section 3.2.
2. Refer to Subsection 3.4.1.
61
Block 2 Evolutionary Biology-II
UNIT 4
MOLECULAR PHYLOGENCY
Structure
4.1 Introduction 4.6 Construction of
Phylogenetic Tree Using
Objectives
Molecular Data
4.2 History
4.7 Construction of
4.3 Definitions and Phylogenetic Tree by
Terminology of Using 16S rRNA Gene
Phylogenetic Tree Sequence
4.4 Representation and 4.8 Concept of Speciation in
Characteristics of Bacteria
Phylogenetic Tree
4.9 Summary
4.5 Limitations of Phylogenetic
4.10 Terminal Questions
Trees
4.11 Answers
4.1 INTRODUCTION
You have read about the types and the importance of variations in evolution.
The underlying mechanism of evolution is genetic mutations which play a very
important role in introducing variations in populations. The gradual
accumulation of mutations or differences in nucleotides and protein sequences
between genomes and proteomes of different species gives clues about the
biological diversity and pattern of genome evolution. This helps in inferring the
common ancestor and time of divergence of species. Furthermore, molecular
phylogenetics is an approach in systematics to study evolutionary
relationships among biological entities using molecular data, wherein
homologous nucleotide and protein sequence data of different species along
with DNA markers is used to study the evolutionary history of human and other
organisms. Molecular sequences provide large and unbiased comparable data
sets/evidences which can be quantified or statistically analysed to infer
phylogeny. The similarities and dissimilarities among related biological
sequences of genomes of different species provide a means to find out the
evolutionary relationships between them. The principle behind phylogenetic
study based on molecular sequence similarity is that the greater is sequence
62
Unit 4 Molecular Phylogency
similarity between two species, lesser are the variations and hence, they share
a recent common ancestor. Phylogenetic analysis is also used for
understanding the adaptive evolution and evolutionary pattern of multigene
family. It is possible to study the evolutionary relationships of the organisms at
varied levels of their classification because the rate of sequence evolution
changes with gene or genomic segment.
Objectives
After studying this unit, you would be able to:
understand the importance of molecular markers in phylogenetics;
describe the key features of a phylogenetic tree;
identify and use the terminology required to describe and interpret a
phylogenetic tree;
explain method of construction of phylogenetic tree using molecular data
and 16s rRNA gene sequence; and
elucidate concept of speciation in bacteria.
4.2 HISTORY
The phylogenetic tree traditionally was used to represent the evolutionary
relationship among species. However, in molecular phylogenetics, sequences
of homologous genes are aligned to construct a tree. Nuttall’s work in 1904 on
immunological assays showed that the molecular sequences could be used for
studying evolutionary relationships. Earlier, only morphological characters
were compared to know the similarities and dissimilarities among organisms.
Further, with the introduction of phenetics (distance based) and cladistics
(character based), the importance of large datasets for comparisons was
established.
Further, Zuckerkandl and Pauling observed that the change of the number of
amino acids of hemoglobin protein among two different species is proportional
to the divergence time of the species. Table 4.1 represents the alignment of
beta haemoglobin gene sequence among different species. The advantages of
using molecular sequence over morphological characters in molecular
phylogenetics are that the genes and amino acid sequences are heritable and
also describe the molecular characters. Molecular phylogenetics also allows
quantitative and homology analysis which can be used to reveal the distant
evolutionary relationships. In addition, molecular phylogenetics is important to
understand the insights into molecular evolution, and can be used to predict
the functions of genes and their diversification. It is also used for detecting
different regimes of selective pressures and epidemiology.
Table 4.1: Beta hemoglobin gene sequence among different organisms.
Organisms Beta Hemoglobin Gene Sequence
Human STPDAVMGNPKVKAHGKKVLGAFSDGLAHLDNLKGTFATLSELHC
DKLHVDPENFRL
63
Block 2 Evolutionary Biology-II
Gorilla STPDAVMGNPKVKAHGKKVLGAFSDGLAHLDNLKGTFATLSELHC
DKLHVDPENFKL
Horse SNPGAVMGNPKVKAHGKKVLHSFGEGVHHLDNLKGTFAALSELHC
DKLHVDPENFRL
Pig SNADAVMGNPKVKAHGKKVLQSFSDGLKHLDNLKGTFAKLSELHC
DQLHVDPENFR
Cow STADAVMNNPKVKAHGKKVLDSFSNGMKHLDDLKGTFAALSELHC
DKLHVDPENFKL
Deer SSAGAVMNNPKVKAHGKRVLDAFTQGLKHLDDLKGAFAQLSGLHC
NKLHVNPQNFRL
Gull SSPTAINGNPMVRAHGKKVLTSFGEAVKNLDNIKNTFAQLSELHCD
KLHVDPENFRL
4.3 DEFINITIONS
● Phylogeny: The term phylogeny has been derived from two Greek
words, phylon and genesis. Phylon means ‘stem’ and genesis means
‘origin’. It gives an idea about evolution or origin of an organism. In other
words, phylogeny is the evolutionary history of a species or group of
related species represented by tree branching patterns.
● Phylogenetics is the study of the evolutionary history of living
organisms or species or groups of related species using tree-like
diagrams to represent the pedigree of these organisms.
● Phylogenetic tree, also called an evolutionary tree, is a
representation to show the genetic relationship between species and
their hypothetical common ancestors i.e., give us the idea of an
evolutionary relationship between different species emerging in different
time durations. We try to see the order in which they share a more
recent common ancestor with one another. The phylogenetic tree is also
known as “Tree of Life” or Dendrogram. Theoretically, we can construct
the phylogenetic tree of all the living organisms.
● Molecular phylogenetics is a fundamental aspect of bioinformatics, the
new scientific discipline wherein homologous nucleotide and protein
sequence data of different species along with DNA markers such as
RFLPs, SSLPs, and SNPs are used to study the evolutionary history and
relationship of genes and other biological macromolecules by analysing
mutations at various positions in their sequences to give hypothesis
about evolutionary affinities.
● Homologous genes that evolved from the common ancestor.
● Orthologous genes that are evolved from common ancestors by
speciation event.
● Paralogous genes that evolved from common ancestor by duplication
event.
● Molecular markers are the sets of nucleotide (DNA, RNA) or amino
acid sequences that show polymorphism, such as nuclear ribosomal
64
Unit 4 Molecular Phylogency
genes, mitochondrial genes, and chloroplast genes that can be used to
determine/find and compare phylogenetic relationships between different
sequences/organisms.
● Speciation is defined as evolution of reproductive isolation among two
different populations resulting from either the by-product of evolutionary
divergence among populations or directly occurring due to
reinforcement.
4.4 Representation and Characteristics of
Phylogenetic Tree
Phylogenetic tree depicts evolutionary relationships among groups of
organisms, composed of nodes (represents taxonomic units, species,
population, individuals) and branches, which defines relationships between
taxonomic units in terms of descent and ancestry. The branching pattern of a
tree is called topology and branch length represents the number of changes
that have occurred in the branch. A typical phylogenetic tree is shown in figure
4.1. The lines in the tree are called branches. At the tips of the branches are
present-day species or sequences known as taxa (the singular form is taxon)
or each species or sequence represented at the tip of each branch of a
phylogenetic tree is also called an operational taxonomic unit (OTU). The
OTUs are the available nucleic acid or protein sequences that we are
analysing in a tree, sufficiently distinct from each other and treated as
separate units, termed as Terminal nodes (external and internal) or leaves.
The connecting point where two adjacent branches join is called a node (from
the Latin for “knot”), which represents an inferred ancestor of extant taxa. The
bifurcation point at the very bottom of the tree is the root node, which
represents the common ancestor of all members of the tree. In a phylogenetic
tree, each node with descendent represents the most common ancestor of the
descendants and the edge length in some trees corresponds to time
estimates. In figure 4.1, the tree consists of seven OTUs (labelled X, A, B, C,
D, E and F). These seven OTUs define seven external nodes/taxa/terminal
nodes. In addition, there are internal nodes at positions G, H, I and J. The
branches that connect internal nodes are known as internal branches (G-H, I-
J). The branches that connect internal nodes with external nodes are known
as external branches (e.g., A-G, E-J).
Fig. 4.1: Representation of phylogenetic tree and its parts- root, internal node,
branch, terminal node and outgroup. 65
Block 2 Evolutionary Biology-II
Outgroup, taxon or a group of taxa in a phylogenetic tree known to have
diverged earlier with maximum sequence dissimilarity than the rest of the taxa
in the tree and it is used to determine the position of the root (Fig. 4.1).
Forms of Tree Representation
Tree topology can be represented in different ways wherein branches of tree
freely rotate but without changing the relationship among taxa, such as
Phylogram or a cladogram.
Phylogram depicts the amount of evolutionary change (Y axis) that has
occurred along different branches of a phylogenetic tree in which branch
length is proportional to the amount of evolutionary divergence between
sequences and their ancestors. Branch length is scaled in phylogram, such
trees are said to be scaled and include a scale bar to indicate how much
evolutionary change is reflected in branch length (Fig. 4.2).
Fig. 4.2: Phylogram showing amount of change.
Phylogram is a Scaled tree in which the branch length is proportional to
number of substitutions or amount of change. Scaled trees indicate the
relationship between taxa, as well as information about relative divergence of
the branches thus conveying a sense of time or rate of evolution. The X-axis
represents distance and/or time (In millions of years) (Fig. 4.3, Panel a). In a
phylogram (additive tree) we can determine total evolutionary divergence
between two taxa on a phylogram by adding up the complete length of
branches separating them: from one taxa to the common ancestor and then
from the ancestor to the second taxa. The scaled trees have the advantage of
showing both the evolutionary relationships and information about the relative
divergence time of the branches.
Cladogram is an unscaled tree in which, the branch length is not proportional
to the number of evolutionary changes (amino acid or nucleotide) and thus
have no phylogenetic meaning, for example branch FA (two units) and FB
(one unit) have same apparent length, the external taxa OTUs (ABCDE) are
lined up neatly in a row or column at the tip of the tree (Fig. 4.3, Panel b).
66
Unit 4 Molecular Phylogency
Fig. 4.3: Representation of a) Scaled phylogenetic tree; and b) Unscaled
(Cladogram) phylogenetic tree.
In such a Cladogram, only the topology of the tree matters, which shows the
relative ordering of the taxa and relative recentness of a common ancestor. In
Cladistics trees are calculated by considering the possible pathways of
evolution and are based on parsimony and likelihood method. The resulting
tree is called a cladogram which is a branching diagram depicting hierarchical
arrangement of taxa. (Fig. 4.4, Panel a).
Dendrogram (Ultrametric tree): It depicts time of divergence. They are
branching diagrams used to represent degrees of relationship or resemblance
(Fig. 4.4, Panel b).
Fig. 4.4: Representation of phylogeny a) Cladogram; and (b) Dendrogram.
Newick Format: Tree topology is written in a special format called Newick
format given by Arthur Cayley. In this format trees are shown by taxa
separated by comma included in the linear form of nested parentheses. Each
internal node is represented by a pair of parentheses that enclose all members
of a monophyletic group. Newcik format for an unscaled tree is (((B,C,A), (D,E)
in figure 4.5, Panel a.
Fig. 4.5: A simple scaled tree represented in Newick format ((((A:2, B:3), C:5,
D:3), E:4).
67
Block 2 Evolutionary Biology-II
For a scaled tree, branch lengths are written immediately after the taxon name
separated by a colon. These numbers are arbitrary units that represent
divergent times (Fig. 4.5).
A phylogenetic tree can be either rooted or unrooted (Fig. 4.6). An unrooted
phylogenetic tree just arranges the taxa to demonstrate their relative
relationships; it does not reveal the identity of a common ancestor. It is the
timeless depiction of the branching relation between taxa. Since, there is no
indication of which node represents an ancestor, there is no direction of an
evolutionary path in an unrooted tree. It gives less information about
phylogenetic relations and shows no ancestral relation. To define the direction
of an evolutionary path, a tree must be rooted. In a rooted tree, all the
sequences under study have a common ancestor or root node from which a
unique evolutionary path leads to all other nodes. Obviously, a rooted tree is
more informative than an unrooted one. To convert an unrooted tree to a
rooted tree, one needs to first determine the origin of the root.
Fig. 4.6: Representation of a) rooted phylogenetic tree and b) unrooted
phylogenetic tree.
Sometimes, a branch point on a phylogenetic tree may have more than two
descendents, resulting in a multifurcating node. The phylogeny with
multifurcating branches is called polytomy (Fig. 4.6).
Fig. 4.7: Representations of polytomy in a phylogenetic tree.
The figure 4.7 represents rooted tree, which means that a branch point within
the tree (typically, the one farthest to the left) that is branch point b represents
the last common ancestor of all taxa (X-C) in the tree. Branch point (node) c
represents the common ancestor of taxa A, B, and C. Taxa Y and Z are sister
taxa, groups of organisms that share an immediate common ancestor (branch
point d) and hence are each other’s closest relatives. Finally, the lineage
68
Unit 4 Molecular Phylogency
leading to taxa A-C includes a polytomy, a branch point from which more than
two descendants groups emerge. A polytomy indicates that evolutionary
relationships among the descendant taxa are not yet clear.
Monophyletic/Paraphyletic/Polyphyletic
Monophyletic group includes the common ancestor (a node) and all its
descendants represented by both nodes and terminal taxa. A group of taxa
descended from a single common ancestor is defined as a clade or
monophyletic group. In monophyletic group, two taxa share a unique
common ancestor not shared by any other taxa. Clad ABC (Species A, B, C)
and Clad DE (Species D, E) are monophyletic (Fig. 4.8).
The branch path depicting an ancestor–descendant relationship on a tree is
called a lineage, which is often synonymous with a tree branch leading to a
defined monophyletic group. When a number of taxa share more than one
closest common ancestor, they do not fit the definition of a clade. In this case,
they are referred to as paraphyletic. It includes the common ancestor and
some but not all descendants of the common ancestor. In the tree (Fig. 4.8),
Species A, B, C, D, E and Node c,d,e and f are included in the group, but
species K is excluded. The group is paraphyletic because it does not include
all the descendants of the common ancestor represented by Node c (i.e.,
species K is missing from the grouping).
A Polyphyletic group consists of distant taxons who do not share recent
common ancestor. It is not defined by a single common ancestor. In
polyphyletic group the most recent common ancestor of all the included
organisms of the group is not included. Species A, B, C and D collectively
represent a polyphyletic group because the common ancestor of the group
represented by node e is not included (Fig. 4.8).
Fig. 4.8: Phylogenetic tree showing monophyletic, paraphyletic, and
polyphyletic groups.
4.5 Limitations of Phylogenetic Trees
It may be easy to assume that more closely related organisms look more alike,
and while this is often the case, it is not always true. If two closely related
lineages evolved under significantly varied surroundings or after the evolution
of a major new adaptation, it is possible for the two groups to appear more
different than other groups that are not as closely related.
69
Block 2 Evolutionary Biology-II
Another aspect of phylogenetic trees is that, unless otherwise indicated, the
branches do not account for length of time; they exhibit only the evolutionary
order. In other words, the length of a branch does not typically mean more
time passed, nor does a short branch mean less time passed-unless specified
on the diagram.
SAQ 1
State True or False:
a) A rooted phylogenetic tree is a directed tree with a unique node.
b) In a Phylogenetic tree, the branching pattern of a tree is called topology
c) Homologous genes that are evolved from common ancestors by
speciation event.
d) A group of taxa descended from a single common ancestor is defined as
a paraphyletic group.
e) Genes that are evolved after the speciation event comprise of
Orthologous genes.
4.6 METHODS TO STUDY PHYLOGENETICS
Phylogeny is inferred using the shared character of species that could be
morphological characters in living and fossil species, biochemical properties
and/or macromolecule sequences, especially DNA and Proteins.
4.6.1 Using Fossil Records
which are remains of dead organisms containing ancestor phenotype
information. Its limitation is that it is available only for certain species. For
microorganisms, fossils are non-existent.
4.6.2 Construction of a Phylogenetic Tree Using
Molecular Data
Molecular sequences, particularly nucleotide sequences of DNA in different
species, are used to study the evolutionary relationships of genes and other
biological macromolecules. In this method, relationships among organisms are
studied by comparing the homologues of DNA or protein sequences
(molecular markers). It helps to establish a natural classification of
microorganisms. In addition, it also helps to infer gene functions of newly
discovered genes and elucidate mechanisms of viral outbreaks.
70
Unit 4 Molecular Phylogency
Steps of Molecular phylogeny
Select appropriate molecular markers from data
Identify and retrieve homologous DNA or protein sequence from database or
amplification, sequencing, and assembly of sequence from lab
Perform multiple sequence alignment of those sequences
Select substitution model or evolutionary model
Select tree building methods to construct phylogenetic tree
Represent trees that give relevant and reliable information
I. Choice of data
In phylogeny, different genes or combinations of genes or DNA regions may
be used to infer phylogenetic trees while addressing groups of organisms.
These are called phylogenetic markers. Sequences that are highly conserved
are accepted as phylogenetic markers.
● Ribosomal RNAs are highly conserved and have long sequences. These
sequences are used in distantly related taxa wherein we require slowly
evolving sequences which accumulate mutations at a very slow rate.
● Chromosomal DNA for comparing homologous genes.
● Mitochondrial DNA has an increased mutation rate that is favoured for
comparing taxa that are closely related like individuals of a population.
● cDNA/RNA
● Protein sequences are more conserved and provide sensitive alignment
in comparison to nucleotide sequences due to degeneracy of codons
and are used for divergent groups of organisms.
● For rapidly evolving microbes which lack fossils, viral DNA sequences
with high mutation rates and amino acid sequences of protein which play
an important role in infection are used to reconstruct virus phylogenies.
For example, for HIV, gp120 protein sequence and for COVID-19, spike
protein sequences are used to reconstruct viral phylogenies.
II. Source: Phylogenetic marker can be obtained from two sources.
i) By sequencing in lab
ii) From available database 71
Block 2 Evolutionary Biology-II
Three publicly accessible databases store large amounts of nucleotide and
protein sequence data: GenBank, which is built by the National Center for
Biotechnology Information (NCBI) ([Link] is part of the
International Nucleotide Sequence database Collaboration, along with its two
partners, the DNA Data Bank of Japan (DDBJ, Mishima, Japan) and the
European Molecular Biology Laboratory (EMBL) nucleotide database from the
European Bioinformatics Institute (EBI, Hinxton, UK). All three centres provide
separate points of data submission, yet all three centres exchange this
information daily, making the same database (albeit in slightly different format
and with different information systems) available to the community at-large.
Accession number (for the DNA sequence): It originally consisted of one
uppercase letter followed by five digits e.g., A024567. New accessions consist
of two uppercase letters followed by six digits e.g., AF024567.
Accession number (for the protein sequence): It consists of three
uppercase letters followed by five digits e.g. BBB12345.
A GenBank Flat File Format is shown in figure 4.9.
Fig. 4.9: GenBank Flat File Format. Selected parts of the GenBank entry
DQ408531. The complete entry can be found at [Link]
[Link]/nuccore/DQ408531.
FASTA format stands for FAST-All, referring to its ability to perform a fast
alignment of all sequences (i.e., proteins or nucleotides). Sequences are
required for construction of a phylogenetic tree with the help of software. For
this, a standard format for inputting nucleic acid and protein sequence
information is FASTA format. A sequence in FASTA format begins with a
72
Unit 4 Molecular Phylogency
single-line description followed by lines of sequence data. The description line
is distinguished from the sequence data by a greater than (“>”) symbol in the
first column. It is recommended that all lines of text be shorter than 80
characters in length (Fig. 4.10).
Fig. 4.10: FASTA format.
BLAST
The Basic Local Alignment Search Tool (BLAST; Altschul et al., 1990) is a
popular method of ascertaining sequence similarity. Basic Local Alignment
Search Tool (BLAST) is the main tool of the National Center for Biotechnology
Information (NCBI) for comparing a protein or DNA sequence to other
sequences in various databases (Altschul et al., 1990, 1997).
III. Sequence Alignment
It is defined as aligning two sequences, one below the other and scoring the
similarities and differences at each and every nucleotide or amino acid level.
It is useful for discovering functional, structural, and evolutionary information in
biological sequences. Sequences that are alike and similar according to
sequence analysis and having a common ancestor are defined as
homologous sequences. Alignment is the necessary step to conclude the
homology between the sequences (genes). To align the sequences, it is
imperative to select an appropriate site and choose a suitable macromolecule
(DNA or protein).
The alignments are of two types:
1. Pair-wise sequence alignment compares only two sequences, one of
the pairs is new or unknown, while the other sequence is of known
structure and function. It helps in finding the evolutionary history of the
sequences.
Seq A: L I A H G S V M L N
Seq B: L I G H G S A M L P 73
Block 2 Evolutionary Biology-II
2. Multiple sequence alignment compares more than two sequences to
find out gene families, protein families, enzyme active sites, functional
and evolutionary relatedness etc. Figure 4.11 shows a representative
picture of multiple sequence alignment using CLUSTAL.
A: L I – H G S V M L –
B: L I – H G S A – L P
C: - I G H – S A - L P
D: L I G H - S A - L -
E: - I G H G S A M L P
Fig. 4.11: Multiple sequence alignment using CLUSTAL.
The unique advantage of MSA is that it reveals more biological information,
allows identification of conserved sequence patterns and motifs in the whole
sequence family which is not possible in pairwise sequence alignment. MSA is
an essential prerequisite to carry out phylogenetic analysis of sequence
families and prediction of protein secondary and tertiary structures. MSA also
has applications in designing PCR primers based on multiple related
sequences.
Trimming Sequence
Sequences must be aligned before they can be utilised for phylogeny
reconstruction. The quality of the multiple sequence alignment has been found
to affect the quality of the derived phylogeny. The elimination of poorly aligned
regions from an alignment has has been demonstrated to improve the quality of
subsequent analyses. Figure 4.12 shows a representative picture of Trimmed
aln file.
74 Fig. 4.12: Trimmed aln file.
Unit 4 Molecular Phylogency
IV. Substitution model: Two types of sequence evolution model
1. Jukes cantor one parameter model: This is the simplest substitution
model for correcting multiple substitutions in a molecular sequence and
this model assumes that the rate of substitution of all nucleotides (A, T,
G and C) is the same. It is based on the hypothesis that the rate of
evolution is constant and hence it is called one parameter model.
2. Kimura two parameter model: Substitution model for correcting
multiple substitutions in molecular sequences and this model assumes
there are two different substitution rates, one for transition and the other
for transversion leading to two parameter model (Fig. 4.13).
Fig. 4.13: Type of substitutions.
V. Choice of methods to construct phylogenetic tree
There are two different methods based on molecular phylogeny viz. Distance
based method and Character based method. Distance based methods are
based directly on pairwise distances between the two aligned sequences. It
uses a substitution model to correct distances. The disadvantage of this
method is that it does not preserve actual sequence information when a
character is converted into a single value distance. Character based methods
are based directly on sequence characters. This method counts mutational
events thus preserving character information when characters are converted
into distances. Both methods are related since the genome strongly
contributes to the phenotype of the organisms. Therefore, organisms with
similar genes are more closely related.
A. Distance-based methods (Phentics)
Neighbor joining: It is aphylogenetic tree-building method that builds a tree
by using stepwise reduced distance matrices. It first corrects unequal
evolutionary rates of raw distances and uses the corrected distances to build a
matrix. The tree construction begins from a completely unresolved star tree by
joining all taxa onto a single node. Then taxa with the shortest distances
sharing more recent common ancestors are joined first as a node followed by
joining next most closely related taxon. This cycle is repeated and then
reduces the tree in a stepwise fashion until all taxa are resolved called star
decomposition. In the neighbour joining method, the outgroup is not defined
and produces an unrooted tree. Outgroup positions need to be determined
based on sequence knowledge. Figure 4.14 shows a representative picture of
Phylogenetic tree construction using Neighbor joining method.
75
Block 2 Evolutionary Biology-II
Fig. 4.14: Phylogenetic tree construction using Neighbor joining method.
B. Character based methods (Cladistics)
i) Maximum likelihood: It is astatistical method of choosing a tree based
on the highest probability (maximum likelihood) of reproducing the
observed data. It selects trees which reflect the actual evolutionary
process. It searches every tree topology and uses all sites in probability
calculation for all combinations of possible ancestral sequences to give a
best tree with highest likelihood values of character distribution. It is a
most useful, robust and time-consuming method. The advantage is that
it can provide information about rates of evolution.
ii) Maximum parsimony: Phylogenetic inference is based on the principle
of choosing a tree with the simplest explanation requiring the smallest
number of evolutionary changes. It is the most useful method to give a
good estimation of a true tree based on the least number of substitution
mutations and shortest overall branch length. The tree with the fewest
assumptions and fewest logical steps has maximum parsimony and is
the most parsimonious tree.
VI. Statistical analysis of obtained phylogenetic tree
This is the last step of phylogenetic tree construction. The tree that is
generated needs to be tested or evaluated statistically. The most well-known
methods for statistical analyses are bootstrap and jackknifing methods.
Bootstrap analysis: It is a statistical method for assessing the consistency of
phylogenetic tree topologies. It tests sampling error in a phylogenetic tree
based on the generation of a large number of replicates with slight
modifications in input data. The consensus trees constructed from the
datasets with random modifications give a distribution of tree topologies that
allow statistical assessment of each individual clade on the trees (Fig. 4.15).
76
Unit 4 Molecular Phylogency
Fig. 4.15: Bootstrap consensus tree.
Jackknifing: This is another resampling technique in which half of the sites
are randomly deleted and the new dataset is half as long as the original which
is subjected to construction of a phylogenetic tree using the same method as
the original. The advantage of jackknifing is that it takes less time and sites are
not duplicated relative to the original dataset. However, deleting one half of the
sites creates datasets which are not replicated, thus the results are not
comparable with bootstrap.
VII. Software for phylogenetic tree reconstruction
The Clustal programs are widely used to carry out automatic multiple
alignments of both nucleotides and amino acid sequences to prepare a
phylogenetic tree.
The most common multiple sequence alignment software is ClustalW and
ClustalX.
ClustalW is a general-purpose multiple sequence alignment program for DNA
or protein sequences. It is available over the web at the European
bioinformatics institute. It produces biologically meaningful multiple sequence
alignments of divergent sequences. It calculates the best match for the
selected sequences and lines them up so that the identities, similarities, and
differences can be seen. Evolutionary relationships can be seen via viewing
cladograms or phylograms. It is useful in comparing sequences from different
sources. For example, to study the differences and similarities between
molecular sequences of different organisms, it identifies and highlights the
regions that are conserved in different species.
Online at [Link]
ClustalX is the X-Window friendly version of ClustalW, which can be easily
downloaded.
The alignment produced by two programs are exactly the same, the only
difference between ClustalW and ClustalX is the way the user interacts with
the program. Multiple sequence format (MSF) is used by several software
tools. For instance, Clustal W/X has its own ALN format. The result output is
coloured, which makes assessment of alignment easier. The histogram below
the ruler indicates the degree of similarity. Peaks indicate position of high
similarity and valleys indicate position of low similarity. The grey line just 77
Block 2 Evolutionary Biology-II
above the sequence is used to mark strongly conserved positions.
Here, “ * ” means that the residues or nucleotides in that column are identical
in all sequences in the alignment.
“ : ” means that the conserved substitution has been observed according to the
colour table which indicates that amino acid is substituted with amino acid
having strongly similar biochemical properties
“ . ” means that semi conserved substitutions are observed which indicates
that the amino acid is replaced with amino acid having weakly similar chemical
properties.
The last thing to modify before alignment is output format. For construction of
a phylogenetic tree using the phylip package, choose the phylip as an output
format.
Phylip (Phylogenetic inference package; by Joe Felsenstein) is a free
multiplatform comprehensive package containing thirty-five subprograms for
performing distance, parsimony, and likelihood analysis, as well as
bootstrapping for both nucleotide and amino acid sequences.
Tree view: This software is used to view consensus out tree files made in
Phylip.
CONSTRUCTION OF PHYLOGENETIC TREE BY USING MOLECULAR
DATA
Protocol
1. Identification and retrieval of homologous molecular sequence from
database
2. Paste to notepad
3. Remove everything before accession number leaving (> Accession
Number)
4. Open Clustal X2
5. Load or paste sequence
6. Do complete alignment - Complete align
7. Three files aln, dnd, and phylip will form
8. Specify destination folder and click OK
9. The .aln file will be saved in destination folder
10. Open aln file
11. Do trimming and save
12. Open Clustal X2
13. Load trimmed sequence
78
Unit 4 Molecular Phylogency
14. Select Output Format - Phylip Complete alignment
15. Phy folder will be formed (Phylip file will be used for the construction of
phylogenetic tree with the help of Phylip software)
16. Open the phylip software
17. Open the exe folder
18. Copy and paste the phylip file (.phy) into exe
A. Neighbor Joining method
• Open Seq
• File name type .Phy
• Press enter
• Press Y
• Press 5
• Press Enter
• Quit
• Outfile will be formed rename it (01)
• Open and start DNA dist (Distance Estimation)
• Type the previously formed outfile name (O1)
• Enter
• Y to accept
• D
• Enter
• M
• Multiple data sets or Multiple Weights-d
• How many data sets type (100)
• Enter
• Press ‘Y’
• Press Enter
• Quit
• Outfile will be formed rename it (O2) ● Open Neighbour
• Type the outfile name
• Press Enter
• Type O 79
Block 2 Evolutionary Biology-II
• Outgroup at (8)
• Enter
• Type M
• Enter
• Multiple Data sets
• 100
• Enter
• Random number seed
• 5
• Type Y and enter
• Quit
• Outfile will be formed rename it (example - O3)
• Outtree will be formed rename it (OT1)
• Open Consensus
• Enter the last outtree formed
• Enter
• Type O
• Enter
• Press 8
• Type R
• Enter
• Type Y
• Enter
• Outfile (O4) and Outtree formed (OT2)
• Tree View Open (OT2)
• Rectangular Cladogram
• Show Internal nodes
B. Maximum Parsimony Method
• Same step till Seqboot
• Open Dna Pars
• Type file name (last formed)
80
Unit 4 Molecular Phylogency
• Enter
• Y
• Number of Tree
• 1000
• Type O
• Enter Number of Outgroup (8)
• Enter
• M
• Enter
• Multiple Data sets
• 100
• Enter
• Random Seed Number
• 5
• Type Y and Enter
• Quit
• Outfile and Out tree will be generated (rename it)
• Same as joining method (Neighbour)
C. Maximum Likelihood Method
• Same till seqboot
• Open DNAml
• Open file (Last formed)
• Enter O (8) (Outgroup)
• Enter
• Multiple Data sets
• 100
• Enter
• How many times Jumble
• 5
• Enter
• Random Seed Number
81
Block 2 Evolutionary Biology-II
• 5
• Type Y and Enter
• Quit
• Outfile and Out tree will be generated (rename it)
• Same as joining method (Neighbour)
4.7 Construction of Phylogenetic Tree by Using
16S rRNA Gene Sequence
Since time immemorial and with the better understanding about the gene
sequences, 16S rRNA gene sequence has widely served as a pivotal tool for
understanding bacterial phylogeny and taxonomy. It is present in all the
bacteria and its function has not changed over the period of time. The 16S
rRNA gene sequence is important in identification of genus and species of
isolates which is not possible using any biochemical profiling. In most cases,
the identification of 16S rRNA gene sequencing aids in genus identification
(>90%). Although the 16S rRNA gene sequence has the potential to identify
bacteria in both clinical and public health research, there are evidences which
show that it is not applicable in each and every place. 16S rRNA gene
sequencing method has been employed for several decades in clinical
laboratories; however, its utilization has suffered a setback, in part, by the time
and cost associated with sequencing along with lack of universality and the
inability to differentiate between true pathogenic organisms, opportunists, and
commensals. Also, 16S rRNA gene sequencing is less sensitive and slower
than pathogen-specific polymerase chain reaction (PCR) assays using panels
of primers.
To identify the bacterial strains on next generation sequencing using the 16S
rRNA gene sequencing comprises amplification of a region of the 16S rRNA
gene. Further, steps involve library preparation; sequencing followed by
analysis of the sequence data. In addition, 16S rRNA sequence diversity
among bacterial species is important to discriminate between bacterial
species. Also, apart from evolution related differences among species, 16S
rRNA gene sequence can differ within the same genome because most of the
organism or species have multiple copies.
Phylogenetic tree is a presentation depicting the lines of evolutionary descent
of different genes, species, and organisms of a common ancestor. It is a
useful tool for understanding evolutionary relationships, time scale and
classifications. A brief overview of the method to construct a phylogenetic tree
using 16S rRNA for bacterial species involves the following steps:
Isolation of bacteria
Extraction of bacterial genome using suitable extraction methods
82
Unit 4 Molecular Phylogency
Amplification of 16S rRNA using PCR
Sequencing of 16S rRNA
Comparison of the sequenced gene with the available data on GenBank
Construction of Phylogenetic tree using methods such as distance-based
methods (Neighbour Joining, UPGMA, WPGMA) or character-based methods
(maximum likelihood and maximum parsimony).
The use of 16S rRNA has been proven to be important in understanding
bacterial classification, identification of sequence similarity and evolution as
16S rRNA is the conserved gene sequence across bacterial species.
4.8 CONCEPT OF SPECIATION IN BACTERIA
A species can be stated as a population that can interbreed and is
reproductively isolated from other groups. For instance, a group of monkeys,
chimpanzees, and human beings can easily be differentiated. But what if one
is asked to classify bacteria into different species? It is not at all surprising to
hear no as an answer because it is indeed a very difficult task. So then, how to
define bacterial species?
Since bacteria do not reproduce sexually and lack a fossil record, taxonomists
face a great disadvantage when defining bacterial species. Although, the most
common definition of bacterial species is that they are the collection of strains
which have stable properties and significantly differ from other strains. Strains
within a species can also be described in a number of ways. Biovars are
strains with different biochemical and physiological properties, morphovars are
strains with different morphology, serovars have distinct serological
differences, and pathovars are the pathogenic strains distinguished based on
the host in which they cause the disease.
A species can have strains that are metabolically and genetically so diverse,
that it seems likely that such strains belong to different species altogether.
Alternatively, two different species can be less different and more similar to
each other. For example, the strains of Bacillus anthracis and Bacillus cereus
are almost alike and it is argued that they come under a common species.
Therefore, the International Committee on Systematics of Prokaryotes (ICSP),
suggested four different criteria for species assignment.
• At least 70% similarity in the whole genome as determined by DNA-
DNA hybridization.
• A minimum of 97% 16S rRNA homology.
• A difference of no more than 5° in the melting temperature of DNA (a
reflection of G + C content). 83
Block 2 Evolutionary Biology-II
• Physiological, morphological, and ecological similarity.
Here, it should be noted that if these criteria were applied to eukaryotes, for
instance, with a similarity of 75% in DNA and 98% rRNA homology, all
primates (monkeys, apes, and humans) would come under a single species!
As opposed to all the confusion in defining a specific criterion for bacterial
species, taxonomists agree that the concept of bacterial species is based on
evolution and natural selection. But since the genetic diversity in bacteria
occurs asexually, mutation and horizontal gene transfer are the two
mechanisms that can explain the evolution of bacteria. The fundamental units
of biotic organization are species which play such a pivotal role in
understanding diversity of organisms. According to Dobzhansky and Mayr,
biological species are reproductively isolated units; however, recent evidence
shows that reproductive boundaries can be leaky. The members of a species
share genotypic and phenotypic connectivity. Also, in eukaryotes gene
exchange occurs by the process of syngamy where two haploid genomes from
different parents result in diploid offspring.
For bacteria and archaea, the situation has been different under a plethora of
questions about the concept of speciation. Bacteria grow clonally whereby
asexual cell division produces two daughter cells from a single parent.
Therefore, it may not be employed to the Mayrian Biological Species Concept,
which was formulated for eukaryotic species wherein obligate genetic
exchange occurs during reproduction. Yet bacteria are not strictly asexual and
may exchange DNA with very distantly related partners through several
mechanisms such as conjugation (direct DNA transfer via cytoplasmic
bridges), transduction (movement by DNA encapsulation within virus particles)
and transformation (direct uptake of naked DNA). Molecular analysis of
bacterial diversity has revealed the presence of numerous species, both within
a single community and worldwide which challenge microbial ecologists and
evolutionary biologists to explain such origins and existence in thousands and
millions.
The mode of gene exchange in eukaryotes is different from prokaryotes.
Though recombination has long been thought to be a powerful force of
cohesion within eukaryotic species, it is not the case in bacteria.
Recombination in bacteria is an extremely rare phenomenon. Bacterial entities
do not exchange genes as frequently as eukaryotes. Recombination events
are highly localised to a very small fraction of the bacterial genome. Bacterial
recombination does not involve transfer of homologous segments.
But the question arises: how does the bacterial microcosm exhibit a diverse
array of species? The answer to this may be directed towards the concept of
cladogenesis (the splitting of one population into multiple ecologically distinct
populations). It was found that in the absence of competing species among the
descendants of a Bacillus subtilis clone, in at least 7 of 10 replicated
microcosm communities, the original population founded one or more new,
ecologically distinct populations (ecotypes). Due to the infrequent occurrence
of recombination in bacteria, a population may diverge into two ecologically
distinct populations that can exhibit coexistence and indefinite divergence,
84
Unit 4 Molecular Phylogency
without sexual isolation. There is also growing documentation of sympatric
splitting of bacterial lineages in nature and in laboratory microcosms.
Defining bacterial species as ecotypes has several advantages because such
species will be ecologically distinct. Henceforth, divergence of animals into
species-like populations (that are ecologically distinct and irreversibly
separate) requires both ecological and sexual divergence whereas irreversible
divergence in bacteria requires only the divergence of ecological features;
moreover, geographic isolation is not required.
SAQ 2
a) In the given phylogenetic tree:
i) the most common ancestor of taxa A, B and
C is?
ii) Is it a rooted tree?
iii) State True or False: A and B are sister taxon.
b) State True or False: Mitochondrial DNA (mtDNA) is used to study the
relatedness of animal populations because mtDNA mutates at a slower
rate than nuclear DNA.
c) Fill in the blanks:
i) Sequences that are highly …………… are accepted as
phylogenetic markers.
ii) Protein sequences are more conserved and provide sensitive
alignment in comparison to nucleotide sequences due to
…………….
iii) The …………… is astatistical method of choosing a tree based on
the highest probability of reproducing the observed data.
iv) Phylogenetic inference in case of ……………. is based on the
principle of choosing a tree with the simplest explanation requiring
the smallest number of evolutionary changes.
v) ……………. is a general-purpose multiple sequence alignment
program for DNA or protein sequences.
vi) ………….. is the X-Window friendly version of ClustalW.
vii) The full form of Phylip is ………………. .
viii) The ………. gene sequence is important in identification of genus
and species of isolates which is not possible using any
biochemical profiling.
ix) The …………. can be defined as the collection of strains which
have stable properties and significantly differ from other strains.
x) A …………. indicates the hypothetical common ancestor, or
ancestral lineage, of the tree.
85
Block 2 Evolutionary Biology-II
4.9 SUMMARY
• Variations and similarity between sequences provide a means to study
the relationship of various organisms and taxa that is phylogeny. Two
closely related species will share a more recent common ancestor and
vice versa.
• In order to study phylogeny, the shared character of species that could
be morphological characters in living and fossil species, biochemical
properties and/or macromolecule sequences, especially DNA and
Proteins are used.
• Molecular data is advantageous over morphological data because of
ease of access and comparative purpose.
• Evolutionary relationships are depicted by a phylogenetic tree which is a
graph composed of branches and nodes. In a phylogenetic tree, each
node with descendents represents the most common ancestor of the
descendants and the edge length in some trees corresponds to time
estimates.
• A phylogenetic tree can be rooted or unrooted. Phylogenetic tree based
on molecular sequence diversity is Molecular phylogeny. The greater is
divergence between sequences, more distant is their last common
ancestor.
• A phylogenetic tree that shows evolutionary relationships as an array of
bifurcating branches is a cladogram or binary tree.
• Useful macromolecular sequences for phylogenetic study are
mitochondrial DNA, ribosomal RNA sequence, chloroplast, and
homologous genes. To study distantly related organisms, a slowly
evolving highly conserved and ubiquitous sequence is preferred like
16srRNA. Mitochondrial DNA is a preferred method to study closely
related organisms.
• There are various methods to construct phylogeny based on Distance
matrix, maximum parsimony and maximum likelihood method. In
addition, two types of evolutionary models are, Jukes cantor (One
parameter) and Kimura based (Two parameter) model.
• The construction of phylogenetic trees using 16S rRNA has been proven
to be important in understanding bacterial classification, identification of
sequence similarity, and evolution as 16S rRNA is the conserved gene
sequence across bacterial species.
• Bacteria species are a collection of strains that share many stable
properties and differ significantly from other groups of strains. Although,
two different bacterial species can be very similar and differ quite little
from each other. For example, the strains of Bacillus anthracis and
Bacillus cereus are almost alike and it is argued that they come under a
common species.
86
Unit 4 Molecular Phylogency
• To standardize bacterial taxonomy, the International Committee on
Systematics of Prokaryotes (ICSP) has recommended four criteria for
species assignment that includes (a) At least 70% similarity in the whole
genome as determined by DNA-DNA hybridization. (b) A minimum of
97% 16S rRNA homology. (c) A difference of no more than 5° in the
melting temperature of DNA (a reflection of G + C content). (d)
Physiological, morphological, and ecological similarity.
4.10 TERMINAL QUESTIONS
1. Differentiate between the following:
a) Orthologous and Paralogous sequence
b) Dendrogram and Phylogram
c) Maximum parsimony and Maximum likelihood method
d) Rooted tree and Unrooted phylogenetic tree
e) Monophyletic and Paraphyletic
2. Define the following
a) Molecular Markers
b) Polyphyletic group
c) Bootstrap
d) Homologous sequence
3.
a) In the rooted tree depicted above, the group A+B represents:
i) Clade;
ii) Paraphyletic group;
iii) Monophyletic group;
iv) A and C
b) In the rooted tree depicted above the group Y+D+C represents:
i) Clade;
ii) Polyphyletic group
87
Block 2 Evolutionary Biology-II
iii) Monophyletic group
iv) A and C
4. In the given rooted tree:
a) Which species diverged earliest?
b) Which species evolved more recently?
c) What do the numbers on the nodes represent? Explain their
significance.
5. Describe a distance-based method of phylogenetic tree construction?
6. What are the advantages of molecular phylogeny?
7. What are evolutionary/substitution models?
8. Describe the construction of phylogenetic tree by using 16S rRNA gene
sequence.
4.11 ANSWERS
Self-Assessment Questions
1. a) True; b) True; c) False
d) False; e) True
2. a) i) M
ii) Rooted tree
iii) False
b) False
c) i) conserved
ii) degeneracy of codons
iii) Maximum likelihood
iv) Maximum parsimony
v) ClustalW
88
Unit 4 Molecular Phylogency
vi) ClustalX
vii) Phylogenetic inference package
viii) 16S rRNA
ix) bacterial species
x) Rooted Tree diagram
Terminal Question
1. a) Orthologous and Paralogous gene
Orthologous gene Paralogous gene
Homologous sequences from Homologous sequences from the same
different organisms derived from a organism, which are derived from gene
common ancestor due to duplication events rather than speciation
speciation events rather than events.
gene duplication events.
Orthologous gene may or may not They have similar functions.
have similar functions.
Example : Alpha Globin gene of Example: Human Alpha and Beta Globin
Mouse and Human gene
b) Dendogram and Phylogram
Dendogram Phylogram
It is a branching diagram in the In a phylogram, the branch lengths
form of a tree used to represent represent the amount of evolutionary
the evolutionary relationship. divergence.
It depicts the time of divergence. These are scaled trees that show both
the evolutionary relationships and
information about the relative time/rate
of divergence of the branches.
c) Maximum parsimony and Maximum likelihood method
Maximum Parsimony method Maximum Likelihood method
89
Block 2 Evolutionary Biology-II
Tree topology is based on Tree topology is based on the highest
minimizing the number of changes probability of having produced the
in character states. Tree with the observed sequences or likelihood
least number of mutations is the values.
most parsimonious tree.
It is a character-based method This is the most statistically suitable
method for phylogenetic analysis.
Employs a subset of alignment Employs all the data
positions of sequences
Assumptions fail when there is Result depends on selection of
fast evolution evolutionary model
A good choice for less than 30 A good choice for small datasets and for
sequences and when homoplasy
validating trees constructed by other
is rare.
methods.
d) Refer to Section 4.4.
e) Refer to Section 4.4.
2. a) Refer to Section 4.3.
b) Refer to Section 4.4.
c) Refer to Sub-section 4.6.2.
d) Refer to Sub-section 4.6.2.
3. a) A and C
b) Polyphyletic group
4. a) Species X
b) Species H
c) Refer to Sub-section 4.6.2.
5. Refer to Section 4.6.
6. Refer to Sub-section 4.6.2.
7. Refer to Section 4.6.
8. Refer to Section 4.7.
Acknowledgment
Fig. 4.9 : [Link]
0032-5/figures/5
90
Unit 5 Molecular Clock
UNIT 5
MOLECULAR CLOCK
Structure
5.1 Introduction 5.6 Molecular Clock
Objectives 5.7 Complications in Inferring
Phylogenetic Tree
5.2 History
5.9 Summary
5.3 Definitions and
Terminology 5.10 Terminal Questions
5.4 Molecular Drive 5.11 Answers
5.5 Molecular Divergence
5.1 INTRODUCTION
Macromolecules like protein and nucleic acid sequence accumulate changes
with time as they pass from generation to generation. This is the basis for
molecular divergence. By looking at the number and rate of variations among
biomolecules one can find the timing of evolutionary change. Molecular clocks
help in estimating the timing of divergence among species. It explains how
species evolve and fix the time of divergence on evolutionary timeline.
Evolution is a result of genetic mutations however not all mutations are
important for affecting the organism’s survival or evolutionary fitness, these
mutations are known as neutral mutations. These neutral mutations play an
important role in studying evolution as they generally occur at a constant rate
over time. Around 50 years ago, it was accepted that these neutral mutations
would also be used in molecular clock estimation. For instance, how humans
diverged from bonobos and chimpanzees. Molecular clock-based dating
methods, or molecular divergence, are the primary tool for estimating the
timely origin of lineages. Divergence dates estimated from molecular
phylogenies provide critical information on the timing of historical evolutionary
events, including the temporal origins of clades. Evolutionary periods can be
estimated from molecular data using molecular clocks. Phylogenetic tree
depicts evolutionary relationships based on molecular divergence data.
However, there are several problems in inferring phylogenetic trees.
91
Block 2 Evolutionary Biology-II
Objectives
After studying this unit, you would be able to:
understand the relationship between the establishment of the timing of
evolutionary divergence of species with the changes in nucleotide and
amino acids;
describe molecular drive, molecular divergence and reasons for its
occurrence;
explain the origins of molecular drive, molecular divergence, and the
neutral theory;
comprehend the neutral theory for variation within and between species.;
describe molecular clock, the method of estimating molecular
divergence, its applications, and disadvantages; and
understand the complications associated with inferring phylogenetic tree.
5.2 HISTORY
Before the 1960s, population geneticists believed that natural selection
maintains homozygosity by removing any deleterious allele and fixing
advantageous allele in the population. People in favor of balance theory
asserted that balancing selection maintains polymorphism in the population.
Both schools of thought asserted that molecular evolution is driven by natural
selection. With the invention of protein separation by electrophoresis, the
results of enzyme polymorphism showed that there is a high level of
polymorphism favouring balancing selection. But it was suggested that
balancing selection alone cannot maintain so much polymorphism, some other
possible mechanisms of action are involved in it. Further, in the early 1960s to
study protein evolution among different species, the amino acid sequence of
globin gene, cytochrome C and fibrinopeptide were compared. It was found
that the number of differences between protein sequences of different species
was roughly proportional to the time since species’ divergence. The difference
in sequence among different species increases linearly with the divergence
time of the two species. Zuckerkandl and Pauling (1965) hypothesized that the
rate of evolution for any given protein is constant over time and given the
concept of molecular clock. It means that the constant accumulation of amino
acid substitutions in a given stretch of proteins with time could be compared to
‘ticks of clocks’ and stated that ‘there may exist a molecular evolutionary clock'
to measure divergence time of two species. Molecular clocks predict a
constant rate of evolutionary change among organisms. This suggestion
implies the existence of a sort of molecular clock ticking faster or slower for
different genes but at a more or less constant rate for any given gene among
different phylogenetic lineages. The idea of molecular clocks was
advantageous as it helps in quick and easier reconstruction of phylogeny. In
addition, Motoo Kimura, who gave the neutral theory of evolution in 1968
reinforced the idea of molecular clocks. However, he believed that the rate
was constant throughout time and across different species. This idea was too
92
Unit 5 Molecular Clock
strict, because molecular evolution rates are different in different organisms.
Today, we use a more relaxed model that allows the rate of molecular
evolution to vary slightly among species.
5.3 DEFINITIONS AND TERMINOLOGY
Fixation: It is a new mutation that increases in frequency and reaches 100
percent or replaces other variants either by a process of natural selection or
genetic drift.
Molecular Clock: Molecular clock is a figurative term or tool that uses
mutation rate of biomolecules to deduce time in prehistory when two or more
life forms diverged (estimated time from fossil record).
Molecular drive: Molecular drive is independent evolutionary process that
changes the genetic material of a population via DNA turnover mechanisms
from one lineage to another lineage.
Molecular Evolution: Changes in biomolecules (DNA, RNA and proteins)
over generations due to mutation, natural selection, genetic drift subsequently
leads to changes in sequence of these molecules in different lineages. It is the
study of patterns and processes that result in different lineages.
Mutation: Change in genetic material due to base insertion, deletion,
substitution, and rearrangements, which is inheritable and eventually results in
new variation in population.
Neutral Theory: It proposes that the number of mutations does not have any
role in survival or evolutionary fitness of an organism. Therefore, these
mutations are not favored or disfavored by natural selection.
5.4 MOLECULAR DRIVE
It is an evolutionary process that changes the genome of the population
through various generations. The term was first coined by Gabriel Dover in
1982. It is distinct from natural selection and neutral drift in that it emerges
from the activities of a number of ubiquitous mechanisms of DNA turnover
(MOT), such as gene conversion, unequal crossing over, slippage replication,
transposition, retro transposition, and RNA mediated exchanges (Fig. 5.1).
This mechanism of turnover can result in convergence or divergence of
sequence subsequently resulting in spread of mutation through a family and
later in population. These mechanisms induce a gain or loss of variant gene in
organism’s life resulting in non-Mendelian segregation ratio. This has been
observed in ribosomal RNA and egg shell egg chorion protein. This is
responsible for structure and function as a cohesive unit of evolutionary
changes in the multigene family. This is a dual process of homogenization and
fixation. Moreover, mutations changing the sequence of a gene are less
common than duplications, deletions, and replacement by another; the gene
copies look like each other more than if they had been evolving autonomously.
93
Block 2 Evolutionary Biology-II
Fig. 5.1: Types of DNA turnover mechanisms.
SAQ 1
Fill in the blanks:
a) The …………… predict a constant rate of evolutionary change among
organisms.
b) The ………..mutations do not have an impact on an organism's ability to
survive or evolve.
c) The …………… is an evolutionary process that changes the genome of
the population through various generations.
5.5 MOLECULAR DIVERGENCE
Molecular Divergence is changes in sequence of biomolecules (DNA, RNA,
and proteins) with time in population. The change in these molecules
demarcates the evolution of an organism's form and function. Also, it
delineates the field of molecular divergence which discovered how genomes
evolve and how these molecular changes result in biological diversity. The
molecular clock estimates the constant rate of change in biomolecules of an
organism over time. The constant rate of change about biomolecules is
different among different species or organisms. For instance, it is different for
plants, animals, fungi and viruses. By calibrating these changes researchers
can make a phylogenetic tree representing the divergence of one species from
another.
94
Unit 5 Molecular Clock
5.6 MOLECULAR CLOCKS OR EVOLUTIONARY
CLOCKS
Molecular clock estimates molecular divergence, the evolutionary differences
and variation among organisms due to mutations. It is a well-known fact that
the number of mutational differences between organisms is equivalent to the
evolutionary distance among them. Greater is the number of mutational
differences between organisms, greater will be the evolutionary distance
among them (Fig. 5.2). The rate at which mutations become fixed is
determined by the molecular clock. In other words, molecular clock is a
figurative term or tool that uses mutation rate of biomolecules to deduce time
in prehistory when two or more life forms diverged. It is also known as gene
clock. Kimura’s hypothesis on molecular clock stated that sequences of
biomolecules evolve at a rate that is constant over time among different
organisms. The consequence of this constant time is that mutational/genetic
differences between any organisms are directly proportional to the
evolutionary time since these species diverge from common ancestor.
Therefore, genetic differences provide time scales which are further used for
making branch length of a phylogenetic tree that is used to estimate the
divergence time or evolutionary time of lineage. Since the present assumption
on the rate of change of molecular sequences is constant, the number of
accumulated mutations is proportional to the divergence time. However, this
assumption does not hold true in reality. If it would have been true then this
hypothesis would have been useful in calculating the evolutionary time of any
species.
Fig. 5.2: Divergence of species due to accumulation of mutation. Species that
diverged recently and after long time showed few and many nucleotides
differences, respectively.
Molecular Clock Calibration
❖ The rate of protein evolution can be calculated as the number of amino
acid differences between the proteins of the species divided by 2xt to
their common ancestor.
❖ Similarly, the rate of nucleotide change between the species can be
calculated by using the formula given below. Researchers can determine
this with the following information
● When ancestor diverges into two separate species
95
Block 2 Evolutionary Biology-II
● Fossils records
D = 2rt
● Where, D is the proportion of base pairs that differ between the two
sequences
● r is the rate of divergence per base pair per million years, t is the time (in
million years) since the species common ancestor.
● The factor 2 represents the two diverging lineages.
Cytochrome C as an example of Molecular Clock
● Cytochrome c, a protein in respiratory chain, is present in most
organisms.
● Changes in amino acid sequences of cytochrome c in different
organisms presents uniform rate of evolution.
● The gene coding for this protein probably appeared very early in
evolution and was favored by natural selection.
● As a result of few mutations, the sequence of some of its amino acids is
altered.
● Cytochrome c from horses and other mammals differs in 5.1 amino acids
and both diverge from other mammals about 90 million years ago. It
means that on an average 1 amino acid change has occurred every 17.6
million years.
● The average amino acid difference between reptilian and mammalian
cytochrome c is 14.8 and these two groups have diverged about 300
million years ago.
● At this rate of amino acid per 17.6 million years, the plants and the
animals are estimated to have diverged at least 792 million years ago.
This figure is close to 800 million years estimated by paleontologists.
Globin gene family example of molecular clock
● The globins are a superfamily of heme-containing globular proteins,
involved in binding and/or transporting oxygen. The globin proteins
composed of globin fold, a sequence of eight alpha helical sections. Two
important examples include hemoglobin and myoglobin.
● One of the most important processes by which genomes have increased
in size is gene duplication. A new copy of a locus (say, b) arises by
duplication of a pre-existing gene (a), so that a single gene locus in an
ancestor is represented by two loci in the descendant.
● These two genes will subsequently undergo different evolutionary
changes in sequence and can therefore be distinguished. If two species
(1 and 2) both inherit the duplicated pair b and a from their common
ancestor, the relationships among the genes represent two forms of
homology, and so warrant different terms.
96
Unit 5 Molecular Clock
● The genes that originate from ancestral gene duplication are paralogous,
whereas the genes that diverge from a common ancestral gene by
phylogenetic splitting at the organismal level are orthologous.
The hypothesis of Molecular Clock has been refined
The notion of presence of molecular clock was first observed by Emile
Zuckerkandl and Linus Pauling in 1962, when they observed the changes in
number of amino acids of hemoglobin among different lineages with timescale.
They found linear relationship between the changes in amino acids of different
lineages with time based on fossils records. Therefore, they proposed the
molecular clock hypothesis based on empirical observations. However soon in
1968, Motoo Kimura challenged their hypothesis based on neutral theory of
molecular evolution. According to Kimura, plenty of mutations do not have any
role in evolutionary survival/fitness. As a result natural selection would not
favor or disfavor these neutral mutations. Ultimately, it was found that either
these neutral mutations spread throughout the population or become fixed or
they would be lost in a stochastic process, also referred as genetic drift.
Further, Kimura demonstrated that the rate of neutral mutations that become
fixed in population is equivalent to the rate of appearance of new mutations in
the population. However, Kimura’s assumption of molecular clock is too simple
because rate of molecular evolution can significantly vary among the
organisms. Never the less, various efforts have been taken by different
researchers for some reluctance to assumption of molecular clock.
Subsequently, such efforts have led the development of “relaxed” molecular
clock which allow molecular clock to vary among the organisms of the
population. The two major types of relaxed molecular clock models are
currently followed. The first model assumes that rate of mutation varies over
time and between the organisms but these mutations/variations occur around
an average value. The second model assumes that rate of molecular evolution
is associated with other biological characteristics that also undergo evolution.
For instance, there are some evidences that show mutation rates are
influenced by an organism's metabolic rate.
Applications of Molecular Clock/Evolutionary Clock
1. It is an essential tool that is used to estimate the time in prehistory
divergence of two species from common ancestor based on the mutation
rate of biomolecules.
2. In addition, it uses fossil constraints and rates of molecular change to
measure the time in geologic history when two species evolved.
3. It predicts time from molecular divergence. Molecular divergence is
roughly correlated with time. It is used to estimate the time of occurrence
of events called speciation or radiation.
4. It reflects the rate at which number of mutations becomes fixed. For
instance, the nucleotide sequence and amino acids sequence of
biomolecules (DNA, RNA, and proteins) is used for such calculations.
5. Molecular information about changes in the sequence of nucleotide in
DNA and RNA or in the sequence of amino acids in different proteins 97
Block 2 Evolutionary Biology-II
and the comparison of their rates helps in evaluating relationship
between distantly related organisms. This information helps in
preparation of phylogenetic trees and estimating proximate time of each
group from the other.
6. The number of amino acids modifications in the line of descent can be
used to estimate the time of divergence of two species from a common
ancestor.
• For example: the rate of amino acid replacement in all species over
evolutionary time is approximately constant. Based on that the amino
acids composition of protein could be compared in present day
organisms to measure the molecular changes that occurred in past
among them.
• The earlier in the past, an ancestor diverged into two present day
species means the more changes accumulated in the polypeptides of
these two species.
Disadvantage of Molecular Clock
Molecular clock gives better indications of evolutionary relationships than
structural changes between groups of related organisms. For instance,
Molecular clock data showed that dolphins are evolutionarily more closely
related to bats than shark and tuna fish.
SAQ 2
a) Who gave the neutral theory of evolution?
b) Choose the correct option for the statements given below:
i) What is one disadvantage with using molecular clocks?
1. The fossil record only goes back to 2 billion years ago.
2. Rates of change may be different in different organisms.
3. The fossil record has too many gaps.
4. They cannot be calibrated.
ii) Molecular clocks are calibrated using:
1. Phylogenetic trees
2. Known events in the fossil record
3. Genome sequences of organisms that do not create fossils
4. Unknown events in the fossil record
iii) Molecular clocks measure which of the following?
1. number of changes in nucleotide sequence
2. change in rate of time
98
Unit 5 Molecular Clock
3. change in species composition over time
4. number of changes in fossils record
iv) A molecular clock is
1. A piece of DNA that accumulates neutral mutation at
consistent pace
2. An organelle that provides energy and protein
3. Any microscopic biomolecules used to estimate evolutionary
relationship
4. A gene present in eukaryotes that initiate cell division
v) Which of the following is NOT a characteristic of a molecular
clock?
1. It acquires good and bad mutations at a constant rate
2. It acquires neutral mutations at a constant rate
3. It has the same function across many organisms
4. It is present in a large number of organisms
vi) Using molecular clock, it was estimated that two species A and B
must have diverged from their common ancestor about 9×106
years ago. If the rate of divergence per bases pair estimated to be
0.0015 per million years. What is the proportion of base pair that
differ between two species now?
1. 0.0270
2. 0.0135
3. 0.00017
4. 0.0035
5.7 COMPLICATION IN INFERRING
PHYLOGENETIC TREES
Inference of phylogenetic relationships is critical with multiple problems. To
find a true phylogenetic tree which accurately shows true evolutionary
relationships, multiple factors need to be carefully reviewed. Some of the
common problems are mentioned below:
1. Non-Phylogenetic Signal: There can be various reasons for non-
phylogenetic signal
• Speciation events which are closely spaced result in poor
phylogenetic signals resulting in short branches in the phylogenetic
tree. 99
Block 2 Evolutionary Biology-II
• Speciation events which occurred a long time ago leads to longer
terminal branches having multiple substitutions occurring at the
same position. These long branches show merged lineages which
are not evolutionarily related.
• Ingroup species which are fast evolving merge without group
which is the longest branch in the phylogenetic tree depicting false
relationships. These long branches can be wrongly interpreted or
not detected in phylogenetic trees known as long branch artifacts
resulting in spurious phylogenetic signals.
• Insufficient or incorrect signals are generated from long gene
sequences.
2. Substitution model: Amino acid/nucleotide of protein and DNA changes
with each generation however, rate of evolutionary divergence of a given
amino acid/nucleotide site of a gene and protein is not constant
(Heterotrachy). Site specific evolutionary rate differ among and
accumulation of variations or changes between different lineages is not
uniform. It varies because of random chance, variable mutation rate or
selection pressure. Evolutionary models fail to capture variation in
mutation rates. Tree reconstruction should take into account variations
in rates of divergence between lineages.
3. Main reason for Orthologous sequence artifacts (Non-vertical
evolution)
Evolution represented by branching is an over simplification which can
be misinterpreted. Recombination brings variation in lineage which is
not always vertical and that is difficult to be represented as a tree.
Lateral gene transfer is the process which moves DNA across large
evolutionary distances. Similarly, it is difficult to consider duplication,
hybridization, genetic drift, deletion, domain shuffling and gene
conversion while inferring trees. All these can bring complexity. To study
phylogenetic relationships, it is very important to select the suitable
OTUs (gene) for phylogenetic analysis.
4. Incomplete lineage sorting leads to error in phylogenetic relationships
which is the result of retention of ancestral polymorphism during
recurring speciation events. There are a number of reasons for
incomplete lineage sorting.
• Inclusion of sequences which are not showing true species
phylogeny.
• Wrong method of sequence evolution is used to assess multiple
substitutions.
• Inaccurate identification of orthologous sequences.
• Incorrect sequence alignments.
• Erroneous reconstruction of multiple substitutions which happen
due to evolutionary method models.
100
Unit 5 Molecular Clock
5. Random sampling error due to usage of a large multi gene dataset.
Two ways to form trees, one approach uses a super matrix which
combines orthologous genes into a single supergene. Another way is the
super tree approach which infers an individual tree for each gene in the
dataset, and then these individual trees are combined to form a single
supertree.
6. The development of organs or other body parts among distinct species
that are similar to one another and serve the same functions but did not
share a common ancestor is known as homoplasy. Homoplasy brings
difficulty in inferring trees. For this, character states need to be identified
which depicts true evolutionary relationships.
In addition to improve the results of phylogenetic trees which show true
evolutionary relationships, it is very important to improve quality of alignment
wherein an important step is selection of orthologous genes which are not
saturated. Saturation of aligned sequences is due to multiple substitutions
which results in distances that underestimate the real genetic distances. To
avoid this, datasets need to be selected which are not completely saturated.
Nucleotide sequences saturate faster than protein sequences. This can be
done by selecting single copy genes like mitochondrial DNA or pre-selected
slowly evolving genes like ribosomal RNA. These problems in inferring trees
can be resolved by inclusion of a bigger dataset having a large number of
species and a real method of sequence evolutionary model.
SAQ 3
State if the statements given below are ‘True’ or ‘False’:
a) In case of non-phylogenetic signal, insufficient or incorrect signals are
generated from short gene sequences.
b) Homoplasy is the result of convergent evolution.
c) Lateral gene transfer is the process which moves DNA across large
evolutionary distances.
d) Protein sequences saturate faster than nucleotide sequences.
5.8 SUMMARY
• Molecular evolution has a constant rate per unit time i.e. it shows a
molecular clock.
• Rate of accumulation of mutational changes in genome and proteins is
constant over a period of time.
• The rate of change in nucleotide and amino acids sequence of DNA,
RNA, and proteins respectively, is proportional to evolutionary
divergence among the species.
101
Block 2 Evolutionary Biology-II
• Greater the mutational differences between organisms, the greater is
their evolutionary divergence.
• Assumptions of molecular clock.
• The mutations at molecular level are incorporated at fixed or regular
rates over a time.
• The proportion rate of fixation of mutation in one gene relative to the
rates of fixation of mutation in other genes remains the same throughout
any line of descent.
• The line of descent leading from a common ancestor to all its
descendants has similar rates of fixed mutation.
5.9 TERMINAL QUESTIONS
1. Define molecular clock and DNA sequence divergence. What is their
relationship to each other?
2. A new insect species is discovered on an island. The species features
are compared to mainland species to determine the most closely related
species. Genetic analysis is performed on a chitin gene of both the
island species and its mainland relative. The gene in both species is
4,000 bases long and there are 10 differences in the bases between the
two species. If the estimated mutation rate for this family of insects is
.000001 per nucleotide per year, how long ago did the common ancestor
of both species exist?
a) 10000 years
b) 2500 years
c) 400 million years
d) 10 million years
3. Discuss the complication in inferring phylogenetic tree.
4. Discuss the difference between molecular drive and molecular
divergence.
5. Based on the DNA sequence ATTCGCATT, which of the following
sequences most likely came from its closest relative?
a) ATTCGCACT
b) TAACGCATT
c) CTTCCCATA
d) CCCGCGATT
102
Unit 5 Molecular Clock
5.10 ANSWERS
Self-Assessment Questions
1. a) Molecular clocks
b) neutral
c) molecular drive
2. a) Motoo Kimura
b) i) Rates of change may be different in different organisms
ii) Genome sequences of organisms that do not create fossils
iii) number of changes in nucleotide sequence
iv) Any microscopic biomolecules used to estimate evolutionary
relationship
v) It has the same function across many organisms
vi) 0.0270
3. a) False
b) True
c) True
d) False
Terminal Question
1. Refer to Sections 5.3 and 5.4.
2. 10000 years
3. Refer to Section 5.7.
4. Refer to Sections 5.4 and 5.5.
5. ATTCGCACT
103
Block 2 Evolutionary Biology-II
UNIT 6
DIVERGENCE OF BACTERIAL
AND ARCHAEAL GENOMES
Structure
6.1 Introduction 6.4 Divergence of Baceria and
Archaea
Objectives
Divergence of Replication
6.2 Origin and Diversification
System
of Bacteria and Archaea
Divergence of Transcription
Bacterial and Archaeal
System
Diversity Based on
Morphological Attributes Divergence of Translation
System
Bacterial and Archaeal
Diversity Based on Comparison of
Preferred Range of Characteristics of Bacteria
Environmental Conditions and Archaea
Bacterial and Archaeal 6.5 Applications
Diversity Based on
Medicine and Industry
Metabolic Properties and
Production of Secondary Biotechnology and Research
Metabolites 6.6 Summary
Bacterial and Archaeal 6.7 Terminal Questions
Diversity Based on
6.8 Answers
Molecular and Genetic
Features
6.3 Nature of Genomes
Bacterial Genomes
Archaeal Genomes
6.1 INTRODUCTION
There is general agreement among scientists that the Last Universal Common
Ancestor (LUCA) probably lived some 3.8-3.5 billion years ago (bya) and
represents the most primitive organism from which all organisms now living
104 might have evolved. Prokaryotes appeared around 3.5 billion years ago while,
Unit 6 Divergence of Bacterial and Archaeal Genomes
eukaryotes appeared much later. Though all prokaryotes fall in the size range
of 0.5-5.0 micrometres and share many characteristics, analysis of ribosomal
RNA from several species of prokaryotes led Carl Woese to split prokaryotes
into two separate lineages, eubacteria and archaebacteria, which were later
renamed as bacteria and archaea respectively. Biologists now divide all living
organisms into three domains: Bacteria, Archaea, and Eukarya.
Both bacteria and archaea are cosmopolitan in distribution. Though majority of
the species are mesophiles, there are several others which can be grouped
under the category of extremophiles, occupying environments which present
extremes of conditions be it temperature, salinity, hydrostatic pressure, or
anoxic conditions. Prokaryotic diversity has been described using
morphological and biochemical features as well as molecular characteristics.
Studies involving gene- and whole genome-sequencing have been very useful
in elucidating not only evolutionary relationships but also divergence times of
different lineages.
Prokaryotes have played key roles in shaping the evolutionary history of life on
earth, three events being the most significant. One such event involved
cyanobacteria, which by their photosynthetic activity, converted the primitive,
reducing atmosphere of the earth to an oxidising one, which further led to the
evolution of aerobic type of metabolism in organisms.
Another crucial event occurred with the engulfment of a facultative anaerobe
by a eukaryotic ancestor, which ultimately gave rise to mitochondria
responsible for energy production by oxidative phosphorylation. Third event,
equally significant, was the uptake of cyanobacteria by eukaryotic ancestors,
giving rise to chloroplast found in all plants and algae.
Prokaryotes, once thought to be small, simple organisms, exhibit extraordinary
biochemical capabilities (presence of specialised enzymes). They are able to
produce a variety of specialised, complex metabolites, using the ordinary
primary metabolites, which are finding various applications.
Objectives
After studying this unit, you would be able to:
comprehend the division of prokaryotes into bacteria and archaea;
understand the important differences between the bacteria and archaea
as revealed by molecular techniques;
explain the extensive diversity in bacteria and archaea in terms of
morphological features, metabolic capabilities, and occupation of a
variety of habitats;
comprehend the divergence in information processing systems that led
to a split in the prokaryotes, separating bacteria from archaea;
explain the specialisation of ribosomal proteins that occurred early in
evolution leading to bacterial and the archaeal lineage;
describe the key role played by the prokaryotes in the evolution of
105
Block 2 Evolutionary Biology-II
aerobic metabolism in organisms, photosynthesis and biogeochemical
cycles of the earth; and
explain the role of prokaryotes as the source of a number of important
enzymes and secondary metabolites which find applications in medicine,
biotechnology, and industry.
6.2 ORIGIN AND DIVERSIFICATIONOF
BACTERIA AND ARCHAEA
6.2.1 Bacterial and Archaeal Diversity Based on
Morphological Attributes
Morphological features of prokaryotes enable us to identify species which can
not only reveal microbial diversity but also help in diagnosis of disease and
prevention of infection. Characteristics studied mostly include cell shape, cell
wall, movement, flagella, Gram staining etc. Four basic shapes can be seen in
bacteria: spherical (cocci), rod-shaped (bacilli), arc-shaped (vibrio), and spiral
(spirochaete) as first described by Antony van Leeuwenhoek in 1675 using
light microscopy. Observations of cell morphology at a higher resolution (of the
order of 2 Angstroms) became available with the development of electron
microscopy some 300 years later. Bacteria exhibit enormous diversity in terms
of their physical features as can be appreciated from Table 6.1 below, which
lists different phyla and classes with their representative examples.
Table 6.1: List of different phyla and classes of bacteria with their
representative examples and their mode of life.
Bacteria Example Micrograph
Phylum Proteobacteria Rhizobium:
Class Alpha
Endosymbiont in roots
Proteobacteria
of leguminous plants
May be photoautotrophic, and responsible for
symbionts or parasitic. fixation of atmospheric
nitrogen.
It has been postulated that
aerobic rickettsial ancestor Rickettsia prowazekii: (i) Rickettsia
parasitized eukaryotic cells
Obligate, intracellular
and gave rise to
parasite, responsible
mitochondria.
for epidemic typhus.
106
Unit 6 Divergence of Bacterial and Archaeal Genomes
Beta proteobacteria Nitrosomonas: Some
species oxidise
Some species of bacteria
ammonia into nitrite.
of this group play a role in
the nitrogen cycle. Spirillium minus is
responsible for rat-bite
fever.
(ii) Spirillum
Gamma proteobacteria E. coli: normally has a
beneficial role in the
Some species are
human gut though
pathogenic, while others
some strains are
play a beneficial symbiotic
pathogenic.
role in the human gut.
(iii) Escherichia coli
Vibrio cholerae is the
causative agent of
cholera
(iv) Vibrio cholerae
Salmonella: some
strains cause typhoid
fever, while some other
strains can cause food
poisoning.
(v) Salmonella
Azotobacter spp. are
soil/water/sediment
inhabiting bacteria with
the ability to fix
nitrogen non-
symbiotically and have
been in use as a
biofertilizer for over a
(vi) Azotobacter
century.
107
Block 2 Evolutionary Biology-II
Delta proteobacteria Desulphovibrio
vulgaris: anaerobic,
Some species can produce
capable of reducing
multicellular fruiting bodies
sulphate.
for spore formation to
overcome unfavourable
conditions.
Also includes sulphate-
(vii) Desulfovibrio vulgaris
reducing species.
Myxobacteria include Sorangium cellulosum
motile bacteria capable of is a soil-dwelling
gliding movements bacterium.
This grouping includes
bacteria which can digest
cellulose and have been
found to reside in a number
(viii) Sorangium cellulosum
of niches such as soil,
decaying vegetation, hot
springs, organic matter and
faeces of ruminants. Micrococcus roseus
resides within the gut
of termite and is
capable of producing
several hydrolytic
enzymes which can
digest cellulose.
(ix) Micrococcus luteus
Micrococcus luteus is
Gram-positive and
non-motile coccus, can
be seen arranged in
tetrads. Shows various
modes of life such as
saprophytic,
commensal or
opportunistic
pathogens
Epsilon proteobacteria Campylobacter: Some
species are
Includes members
responsible for causing
occupying deep sea
blood poisoning and
hydrothermal vents (where
intestinal inflammation.
they can oxidise sulphur
and reduce nitrogen oxide) (x) Campylobacter
as well as species which
108
Unit 6 Divergence of Bacterial and Archaeal Genomes
are opportunistic Helicobacter pylori:
pathogens.
Colonises the lower
part of the stomach
and duodenum, and
plays a role in the
development of (xi) Helicobacter pylori
gastritis,
peptic/duodenal ulcers.
Phylum Chlamydias Chlamydia psittaci
causes psittacosis, C.
All members of this group
pneumoniae causes
lack peptidoglycan in their
respiratory tract
cell walls and lead obligate
infections, and C.
intracellular parasitic
trachomatis strains are
modes of life.
responsible for (xii) Chlamydia
chlamydia,
conjunctivitis,
trachoma,
lymphogranuloma
venereum.
Phylum Spirochetes Treponema pallidum
syphilis is responsible
Majority of the members
for causing syphilis.
are free living (some are
pathogenic), having Borrellia burgdorferi is
anaerobic type of the causative agent of
metabolism. Cells are Lyme disease.
spiral shaped and flagella (xiii) Treponema pallidum
run lengthwise in the syphilis
periplasmic space between
the inner and outer
membrane.
Phylum Cyanobacteria Prochlorococcus being
photosynthetic is
Also known as blue-green
responsible for
algae, occur abundantly in
producing enormous
all kinds of environments
amounts of oxygen.
such as terrestrial, marine,
and fresh water and are Nostoc includes a
photosynthetic. number of species
which have different
morphology, habitat (xv) Scytonema
distribution and
ecological roles and
can be found in both
109
Block 2 Evolutionary Biology-II
terrestrial and aquatic
environments as well
as existing
symbiotically inside
plant tissues and serve
to provide nitrogen to
its host plant.
Phylum Firmicutes Clostridium botulinum
contains is rod shaped,
anaerobic, spore-
Gram-positive bacteria
forming, found in soil,
Members of this group dust and sediments of
have a thick cell wall water bodies. It
(a) (b)
without an outer produces a neurotoxin
membrane. botulinum which
Class Clostridia causes botulism. (xvi)
They are anaerobic (some (a) Clostridium
members are capable of C. tetani, commonly Botulinum (b) Clostridium
forming highly resistant found in soil, is Tetani
spores). Species may be responsible for causing
found inhabiting soil, tetanus.
aquatic bodies or raw
meat/poultry.
C. perfringens can be
found on raw meat and
poultry and in the
intestines of animals
and is the causative
agent of food
poisoning.
Streptomyces has over
500 species of
filamentous bacteria
occurring in soil. They
are the source of a
large number of
bioactive compounds
such as antibacterial,
antifungal,
(xvii) Streptomyces
antiparasitic,
immunosuppressants
and extracellular
enzymes. Some
species are pathogenic
in plants.
110
Unit 6 Divergence of Bacterial and Archaeal Genomes
Phylum Firmicutes Mycoplasma
pneumoniae, causes
Class Mollicutes
pneumonia, and M.
Mycoplasma represent the hominis and
smallest bacteria which Ureaplasma
lack a cell wall. It includes urealyticum are (xviii) Mycoplasma
free-living species as well responsible for some
as commensal and genitourinary diseases.
pathogenic species.
Lactobacillus is a
Class Bacilli
probiotic species,
Includes members which found in fermented
are facultative anaerobes dairy products and
and some species can form contributes to gut
spores. health and boosts
immunity.
Streptococcus,
Listeria,
Staphylococcus,
Bacillus are some of
the pathogenic species
capable of causing
infection.
Planctomycetes In Gemmata
obscuriglobus
These bacteria occupy a
membranes divide the
variety of habitats such as
interior of the cell into
marine and freshwater
separate
ecosystems, soils, even
compartments,
extremely dry soils of
condensed nucleoid (xix) Gemmata
Atacama desert and form
surrounded by obscuriglobus
living stromatolite microbial
membranes.
mats in the marine shallow
waters with extreme Anammox group of
salinity. All planctomycetes planctomycetes exhibit
lack peptidoglycan in their the process of
cell wall, making them generating energy by
resistant to beta-lactam anaerobic oxidation of
antibiotics, including ammonia and plays an
penicillins. They reproduce important role in
by budding rather than by nitrogen cycling
binary fission. globally.
Though majority of bacterial species show one of the basic four shapes and
are unicellular without intracellular compartmentalisation, there are some
exceptions too. Multicellularity encountered in Actinobacteria and
Myxobacteria is simple type of cell aggregation. In Streptomyces genus,
hyphae are formed by germinating spores. Hyphae have multinuclear aerial
111
Block 2 Evolutionary Biology-II
mycelium, having septa at intervals. This creates a chain of uninucleated
spores (Table 6.1 xvii). Multicellularity encountered in Cyanobacteria,
Chloroflexi and some proteobacteria such as Beggiatoa is in the filamentous
form arising due to cell division and adhesion (Table 6.1 xv). In phylum
Planctomycetes, membranes divide the interior of the cell into separate
compartments, condensed nucleoid surrounded by membranes (Table 6.1
xviii). Like Planctomycetes, phyla Verrucomicrobia and Chlamydia too show
the presence of internal membranes which, though superficially similar to the
endomembranes possessed by eukaryotes, bear no homology to the latter
and represent a case of convergent evolution.
Based on the cell wall structure and the ability to take up the Gram stain,
bacteria are divided into two groups, Gram-positive and Gram-negative (Fig.
6.1). Table 6.2 below compares some of the attributes of Gram-positive and
Gram-negative bacteria.
Fig. 6.1: Diagrammatic representation showing cell wall differences
between Gram-positive and Gram-negative bacteria.
Table 6.2: Differences between characteristics of cell wall of Gram-
positive and Gram-negative Bacteria.
Characteristic of cell Gram-positive bacteria Gram-negative bacteria
wall
Gram staining Retain primary stain after Lose primary stain after
alcohol treatment alcohol treatment
Number of layers Single; outer membrane Double; outer membrane
absent present to the outside of
peptidoglycan layer
Peptidoglycan layer Thick, about 20-80 nm Thin, about 5-10 nm
Periplasmic space Small, if present Large
between cytoplasmic
and outer membrane
Teichoic acid Present Absent
112
Unit 6 Divergence of Bacterial and Archaeal Genomes
#Porins Absent Present
Lipid content Very low (2-5%) Very high (15-20%)
Lipopolysaccharide Absent (except for Listeria Present
(LPS) monocytogenes)
Susceptibility to Sensitive (lysis of cells) Resistant (cells protected
lysozyme by the outer membrane)
Susceptibility to *Sensitive *Intrinsically more resistant
penicillin and
detergents
Examples of bacteria Bacillus anthracis, Salmonella species,
Staphylococcus aureus, Shigella species,
Streptococcus pneumoniae, Haemophilus influenzae,
Enterococcus faecalis, Escherichia coli, Neisseria
Mycobacterium tuberculosis, gonorrhoeae, Neisseria
Clostridium botulinum meningitidis, Pseudomonas
aeruginosa
#Porins are surface proteins which form channels to allow influx of nutrient
molecules (excluding antibiotics and inhibitors) and efflux of waste products.
* Although antibiotic-resistance genes in members of both the groups can be
acquired through HGT, it is the cell wall architecture of Gram-negative bacteria
which does not allow antibiotics to enter the cell easily.
Name Archaea has been derived from the Greek word archaios which means
ancient or primitive and appropriately so as some members belonging to
archaea do indeed show primitive characteristics. List of different phyla of
archaea with their representative examples and their mode of life is shown in
Table 6.3.
Table 6.3: List of different phyla of archaea with their representative
examples and their mode of life.
Archaea Example Micrograph
Phylum Euryarchaeota Methanogens: Produce
includes methanogens and methane as a by-
halophiles. product of metabolism,
which causes flatulence
in humans and other
animals.
Methanosaeta is
Methanogenic
acetotrophic, using
microorganisms
acetate fermentation to
produce methane
113
Block 2 Evolutionary Biology-II
Methanogens are the .
largest microorganisms
belonging to archaea and
show a cosmopolitan
distribution, occurring
mostly in anoxic
environments such as
aquatic sediments, rice
paddies, anaerobic
digesters, and the
gastrointestinal tract of
animals. They can grow at
a wide range of
o
temperatures (4 C to
o
100 C), salinities
(freshwater to brine) and
pH (3 to 9). They possess
several unique coenzymes.
Because of the unique cell
wall construction, they are
not susceptible to penicillin
and other antibiotics. They
are chemoautotrophic, H2
serves as the source of
both energy and electrons,
and CO2 works both as an
electron sink and the Halobacterium is found
source of cellular carbon. in waters of extreme
salinity and contains
bacteriorhodopsin in the
Halophiles grow best in membrane which
environments having salt imparts red colour to
concentrations between their blooms.
10% and 35% due to the
presence of adaptations
like efficient ion pumps, UV
absorbing pigments and
specific proteins to Halobacterium
withstand osmotic stress.
Phylum Crenarchaeota Sulfolobus: Several
species grow in volcanic
This includes members
springs at temperatures
which are sulphur- o
of 75 to 80 C and at a
dependent extremophiles
pH between 2 and 3.
and perform an important
function of carbon fixation.
Sulfolobus archaea
114
Unit 6 Divergence of Bacterial and Archaeal Genomes
Phylum Nanoarchaeota Nanoarchaeum equitans
is an obligate symbiont
contains only one species,
on Ignicoccus, a genus
Nanoarchaeum equitans,
of hyperthermophilic
the smallest known
archaea found in marine
organism with a diameter
hydrothermal vents.
of 400 nm (1/100th the size
of E. coli). It has the
smallest genome and lacks
the genes required to Nanoarchaeum equitans
synthesise amino acids,
Small dark spheres of
nucleotides, lipids and
Nanoarchaeum equitans are
cofactors.
in contact with the host cell,
Ignicoccus.
Phylum Korarchaeota Candidatus
Korarchaeum
cryptofilum, is an
It represents one of the ultrathin, filamentous
most primitive forms of life heterotroph,
and its members have metabolising peptides
been found to occur only in and protein into H2
the Obsidian Pool, a hot anaerobically.
spring at Yellowstone Korarchaeum cryptofilum
National Park.
BOX 6.1: Contributions of cyanobacteria on the evolution of life forms
on earth.
Cyanobacteria, one of the oldest among prokaryotic phyla (was already in existence
around 2.45-2.32 billion years ago [bya]), include a very large and diverse assemblage
of 374 genera and 1468 species and have been found to occupy marine, freshwater,
or terrestrial environments as well as varied climatic zones. Their evolution had a very
profound influence on the evolution of subsequent life forms on the earth due to
various reasons such as:
● Earth’s climate: To begin with, the earth had a reducing atmosphere consisting of
methane, carbon dioxide and water vapour. It is hypothesised that photosynthetic
activity of cyanobacteria led to the production of oxygen as a by-product which was
released into the sea water (Cyanobacteria have no membrane-bound organelles
and folds in the outer membrane of the cell are the site of photosynthesis).
Gradually as the amount of oxygen produced increased over a span of 200-300
million years, oxygen started escaping into the atmosphere where it could react
with methane, gradually displacing it and converting earth’s atmosphere into an
oxidising one. This event occurred probably between 2.4- 2.1 bya and is referred
to as the “Great Oxidation Event”.
● Ice age: Displacement of the greenhouse gas, methane, to a large extent, by
oxygen led to cooling down of temperatures globally resulting in the formation of
widespread ice sheets (one of the earliest ice ages) on earth.
● Ozone layer: UV radiation from the sun by acting on the atmospheric oxygen led to
the generation of an ozone layer near the upper part of the atmosphere.
Protective effect of the ozone layer by blocking harmful UV radiation from reaching
the earth allowed colonisation of the surface of the ocean and ultimately the land
by living forms.
115
Block 2 Evolutionary Biology-II
● Aerobic metabolism: It is postulated that anaerobic organisms already inhabiting
the earth were adversely affected by the presence of oxygen contributed by the
photosynthetic activity of cyanobacteria, around 2.4-2.1 bya, wiping them out in
large numbers. This was followed by the evolution of aerobic metabolism by
organisms, oxygen serving as the final electron acceptor to generate energy after
nutrient breakdown as well as evolution of enzymes to detoxify reactive oxidative
species arising due to aerobic metabolism. Aerobic respiration paved the way for
the evolution of more complex life forms.
● Biogeochemical flux: Being primary producers, and present in great abundance
and diversity, various species of cyanobacteria play an important role in the
biogeochemical cycling of carbon, nitrogen, and phosphorus in a very
significant way besides being an important link in food webs.
● Evolution of chloroplasts: Endosymbiosis of a cyanobacterium within a unicellular
eukaryote is postulated to have given rise to chloroplasts found in algae and
plants.
On the basis of analysis of dataset of 16S rRNA sequences of several cyanobacterial
species, researchers have postulated that evolution of multicellularity occurred
around the time of “Great Oxidation event”. Multicellular level of organisation in the
form of filaments might have proved advantageous by improving motility and metabolic
fitness as compared to single cells and could have contributed to abundance and
diversification of cyanobacteria into newer ecological niches. Filamentous
cyanobacteria exhibit directional growth as well as intercellular communication and
exchange of resources indicative of evolution of division of labour and terminal cell
differentiation.
6.2.2 Bacterial and Archaeal Diversity Based on
Preferred Range of Environmental Conditions
According to the preferred habitats, bacteria and archaea can be classified
into various groups as shown in Table 6.4 below.
Table 6.4: Classification of bacteria and archaea into various groups based on
their preferred habitats
Category Preferred Examples of Bacteria Examples of
range Archaea
Mesophiles 25°C-40°C and Escherichia coli, Methanobrevibact
Listeria er smithii
pH 5-9
monocytogenes,
Staphylococcus aureus
Thermophiles 50-55°C Thermus aquaticus
Thermus thermophilus
Hyperthermophiles 80°C Thermotoga maritima, Methanopyrus
Chloroflexus sp., kandleri
Thiobacillus sp.
Sulfolobus
solfataricus
Pyrodictium
abyssii
116
Unit 6 Divergence of Bacterial and Archaeal Genomes
3
Pyrodictium
occultum
3
Pyrolobus fumarii
1 5
Psychrophiles 10-15°C Moraxella sp., Methanococcoide
Pseudomonas sp., s burtonii,
Vibrio sp., Methanogenium
Flavobacterium sp., frigidum
Bacillus sp.
1
Psychrotrophs 15-25°C Pseudomonas sp.,
Acinetobacter sp.,
Listeria
monocytogenes
2 4
Piezophiles high Shewanella benthica, Pyrococcus
hydrostatic Photobacterium yayanosii
pressure in the profundum,
depths of Psychromonas sp.
ocean, 38
megapascals
(MPs)
8
Acidophiles and pH 0.5-5.0 Thiobacillus Picrophilus
thermoacidophiles ferrooxidans oshimae
Pyrodictium
abyssi,
Ferroplasma
acidiphilum
6
Halophiles 15-23% salt Halorhodospira Halobacterium
6
concentration halophila, Salinibacter salinarum,
ruber Haloferax
mediterranei
7
Alkaliphiles and High salinity Bacillus halodurans Natrialba
Haloalkaliphiles and pH 9-12 magadii,
7
Halorubrum
vacuolatum
Hyperthermophiles show optimal growth at 80°C or higher but are also
capable of growing at temperatures up to 105°C while thermophiles grow
optimally at temperatures ranging from 50° to 70°C.
1
can grow even at 0°C.
2
can be accompanied by either high temperature as in hydrothermal vents
(thermophiles) or permanently cold conditions as in the deep oceans
(psychrophiles).
3
grow fastest at 105oC and can withstand autoclaving for one hour at 121°C.
117
Block 2 Evolutionary Biology-II
As most acid environments have high temperatures, thermophiles and
acidophiles may be grouped together.
6
grows in saltern crystallizer ponds.
7
grow in Lake Magadi, pH of the brine being 10.
8
can grow in 1.2M sulphuric acid at 60°C.
4
can tolerate pressures up to 150MPs (as compared to a pressure of 0.1 MPs
at sea level) and is a thermophile, found near hydrothermal vents.
5
resides in Ace lake in Antarctica at temperature of 1 to 2°C.
Not all archaea are extremophiles, and are found to play important ecological
roles in a variety of habitats. For instance, organisms belonging to
Crenarchaeota are quite abundant in soil (where they function to oxidise
ammonia) as well as in the oceans of the world constituting a good proportion
of the planktonic organisms.
Organisms belonging to Euryarchaeota include those dwelling in deep-sea
sediments where they function to oxidise anaerobically methane stored in
these sediments as well as methanogens living in terrestrial anaerobic
environmental conditions which are responsible for a large percentage of
methane emissions globally.
Box 6.2: Prokaryotic adaptations under different environmental
conditions.
Biomolecules, especially proteins, generally remain stable and functional within a
narrow range of temperature and pH conditions. As a number of prokaryotic species
are extremophiles, several adaptations have arisen which allow the cellular proteins to
remain stable and active under the extreme conditions.
● Thermophilic adaptations of proteins: At high temperatures, proteins without
appropriate modifications show changes in their folding patterns which expose the
hydrophobic cores resulting in aggregation and loss of activity. Proteins in these
organisms show higher numbers of hydrophobic amino acids, disulphide bonds
and ionic interactions so as to retain protein structure and activity at high
temperatures.
● Acidophilic adaptations of proteins: Species inhabiting acidic environments
maintain cytoplasmic pH of 5.0 to 6.5 by pumping protons out of the cell. As under
conditions of low pH many charged polar amino acids become protonated
changing their charges, enzyme structure can get disrupted making the enzymes
non-functional. Acidophiles possess enzymes and other proteins which can retain
optimal structure and catalytic function under highly acidic conditions due to the
predominance of acidic amino acids (glutamic and aspartic acid) on the surface of
these enzymes and proteins. As most of the acidophiles experience high
temperature, their proteins show thermophilic adaptations too.
● Halophilic adaptations of proteins: Proteins show a higher content of acidic amino
acids which confer increased negative charge on the surface of proteins. Thus,
extreme ionic conditions can be compensated and stability and activity of proteins
is maintained.
• Membrane system of halophiles allows them to pump potassium in and to pump
sodium out to cope up with extreme osmotic stress and maintain osmotic balance.
In different species of halobacteria, intracellular concentration of potassium may
118
Unit 6 Divergence of Bacterial and Archaeal Genomes
have a range of 1.2 to 4.5 M.
● Psychrophilic adaptations of proteins: Reduced hydrophobic core and less charged
surface of proteins ensures flexibility and functionality at low temperatures.
● Piezophilic adaptations of proteins: Piezophilic adaptations of proteins include
possession of a dense, compact hydrophobic core, predominance of smaller amino
acids capable of forming hydrogen bonds with a concomitant decrease in large
hydrophobic residues and a multi-unit organisation with individual monomers tightly
packed helping to protect hydrogen bonds. This pattern of amino acid residues
allows the protein to pack very tightly conferring stability to the protein structure
under high pressure. Thermophilic adaptation of proteins employs higher content
of basic amino acids like arginine.
Temperature, pH, and hydrostatic pressure also affect the function of the cell
membrane by impacting the fluidity and permeability of the cell membrane for diffusion
of nutrients. Therefore, organisms try to maintain membrane integrity continuously
despite variations in external environmental conditions by modifying lipid composition
of their membranes. For example, psychrophilic bacteria tend to have higher
proportion of unsaturated and shorter-chain fatty acids in their plasma membrane so
as to maintain flexibility at low temperature while thermophilic bacteria contain higher
proportion of saturated fatty acids in their cell membrane so as to remain stable at high
temperature. Thermophilic and hyperthermophilic archaea achieve higher
thermostability of their cell membranes by having repeating subunits of the C5
compound, phytane (branched saturated isoprenoid) joined by ether linkage (as
against ester linkage of phospholipids found in bacterial cell membranes) (Fig. 6.2) as
well as monolayer structure constituted by membrane-spanning tetraether lipids in
order to avoid separation of inner and outer layers of a membrane bilayer which can
occur at very high temperature. This structure is almost impermeable to ions and
protons helping to maintain membrane integrity under conditions of extreme
temperature and acidity.
Fig. 6.2: Comparison between membrane lipids of archaea and bacteria.
6.2.3 Bacterial and Archaeal Diversity Based on
Metabolic Properties and Production of Secondary
Metabolites
Metabolic properties such as the ability to use different sources of energy,
specific requirements of growth factors, ranges of various environmental
factors conducive to growth and susceptibility to different classes of antibiotics
have been used to distinguish different species of microorganisms. Production
of secondary metabolites is employed by several microbial species in order to
compete with other microorganisms in their surroundings. As production of
most secondary metabolites has been found to be species- and/or strain-
specific, it can be used to classify different species (Table 6.5).
119
Block 2 Evolutionary Biology-II
Table 6.5: Secondary metabolites produced by bacteria.
Organism Substance/metabolite produced
Streptomyces avermitilis antibiotic Avermictin
S. griseus antibiotic streptomycin
S. bingchenggensis anthelmintic Milbemicin
S. venezuelae antibiotic Chloramphenicol
S. kanamyceticus antibiotic Kanamycin
Clostridium botulinum neurotoxin botulinum
C. tetani exotoxins, tetanolysin and tetanospasmin
C. perfringens enterotoxin CPE)
Sorangium cellulosum Soce56 epothilone
Angicoccus disciformis myxochelin A
6.2.4 Bacterial and Archaeal Diversity Based on
Molecular and Genetic Features
Molecular and genetic features have been studied to investigate microbial
diversity using several approaches such as nucleic acid hybridisation, DNA
cloning and sequencing, RFLP and other PCR-based methods. Molecular
approach for studying diversity, being sequence-specific, also allows us to
gain an insight into the evolutionary relationships among the species sampled.
Phylogenetic analysis using the 16S rRNA gene has been very extensively
used in the classification of different species of prokaryotes (Fig. 6.3). Several
reasons support the use of the 16S rRNA gene: every prokaryotic genome
has at least one copy of this gene, sample identification using PCR is possible
due to the presence of conserved regions in this gene and information
provided by its sequence has been found to be reliable in microbial
classification.
Woese’s delineation of Archaea as a third domain of life, bacteria and eukarya
being the other domains, is supported by molecular genomics and
phylogenetics data. According to these studies, hyperthermophilic and non-
methanogenic ancestors could have given rise to present-day archaeal
lineages followed by divergence into two major phyla, the Crenarchaeota and
the Euryarchaeota (Table 6.3).
Based on analysis of 16S rRNA gene homologies, one study estimates that
there may be 1.4-1.9 million extant bacterial lineages. Although single gene
comparisons have been useful, new approaches such as whole genome
120
Unit 6 Divergence of Bacterial and Archaeal Genomes
sequencing and generation of metagenomic data (analysis of all the DNA
present in a given sample), are being increasingly employed to gain
comprehensive knowledge of prokaryotic diversity and genetic relationships.
Availability of whole genome sequences not only allows determination of
orthologous genes but also presence or absence of any specific genes in any
particular genome (giving an insight of metabolic functions of the organism).
For instance, information about functional diversity and identification of
species can be elucidated by using specific genes coding for enzymes such as
nitrogenase, nitrate reductase etc. in DNA microarrays combined with DNA-
DNA hybridisation.
Fig. 6.3: Phylogenetic tree of life based on rRNA sequences.
SAQ 1
a) Write true or false against the following statements.
i) Bacteria have two layers of phospholipids in their cell membranes.
ii) Gram-positive bacteria have a single cell wall surrounded by an
outer membrane containing lipopolysaccharides.
iii) The cell wall of Gram-positive bacteria is thick and composed of
peptidoglycan.
iv) Porins allow entry of substances into only Gram-positive bacteria.
v) Cell wall of Gram-positive bacteria is anchored to the cell
membrane by lipoteichoic acid.
b) State the appropriate reason for the following statements
i) Psychrophiles seldom cause disease.
ii) Food can get spoiled even under refrigerated and freezing
conditions.
iii) Pyrolobus fumarii can survive at temperatures as high as 113oC.
iv) Presence of multiple origins of replication in several species of
archaea is advantageous.
121
Block 2 Evolutionary Biology-II
v) Infections caused by Mycoplasma species continue to persist
despite prolonged treatment with beta-lactam antibiotics.
6.3 NATURE OF GENOMES
6.3.1 Bacterial Genomes
Complete set of genetic information contained in the DNA of the chromosome
of an organism constitutes its genome. Clear understanding of the biology and
evolutionary relationships of various bacterial species can be achieved from
the sequencing data of bacterial genomes. Additionally, genome sequencing
helps in clarifying the basis of antibiotic-resistance and pathogenicity of
different species of bacteria as well as in identifying novel targets of
antibiotics. This is of immense importance as the majority of pathogenic
bacteria have evolved resistance to most of the commonly used antibiotics.
Haemophilus influenzae was the first bacterium whose genome was
sequenced by J. Craig Venter in 1995. Subsequently, as sequencing
technologies improved, more and more genomes could be tested with the
result that at present nearly 30,000 sequenced bacterial genomes from 50
different phyla are publicly available. Some of the medically important species
whose genomes have been completely sequenced include Streptococcus
pneumoniae, Mycobacterium tuberculosis, Escherichia coli O157:H7, Vibrio
cholerae, Clostridium difficile and Staphylococcus aureus. Rapid development
of newer technologies for next generation sequencing is also greatly
contributing to the bacterial genome data. Some of the main features that have
emerged from bacterial genome data are briefly described below.
Majority of bacteria have a single, circular chromosome consisting of a double-
stranded DNA molecule which is compacted into a structure called nucleoid,
there being no nuclear membrane and hence bacteria are categorised as
prokaryotes. Proteins facilitate packing of the chromosome into a nucleoid
which occupies nearly ⅓ space of the cell’s interior. Since the length of the
chromosome (approximately 1.5 millimetres) is nearly 500 times the size of the
bacterial cell (approximately 1-2 micrometres in length), supercoiling and tight
packing of the chromosome into the nucleoid area leave space within the cell
for other processes like cell metabolism and protein synthesis to take place.
However, there are a few exceptions too as some bacteria such as Vibrio,
Burkholderia, Leptospira, Brucella and Deinococcus species have two or more
chromosomes and Borrelia burgdorferi has a linear chromosome rather than a
circular one.
As there is only a single chromosome in the majority of bacterial species, they
are haploid and there is no possibility of masking the recessive alleles.
However, about 10% of bacterial species have multiple copies of individual
chromosomes i.e., they are polyploids e.g., Synechococcus elongatus,
Azotobacter vinelandii, Deinococcus radiodurans, Sinorhizobium meliloti.
122
Unit 6 Divergence of Bacterial and Archaeal Genomes
Typical genome size of bacteria is 5MB, though the size varies greatly in
different bacterial species and even among different strains within a species;
5.1Mbp found in Bacillus megaterium, 13,033,799 bp in Sorangium cellulosum
So ce56 and 14,782,125 bp in Sorangium cellulosum strain So0157-2;
Haemophilus influenzae HK1212 has 1.0Mbp as against the strain F3047 with
2.0 Mbp; Burkholderia pseudomallei THE has 6.3 Mbp as against the strain
MSHR520 with 7.6 Mbp. Deinococcus radiodurans R1 has two chromosomes
with 2,648,638 and 412,348 bp, occurring in multiple copies and a
megaplasmid composed of 177,466 bp and a small plasmid having 45,704 bp,
total genome amounting to 3,284,156 bp.
An obligate symbiont, Nasuia deltocephalinicola strain NAS-ALF has the
smallest genome with only 112,091 bp. Tremblaya princeps has the second
smallest genome consisting of 139,000 bp.
Thus, genome size in bacteria varies due to acquisition of genes by horizontal
gene transfer or loss of functional accessory genes due to long term
association with the hosts.
BOX 6.3: Plasticity of Bacterial Genomes.
Examination of data on bacterial genomes indicates a major pattern in genome size
across different species of bacteria. In general, genome size tends to be larger in free-
living species than parasitic species which in turn have larger genomes than obligate
pathogens. For instance, studies have shown a progressive reduction in genome size
of rickettsial spp. from 1.5 to 1.1 Mb, with the greatest reduction and degradation of
genome encountered in the most virulent species when compared with the closely
related non-pathogenic species. This trend of convergent evolution shown by a
number of bacterial species pathogenic in humans is associated with a selective loss
of non-essential genes encoding enzymes required for amino acid synthesis and other
pathways involved in ATP, LPS and cell wall component biosynthesis but selective
retention of genes coding for recombination and DNA repair proteins and toxin-
antitoxin modules needed for evasion of host immune response as well as selective
expansion due to plasmids, short palindromic elements, ADP-ATP translocases, type
IV secretion system and duplication of gene families, enabling better adaptation to the
host-environmental conditions.
However, several studies document the dynamic nature of bacterial genomes as loss
or gain of genetic material through various mechanisms (unit 5) can affect genome
size even within related strains. Genomes of pathogenic strains of Enterococcus
faecalis were reported to be 25% larger than those of commensal strains due to the
presence of plasmids, phages and pathogenicity islands (acquired through HGT) (see
unit 5 for HGT mechanisms), responsible for their rapid evolution in an antibiotic-
treated patient. Thus, genetic content even within a species varies between non-
pathogenic strains and clinical isolates. This has also been confirmed by a study
comparing the genome sizes of 3 different [Link] strains: E. coli K-12, non-pathogenic,
4,639,221 bp.; E. coli O157:H7 strain EDL933, enterohemorrhagic, 5,528,445 bp. and
E. coli strain CFT073, uropathogenic, 5,231,428 bp.
In the large genome (>2 Mb) of Rickettsia endosymbiont of Ixodes scapularis (REIS),
transposons, insertion sequences and other mobile genetic elements constitute nearly
35% of the total genome.
If multiple genomes of the same species are available, calculation of pan genome
(entire set of all the genes from all strains of a clade) and core genome (set of
homologous genes present in all the tested genomes) can be achieved to derive
information and clearer understanding about species’ relatedness and evolution.
123
Block 2 Evolutionary Biology-II
Comparative analysis of 2000 E. coli genomes revealed that the core genome of E.
coli had nearly 3100 gene families (found in all E. coli genomes) and about 89,000
different gene families encountered in various strains and any one E. coli bacterium
contains less than 10% of the total number of E. coli genes in the E. coli pan-genome.
Nucleotide composition also shows variation between species: the G+C
(guanosine-cytosine) content is relatively uniform within a bacterial genus or
species and shows a range of about 25% in Mycoplasma species to about
75% in some Micrococcus species. In general, analysis of bacterial genomes
indicates that bacteria occupying complex environmental niches tend to have
larger genomes with a higher GC content as compared to host-associated
bacteria.
There are about 2,500 genes in a typical bacterial genome which code for all
the metabolic requirements needed for survival; in addition to chromosome,
bacteria also have plasmids which confer survival advantage to bacteria
against one or more antibiotics, (Bacillus megaterium strain QMB1551
harbours seven plasmids, the largest plasmid array in a single bacterial strain).
B. megaterium, strain QMB1551 while carrying 5300 genes on the
chromosome has additional 523 genes located on plasmids. Sorangium
cellulosum So ce56 with a very large genome contains 9367 protein-coding
genes while Sorangium cellulosum strain So0157-2 with the largest genome
analysed so far contains 11,599 genes. Small genomes are found in
mycoplasmas with only about 500 to 1000 genes. It has been postulated that
mycoplasmas evolved by degenerative evolution from Gram-positive bacteria
and show phylogenetic relationship with some clostridia. Nasuia
deltocephalinicola strain NAS-ALF with the smallest genome encodes just 137
proteins.
Genes in prokaryotes are arranged in operons which have evolved over time
due to the working of natural selection. Coordinate control of all the genes
which code for proteins that are related functionally serves to conserve energy
for the cell. For example, most of the bacterial genes coding for ribosomal
proteins are clustered in a few operons (present near the origin of replication)
which allows for coordinated regulation. Placed next to the large RP cluster in
the genome are several genes involved in synthesis of proteins and
chaperones.
Gene order and gene content are found to remain quite stable among related
species though mobile genetic elements may bring about changes in genome
architecture by introducing insertions, inversions, duplications, and
translocations.
Bacterial genomes are gene rich, 88% of the genome being represented by
protein-coding regions and there are also intergenic spaces (though very
limited in extent as compared to eukaryotic genome), repeated elements and
inactivated or otherwise functionless genes.
Whole genome sequencing allows to decipher completely the biology of an
organism. For instance, the genome of Streptococcus pneumoniae TIGR4, a
virulent strain (serotype 4, ST 205) has been completely sequenced and has
revealed important information about this pathogen. The genome of this strain
124
Unit 6 Divergence of Bacterial and Archaeal Genomes
of S. pneumoniae consists of 2,160,837 base pairs and 2,236 putative genes
which include several ATP dependent transporters. Some of these
transporters facilitate sugar transport indicative of ecological adaptation of the
bacteria to sugar rich environments of the oral cavity. Sequencing also
revealed a 13-gene cluster responsible for capsular biosynthesis, important for
virulence of the bacterium.
BOX 6.4: Deinococcus radiodurans, the bacterium with extraordinary
capabilities
Arthur W. Anderson first discovered Deinococcus radiodurans in 1956 when
he observed red bacteria in a can of ground meat which despite having been
sterilised with a megarad range of radiation, had got spoiled. Colonies appear
red because of the presence of carotenoid pigments within the bacteria.
D. radiodurans is not only capable of tolerating ionising radiation of very high
intensity at a dose of 5,000 Gy (1000 times the lethal dose for a human, and
which can kill virtually all other microorganisms), but also desiccation, UV
radiation, oxidising agents and electrophilic mutagens.
As there have not been any highly radioactive habitats on the earth over large
geologic times, radiation does not seem to be responsible for the evolution of
resistance to continuous exposure to radiation in organisms. However, other
factors such as UV radiation, oxidising agents or alternating periods of
desiccation and hydration or high and low temperatures are quite effective in
causing DNA damage and it is postulated that these organisms evolved
efficient mechanisms to repair damage to DNA in response to these factors
which incidentally also happened to bestow protection against radiation.
Due to the presence of highly expressed genes for several unique proteases
and for detoxification and chaperones, D. radiodurans is able to maintain
integrity of its essential macromolecules like nucleic acids, enzymes and
proteins and repair damaged DNA very efficiently. Large chromosome of D.
radiodurans contains duplicate copies of most of the highly expressed genes
of the small chromosome.
16S rRNA analysis reveals that Deinococcus along with Thermus represents
an ancient group distinct from other bacterial lineages. D. radiodurans lacks
the conventional phospholipids found in other bacteria; instead, its membrane
has phosphoglycolipids containing alkylamines, not found in other bacterial
lineages; a small set of genes found only in Thermus and Deinococcus (and
not in majority of bacteria) have subunits which are found in archaeal genes.
Similarities between members of phyla Thermus-Deinococcus and the
archaeal members in terms of genetics as well as mode of life in extreme
conditions indicate that Thermus-Deinococcus members may be the closest
living relatives of the archaeal species.
By engineering D. radiodurans to express the functions of
detoxification/degradation of metal and organic compounds, scientists are
trying to explore the possible use of this bacterium for cleaning the radioactive
waste sites.
125
Block 2 Evolutionary Biology-II
In prokaryotic genomes, a gene whose codon usage is very similar to that of
the genes encoding ribosomal proteins, translation and transcription
processing factors, principal energy metabolism proteins and the chaperone-
degradation proteins, and deviates from that of the average genes is referred
to as predicted highly expressed (PHX).
Synechocystis genome contains more than 30 PHX genes essential for
photosynthesis. Helicobacter pylori genome contains the PHX genes encoding
urease alpha, urease beta and urease I, which enable the bacterium to
convert urea from gastric juices into bicarbonate and ammonia, neutralising
highly acidic environment that prevails in the stomach. The H. pylori genome is
also rich in PHX genes which code for a family of outer membrane proteins
(nearly 32 members).
Chlamydia trachomatis and C. pneumoniae are mammalian obligate
intracellular parasites and contain PHX gene coding for ATP-ADP translocase
which is also present in Rickettsia but absent from all other bacteria.
Genome of Treponema pallidum, (which causes syphilis in man), has the
highest number of PHX flagellar genes whose products ensure its survival by
facilitating its movement and spreading to all parts of the body including the
brain.
Each region seems to follow its own evolutionary rate (gene-coding regions
may be slow to change as against intergenic and repeated sequences).
Gene order and nucleotide sequences of the genome may be conserved
among different strains as shown for some bacteria or there may be variation
among even related strains due to various factors such as transposon-
mediated deletions/insertions, rearrangements and through horizontal gene
transfer (you will be studying this in detail in unit 5).
Genomes of a large number of bacterial species contain genes with defence
functions, especially the restriction-modification system which provides innate
immunity against viral infections and CRISPR-associated proteins (CRISPR-
Cas) (found in nearly 40% of bacterial species) which confer upon bacteria
adaptive immunity against viral infections (described in more detail in Unit 7).
6.3.2 Archaeal Genomes
Methanocaldococcus. jannaschii DSM2661 was the first archaeal genome to
be sequenced in 1996. Subsequently genomes of other archaeal species from
11 different phyla have also been sequenced revealing important features of
the archaeal genomes. Extensive variation of genome size and composition
has been documented in archaea.
Genomes of archaeal species consist of a single copy of a circular
chromosome with size ranging from 0.5 to 5.8 Mbp. Organisms belonging to
Nanoarchaea have the smallest known genome as exemplified by the genome
of Nanoarchaeum equitans which consists of only 490885 bp., is very compact
and although it codes for complete machinery required for information
processing, there are no genes for lipid, cofactor, amino acid and nucleotide
biosynthesis due to its symbiotic mode of life.
126
Unit 6 Divergence of Bacterial and Archaeal Genomes
Some species are polyploid and may contain many copies of their
chromosome. For example, Methanocaldococcus jannaschii, Methanococcus
maripaludis, Haloferax volcanii and haloarchaeal isolates from salt deposits.
Some of the postulated possible advantages of polyploidy in these species
include repair of double-strand breaks, improved survival under prolonged
stressful environment, genome equalisation by gene conversion, genomic
DNA serving as a phosphate reservoir.
Symbionts which derive their nutrition from a host tend to have smaller
genomes due to loss of genes which are not essential for this mode of life.
In the genomes of the majority of archaeal species there is a single origin of
replication (regions which are particularly rich in AT nucleotides with
conserved sequences termed origin of recognition boxes) which proceeds
bidirectionally as in bacteria, but there are exceptions. For example, the
genome of Sulfolobus solfataricus has three, Aeropyrum pernix has two and
Pyrobaculum calidifontis has four and Halobacterium sp. NRC-1, Haloferax
volcanii and Haloarcula hispanica have multiple origins of replication.
Presence of multiple origins of replication in the genome is advantageous as it
reduces the replication time.
G+C content in archaeal genomes varies from 28 to 66 mol.% but seems to
show no correlation with the temperatures conducive to optimal growth of the
organisms. Instead, DNA stability is achieved by counterions and DNA-binding
proteins and histones (in most archaeal species) which can compactly pack
the genome and overcome the effect of environmental temperature. Some
archaeal species show acetylation-dependent compaction and organisation of
the genome.
Like bacterial genomes, archaeal genomes are also gene-rich, (nearly 2000
genes) with only minimal noncoding regions in between. Short introns have
been found to occur in some protein-coding genes and tRNA genes in
crenarchaeota species but none have been reported from genes in species
belonging to euryarchaeota.
Just like bacterial genomes, insertion sequences, transposable elements and
other mobile genetic elements are present in almost all archaeal genomes. For
instance, Sulfolobus solfataricus shows a lot of metabolic diversity due to the
production of a variety of catabolic enzymes attributed to several hundred
mobile elements detected in its genome.
Genomes of most of the archaeal species encode histone proteins, which
assemble as nucleosome-like structures and are different from the canonical
histone octamers found in eukaryotes.
Genomes of a large number of archaeal species contain a number of genes
with defence functions, especially the restriction-modification system which
provides innate immunity against viral infections and CRISPR-associated
proteins (CRISPR-Cas) (found in nearly 80% of archaeal species) which
confer upon archaea adaptive immunity against viral infections (described in
more detail in Unit 5) and toxin-antitoxin modules which are responsible for
induction of apoptosis following phage infection. Such defence genes tend to
127
Block 2 Evolutionary Biology-II
be organised as genomic islands and horizontal gene transfer seems to play
an important role in the maintenance and evolution of these defence islands.
Sequencing of the genome of M. jannaschii revealed that while genes with a
function in energy production, cell division, and metabolism bore similarity to
their counterparts in bacteria, genes involved in replication, transcription and
translation were more similar to those found in eukaryotes than to those of
bacteria. Its genome houses more than 20 highly expressed genes which are
involved in methanogenesis.
SAQ 2
Write true or false against the following statements:
a) Synechococcus elongatus is a polyploid archaeal species.
b) In general, bacteria occupying complex environmental niches tend to
have larger genomes with a higher GC content as compared to host-
associated bacteria.
c) Genome architecture (gene order and gene content) in bacteria remain
unaffected by mobile genetic elements.
d) Genomes of most of the archaeal species encode histone proteins.
e) Genomes of 80% of archaeal species contain a number of genes
encoding the restriction-modification system and CRISPR-associated
proteins which confer immunity against viral infections.
6.4 DIVERGENCE OF BACERIA AND ARCHAEA
6.4.1 Divergence of Replication System
The various steps in the process of DNA replication such as origin of
synthesis, unwinding and stabilisation of DNA, initiation and priming, relaxation
of supercoiling, repair and ligation rely on a number of products encoded by
several genes. Thus, the replication process is quite complex even in
seemingly simple prokaryotes.
Though all the three domains of life show a conserved process of DNA
replication, modifications of certain proteins involved in the process are
evident in each domain of life as shown in Table 6.4. Archaea present a
mosaic nature of the information processing system, resembling bacteria in
some features while sharing certain other characteristics with eukaryotes.
Table 6.4: Comparison of various components of DNA replication in
bacteria, archaea, and eukaryotes.
Attribute Bacteria Archaea Eukarya
Origin of replication Single Single/more than one Multiple
128
Unit 6 Divergence of Bacterial and Archaeal Genomes
Recognition of DnaA protein ^Orc1/Cdc6 proteins origin recognition
origin of replication complex formed
of 6 different
proteins
Replicative Homohexamer *MCM complex, *MCM complex, a
helicase to unwind DnaB with 5’-3’ dimer of hexamers, heterohexamer,
dsDNA unwinding with 3’-5’ unwinding with 3’-5’
polarity polarity and GINS unwinding polarity
complex (in some) and GINS
complex
ssDNA binding homodimers or Diverse across Heterotrimers of
proteins homotetramers of different species. **RPA
SSB Pyrococcus furiosus
has heterotrimers of
**RPA.
Primase (DNA- Single subunit homologs of Two-subunit
dependent RNA protein #DnaG bacterial DnaG protein, (small
polymerase) primase as well as catalytic PriS and
two-subunit protein large PriL)
(small catalytic *^PriS complexed with
and large PriL) DNA polymerase
alpha
^bear homology with the corresponding proteins in eukaryotes
* minichromosome maintenance complex
** replication protein A
# DnaG functions in RNA degradation
*^PriS is involved in primer synthesis and DNA repair in archaea.
Though DNA binding proteins, HU and H1, rich in positively charged amino
acids, are associated with the DNA in bacterial chromosomes, they do not
function to compact the DNA and the bacterial chromosomes are fully
functional, capable of being easily replicated and transcribed. Table 6.5 lists
the DNA polymerases associated with the process of DNA replication in
bacteria.
Table 6.5: DNA polymerases associated with the process of DNA
replication in bacteria.
DNA Properties Function
Polymerase
I 5’-3’ polymerisation, 3’-5’ Removal of the primer and
exonuclease activity, 5’- filling the gaps by DNA
3’exonuclease activity synthesis and DNA repair
129
Block 2 Evolutionary Biology-II
II 5’-3’ polymerisation, 3’-5’ DNA repair
exonuclease activity, no 5’-3’
(family B)
exonuclease activity
III 5’-3’ polymerisation, 3’-5’ Essential for replication. Also
exonuclease activity, no 5’-3’ functions in proofreading.
(family C)
exonuclease activity.
(Holoenzyme consists of ten
polypeptides).
IV DNA repair
V DNA repair
Taq DNA polymerase obtained from Thermus aquaticus is a thermostable
enzyme which shows optimal polymerisation activity at 75-80o C. There are
studies to show that Taq DNA polymerase bears 50% homology with the E.
coli DNA Pol I and is placed in family A DNA polymerases. Though it shows 5’-
3’ exonuclease activity it lacks proofreading activity due to the absence of 3’-5’
exonuclease activity. When compared with Klenow Pol (large subunit of E. coli
DNA polymerase 1), Klentaq1 (fragment of taq DNA polymerase) has 19
opposite-charge substitutions indicating the role played by redistribution of
charges in imparting thermostability to the enzyme.
Table 6.6 shows the DNA polymerases encountered in archaea.
Polymerase D (Pol D), present in most of the archaeal species, consists of two
subunits, DP1 with exonuclease activity and DP2 with polymerase catalytic
activity.
Catalytic core of archaeal replicative DNA polymerase (Pol D) (responsible for
synthesis of both DNA strands) and the large subunits of the universal RNA
polymerase (RNAP) required for transcription (in all three domains of life and
several large DNA viruses) show homologous sequences. One model
suggests that evolution of RNA polymerases and replicative DNA polymerases
occurred from a common ancestor which had RNA-dependent RNA
polymerase in the RNA-protein world that existed before the start of DNA
replication. According to a number of studies, archaeal Pol D might have
evolved from the replicative DNA polymerase of the Last Universal cellular
Ancestor (LUCA) and subsequently other types of DNA polymerases like Pol
A, Pol B and Pol C might have evolved.
Table 6.6: DNA polymerases encountered in archaea.
DNA Polymerase Property Function
Family B. Core consists of 3 domains- PolB3 involved in the
palm, fingers and thumb, N- synthesis of the leading
PolB3, present in all
terminal 3’-5’exonuclease strand in euryarchaea.
except
domain and a uracil-recognition
Thaumarchaeota. In Crenarchaeota, two
domain
separate enzymes
PolB1, present in
130
Unit 6 Divergence of Bacterial and Archaeal Genomes
Thaumarchaeota, replicate the leading and
Aigarchaeota, lagging strands.
Crenarchaeota and
PolB1 polymerases are
Korarchaeota.
replicative
PolB2 (inactive).
(Encoded by
archaeal
chromosome as well
as several mobile
genetic elements
including
casposons)
Family D (except Consist of 2 subunits-small DP2 has polymerase
Crenarchaeota) subunit DP1 having 2 domains activity, especially for the
(one binds ssDNA and second synthesis of the lagging
has 3'-5'exonuclease activity, strand in euryarchaea. In
and a multidomain large subunit Thermococcus
DP2. kodakarensis, it can
replicate both strands.
The components and the process of replication in archaea show some degree
of similarity with those in eukaryotes.
6.4.2 Divergence of Transcription System
Transcription system makes use of transcription factors and DNA-dependent
RNA polymerase (RNAP).
In archaea, general transcription factors for recognition of core promoter
include TATA-box binding protein (TBP), transcription factor B (TFB) and
transcription factor E (TFE). Presence of homologous regions of helix-turn-
helix units in archaeal TFB and bacterial sigma factors has led researchers to
postulate that bacterial sigma factors might have evolved from TFB with some
modification while TBP and TFE were lost during the course of bacterial
evolution. Studies revealing homology between sigma factors and TFB have
led to a better understanding of events leading to divergence between archaea
and bacteria. Researchers have shown that in the archaeal species Sulfolobus
solfataricus promoters the promoter-proximal element (PPE) present upstream
of the transcription start, that is (-11)-AATATTAA-(-4) bears similarity with
regard to its sequence and placement to the bacterial Pribnow box, that is, (-
12)-TATAAT-(-7) (Fig.6.4) and proposed that the Pribnow box of bacterial
promoters might have evolved from the PPE sequence of archaeal promoters.
The initiator element in both archaea and bacteria can be identified by RNA
polymerase.
131
Block 2 Evolutionary Biology-II
Fig. 6.4: Comparison between bacterial and archaeal promoters. Sso:
Sulfolobus solfataricus; HTH: helix turn helix; PPE: promoter proximal
element; TBP: TATA-box binding protein; TFB: transcription factor B;
Inr: initiator element.
Transcription in bacteria is accomplished by a single form of DNA-dependent
RNA polymerase (RNAP) consisting of five subunits, two alpha, beta, beta’
and omega. The catalytic mechanism and active site for transcription are
provided by beta and beta’ polypeptides while two alpha subunits function in
RNAP assembly and transcriptional regulation. Binding of the core enzyme
with a sigma factor constitutes the holoenzyme. Variations in the polymerase
holoenzyme are created due to the presence of several different forms of
sigma factor. Sigma factor plays the key role of directing the polymerase to
specific promoters.
Archaeal DNA-dependent RNA polymerase core enzyme consists of 10 to 12
subunits, designated as A’/A’’, B’/B’’ (for catalysis), L, N, D and P (for
assembly) and F, E, H and K (for auxiliary functions). Several studies have
shown that there is homology in sequence and structure between the subunits
of bacterial RNAP and the subunits of archaeal RNAP. For instance, beta’
subunit of bacterial RNAP shows homology to subunit A’/A’’ of archaeal
RNAP; subunit beta of bacterial RNAP bears homology with subunit B’/B’’ of
archaeal RNAP; alphaI and alphaII of bacterial RNAP shows similarity with
subunits D and L respectively of archaeal RNAP and omega of bacterial
RNAP shows homology to K subunit of archaeal RNAP.
Histone-based chromatin has been shown to play an important role in
packaging of the DNA and regulation of transcriptional activity for various
cellular processes in eukaryotes. Genomes of several archaeal species
encode either histone proteins (in most of the archaea and can form
nucleosome-like structures, homologous to those found in eukaryotes) or
small basic proteins (analogous to those found in bacteria). However, unlike
eukaryotes, linker histones or chromatin-remodelling complexes are not
encoded by archaeal genomes nor do archaea show post-translational
modification of histones.
6.4.3 Divergence of Translation System
Size and composition of archaeal ribosomes are similar to those of bacterial
ribosomes, i.e., they consist of three RNA molecules, 16S, 23S and 5S RNA
and 50-70 proteins (varies with species) which join to form a 70S ribosome
particle. However, primary structures of ribosomal RNA (rRNA) and ribosomal
132 proteins (RP) bear similarity to those found in eukaryotes rather than to those
Unit 6 Divergence of Bacterial and Archaeal Genomes
occurring in bacteria. In general, extremophilic archaea tend to have a more
rigid structure of ribosomes as compared to ribosomes found in mesophiles.
Halophilic species have acidic ribosomal proteins in order to increase their
hydration capacity.
The archaeal genome generally encodes more RP genes as compared to
bacterial genomes. The increase in RPs in archaea compared with that in
bacteria may reflect a more complex set of interactions in archaea in
regulating translation, e.g., differences in structure requiring scaffolding of
longer rRNA molecules and expanded interactions with the chaperone
machinery. Many genes in the translation category are shared by bacteria and
archaea and these genes have been transmitted vertically, and not
horizontally. Thus, translation proteins are highly conserved between bacteria
and archaea. Bacteria and archaea also share genes for nucleotide transport
and metabolism.
A study investigating ribosomal genes in complete genomes from a number of
species reported that while thirty-four ribosomal (r) - protein families were
represented in all domains of life, thirty-three families of RPs were found only
in archaea and eukaryotes and with regard to protein composition, archaeal
ribosome appeared to be a miniature version of eukaryotic ribosome. Another
important observation of this study was that archaea had plasticity of the
translation apparatus as during the course of archaeal evolution, there was a
progressive reduction of ribosomal genes.
Researchers engaged in comparative studies of genetic code as well as the
various nucleic acids and proteins involved in transcription and translation
processes between bacteria and archaea have postulated that archaea
represent a lineage that appeared on the earth earlier and from which arose
more flexible, successful, and evolutionarily derived organisms constituting the
bacterial domain.
6.4.4 Comparison of Characteristics of Bacteria and
Archaea
Characteristic Bacteria Archaea
Occurrence Cosmopolitan, found in Cosmopolitan, found in
terrestrial as well as aquatic terrestrial as well as
habitats, some members are aquatic habitats, some
extremophiles (living in hot members are
springs, radioactive waste water, extremophiles (living in
organic matter, bodies of plants extreme conditions such
and animals. as hot springs, salt
lakes, marshlands,
oceans, Antarctica);
some live within the gut
of ruminants and
humans.
Cell morphology rod- rod-
shaped/spherical/spirals/coils shaped/spherical/spiral/c
oils
133
Block 2 Evolutionary Biology-II
Cell wall composition Peptidoglycan present (except No peptidoglycan; may
Chlamydias, Planctomycete spp. have
and Mycoplasma spp. pseudopeptidoglycan,
polysaccharides,
glycoproteins or proteins
Cell membrane Fatty acids linked to sn-glycerol- Isoprenoid units linked
3-phosphate by ester linkage to sn-glycerol-1-
and lipid bilayers phosphate by ether
linkage, lipid bilayers,
though some species
have lipid monolayers
Membrane-bound Absent Absent
organelles
Nucleus Absent Absent
Chromosome Circular Circular
Origins of replication Single Multiple in some species
RNA polymerase Single type, consisting of 5 Single type, consisting
polypeptides of 10-12 polypeptides
Initiator tRNA Formyl-methionine Methionine
Response to Susceptible Resistant
antimicrobial agents
(which interfere with
peptidoglycan
biosynthesis). Susceptible Susceptible
Novobiocin (targets
DNA gyrase)
Classical Some species can do Absent
photosynthesis using
chlorophyll
Methanogenesis Absent Occurs in Euryarchaeota
Nitrogen fixation, Present Present
denitrification,
chemolithotrophy, and
hyperthermophilic
growth
Pentose phosphate Present Mostly absent
shunt
Reproduction Asexual, by means of binary Asexual, by means of
fission, budding and binary fission, budding
fragmentation. and fragmentation.
Some species are capable of Do not show sporulation.
sporulation and spores can
remain dormant for several
years.
134
Unit 6 Divergence of Bacterial and Archaeal Genomes
BOX 6.5 : Contributions of prokaryotes in the evolution of eukaryotes.
The contributions made by cyanobacteria in the evolutionary history of various life
forms on the earth is discussed in Box 6.1. Genomes record their own history and
investigations on comparisons of prokaryotic genomes with eukaryotic genomes have
revealed that prokaryotes have contributed to eukaryotic genomes to a very large
extent.
Eukaryotic genes for biochemical and metabolic functions are indicative of bacterial
origin. Bacteria have also contributed building blocks of the endomembrane system to
the eukaryotic lineage.
Archaeal genes involved in information processing, i.e., the RNA polymerase, the
ribosome and aminoacyl tRNA synthetases, and the genes for histones (which enable
formation of nucleosomes and exert epigenetic control of gene expression) are
conserved in eukaryotic lineage. Chromatin organisation based on histones was
instrumental in generating the genomic complexity so characteristic of eukaryotes.
According to one study, archaeal contribution to the eukaryotic genome is about 44%
while bacterial genes are nearly 56%. Land plants have a very high proportion of
genes derived from bacteria (nearly 67%) and bacterial homologs in the yeast genome
comprise nearly 75%.
By their association with ruminants, bacteria have influenced their evolutionary
success.
SAQ 3
Fill in the blanks:
a) Sporulation is seen in …….. but not in………..
b) Both bacteria and archaea are susceptible to novobiocin as it targets
………..
c) Binding of the core enzyme RNAP in bacteria with …………constitutes
the holoenzyme.
d) Archaeal ribosomes consist of RNA molecules ……… , …….…… and
……………
e) In archaea, general transcription factors for recognition of core promoter
include ………., ………. and ……………
6.5 APPLICATIONS
6.5.1 Medicine and Industry
• Several species of killed- or attenuated-bacteria are used in vaccines
while tetanus vaccine is produced from the toxoid of Clostridium tetani.
• Streptomyces spp. (over 500) are the source of a large number of
bioactive compounds such as antibacterial, antifungal, antiparasitic,
immunosuppressants and extracellular enzymes.
135
Block 2 Evolutionary Biology-II
• Secondary metabolites produced by Sorangium cellulosum Soce56 have
been found to be of medicinal value, for example epothilone has anti-
cancer properties, carolacton has been shown to be effective for the
treatment of COVID-19 and myxochelin A produced by Angicoccus
disciformis has been reported to exhibit activity against certain human
cancer cell lines.
• Archaeal membrane lipids are useful in synthesising liposomes for drug
delivery.
• Methanogens are added in sewage treatment plants where they use
organic pollutants as sources of energy forming methane gas as a
byproduct which finds use in cooking.
• Enzymes of Sulfolobus spp exhibit not only great catalytic diversity, (for
instance, starch-hydrolysing, cellulolytic, pectinolytic, chitinolytic,
proteolytic, and lipolytic) but also stability at high temperatures and low
pH. Hence, these enzymes are being used in the food and feed industry,
textile and cleaning industry, pulp, and paper industry.
• Micrococcus luteus has applications in bioremediation due to its ability to
degrade olefinic compounds and hydrocarbons as well as the property to
remediate the regions which contain organic pollutants along with
metals.
• Anammox group of planctomycetes, due to their ability to oxidise
ammonia anaerobically, may find an application in clean-up of nitrogen
in wastewater remediation plants.
6.5.2 Biotechnology and Research
• Taq polymerase from Thermus aquaticus is used in PCR technique due
to its stable nature even at high temperature (70o to 80oC).
• Sulfolobus species are used as a model to study molecular mechanisms
of DNA replication in archaea.
6.6 SUMMARY
● Woese’s designation of three domains of life is supported by molecular
data.
● Bacteria and archaea show similarities in overall morphology and basic
biochemical features, occupation of a variety of ecological niches as well
as use of a common genetic code; however, the two groups show
marked differences in the composition of their cell membranes and cell
walls. Membrane lipids in bacteria consist of fatty acids joined to sn-
glycerol-3-phosphate backbone by ester linkage and in archaea,
isoprenoid units are joined to sn-glycerol-1-phosphate backbone by
ether linkage.
● In most of the species of bacteria and archaea, the genome consists of a
single, circular chromosome and several polycistronic genes organised
136
Unit 6 Divergence of Bacterial and Archaeal Genomes
as operons. However, there are examples of bacterial and archaeal
species which exhibit polyploidy. Possible advantages of polyploidy
include resistance against double-strand breaks, long-term survival and
DNA serving as a phosphate storage polymer.
● Both bacteria and archaea show a remarkable biochemical diversity.
● Prokaryotic evolution is driven by horizontal gene transfer to a great
extent, which may result in expansion or contraction of genome size.
● Enzymes involved in the processes of nucleotide biosynthesis,
transcription and translation are quite conserved among the three
domains.
● Several studies have shown that members belonging to archaea exhibit
a high degree of diversity not only in terms of chromosome copy
numbers but also in systems controlling DNA duplication and cell
division.
● One model suggests that evolution of RNA polymerases and replicative
DNA polymerases occurred from a common ancestor which had RNA-
dependent RNA polymerase in the RNA-protein world that existed before
the start of DNA replication. Archaeal polymerase D might have evolved
from the replicative DNA polymerase of the Last Universal cellular
Ancestor (LUCA).
● Several studies have shown that there is homology in sequence and
structure between the subunits of bacterial RNAP and the subunits of
archaeal RNAP.
● The increase in RPs in archaea compared with that in bacteria may
reflect a more complex set of interactions in archaea and eukaryotes in
regulating translation, e.g., differences in structure requiring scaffolding
of longer rRNA molecules, expanded interactions with the chaperone
machinery.
● Cyanobacteria have been responsible for converting the reducing
atmosphere of the primitive earth into an oxidising one, paving the way
for the evolution of aerobic metabolism in organisms around 2.4-2.1bya.
● Rickettsia belonging to Proteobacteria have contributed mitochondria to
the eukaryotes and cyanobacteria have given rise to chloroplasts of the
photosynthetic eukaryotes.
● Some of the archeal species exhibit metabolic features that allow them
to cope up with the environmental conditions which seem to have
existed on primitive earth in ancient times. Based on this property,
scientists postulate that archaea species arose earlier than bacteria.
6.7 TERMINAL QUESTIONS
1. Fill in the blanks:
a) Members of the bacterial phylum ………. lack peptidoglycan in
their cell walls. 137
Block 2 Evolutionary Biology-II
b) ………….. is the causative agent of food poisoning.
c) …………. are the source of a large number of antibacterial and
antifungal compounds.
d) Engineered bacterium ………….. can be exploited for cleaning up
the radioactive waste sites.
e) Genome of …………. has three origins of replication.
f) ……………has a linear chromosome rather than a circular one.
g) ……………. was the first archaeal genome to be sequenced.
h) Pentose phosphate shunt is present in………. but mostly absent
in………
i) Initiator tRNA carries …………in archaea.
j) Archaea are better adapted to survive in extreme environmental
conditions due to the presence of ………….. in their cell
membranes.
2. Write true (T) or false (F) against the following statements:
a) First photosynthetic organisms belonged to archaea.
b) Extremophiles are encountered in both archaea and bacteria.
c) Aerobic rickettsial ancestor gave rise to mitochondria.
d) Thermus aquaticus, being a thermophile, belongs to archaea.
e) Archaea are not susceptible to beta-lactam antibiotics.
f) Archaeal genomes do not encode linker histones.
g) Many genes of both bacteria and archaea are organised within
operons.
h) The origins of replication in archaea are generally GC-rich regions.
i) Sigma factor in archaea plays the key role of directing the
polymerase to specific promoters.
j) Eukaryotic genes for biochemical and metabolic functions are
indicative of bacterial origin.
3. Match items in column I with the items in column II
Column I Column II
a) Prochlorococcus i) Used as a nitrogen fixing
biofertilizer
b) Borrelia burgdorferi ii) Highly resistant to ionising
radiation
138
Unit 6 Divergence of Bacterial and Archaeal Genomes
c) Azotobacter spp iii) Can withstand temperature of
113oC
d) Deinococcus iv) Produces enormous amounts
radiodurans of oxygen
e) Pyrolobus fumarii v) Can withstand highly acidic
soils (even pH 0)
f) Picrophilus sp. vi) Causes Lyme disease
4. Distinguish between the following pairs of terms:
a) Core genome and pan genome
b) Bacterial cell membrane and archaeal cell membrane
c) Thermophilic proteins and Psychrophilic proteins
5. Enumerate salient features of bacterial genomes briefly.
6. Write short notes on:
a) methanogens;
b) applications of thermophiles
7. Describe briefly the role played by cyanobacteria in the evolutionary
history of life on earth.
6.8 ANSWERS
Self-Assessment Questions
1. a) i) True; ii) False; iii) True;
iv) False; v) True
b) i) Proteins of psychrophilic organisms denature at
temperatures higher than their optima and thus are not able
to function at body temperatures of warm-blooded animals
(37oC) and hence these organisms do not cause disease in
warm-blooded animals.
ii) Listeria monocytogenes is a psychrotroph. Though 30-37oC
is the optimal temperature range for growth, it can grow at a
wide temperature range, from 1 to 45oC and survives
refrigeration, freezing and drying. This is the most important
bacterium responsible for causing food spoilage and
poisoning even under refrigerated conditions.
iii) It has salt concentration 10 to 20 times higher than what is
found in normal cells leaving very little free water within the
cells. This helps in stabilising molecules such as proteins 139
Block 2 Evolutionary Biology-II
and nucleic acids and ensuring survival of these organisms
at such high temperatures.
iv) Presence of multiple origins of replication greatly reduces the
time required for replication.
v) Beta-lactam antibiotics affect bacterial growth by inhibiting
cell wall synthesis. However, as Mycoplasma species lack a
cell wall, they remain unaffected by the treatment using beta-
lactam antibiotics.
2. a) i) False; b) True; c) False;
d) True; e) True
3. a) bacteria; archaea
b) DNA gyrase
c) Sigma factor
d) 16S; 23S; 5S
e) TBP; TFB; TFE
Terminal Questions
1. a) Chlamydia
b) Clostridium perfringens
c) Streptomyces spp.
d) Deinococcus radiodurans
e) Sulfolobus
f) Borrelia burgdorferi
g) Methanocaldococcus jannaschii DSM2661
h) bacteria, archaea
i) methionine
j) tetraether lipids
2. a) False; b) True; c) True; d) False;
e) True; f) True; g) True; h) False;
i) False; j) True.
3. a) Produces enormous amounts of oxygen
b) Causes Lyme disease
c) Used as a nitrogen fixing biofertilizer
d) Highly resistant to ionising radiation
140
Unit 6 Divergence of Bacterial and Archaeal Genomes
e) Can withstand temperature of 113°C
f) Can withstand highly acidic soils (even pH 0)
4. a) Pan genome refers to the entire set of all the genes from all strains
of a clade while, core genome represents a set of homologous
genes present in all the tested genomes.
b) Membrane lipids in bacteria consist of fatty acids joined to sn-
glycerol-3-phosphate backbone by ester linkage while, in archaea,
isoprenoid units are joined to sn-glycerol-1-phosphate backbone
by ether linkage.
c) Proteins in thermophilic organisms show higher numbers of
hydrophobic amino acids, disulphide bonds and ionic interactions
so as to retain protein structure and activity at high temperatures.
On the other hand, proteins in psychrophilic organisms have
reduced hydrophobic core and less charged surface which
ensures flexibility and functionality at low temperatures.
5. Refer to Subsection 6.3.1.
6. a) Methanogens: Methanogens belong to the phylum Euryarchaeota
and are the largest microorganisms belonging to archaea and
show a cosmopolitan distribution, occurring mostly in anoxic
environments such as aquatic sediments, rice paddies, anaerobic
digesters, and the gastrointestinal tract of animals. They can grow
at a wide range of temperatures (4oC to 100oC), salinities
(freshwater to brine) and pH (3 to 9). They possess several unique
coenzymes. Because of the unique cell wall construction, they are
not susceptible to penicillin and other antibiotics. As they are
chemoautotrophic, H2 serves as the source of both energy and
electrons, and CO2 works both as an electron sink and the source
of cellular carbon. Those residing within the gastrointestinal tract of
animals and humans produce methane as a by-product of
metabolism, which causes flatulence in humans and other animals.
Methanosaeta is acetotrophic, using acetate fermentation to
produce methane. Methanogens are added in sewage treatment
plants where they use organic pollutants as sources of energy
forming methane gas as a byproduct which finds use in cooking.
b) Applications of thermophiles: Thermophiles have several
adaptations which enable them to survive at high temperatures.
Enzymes of Sulfolobus spp. exhibit not only great catalytic
diversity, (for instance, starch-hydrolysing, cellulolytic, pectinolytic,
chitinolytic, proteolytic, and lipolytic) but also stability at high
temperatures and low pH. Hence, these enzymes are being used
in the food and feed industry, textile and cleaning industry, pulp,
and paper industry. Taq polymerase from Thermus aquaticus is
used in PCR technique due to its stable nature even at high
temperature (70° to 80°C).
7. Refer BOX 6.1: Role played by cyanobacteria in the evolutionary history
of life on earth.
141
Block 2 Evolutionary Biology-II
UNIT 7
BACTERIAL GENOME
EVOLUTION
Structure
7.1 Introduction Sporulation
Objectives Evolution of Antiviral
Defence
7.2 Agents of Horizontal Gene
Transfer in Bacteria Formation of Biofilms
Plasmids via Conjugation Spread of Nitrogen-Fixing
Property
Exogenous DNA via
Transformation Biotransformation of
Xenobiotics
Bacteriophages via
Transduction 7.4 Consequences of
Bacterial Evolution on
Transposons
Human Health Care
Integrons
7.5 Applications
Genomic Islands
Bioremediation
7.3 Consequences of
Genetic Research
Horizontal Gene Transfer
in Bacteria 7.6 Summary
Development and Spread 7.7 Terminal Questions
of Antibiotic-Resistance
7.8 Answers
and Virulence
Competitive Advantage
Over Other Bacteria
7.1 INTRODUCTION
Bacteria have been in our earth around for four billion years despite extensive
changes in the environment, including introduction of a large number of
antibiotics. Several bacteria adapt to multiple environments and can occupy
142
Unit 7 Bacterial Genome Evolution
diverse ecological niches. For example, Vibrio cholerae residing normally in
aquatic ecosystems, can cause severe diarrhoea in humans and Bacillus
anthracis, found in soil as dormant spores, is the causative agent of anthrax
disease in humans and animals. So, the question which often arises in our
mind is that what makes them so successful? Research in the field of genetics
is providing an insight into their success story. For a long time, the common
belief was that genomes of organisms are fixed and passed only vertically
from parental generation to offspring (Fig.7.1A). A large number of studies
have shown that genomes are not static, but exist in a state of incessant flux.
This property enables them to survive in a variety of habitats and tolerate
chemical onslaughts of all kinds. While the bacterial chromosome codes for all
the basic requirements (mutations and genomic rearrangements providing the
necessary variations for adaptability), bacteria are capable of expanding their
genomes through horizontal gene transfer (HGT) and gene duplication.
Plasmids, viruses, transposable elements, integrons and genomic islands, all
play key roles in HGT (Fig.7.1B). As the origin of new beneficial genes by the
process of spontaneous mutation is very rare, HGT enables a cell to acquire a
functional gene which arose in another cell rather quickly.
Fig. 7.1: Comparisons between vertical transmission (A) and horizontal
transmission, (B) of plasmid in bacteria.
By employing these strategies, bacteria become resistant to most of the
available antibiotics and some become more virulent. Hence, providing
adequate healthcare against bacterial diseases becomes a major challenge.
Therefore, a clear understanding of bacterial evolution at the genetic level
becomes very important. This will enable public health officials to formulate
strategies to deal with infectious diseases, caused by bacteria, more
effectively. In addition, these studies may assist in bioremediation and genetic
research.
143
Block 2 Evolutionary Biology-II
Objectives
After studying this unit, you would be able to:
appreciate that the bacterial genome can show a great deal of plasticity;
comprehend the survival advantage gained by the bacteria by acquiring
new genes with novel functions in varied environmental conditions;
understand the concept of horizontal gene transfer (HGT) in bacteria;
understand about the agents which are capable of expanding the
bacterial genome in a horizontal fashion;
describe the role of HGT in speeding up the rate of evolution;
appreciate the consequences of evolution of antibiotic resistance and
greater virulence in bacteria on human health care system;
understand the possibilities of exploiting bacterial evolution of resistance
to chemicals and heavy metals for the purpose of bioremediation; and
appreciate the application of HGT strategies in biotechnological
research.
7.2 AGENTS OF HORIZONTAL GENE TRANSFER
In 1956 in Japan, six years after the clinical use of antibiotics, the dysentery-
causing pathogen Shigella dysenteriae was found to be resistant to up to four
antibiotics simultaneously. This emergence of multiple resistant strains could
not be attributed to co-appearance of multiple mutations in bacteria. It was
soon established that bacteria were acquiring genes that confer resistance,
rather than mutations arising in resident genes as rate of spontaneous
mutation is very low and change is acquired for a single character. Agents
such as plasmids, exogenous DNA and bacteriophages can mediate transfer
of a number of genes in a single event, achieving a much faster rate of
evolution of bacteria.
7.2.1 Plasmids via Conjugation
Though bacteria don’t show sexual reproduction and recombination,
processes as seen in eukaryotes, they have been shown to engage in sex-like
process as suggested by the work of Joshua Lederberg and Edward Tatum
initially. They were working with two multiple auxotrophs of E. coli strain K12
as can be seen in the Figure 7.2 A. When plated independently on minimal
culture medium, each auxotroph failed to form any colony. However, when
cells of the two auxotrophs were mixed and allowed to grow together,
prototrophs were recovered at a rate of 1/107 which could grow on minimal
culture medium. This showed that some kind of exchange had occurred
between the two kinds of cells, though the physical nature and genetic basis of
this exchange became known only after further experimentation.
In Davis U tube experiment (Fig. 7.2 B), two arms of the U tube were
144 separated by a glass filter with pore size such that bacteria in one arm could
Unit 7 Bacterial Genome Evolution
not come in contact with bacteria present in the other arm, though they shared
a common growth medium suggesting the need for a physical contact
between the two kinds of cells to lead to the production of prototrophs.
(A)
(B)
Fig. 7.2: Diagrammatic illustration showing outlines of the experiments
performed by Lederberg and Tatum (A) and Davis (B).
Hayes reported that if strain A cells were inactivated by exposing them to
streptomycin (inhibitor of protein synthesis) prior to the cross the number of
prototrophs recovered was not affected. However, when strain B was treated
with streptomycin, no prototrophs were recovered. Streptomycin, by inhibiting
growth and division of cells of B strain, stopped the process leading to the
recovery of prototrophs without affecting the ability of strain A cells to transfer
their genetic material. This suggested that genetic material transferred in a
non-reciprocal/unidirectional fashion.
Cells capable of working as donors of their genetic material were designated
as F+ cells (F for fertility). Bacteria receiving the genetic material from the
donor were designated as F-. Subsequent experimentation revealed the
details of the conjugation process as can be seen in fig. 7.3.
145
Block 2 Evolutionary Biology-II
It became established that F+ cells contain a fertility factor (F factor) that
confers the ability to donate part of their chromosome during conjugation. F
factor exhibits following features:
a) It consists of a circular, double stranded DNA molecule with about 40
genes (Fig. 7.3 A) and shows autonomous replication.
b) The transfer (tra) genes present on the F factor are needed to establish
a stable mating pair and initiate DNA transport from the donor to the
recipient cell through a pore formed at the point of contact between a
mating pair.
c) Conjugative system also includes a relaxase that nicks DNA to give a
single-stranded DNA that is suitable for transfer, found in both Gram-
positive (G+) and Gram-negative (G-) systems. However, some species
show transfer of double-stranded DNA, e.g., G+ actinomyces.
Process of conjugation begins with pilus assembly which is a function of the
type IV secretion system (T4SS) in which a coupling protein links a trans
envelope protein complex (a transferosome) to a nucleoprotein complex (a
relaxosome), that is bound at the plasmid’s origin of transfer (ori T).
While G- bacteria establish contact between donor and recipient by
conjugative pili (Fig. 7.3 B; 7.4), G+ bacteria use surface adhesins to form
contact between the mating pair.
Only one of the two strands of plasmid DNA passes through the fused
membranes into the recipient cell. This is followed by DNA synthesis in both
donor and recipient to replace the missing strand in each (Fig. 7.3 B). The
genes encoding the enzymes responsible for this part of the conjugative
process are also found on the plasmid. After completion of the process now
there are two donor cells each with a whole, double stranded, circular,
conjugative plasmid, i.e., F- cells become F+ cells. This process is so efficient
that it can quickly change an entire recipient population to donor cells. Some
types of conjugative plasmids are transferred only between cells of the same
species. Other types can be transferred across species, known as
promiscuous plasmids.
(A)
146
Unit 7 Bacterial Genome Evolution
(B)
Fig. 7.3: Diagrammatic illustration showing the composition of F-plasmid (A)
and the steps involved in the transfer of plasmid during the conjugation
process in bacteria (B).
Fig 7.4: Photomicrograph depicting the formation of a sex-pilus between two
bacteria taking part in the process of conjugation.
Although these experiments showed the transfer of plasmid from donor to
recipient bacterium every time conjugation occurred, they could not explain the
very low rate of, as well as, the mechanism of genetic recombination.
Subsequent investigations by a number of workers clarified the process of
recombination in bacteria.
Cavalli-Sforza treated an F+ strain of E. coli K12 with nitrogen mustard, a
potent mutagen. This enabled him to recover a genetically altered strain of
donor bacteria which showed a much higher rate of recombination (1/104).
Later a similar strain showing a higher rate of recombination was isolated by
William Hayes. Both strains were designated Hfr, for high-frequency
recombination. Important differences were observed between the two types of
mating such as: 147
Block 2 Evolutionary Biology-II
● F+ X F- recipient cell becomes F+ (low rate of recombination), pattern of
gene transfer random
● Hfr X F- recipient cell remains F- (high rate of recombination), pattern of
gene transfer non-random, and varies from strain to strain.
These differences could not be explained until the experiments conducted by
Wollman and Jacob clarified the genesis of Hfr bacteria. They further
postulated that integration of the F factor into the chromosome determined the
O site (Fig.7.5). When Hfr and F- strain underwent conjugation, the initial point
of transfer was determined by the site of integration of the F factor. Genes
next to O enter the recipient first and the last to enter is F factor. As
conjugation generally does not persist for too long a period, the recipient cell,
though acquiring bacterial genes, remains F-. At the site O, the DNA molecule
of the donor opens up, allowing the transfer of one strand.
Fig. 7.5: Diagrammatic illustration
illustration showing formation of an Hfr cell by
integration of F factor into the E. coli chromosome by a single
crossover and how homologous recombination occurs between the
-
chromosome of F recipient and the DNA fragment transferred from the
Hfr, the donor bacterium. After recombination, which involves a double
+ +
crossover, markers (in this case pro and lac ) become a permanent
-
148 feature of the recipient bacterium though it still remains F .
Unit 7 Bacterial Genome Evolution
Subsequent studies further clarified that Hfr’s result when homologous
recombination occurs between an insertion sequence (IS) on the F plasmid
and identical IS element on the chromosome of host the bacterium (Fig.7.6).
Chromosomes of many bacterial species contain several IS elements. For
instance, there are 8 IS1, 6 IS2 and 5 IS3 elements in the chromosome of wild
type E. coli K-12. IS elements being transposable, show variable numbers in
different strains. Because of the presence of several IS elements in the
plasmid, homologous recombination can take place at a number of sites in the
E. coli chromosome as also in different orientations with respect to the
chromosome. However, not all bacterial species have IS elements in their
plasmids. For example, plasmids in Salmonella typhimurium do not contain
many of the IS elements found in E. coli and thus only rarely show the
formation of Hfr which arise by homologous recombination between IS
elements on the plasmid and the bacterial chromosome.
Fig. 7.6: Diagrammatic illustration showing integration of F factor by
homologous recombination involving a single crossover.
All these experiments and observations have helped clarify the process of
recombination in F+ and F- matings. The experiments showed that in F+ × F-
matings, very low rate of recombination is due to rare integration of F factor
into the bacterial chromosome to convert it into Hfr cell which is then capable
of transferring its DNA strand into the recipient bacterium.
Jacob and Adelberg discovered in one of the crosses involving Hfr lac+ strain
and F- lac- strain, lac+ was transferred to lac- at a very high frequency. It was
also found that the transferred lac+ was not integrated into the recipient’s
chromosome because these F+ lac+ exconjugants occasionally gave rise to F-
lac- daughter cells at a frequency of 1X10-3, suggesting their genotype to be F+
lac+/F- lac-. Occasionally, the integrated F factor could excise out of the
chromosome and the cell again becomes F+. This excision process happens
to be faulty sometimes, (Fig. 7.7) such that the excised F factor carries some
bacterial genes lying adjacent to the point of integration. Faulty excision
occurs because there is another homologous region nearby that pairs with the
original. F factors carrying bacterial genes are designated F’. Since F’ has all
the genes to engage in conjugation process with an F- cell, it transfers not only
a strand of F factor but also the bacterial genes that are part of it converting
the recipient bacterium partially diploid for these genes, called merozygote,
149
Block 2 Evolutionary Biology-II
conferring new properties on the recipient bacterium. As F’ plasmids can
confer a state of partial diploidy, they are being increasingly used in the study
of genetic regulation in bacteria.
Fig. 7.7: Diagrammatic illustration showing formation of F plasmid due to faulty
excision of the integrated F factor which now carries a part of bacterial
chromosome.
Although various experiments clarified the conjugation process, it was still not
clear as to what was the mechanism of recombination in the recipient cell.
Answer to this question came with the discovery of several mutants which
greatly reduced recombination in bacteria. Mutant gene, recA nearly abolished
recombination in bacteria; recB, recC and recD mutant genes reduced
recombination by 100 times, all pointing to important roles played by wild-type
products of these genes in recombination. Further investigations revealed that
RecA protein has an important role in recombination involving single stranded
DNA. As the double stranded DNA enters the cell, one strand is degraded,
leaving a single strand which pairs up with its homologous region along the
host chromosome and then RecA facilitates recombination. The recB, recC
and recD are other three genes, whose product makes a complex enzyme,
RecBCD protein; it is capable of unwinding double-stranded DNA, to facilitate
recombination by RecA.
F’ factors are able to alter characteristics of even those bacteria which have
mutant genes coding for RecA and RecBCD proteins.
150
Unit 7 Bacterial Genome Evolution
Table 7.1: Different types of plasmids along with their characteristic
properties.
Types of Property Example
plasmids
i) Conjugative/ Allow synthesis of sex F, P, R, Col (only the large
fertility pilus and can be sized) and virulence plasmids
transferred from one cell
to another
ii) Degradative Impart bacteria with the TOL plasmid of Pseudomonas
plasmids ability to break down putida
unusual organic
compounds like
camphor, salicylic acid,
toluene etc
iii) Resistance Bacteria acquire Plasmid R100 contains genes
(R plasmids) resistance to various which encode resistance against
antibiotics and inhibitors sulfonamides, streptomycin,
of growth tetracycline, chloramphenicol,
fusidic acid. Transfer of R100
can occur among Escherichia,
Klebsiella, Salmonella and
Shigella
iv) Virulence Code for proteins that Colonisation factor antigen of E.
plasmids facilitate adhesion and coli.
colonisation by bacteria
Production of hemolysin and
of specific sites within
enterotoxin by E. coli.
the host or allow
production of toxins Synthesis and secretion of
coagulase, hemolysin,
fibrinolysin and enterotoxin by
Staphylococcus aureus.
v) Col plasmids Encode peptides E. coli
(bacteriocins) which can
kill closely related
bacterial cells
BOX 7.1: Plasmid characteristics.
Plasmids found in bacteria are autonomously replicating genetic elements, whose size
can range from 10 kilobase pairs (kbp) to more than 400kbp. Features shared by all
plasmids, which can multiply only within a host cell, include the presence of origin of
replication where DNA polymerase binds to replicate plasmid DNA and a set of genes
whose products ensure stable maintenance of plasmids in host bacteria.
Although some plasmids occur only as one or two copies per bacterial cell, for
example F plasmid of E. coli, there are examples of plasmids that occur in multiple
copies like Col plasmids of E. coli (more than 50 per cell).
Various plasmids also differ regarding their host range. F plasmids are found only in
151
Block 2 Evolutionary Biology-II
related enteric bacteria E. coli, Shigella and Salmonella. On the other hand, plasmids
of the P-family (P stands for promiscuous) (originally discovered in Pseudomonas)
exhibit an extended host range and can inhabit several hundred species of bacteria
and confer resistance against a number of antibiotics including penicillin.
As transferability of plasmids is dependent on the participation of a number of genes
(approximately over 30 genes), only medium to large sized plasmids have the ability to
move themselves from one bacterial cell to another and are referred to as conjugative
+ -
plasmids or Tra . Some plasmids, though very small and Tra , for example ColE
plasmids, are mobilizable and can be transferred by self-transferable plasmids.
Plasmids exhibit another property known as incompatibility, when two different
plasmids, belonging to the same family and having similar DNA sequences in their
replication genes, fail to co-exist in the same bacterial cell. Plasmids belonging to
different families can share the same bacterial cell, for example plasmid P can co-exist
with plasmid F.
In general, most of the plasmids found in bacteria consist of double stranded,
covalently closed, circular, DNA molecules; however, there are some examples of
bacteria, such as Borrelia (causes Lyme’s disease) and Streptomyces, which contain
plasmids that are linear, double stranded DNA molecules and encode hemolysins
which cause damage to blood cells.
Plasmids, though not essential for the normal metabolic functions of the bacteria which
are taken care of by bacterial chromosome, confer additional attributes to bacteria
which greatly improve their survival under challenging environmental conditions such
as presence of various antibiotics, heavy metals, unusual compounds like
hexachlorophene and quaternary ammonium compounds (Table 7.1).
Bacteria harbouring R plasmids, which encode resistance genes against various
antibiotics, have been shown to undergo conjugation at an exceedingly high rate,
effectively transferring R plasmids to a large number of bacteria in the gut microbiome
which are sensitive to antibiotics. This leads to persistence of a large pool of antibiotic-
resistant bacteria within the gut even after the discontinuation of the antibiotic therapy.
SAQ 1
a) Distinguish between the following:
i) Horizontal transmission and vertical transmission.
ii) F plasmid, Hfr and F’
iii) Virulence plasmid and resistance plasmid
iv) Conjugative plasmid and mobilizable plasmid
v) Prototroph and auxotroph
b) Match the items in column A with the items in column B.
Column A Column B
i) Davis a) F factor carries some bacterial
genes
ii) Hayes b) Hfr strain of donor bacterium
iii) Lederberg and c) Transfer of genetic material
Tatum between conjugating pair is
152 unidirectional
Unit 7 Bacterial Genome Evolution
iv) Cavalli-Sforza d) Physical contact is required for
transfer of genetic material during
conjugation
v) Jacob and Adelberg e) Transfer of genetic material occurs
between two bacteria during
conjugation
7.2.2 Exogenous DNA via Transformation
There is yet another mechanism by which bacteria can show altered
characteristics. Acquisition of genes by cells from free DNA molecules in the
surrounding medium, resulting in a phenotypic change in the recipient, is
called transformation. In 1928, Frederick Griffith, working with Streptococcus
pneumoniae, a bacterium that causes pneumonia, noticed a change in the
property of the avirulent strain when injected into mice along with the heat
killed virulent strain as shown in the Figure 7.8 A.
(A)
153
Block 2 Evolutionary Biology-II
(B)
Fig. 7.8: Outlines of Griffith’s experiment (A) and experiment conducted by
Avery, Macleod and McCarty (B).
Griffith interpreted this as transformation of avirulent bacteria into virulent type
by the virulent strain though the exact mechanism involved was not known at
the time.
Later in vitro studies conducted by Avery, Macleod, and McCarty (Fig. 7.8 B)
conclusively proved that DNA was indeed the transforming principle. In natural
settings, such as soil, free DNA can become available by spontaneous
breakage of donor cells.
(i)
154
Unit 7 Bacterial Genome Evolution
(ii)
Fig. 7.9: Diagrammatic illustration showing uptake of exogenous DNA by Gram
negative competent bacteria i) A: binding of dsDNA at the cell surface;
B: retraction of type 4 pilus pulls DNA through the type II secretin pore;
C: translocation of one strand into the cytoplasm by the Rec2/ComEC
protein; D: recombination between the new strand, ii) the homologous
sequence in the recipient’s chromosome and replacement of locus b in
recipient’s chromosome by locus B present on exogenous DNA in the
transformation process.
o.m. refers to outer membrane and i.m. refers to inner membrane
In a population of bacterial cells, only those in a particular physiological state
called competence, take up exogenous DNA (Fig. 7. 9 A). In Gram-negative
bacteria, DNA bound to the cell surface is pulled through the type II secretin
pore and one strand is translocated into the cytoplasm. Though the detailed
mechanism of DNA uptake and translocation is not completely understood,
studies have shown that once the DNA crosses the barrier of the outer
membrane or cell wall, protein Rec2/ComEC translocates single-stranded
DNA into the cytoplasm. This is supported by the observation that bacteria
bearing mutations in the gene coding for Rec2/ComEC fail to show
transformation.
If the sequence of exogenous DNA bears similarity to any segment in
recipient’s chromosome, it is incorporated into the chromosome by
homologous recombination mediated by RecA protein as discussed earlier in
the section on conjugation; otherwise, it is broken down into nucleotide
subunits which serve as a very good source of deoxyribonucleotides for DNA
replication saving cell’s energy for de novo synthesis of nucleotides.
Homologous recombination between the incoming DNA and the chromosome
of the recipient bacterium may result in a change in the genotype of the cell if
alleles carried by the two DNA segments are different (Fig. 7.9 B). A large
number of species of bacteria have been shown to undergo transformation
naturally including Haemophilus influenzae, Bacillus subtilis, Shigella
paradysenteriae, Streptococcus pneumoniae. Others, like E. coli, can be
induced in the laboratory to become competent. 155
Block 2 Evolutionary Biology-II
Several studies have demonstrated that natural transformation also has the
potential to bring about transfer of transposons and class1 integrons
(described in 7.2.4, 7.2.5), (which frequently carry antibiotic-resistance genes)
between unrelated bacterial species. Thus, following transformation, the
recipient bacterium can acquire new properties such as resistance to
antimicrobials carried by such genetic elements.
7.2.3 Bacteriophages via Transduction
It was an absolutely stunning surprise to us that something as strange as
viruses carrying genes from one cell to another can happen - Joshua
Lederberg
Lederberg and Zinder discovered yet another method by which genes from
one bacterium could be passed to another bacterium. Their experiment
consisted of mixing two multiple auxotrophic strains of Salmonella, LA-22 and
LA-2 and plating them on minimal culture medium, and recovering the
prototrophs at a rate of about 1/105. This observation suggested that this
process was similar to conjugation described earlier for E. coli. In order to
understand the mechanism of gene transfer between bacteria, they placed the
two strains of bacteria in two arms of the Davis tube separated by a filter of
pore size that allowed medium to pass through but not the cells, and thus
allowed the bacteria to grow in common medium (Fig. 7.10). When samples
were removed from both sides of the filter and plated independently on
minimal culture medium, prototrophs were recovered, though only from the
side of the tube containing LA-22 bacteria. As the filter prevented cell contact,
recovery of prototrophs could not have resulted from the process of
conjugation. It was speculated that the genes phe+ and trp+ from LA-2 could
have reached the other side of the tube to convert LA-22 into prototrophs but
the mechanism was not clear and it was called filterable agent (FA).
(A)
156
Unit 7 Bacterial Genome Evolution
(B)
Fig. 7.10: Diagrammatic illustration of the experiment conducted by Lederberg
and Zinder (A) and the comparison of generalised and specialised
types of transduction (B).
Following experiments and observations helped clarify some aspects of this
type of recombination.
• LA-2 cells produced FA only when grown in association with LA-22 cells
but not when grown alone in culture medium and later added to LA-22
cells. This suggested some role played by LA-22 cells in the production
of FA which appeared only when the two strains were grown in a
common medium.
• Digestion with DNase did not destroy the activity of FA, thus indicating
that FA was not exogenous DNA suggesting that recombination
observed was not due to transformation.
• When the pore size of the filter was reduced below the size of
bacteriophages, the FA could not pass across the filter.
Thus, researchers proposed that LA-22 harboured a prophage P22, which
upon entering lytic phase, reproduced, and was released into the medium.
P22 phages crossed the filter, infected, and lysed some of the LA-2 cells. 157
Block 2 Evolutionary Biology-II
Sometimes, P22 phages packaged a region of the LA-2 chromosome in their
heads. If this region contains phe+ and trp+ genes, and if phages pass back
across the filter and infect LA-22 cells, the cells can become prototrophs upon
recombination.
Another set of experiments were conducted to test the possibility of this
hypothesis. Bacteria were grown for several generations in a medium
containing a precursor of DNA that is heavier than normal (15N instead of 14N),
then transferred to light medium (with normal precursors of DNA), containing
radioactive DNA precursor (such as 32P), and coliphage P1. Transducing
particles (containing bacterial DNA) were found to band at a heavier than
normal position and no radioactive phage DNA (DNA synthesised after
infection) was found associated with the band of transducing particles. This
experiment showed that transducing particles produced by certain viruses
carry only bacterial DNA.
When genes are transferred from one bacterium to another by the agency of
bacteriophages, the process is termed transduction. Bacteriophages show
either a lytic or a lysogenic cycle. Figure 7.10 compares the two types of
cycles and their consequences. In lytic cycle, bacteriophage adsorbs to a
specific receptor on the bacterial cell surface injecting its DNA into the
bacterial cytoplasm. This is followed by transcription of phage genes,
replication of phage genome, synthesis of viral proteins and packaging of viral
genomes into capsids to form virions. At the end of the lytic cycle, phage
produces two proteins, holin, to create holes in the cytoplasmic membrane and
endolysin to hydrolyse peptidoglycan layer and thus hundreds to thousands of
virus particles are released into the surrounding medium. Phage, upon
entering the bacterium, causes breakage of the bacterial chromosome and
during the packaging stage of the virus, any fragment of bacterial chromosome
which can fit in the viral head can get packaged. As the infective property
resides with the protein coat, this aberrant phage can infect another bacterium,
transferring the bacterial genes carried by it. Since any random sample of
genes can thus get transferred, the process is termed generalised
transduction. All virulent bacteriophages do not cause transduction such as T-
even phages which degrade the host DNA and reutilize the mononucleotides
thus produced for phage-DNA synthesis.
However, if the bacteriophage DNA gets integrated into bacterial DNA to
become a prophage. And, it does not cause lysis of the cell and is replicated
along with the bacterial chromosome. Transcription of most of the phage
genes, including those needed for lytic cycle are repressed by a specific
repressor.
Once in a while, under any stressful condition, there may be a proteolytic
cleavage and displacement of the phage repressor bound to the early
promoter, the prophage loses its integrated state and enters lytic cycle when
its chromosome loops out of the bacterial chromosome. For example, phage
lambda has att region in its DNA which is homologous to att region of E. coli
chromosome and it is in this region phage DNA integrates into bacterial DNA.
Markers flanking att region in the E. coli chromosome are gal and bio (for
galactose and biotin). If excision is imprecise, then either of the markers can
158
Unit 7 Bacterial Genome Evolution
go along with the phage chromosome, leaving a little portion of the phage
chromosome behind. Transduction achieved by such phages is called
specialised transduction. This is of considerable significance, as virulence
factors encoded by pathogenicity island SaPI1 of Staphylococcus aureus are
transduced between bacterial cells by bacteriophages, greatly speeding up the
evolution of this species.
Phages, which have shown the production of generalised transducing
particles, include E. coli phage P1, Salmonella phage P22, and Bacillus
subtilis phages PBS1 and SP10.
Transduction by bacteriophages can include any bacterial DNA such as linear
chromosome fragments, plasmids, transposons, insertion elements. In
addition to generalised and specialised transduction by bacteriophages,
another method by which phages can contribute to HGT is by release of intact
plasmids following bacterial lysis which then take part in transformation of new
bacterial hosts in the vicinity. An important feature of transduction is that
“donor” and “recipient” cells need not occur at the same place or at the same
time.
In most of the natural environments such as oceans, lakes, soil, even within
the bodies of humans and animals as well as managed environments such as
sewage treatment plants, bacteria as well as bacteriophages occur most
abundantly. This relative high concentration of bacteria and bacteriophages
allows frequent infections and consequent transductions to proceed.
Commensal bacteria have been shown to contain a number of antibiotic
resistance genes which can be acquired by pathogenic members of the gut
microbial community. Other reasons contributing to the higher frequency of the
transduction process could be the polyvalent nature of certain bacteriophages
exhibiting a wide host range (Transduction has been described between
bacteria belonging to different taxa) and the basic structural attribute of
phages which allows them to persist for prolonged periods of time in their
extracellular phase quite resistant to various environmental stresses such as
temperature and radiation. This contrasts with the process of transformation
mediated by naked DNA, which is quite susceptible to environmental stresses.
Some phages, harbouring mobile genetic elements, may be able to transfer
plasmids to bacterial hosts which are unable to receive the same through
conjugation. With this, simultaneous transmission of antibiotic resistance
genes and some virulence determinants, found in mega plasmids, may be
achieved. Viral communities from different biomes, such as natural
environments, wastewater treatment plants, human and animal bodies, have
been studied by metagenomic analysis. These studies have shown that up to
50-60% of the bacteriophage particles contain bacterial genes involved in core
cellular functions as well as prophages, mobile genetic elements, integrases,
transposases and recombinases and thus indicate rather frequent occurrence
of transductions, especially generalised transduction. Viral communities of the
human gut, human lungs and wastewater treatment plants show sequences
corresponding to antibiotic resistant genes in high numbers and DNA
extracted from these bacteriophage particles is able to transform bacteria for
antibiotic resistance.
159
Block 2 Evolutionary Biology-II
SAQ 2
a) Define the following terms:
i) Transducing particle
ii) Competence
iii) Prophage
iv) Lysogenic conversion
v) Phage morons
b) Distinguish between the following pairs of terms:
i) Transformation and transduction
ii) Generalised transduction and specialised transduction
7.2.4 Transposon-Mediated
Transposable elements (TEs) are DNA segments that can move around and
insert themselves into different sites within the genome as well as in plasmids,
using a transposition mechanism which is not dependent on large regions of
homology between TE and the target site. Unlike plasmids and
bacteriophages which can be transferred from cell to cell (intercellular MGEs),
TEs belong to the category of intracellular MGEs. However, upon their
movement and insertion into plasmids or phages, they can also get transferred
to other cells.
Bacterial transposable elements which use DNA intermediates in majority of
cases are of two types: insertion sequences (IS elements) and transposons
(Fig. 7.11). Discovery of reversion of mutants of gal operon in E. coli to wild
type phenotype by spontaneous excision of DNA segment present in the gal
operon, led to the identification of first IS elements. Subsequent studies
showed that bacteria harboured a number of other elements which also
behaved in a similar manner. IS elements, consisting of 800 to 2000 bp, have
inverted repeats, about 10 to 40 bp, at their termini and encode the enzyme
transposase which catalyses the transposition. Bacterial genomes can have
multiple copies of IS elements. For example, there may be 5 to 8 copies of
IS1, 5 copies each of IS2 and IS3 in the E. coli chromosome as well as
multiple copies of different IS elements on F plasmids. R plasmids (Fig. 7.11
B) also contain several IS elements. Analysis of databases shows that IS
elements occur in higher density in bacterial plasmids than in the host
chromosome and plasmids serve as most important agents/vectors in their
transmission.
Transposons show a size range of 2500 to 21,000 bp and possess a gene (not
related to their transposition) coding for antibiotic resistance or some other
function in addition to the transposase-coding gene. It is postulated that
160
Unit 7 Bacterial Genome Evolution
evolution of composite transposon occurred with the insertion of two IS
elements on both the ends of a gene (Fig. 7.11 A).
(A)
(B)
Fig. 7.11: Structure of insertion sequences and transposons (A). Green triangles
represent Flanking direct repeats (FDRs), red or purple triangles
indicate inverted repeats (IRs) and insertion sequences (ISs) are
indicated by yellow boxes with red triangles at the end. An IS50
element is present on the left side (IS50L) as well as on the right side
(IS50R) in an inverted orientation in Tn5. Curly lines represent
transcripts with an arrowhead pointing in the direction of
transcription. ISR codes for transposase for Tn5. Members belonging
to TnA family of transposons possess terminal inverted repeats but
are without terminal IS elements. The tnpA gene codes for
transposase and the tnpR gene encodes a resolvase. ApR codes for
beta-lactamase conferring ampicillin resistance to bacteria.
161
Block 2 Evolutionary Biology-II
Transposable elements (TEs) play a very important role in genome evolution
by introducing changes such as deletions, insertions, or other rearrangements
in bacterial genetic material as well as by transferring genes which may confer
resistance to some antibiotic. Insertion of a TE within a functional gene may
inactivate it or alter its expression level when inserted into the regulatory
region. Genes which are, though, not a part of transposon but present
between two copies of the TE, may also be carried along with them to a new
location either within the chromosome or onto the plasmid.
Sometimes, a transposon carries a drug resistance allele to a plasmid,
creating an R plasmid (Fig.7.11 B). As many R plasmids are conjugative, they
are effectively transmitted to a recipient cell during conjugation. Even R
plasmids that are not conjugative can donate their R alleles to a conjugative
plasmid by transposition. A transposon encoding tetracycline (tetM) resistance
has been reported to be inserted within a 24.5-mDa conjugative plasmid
residing in gonococci. TetM locus encodes a protein which functions to protect
bacterial ribosomes from the effect of tetracycline. Since tetM determinant is
associated with a conjugative plasmid, it is very easily transferred to other
members of gonococcus as well as to other genital organisms such as
Gardnerella vaginalis, Ureaplasma urealyticum.
Resistance to mercuric ions and to more than one antibiotic can be conferred
by large transposons Tn1696 and Tn21. Plasmid R1033 isolated from
Pseudomonas aeruginosa strain was found to possess Tn1696 and plasmid
NR1 (R100) isolated from a strain of Shigella flexneri contained Tn21. Both
Tn1696 and Tn21 contain class1 integrons conferring antibiotic resistance.
Conjugative transposons have the ability to move to another bacterium via
conjugation as well as to transpose. They are present in a number of gram-
positive and gram-negative bacteria and consist of genetic elements
integrated within the genome which can excise themselves and form a
covalently closed circular DNA molecule that can be transferred to a recipient
bacterium during conjugation. Many plasmids residing within the bacteria can
be mobilised by conjugative transposons. Examples of conjugative
transposons include Tn916 from gram-positive Enterococcus faecalis and
pSAM2 from Streptomyces ambofaciens. Though Tn916 does not have a
gene encoding sex pilus, it carries genes which code for all the proteins
required for the formation of mating pore to facilitate its transfer.
7.2.5 Integron-Mediated
Bacteria possess genetic elements called integrons, that code for site-specific
recombinases which allow them to integrate, express and exchange gene
cassettes.
All known integrons are composed of three essential elements for procuring
exogenous genes (Fig.7.12)
● a gene coding for an integrase (intI),
● a primary recombination site (attI), and
● a strong promoter (Pc).
162
Unit 7 Bacterial Genome Evolution
Integron integrases recombine discrete units of circularised DNA known as
gene cassettes, downstream of the resident Pc promoter at the proximal attI
site, permitting expression of their encoded proteins. Figure 7.12 shows how
an integron captures a cassette by site-specific recombination between attI
present in integron and attC (originally called a 59-base element) present in
cassette. In general, cassettes are promoter-less genes which can be
transcribed only by read-through transcription from an adjacent promoter.
Once a cassette has been captured, using the same attI site, another with the
attC site can be integrated immediately adjacent to attI and mRNA produced
from the P ant promoter has the coding sequences for both.
Fig. 7.12: Diagrammatic illustration showing organisation and mechanism of
recombination of an integron and gene cassette (GC). Integration of
the GC occurs at the attI site in the integron (A) and excision (B) of the
GC from the integron are catalysed by the enzyme IntI1. When two attC
sites in a cassette array participate in recombination, a cassette gets
excised as a DNA circle, leading to deletion or rearrangement of a
block of linked cassettes.
Integrons have played a key role in enabling bacteria to accumulate a battery
of exogenous genetic loci to fight off a number of antimicrobials. Some
bacteria have up to eight different resistance cassettes in a single integron.
Integrons represent very ancient elements, present in all environments and
thus can capture novel genes as part of gene cassettes from a very large pool.
There are reports of hundreds of different integron families and more than
17% of bacterial species, whose genomes have been sequenced, carry an
integron based on the presence of IntI gene. Integrons do not exhibit mobility
on their own as the integrase enzyme is not able to excise its own gene from
the chromosome. However, association of integrons with transposons and
plasmids allows easy mobility (by HGT) [Class 1 integrons have been reported
from plasmids R46, R388, R751 and pVS1, and from the transposons Tn21
and Tn1696 and Tn402 is an active transposon] and they can be transferred
between species and across lineages as evidenced by the observation that
163
Block 2 Evolutionary Biology-II
closely related integrons can be detected in distantly related lineages of
bacteria.
Chromosomal location of integrons has been described for a number of
species and can contribute to genomic diversity within serotypes, strains and
species by capture, loss, and rearrangement of gene cassettes, for example,
in Vibrio cholerae, chromosomal integron contains hundreds of gene cassettes
coding for novel proteins and loss or gain of individual gene cassettes occurs
frequently. Several studies suggest that chromosomal integrons are more
prevalent in environmental bacteria while integrons associated with MGEs are
more common in clinical settings subjected to antibiotic selection. It is for this
reason that integrons have, and continue to, play an important role in genome
evolution by generating genotypic and thus phenotypic diversity in bacteria
enabling them to cope with various environmental stresses.
Integron-mediated genomic innovation has been particularly advantageous for
the reason that integration of the new genetic material into the bacterial
genome occurs at a specific recombination site without disrupting the function
of the existing genes. Additionally, promoter (Pc) leads to the expression of
the newly introduced gene making it instantly available to the scrutiny of
natural selection. Use of antibiotics led to selection of particular integrons from
among the vast pool of these elements in the environment with the result that
the majority of Gram-negative bacteria now carry genes conferring resistance
to antimicrobials.
Promoter region of the integrase gene contains a binding site for LexA which
is a regulator of SOS response. In the event of induction of the SOS response,
as seen during transformation with exogenous DNA, bacterial conjugation,
stress and exposure to antimicrobials, there is transcriptional activation of
integron integrase greatly increasing cassette excision and integration. Thus,
the machinery involved in acquisition and rearrangement of gene cassettes is
up regulated at a time when it is going to be most advantageous to acquire
new functions and genomic diversity. However, under stable environmental
conditions, rearranging an already optimum order of cassette array would not
be desirable, hence, activity of the integrase is kept downregulated in the
cells. Thus, integrons contribute immensely in bacterial adaptive response by
rapidly generating diversity in gene content and order or maintaining status
quo depending on the environmental demands.
7.2.6 Genomic Islands
Genomic islands (GIs), acquired by HGT, represent evolutionarily ancient
elements which are not related to plasmids, integrons and integrative
conjugative elements. In 1990, Hacker and his co-workers, while trying to
determine the genetic basis of virulence of E. coli strains which were
pathogenic in the urinary tract, were able to identify gene sets responsible for
coding virulence factors and termed them as pathogenicity islands (PAI).
Researchers also reported that the same group of genes was not present in E.
coli strains which were existing as commensals. Subsequent studies revealed
that different species of bacteria harbour several other classes of GIs and
acquisition of GIs contributes immensely in bacterial evolution by bringing
164
Unit 7 Bacterial Genome Evolution
about diversification and adaptation to varied environmental conditions. In
general, GIs exhibit the following characters (Fig. 7.13):
● size varies between 10 to 200 kb (elements less than 10 kb are called
genomic islets).
● have sequence composition, specific GC percent content and
dinucleotide frequency, different from that of the main bacterial
chromosome.
● To be able to capture gene cassettes, GIs contain a recombination
module consisting of a gene coding for integrase (intl), associated
attachment site (attl) (in some cases a factor governing directionality of
recombination also present),
● Frequently integrate immediately downstream of a tRNA gene.
● Are found upstream of direct repeats (DR). (3’ 16-20 bp of the tRNA
gene is usually present as a direct repeat at the other end of the island).
● Some GIs may carry genes which code for: a) integrases, b) factors
involved in conjugation, and c) genes derived from phages which enable
island transfer between organisms.
● some GIs may contain insertion element (IS) and transposons enabling
GI mobilisation.
● are highly diverse, encoding various functional characteristics, and
hence named accordingly such as:
i) pathogenicity islands (PAI) whose genes code for virulence
factors/toxins, have been shown to include remnants of
bacteriophages, plasmids, and integrative conjugative elements;
ii) resistance islands (RIs) whose genes encode proteins which
confer resistance to antibiotics;
iii) metabolic islands (MIs), whose genes code for proteins involved in
metabolic functions;
iv) symbiosis islands (SIs) whose genes encode proteins needed for
symbiotic mode of life.
● Although characteristics of GIs suggest that they evolved as mobile
genetic elements, most of the currently encountered GIs in bacteria no
longer show mobility and are fixed in the genome. However, they can be
excised and transferred to a new host by any one of the processes of
HGT, such as conjugation, transformation, or transduction. GIs
transferred by conjugation are designated as integrative and conjugative
elements or ICEs.
Some examples of GIs include:
* SXT element in Vibrio cholerae, which provides resistance to antibiotics
sulfamethoxazole, trimethoprim, chloramphenicol, and streptomycin;
165
Block 2 Evolutionary Biology-II
* clc element of Pseudomonas knackmussi strainB13 enables bacteria to
degrade chlorocatechols;
*ICEMISymR7A of Mesorhizobium loti strain R7A allows a saprophytic soil
bacterium to assume a mutualistic relationship (nitrogen fixing capacity) with
leguminous plants belonging to the genus Lotus.
Fig. 7.13: Main features and functions of genomic islands.
SAQ 3
a) Define the following terms:
i) Transposable element
ii) Integron
iii) Genomic island
b) Fill in the blanks:
i) Transposon Tn1696 confers resistance to ………….in
Pseudomonas aeruginosa.
ii) Composite transposons have flanking …………….at each end.
iii) Integrons are composed of a ………, a ………….. and a gene
coding for………….
iv) Genomic islands frequently integrate immediately downstream of a
……………
v) …………….in Vibrio cholerae provides resistance to several
antibiotics.
166
Unit 7 Bacterial Genome Evolution
7.3 CONSEQUENCES OF HORIZONTAL
GENE TRANSFER IN BACTERIA
Horizontal gene transfer (HGT) has been instrumental in greatly speeding up
the process of evolution in bacteria due to transfer of a number of genes not
only among members of one species but also from across different species
aided by plasmids, bacteriophages, transposons and integrons. Genes
transferred by HGT include antibiotic resistance genes, virulence factors,
genes required for nitrogen fixation and biotransformation of xenobiotics and
anti-viral defence as well as genes which allow survival in different habitats
and are described below in brief.
7.3.1 Development of Antibiotic Resistance,
Virulence and Spread
Bacterial-genome sequencing is revealing the important role played by phages
and plasmids in bacterial evolution leading to emergence of new strains and
species. Through plasmids, antibiotic resistance alleles can spread rapidly
throughout a population of bacteria, a very effective strategy for the survival of
bacteria. Research has shown there is frequent horizontal gene transfer
among bacteria in the human body, which contains over 1,000 species of such
microorganisms amounting to over 100 billion in number. Some strains of
bacteria have been discovered which are able to make an enzyme called
NDM-1, or New Delhi metallo-beta lactamase 1. This enzyme enables
bacteria to grow in the presence of nearly all known antibiotics. The resistant
strains have been found in 180 people in India, Pakistan, U.K., Australia,
Canada, Netherlands, U.S., and Sweden.
Therapeutic treatment of farm animals with antibiotics has selected for
resistant bacteria that are found in food when they survive the production
processes, as in the production of raw milk cheeses, e.g., plasmid pK214 in
bacterium Lactococcus lactis (Fig. 7.14) carries genes which confer resistance
against several antibiotics.
Fig. 7.14: Diagrammatic representation showing a plasmid pK214 harbouring
several genes acquired from various sources which confer resistance
to a number of antibiotics. 167
Block 2 Evolutionary Biology-II
Researchers from China have recently reported the presence of bacteria,
which are resistant to colistin, in pigs, meat and a small number of patients.
Integrons have been immensely beneficial to bacteria in overcoming
challenges imposed by antibiotics. More than 70 different resistance cassettes
have been described which can be captured by integrons and they confer
resistance to all beta-lactams, all aminoglycosides, chloramphenicol,
trimethoprim, streptothricin, rifampin, erythromycin, and antiseptics of the
quaternary ammonium compound fam. Association of integrons with
transposons and plasmids allows easy mobility (by HGT).
Transposable elements (TEs) play a very important role in genome evolution
by introducing changes such as deletions, insertions, or other rearrangements
in bacterial genetic material as well as by transferring genes which may confer
resistance to some antibiotic. Insertion of a TE within a functional gene may
inactivate it or alter its expression level when inserted into the regulatory
region. Genes which are, though, not a part of transposon but present
between two copies of the TE, may also be carried along with them to a new
location either within the chromosome or onto the plasmid. Sometimes, a
transposon carries a drug resistance allele to a plasmid, creating an R plasmid
(Fig. 7.11 B). As many R plasmids are conjugative, they are effectively
transmitted to a recipient cell during conjugation.
Presence of IS elements in the F plasmid as well as in the chromosome of E.
coli enables the formation of Hfr strains (Fig. 7.3; 7.5), leading to high
frequency of gene transfer. Insertions of IS elements have also been shown to
produce genomic changes leading to antibiotic resistance and virulence in
bacteria. For example;
● Porin OprD is required in Pseudomonas aeruginosa for diffusion of basic
amino acids and carbapenems into the cell. IS insertion into the gene
oprD causes its inactivation leading to decreased production of porin
OprD, thus conferring carbapenem-resistance to P. aeruginosa.
● Rot is a repressor of toxin production in Staphylococcus aureus. IS256
insertion into the rot promoter caused derepression of cytotoxin
expression and increased virulence.
● IS21 insertion into the mexR repressor in P. aeruginosa causes
derepression of the efflux pump mexAB-oprM whose enhanced
transcription results in increased resistance to beta-lactam antibiotics.
● IS elements having outward oriented -35 promoter components in their
ends upon insertion at the correct distance from a suitable -10 box can
generate a strong hybrid promoter. This has been documented from at
least 17 bacterial species.
R plasmid of a strain of Salmonella enteritidis (Fig. 7.11 B) contains Tn21
which carries resistance genes for mercuric ions, sulfonamide and
streptomycin antibiotics, and Tn10 which encodes tetracycline resistance.
Since R plasmids undergo conjugation at an exceedingly high rate (BOX 7.1),
they also mediate the transfer of Tn21 and Tn10, contributing to dissemination
168
of genes coding for antibiotic resistance. Resistance to mercuric ions and to
Unit 7 Bacterial Genome Evolution
more than one antibiotic is also encoded by large transposons Tn1696 and
Tn21. Plasmid R1033 isolated from a Pseudomonas aeruginosa strain was
found to possess Tn1696 and plasmid NR1 (R100) isolated from a strain of
Shigella flexneri contained Tn21. Antibiotic resistance genes in both the
transposons, Tn1696 and Tn21, are carried by class1 integrons. Transposable
element Tn1546, found in Enterococcus, carries a gene which confers
resistance to vancomycin.
As bacteriophages can transfer any bacterial DNA such as linear chromosome
fragments, plasmids, transposons, insertion elements, CRISPR-Cas system,
by generalised transduction, they have and continue to play a very important
role in not only dissemination of genes encoding antibiotic resistance but also
in genome evolution. Additionally, prophages have been linked to a
phenomenon of lysogenic conversion, i.e., non-virulent bacterial strain
becoming virulent. Prophages have contributed to virulence of a number of
bacterial pathogens (Table 7.2) such as E. coli, Streptococcus pyogenes,
Salmonella enterica and Staphylococcus aureus. While most of the genes
carried by prophages are repressed, virulence genes are organised as
“morons”, representing discrete, autonomous genetic elements, flanked on
one side by sigma70-like promoter and factor-independent transcriptional
terminator on the opposite side, allowing their expression.
Table 7.2: Few bacterial species whose functions/properties have been
modulated by the presence of temperate bacteriophages.
Bacterial species Temperate Function
Bacteriophage
E. coli 0157: H7 Sp5 and Sp15 Shiga toxin
Vibrio cholerae CTX phi Cholera toxin
Corynebacterium diphtheriae phage beta Diphtheria toxin
Clostridium botulinum CE beta Botulinum neurotoxin
type C1
Clostridium difficile pac-type Siphoviridae TcdA and TcdB
exotoxins.
Pseudomonas aeruginosa Inoviridae family Facilitates biofilm
formation
Bacillus subtilis PMB 12 and SP 10 convert sporulation-
negative cells to
sporulating cells
Salmonella enterica serovar SopEphi Facilitate entry of
Typhimurium Salmonella into
intestinal epithelial cells
169
Block 2 Evolutionary Biology-II
Salmonella enterica serovar Gifsy-1 Mediates survival of
Typhimurium Salmonella in Peyer’s
patches
Thus, temperate phages can affect fitness and evolution of their bacterial
hosts, in a number of ways such as:
● Specialised transduction.
● Lysogenic conversion.
● Protection from further infection by virulent phages.
● Modulation of sporulation and biofilm formation.
● Contributing phage genetic material which becomes an integral part of
the bacterial genome (for example SaPIs of S. aureus and PaLoc of C.
difficile) (BOX 7.2).
BOX 7.2: Origin of pathogenicity Islands
Staphylococcus aureus bacteria carry in their genome Staphylococcal pathogenicity
islands (SaPIs) which contain a set of genes that encode superantigens and other
virulence factors including toxic shock syndrome toxin (TSST-1), enterotoxin B (SEB).
SaPIs, while integrated within the genome, express the virulence factors resulting in
pathogenicity. Studies have revealed that SaPIs in the past split from a protophage
lineage, evolved their distinctive features and became a constituent of the
staphylococcal genome. Though capable of undergoing a replication cycle and
forming capsid-enclosed phage-like particles which can be transferred to other
bacteria with high frequency, they are normally repressed from replicating their genes.
If the cells get infected by helper phages which derepress the SaPI, induction of the
replication cycle occurs leading to the formation of capsid-enclosed phage-like
particles which are mobile and are transferred to other bacteria. SaPI-like elements
have been identified in the majority of the Gram-positive cocci.
Similarity between the genes of pathogenicity locus (PaLoc) of Clostridium difficile
(which encode two large exotoxins TcdA and TcdB) and phage phiCD119 as well as
the report that a protein encoded by phiCD119 can bind to the regulatory region of
PaLoc of C. difficile suggest that PaLoc may have evolved from an ancient prophage.
Integrons have, and continue to, play an important role in genome evolution by
generating genotypic and thus phenotypic diversity in bacteria enabling them
to cope with various environmental stresses. Promoter region of the integrase
gene contains a binding site for LexA which is a regulator of SOS response. In
the event of induction of the SOS response, as seen during transformation
with exogenous DNA, bacterial conjugation, stress and exposure to
antimicrobials, there is transcriptional activation of integron integrase which
greatly increasescassette excision and integration. Thus, the machinery
involved in acquisition and rearrangement of gene cassettes is up regulated at
a time when it is going to be most advantageous to acquire new functions and
genomic diversity. However, under stable environmental conditions,
rearranging an already optimum order of cassette array would not be
desirable, hence, activity of the integrase is kept downregulated in the cells.
Thus, integrons contribute immensely in bacterial adaptive response by rapidly
170
Unit 7 Bacterial Genome Evolution
generating diversity in gene content and order or maintaining status quo
depending on the environmental demands.
7.3.2 Competitive Advantage Over Other Bacteria
By producing colicins, the producer bacterial cell can kill other bacteria in the
vicinity gaining competitive edge and improving its fitness.
7.3.3 Sporulation
A number of bacterial species form spores to overcome unfavourable
environmental conditions including antibiotics and disinfectants. Spores are
metabolically dormant and can remain viable for very prolonged periods of
time. Spores are an issue of great concern as they are resistant to most of the
antimicrobials and can cause food poisoning (Bacillus cereus and Clostridium
botulinum), wound infections (C. perfringens and C. tetani), anthrax (B.
anthracis) and intestinal diarrhoea and colitis (C. difficile).
A number of phages have been shown to influence the rate and extent of
sporulation in several bacterial species; for example, sporulation in B. subtilis
is facilitated by phages PMB 12 and SP 10 (Table 7.2).
7.3.4 Evolution of Antiviral Defence
Bacteriophages, which occur in extremely large numbers, attack and lyse
bacteria extensively in all ecosystems (according to one estimate, phages
might be killing 25% of the planet's bacteria every day). Under selection
pressure by phage-induced lysis, bacteria have evolved several defence
strategies to limit invasion by bacteriophages, (employing nearly 10% of
bacterial genome in antiviral defence), which may work by affecting different
stages of the life cycle of viruses such as:
a) Restriction of phage entry
Bacteriophages initiate infection of bacteria by recognising and binding to
specific receptors on the cell surface, like lipopolysaccharides, membrane
proteins, pili etc. followed by DNA injection. Blockade of viral adsorption to
receptors is achieved by such strategies as mutating or masking phage
receptors or preventing access to receptors by synthesis of extracellular
matrix. For example, T-even-like E. coli phages gain entry by binding to outer-
membrane protein A (OmpA). F-plasmid of E. coli encodes outer-membrane
lipoprotein, TraT, which interacts with OmpA, preventing phage attachment.
Some plasmids, which can be acquired by HGT, carry genes for the
production of exopolysaccharides that constitute bacterial capsules, effectively
blocking viral adsorption.
b) Evolution of R-M systems
Investigations in several laboratories have revealed that bacteria restrict
phage propagation by cleaving the phage DNA (foreign) using restriction
endonuclease (REase) while protecting their own genome due to methylation
by cognate methyl transferase (MTase) which together constitute the
restriction-modification (R-M) systems. As bacteria are able to discriminate self
171
Block 2 Evolutionary Biology-II
from non-self, like the property shown by the immune systems of higher
organisms, R-M systems seem to function as primitive immune systems.
According to several studies, there are different types of R-M systems (Table
7.3) and can confer 10- to 108-fold protection to host cell against phage
infection and occur widely in eubacteria and archaea (nearly 90% of the
genomes sequenced show the presence of at least one R-M system and many
bacterial species contain multiple R-M systems per genome) and show a lot of
diversity (nearly 4000 enzymes have been described so far). Specific sites on
the foreign DNA are recognized and cleaved by REase while MTase
recognizes and transfers methyl group to the same specific DNA sequence
within the chromosome of the host bacterium. Cleavage by REs occurs
endonucleolytically at phosphodiester bonds which generate 5’ or 3’
overhangs or blunt ends. MTases transfer the methyl group from S-adenosyl
methionine to the C-5 carbon or the N4 amino group of cytosine or to the N6
amino group of adenines.
Table 7.3: Different types of R-M systems and their characteristics.
Type of R-M system Characteristics
Type I Hetero-oligomeric protein complex having both restriction
and modification activities. Cleave from 100 to tens of
thousands of base pairs away from the target e.g., EcoK1.
Type II Homodimeric or homotetrameric with separate REase and
MTase enzymes. Cleave DNA within or near their target site
e.g., R.EcoR1.
Type III Heterotrimers (M2R1) or heterotetramers (M2R2), contain
restriction-, methylation-, and DNA-dependent NTPase
activities. Recognize short, asymmetric sequences of 5 to 6
bp, translocate along DNA, and cleave 3’ side of the target
site at a distance of about 25 bp only when two recognition
sequences are in inverse orientation with respect to each
other e.g., EcoP1I.
Type IV Cleave only those DNA sequences which contain
methylated, hydroxy methylated or glucosyl-hydroxy
methylated bases at specific sequences e.g., EcoKMcrBC
BOX 7.3: Functions of R-M systems
In order to maintain a functional R-M defence system, bacteria have to incur a cost as
it requires extensive hydrolysis of ATP (except class II R-M system). Additionally,
bacteria show a decreased frequency of restriction sites in their genomes for that
particular recognition sequence so as to protect its genome from attack by its own
restriction enzyme. However, this restriction-site avoidance by mutations may have an
impact on other cellular functions.
Despite the fitness costs involved in maintaining functional R-M systems, bacteria
continue to have R-M systems because of the roles they seem to play in addition to
providing innate immunity. One such function involves stabilisation of the host
chromosome and the plasmid, and helps the bacterial cell to retain the genomic
islands acquired through HGT.
R-M systems may reside on the bacterial chromosome as well as on mobile
elements and thus can invade new genomes through HGT while being carried
by plasmids, phages, transposons and integrons. This is indicated by the
172
Unit 7 Bacterial Genome Evolution
observations that close homologs of R-M systems are found in distantly related
organisms and R-M genes show GC content and codon usage different from the
majority of other genes in the genome.
Restriction endonucleases, by restricting the entry of foreign DNA into the cell,
facilitate genetic isolation and thus serve to maintain species identity and
integrity. Each strain possesses a specific methylation pattern which is distinct from
other closely related species, allowing for distinction between self and non-self.
Species are further divided into different strains, called biotypes, due to the presence
of different recognition specificities of the various enzymes. For example, different
strains of Staphylococcus aureus show variations in the specificities of the type 1
enzymes due to differences in the HsdS subunit and this prevents transfer of mobile
genetic elements among different strains, thus playing a role in the evolution of S.
aureus strains. Type IV enzymes have been reported to act as a barrier to HGT in
clinical strains of methicillin-resistant S. aureus. Strains lacking these enzymes readily
accept DNA from other species. R-M systems have led to the generation of
phylogenetically distinct clades in Neisseria meningitidis as revealed by analysis of
whole genome sequencing. Exchange of genetic material cannot take place among
different strains due to differences in the methylation patterns. In due course of time,
after accumulation of sufficient genetic variation, each biotype may give rise to a
different species.
Foreign DNA, cleaved by restriction endonucleases, may either be further degraded
by exonucleases inside the host, producing small fragments or used for homologous
or non-homologous recombination and integrated into the host chromosome. This may
enable the recipient bacterium to regain the genetic information that may have been
lost accidentally or acquire novel trait/genetic variation and this may lead to improved
fitness of the host bacterium. Some studies have also reported that R-M systems
could lead to the production of genome rearrangements. Thus, by playing a role in
generating genomic diversity, R-M systems can modulate the rate of bacterial
evolution.
As we have discussed in earlier sections that MGEs and bacteriophages play
important roles in bacterial evolution via HGT, the question that comes to our mind is
whether the R-M system is really efficient in checking the entry of foreign DNA into the
cell. The fact of the matter is that evolution of the capacity to inhibit entry of foreign
DNA in one system, (by restriction enzymes), drives the evolution of a counter strategy
in the bacteriophages and other MGEs to escape digestion by restriction enzymes.
Such counter adaptations include (a) Production of methyltransferases by phages to
methylate their genomes; (b) modification of DNA bases such as glycosylation,
hydroxy methylation and acetamidation; (c) co injection of anti-restriction proteins
along with the phages and conjugative plasmids during the process of transfer; (d)
evolution of single-stranded DNA genome of phages (except for a few examples,
restriction endonucleases mostly recognize and cleave double-stranded DNA).
Thus, bacteriophages and other MGEs respond to bacterial evolutionary strategies
which undergo further change in order to remain protected from attack of
bacteriophages and MGEs and this process continues exhibiting coevolutionary arms
race.
c) Evolution of CRISPR-Cas system
Clustered Regularly Interspersed Short Palindromic Repeats (CRISPR)-
CRISPR-associated (Cas) system occurs in the genome of nearly 87% of
species belonging to Archaea and approximately 45% eubacterial species and
confers adaptive immunity against invasion by foreign DNA in prokaryotes.
Although researchers have discovered several types of CRISPR-Cas systems
173
Block 2 Evolutionary Biology-II
which have different sequences and are associated with different Cas
proteins, there are some features which are shared by all CRISPR-Cas
systems as they all depend on DNA-encoded, RNA mediated activity. These
features include:
● A leader sequence of about 500 base pairs is present upstream of
CRISPR locus. It is especially rich in bases adenine and thymine. The
promoter elements and signals carried by the leader sequence are
required for transcription of CRISPR RNA (crRNA) (expression stage)
and integration of foreign genetic material into CRISPR array.
● Each CRISPR locus contains short repeat sequences, 28 to 37 base
pair-long.
● Each repeat exists in a palindromic fashion, i.e., the sequence on one
strand is identical to the sequence on the second strand when read in 5’
to 3’ direction in each strand.
● Spacers, which separate the repeat sequences, contain unique genetic
material which has been derived from either a bacteriophage or a
plasmid in previous encounters,
● Spacers are responsible for conferring specificity to the CRISPR-
mediated defence response by providing immunological memory and
can vary in number from one to several hundred in different species.
● CRISPR-associated genes (cas genes) precede the leader sequence
and code for Cas proteins.
● Cas proteins and crRNA associate into effector complexes which bring
about silencing and cleavage of the invading foreign genetic material
from bacteriophage or plasmid (whose sequence is identical to the
spacer) (interference stage) as well as incorporation into the CRISPR
array referred to as spacer acquisition or adaptation.
Researchers have reported evolutionary relationships between CRISPR-Cas
systems and different classes of mobile genetic elements. This is based on the
following observations:
* A self-synthesising bacterial transposon uses a recombinase which bears
homology to Cas-1 endonuclease of the CRISPR-Cas system. Mechanism of
reaction catalysed by and the target site specificity of, the integrases during
transposon integration are also similar to Cas1-mediated spacer integration
into CRISPR arrays, hence the name casposase for the enzyme and
casposon for the transposon.
* Reverse transcriptase from Group II intron may have been recruited by a
subset of type III CRISPR-Cas systems to provide for spacer acquisition from
RNA.
* Transposon-encoded TnpB nucleases bear homology with the effector
nucleases of Class 2 CRISPR-Cas systems which recognize and cleave the
target DNA.
174
Unit 7 Bacterial Genome Evolution
7.3.5 Formation of Biofilms
According to NIH estimates up to 80% of the bacterial infections affecting
humans are mediated by biofilm (self-secreted matrix) -associated
microorganisms. Biofilm formation occurs in a number of stages as shown in
figure 7.15 A. As these microbes have been found to be resistant not only to
the commonly used antibiotics (they can withstand antibiotics at
concentrations 10 to 1000 times that required to effectively eliminate
planktonic bacteria) but also to phagocytosis by macrophages, biofilm forming
microbes generally lead to persistent infections (Fig. 7.15 B, C), as seen in
native valve endocarditis, osteomyelitis, dental caries, recurrent urinary tract
infections, chronic rhino-sinusitis, lung infection that is associated with cystic
fibrosis. Biofilm formation may also be encountered on devices such as
catheters, stents, orthopaedic implants, contact lenses etc. Besides resistance
to antimicrobials and the ability to evade host immune defences, biofilm mode
of life of some microbes confers other advantages too like the ability to tolerate
free oxygen radicals, changes in pH and nutrient deprivation.
Biofilm formation occurs ultimately secreting a thick matrix of polysaccharide
which prevents accessibility of antimicrobial agents to microbes. Transition of
microbes from free-living, independent existence in the environment to sessile,
community-based mode of life, (bacterial cells surrounded by extracellular
matrix consisting of polysaccharide, extracellular DNA, and proteins) has been
made possible due to a set of a large number of genes and their differential
[Link] phages have been implicated in the formation of
biofilms by some bacterial species, (as listed in Table 7.2) a property greatly
contributing to improved survival of the cells involved. Some of the biofilm-
forming species include Pseudomonas aeruginosa, Staphylococcus aureus, E.
coli, Enterococcus faecalis, Bacillus subtilis and Clostridium difficile.
(A)
175
Block 2 Evolutionary Biology-II
(B)
(C)
Fig. 7.15: Different stages of biofilm formation (A), some examples of possible
infections related to biofilms (B) and biofilm dental plaque on the
surface of teeth (C).
BOX 7.4: Advantages and disadvantages of HGT to bacteria.
In addition to vertical transmission of genetic information, prokaryotes also experience
a great deal of horizontal gene transfer (HGT) from unrelated individuals. When the
genes are transferred via vertical inheritance, there is no risk involved as the set of
genes have already worked optimally for the parents and have evolved by the working
of natural selection. However, in case of HGT, though there is a possibility of gaining a
beneficial gene, the gene transferred can also be useless or even harmful and thus
HGT may not be a safe evolutionary strategy. Outcome of the gene thus acquired may
lead to the evolution of increased or decreased rate of acceptance of genes
transferred horizontally.
There are also some potential disadvantages associated with HGT. For instance,
when a cell gains genetic material, increased genome size increases an organism's
replication time which may not be favoured by natural selection. It is also possible that
the transferred genetic material is non-coding and thus non-functional or it is non-
functional in its new location and thus become a drain on the energy resources of the
cell in terms of transcription and translation of the acquired genetic material.
Horizontally acquired genetic material may disrupt a functioning gene if inserted
randomly within the genome or it may be a transposon capable of replicating several
176 times or it is a lytic virus capable of killing the cell. Thus, cells have evolved ways by
Unit 7 Bacterial Genome Evolution
which to control or prevent the entry and recombination of foreign genetic material into
the cells’ genome.
Recipient prokaryotic cells have evolved various strategies to avoid the ill effects that
may be brought on by HGT, such as development of innate immunity by restriction
enzymes and adaptive immunity by CRISPR-Cas system or silence the expression of
foreign genes which have become inserted in the genome. Several studies have
reported that CRISPR-Cas systems effectively inhibit DNA uptake by phage infection,
plasmid transfer by conjugation and artificial transformation. Maintenance of CRISPR-
Cas systems in nearly 90% of archaea and about 50% of bacteria are suggestive of
the fact that it is more beneficial to prevent uptake of DNA than to obtain it.
However, when faced with an environment containing antibiotics, bacterial acquisition
of genes conferring antibiotic resistance becomes crucial for survival and this is made
possible with the loss of CRISPR-Cas system. This has been clearly demonstrated for
Enterococcus faecalis (E. faecalis), which resides in the gastrointestinal tracts (GIT) of
animals including man. As GIT also contains enterococcal phages, presence of
functional CRISPR-Cas system in the bacteria would appear to be very beneficial.
However, genomic studies revealed the presence of antibiotic-resistance genes
concentrated in hospital-adapted E, faecalis strains which had lost CRISPR-Cas
system. Genomes of these strains were 25% larger than those of commensal strains
and showed the presence of plasmids, phages and pathogenicity islands, responsible
for their rapid evolution in an antibiotic-treated patient. Thus, genetic content even
within a species varies between non-pathogenic strains and clinical isolates. This has
also been confirmed by a study comparing the genome sizes of 3 different [Link]
strains:
E. coli K-12, non-pathogenic, 4,639,221 bp.
E. coli O157:H7 strain EDL933, enterohemorrhagic, 5,528,445 bp.
E. coli strain CFT073, uropathogenic, 5,231,428 bp.
Strains of Streptococcus pneumoniae (S. pneumoniae), which were shown to undergo
transformation by Griffith (see section 5.8) have been shown to be naturally lacking
CRISPR-Cas loci. Researchers were able to demonstrate that when a functional
CRISPR-Cas locus was grafted into S. pneumoniae from a related species S.
pyogenes, both in vitro and in vivo transformation was blocked.
Thus, bacteria benefit by maintaining CRISPR-Cas system when exposed to harmful
DNA, (e.g., lytic phages) but when acquisition of foreign genes (e.g., antibiotic
resistance) becomes crucial for survival, bacteria may show either loss of this defence
system or continue to maintain this in a repressed state using genetic switches, along
with functional HGT as has been shown for E. coli. In the event of HGT becoming
harmful, bacteria switch on the silent loci to obtain the desired immunity from the
CRISPR-Cas array.
Thus, bacteria are very adaptable to varying environmental conditions because their
genome is highly dynamic, always in a state of flux.
7.3.6 Spread of Nitrogen Fixing Property
Bacterium Rhizobium is very important both ecologically and agriculturally.
Very large (>250 kb) conjugative plasmids in Rhizobium contain genes which
allow bacteria to invade the host- plant root cells and carry out conversion of
atmospheric nitrogen to ammonia, thus fulfilling the nitrogen needs of the
plant. Mobile genetic elements have also been shown to carry genes important
in global carbon cycling. Phages of photosynthetic, marine blue-green bacteria
carry genes involved in photosynthesis as also the genes which enable host
bacteria to survive in nutrient poor conditions of the ocean.
177
Block 2 Evolutionary Biology-II
7.3.7 Biotransformation of Xenobiotics
Bacterial species have been described containing naturally occurring plasmid-
encoded genes which allow biotransformation of hydrocarbons, (thus
indicating their potential role in bioremediation applications). Such plasmid-
encoded genes have been found to be organised either in large operons or on
genomic islands, some of which are conjugative.
Soil bacteria show abundance of transposons carrying genes which can
degrade pollutants such as radioactive and metallic substances, and thus play
an important role in soil transformation. There are reports that Bacillus subtilis
has acquired genes, through the agency of bacteriophages, to resist heavy
metal in its environment.
SAQ 4
a) What are bacterial biofilms? What advantages do they confer on
bacteria?
b) Fill in the blanks:
i) Insertion of IS256 into the rot promoter caused ……………. of
cytotoxin expression and ……………virulence.
ii) ………………is responsible for the production of botulinum toxin
type C1 from Clostridium botulinum.
iii) Spores of ………………are resistant to most of the antibiotics and
can cause ……………
iv) ………………. facilitate genetic isolation to maintain species
identity and integrity.
v) Conjugative plasmids in ………….………allow bacteria to convert
……….…….to ammonia.
7.4 CONSEQUENCES OF BACTERIAL
EVOLUTION ON HUMAN HEALTH CARE
As would be clear from the above discussion, bacteria have acquired through
HGT, a number of genes which either confer resistance to a number of
antibiotics simultaneously or enable bacteria to produce virulence factors
increasing their pathogenicity. One study engaged in analysis of sewage and
tap water in New Delhi showed the presence of blaNDM-1 (a metallo-beta
lactamase which confers resistance to carbapenem drugs) in 20 diverse
bacterial strains and each one of these was capable of transferring its plasmid-
encoded resistance-genes to other bacteria via conjugation. Vancomycin is
employed as the last resort drug against infections caused by methicillin-
resistant Staphylococcus aureus (S. aureus). However, researchers analysing
clinical samples, have reported that strains of vancomycin-resistant S. aureus
178 have begun to emerge due to transfer of Tn1546 from Enterococcus to S.
Unit 7 Bacterial Genome Evolution
aureus by conjugation. These examples as well as several others have amply
demonstrated that bacteria have very successfully acquired antibiotic
resistance at a global level. Thus, providing adequate health care has become
a serious challenge.
Two approaches may be required to deal with this situation:
a) One is to avoid overuse and misuse of antibiotics in medical applications
and treatment of farm animals, and to employ proper disposal of
antibiotics.
b) Second is to investigate the mechanisms of accelerated spread of
antibiotic resistance and pathogenicity of bacteria so that effective
strategies could be developed to minimise the spread of genes involved
in antibiotic resistance and virulence. Table 7.4 lists a few common
bacterial diseases with their causative agent along with the site of
virulence gene.
Researchers are testing the efficacy of several compounds which can inhibit
the process of conjugation, thereby slowing/preventing the spread of antibiotic
resistance genes. One study has reported the discovery of two drugs, rotterlin
and the red compound which showed inhibition of conjugal transfer of
plasmids pKM101, TP114, pUB307 and R6K by affecting plasmid replication,
without affecting bacterial growth. Another study was engaged in exploring
relaxase-specific inhibitors to prevent bacterial conjugation. Synthetic fatty
acid 2-hexadecynoic acid (2-HDA), when added to the food for mice, was
found to reduce frequency of conjugation 50-fold in the mouse gut flora. In one
approach, it was found that a newly synthesised short-chain peptide could
block the transposase from acquiring activated conformation, effectively
preventing its movement. Further research may yield some effective
compounds which might be able to arrest the spread of antibiotic resistance in
bacteria.
Researchers are also trying to explore strategies such as virotherapy (by lytic
phages) and use of colicins and related compounds as an alternative to
traditional antibiotics for the treatment of bacterial infections as well as food
preservatives.
Table 7.4: Causative agents of a few common bacterial diseases
affecting human population along with the site of virulence
gene and mode of action.
Disease Causative agent Site of virulence Mode of action
gene
Typhoid Salmonella typhi Plasmids, Cell invasion and
pathogenicity intracellular survival
islands
plague Yersinia pestis Plasmids, Siderophores, toxin
pathogenicity delivery system
islands
179
Block 2 Evolutionary Biology-II
cholera Vibrio cholerae Bacteriophage Toxin
anthrax Bacillus anthracis Plasmids toxin
tetanus Clostridium tetani Plasmid Tetanus toxin
Several infections E. coli Plasmids, Adhesion, toxins
pathogenicity island
Several infections S. aureus Bacteriophage, Toxins
pathogenicity island
7.5 APPLICATIONS
7.5.1 Bioremediation
Bacteria capable of degrading heavy metals and hydrocarbons have these
genes either on plasmids or transposons and thus have the potential to be
used for remediation of degraded soils.
7.5.2 Genetic Research
Plasmids and transposons have been used extensively to determine functions
of genes. Engineered plasmids are being used in DNA cloning and the
phenomenon of transformation allows plasmid vectors to be introduced into
and expressed by E. coli cells. Researchers have used plasmid ColE1 to study
the mechanism of recombination and to test Holliday’s model.
7.6 SUMMARY
● Genomes of bacteria though very stable from one generation to the next,
show high plasticity over evolutionary time scale and can accommodate
enormous variability generated by mutations, genome rearrangement or
acquisition of genetic information by horizontal gene transfer.
● Natural selection then selects the fittest variants thereby resulting in
bacterial adaptation and colonisation of the habitat.
● HGT has the potential to bring novel, functional genes into the recipient,
greatly speeding the evolutionary process as the evolution of new
functional genes may be a slow and rare event. Insertion of new genetic
material may not only generate chromosomal rearrangement but also
have an effect on the level of expression of genes at the site of
integration.
● Mechanisms of HGT occurs through conjugation (by plasmids),
transformation (by exogenous DNA) and transduction (by
bacteriophages), where conjugation contributes the most in HGT.
● Transposable elements represent intracellular mobile genetic elements,
capable of excision from one site to integration into another location
180
Unit 7 Bacterial Genome Evolution
within the genome. However, they can be transferred to other cells when
present on plasmids or by transduction.
● Integrons are found in bacterial chromosomes as well as in transposons
and plasmids, which allow their easy mobility. Because of their ability to
capture gene cassettes, integrons have played a very important role in
the evolution of antibiotic resistance in bacteria.
● Several species of bacteria have benefitted immensely by acquisition of
genomic islands which are of different types such as pathogenicity
islands, resistance islands, metabolic islands, symbiosis islands.
● Restriction endonucleases, by restricting the entry of foreign DNA into
the cell, facilitate genetic isolation and thus serve to maintain species
identity and integrity. Each strain possesses a specific methylation
pattern which is distinct from other closely related species.
● Bacteria benefit by maintaining CRISPR-Cas system when exposed to
harmful DNA, (e.g., lytic phages) but when acquisition of foreign genes
(e.g., antibiotic resistance) becomes crucial for survival, bacteria may
show either loss of this defence system or continue to maintain this in a
repressed state using genetic switches, along with functional HGT.
7.7 TERMINAL QUESTIONS
1. a) Write true or false for the following statements.
i) All gram-positive bacteria form sex pilus for conjugation.
ii) R plasmids are responsible for the production of
bacteriocins.
iii) Col plasmids can be conjugative or non-conjugative.
iv) Gene cassettes do not contain their own promoters.
v) Transformation process is DNase-resistant.
vi) Transduction requires physical contact between the cell
donating DNA and the recipient cell.
vii) Integrons disrupt the functioning of a gene by insertion into
the coding region.
viii) Transposable elements integrate into a new site in the
chromosome by non-homologous recombination.
b) Fill in the blanks:
i) Cleavage by restriction enzymes occurs endonucleolytically
at ……………..……which generate 5’ or 3’overhangs or blunt
ends.
ii) ………….…system provides bacteria with an adaptive
immune system which has memory.
181
Block 2 Evolutionary Biology-II
iii) R-M systems can invade new genomes through ……..………
iv) ……………………microbes lead to persistent infections as
seen in dental caries and lung infection associated with
cystic fibrosis.
v) Transfer of …………….…. Staphylococcus aureus is
responsible for emergence of resistance to vancomycin.
2. What are bacterial plasmids? Enumerate different types of plasmids with
their functions.
3. Describe the role played by transposable elements in bacterial evolution
giving suitable examples.
4. Match the items in column I with the correct option in column II.
Column I Column II
a) Shiga toxin i) plasmid of Clostridium tetani
b) Mediates survival in ii) Temperate bacteriophage of Bacillus
Peyer’s patches subtilis
c) Plague iii) Temperate bacteriophage of E. coli
d) Sporulation iv) Temperate bacteriophage of
Pseudomonas aeruginosa
e) Biofilm formation v) Plasmid and pathogenicity islands of
Yersinia pestis
f) Tetanus vi) Temperate bacteriophage of
Salmonella serovar
5. Explain the contribution of pathogenicity islands in the development of
disease giving suitable examples.
6. Show diagrammatically the process of conjugation and recombination
between an Hfr bacterium and an F- bacterium.
7. Show diagrammatically the organisation and mechanism of
recombination of an integron and gene cassette.
7.8 ANSWERS
Self-Assessment Questions
1. a) i) Horizontal transmission refers to non-reproductive transfer of
genetic material from one organism to another organism in
the same generation and can cross phylogenetic boundaries.
182
Unit 7 Bacterial Genome Evolution
Vertical transmission refers to transfer of genetic material
from parents to offspring generation after generation.
ii) F factor is an autonomously replicating plasmid in a bacterial
cell which allows the bacterial cell to function as a donor
during bacterial conjugation.
Hfr refers to a strain of bacteria capable of showing high
frequency of recombination. F factor is integrated within the
bacterial chromosome, which upon mobilisation transfers a
part of the chromosome to a recipient F- cell.
F’ is the F factor present in a bacterial cell that contains a
portion of the bacterial chromosome.
iii) Virulence plasmids, generally found in enteric bacteria, carry
genes which code for virulence factors which may act as
toxins killing host cells or may enable bacteria to adhere to
and invade host cells or may protect bacteria from host’s
immune defences.
Resistance plasmids carry genes whose products confer
resistance to one or more antibiotics.
iv) Conjugative plasmids contain tra genes that encode proteins
required for the formation of mating pore to allow transfer of
plasmids to another bacterium.
Mobilizable plasmids do not contain tra genes and thus
cannot initiate the process of conjugation by mating pore
formation. However, they do carry genes needed for the
formation of relaxosome and can be easily mobilised if a
conjugative plasmid co-exists in the same bacterial cell.
v) Prototroph is a strain of a microorganism which can grow on
a defined minimal culture medium consisting only of an
organic carbon source and various inorganic ions, including
Na+, K+, Mg+,Ca+ and NH4+ present as inorganic salts. It is
able to synthesise all essential organic compounds required
for its growth.
Auxotroph is a strain of a microorganism which cannot grow
on minimal culture medium and requires supplementation of
culture medium with a vitamin or amino acid which this strain
is not able to synthesise.
b) i) Physical contact is required for transfer of genetic material
during conjugation.
ii) Transfer of genetic material between conjugating pair is
unidirectional.
iii) Transfer of genetic material occurs between two bacteria
during conjugation.
183
Block 2 Evolutionary Biology-II
iv) Hfr strain of donor bacterium.
v) F factor carries some bacterial genes
2. a) i) Transducing particle refers to an aberrant phage which
contains a part of host bacterial genome in place of phage
genome encased in phage capsid.
ii) Competence is a physiological condition during which a
bacterium is capable of taking up and internalising
exogenous DNA molecule.
iii) Prophage is a bacteriophage whose genome is integrated
into a bacterial chromosome and replicates along with the
bacterial chromosome.
iv) Lysogenic conversion occurs when non-virulent bacterial
strain becomes virulent due to the presence of prophages.
v) Phage morons: Virulence genes carried by a prophage are
organised as “morons”, representing discrete, autonomous
genetic elements, flanked on one side by sigma70-like
promoter and factor-independent transcriptional terminator
on the opposite side, allowing their expression.
b) i) Transformation refers to heritable change in a bacterial cell
brought about by homologous recombination of the
internalised exogenous naked DNA (susceptible to DNase
digestion) with a homologous region on the recipient’s
chromosome.
Transduction refers to virus-mediated bacterial
recombination brought about by transducing particles which
are not susceptible to DNase digestion.
ii) Generalised transduction refers to virus-mediated transfer of
any non-viral gene from host bacterium to a recipient
bacterial cell and is brought about by transducing particles
which can carry any set of bacterial genes (and devoid of
phage genes) encased in phage capsid.
Specialised transduction refers to virus-mediated transfer of
a specific gene from host bacterium to a recipient bacterial
cell and is brought about by a transducing particle which
carries a bacterial gene along with phage genetic material
encased in phage capsid.
3. a) i) Transposable element is a DNA segment which can move
from one location to another within the genome independent
of sequence homology.
ii) Integrons represent genetic elements which have the ability
to integrate, express and exchange gene cassettes
employing a site-specific recombination system.
184
Unit 7 Bacterial Genome Evolution
iii) Genomic islands are blocks of DNA, mostly inserted into
tRNA genes, and may carry phage-derived and/or plasmid-
derived sequences, including transfer genes or integrases
and IS elements in addition to gene sets with specific
functions.
b) i) mercuric ions
ii) IS elements
iii) recombination site; promoter; integrase
iv) tRNA gene
v) SXT element
4. a) Bacterial biofilms consist of cells surrounded by an extracellular
matrix consisting of polysaccharide, extracellular DNA, and
proteins. Besides resistance to antimicrobials and the ability to
evade host immune defences, biofilm mode of life of some
microbes confers other advantages too like the ability to tolerate
free oxygen radicals, changes in pH and nutrient deprivation.
b) i) derepression; increased
ii) CEbeta phage
iii) Bacillus cereus; food poisoning
iv) Restriction endonucleases
v) Rhizobium; atmospheric nitrogen
Terminal Question
1. a) i) False, ii) False, iii) True, iv) True
v) False, vi) False, vii) False, viii) True
b) i) phosphodiester bonds
ii) CRISPR-Cas
iii) HGT
iv) Biofilm-forming
v) Tn1546
2. Plasmids found in bacteria are autonomously replicating genetic
elements, whose size can range from 10 kilobase pairs (kbp) to more
than 400kbp. Features shared by all plasmids, which can multiply only
within a host cell, include the presence of origin of replication where
DNA polymerase binds to replicate plasmid DNA and a set of genes
whose products ensure stable maintenance of plasmids in host bacteria
(Table 7.1).
185
Block 2 Evolutionary Biology-II
3. Transposable elements (TEs) play a very important role in genome
evolution by introducing changes such as deletions, insertions, or other
rearrangements in bacterial genetic material as well as by transferring
genes which may confer resistance to some antibiotic. Insertion of a TE
within a functional gene may inactivate it or alter its expression level
when inserted into the regulatory region. Genes which are, though, not a
part of transposon but present between two copies of the TE, may also
be carried along with them to a new location either within the
chromosome or onto the plasmid.
Sometimes, a transposon carries a drug resistance allele to a plasmid,
creating an R plasmid (Fig.7.11 B). As many R plasmids are conjugative,
they are effectively transmitted to a recipient cell during conjugation.
Presence of IS elements in the F plasmid as well as in the chromosome
of E. coli enables the formation of Hfr strains (Fig. 7.3; 7.5), leading to
high frequency of gene transfer. Insertions of IS elements have also
been shown to produce genomic changes leading to antibiotic resistance
and virulence in bacteria. For example;
● Porin OprD is required in Pseudomonas aeruginosa for diffusion of
basic amino acids and carbapenems into the cell. IS insertion into
the gene oprD causes its inactivation leading to decreased
production of porin OprD, thus conferring carbapenem-resistance
to P. aeruginosa.
● Rot is a repressor of toxin production in Staphylococcus aureus.
IS256 insertion into the rot promoter caused derepression of
cytotoxin expression and increased virulence.
● IS21 insertion into the mexR repressor in P. aeruginosa causes
derepression of the efflux pump mexAB-oprM whose enhanced
transcription results in increased resistance to beta-lactam
antibiotics.
● IS elements having outward oriented -35 promoter components in
their ends upon insertion at the correct distance from a suitable -10
box can generate a strong hybrid promoter. This has been
documented from at least 17 bacterial species.
R plasmid of a strain of Salmonella enteritidis (Fig. 7.11 B) contains
Tn21 which carries resistance genes for mercuric ions, sulfonamide and
streptomycin antibiotics, and Tn10 which encodes tetracycline
resistance. Since R plasmids undergo conjugation at an exceedingly
high rate (BOX 7.1), they also mediate the transfer of Tn21 and Tn10,
contributing to dissemination of genes coding for antibiotic resistance.
Resistance to mercuric ions and to more than one antibiotic is encoded
by large transposons Tn1696 and Tn21. Plasmid R1033 isolated from a
Pseudomonas aeruginosa strain was found to possess Tn1696 and
plasmid NR1 (R100) isolated from a strain of Shigella flexneri contained
Tn21. Antibiotic resistance genes in both the transposons, Tn1696 and
186
Tn21, are carried by class1 integrons. Transposable element Tn1546,
Unit 7 Bacterial Genome Evolution
found in Enterococcus, carries a gene which confers resistance to
vancomycin.
4. a) Temperate bacteriophage of E. coli.
b) Temperate bacteriophage of Salmonella serovar
c) Plasmid and pathogenicity islands of Yersinia pestis.
d) Temperate bacteriophage of Bacillus subtilis.
e) Temperate bacteriophage of Pseudomonas aeruginosa
f) plasmid of Clostridium tetani.
5. Pathogenicity islands represent a group of mobile genetic elements
which have and continue to play an important role in the development of
pathogenicity/virulence of strains of several species of bacteria. These
islands are absent from the non-pathogenic strains of the same species.
For general structure see section 7.2.6 and figure 7.13
Some examples of pathogenicity islands are:
*PAI I and II in E. coli, responsible for hemolysin production.
*SaPI in Staphylococcus aureus which encodes superantigens, toxic
shock syndrome toxin and enterotoxin.
*PaLoc of Clostridium difficile which encodes exotoxins TcdA and TcdB.
6. Refer to Subsection 7.2.1 (Fig. 7.6).
7. Refer to Subsection 7.2.5 (Fig. 7.12).
Acknowledgment
Fig. 7.4 : [Link]
[Link]
Fig. 7.15 (C) : [Link]
187