DNA BARCODING FOR IDENTIFICATION AND DETECTION OF
SPECIES
Abstract:
Identification is a very important part of the taxonomy. Since a species
represents the basic unit of biological classification, identifying species is
important to understand the systematics and the precise phylogenetic position of
particular species. In recent years, species identification and delimitation have
seen major improvements because of the incorporation of DNA sequence data.
This review provides a comprehensive list of commonly employed nuclear and
chloroplast regions used for the barcoding of plants.
1. Introduction DNA barcoding is a technology for species-level identification
and detection. It relies on DNA sequence variations in selected and small
regions of nuclear and/or cytoplasmic genomes to provide unique molecular
recognition tags to species. Thus, DNA barcodes are short sequences of DNA
from standardized and globally agreed-upon locus/loci of either nuclear or
cytoplasmic genome or both. These can be from coding or non-coding regions.
The concept of DNA barcoding was introduced by Paul Hebert of the
University of Guelph in 2003, based on his pioneering study on 200 closely
allied Lepidopteran species and subsequent investigations on birds, fishes, and
insects .
DNA sequence that was found to be effective in his pioneering and subsequent
studies on insects, fishes, and birds, "Folmer's region" at 5' end of cytochrome C
oxidase 1 (Cox1) having 658 base pairs was proposed as the universal barcode
for all eukaryotes. Short standardized gene regions (DNA barcodes) such as the
5.8S ribosomal RNA gene and flanking internal transcribed spacers 1 & 2 (ITS)
region have been employed for the rapid and accurate identification of many
species. DNA barcoding had shown tremendous progress in global research
programs since its beginning in 2003. Initially, the major drawback in applying
DNA barcoding was its dependence on reference databases which was limited
primarily. Now millions of barcode sequences have been produced, and a good
amount of reference databases is available for researchers. The advantage of
DNA barcoding over the current taxonomic identification methods is that a
species can be identified even if a small amount of its tissue/DNA is available.
Within two years of promulgating the concept, it was realized that because of
the low substitution rate of nucleotides present in plant mitochondria, cox 1 and
other regions in the mitochondrial genome cannot be used for DNA barcoding
of plants, except for some macroalgae. Thus, the search for a barcode for plants
began earnestly in 2005. Therefore, many such regions from the chloroplast as
well as from the nuclear genomes were tested as possible barcodes for plants.
However, before commenting on the tested loci, it would be worthwhile to list
characteristics, which a locus should possess to become a suitable barcode. The
desirable attributes of a barcode locus are that it should be (i) short (-800bp), so
that it can be easily sequenced in one reaction, (ii) the one that evolves fast to
provide sufficient sequence variations at the species level, (iii) having
sufficiently conserved flanking regions so that universal primers could be used
for their amplification, (iv) variable enough to allow species distinction but with
a little intraspecific variation, to provide distinct barcode gap (v) technically
simple to sequence (vi) easy to align for development of effective
bioinformatics tools (vi) recoverable from herbarium specimens and other
degraded DNA samples
2. Nuclear and Chloroplast regions used for barcoding of plants
2.1. Nuclear Ribosomal Internal Transcribed Spacer (nrITS). ITS part Internal
Transcribed Spacer 2 (ITS 2), are the most commonly sequenced loci, with the
size of the former ranging from 400 to > 1000 bp. 2.2. tRNA for histidine and
photosystem II protein D1 protein. It is an intergenic spacer between the genes
coding for tRNA for histidine and photosystem II protein D1 protein (trnH-
psbA). This is the most variable region in angiosperms. Its length varies from
100-800bp 2.3. Rubisco Large subunit (rbcL). The enzyme Rubisco (Ribulose
bisphosphate carboxylase) is the most widespread and commonly known to
participate in catalysis reactions in carbon assimilation. It comprises both small
and large subunits, where a large subunit (rbcL) carries the site for carbon
fixation exhibiting greater than 20% amino acid conserveness among plant
species. The length of rbcL is approximately 650 bp. It is easy to amplify,
sequence, and align in most land plants and is a widely used phylogenetic
marker. Although this locus is not suitable for species-level identification, it
evolves too slowly to broadly use intra-species delimitation. 2.4. Maturase K
(matK). A portion (-950 bp) of Maturase K (mat K), a group II intron having a
total length of 1536 base pairs, present in the trnK gene, was recommended as
the single locus barcode for plants. CBOL recommended it as one of the loci in
the two-locus barcode. 2.5. Several other loci (atpB-rbcL, atpB, ndhF, psbM-
trnB, rpl36-rps8, rpoB, rpoC1, rps16, trnC-ycf6, trnK-rps16, trnL, trnL-F, trnV-
atpE, ycf6-psbM intron). Have also been tested from the chloroplast DNA.
However, none of these were found to be as suitable as the one described earlier
across a large group of plants.
3. Applications of DNA barcoding DNA barcoding has several applications,
such as:
In assigning unidentified individuals to a species . ▪
In hastening the biodiversity incentivisation and analysis even if the organisms
or their cells/tissue are not available using metabarcoding
. ▪ As genetic resource tags [13], it could enhance the discovery of new species
▪ In identifying cryptic and polymorphic species
. ▪ In connecting distinct stages of the life cycle, which otherwise would be
difficult to identify based on their lack of any common feature
. ▪ In establishing botanical identities of herbals by detecting the various
constituents involved in herbal formulations and foodstuffs, and their other
substitutes [
▪ Paleo Barcoding, a method to study the effect of climate change on life forms
[ ▪ In forensic investigations
. ▪ In screening of propagules of invasive species just during its confinement . ▪
In identification of endangered species of both plants and animals to tackle
illegal trade, even if fragments of these are traded
. ▪ In recognizing complex food webs by DNA analysis of animal gut
. ▪ In Bio-surveillance of habitats for Biosecurity
. ▪ Provides additional evidence for taxonomic circumscription
4. Laboratory Protocol
The steps involved for DNA barcoding are as follows:
4.1. Collection, documentation, identification, preservation, and deposition of
the samples. This is the most critical step. Each specimen needs to be numbered,
and all details about the site of collection (GPS coordinates), habit, and habitat
are to be recorded precisely with a number assigned to each. The digital images
of the specimens are taken as soon as possible. The specimens are identified
correctly by the expert if needed. The plant is preserved, so that tissue for DNA
isolation remains available in the distant future and deposited in a reputed
repository. The accession number of each deposited specimen is obtained
. 4.2. Isolation of genomic DNA. Total genomic DNA is isolated using any
methods, such as original C-TAB protocol [34], modified C-TAB methods, or
Genomic DNA extraction kits (e.g., Qiagen, Fermentas, etc.). The quality and
quantity of extracted DNA are estimated either spectrophotometrically by
taking the absorbance at 260/280, or electrophoretically by resolving DNA
fragments using 0.8% agarose gel. 4.3. Selection/designing of primers of the
selected barcode loci. Primers for the selected loci can be made from the genes'
flanking regions on an allied taxon's available chloroplast genome sequence.
Complete chloroplast genome sequences of many plant species are available on
the NCBI database.
.4. PCR amplification, cleaning of PCR products, and sequencing of the
amplicons. The selected loci are amplified using any DNA polymerase that also
possesses proof reading ability. The thermal cycle varies according to the
primer used, the basic steps of which involve denaturation, annealing, and
extension of the target region
. The PCR products are cleaned using the enzymatic method or the gel
extraction method. The enzymatic purification method involves the use of
Exonuclease-I and shrimp Alkaline Phosphatase (Exo/SAP). Exonuclease I
degrade single-stranded DNA molecules (remaining primers), while Shrimp
Alkaline Phosphatase removes phosphate groups from the remaining dNTPs.
After purification, amplicons are sequenced bi-directionally using Sanger's
sequencing method.
4.5. Analysis and calculating divergence values. For assembling and having
accurate base calling (PHRED score> 20, base calling is more than 99%
correct), codon-code aligner or any other software is used. 6. Limitations of the
DNA barcoding DNA barcoding, like other technologies, is not 100% perfect,
though less than 100% species resolution in some instances could be due to the
similar less than perfect nature of taxonomic delimitation methods. Thus, the
conflicts arising out of the DNA barcode data may require reexamination by
taxonomists. Furthermore, the utility of an individual locus as the barcode
depends not only on its species delimitation ability but also on its success in
easy amplification and sequencing reactions
. Many researchers have signified ITS as a potential universal DNA marker due
to its ability to evolve rapidly and discriminate closely related species with
ease . However, it imposes a few inherent limitations, such as incidences of
intra-genomic variability due to the presence of divergent paralogous copies
within the individuals [36]; pseudogenes [37], which could hinder getting good
quality sequences. Based on these limitations, CBOL Plant Working Group [20]
has not included it in plants' core barcode (matK+ rbcL). Major problems for
authenticating the herbal material using DNA barcoding are the absence of
authentic reference sequences in the GenBank associated with vouchered
specimens submitted in herbaria. The isolation of good-quality DNA is
important for the successful application of molecular methods. However,
sometimes it becomes quite challenging to extract high molecular weight DNA
being in a highly degraded state or due to the presence of higher amounts of
polysaccharides, polyphenols, secondary metabolites in the processed medicinal
plant material [38].
Primarily, most of the DNA barcode-based studies were based on Sanger's
sequencing. These days, Next-Generation Sequencing (NGS) is also a preferred
technique particularly for analyzing samples: showing varying levels of DNA
degradation, and/or acquired from multiple species, and/or containing fillers or
contaminants. In addition, it offers numerous advantages over Sanger
sequencing, including multiple parallel sequencing reactions at a time, clonal
templates separation, superior sensitivity, and faster turnaround time [39]. 7.
Conclusions Nowadays, DNA barcoding is the central molecular technique for
species-level identification. DNA sequence variations in selected and small
regions of nuclear and/or cytoplasmic genomes are employed to provide unique
molecular recognition tags to species. A list of DNA barcodes (about 1.3 M
public records) is available in the BOLD system specific to animal, fungal, and
plant species with a data retrieval interface. Such a system would ease the
application of DNA barcoding to identify species apart from routine taxonomic
methods
Geological Time Scale | Palaeobotany
Geological time scale is a record of earth’s history based on the organisms
that lived at different times.
The geological time scale is a system of chronological measurement that
related stratigraphy (the study of rock strata, especially the distribution,
deposition and age of sedimentary rocks) to time, and is used by the
geologists, palentologists and other earth scientists to describe the time and
relationship between the events that have occurred throughout earth’s
history.
The first geological time scale was proposed in 1913 by the British geologist
Arthur Holmes (1890-1965). This was soon after the discovery of the
radioactivity and using it Holmes estimated that the earth was about 4 billion
years old (evidence from radioactive dating indicates that earth is about 4.5
million years old). This was much greater than previously believed.
The geological time scale is divided into five main eras: Coenozoic,
Mesozoic, paleozoic, Proterozoic and Archezoic. Each era is divided into
periods and each period is divided into epochs.
It is as follows:
There is another kind of time division used – the eon. The entire interval of
the existence of visible life is called the Phanerozoic eon. The great
Precambrian expanse of time is divided into the Proterozoic, Archean and
Hadean eons in order of increasing age.
The names of the eras in the Phanerozoic eon (the eon of visible life) are the
Cenozoic (recent life), Mesozoic (middle life) and Paleozoic (ancient life).
The further subdivision of the eras into 12 periods is based on in-identifiable
but less profound changes in life-forms.
Geological time scale
The duration of the earth’s history has been divided into eras that include
the Paleozoic, Mesozoic, and Cenozoic. Recent eras are further divided
into periods, which are split into epochs. The geological time scale with the
duration of the eras and periods with the dominant forms of life is shown in
Table
The Paleozoic era is characterized by abundance of fossils of marine
invertebrates. Towards the later half, other vertebrates (marine and terrestrial)
except birds and mammals appeared. The seven periods of Paleozoic era in
order from oldest to the youngest are Cambrian (Age of invertebrates),
Ordovician (fresh water fishes, Ostracoderms, various types of Molluscs),
Silurian (origin of fishes), Devonian (Age of fishes, many types of fishes such
as lung fishes, lobe finned fishes and ray finned fishes), Mississippian (earliest
amphibians, Echinoderms), Pennsylvanian (earliest reptiles), Permian (mammal
like reptiles).
Mesozoic era (dominance of reptiles) called the Golden age of reptiles, is
divided into three periods namely Triassic (origin of egg laying mammals),
Jurassic (Dinosaurs were dominant on the earth, fossil bird – Archaeopteryx)
and Cretaceous (extinction of toothed birds and dinosaurs, emergence of
modern birds).
Cenozoic era (Age of mammals) is subdivided into two periods namely
Tertiary and Quaternary. Tertiary period is characterized by abundant
mammalian fauna. This period is subdivided into five epochs namely,
Paleocene (placental mammals, Eocene (Monotremes except duck
billed Platypus and Echidna, hoofed mammals and carnivores), Oligocene
(higher placental mammals appeared), Miocene (origin of first man like apes)
and Pliocene (origin of man from man like apes).Quaternary period witnessed
decline of mammals and beginning of human social life.
The age of fossils can be determined using two methods namely, relative dating
and absolute dating. Relative dating is used to determine a fossil by comparing
it to similar rocks and fossils of known age. Absolute dating is used to
determine the precise age of a fossil by using radiometric dating to measure the
decay of isotopes.
1. Glycolysis:
Glycolysis (Gk. glykys = sweet, lysis = splitting), also called glycolytic
pathway or Embden-Meyerhof-Parnas (EMP) pathway, is the sequence of
reactions that metabolises one molecule of glucose to two molecules of
pyruvate with the concomitant net production of two molecules of ATP.
Glycolysis is almost an universal central pathway of glucose catabolism, and
the complete pathway of glycolysis was elucidated by 1940, largely through
the pioneering contributions of G. Embden, O. Meyerhof, J. Parnas, C.
Neuberg, O. Warburg, G. Cori, and C. Cori. However, glycolysis occurs in
all major groups of microorganisms and functions in the presence or absence
of oxygen. It is located in the cytoplasmic matrix of the cells of an organism.
The whole process of glycolysis (i.e., the breakdown of the 6-carbon glucose
molecule into two molecules of the 3-carbon pyruvate) occurs in ten steps
(Fig. 24.1). The first five-steps constitute the preparatory phase while the
rest live-steps represent the payoff phase (oxidation phase).
In preparatory phase there is phosphorylation of glucose and its conversion
to glyceraldehyde 3-phosphate at the expense of two molecules of ATP.
Oxidative conversion of glyceraldehyde 3-phosphate to pyruvate and the
coupled formation of ATP and NADH is the feature of payoff phase.
The step-wise concise account of glycolysis is the following:
1. Glucose (hexose sugar) is activated for subsequent reactions by its
phosphorylation to yield glucose 6-phosphate, with ATP as the phosphoryl
donor. This reaction, which is irreversible under intracellular conditions, is
catalyzed by enzyme hexokinase, which requires Mg 2+ for its activity.
2. Enzyme phosphohexose isomerase (phosphoglucose isomerase) catalyzes
the reversible isomerization of glucose 6-phosphate (an aldose) to fructose
6- phosphate (a ketose). Phosphohexose isomerase requires Mg 2+ and is
specific for glucose 6-phosphate and fructose 6-phosphate.
3. Enzyme phosphofructokinase catalyses the transfer of a phosphoryl group
from ATP to fructose 6-phosphate to yield fructose 1, 6-bisphosphate. This
reaction is essentially irreversible under cellular conditions.
Phosphofructokinase also requires Mg 2+ for its activity.
4. The enzyme fructose 1, 6-bisphosphate aldolase, often called simply
aldolase catalyses the cleavage of fructose 1,6-bisphosphate to yield two
different triose sugar phosphates, glyceraldehyde 3-phosphate (an aldose)
and dihydroxyacetone phosphate (a ketose).
5. Glyceraldehyde 3-phosphate and dihydroxyacetone phosphate are inter-
convertible. Only glyceraldehyde 3-phosphate is directly degraded in the
subsequent steps and, therefore, dihydorxyacetone phosphate is rapidly and
reversibly converted to glyceraldehyde 3-phosphate by the enzyme triose
phosphate isomerase. This reaction completes the preparatory phase of
glycolysis.
6. This step is the first step of payoff phase of glycolysis, Glyceraldehyde 3-
phosphate oxidises to 1, 3- bisphosphoglycerate with the involvement of
enzyme glyceraldehyde 3- phosphate dehydrogenase. During this reaction
NAD+ is reduced yielding NADH (oxidative phosphorylation).
7. 1, 3-bisphosphoglyceratc is converted to 3-phosphoglycerate. In this
reaction the enzyme [Link] transfers the high-energy
phosphoryl group from 1,3-bisphosphoglycerate to ADP yielding ATP and
3-phosphoglycerate. The formation of ATP by phosphoryl group transfer
from a substrate (1,3-bisphosphoglycerate) is called substrate level
phosphorylation.
8. 3-phosphoglycerate is now converted to 2-phosphoglycerate. In this
reaction the enzyme phosphoglycerate mutase catalyses a reversible shift of
the phosphoryl group between C-2 and C-3 of glycerate; Mg 2+ is essential for
this reaction.
9. In this step the enzyme enalase promotes reversible removal of a molecule
of water from 2-phosphoglycerate to yield phosphoenolpyruvate.
10. This is the last step in glycolysis. Phosphoryl group from
phosphoenolpyruvate is transferred to ADP by enzyme pyruvate kinase to
yield ATP and pyruvate via substrate level phosphorylation. The enzyme
pyruvate kinase requires K and cither Mg 2+ or Mn2+ for its activity.
Crassulacean Acid Metabolism (CAM) | Photosynthesis
This type of metabolism, refers to a mechanism of photosynthesis, that is,
different from C 3 and C4 pathways. Crassulacean acid metabolism (CAM) is
found only in succulents and other xerophytes or plants that grow in dry
conditions.
In this type of metabolism, CO 2 is taken up by the leaves on green stems
through stomata which remain open during night. However, during day time,
stomata in such plants remain closed to conserve moisture.
The CO2 taken up by succulent plants in night is fixed in the similar way as
it takes place in C 4 plants to form malic acid, which is being stored in
vacuole.
Hence, malic acid formed during night is used during day time as a source of
CO2 for photosynthesis to proceed through C 3 pathway.
Crassulacean metabolism is a kind of adaptation found in certain succulent
plants such as pineapple to proceed photosynthesis without much loss of
water, which generally occurs in plants with C 3 and C4 pathways.