DNA Barcoding for Species Identification
DNA Barcoding for Species Identification
IDENTIFICATION
Mathew Osagie Lawani
Author Details
Mathew Osagie Lawani just concluded his masters in Cell and Molecular Biology at the University of Benin, Benin City, Edo State, Nigeria. Email:
lawanimathew@[Link]. Phone: +2348115642744
Keywords
Amplification of Barcode, Bioinformatics tools, Databases, DNA Barcoding, Isolation of DNA, Sequencing the Barcode, Taxonomy.
ABSTRACT
DNA barcoding is a taxonomic method used to identify organisms based on short DNA sequences in their
genome. It was first proposed by Paul Herbert in 2003. CO1, ITS, rbcL, and matK are examples of commonly
used short DNA sequences or barcodes used to identify animals, plants, and fungi. DNA barcoding involves four
basic steps which require expertise and discretion. Numerous projects and databases have long emerged since its
inception, as well as novel works and methods. DNA barcoding has a lot of benefits and applications in the areas
of Medicine, Agriculture, Water quality, etc. There are however some limitations to this biotechnological method
of species identification.
INTRODUCTION
In the past scientists had to use visible morphological features and rigorously study different species under a microscope in
an attempt to distinguish them by their physical features. However, this did not work accurately because two vastly
different species can look the same physiologically under a microscope or petri-dish.
It is the analysis of their DNA sequences that now aids scientists in grouping species based on their different molecular
biodiversity. DNA Barcoding is a vital task used in microbiome analysis. You can identify an organism based on its ribosomal,
1
DEFINITION
DNA barcoding is coined from two words, which are: DNA and barcode. DNA (Deoxyribonucleic acid) is now the
predominant genetic material in the living world (A. Travers, 2015). It is a double helix structure that stores genetic
information. It is also a biopolymer made up of nucleotides as its building block. Each nucleotide in turn is made from a
nitrogenous base (Adenine, Guanine, Thymine, and Cytosine) chemically bonded with a pentose sugar and esterified to a
phosphate group. In DNA, the base sequence is of paramount importance. The genetic information of every living organism
is encoded in a specific sequence of bases. Being a double helix structure, the two strands of the DNA are always
complementary to each other. So, the Adenine of one strand will pair with the Thymine of the opposite strand, while
Figure 1: The DNA base pairing rule (Travers and Muskhelishvili 2015)
2
Figure 2: DNA structure showing the different base pairs A with T ; G with C (Vasudevan et al 2016).
Barcodes on the other hand are applied to products for quick identification. They are series of parallel bars or lines of
varying widths used to enter data into a computer system. The bars are typically black on a white background, and their
width and quantity vary according to application. The bars are used to represent the binary digits 0 and 1 sequences which
in turn can represent numbers from 0 to 9 and be processed by a digital computer to identify a particular product.
Figure 3: A barcode or Universal Product Code (UPC) showing the sequence of binary digits (Britannica image, 2019)
DNA barcoding is similar to the barcoding of products, both are employed for an easier means of identification. The table
below gives some major differences between DNA barcoding and Product barcoding.
Used to identify products found in stores and Used to identify living organisms found in nature
supermarkets
Data is interpreted by optical scanners Data is interpreted via databases such as BOLD and
3
GenBank
From the information gathered thus far, DNA Barcoding can therefore be defined as a taxonomic method used to identify
living organisms based on the peculiarity of their base sequence which is interpreted via bioinformatics.
N.B: Short DNA sequences called a genetic marker, candidate gene, or Gene reference instead of the whole genome are
chosen for identification. This standardized short sequence used for species identification is called a DNA barcode.
4
HISTORY AND DEVELOPMENT OF DNA BARCODING
The taxonomic impediment that exists today for many systematists, field ecologists, and evolutionary biologists, i.e.,
determining the correct identification for any plant or animal sample in a rapid, repeatable, and reliable fashion is a reality
we all must accept. This taxonomic problem was a major reason for the development of a new method for the quick
identification of any species based on extracting a DNA sequence from a tiny tissue sample of any organism. Appropriately
called “DNA barcoding,” referring to the UPC labels one finds on commercial products, DNA barcodes consist of a
standardized short sequence of DNA between 400 and 800bp long that, in theory, can be easily isolated and characterized
The use of such short DNA sequences for biological identifications was first proposed by Paul Hebert and colleagues in 2003
( in his book Biological Identification through DNA barcodes) with the ultimate goal of quick and reliable species-level
identifications across all forms of life, including animals, plants, and microorganisms. The concept of a universally
recoverable segment of DNA that can be applied as an identification marker across species was initially applied to animals
which was the cytochrome oxidase 1 gene (CO1) found in the mitochondria. However, a standard DNA barcode locus for
plants was not accepted by the botanical community until 2009, which was 6 years after Hebert published his first paper on
barcoding animals. After several broad screenings of gene regions in the plant genome, three plastids ( rbcL, matK, and
trnH-psbA ) and one nuclear (ITS) gene region have become the standard barcode of choice in most applications for plants
and fungi.
Key Acronyms rbcL= large subunit of ribulose-biphosphate carboxylase gene. matK= Mutarase K gene. ITS= Internal
transcribed spacer. trnH-psbA = Photo system Q(B) protein – tRNA – His gene.
It was not a coincidence that DNA barcoding developed in concert with genomics-based investigations in the first decade of
the twenty-first century. DNA barcoding (a rapid tool for species identification based on DNA sequences) and genomics (a
broad-based comparative approach to entire genome structure and expression) share an emphasis on large-scale genetic
data acquisition that offers new answers to questions previously beyond the reach of traditional disciplines. DNA barcodes,
5
which in principle will eventually be generated and characterized for all species on the planet, are intended to be stored in
an online digital library of sequences for matching and recognizing unidentified biological samples.
Genomics has accelerated the process of recognizing novel genes and gene functions through the comparisons of vast
amounts of sequence data of the entire genomes of a limited number of taxa. In other words, DNA barcoding aims to utilize
the information of ONE OR A FEW gene regions to identify ALL species of life whereas genomics, the inverse of barcoding,
describes in ONE OR A FEW (but eventually many) selected species the function and interactions across ALL genes. All other
types of DNA sequence-based investigations of organisms, including population genetics and phylogenetics, fall between
(2) Matching, or assigning the barcode sequence of the unknown sample against the barcode library for identification.
The first step requires taxonomic expertise in selecting one or preferably several individuals per species to serve as
Tissue samples that yield high-quality DNA extractions in some cases can be obtained from specimens already housed in
museum collections and herbaria. However, in most cases, new tissues will be taken directly from live specimens in the field
before they are prepared, labeled, and stored as voucher specimens in museum collections. These vouchers then serve as
the permanent record that connects the DNA barcode to a particular species of plant, fungus, or animal. Once the reference
barcode library is complete for the organisms under study, whether they comprise a geographic region, a taxonomic group,
or a target assemblage (e.g., medicinal plants, timber trees, etc.), then the DNA barcodes generated from the unidentified
samples are compared to the known barcodes using some type of matching algorithm. Most practical algorithms for
species assignment start by comparing two DNA sequences to produce a distance measure between the sequences.
In DNA barcoding, a sequence alignment algorithm is usually employed to assign an unknown sample to a known species by
finding the closest database sequence to the sample sequence. Basic local alignment search tool (BLAST) is a matching tool
6
that is provided through GenBank to search for correspondence between a query sequence and a sequence library. Two
additional commonly used distance measures are the Kimura-2-Parameter Distance and the Smith-Waterman Algorithm
7
DNA BARCODE
(2) Possess conserved flanking sites for developing universal PCR primers for the widest taxonomic application.
(3) Be of appropriate sequence length to facilitate current capabilities of DNA extraction and sequencing.
A short DNA sequence of 600bp in the mitochondrial gene for cytochrome c oxidase subunit 1 (CO1) generally fits these
criteria and was accepted early on as a practical, standardized species-level barcode for many animals.
The inability of CO1 to work as a barcode in plants and fungi required that botanists find a more appropriate marker.
Several candidate gene regions were immediately suggested as possible barcodes for plants, but until 2009 none were
universally accepted by the plant taxonomic community. This lack of consensus was for the most part because plants have a
low level of variability in mitochondria DNA so (did not meet the first criterion above), CO1 could not be used as a barcode
for plants. Plants also have a slow evolutionary rate of chloroplast DNA. The candidate genes or barcodes for plants,
therefore, vary with rbcL, matk being the ones often used, while ITS is used for fungi.
8
9
TYPES OF BARCODES
Cytochrome c oxidase I (COX1) also known as mitochondrial encoded cytochrome c oxidase I (MT-CO1) is a protein that
in humans is encoded by the MT-CO1 gene. In other eukaryotes, the gene is called COX1, CO1, or COI Cytochrome c
oxidase I. it is the main subunit of the cytochrome c oxidase complex It is a gene that is often used as a DNA barcode to
identify animal species. MT-CO1 gene sequence is suitable for this role because its mutation rate is often fast enough to
distinguish closely related species and also because its sequence is conserved among conspecifics. Contrary to the primary
objection raised by skeptics that MT-CO1 sequence differences are too small to be detected between closely related
species, more than 2% sequence divergence is typically detected between such organisms, suggesting that the
barcode is effective.
Figure 4: The location of the CO1 gene in mitochondrial genome (L. Sarvananda, 2018)
10
2. Maturase K gene (matK)
The chloroplast maturase K gene (matK) is, except for some ferns, situated within an intron of the trnK gene. The gene
is approximately 1535 basepair long in monocots and is the only chloroplast-encoded group II intron maturase. Universal
primers situated in the trnK gene are used to amplify the entire gene region for phylogenetic studies in orders or families
but are sometimes effectively used on the genus or species level, i.e. in the genus Paeonia (Paeoniaceae). Only 600 to 800
base pair regions of the matK gene are utilized for DNA-barcoding purposes. The matK gene evolves fast (three times
faster than rbcL and trnH-psbA) and some studies suggest it can effectively discriminate between species in the
angiosperm.
Figure 5: The matK chloroplast coding region based on the schematic drawing of L. Sarvananda
2018
11
3. The large subunit of the ribulose-biphosphate carboxylase gene (RBCL)
The chloroplast gene rbcL, which codes for the larger unit of Ribulose-1, 5-bisphosphate carboxylase (RuBisCO) is broadly
RuBPco, is an enzyme involved in the first major step of carbon fixation, a process by which atmospheric carbon dioxide is
converted by plants and other photosynthetic organisms to energy-rich molecules such as glucose. In chemical terms, it is
the enzyme responsible for the carboxylation of ribulose-1, 5 bisphosphates (also known as RuBP). It is probably the most
Within the eukaryotic primary rRNA transcript, the mature 18S, 5.8S, and 25S/28S rRNAs are separated by the internal
transcribed spacers 1 (ITS1) and 2 (ITS2) and flanked by the 5′ and 3′ external transcribed spacers (5′-ETS and 3′-ETS).
In bacteria and archaea, ITS is situated between the 16S and 23S rRNA genes in the rDNA. Sequence evaluation of the
ITS region is widely used in taxonomy and molecular phylogeny, it has a lot of flanking sites for primers to bind to and has a
12
Figure 6: Organization of the eukaryotic nuclear ribosomal DNA tandem repeats (Henras et al, 2015).
13
LITERATURE REVIEW
Hebert et al., (2003) described the employment of sequences of DNA for the identification of a species. In their research,
they emphasized the use of mitochondrial gene cytochrome c oxidase I in the global identification system, to understand
the diversity of life and also to study molecular evolution. They designed three COI profiles, for seven phyla of animals, for
eight largest order of insects, and for two hundred closely allied species of lepidopterans to provide an overview of COI
diversity. They demonstrated that differences in COI sequences were sufficient to assign organisms to their taxonomic
Hebert et al., (2004) tested the effectiveness of barcoding using the COI gene in 260 bird species of North America for
identification and discrimination. The large COI sequence variation concluded that the variation within closely related
Spooner, (2009) stated that DNA barcoding could not identify species within a complicated plant group, Solanum sect.
Petota (wild potatoes) even with the use of different genetic markers such as ITS, trnH-psbA, and matK. It was concluded
that DNA barcoding is a retroactive procedure that relies on well–defined species to function.
Costion, C. et al. (2011) demonstrated the potential of using plant DNA barcodes for the rapid estimation of species richness
in taxonomically poor known areas or cryptic populations, thus revealing a powerful new tool for rapid biodiversity
assessment. The study showed that although DNA barcodes fail to discriminate all species of plants, new perspectives and
methods on biodiversity value and quantification may overshadow some of these shortcomings by applying barcode data in
new ways.
Hajibabaei & McKenna (2012) reported the use of shorter DNA sequences called mini-barcodes in identifying older
museum specimens and samples which have been preserved in formalin or similar DNA unfriendly preservatives.
Eberhardt, U. (2012) described methods currently used for DNA barcoding of fungi, including some comments on the
barcoding of aged herbarium material. His work also outlines the amplification and sequencing of nuclear ribosomal genes:
14
Ng'endo et al., (2013) sequenced the mitochondrial DNA Cytochrome oxidase subunit 1, COI gene from 47 ants of the
genus Pheidole. Their work resulted in significant findings where most sequences clustered into well-differentiated groups
Garcıa-Robledo, C. et al. (2013) used DNA barcoding to identify unknown herbivore eggs as leaf beetles (Order: Coleoptera,
Family: Chrysomelidae) in a premontane tropical forest in Costa Rica. The DNA barcode CO1 accurately identified all the
Pecnikar, Z.F., and Buzan, E.V (2014) reported that DNA barcoding of animals, as well as plants and other organisms, will
improve with advances in PCR amplification and DNA sequencing. It was further added that the technology of DNA
sequencing in the last 25 years has greatly improved and, most recently, next-generation sequencing systems have become
available, enabling the production of large amounts of DNA sequences in a very short time and also, in those cases where a
single DNA region is not enough for barcoding, a combination of two or more regions should be applied, such as in the case
of plants.
Joly et al., (2014) considered several reviews on DNA barcoding and presented the potential uses of DNA barcoding in
eco-informatics, community ecology, invasive species, macroevolution, trait evolution, food webs, trophic interactions,
and spatial ecology. They suggested that DNA barcoding would also lead us to understand interactions between species
Ude et al., (2019) used DNA barcoding to facilitate the identification and biodiversity studies of yam species from Southern
Nigeria by making use of the rbcL gene as the barcode or gene marker. Seventy-five yam accessions were collected from
Enugu and Ebonyi States. It was discovered that the rbcL gene could not resolve the yam accessions well and they further
demonstrated that rbcL is not an effective gene marker for DNA barcoding and as such it should not be recommended as a
15
DNA BARCODING PROCEDURE
According to the International Barcode of Life (IBOL), there are four (4) steps involved in DNA barcoding which are:
Step 4:Compare the resulting sequences against reference databases to find the matching species
16
EXPLANATION OF PROCEDURE
DNA ISOLATION
The first step in most molecular biological research, especially the ones involving gene and expression studies, is obtaining
nucleic acid from the tissues of organisms (Shittu, 2012). The extraction of genomic DNA requires careful sample
preparation, followed by tissue lysis and isolation of the nucleic acids which are separated from all other remaining cellular
components. The condition of the biological source material plays a pivotal role in the quality, quantity, and purity of the
extracted DNA. Therefore, appropriate tissue or sample storage after collecting the biological source material in the field is
required.
N.B: For DNA barcoding, standardized DNA extraction protocols have been established for different taxon groups.
Equipment, Reagents, and Safety: Pestles, scissors, and forceps, micro-pipettes (p1000, p100, p10) with tips, waste beaker
for effluent, napkins, tube rack with 1.5mL microfuge tubes, a waste beaker filled with ice, water bath, micron centrifuge,
Reagents include Nuclear lysis solution placed on ice, protein precipitation solution, RNase solution, DNA degradation
For safety purposes, a pair of latex gloves changed frequently is needed to reduce sample contamination. Make sure to
17
SAMPLE COLLECTION AND PREPARATION
Samples can be collected from a museum, herbariums, etc. Fresh samples collected from the field are still the best source
for DNA extraction. Important data such as the date collected, the name of who collected it, and where it was collected
should be recorded.
Only a small amount of the sample (50mg for plants, 10-20mg for fish/insects) is needed for DNA extraction, the rest should
Plate 1: Students collecting samples for DNA extraction (Google image, 2015)
EXTRACTING DNA
Add 100 uL of nuclear lysis solution to a microfuge tube containing the prepared sample and grind for one minute
18
Add 500uL of another nuclear lysis solution to the grinded sample and incubate in the water bath for 15 minutes at
65o C.
Add 3uL of RNase solution, shake the tube, and place in the water bath at 37o for 15 minutes
Add 200uL of protein precipitation solution to the incubated sample and place on ice for 4 minutes
Place the tube in a micro-centrifuge and spin for 4 minutes at high speed
Remove 600uL of supernatant from the debris and place it into another microfuge tube
Carefully pour out the supernatant of isopropanol into the waste beaker
Add 600uL ethanol to wash excess impurities, and spin again for one minute at maximum speed
Carefully pour out the excess ethanol and leave the tube in the open air for 10-20 minutes or dry with a hair dryer
at low speed.
Add 100uL of DNA dehydration solution to the tube and place in the water bath at 65o C For 45-60 minutes
Keep on ice for PCR analysis for immediate use or store in -20 o C refrigerator for future use. Chemical trehalose can
19
AMPLIFICATION OF BARCODE USING PCR
PCR (Polymerase Chain Reaction) is a molecular technique used to amplify or make multiple copies of DNA segments. This
technique is relevant in DNA barcoding because it can be used to amplify barcode regions such as CO1, rbcL, mat K, ITS, etc.
The use of suitable primers is essential for amplification success. Barcoding primers should correspond to rather
conservative sites with low substitution rates to apply them to a broad range of taxa. Such “universal” primers amplifying
an approximately 650-bp-long fragment of the mitochondrial cytochrome oxidase subunit I (COI) gene were first defined by
Folmer et al. In other words, each barcode regions have its specific primers (oligonucleotides 18-20bp long that is
complementary to the 3 prime end of the sense and antisense strand of the barcode region to be amplified.
In the first step, the DNA double helix is denaturized by heating it into single-stranded template DNA, where the primers
can bind them. The thermostable enzyme DNA Polymerase starts to extend the primers by adding single Deoxynucleotide
triphosphates (dNTPs) producing new double-stranded DNA. This process is performed in a Thermocycler and has to be
repeated several times to increase the number of the target fragments exponentially. The quality of the PCR products is
commonly checked by agarose gel electrophoresis. Before sequencing, PCR products have to be purified to eliminate the
1. DNA Polymerase: Recombinant Taq DNA Polymerase (e.g., Qiagen) is commonly used for standard PCR. It is a
thermostable enzyme of the thermophilic bacterium Thermus aquaticus and is, therefore, able to synthesize DNA at high
2. PCR buffer: For optimal DNA Polymerase reaction activity, PCR buffers are used containing Tris–HCl, KCl, and,
optional, MgCl buffers are provided by the supplier together with Taq Polymerase. It is important to use Polymerase and
20
3. Oligonucleotide primers: PCR primers are short, single-stranded DNA fragments (usually, 20–30 nucleotides).
PCR requires one forward and one reverse primer to assign the favored fragment of the DNA.
4. Deoxynucleotide triphosphates (dNTPs): dNTPs (dATP, dTTP, dGTP, and dCTP) are the nucleotide bases added by
the DNA Polymerase during the synthesis of the template strand. They are available as single ingredients or as a dNTP mix
(e.g., Fermentas). There should always be a slight surplus of dNTPs in the reaction mix. For PCR, a final concentration of 2
5. Additives: Additives, like MgCl 2, trehalose, DMSO, Q-solution, etc., can enhance PCR efficiency. Use additives only
if standard protocols do not work. Too high concentrations of MgCl 2, for instance, increase the amount of unspecific
6. Molecular water: Use only ultra-pure and nuclease-free water for PCR. Water is used to fill the mix of ingredients
7. Template DNA: This is the original genomic DNA material. Use 1–2ml DNA solution (obtained from extraction)
with a concentration between 20 and 100 ng/ ml. Usually, PCR also works well with lower concentrations (below 2 ng/ ml)
PCR PROCEDURE
1. Initial denaturation: Melting of double-stranded DNA in two single-stranded templates by disrupting the hydrogen
bonds between complementary nucleotides. This step is usually performed at a temperature of 94°C for about 5 min. If the
template DNA is GC rich, the interval should be extended up to 10 min. Heating the lid is recommended and normally an
2. Denaturation: Similar to the initial denaturation, this step leads to the melting of the double-stranded DNA into single
strands for primer annealing. Amplified DNA with high GC content needs increased denaturation time (3–4 min).
21
3. Annealing: In most cases, temperatures between 50 and 65°C allow successful annealing of primers to the single-
template DNA strands. Typically, the optimal annealing temperature (T a ) is 3–5°C below the melting temperature of the
4. Elongation: In this step, the DNA Polymerase synthesizes a new DNA strand complementary to the template strand by
adding dNTPs. The optimal elongation temperature is dependent on the Polymerase itself and the length of the desired
fragment. In the case of Taq DNA Polymerase, the highest synthesis rates can be performed at 70–75°C. For fragments up
For longer fragments, more elongation time is needed (and vice versa for smaller fragments).
5. Number of cycles: Now, steps 2–4 are repeated several times (cycles). The number of cycles depends on the amount of
template DNA. If the initial DNA quantity is low, up to 40 cycles can be performed. For higher amounts of template, 30–35
6. Final elongation: After the last PCR cycle, a final elongation is performed to ensure that all remaining single DNA
strands are fully extended. It is usually performed at 72°C for 5–10 min.
7. Cooling (optional): After the final elongation step, samples can remain in the Thermocycler if reactions are performed
overnight. For cooling overnight, use a temperature of 15°C. This temperature neither damages PCR products nor strains
the heating block too much. Subsequently to amplification, the PCR products can be stored for a while in the fridge (4°C)
PROTOCOL FOR ONE SAMPLE MATER MIX USING TAQ DNA POLYMERASE
With a 25 ml PCR reaction volume; dispense 24 ml of the master mix to 1ml of the template DNA:
22
1. Molecular-grade water: 15.875 ml.
The reaction volume containing the master mix and DNA template is vortexed using a vortex mixer and then centrifuge
briefly before placing in the thermocycler to kick start the PCR Procedure as stated above.
1. Prepare loading dye and Molecular Size Marker (100 bp DNA Ladder Plus, Fermentas) according to the
3. In case of a usual 100-ml gel, weigh 1.0 g agarose powder and add 100 ml 1× TBE buffer. Boil the mixture in a
microwave until the agarose powder is completely dissolved. Add 2 ml (or one drop) ethidium bromide or 10 m l GelRed
and shake carefully. Immediately pour the mix into the prepared tray and wait until the agarose gel is solid which takes
23
4. Apply the gel to an adequate electrophoresis chamber filled with 1× TBE buffer. The gel should be completely dipped.
Remove the combs. Mix 2 ml of each PCR product with 2 m l loading dye (prepared in a microtiter plate according to the
number of samples) and pipette up and down a few times to mix. Load the PCR samples into the pockets or Wells in the gel.
5. Connect voltage (90 V) and let samples run for about 30 min. Afterward, the double-stranded PCR products can be
viewed in ultraviolet light. Take a photo to select samples for the cleanup. Sharp bands indicate successful amplification of
the desired DNA fragment (Fig. 8 ). When PCR fails, no bands are present.
The molecular ladder is in the middle. Lanes 1 – 4 show very intense and sharp PCR products. In lanes 5 – 8, DNA
amplification failed, and only unconsumed primers are visible. (Thomas and Isabella, 2012).
Upon successful PCR amplification, the PCR product is purified from the PCR ingredients via ethanol precipitation, enzymatic
digestion, or the use of commercial kits before it can proceed into step 3 (sequencing).
1. Put ethanol (100% and 70%) into the freezer: both need a temperature of −20°C when applied. Mark the appropriate
24
2. Add one-tenth of the amount of the PCR product of 3 M sodium acetate to the PCR product (e.g., if you have 10 ml
PCR product, then add 1 ml 3 M sodium acetate to the complete PCR product).
3. Add two volumes of 100% ethanol (−20°C) to one volume of the PCR product (e.g., if you use 10 ml PCR product, then
5. Discard supernatant without discarding the pellet that might be visible or not.
6. Add 200 ml 70% ethanol (−20°C) onto the pellet to wash it.
8. Discard supernatant.
9. Dry the pellet at room temperature or in a Thermomixer at 37°C to remove the residual ethanol.
10. Dissolve the pellet in 30ml molecular water and the cleaned-up PCR product is again ready to use for the sequencing
reaction.
The process of working out the order of the building blocks or bases in a strand of DNA. There are different Sequencing
2 Maxam-Gilbert method
25
Automated DNA sequencing gives a better visual representation of the base sequence in a DNA segment. The PCR-
cleaned DNA is transferred to a plate mixed with free DNA bases, DNA polymerase enzyme, DNA primers, and modified
DNA bases labelled with coloured florescent (called Terminator bases). These reactants are placed in a sequencing
machine. The following steps explain how Automated Sequencing is carried out:
Step 2: The temperature is lower to 500 C enabling the DNA primers to bind to the template strand.
Step 3: The temperature is increased to 600 C enabling the enzyme (DNA Polymerase) to add base pairs to the primers until
Through Gel Electrophoresis the Terminator base of the shorter DNA fragments reach the positive (+) end quicker than the
larger fragments- A laser makes the Terminator bases light up ( ddA lights green, ddC blue, ddG orange, ddT red). The
colour of the fluorescent light is analyzed by a detector and the information is transmitted to a programmed software that
26
Figure 9: Automated DNA Sequencing method, Source: Google, 2019
Before using a database to compare a query barcode with other species the first step is to validate the quality of the
Barcode data to certify that the sequences have been properly assembled, edited, and aligned. To do that, users will need
to have some basic knowledge about sequence editing and what to look for when editing sequences. GenBeans, UGENE,
BioEdit, and Serial Cloners are examples of DNA sequence editing software.
After editing the DNA sequence, a database is selected to identify species. The choice of database affects conclusions, so
care must be taken that the database reflects the scientific aims of a study. Once an appropriate database has been
selected e.g The Barcode of Life Data System (BOLD) database, the computer must assign a species from the database to
The next step, therefore, is to select a computer algorithm for assigning each specimen and its barcode sequence to the
unknown species. No algorithm seems to improve noticeably on assigning to a specimen the species of its nearest neighbor
within a barcode database. Thus, many algorithms begin by estimating a “separation” between the barcode sequences in
two specimens. (The term “separation” is preferable to “distance”, which connotes some specific mathematical properties
(3) Evolutionary distances (which usually require prior alignment of the barcode
sequences)
27
Studies have compared different measures of separation, but they are too limited to draw definitive conclusions about
which separation provides the best species assignments. There are, however, some distinctly bad measures of separation.
Like any assignment method, species assignment should use all available information. BLAST (Basic local alignment search
tool) is a popular sequence comparison tool or algorithm, but as a measure of separation, it can mislead, because it
compares two sequences with local alignment, which matches and scores only the two most similar subsequences within
two sequences.
Global alignment, which matches the entire length of sequences, is better for measuring the separation of barcode marker
sequences. So on the one hand a BLAST local alignment might make distant species appear spuriously or falsely close. On
the other hand, a global alignment might resolve the species by highlighting dissimilarities across the whole marker. In the
context of barcodes, therefore, a global alignment (e.g., with some close relative of the Needleman–Wunsch Algorithm) is
generally preferable to a local alignment (e.g., with the Smith-Waterman Algorithm or BLAST). Other types of alignments
exist, but there is little reason to expect them to assign species notably better than global alignment.
With an appropriate database and species assignment algorithm in hand, a scientist interested in barcode efficacy must
measure the algorithm’s success in identifying species. Any reasonable measure of barcode efficacy should reflect the
probability that a database based on the prospective barcode identifies a specimen’s species correctly. Consensus has
therefore emerged on “the probability of correct identification” (PCI) as the appropriate measurement of barcode efficacy.
Consider a particular data set, and assume that PCI can be defined for each species within the data set. The overall PCI for
the data set is the average of the species PCIs, taken over all species in the data set. On success, the species PCI is 1; on
failure, it is 0.
PCI should estimate the success in correctly identifying a known species. Under present technology, species identification
28
3. PCR primers must amplify it.
4. It must be sequenced.
29
Figure 10: Species identification using BOLD database Source: Google image, 2019
30
DNA BARCODING TODAY
3. Can process a great number of specimens at a time, thus is useful for example in biodiversity surveys.
1. It is not always true that intraspecific variability is negligible, or at least lower than interspecific values
3. The validity of DNA bar coding depends on establishing reference sequences from taxonomically confirmed
4. Its application for finished herbal/botanical dietary supplements is limited due to the generally low quality of DNA
in those products.
5. Experts are required from the DNA extraction step to Sequencing Step
1. Biodiversity studies
31
8. Tracking adulterations
9. Education
32
3. National Centre for Biotechnology Information GenBank (NCBI GenBank)
33
CONCLUSION/RECOMMENDATIONS
DNA Barcoding provides a faster, and better method of species identification than classical taxonomy, however, it requires
a huge amount of expertise, equipment, and discretion. In experiments geared towards the identification of plant species,
more than one candidate gene should be selected as the barcode. More research should be done to enable scientists to
generate novel methods to reduce the dependency of DNA barcoding on classical taxonomy. The Nigerian government
should train plant and animal scientists in the area of DNA barcoding as the benefits of this biotechnology tool are arguably
limitless.
34
ACKNOWLEDGEMENT
I acknowledge God Almighty for guiding me through this work. To my supervisor and postgraduate coordinator, Prof. H. O.
Shittu, who has been and will always be a driving force behind my academic passion. God bless you sir for your
understanding, tolerance, and benevolence. I thank the Head of the Department, Prof. F. Okungbowa for her role in making
sure things are done as and when due in the department. My utmost appreciation goes to my boss, the owner of Charryville
Academy for her tolerance, encouragement, and support, and all my lecturers and the non-academic staff for their
mentorship and advice, God bless you all. My acknowledgment also goes to my colleagues and friends. My family members,
In reality, this work would not have been possible without the aforementioned personalities. Dreams do come through with
God’s help and a cluster of those who have your best interest at heart.
35
REFERENCES
[Link]
[2] C. Costion, A. Ford, H. Cross, D. Crayn, M. Harrington, and A. Lowe, “Plant DNA barcodes can accurately estimate species richness
[3] U. Eberhardt, “Methods for DNA barcoding of fungi,” Methods Mol. Biol, Vol 858 (2012), pp.183-205.
[4]C. Garcia-Robledo, D. Erickson, C. Staines, and T.K. Erwin, “Tropical plant-herbivore networks: reconstructing species interactions using
[5] M. Hajibabaei, and C. McKenna, “DNA mini-barcodes,” Methods in molecular biology, Vol 858 (2012), pp. 53-339.
[6] P. Hebert, A. Cywinska, S. Ball, and J. Deward, “Biological identifications through DNA barcodes,” Proc Biol Sci, Vol 270 No 1512 (2003), pp.21-313.
[7] P.
Hebert, Y. Mark, Z. Tyler, Z. and M. Charles, "Identification of birds through DNA barcodes," PLoS Biology, Vol 2 No 10 (2004), pp.1657-1663.
[8] A.
Henras, C. Pilisson-Chastang, M. O'Donohue, A. Chakraborty, and P. Gleizes, "An overview of pre-ribosomal RNA processing in eukaryotes,"
[9] DNA Barcoding: A Tool For Specimen Identification And Species Discovery. March 12, 2021. Retrieved from
[Link]
[10] S. Joly, J.A. Davies, A. Bruneau, A. Derry, W. Kembel, P. Peres-Neto, and A. Wheeler, "Ecology in the age of DNA barcoding: the resource,
the promise and the challenges ahead," Molecular Ecology Resources, Vol 14 No 2 (2014), pp. 221-232.
[11] R. Ng'endo, Z. Osiemo, and R. Brandl, "DNA barcodes for species identification in the hyperdiverse ant genus pheiodole
(Formicidae: Myrmicinae)," Journal of Insect Science, Vol 13 Nov 27 (2013), pp. 13-14.
[12] Z. Pecnikar, and E. Buzan, "20 years since the introduction of DNA barcoding: from theory to application," Journal of
[13] L.
Sarvananda, " Short introduction of DNA Barcoding," International Journal of Research, Vol 5 No 4 (2018), pp. 673-686.
[14] D.M. Spooner, " DNA barcoding will frequently fail in complicated groups: An example in wild potatoes," Am J Bot, Vol 96 N0 6 (2009), pp.89-1177.
[15] K.
Thomas, and S. Isabella, "DNA Extraction, Preservation, and Amplification," Methods in molecular biology, Vol 858 (2012), pp.38-311.
[16] A.
Travers, and G. Muskhelishvili, "DNA Structure and function," FEBS, Vol 58 (2015), pp. 2279-2295.
[17] G. Ude, O. Igwe, J. Cormick, and O. Ozokonkwo, "Genetic Diversity and DNA Barcoding of Yam Accessions from Southern Nigeria,"
36
[18] S.]Vasudevan, M. Garcia-Blanco, S. Bradrick, and C. Nicchitta, "Flavivirus RNA transactions from viral entry to genome replication,"
37