0% found this document useful (0 votes)
3 views12 pages

Module 4 Sequencing

Uploaded by

drn6v9fc4m
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views12 pages

Module 4 Sequencing

Uploaded by

drn6v9fc4m
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

DNA Sequencing

Sanger Sequencing

The first widely adopted method for determining the sequence of bases in DNA was
developed by double Nobel prize winner, Paul Sanger. Sanger’s method is a type of
sequencing-by-synthesis, because it depends on the production of a new DNA strand from a
single-stranded template. Sanger sequencing is also called dideoxy sequencing, because it
relies on the incorporation of modified nucleotides, called ddNTPs. Natural nucleotides are
deoxy NTPs (dNTPs), and have a 3’OH group on the ribose sugar, to which the next
nucleotide can be joined in a phosphodiester bond. The incorporation of a ddNTP
permanently terminates the elongation of the nascent (i.e. growing) strand. Sanger’s method
takes advantage of this by adding a label (such as a fluorescent dye) to each of the four
ddNTP bases. This means that the researcher can detect which base was the last base added
in the sequence, by simply detecting the colour of the molecule’s fluorescence. This is useful
because DNA molecules can be separated very precisely based on their length, using
electrophoresis. Originally, long slab gels were used for electrophoresis, but capillary
electrophoresis is now the state-of-the art for Sanger sequencing. Capillary sequencers have
a gel in a very long, very thin tube. An electric current pulls DNA through this tube, with
shorter molecules moving faster. All the molecules of a given length with exit the tube at the
same time, and their fluorescent signal can be read by a specialized laser and detector.

Figure 1. dsDNA molecules (millions of copies of the same molecule) are denatured and a primer is annealed
to the site where sequencing will start. A mixture of unlabeled dNTPs, and fluorescent ddNTPs are used in the
synthesis of a new strand on each molecule. Each molecule is randomly terminated with a fluorescently
labeled ddNTPs.

So, putting this all together, a Sanger sequencing reaction contains millions of copies of
identical single-stranded DNA template molecules, as well as a mixture of both normal
dNTPs, and fluorescently labeled ddNTPs. At each position of each growing strand, there is a
possibility that a ddNTP, rather than a dNTP, will be incorporated. Because there are millions
of templates, at the end of the reaction, there will be nascent molecules terminated at each

2020-10-04 BIOL 477/577 page 1


base along sequence, and identity of that base can be read from the fluorescence of the
molecule. When the molecules are separated by length, the sequence (complementary to the
template) can be read by reading the fluorescence of each length of molecule, from shortest
to longest. If the reaction is analyzed on a capillary sequencer, the output is called a
chromatogram, which shows the fluorescent intensity of each of the four colours, at each
position of the sequence. Software is then used to infer the sequence from the
chromatogram, or “call” the bases.

Figure 2. Fluorescently labeled molecules are separated based on length, using electrophoresis.

As described above, the Sanger sequencing reaction contains, dNTPs, labelled ddNTPs, and
template DNA. The template DNA is normally dsDNA, but can be made single stranded by
heating. In addition, DNA polymerase (DNApol), and its co-factor Mg2+ are also required in
the reaction, and finally a primer, which is also required by DNA polymerases. A primer is a
short, single-stranded DNA (ssDNA) molecule, approximately 12-24 nucleotides, which is
made (i.e. chemically synthesized) to be complementary to the position on the template
where sequencing is intended to begin.

A single Sanger sequencing reaction can produce up to 700 bases of output, but only one
template can be sequenced in each reaction. Each reaction costs $1-$5. The human genome
and a few of the other first whole genome sequencing projects were sequenced using Sanger
sequencing.

Next-generation sequencing

Beginning around 2005, several alternatives to Sanger sequencing emerged. These “next-
generation” sequencing (NGS) technologies all differ in their chemistry and detection
methods, but they all have a few things in common: miniaturization and parallelization. By
making the reactions physically smaller, they become less expensive, and sequence millions of
different templates in parallel, within the same reaction tube. As a result, compared to Sanger
sequencing, each next-generation sequencing reaction costs more ($1000-$5000), but can
sequence much, much more DNA. Still, if you are interested in only a single template, it
makes more sense to use Sanger sequencing.

2020-10-04 BIOL 477/577 page 2


Roche 454 pyrosequencing

Pyrosequencing, commercialized as Roche 454, was the first widely-used NGS method. It is
now obsolete, but it is still worth examining this technology to contrast it with other systems.

The basis of 454 sequencing was detection of the pyrophosphate (PPi) that is released when
a nucleotide is adding during DNA synthesis. The release of the PPi was detected in a two-
stage enzymatic reaction that resulted in the emission of a photon by luciferase (an enzyme
isolated from fireflies). So, unlike Sanger sequencing, no ddNTPs or other types of
terminators are used in 454 sequencing. Instead, unmodified dNTPs are added to each
template, one at a time (e.g. dATP), and the emission of light is used to infer the incorporation
of that base into the sequence (if no light is emitted, the base was not required at that position
in the sequence). One outcome of this is that if more than one base is incorporated in a row,
the amount of light released will be proportionately higher. Thus, the detector measures the
light intensity emitted for each nucleotide at each position in the sequence and infers whether
0, 1, or more of those bases is present in the sequence. This highlights the major problem
(and limitation) of pyrosequencing: the relationship between light emission and the number of
nucleotides is only linear up to approximately four of the same nucleotides in a row. Above
this, this system cannot accurately determine the number of bases. This is called the
“homopolymer” problem.

Figure 3. Emission of light as a result of PPi release.

A major challenge in miniaturizing sequencing reactions is getting enough copies of the same
template concentrated in a small space, so that sufficient signal is generated from the reaction
for detection. In the Roche 454 system, sequencing templates are ligated to adapters, which
are complementary to oligonucleotides that cover the surface of small beads. The beads and
templates are mixed in a dilution that ensures most beads have only one template attached to
them. The beads are then separated from each other in a water/oil emulsion. The emulsion is
by taking a solution contain the beads, plus PCR reagents, and shaking them in a volume of oil
under precise conditions. The result is millions of aqueous droplets suspended in oil, where
each droplet contains a different bead and template and serves as an independent reaction
vessel, allowing each template to be amplified by PCR . At the end of this process, there
should be millions of beads, and each bead should be covered by millions of copies of the
same template, with a different template on each bead. The beads are then distributed in tiny
wells on the surface of a “PicoTiter” plate, and the PicoTiter plate is put in an instrument that

2020-10-04 BIOL 477/577 page 3


conducts many cycles of the pyrosequencing reaction: the PicoTiter plate is flooded with a
single nucleotide (e.g. dATP), and all of the reagents for DNA polymerization and detection of
PPi. The emission of light from each bead is detected and measured, the unincorporated
dATP is removed, and the next nucleotide (e.g. dCTP is added). This is repeated for each
nucleotide, and then the cycle is started again, for over 100 cycles.

Figure 4. Emulsion PCR. Source: [Link]

Because there are no chain terminators in pyrosequencing, multiple nucleotides can be


incorporated in a single cycle, if the template allows it. For example, if a template has three T
based in a row, then three dATP nucleotides will be incorporated by the polymerase on the
growing strand, and three times the usual amount of light will be emitted. The sequencing
instrument measures the amount of light produced at each cycles and uses this to determine
the number of nucleotides that were incorporated. However, the relationship between the
number of bases in the template and the amount of light detected is not linear; it begins to
plateau around 4-8 bases. This is called the homopolymer problem: pyrosequencing is not
accurate when the same base is repeated consecutively more than a few times. This was one

2020-10-04 BIOL 477/577 page 4


of the major limitations of pyrosequencing, and part of the reason that it is no longer in use
today.

IonTorrent sequencing

IonTorrent is a sequencing-by-synthesis technology that uses only unlabeled dNTPs and


borrows heavily from semiconductor (computer chip) technology. The templates are
amplified on beads, similar to the emulsion PCR process in 454 sequencing, and the beads are
distributed in wells on the surface of a semiconductor chip. Each one of the four nucleotides is
added, one at a time, and the incorporation of nucleotides is measured by the release of
protons. The change in voltage in a well means that the nucleotide was incorporated. Like
454 sequencing, the voltage increase is proportional to the number of nucleotides added at a
time – so IonTorrent sequencing is susceptible to homopolymer problems as well. Because
the instrument has no special optics and the dNTPs have no special modifications, the
reaction costs are relatively low. However, IonTorrent has failed to capture as much of the
market as Illumina, which has been increasing the accuracy and efficiency of its technology
much more aggressively.

Illumina sequencing

Illumina sequencing is currently the dominant sequencing technology. It is the most like
Sanger sequencing of any of the NGS methods. In Illumina sequencing, all of the nucleotides
in the reaction are modified nucleotides. These are reversible fluorescent dye terminators
used, (rather than irreversible ddNTPs used by Sanger) and the sequencing is conducted on
millions of templates immobilized on the surface of a “flow cell”.

To create the templates on the surface of the Illumina flow cell, short adapter sequences are
ligated to the end of millions of different DNA fragments. These adapters are complementary
to short DNA sequences that are covalently attached to the surface of the flow cell, allowing
different DNA template to be annealed across the surface of the flow cell. In the next step,
each of the different templates is copied (amplified) locally, in a clever process called “bridge
PCR”. The result is that there are millions of copies of the same molecule in a dense cluster,
with many clusters across the surface of the flow cell. Each cluster was made from a different
DNA template. The amplification was necessary to increase the strength of the signal from
each reaction, which is detected in the next step.

2020-10-04 BIOL 477/577 page 5


Figure 5. Bridge PCR. Image source: Illumina.

As with Sanger sequencing, each template is annealed to a primer, and DNA polymerase,
Mg2+, and other co-factors are added to the reaction. Unlike Sanger, in Illumina, all of the
nucleotides in the reaction are modified terminators. This is feasible because the terminators
in Illumina sequencing are reversible. The Illumina flow cell is flooded by these reversible dye
terminators, a single nucleotide is incorporated at each template, and the fluorescent signal is
recorded in a photograph. (Because only one nucleotide is incorporated one at a time, there is

2020-10-04 BIOL 477/577 page 6


no problem with homopolymers.) After the image is recorded, the fluorescent dye and
terminator (but not the nucleotide) are removed from each nascent strand, and a new set of
modified terminators are added, and the cycle repeats again. There may be several hundred
cycles in a typical Illumina run, resulting in a series of several hundred images that must be
analyzed to determine the color of fluorescence at each cluster in each image.

Figure 6. Illumina sequencing. Source: Illumina

2020-10-04 BIOL 477/577 page 7


DNBSEQ

BGI is one of the world’s biggest DNA sequencing service providers, and it recently released
its own DNBSEQ sequencing instruments, which are based on the NGS technology of a
company they purchased, Complete Genomics. In this technology, genomic DNA is
fragmented and circularized, and then amplified by rolling circle replication (RCR) to produce
a concatemer of hundreds of identical templates, which then fold into a tangled “DNA
nanoball”. Separately, during template preparation, Type IIS restriction endonucleases are
used to introduce known adapter sequences at regular intervals throughout the template. The
DNA nanoballs are attached within wells of a flow cell, where they serve as templates for
combinatorial probe anchor sequencing (cPAS).

Image: Wikipedia user Tanxt

2020-10-04 BIOL 477/577 page 8


MGI

MGI

2020-10-04 BIOL 477/577 page 9


cPAS requires a large set of oligonucleotide probes for each base in the template. Nearly
every possible sequence is represented in each probe set, so at least one probe molecule will
be complementary to every position in the sequencing template. The probes are labelled with
a fluorescent dye corresponding to the base that is being interrogated by that probe set.
Binding of the correct probe is stabilized by ligation to an anchor sequence, and unbound
probes are washed away. The wavelength of fluorescence is recorded at each position on the
flowcell, and then the anchor and probe are removed, and a new anchor and probe set are
introduced, this time with fluorescent tags matching the base at the next position of the probe
set. The process continues for each base in the sequence, with increasingly longer anchors
allowing hybridization and ligation to occur at bases increasingly far from the anchor
sequence.

Nature Rev Genetics (2016) 17:333-351

PacBio sequencing

PacBio (and Oxford Nanopore) are sometimes called next-next-generation, or third


generation sequencing technologies, since they are fundamentally different from Illumina,
IonTorrent, and 454. One of the main differences is that PacBio sequences only a single
molecule in each reaction, although there are thousands of independent reactions per plate.
To do this, PacBio relies on highly sensitive optics that detect the release of fluorescence from

2020-10-04 BIOL 477/577 page 10


a labelled dNTP when a nucleotide is incorporated during DNA synthesis. This relies on both
a sensitive camera, and an immobilized polymerase at the bottom of special optical plate
(“zero-mode-waveguide”) that is released when dNTP is incorporated. This results in very
long reads; but with low accuracy (~87%) per read. However, many reads can be averaged to
increase accuracy; templates are often circularized so that the same template can pass
through the sequencing site dozens of times.

PacBio

Oxford Nanopore

Each of the hundreds of wells on an Oxford Nanopore sequences a single molecule at a time.
This is a type of direct sequencing: a single molecule of the template is passed through an
engineered protein nanopore embedded in a membrane. The shape (and therefore the
sequence) of the DNA changes the conductivity of the nanopore. By detecting changes in the
current as the DNA is threaded through the molecule, the sequence can be inferred. The DNA
is unmodified, except for that ligation of adapters that help tether and thread the DNA
through the nanopore. The Oxford Nanopore system can sequence molecules that are up to
millions of bases long. As with PacBio, the low accuracy of many individual reads can be
partly overcome by averaging the results of many reads. Oxford Nanopore sequencing

2020-10-04 BIOL 477/577 page 11


instruments are very small and inexpensive and can attach to laptops or even smartphones,
making them particularly useful for sequencing in the field or clinic.

Oxford Nanopore

2020-10-04 BIOL 477/577 page 12

You might also like