0% found this document useful (0 votes)
12 views14 pages

DNA Sequence Analysis and Methods

The document provides an overview of DNA sequence analysis, detailing the structure of genes, the process of DNA sequencing, and various sequencing methods such as Maxam-Gilbert and Sanger sequencing. It also discusses the importance of DNA sequencing in applications like genetic disease diagnosis, personalized medicine, and forensic science, along with techniques like PCR and hybridization. Additionally, it covers concepts like sequence alignment and the algorithms used for comparing biological sequences.

Uploaded by

ritikajha0604
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
12 views14 pages

DNA Sequence Analysis and Methods

The document provides an overview of DNA sequence analysis, detailing the structure of genes, the process of DNA sequencing, and various sequencing methods such as Maxam-Gilbert and Sanger sequencing. It also discusses the importance of DNA sequencing in applications like genetic disease diagnosis, personalized medicine, and forensic science, along with techniques like PCR and hybridization. Additionally, it covers concepts like sequence alignment and the algorithms used for comparing biological sequences.

Uploaded by

ritikajha0604
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

DNA SEQUENCE ANALYSIS

UNIT – 3
Q. What is a Gene made of?
A gene is made up of a chemical molecule called DNA (Deoxyribonucleic
Acid).
DNA is like a long-twisted ladder — also known as a double helix — and it
carries all the information needed for the growth, development, and
functioning of living things.

Main components of a Gene:

 DNA – The main substance that makes up genes.


 Nucleotides – Building blocks of DNA.
Each nucleotide has:
o a sugar (deoxyribose),
o a phosphate group, and
o a nitrogen base.
 Nitrogen Bases – These are the “letters” of the genetic code:
o A = Adenine
o T = Thymine
o C = Cytosine
o G = Guanine

These bases always pair as:

 A with T
 C with G
 Base Pairs – The order (sequence) of these bases forms the genetic
code, which determines traits like eye color or blood type.

Q. What is DNA Sequencing?


DNA sequencing is the process of determining the exact order of the
nucleotides (A, T, C, G) in a DNA molecule.

Q. Why is DNA Sequencing Important?


 To identify genes and their functions
 To detect genetic mutations that may cause diseases
 To study evolutionary relationships between species
 To help in medicine, like personalized treatments or drug
development

DNA Sequencing Methods


1. Maxam-Gilbert Sequencing (Chemical Cleavage Method)
Maxam-Gilbert sequencing was one of the first methods
developed to determine the order of bases in DNA.
It was created by Allan Maxam and Walter Gilbert in the 1970s.
 A radioactively labeled DNA fragment is prepared.
 The DNA is chemically treated so that it breaks at specific bases:
o One reaction cut at G (Guanine)
o Another at A+G (Adenine and Guanine)
o Another at C (Cytosine)
o Another at C+T (Cytosine and Thymine)
 The resulting fragments are separated by size using gel
electrophoresis.
 By reading the pattern of fragments on the gel, the DNA
sequence can be determined base by base.

2. Sanger Sequencing (Chain-Termination Method)


Sanger sequencing was developed by Frederick Sanger in 1977.
It is one of the most widely used methods for determining the order
of nucleotides (A, T, C, G) in DNA.

1. DNA Preparation

 Take a single-stranded DNA to use as a template.


 Attach a primer to start copying the DNA.

2. DNA Copying with Special Nucleotides

 Add normal nucleotides (A, T, C, G) and special ddNTPs.


 ddNTPs stop the DNA from growing, making fragments of
different lengths.
3. Separating DNA Fragments

 Separate the DNA pieces by size using a gel.


 Each fragment ends at a specific base (A, T, C, or G).

4. Reading the Sequence

 Look at the fragments from shortest to longest.


 This tells the order of the DNA letters (A, T, C, G).

Difference between them:

Feature Maxam-Gilbert (Chemical Sanger (Chain-Termination


Method) Method)
Developed by Maxam & Gilbert Frederick Sanger

Method Uses chemicals to cut DNA at Uses special nucleotides


specific bases (ddNTPs) to stop DNA copying
DNA Needed Radioactively labeled DNA Single-stranded DNA + primer

How DNA is By looking at fragment pattern By looking at fragment lengths


read on a gel on a gel

Best for Small DNA fragments Small to medium DNA sequences

Safety & Ease Uses toxic chemicals, more Safer and easier to use
complicated

Q. Explain Automated DNA Sequencing.


Automated DNA sequencing is a modern, machine-based method to
determine the exact order of nucleotides (A, T, C, G) in DNA.
It is an improved version of Sanger sequencing, using fluorescent
dyes and computers for faster and more accurate results. It is fast,
accurate, and widely used in research, medicine, and genetics.
Advantage of Automated DNA Sequencing over manual
DNA Sequencing
Feature Automated DNA Manual DNA Sequencing
Sequencing (Traditional
Sanger/Maxam-Gilbert)
Speed Very fast, can sequence Slow, sequences one
thousands of fragments at fragment at a time
once
Accuracy Highly accurate, fewer More prone to human errors
mistakes
Detection Uses fluorescent dyes and Uses radioactive labels or
computer detection manual reading, slower and
less safe
Safety Safer, no toxic chemicals Maxam-Gilbert uses
hazardous chemicals
Data Computer automatically Requires manual
Handling reads and stores interpretation, time-
sequences consuming
Throughpu Can handle large-scale Best for small DNA
t projects (like whole fragments only
genomes)
Cost & Less labor-intensive in the Requires more hands-on
Labor long run work

Applications of DNA Sequencing


 Gene Identification:
o Helps scientists find specific genes and understand their
function.
 Genetic Disease Diagnosis:
o Detects mutations in genes that cause diseases like cystic
fibrosis, sickle cell anemia, or cancer.
 Personalized Medicine:
o Doctors can tailor treatments based on a patient’s unique DNA
sequence.
 Evolution and Phylogenetics:
o Helps study relationships between species and evolutionary
history.
 Forensic Science:
o Used in crime investigations to identify people from DNA
evidence.
 Agriculture and Biotechnology:
o Helps develop disease-resistant crops or improve livestock
by understanding their genes.
 Drug Discovery:
o Identifies targets for new drugs by studying genes and
proteins involved in diseases.
 Microbiology and Virology:
o Sequencing viruses and bacteria helps track disease outbreaks
and develop vaccines.

DNA mapping is the process of determining the location of genes or


markers on a DNA molecule.

DNA Assembly is the process of putting small DNA fragments together


to reconstruct the entire DNA sequence of an organism.

Size of Human DNA


 The human genome contains DNA in 23 pairs of chromosomes.
 Total DNA length in a human cell: about 3 billion base pairs.
 If stretched, DNA from one human cell is about 2 meters long.
 The total DNA in the human body (all cells) would be approximately 10
billion miles

Copying DNA (Polymerase Chain Reaction)


PCR is a lab method used to make many copies of a specific DNA piece.

 Developed in 1983 by Kary Mullis.


 Can produce millions of copies of a small DNA segment.
 Widely used in molecular biology and biotechnology labs.

Principle of PCR

PCR works on the principle of DNA replication — the natural process by


which DNA makes copies of itself.
 DNA is heated to separate its two strands (denaturation).
 Short DNA pieces called primers bind to the target sequence
(annealing).
 DNA polymerase makes a new strand of DNA starting from the
primers (extension).
 By repeating these cycles, the specific DNA segment is copied
millions of times.

Components Of PCR

DNA Template: The DNA segment you want to copy.

DNA Polymerase: The enzyme that builds new DNA strands. Taq
Polymerase is used.

Oligonucleotide Primers: Short DNA pieces that bind to the start and end
of the target sequence.

Deoxyribonucleotide triphosphate: The building blocks (A, T, C, G) are


used to make new DNA strands.

Buffer System: Maintains the right pH and salt conditions for the
enzyme to work.

PCR Steps

Denaturation:

 The DNA is heated to 94°C for 30 seconds to 2 minutes.


 This breaks the hydrogen bonds holding the two DNA strands
together.
 The DNA becomes single-stranded, ready to be copied.
 Longer heating ensures complete separation of the strands.

Annealing

 The temperature is lowered to 54–60°C for 20–40 seconds.


 Primers attach (bind) to their matching sequences on the single-
stranded DNA.
 Primers are short DNA pieces (about 20–30 bases long).
 They act as the starting point for new DNA synthesis.
 Because DNA strands run in opposite directions, there are two
primers:
o Forward primer
o Reverse primer

Elongation

 The temperature is raised to 72–80°C.


 Taq polymerase enzyme adds DNA bases to the 3’ end of the
primer.
 DNA is built in the 5’ → 3’ direction, forming a new DNA strand.
 Taq polymerase works fast, adding about 1000 bases per minute
and can tolerate high temperatures.
 At the end of this step, a double-stranded DNA is formed.

Applications of PCR

 Medical Diagnosis
o Detects genetic disorders and diseases (like cystic fibrosis,
sickle cell anemia).
o Identifies viruses and bacteria, e.g., COVID-19, and HIV.
 Forensic Science
o Helps in crime investigations by analyzing DNA from hair,
blood, or other evidence.
 Genetic Research
o Used to study genes, mutations, and hereditary traits.
 DNA Cloning and Sequencing
o Produces millions of DNA copies for cloning or sequencing
experiments.
 Evolutionary Biology
o Helps compare DNA from different species to study evolution.
 Agriculture and Biotechnology
o Detects genetically modified organisms (GMOs) or
plant/animal disease genes.

Hybridization
Hybridization is the process where two complementary single-stranded DNA
or RNA molecules bind together to form a double-stranded molecule.

 Complementary base pairing (A-T and C-G in DNA, A-U and C-G in
RNA) drives the binding.
 Used to detect specific DNA or RNA sequences.
 Works because single-stranded nucleic acids naturally stick to their
matching sequences.

Principal of hybridization
Hybridization works on the principle of complementary base pairing in
nucleic acids:

1. Single-stranded DNA or RNA molecules can bind to


complementary sequences.
2. Adenine (A) pairs with Thymine (T) in DNA or Uracil (U) in RNA,
and Cytosine (C) pairs with Guanine (G).
3. When complementary sequences meet, they form stable double-
stranded structures.
4. This process is specific — only matching sequences can bind tightly.

Microarrays
 Microarrays are tiny DNA chips with many DNA probes fixed on a
solid surface (like a glass slide).
 Probes are short DNA pieces (oligonucleotides) made by printing or
chemical synthesis.
 Sample DNA or RNA is labeled and applied to the microarray.
 The sample sticks (hybridizes) to matching probes on the chip.
 The strength of the signal shows how much of a specific DNA or RNA
is present in the sample.
Why cutting of DNA is essential?
Cutting of DNA is an important step in molecular biology because DNA
molecules are very long and difficult to work with in their natural form. By
cutting DNA into smaller fragments, scientists can study specific genes or
regions more easily. This process is essential for techniques like gene
cloning, sequencing, and recombinant DNA technology, where
fragments of DNA are inserted into vectors or combined with DNA from other
sources. DNA cutting also allows researchers to analyze and manipulate
genetic material for experiments such as PCR, hybridization, or
constructing DNA libraries. Special enzymes called restriction enzymes are
commonly used to cut DNA at precise locations, making the process
accurate and predictable. Overall, cutting DNA is a key step that makes
genetic research and biotechnology possible.

Sequencing Short DNA Molecules:


Sequencing short DNA molecules involves determining the exact order of
nucleotides (A, T, C, G) in a small piece of DNA. These short fragments are
usually 50–1000 base pairs long, making them easier to handle than long
DNA. Techniques like Sanger sequencing are commonly used for short
DNA because they are accurate and reliable. The DNA is first copied
many times using PCR, then special chain-terminating nucleotides
are added to generate fragments of different lengths. These fragments are
separated by size using gel or capillary electrophoresis, and the
sequence is read based on the pattern of fragment lengths.

Mapping Long DNA Molecules


Mapping long DNA molecules is the process of finding the positions of
genes or markers along a long stretch of DNA. Unlike short DNA fragments,
long DNA is too big to read all at once, so it is cut into smaller pieces, and
each piece is sequenced. Then, computational tools are used to
assemble the sequences in the correct order, creating a map of the
entire DNA molecule. This process helps scientists understand the
organization of genes, locate important genetic markers, and study
large genomes, like the human genome.
DeBruijn Graph
 A De Bruijn graph is a tool used to assemble DNA from many short
DNA fragments.
 DNA is broken into small pieces called k-mers (like puzzle pieces).
 Nodes represent overlapping parts of k-mers, and edges connect
them.
 By following the edges, scientists can reconstruct the full DNA
sequence.

Sequence Alignment
Sequence alignment is considered the most essential step in
comparing biological sequences. Sequence alignment is the process of
lining up two or more DNA, RNA, or protein sequences.

The goal is to find similar regions between the sequences. These similarities
help scientists understand their function, structure, and evolutionary
relationships.

Two commonly used Sequence alignment algorithms are:


Global alignment:

 Aligns the entire length of two sequences.


 Maximizes the overall similarity.
 Best for sequences of similar lengths.

Local alignment:
 Aligns only the most similar regions of the sequences.
 Useful for finding short, conserved regions in DNA, RNA, or
proteins.

Types of Sequence Alignment


Pairwise Alignment
 Pairwise sequence alignment compares two sequences to find the
best match between them.
 It uses a scoring system:
o Positive points for matches
o Negative points for mismatches or gaps
 The goal is to get the highest score, showing how similar the two
sequences are.

Multiple Sequence Alignment


 MSA aligns three or more DNA, RNA, or protein sequences to find
the best overall match.
 It helps to identify conserved regions shared among sequences.
 MSA is used to study evolutionary relationships and build
phylogenetic trees.

Difference between:
Feature Pairwise Alignment Multiple Sequence
Alignment (MSA)
Number 2 sequences 3 or more sequences
of
sequence
s
Purpose Find similarity between two Find similarity and conserved
sequences regions among many
sequences
Output Shows the best match and Shows conserved regions
score between 2 sequences across all sequences and can
help build phylogenetic trees
Use Simple comparison Functional and evolutionary
analysis
Dynamic programming
 Dynamic programming helps to find the best (optimal) alignment
between two DNA, RNA, or protein sequences.
 It works by comparing all possible pairs of characters in the
sequences.
 DP can be used for both global and local alignments:
o Global alignment → Needleman-Wunsch algorithm
o Local alignment → Smith-Waterman algorithm

The Needleman Wunsch


algorithm
The problem of finding the best possible alignment for 2 sequences is solved
by the Needleman Wunsch algorithm. The N – W algorithm takes time
proportional to n^2 to find the best alignment of two sequences.

Steps of the Algorithm:

 Initialization
o Create a matrix with one sequence along the top and the other
along the left.
o Fill in the first row and column with gap penalties.
 Matrix Filling
o Fill each cell with the maximum score based on:
 Match → positive score
 Mismatch → negative score
 Gap → penalty
o Each cell score = max (top + gap, left + gap, diagonal +
match/mismatch)
 Traceback
o Start from the bottom-right corner of the matrix.
o Move back to the top-left, following the path that gives the
highest score.
o This gives the optimal global alignment of the two sequences.

The Smith Waterman algorithm


The concept of ‘local alignment’ was introduced by the Smith Waterman
algorithm. S – W is mathematically proven to find the best (highest-scoring)
local alignment of 2 sequences.

Steps of the Algorithm:

 Initialization
o Create a matrix with one sequence along the top and the other
along the left.
o Fill the first row and column with zeros (no gap penalties at the
start).
 Matrix Filling
o Fill each cell with the maximum score based on:
 Match → positive score
 Mismatch → negative score
 Gap → penalty
 Zero → ensures only positive scores are kept
o Each cell score = max (top + gap, left + gap, diagonal +
match/mismatch, 0)
 Traceback
o Start from the cell with the highest score anywhere in the
matrix.
o Move back to a cell with a score of 0, following the path of
maximum scores.
o This gives the best local alignment between the sequences.

You might also like