Genome and Proteome Analysis Overview
Genome and Proteome Analysis Overview
GENOMICS AND
Indira Gandhi PROTEOMICS
National Open University
School of Sciences
Block
4
APPLICATIONS OF GENOMICS AND
PROTEOMICS
UNIT 14
Analysis of the Genome 85
UNIT 15
Manipulation of the Genome 102
UNIT 16
Expression Analysis of Genome 134
UNIT 17
Proteome Analysis and Application of Proteomics 154
COURSE NAME: GENOMICS AND PROTEOMICS COURSE CODE: MZO-005
Course Coordinators : Prof. Aryadeep Roy Choudhury and Dr. Ravi Rajwanshi
Course Editor : Prof. Satheeshkumar P.K.
Center of Advanced Study in Botany, Institute of
Science, Banaras Hindu University, Varanasi
UP-221005, India
All rights reserved. No part of this work may be reproduced in any form, by mimeograph or any other
means, without permission in writing from Indira Gandhi National Open University.
Further information on Indira Gandhi National Open University courses may be obtained from the
University’s office at Maidan Garhi, New Delhi-110 068 or IGNOU website [Link].
Printed and published on behalf of Indira Gandhi National Open University, New Delhi by the Registrar,
MPDD, IGNOU.
BLOCK 4: APPLICATIONS OF GENOMICS AND
PROTEOMICS
Block 4, “Applications of Genomics and Proteomics”, consists of the four Units that
focus on the Genome-proteome analysis and their applications. The genome and proteome
are essential threads that tell the story of an organism's existence in the complex tapestry of
life. The proteome is the dynamic expression of genes that regulates the functions and
behaviors’ of living organisms, whereas the genome is similar to a blueprint that contains the
genetic information passed down through generations. By providing insights into evolution,
disease causes, and possible therapeutic approaches, the analysis and applications of these
genomic landscapes have revolutionised a variety of fields, including agriculture and
medicine. The present block will discuss the topics related to the analysis of the genome of
few important organisms including humans, various techniques for the manipulation of the
genome, expression analysis of the genome followed by the proteome analysis and its
applications for the benefit of mankind.
In Unit 14, you will learn about sequencing of DNA and the steps involved in the analysis of
the sequenced genome. The genome of few important organisms such as Plasmodium
falciparum, Mycobacterium tuberculosis along with humans has been discussed in detail.
Manipulation of genomes via PCR based Site-directed mutagenesis is also explained along
with its protocol.
The Unit 15 of this block focuses on the manipulation of the genome that involves the
techniques like cloning of genes as well as the role of reporter gene in the identification of the
locations of regulatory sequences like enhancer elements that drive a specific pattern of
gene expression. The Unit also discusses the Gene knockout and knockin methods in
transgenics that can alter genes in a selected model system, providing valuable insights into
the functioning of individual genes. In this Unit you will learn about the important genome
sequencing techniques and the role of Restriction Enzymes along with the cloning vectors in
the manipulation of the genome.
Unit 16 titled as “Expression analysis of genome” highlights the concept and function of
reporter genes. In this Unit, you will also learn about the basics of temporal and site-specific
gene expression. Also you will understand the detailed information provided in this Unit about
the Gene Silencing with special reference to the mechanism, biological functions and
applications of RNA interference.
Unit 17 is the last Unit of this block as well as course in which you will learn about the
Proteome analysis and applications of proteomics. This Unit will provide you the information
about proteomics and its various forms. The techniques of proteomics such as Two
dimensional gel electrophoresis, Two dimensional difference gel electrophoresis, Isotope
Coded Affinity Tag, Stable isotope labeling by amino acids in cell culture are discussed in
detail. The Unit also describes the different applications of proteomics in the field of
Pharmaceuticals, drug development and toxicology. Phage antibody as tool and application
of phage display in proteomics is also explained in this unit. You will also learn about the
various high throughput techniques used in proteomics based studies and application of
proteomics for the welfare of society. 83
Objectives
After studying this block, you would be able to:
• comprehend the mechanism of action of the restriction enzymes to cleave DNA and
steps of DNA cloning,
• explain the role of proteomics in drug development and toxicology and enumerate its
application in drug discovery in humans and pharmaceutical industry.
84
UNIT 14
$1$/<6,62)7+(*(120(
$1$/<6,62)7+(*(120(
6WUXFWXUH
6WUXFWXUH
14.1 Introduction Human Genome Project:
Implications for Medical
Objectives
Science
14.2 Genome
14.7 Genome Analysis of
DNA Sequencing Plasmodium Falciparum
Genome Analysis 14.8 Genome Analysis of
The Bulky Genomes in Plants Mycobacterium
Tuberculosis
14.3 Analysis of Genomes
Sequence Analysis
14.4 Steps of Genomic Data
Analysis Genes Encoding Proteins
14.1 INTRODUCTION
Genomic analyses allow researchers and clinicians to learn about differences
and changes in an organism’s or individual’s (be it an animal, a plant, a
bacterium, an archaean, a protist, a fungus, or a virus) genetic makeup,
leading to understand and discover the mechanism of disease progression
Block 4 Applications of Genomics and Proteomics
and help to design the treatment strategies. Genome projects ultimately aim to
determine the complete genome sequence of an organism and annotate
protein-coding genes, non-coding genes, and other important genome-
encoded features. The genome of an organism includes the complete DNA
sequences of each chromosome in the organism.
The present Unit provides a brief description of the steps involved in genome
analysis, features of the whole genome analysis of Plasmodium,
Mycobacterium tuberculosis, and the human genome project. This Unit also
descries the genetic modification methods used to manipulate the gene
sequences, especially by site-directed mutagenesis.
2EMHFWLYHV
2EMHFWLYHV
After studying this Unit, you would be able to:
14.2 GENOME
The complete set of DNA content in an organism is called its genome.
Virtually, every single cell in the human body contains a complete copy of 3
billion DNA base pairs approximately, that make up the human genome. DNA
contains the information needed to build and function the entire organism. A
gene refers to the unit of DNA that carries the information for making a specific
protein or set of proteins. For example, each gene in the human genome
codes for an average of three proteins.
Unit 14 Analysis of the Genome
fluorescently labeled deoxyribonucleoside 5’-triphosphates (dNTPs) into the
new DNA strand. The nucleotide is excited by a light source, resulting in the
emission of a fluorescent signal and detected by a detector.
Some of the questions which biologists want to answer using genome analysis
are:
Block 4 Applications of Genomics and Proteomics
Unit 14 Analysis of the Genome
The key to successful sequence analysis is the alignment of the sequence of
interest with another sequence whose function is known (reference genome).
This will reveal the function of the unknown genes, and also the evolutionary
relation between the sequences/organisms. Furthermore, the sequences can
be analyzed to find the significant matches between the domains of a
sequence that have been previously described which may have a huge impact
on the genes function.
Mostly, data analysis deals with imperfect data such as missing values or
measurements that are noisy. Data quality check and cleaning aims to
identify any data quality issues and rectify them by cleaning them from the
dataset. Identifying low-quality or missing bases and removing them from
the dataset will improve the read-mapping step.
This step involves the processing of data into a suitable format for
exploratory analysis and modeling. Sometimes, the data needs to be
converted into other formats by transforming data points (such as log
transformation, normalization, etc.), or fragment the data into subsets with
some arbitrary or pre-defined conditions. In terms of genomics, processing
includes aligning the sequence reads to the genome and quantification over
genes or regions of interest.
Block 4 Applications of Genomics and Proteomics
In the context of genomics, modeling is predicting the disease status of the
patients from the gene expression values measured from their tissue
samples, if the variable of interest is disease status. This kind of approach is
generally called “predictive modeling”, and involves regression-based
machine learning methods.
6$4
6$4
a) Choose the correct option in the multiple choices given below each
question.
A. Individual genes
B. Chromosomes
C. Whole genomes
D. Genetic disorders
B. Microarray analysis
C. Sanger sequencing
D. Gel electrophoresis
iii) The ……………… allows plants to form hybrids easily when pollen
and ova from different species fertilize.
Unit 14 Analysis of the Genome
The Human Genome Project was designed with the aim of generating a
resource that could be used for a wide range of biomedical studies such as the
study of genetic variations that increase the risk of certain diseases (like
cancer, and cardiac diseases), or to look for specific genetic mutations
frequently observed in cancerous cells. In its initial release in 2003 (Table.
14.2), the human genome project has covered only the euchromatic regions of
the genome, and not covered the important heterochromatic regions. After
multiple updates, the GRC in 2017 released GRCh38, which is the first
coordinate-changing assembly update since 2009; GRCh38 reflects the
resolution of roughly 1000 issues and covers modifications ranging from
thousands of single base changes to megabase-scale range reorganizations,
localization of previously orphaned sequences, and gap closures.
Block 4 Applications of Genomics and Proteomics
200 million bp of sequence containing 1,956 gene predictions, 99 of which are
predicted to be protein-coding (Fig. 14.2). The completed regions include
recent segmental duplications, all centromeric satellite arrays, and the short
arms of all five acrocentric chromosomes, uncovering these complex regions
of the genome and making them available for variational and functional
studies. Twenty years after the initial draft was published, a truly complete
sequence of a human genome reveals what has been missing.
With the vast amount of data about the human genome (generated by the
Human Genome Project and other genomics research consortiums), scientists
and clinicians have more powerful high-end tools, to investigate the role that
multiple genetic factors, which act together with the environmental factors and
play major role in much more complex diseases. The diseases, such as
diabetes, cancer, and cardiovascular disease constitute the majority of health
92 problems in global relevance. Genome-based research is already enabling
Unit 14 Analysis of the Genome
scientists and clinicians to develop improved diagnostics, evidence-based
approaches for improving clinical efficacy, highly effective therapeutic
strategies, and better decision-making technologies for patients and
healthcare providers. Ultimately, it appears inevitable in the future that,
treatments will be tailored to a patient's particular genomic makeup
(personalized medicine/treatment). Thus, the role of genetics in healthcare is
changing profoundly and the human genome sequence has accelerated the
era of genomic medicine.
Block 4 Applications of Genomics and Proteomics
about the genome organization and its’ effect on gene expression. Using a
whole chromosome shotgun sequencing strategy, researchers have
determined the genome sequence of P. falciparum (3D7 clone).
Unit 14 Analysis of the Genome
Block 4 Applications of Genomics and Proteomics
a slight bias in the orientation of the genes with respect to the direction of
replication, as 59% are transcribed with the same polarity as replication,
compared with 75% in B. subtilis. The even distribution of gene polarity
in Mycobacterium tuberculosis may reflect the slower growth and infrequent
replication cycles of this pathogen.
In the circular chromosome map, the outer circle shows the size scale in Mb,
with “0” representing the origin of replication. The first circle from the exterior
represents the positions of stable RNA genes (tRNAs- blue, other RNAs-pink)
and the direct repeat region (pink cube); the second circle inwards represents
the coding sequence by strand (clockwise-dark green; anticlockwise-light
green); the third circle shows repetitive DNA (insertion sequences-orange;
13E12 REP family, dark pink; prophage-blue); the fourth circle depicts the
positions of the PPE family members (green); the fifth circle shows the PE
family members (purple-excluding PGRS). The histogram (centre) shows G +
C content, with <65% G + C in yellow, and >65% G + C in red.
Unit 14 Analysis of the Genome
Their main aims to make specific DNA alterations (insertions, deletions, and
substitutions) are:
• To study the changes in protein activity that occur as a result of the DNA
manipulation (which further changes protein sequence).
Block 4 Applications of Genomics and Proteomics
2. Combine the two independent-primer PCR products from each reaction
in one test tube and denature at 95°C to separate the newly synthesized
DNA from the template DNA (plasmid).
Unit 14 Analysis of the Genome
6$4
6$4
Choose the correct answer from the multiple options given below each
question.
i) 2003
ii) 1990
iii) 1998
iv) 2000
i) T-DNA
14.10 SUMMARY
• Genome analysis of an organism’s whole genome can provide many
predictions about diagnosis, or susceptibilities to conditions, and offers
the possibility of identifying all putative protein-encoding genes of a
given organism. 99
Block 4 Applications of Genomics and Proteomics
• The recent advances in gene sequencing and annotation technology
have allowed high throughput genomic sequencing to be done quickly
and relatively cheaply which propelled the work of genome analysis
forward.
14.12 ANSWERS
Self-Assessment Questions
1. a) i) C) Whole genomes
Unit 14 Analysis of the Genome
c) i) initial version published in 2003
Terminal Questions
1. Refer to Sections 14.1, 14.2.
Acknowledgement of Figures
Fig 14.2: S. Nurk et al., SCIENCE, Vol 376, Issue 6588, pp. 44-53, 2022,
The complete sequence of a human genome, DOI:
10.1126/science.abj6987
101
UNIT 15
0$1,38/$7,212)7+(
*(120(
*(120(
6WUXFWXUH
6WUXFWXUH
15.1 Introduction Frequency of Occurrence of
Restriction Sites in DNA
Objectives
Restriction Sites and Creation
15.2 Converting Genomes into
of Recombinant DNA
Clones, and Clones into
Molecules
Genomes
15.5 Cloning Vectors and DNA
DNA Cloning
Cloning
Reporter Gene
Plasmid Cloning Vectors
15.3 Gene Knockout Method in
Artificial Chromosomes
Transgenics
15.6 Genomic Libraries
Gene Knockouts in Yeast
15.7 Chromosome Libraries
Gene Knockouts in the Mouse
15.8 DNA Sequencing and
Gene Knockin Method in
Analysis of DNA Sequences
Transgenics
15.9 Summary
Gene Knockin in the Mouse
15.10 Terminal Questions
Knockin Mice Mutations
15.11 Answers
15.4 Restriction Enzymes
General Properties of
Restriction Enzymes
15.1 INTRODUCTION
Genomics is the science of obtaining and analyzing the sequences of
complete genomes. At the core of genomics is recombinant DNA technology,
the ability to construct and clone individual fragments of a genome, and to
manipulate the cloned DNA in various ways, including expressing it in a
foreign cell. The development of molecular techniques for analyzing genes
and gene expression has revolutionized experimental biology. The present
Unit describes the techniques used for manipulation of genome.
Unit 15 Manipulation of the Genome
2EMHFWLYHV
After studying this unit you would be able to:
2. Cut the DNA into pieces with a restriction enzyme, (an enzyme that
recognizes and cuts within a specific DNA sequence) and insert (ligate)
each piece individually into a cloning vector, that is, cut with the same
restriction enzyme to construct a recombinant DNA molecule, (a DNA
molecule constructed in vitro containing sequences from two or more
distinct DNA molecules). 103
Block 4 Applications of Genomics and Proteomics
Unit 15 Manipulation of the Genome
15.3.1 Gene Knockouts in Yeast
In yeast, gene function can be deactivated using a polymerase chain reaction
(PCR)-based approach, where PCR primers are designed based on the
known genome sequence to create an artificial linear DNA deletion module,
containing the gene sequence flanked by a selectable marker such as the
kanR (kanamycin) marker for resistance to a specific chemical G418.
Essentially, the kanR marker replaces a major portion of the gene of interest's
coding region, rendering the gene unable to produce its protein. When this
linear DNA is introduced into yeast, colonies resistant to G418 are chosen.
Unlike the previously mentioned plasmids, this linear DNA fragment cannot
replicate in the host cell due to the absence of an origin of replication. Through
a process known as homologous recombination, the linear plasmid integrates
into the yeast chromosome, effectively knocking out the chromosomal copy of
the gene of interest because the selectable marker replaces most of the
coding region.
Block 4 Applications of Genomics and Proteomics
15.3.5 Knockin Mice Mutations
Constitutive knockin mice: This particular model has been designed to
transport a cDNA sequence that encodes a protein to a specific location.
Point mutation knockin mice: This model involves substituting one or a few
DNA bases within the sequences of a designated gene.
Humanized mice: This model includes the replacement of murine gene by its
human counterpart.
More than 400 different restriction enzymes have been isolated, and at least
2,000 more have been characterized partially. They are named for the
organisms from which they are isolated. Conventionally, a three-letter system
106 is used. Commonly the first letter is that of the genus, and the second and
Unit 15 Manipulation of the Genome
third letters are from the species name. The letters are italicized or underlined,
followed by roman numerals that signify a specific restriction enzyme from that
organism. Additional letters sometimes are added just before the number to
signify a particular bacterial strain from which the enzymes were obtained. For
example, EcoRI and EcoRV are both from Escherichia coli strain RY13, but
recognize different restriction sites; HindIII is from Haemophilus influenzae
strain Rd. The Roman numerals indicate the order in which the restriction
enzymes from that strain were identified. Hence, EcoRI and EcoRV are the
first and fifth restriction enzymes identified for E. coli strain RY13. The names
are pronounced in ways that follow no set pattern. For example, BamHI is
“bam-H-one,” BglII is “bagel-two,” EcoRI is “echo-R-one” or “eeko-R-one,”
HindIII is “hin-D-three,” HhaI is “ha-ha-one,” and HpaII is “hepa-two.”
Many restriction sites have an axis of symmetry through the midpoint. Figure
15.2 shows this symmetry for the EcoRI restriction site: the nucleotide
sequence from 5ƍ to 3ƍ on one DNA strand is the same as the nucleotide
sequence from 5ƍ to 3ƍ on the complementary DNA strand. Thus, the
sequences are said to have two fold rotational symmetry. A number of
restriction sites are shown in Table 15.1. The most commonly used restriction
enzymes recognize four nucleotide pairs (for example, HhaI) or six nucleotide
pairs (for example, BamHI, EcoRI). Some enzymes recognize eight-nucleotide
pair sequences (for example, NotI [“not-one”]). Other classes of enzymes do
not fit our model because the restriction site is not symmetrical about the
center. HinfI (“hin-f-one”), for example, recognizes a five-nucleotide pair
sequence in which there is symmetry in the two nucleotide pairs on either side
of the central nucleotide pair, but the central nucleotide pair is obviously
asymmetrical within the sequence. BstXI (“b-s-t-x-one”) is representative of a
number of restriction enzymes with a nonspecific spacer region between
symmetrical sequences. 107
Block 4 Applications of Genomics and Proteomics
Table 15.1: Characteristics of some restriction enzymes
Fig. 15.2: Restriction site in DNA, showing the two fold rotational symmetry of
the sequence. The sequence reads the same from left to right (5ƍ to 3ƍ)
on the top strand (GAATTC, here) as it does from right to left (5ƍ to 3ƍ)
108 on the bottom strand. Shown is the restriction site for EcoRI
Unit 15 Manipulation of the Genome
15.4.2 Frequency of Occurrence of Restriction
Sites in DNA
Since each restriction enzyme cuts DNA at an enzyme-specific sequence, the
number of cuts the enzyme makes in a particular DNA molecule depends on
the number of times that particular restriction site occurs. When you cut a
number of copies of the same genome with a particular restriction enzyme, the
DNA is cleaved at the specific restriction sites by that enzyme, which are
distributed throughout the genome. Although this produces millions of
fragments of different sizes from one genome copy, all copies of the same
genome will be cut at identical places.
Restriction enzymes in the first class cut DNA in different ways. As Table 15.1
indicates, some enzymes, such as SmaI (“sma-one”), cut both strands of DNA
between the same two nucleotide pairs to produce DNA fragments with blunt
ends (Fig. 15.3a). Other enzymes, such as BamHI, make staggered cuts in the
symmetrical nucleotide-pair sequence to produce DNA fragments with sticky
or staggered ends, either 5ƍ overhanging ends, as in the case of cleavage with
BamHI (Fig. 15.3b) or EcoRI, or 3ƍ overhanging ends, as in the case of
cleavage with PstI (“P-S-T-one”; Fig. 15.3c).Restriction enzymes that produce
sticky ends are of particular value in cloning DNA because every DNA
fragment generated by cutting a piece of DNA with the same restriction
enzyme has the same single-stranded nucleotide sequence at the two
overhanging ends. If the ends of two pieces of DNA produced by the action of
the same restriction enzyme (such as EcoRI)—a cloning vector and a
chromosomal DNA fragment, for example, come together in solution, base
pairing occurs between the overhanging ends; the two single-stranded DNA
ends are said to anneal (Fig. 15.4). Using DNA ligase, the two DNAs can be
covalently linked (ligated) to produce a longer DNA molecule with the
restriction sites reconstituted at the junction of the two fragments. Even DNA
fragments with blunt ends can be ligated together by DNA ligase at high
concentrations of the enzyme. The ligation of two DNA fragments is the
principle behind the formation of recombinant DNA molecules. Paul Berg
received part of the 1980 Nobel Prize in Chemistry “for his fundamental
studies of the biochemistry of nucleic acids, with particular regard to
recombinant-DNA.” 109
Block 4 Applications of Genomics and Proteomics
Fig. 15.3: Examples of how restriction enzymes cleave DNA. (a) SmaI results in
blunt ends, (b) BamHI results in 5ƍ overhanging (“sticky”) ends, (c) PstI
results in 3ƍ overhanging (“sticky”) ends.
Fig. 15.4: Cleavage of DNA by the restriction enzyme EcoRI. EcoRI makes
staggered, symmetrical cuts in DNA, leaving “sticky” ends. A DNA
fragment with a sticky end produced by EcoRI digestion can bind by
complementary base pairing (anneal) to any other DNA fragment with
a sticky end produced by EcoRI cleavage. The nicks can then be
110 sealed with DNA ligase
Unit 15 Manipulation of the Genome
6$4
6$4
a) Answer in one word only:
Block 4 Applications of Genomics and Proteomics
genes for the other functions of the plasmid. Plasmid cloning vectors are
derivatives of circular natural plasmids “engineered” to have features useful for
cloning DNA. We focus here on features of E. coli plasmid cloning vectors.
ii) A selectable marker, so that E. coli cells with the plasmid can be
distinguished easily from cells that lack the plasmid. A selectable marker
is a gene that allows us to determine easily if a cell does or does not
contain the cloning vector. For bacterial plasmid cloning vectors, typically
the selectable marker is a gene for resistance to an antibiotic, such as
the ampR gene for ampicillin resistance or the tetR gene for tetracycline
resistance. When plasmids carrying antibiotic-resistance genes are
added to a population of plasmid-free and therefore antibiotic sensitive
E. coli, the cells that take up the plasmid can be selected for by culturing
the cells on a solid medium containing the appropriate antibiotic. Only
bacteria with the plasmid will grow on the medium.
iii) One or more unique restriction enzyme cleavage sites, that is, the sites
present just once in the vector for the insertion of the DNA fragments to
be cloned. Typically, a number of sites are present in the vector, and
these sites tend to be engineered as a multiple cloning site or polylinker.
A multiple cloning site is a region of DNA containing several unique
restriction sites where a fragment of foreign DNA (not originally part of
the vector) can be inserted into the vector. With a number of different
sites available in the multiple cloning site of a vector, an investigator can
use the same vector in different cloning experiments by choosing
different restriction sites for the cloning purpose.
As an example, Fig. 15.5 shows the plasmid cloning vector pBluescript II. This
2,961 bp vector has the following features that make it useful for cloning DNA
in E. coli:
Unit 15 Manipulation of the Genome
functional enzyme present) or pBluescript II with an inserted DNA
fragment (functional enzyme absent).The chemical X-gal, a colorless
artificial substrate for ȕ-galactosidase is included in the medium on
which the cells containing plasmids are plated as an indicator for ȕ-
galactosidase activity in cells of a colony. Cleavage of X-gal by ȕ-
galactosidase leads to the production of a blue dye. Thus, if a functional
enzyme is present (vector with no insert), the colony turns blue, whereas
if nonfunctional ȕ-galactosidase is made (vector with inserted DNA), the
colony is white. This protocol is called blue-white colony screening.
Fig. 15.5: The plasmid cloning vector pBluescript II. This plasmid cloning vector
has an origin of replication (ori), a selectable marker, and a multiple
cloning site located within part of the ȕ-galactosidase gene lacZ+
Figure 15.6 illustrates how a piece of DNA can be inserted into a plasmid
cloning vector such as pBluescript II. In the first step, pBluescript II is cut with
a restriction enzyme that has a site in the multiple cloning site. Next, the piece
of DNA to be cloned is generated by cutting high-molecular-weight DNA with
the same restriction enzyme. Since restriction sites are non uniformly arranged
in DNA, fragments of various sizes are produced. The DNA fragments are
mixed with the cut vector in the presence of DNA ligase; in some cases, the
DNA fragment becomes inserted between the two cut ends of the plasmid and
DNA ligase joins the two molecules covalently. The resulting recombinant
DNA plasmid is introduced into an E. coli host by transformation. This is done
either by incubating the recombinant DNA plasmids with E. coli cells treated
chemically (such as with CaCl2) to take up DNA, or by electroporation, a 113
Block 4 Applications of Genomics and Proteomics
method in which an electric shock is delivered to the cells, causing temporary
disruptions of the cell membrane to let the DNA enter. Transformed cells are
plated onto media containing ampicillin and X-gal. Cells that can grow and
divide on this medium, forming a colony, must have been transformed by a
plasmid. Colonies containing plasmids with an insert can be identified by the
blue–white colony screening method.
Fig. 15.6: Insertion of a piece of DNA into the plasmid cloning vector pBluescript
II to produce a recombinant DNA molecule. The vector pBluescript II
contains several unique restriction enzyme sites localized in a multiple
cloning site that are convenient for constructing recombinant DNA
molecules. The insertion of a DNA fragment into the multiple cloning
site disrupts part of the ȕ-galactosidase (lacZ+) gene, leading to
nonfunctional ȕ-galactosidase in E. coli. The blue-white colony
screening method described in the text can be used to identify vectors
with or without inserts
Unit 15 Manipulation of the Genome
clones would be needed to contain a single genome of a complex multicellular
organism such as a human. To clone larger DNA inserts, different vectors are
used such as cosmids and artificial chromosomes. A cosmid can
accommodate DNA inserts in the range of 40-45 kb for genomics uses. A
cosmid cloning vector is similar to a plasmid cloning vector, with an origin, a
drug resistance marker, and a multiple cloning site, but it is introduced into
host cells differently. Cosmids are frequently used as vectors when libraries
are made because they are able to hold larger inserts.
Yeast artificial chromosomes (YACs) are cloning vectors that enable artificial
chromosomes to be made and replicated in yeast cells. YAC vectors can
accommodate DNA fragments that are several hundred kilobase pairs long,
much longer than the fragments that can be cloned in the plasmid, cosmid, or
BAC vectors. Therefore, YAC vectors have been used to clone very large DNA
fragments (between 0.2 and 2.0 Mb), for example, in creating physical maps of 115
Block 4 Applications of Genomics and Proteomics
large genomes such as the human genome. A YAC (shown in its linear form)
has the following features (Fig. 5.7b):
3. A selectable marker on each arm for detecting and maintaining the YAC
in yeast (for example, TRP1 and URA3 to enable transformed trp1
[tryptophan requiring] ura3 [uracil requiring] mutant yeast to grow on a
medium lacking tryptophan and uracil).
There are two disadvantages associated with these very large YAC-based
clones. First, during the cloning process, a fraction of the YAC vectors accept
two or more inserts, rather than one, creating a chimeric YAC. A second
problem is that portions of the insert DNA are frequently deleted or otherwise
modified by the host cell, or undergo recombination with other DNA in the host
cell. The altered inserts in chimeric and rearranged YACs will confound the
assembly of the genome, because assembly requires that we compare how
different inserts in our library overlap. The alterations in these inserts will
cause us to misinterpret how they overlap with other clones, because a
chimeric clone might contain, for instance, DNA from chromosome 5 ligated to
DNA from chromosome 18. Determining which YACs are modified is often a
very slow and labor-intensive process, making the assembly of a genome
sequence more difficult.
Empty YAC vectors, ones that have yet to contain a DNA insert are
propagated in E. coli as circular plasmids; in this form the two telomeres are
end-to-end. This propagation step makes use of the bacterial origin of
replication and the bacterial selectable marker. Bacterial and eukaryotic
origins of replication are not functionally similar, which means that the yeast
ARS sequence will not work in a bacterial cell, just as the bacterial ori
sequence will not function in a yeast cell. In addition, bacterial and eukaryotic
promoters are different, meaning that the bacterial RNA polymerase cannot
transcribe the yeast TRP1 and URA3 genes, so those selectable markers will
function only in yeast, not in bacteria. Likewise, yeast RNA polymerase II is
unable to transcribe the ampR gene.
For cloning experiments, a circular YAC is cut with one restriction enzyme that
cuts in the multiple cloning site and with another restriction enzyme that cuts
116 between the two TELs. In this way, the left and right arms are produced. High-
Unit 15 Manipulation of the Genome
molecular-weight DNA, cut with the same restriction enzyme used to cut the
YAC multiple cloning site, is ligated to the two arms and the recombinant
molecules are transformed into yeast. By selecting for both TRP1 and URA3, it
can be ensured that the transformants have both the left and right arms.
Block 4 Applications of Genomics and Proteomics
genetic analysis. A genomic library can also be used to isolate and study a
particular clone, such as that for a gene of interest. In this section, we will
focus on the construction of genomic libraries of eukaryotic DNA.
Genomic libraries are made using the basic cloning procedures already
described. A restriction enzyme is used to cut the genomic DNA, and a vector
is chosen so that the entire genome is represented in a manageable number
of clones. You might assume that it is as simple as digesting the genomic DNA
completely with a restriction enzyme and cloning the resulting DNA fragments
in a cloning vector. This will create a genomic library, but this library will have
serious functional limitations for four important reasons:
3. The number of base pairs between adjacent restriction sites can vary
significantly; so, for instance, cutting a 10 kb fragment of DNA with
BamHI might yield fragments of 500, 2,500, and 7,000 base pairs. When
genomic DNA is digested, the resultant fragments will fall in a range of
sizes. Some of these fragments will be too large to clone. As a result,
part of the genome would be unclonable in this type of library.
To deal with these functional limitations, you need to break the genomic DNA
differently. Specifically, you need to break the genomic DNA into fragments
that are of the correct size for your cloning vector and that overlap each other.
To generate these overlapping fragments, you can either mechanically break
(shear) the genomic DNA, or you can use a restriction enzyme under
conditions such that the genomic DNA is digested partially.
Unit 15 Manipulation of the Genome
physical means and not by cutting with restriction enzymes, additional
enzymatic manipulations are necessary to add appropriate ends to the
molecules for their insertion into a restriction site of a cloning vector.
The separated DNA fragments are invisible to the eye. They are made visible
by adding either ethidium bromide or SYBR® Green to stain the DNA. Both
chemicals bind tightly to DNA and emit visible light when excited with the
correct wavelength of light. Ethidium bromide emits visible light after being
excited with ultraviolet light, and SYBR® Green, when bound to DNA, emits
green light after being excited with blue light. The emission of visible light
makes the position of the DNA in the gel obvious. Since the wells are
rectangular, the DNA fragments form “bands” on the gel. In Figure 15.9b, an
agarose gel electrophoresis analysis shows partial digestion of genomic DNA.
The vertical “lanes” of the gel show how the DNA fragments in the samples
loaded into the wells at the top separated during the electrophoresis. Lane 1
contains the DNA ladder, in this case the lambda ladder. Note the discrete set
of bands of known sizes in the lane. Lane 2 shows a sample of genomic DNA
not treated with a restriction enzyme. There is not a highly discrete band, but a
concentrated mass of DNA in a region of the lane corresponding to the large
DNA fragments of the lambda ladder, and a smear of DNA going down the 119
Block 4 Applications of Genomics and Proteomics
lane from that point. The mass of DNA is the large DNA fragments of genomic
DNA that came out of the cell. It is unavoidable to break the genomic DNA
mechanically during isolation, so the size of the large DNA is much smaller
than the sizes of chromosomes. The mechanical shearing during isolation is
also responsible for the many bands of various sizes of DNA fragments that
are seen as a smear down the lane. Lane 3 shows genomic DNA digested
completely with a restriction enzyme. There are no discrete bands of DNA
fragments here either. Instead, a smear of fragments is seen, most of which
are smaller than the smallest visible lambda ladder fragment at 2.0 kb. Lanes
4 and 5 show the results of digesting the genomic DNA partially using the
same restriction enzyme. In both cases the DNA is of much larger size than
that seen in the complete digest lane, this being the expected outcome of
partial digestion. The partial digestion conditions were different for the samples
loaded in the two lanes, with more digestion carried out for the DNA in lane 4
than for the DNA in lane 5. The difference in partial digestion conditions is
reflected in the range of DNA fragment sizes on the gel; that is, larger DNA
fragments are seen in lane 5 than in lane 4. As for the complete digestion of
genomic DNA, partial digestion does not result in discrete bands when the
digested DNA is analyzed by agarose gel electrophoresis. Rather, there is a
smear of DNA fragments of different sizes. Since there is a DNA ladder in the
gel showing where DNA fragments of particular sizes migrated, researchers
can use that information and isolate DNA fragments of the desired size for
cloning from the partial digest lanes. The isolation is done simply by cutting out
a block of agarose containing the DNA fragments of the desired size and then
extracting the DNA from the gel piece.
How many clones are needed to contain all sequences in the genome? The
number of clones needed to include all sequences in the genome depends on
the size of the genome being cloned and the average size of the DNA
fragments inserted into the vector. The probability of having at least one copy
of any DNA sequence in the genomic library can be calculated from the
following formula:
Unit 15 Manipulation of the Genome
Block 4 Applications of Genomics and Proteomics
Unit 15 Manipulation of the Genome
Dideoxy sequencing
Block 4 Applications of Genomics and Proteomics
Fig. 15.10: Primers for DNA sequencing. (a) In a DNA sequencing reaction,
double-stranded DNA is denatured to single strands, and the
sequencing primer anneals to a specific region of one of the two
strands. Extension of the primer by DNA polymerase produces new
DNA that is complementary to DNA to which the primer annealed; this
is the sequencing reaction. The other DNA strand plays no role in the
sequencing reaction. (b) Most commonly used vectors allow the use
of universal sequencing primers. For pBluescript II, the T7 universal
sequencing primer anneals near the KpnI site of the multiple cloning
site, and the SP6 universal sequencing primer anneals near the SacI
site at the other end of the multiple cloning site. The binding sites for
the primers are positioned so that, when a sequencing primer
anneals, extension of the primer by DNA polymerase produces a DNA
strand complementary to that of the DNA insert
Unit 15 Manipulation of the Genome
precursors (dNTPs, that is dATP, dTTP,dCTP, and dGTP; Fig. 15.11a), and a
small amount of modified nucleotide precursors called dideoxynucleotides
(ddNTPs, that is ddATP, ddTTP, ddCTP, and ddGTP; Fig. 15.11b) are then
added. A dideoxynucleotide differs from a normal deoxynucleotide in that it
has a 3ƍ-H rather than a 3ƍ-OH on the deoxyribose sugar. Furthermore,
different fluorescent dye molecules are linked covalently to each of the four
dideoxynucleotides. These dyes absorb certain wavelengths of light, causing
them to emit very specific wavelengths of light. For instance, the ddGTP
appears blue-green because a dye is bound to it that emits light with a
wavelength of 520 nm (blue-green), while the ddATP appears green, the
ddCTP appears a different shade of green, and the ddTTP appears greenish
yellow. Generally, the dideoxynucleotide (ddNTP) precursors are present in
the reaction mixture at about one by one-hundredth the amount of the normal
deoxynucleotide (dNTP) precursors so that some DNA synthesis occurs in the
dideoxy sequencing reactions. When the dideoxysequencing reaction starts,
DNA polymerase adds a nucleotide to the 3ƍ-OH at the end of the primer. In
the example shown in Figure 15.12a, the template has an A nucleotide, so the
primer is extended by a T nucleotide. Since most of the DNA precursors in the
reaction are dNTPs, the probability is great that a dTTP will be used for this
extension step. However, there is a small chance that DNA polymerase will
use the ddTTP precursor for this extension step. If the normal dTTP precursor
is used, the extended DNA chain has a 3ƍ-OH at its end and, therefore,
another nucleotide can be added by DNA polymerase. However, if the dideoxy
ddTTP precursor is used, the extended DNA chain has a 3ƍ-H at its end and,
therefore, another nucleotide cannot be added by DNA polymerase. In other
words, the addition of a dideoxy nucleotide to a DNA chain being synthesized
terminates the DNA synthesis reaction. Therefore, in the example in Figure
15.12a, the addition of the normal T nucleotide leads to the next extension
step, during which again there is a choice of nucleotide precursor types, in this
case between dATP and ddATP.
Block 4 Applications of Genomics and Proteomics
The DNA chains in each reaction mixture are separated by a special, very
sensitive type of electrophoresis in a very small capillary, and a laser eye at
the end of the capillary detects the colored fragments as they exit the capillary.
While the dyes emit similar colors, the computer converts the minor color
differences into a far more obvious difference by assigning “false colors” to
each dye, such as using green for A, black for G, red for T, and blue for C. The
output is a series of colored peaks corresponding to each nucleotide position
in the sequence (Fig. 15.12c). The graphic representation is converted to a
sequence of nucleotides by a computer with the oversight of the researcher.
Automated sequencing is of great utility to research teams in determining the
complete sequences of various genomes because a single machine can
analyze 100 or more samples per day.
To sequence more nucleotides than can be read for a single reaction, the first
sequence obtained is used to design a custom primer that will anneal to the
DNA insert near the 3ƍ end of that sequence. The sequencing reaction using
the new primer generates a DNA sequence that partially overlaps the first
sequence. In this way, a researcher can step down a long DNA insert and
126 obtain its complete sequence.
Unit 15 Manipulation of the Genome
Block 4 Applications of Genomics and Proteomics
Pyrosequencing
Unit 15 Manipulation of the Genome
carried out. Thus, the sequencing of many DNA templates is done
simultaneously, making it possible to obtain about 20 million nucleotides of
genome sequence in about 6 hours. The pyrosequencing technique is still
quite new and expensive, but it should become an important technique as the
equipment becomes refined and more affordable.
6$4
6$4
a). Answer in one word only:
ii) How the separated DNA fragments in the agarose gel are made
visible?
iv) Name the cloning vector that can carry the maximum size of DNA
insert.
Block 4 Applications of Genomics and Proteomics
b). Fill in the blanks with appropriate words:
ii) The selectable marker for BAC is camR for ……………. resistance.
Column A Column B
15.9 SUMMARY
• Genomics is the study of an organism's complete DNA sequence. The
process begins with cloning the organism's DNA into various types of
vectors. The exact nucleotide sequence of these clones is then
determined. These sequence data can be utilized in numerous analyses,
such as identifying regions that encode genes.
Unit 15 Manipulation of the Genome
• Reporter genes are a type of protein-coding gene that are often tagged
to a gene of interest. A reporter gene is an exogenous coding region that
is joined to an expression vector containing a promoter sequence or
element in order to enable the measurement of promoter activity in cells.
Block 4 Applications of Genomics and Proteomics
• Once a genomic library is completed, the DNA within that library can be
sequenced. One popular method of DNA sequencing involves the use of
dideoxynucleotides (ddNTPs) to terminate chain extension in a modified
version of DNA replication. These terminated fragments are detectable
because the individual ddNTPs are linked to colored dyes. The dyes
allow the fragments to be visualized and provide information on which
ddNTP terminated the fragment.
15.11 ANSWERS
Self-Assessment Questions
1. a) i) Escherichia coli strain RY13, ii) Blunt ends, iii) Gene
knockout, iv) knock-in animal
Unit 15 Manipulation of the Genome
®
2. a) i) Dideoxy sequencing, ii) Ethidium bromide or SYBR Green,
iii) up to 300 kb, iv) Yeast artificial chromosomes,
v) blue dye
Terminal Questions
1. Refer to Section 15.2.
133
UNIT 16
(;35(66,21$1$/<6,62)
*(120(
*(120(
6WUXFWXUH
6WUXFWXUH
16.1 Introduction 16.5 RNA Interference
Components of Gene
Silencing
16.1 INTRODUCTION
You learned about the analysis and modification of the genome in Unit 15. In
the present Unit, you will learn about the principles of genomic expression
analysis. The simplest way to define gene expression analysis is the study of
how genes are transcribed to produce functional gene products, such as
functional RNA species or proteins. Understanding gene control helps
distinguish between abnormal or unhealthy cellular processes and normal
ones such as differentiation.
Unit 16 Expression Analysis of Genome
2EMHFWLYHV
The gene of interest and the reporter gene are cloned in a DNA construct
before being transferred into the cell or organism. This construct typically
takes the shape of a plasmid, a circular DNA molecule that is present in
prokaryotic or bacterial cells in a culture. The expression of the reporter gene
is considered as a signal for successful uptake of the gene of interest by the
cell or organism and hence its study is essential.
One of the most commonly utilised reporter genes in bacteria is the lacZ gene
of E. coli, which codes for the protein beta-galactosidase. The enzyme
expressed by this gene gives bacteria a blue colour when they are grown on
medium containing the substrate analogue X-gal. Another example of a
bacterium-selectable marker that also serves as a reporter is the
chloramphenicol acetyltransferase (CAT) gene, which confers resistance to
the antibiotic chloramphenicol. 135
Block 4 Applications of Genomics and Proteomics
Reporter genes can result in the production of a protein that has minimal direct
effects on the organism or cell culture. Reporter genes can be activated, that
is, they can be expressed constitutively by cloning the gene of interest
together with the reporter gene where both the genes are regulated by the
same promoter. This results in the transcription of a messenger RNA that code
for two protein-coding sequences (fusion protein). Despite being fused, it is
crucial that both proteins are able to correctly fold into their active
conformations and engage in interactions with their substrates. In order to
ensure that the reporter and gene product do not restrict each others function,
a piece of DNA encoding a flexible polypeptide linker region is typically
inserted while creating the DNA construct. Additionally, reporter genes may be
induced to express throughout growth. In these situations, the reporter gene is
expressed by the use of trans-acting elements, such as transcription factors.
Unit 16 Expression Analysis of Genome
16.2.3 Promoter Assays
6$4
6$4
Fill in the blanks:
Block 4 Applications of Genomics and Proteomics
16.3.1 Temporal Analysis
For example, the wingless gene, a member of the wnt gene family, is
expressed in alternating stripes separated by three cells in the fruit fly
Drosophila melanogaster during the early stages of embryonic development.
This pattern disappears by the time the embryo becomes a larva, but the
wingless gene is still present in certain tissues, such as the imaginal discs of
the wings, which are patches of tissue that eventually develop into the adult
wings. The spatiotemporal pattern of wingless gene expression is determined
by a network of regulatory interactions, which includes the effects of multiple
unique genes such as even-skipped and Krüppel.
Unit 16 Expression Analysis of Genome
specific gene expression, or genome analysis approaches for single cells have
been described. More information on the corresponding functions is now
available than was previously possible, due to the combination of tissue
microdissections and high throughput examination of single cells. To fully use
analytical technologies, a technique for site-specific cell or microdissection
collection from a tissue should be established.
6$4
6$4
Fill in the blanks:
a) The action that converts a gene's informational content into the creation
of a functional product, often a protein, is referred to as the ……………..
process.
The first PTGS/RNAi studies were made in plants, but afterward, nearly all
eukaryotic species, including parasites, insects, protozoa, nematodes, flies,
mouse, and human cell lines, were shown to be experiencing RNAi-related
events. Quelling in fungi, co-suppression or PTGS in plants, and RNAi in the
animal kingdom are the phenotypically diverse but mechanistically related
forms of RNA interference. Recently, it has been discovered that other aspects
of eukaryotic cells' naturally occurring RNAi processes, such as microRNA
production and heterochromatinization, also exist.
Block 4 Applications of Genomics and Proteomics
manage the maturation of multiple species by processing a significant number
of non-coding RNAs, or microRNAs. MicroRNA biosynthesis and function
share common characteristics with RNAi processes. Because of its
exceptional effectiveness and specificity, RNAi is being viewed as a crucial
tool for gene-specific therapeutic activities that affect the mRNAs of disease-
related genes, as well as for functional genomics.
RNA silencing in plants came to light accidentally while looking for transgenic
petunia blooms that were supposed to be more purple. The goal of R.
Jorgensen's lab in 1990 was to increase the activity of the chalcone synthase
(chsA) gene, an enzyme essential for the synthesis of anthocyanin.
Unexpectedly, few transgenic petunia plants carrying the chsA coding area
controlled by a 35S promoter lost both transgene and endogenous chalcone
synthase activity, leading to the development of white or variegated sectors in
many of the flowers. Run-on transcription assays in isolated nuclei showed
that the decrease in cytosolic chsA mRNA was not associated with decreased
transcription. The term "co-suppression" was created by Jorgensen to
characterize the loss of mRNAs from both the transgene and the endogene.
Both sense and antisense transgenes have the potential to cause PTGS, and
biochemical data show that comparable processes may be at play in both
situations. It is important to note that the co-suppression phenomena have
been proven in metazoans and mammals in addition to plants.
Owing to the discovery of Fire et al., who firmly showed the biological nature of
inducers in gene silencing by administering pure dsRNA directly into the body
of Caenorhabditis worms, the phenomena of RNAi initially gained attention.
Unit 16 Expression Analysis of Genome
have independently led to the discovery of a universal paradigm for gene
control. The dsRNA serves as the inducer and the target RNA is destroyed in
a homology-dependent manner. Additionally, the degradative machinery
needs a set of proteins that are common to most species in terms of both
structure and function. siRNA production and systemic transmission of
silencing from its site of initiation are two characteristics that are present in the
majority of these activities.
Dicer
Members of the RNase III family are one of the few nucleases which
specifically target dsRNAs and cleave them with 3ƍ-hydroxyl and 5ƍ-phosphate
termini and 3ƍ overhangs of 2 to 3 nucleotides. This enzyme was given the
name Dicer because it can convert dsRNA into evenly sized short RNAs
(siRNA). The nucleases in question have been preserved throughout evolution
in flies, worms, fungi, animals, and plants. Dicer has four diverse domains,
including an amino-terminal helicase, a dsRNA binding domain, two RNase III
motifs, and a PAZ domain (a 110-amino-acid domain found in proteins like
Argo, Piwi, and Pinhead/Zwille). It also shares this domain with the
QDE2/RDE1/Argonaute family of proteins, which has been genetically linked
to RNAi by separate studies. Dicer's tandem RNase III domains are suggested
to catalyze cleavage.
Block 4 Applications of Genomics and Proteomics
RNA-Induced Silencing (RISC) Complex and the Guide RNAs
The inability of cellular isolates exposed to a Ca2+ dependent nuclease
(micrococcal nuclease, that can digest both RNA and DNA) to degrade the
homologous mRNAs and the lack of this impact with DNase I treatment served
as evidence that RNA was a critical element of the nuclease activity. The
RNA-induced silencing complex (RISC) was named after the sequence-
specific nuclease activity that was seen in the cellular extracts and was
responsible for abasing target mRNAs.
6$4
6$4
Fill in the blanks:
Unit 16 Expression Analysis of Genome
Numerous eukaryotic and animal cells naturally include the RNAi pathway. It is
started by the enzyme Dicer, which breaks lengthy dsRNA molecules into
short double-stranded pieces called small interfering RNAs (siRNAs), which
have 21 to 23 nucleotides. The sense, that is, passenger strand and the
antisense, that is, guide strand of each siRNA are unwound to form two single-
stranded RNAs (ssRNAs), respectively. The protein Argonaute 2 then cleaves
the passenger strand (Ago2). The guide strand is integrated into the RISC
while the passenger strand is destroyed. The target mRNA is then bound and
degraded by the RISC assembly. The guide strand interacts with a
complementary strand in an mRNA molecule, which activates Ago2, a catalytic
subunit of the RISC, to initiate cleavage.
• To control the target sequence, the Argonaute protein either cleaves the
mRNA or enlists the help of other agents.
Block 4 Applications of Genomics and Proteomics
By making the ribonuclease Dicer more active, which binds to and cleaves
human short hairpin RNAs (shRNAs) or exogenous dsRNAs into double-
stranded fragments with 20 –25 base pairs and a 2-nucleotide overhang at the
3ƍ end, exogenous dsRNA triggers RNAi. The RISC-Loading Complex
subsequently divides these siRNAs into single strands and incorporates them
into an active RISC (RLC). Dicer-2 and R2D2 are parts of RLC, which is
necessary to connect RISC and Ago2. TATA-binding protein-associated factor
11 (TAF11) induces Dcr-2-R2D2 tetramerization, which increases the binding
affinity to siRNA and facilitates RLC formation tenfold. The R2-D2-Initiator
(RDI) complex would change into the RLC by association with TAF11. R2D2
possesses tandem double-stranded RNA-binding domains that enable it to
recognise the thermodynamically stable end of siRNA duplexes, while Dicer-2
recognises the opposite, less stable extremity. Asymmetric loading is brought
about by Ago2's MID domain, which locates the thermodynamically stable end
of the siRNA. The "passenger" (sense) strand, whose 5ƍ end is abandoned by
MID, is emitted, and the "guide" (antisense) strand is preserved and
collaborates with AGO to build the RISC. After joining the RISC, siRNAs base-
pair to their target mRNA and cleave it to prevent it from being used as a
translation template. In contrast to siRNA, a miRNA-loaded RISC looks for
possible complementarity among cytoplasmic mRNAs. miRNAs bind to
mRNAs in the 3ƍ untranslated region (UTR) where they normally show weak
complementarity, obstructing ribosome access and preventing translation.
Exogenous dsRNA is recognised and bound by an effector protein called
R2D2 in Drosophila and RDE-4 in C. elegans, which increases dicer activity. It
is uncertain what mechanism results in this length selectivity in this protein,
which only binds to long dsRNAs.
siRNA
miRNA
Unit 16 Expression Analysis of Genome
mechanisms caused by the external dsRNA and the endogenously produced
gene silencing effects of miRNAs. Although mature miRNAs share structural
similarities with siRNAs made from external dsRNA, miRNAs must first go
through a significant amount of post-transcriptional modification. A miRNA is
produced in the cell nucleus from a much longer RNA-coding gene as a
primary transcript called a pri-miRNA, which is then processed by the
microprocessor complex into a 70-nucleotide stem-loop structure called a pre-
miRNA. This complex is made up of the dsRNA-binding protein DGCR8 and
the RNase III enzyme Drosha. Since Dicer binds to and cleaves the dsRNA
part of this pre-miRNA to create the mature miRNA molecule which can be
incorporated into the RISC, siRNA and miRNA, all have similar downstream
biological mechanisms. Epstein-Barr virus (EBV) was the first human virus
shown to express miRNAs. Since then, several microRNAs in viruses have
been identified.
It is unclear how the active RISC finds mRNAs that are compatible in a cell.
Translation of the mRNA target is not necessary for RNAi-mediated
degradation, despite the fact that it has been suggested that the cleavage
process is tied to translation. P-bodies, also known as GW bodies or
cytoplasmic bodies, are areas of the cytoplasm where argonaute proteins are
localized and miRNA activity is likewise concentrated. P-bodies have high
rates of mRNA degradation. P-bodies are thought to be a crucial location in
the RNAi process since their disruption reduces RNAi's effectiveness. 145
Block 4 Applications of Genomics and Proteomics
16.5.3 Transcriptional Silencing
Many eukaryotes employ RNAi pathway components to maintain the structure
and organization of their genomes. Pre-transcriptional downregulation of
genes is achieved by modification of histones and the ensuing induction of
heterochromatin formation; this procedure is known as RNA-induced
transcriptional silencing (RITS), and it is carried out by a protein complex
known as the RITS complex. It is unclear how the RITS complex influence the
development and organization of heterochromatin. To sustain the existing
heterochromatin regions, RITS assembles a complex of siRNAs that are
homologous to the localized genes and firmly attach to the methylated
histones. This complex also acts co-transcriptionally to damage any nascent
pre-mRNA transcripts that are started by RNA polymerase.
Dicer is necessary to produce the first complement of siRNAs that target future
transcripts; therefore, it makes sense that while its maintenance is not dicer-
dependent, the development of such a heterochromatin area is. It has been
proposed that heterochromatin maintenance works as a self-reinforcing
feedback loop, with new siRNAs being produced from sporadic nascent
transcripts by RdRP and incorporated into regional RITS complexes.
Unit 16 Expression Analysis of Genome
or almost perfectly complementary with their target genes. In animals, miRNAs
typically have more divergent sequences and suppress translation. It is
possible to prevent translation initiation factors from interacting with the
polyadenine tail of the mRNA in order to produce this translational effect.
RNAi can cause an antiviral response in some animals, despite the fact that
plants typically express more variations of the dicer enzyme than mammals.
RNAi is crucial for antiviral innate immunity in juvenile and adult Drosophila
and is effective against viruses like the Drosophila X virus. Worms that
overexpress components of the RNAi process are immune to viral infection
and produce higher levels of argonaute proteins in response to viruses.
Downregulation of genes
Block 4 Applications of Genomics and Proteomics
F-box proteins, to regulate whole gene networks during development. miRNAs
are connected to the development of cancers and the disruption of the cell
cycle in many species, including humans. Here, miRNAs can act as both
tumor suppressors and oncogenes.
Evolution
Parsimony-based phylogenetic analysis suggests that an early RNAi pathway
was likely already present in the most recent common ancestor of all
eukaryotes; the absence of the process in some eukaryotes is regarded to be
a derived trait. This primitive RNAi system contained at least one dicer-like
protein, one argonaute, one PIWI protein, and one RNA-dependent RNA
polymerase, possibly with other biological functions. These elements were
most likely present in the eukaryotic crown group and may have closer
functional connections with RNA degradation mechanisms like the exosome.
This is supported by a large-scale comparative genomics investigation.
Additionally, this research reveals that the RNA-binding argonaute protein
family is homologous to and originally descended from components of the
translation initiation machinery and is found in eukaryotes, the majority of
archaea, and some bacteria (including Aquifex aeolicus).
Medications
The strategy of using RNAi treatments to shut down genes has proven to be
effective, as demonstrated by randomised controlled clinical trials. The
treatments in this class, which are expanding, work by reducing the expression
of the proteins that particular genes are able to encode using siRNA. To date,
regulatory agencies in the US and Europe have authorized four RNAi drugs:
patisiran (2018), givosiran (2019), lumasiran (2020), and inclisiran (2020 in
Europe with anticipated US approval in 2021).
While all of the RNAi therapeutics that have been currently approved by
regulatory bodies targeting liver-related illnesses; other drugs that are still in
the research phase focus on a variety of other conditions, such as cystic
fibrosis, cardiovascular problems, carcinoma, bleeding issues, gout, alcohol
use disorders, and eye problems.
Delivery mechanisms
For RNA interference to achieve its therapeutic potential, siRNA must be
148 efficiently delivered to the cells of the target tissues. However, a number of
Unit 16 Expression Analysis of Genome
obstacles need to be removed before it may be applied therapeutically. For
instance, "naked" siRNA is prone to a number of challenges that lower its
therapeutic efficiency. Additionally, bare RNA can activate the innate immune
system and be destroyed by serum nucleases once siRNA has reached the
circulation. Unmodified siRNA molecules cannot easily cross the cell
membrane because of their size and extremely polyanionic (carrying negative
charges at several locations) nature. Therefore, siRNA needs to be
synthetic or enclosed in nanoparticles. If therapeutic dosages are not adjusted,
siRNA transport across the cell membrane may result in unexpected toxicities,
and siRNAs may have off-target effects (e.g., unexpected suppression of
genes with partial sequence complementarity). Since their effects are
diminished with each cell division, frequent treatment is necessary even after
they have entered the cells. Lipid nanoparticles and conjugates are two
strategies that aid in siRNA distribution to target cells in response to these
possible problems and hurdles.
Lipid nanoparticles
The core of lipid nanoparticles (LNPs) is modelled following liposomes, which
are lipid shell-encased aqueous cores. Large unilamellar vesicles (LUVs),
which can be 100 nm in size, are the resting place for a subset of liposomal
structures utilized to carry medications to the target tissues. Plasmids,
CRISPR, and mRNA are examples of LNP delivery systems that may be used
to encase nucleic acids.
Conjugates
Targeted delivery for RNAi therapies using siRNA conjugates is an alternative
to LNPs (e.g., aptamers, carbohydrates, GalNAc, peptides, antibodies). In
addition to other cardiometabolic disorders including hypertension and non-
alcoholic steatohepatitis, therapeutics utilizing siRNA conjugates that are
developed for uncommon or inherited diseases such as hemophilia, acute
hepatic porphyria (AHP), hereditary ATTR amyloidosis (NASH), and primary
hyperoxaluria (PH).
Biotechnology
There have been numerous documented other applications for RNAi, such as
the manufacturing of insecticides, crops, and food. The RNAi pathway has
produced a wide range of products, including nutrient-fortified plants, arctic
apples, decaffeinated coffee, nicotine-free tobacco, and hypoallergenic crops.
A variety of new products could be made with the help of RNAi based
technology.
Viral infection
The creation of two unique antiviral therapies was one of the first uses of RNA
interference in medicine. The first type targets viral RNAs. Targeting viral
RNAs has been shown in multiple studies to reduce the replication of multiple
viruses, including adenovirus, hepatitis A, HIV, HPV, hepatitis B, SARS
coronavirus respiratory syncytial virus (RSV), SARS-CoV, influenza virus, and
measles virus. Targeting the host cell's genes is the second strategy used to
stop early viral invasions. For example, blocking the chemokine receptors
(CXCR4 and CCR5) can stop HIV entry. 149
Block 4 Applications of Genomics and Proteomics
Cancer
Conventional chemotherapy can kill cancer cells with effectiveness, but since
it lacks the ability to differentiate between normal and malignant cells, it
frequently results in serious side effects. Various studies have shown that
RNAi can give a more targeted method of preventing tumor growth by
targeting genes relevant to cancer (i.e., oncogene). Additionally, it has been
proposed that RNA interference (RNAi) may increase the susceptibility of
cancer cells to chemotherapeutic agents, providing a complementary
therapeutic approach to chemotherapy. Inhibiting cell invasion and migration is
yet another possible RNAi-based therapy. RNA interference therapies treat
cancer by suppressing particular genes that promote malignancy. By
complementing the cancer genes with RNA interference (RNAi), for example,
by keeping the mRNA sequences consistent with the RNAi drug, this is
achieved. RNA interference (RNAi) sequences should ideally be chemically
altered to enhance their ability to bind to cancer cells. RNAi uptake is
regulated and monitored by the kidneys.
Neurological diseases
Transgenic plants
Transgenic crops express dsRNA that has been carefully chosen to silence
important genes in insect targets. These dsRNAs are exclusively meant to
affect insects that express specific gene sequences. As a proof-of-concept, a
2009 study showed that the dsRNAs could kill any one of four species of fruit
flies while harming none of the others.
Insecticides
Unit 16 Expression Analysis of Genome
Food
RNAi has been used to genetically modify plants such that they produce less
natural plant toxins. These methods make use of the RNAi phenotype that is
persistent and heritable in plant populations. Cotton seeds are a good source
of dietary protein; however, they should not be consumed by humans since
they naturally contain the hazardous terpenoid gossypol. A critical enzyme in
the production of gossypol, delta-cadinene synthase, has been reduced in
cotton stocks using RNA interference (RNAi), without impacting the production
of the enzyme in other parts of the plant, where gossypol is crucial for
protecting plant against pest damage. The amounts of allergens in tomatoes
have been successfully reduced through development efforts, and plants have
been fortified with nutrient-rich antioxidants.
The innate immune system, which may be further separated into acute
inflammatory responses and antiviral responses, is in charge of controlling
siRNA. Small signaling molecules known as cytokines provide messages that
trigger the inflammatory response. Tumor necrosis factor (TNF-), interleukin-6
(IL-6), interleukin-1 (IL-1), and interleukin-12 (IL-12) are a few of them.
Inflammation and antiviral responses produced by the innate immune system
result in the release of pattern recognition receptors (PRRs). These receptors
aid in classifying infections as bacterial, fungal, or viral. More PRRs should be
included in siRNA and the innate immune system in order to assist it in
identifying various RNA structures. In the event of an infection, the siRNA is
therefore more likely to trigger an immunostimulant response.
6$4
6$4
Fill in the blanks:
d) The ………..……… was the first human virus shown to express miRNAs.
16.7 SUMMARY
• A test gene with quantifiable expression is known as a reporter gene. It
can be present on plasmids that have their T-DNA integrated into the
genome of a cell. Examining how those genes are expressed after a cell
has undergone a transformation is important. 151
Block 4 Applications of Genomics and Proteomics
• Reporter genes are those sequences that may be examined to ascertain
how altered genes are expressed. It is possible to do a reporter gene
test by calculating the total amount of protein synthesised. They often
have luminous properties and provide visual clues for precise estimates.
Examples include, Green fluorescent proteins, luciferase, octopine
synthase, etc.
• Both in cell culture and in live animals, the selective and powerful impact
of RNAi on gene expression makes it an invaluable research tool.
Synthetic dsRNA put into cells can cause the suppression of the target
genes of interest.
3. Briefly explain about the type of RNAs that play role in RNA
152 interference?
Unit 16 Expression Analysis of Genome
4. Explain the concept of temporal and site-specific gene expression and
their analysis.
16.9 ANSWERS
Self-Assessment Questions
1. a) selectable markers, b) beta-galactosidase, c) chloramphenicol
acetyltransferase (CAT)
Terminal Questions
1. Refer to Section 16.2 and Subsection 16.2.1.
153
UNIT 17
3527(20($1$/<6,6$1'
$33/,&$7,212)3527(20,&6
$33/,&$7,212)3527(20,&6
6WUXFWXUH
6WUXFWXUH
17.1 Introduction Stable Isotope Labeling by
Amino Acids in Cell Culture
Objectives
(SILAC)
17.2 Origin of Proteomics
17.5 Application of Proteomics
17.3 Classes of Proteomics
Pharmaceutical Field
Profiling Proteomics
Drug Discovery
Functional Proteomics
Drug Development and
Chemoproteomics Toxicology
Phosphoproteomics Phage Antibody as Tool
17.4 Techniques of Proteomics 17.6 Summary
Two Dimensional Gel 17.7 Terminal Questions
Electrophoresis (2DE)
17.8 Answers
Two Dimensional Difference
Gel Electrophoresis (2D-DIGE)
17.1 INTRODUCTION
Proteins are the molecules that play various roles in the biological system and
proteomics is the study of complete set of proteins at a time. With the advent
of technologies in the early seventies, genome sequencing was gaining
attention among scientists. The gene sequence cannot provide information
related to protein function, localization, post-translational modifications,
relative expression in different cell organelles, protein-protein interaction, etc.
The human genome has approximately 31,000 protein-encoding genes but the
protein products generated are estimated to be close to 1 million. This
indicates that the functional information in genes is actually located in the
proteome. Therefore, understanding of ‘proteome’ means the complete set of
proteins within the cell is important. Hence, the term ‘proteomics’ which was
Unit 17 Proteome Analysis and Application of Proteomics
first coined in 1995 refers to the study and characterization of complete set of
proteins in a cell, tissue or organism. This study plays an important role in
biomarker identification, drug discovery, disease pathogenesis, identification of
drug targets for various diseases, and so on.
2EMHFWLYHV
2EMHFWLYHV
After studying this Unit you would be able to:
describe proteomics,
Block 4 Applications of Genomics and Proteomics
crucial biological functions. Studying protein-protein interaction by a two-hybrid
approach is common in functional proteomics. For example, X protein is
interacting with A/B/C/D proteins within the cell. We need to identify which
protein is A/B/C/D. For this immunoprecipitation technique can be used. An
expression construct is generated with gene X along with a tag such as FLAG,
GFP, or c-myc, etc. When it is transfected into the cell (bacteria/fungi/yeast
/mammalian) it expresses fusion protein. This is used as bait or ligand to catch
the prey which is A/B/C/D proteins in this case. Following fusion protein
expression in the host, this protein interacts with A/B/C/D proteins; in order to
decipher the process in which this protein is involved, you must identify this
complex. The cell extract is immunoprecipitated with anti-tag antibodies (anti-
FLAG/anti-GFP/anti-cmyc etc.). The protein components are eluted and
separated by Sodium dodecyl-sulfate polyacrylamide gel electrophoresis
(SDS-PAGE). The protein bands are subjected to in-gel digestion or are
extracted and trypsinized. The interacting proteins A/B/C/D are identified after
the peptide mixtures are separated using the liquid chromatography with
tandem mass spectrometry (LC-MS-MS) approach. Thus, the integration of
molecular biology, protein tagging, immunoprecipitation, and mass
spectrometry facilitates the high-throughput analysis of protein complexes that
are produced within cells.
17.3.3 Chemoproteomics
Understanding the mechanism of action of drugs and small molecules remains
one of the biggest challenges in chemistry and biology sciences. Chemical
proteomics is a new field in chemical biology that aims to design small
molecules to understand protein function. Chemical proteomics is used to
identify the protein binding partners or targets of small molecules in live cells.
In this approach instead of using proteins as bait, small molecules/drugs are
used as bait to look out for interacting proteins.
17.3.4 Phosphoproteomics
Post-translational modifications of proteins for example, phosphorylation,
acetylation, ubiquitination and SUMOylation takes place within the cell after
the translation process. These modifications are essential to regulate protein
activation/inactivation and protein-protein interaction in cells. Phosphorylation
of serine, threonine and tyrosine residues of proteins play a very crucial role
as it can switch on or switch off the function of protein. This is seen in
transcription factors where phosphorylation of some factors activates the
protein and dephosphorylation deactivates it. Understanding the
phosphoproteome is crucial in biology as it provides better insights into protein
function and regulation. It's important to comprehend which proteins within the
cell are phosphorylated as well as how a particular protein site influences
the protein interactions with other proteins. Phosphorylation mapping by mass
spectrometry (MS) has helped in the understanding of phosphoproteome.
6$4
6$4
a) State whether these statements are “True” or “False”:
Unit 17 Proteome Analysis and Application of Proteomics
ii) 2-dimensional electrophoresis is a technique to study proteomics.
Block 4 Applications of Genomics and Proteomics
have a pH gradient. Under the influence of an applied electric field, a protein
migrates towards the electrode with an opposite charge to that of the protein. It
will migrate until it reaches a point on the strip where the pH of the strip is
equal to the isoelectric point of the protein, as the protein shows no migration
after reaching its isoelectric point.
Unit 17 Proteome Analysis and Application of Proteomics
TOF) mass spectrometry (MS), tandem time-of-flight (TOF/TOF) mass
spectrometer (TOF/TOF MS), Electrospray ionization mass spectrometry (ESI-
MS/MS). Software like Melanie, PDQuest, Proteomweaver, Decyder 2D,
Progenesis, REDEFIN etc. are able to identify differentially expressed proteins
in 2DE gel. The differentially expressed proteins are marked on gel and
excised manually. Excised gel with spots are washed and destained. In-Gel
digestion of protein is carried out and samples are further spotted on MALDI
target plate. Sample is analyzed by using mass spectrometry and mass
spectra are obtained. The spectra are submitted to software like MASCOT 1.9
for database search against the National Center for Biotechnology Information
(NCBI) Database and the protein whose expression is changed is identified.
Application and utilities of 2D Gel electrophoresis
It is a powerful technique for proteome analysis and has capability to resolve
thousands of protein at once. Various applications of 2DE are:
• Detection of biomarkers
• Drug discovery and cancer research
• Protein characterization
• Study of post-translational modification
• Protein-protein interaction
CyDye DIGE Fluor Cy2, Cy3, Cy5 Lysine residues Similar to silver
minimal dye staining
CyDye DIGE Fluor Cy3, Cy5 Cysteine residues 100 times to silver
staining
Saturation dye
159
Block 4 Applications of Genomics and Proteomics
Use of internal standard decreases the gel-to-gel variation and is used to
match and normalize the protein patterns across different gels. Two samples
mixed in equal amounts and labeled with a third dye other than the dye used
for labeling samples, serve as an internal standard. The internal standard is
mixed with protein samples and separated on gel. A fluorescence image is
captured on a multiwavelength scanner and image analysis is carried out. The
relative intensity of the labeled sample protein spot of two test samples are
compared to the intensity of the corresponding spot in standard. DeCyder
software allows the identification of spots, co-detection of spots, spot volume
ratio, etc. The steps involved in 2D DIGE are shown in Figure 17.3.
Advantages of 2D-DIGE
• Saturation dye is about 100 times more sensitive than silver staining
Unit 17 Proteome Analysis and Application of Proteomics
The light form is also known as the normal form and the heavy form is known
as the deuterated form. In the heavy form, a hydrogen atom is replaced with a
deuterium atom. The deuterium atom is absent in light form and this results in
mass difference between the light form and heavy form. A mass difference of
at least 5 Da is desirable to allow tagged-peptide ion separation. The isotope-
coded affinity tag reagent consists of three elements, that is, an affinity tag,
linker, and thiol reactive group. The affinity tag (biotin) is used for the isolation
of ICAT-labelled peptides with the help of avidin affinity chromatography. The
linker part helps to form a stable isotope and generates the mass difference
and cysteine residue which is present in protein, covalently forms a complex
with thiol-reactive group and they can be recovered from the mixture of
protein. The ICAT reagent does not change the property of the protein after
labeling.
ICAT Workflow
• Lysis and labeling: Protein samples that contain cysteine chains are
isolated from cells by various methods like cell lysis by freeze-thaw,
sonication, etc. and are labelled. In tagging or labeling, one protein
sample is tagged with the isotopically light form of the ICAT reagent.
Here the cysteinyl residue forms a complex with the thiol reactive group.
Another protein sample is tagged with isotopically heavy reagent.
• Proteolysis: Then both the samples are combined in the ratio of 1:1 and
the proteolysis is done in the presence of proteolytic enzymes like
trypsin etc. for the formation of peptide fragments.
Block 4 Applications of Genomics and Proteomics
the help of micro-capillary high-performance liquid-mass spectrometry.
Protein quantification is accomplished by comparison of integrated peak
intensities. The ratios of the maxima of lower and upper mass
components provide a precise estimate of the relative abundance of
peptides.
Applications
SILAC is a method that involves the incorporation of stable isotopes 13C and
15
N. Two different cell populations are grown in two different culture media
162 which are referred to as light medium and heavy medium respectively. The
Unit 17 Proteome Analysis and Application of Proteomics
light medium contains amino acids that are labeled with natural isotopes (12C,
14
N) whereas the heavy medium contains amino acids that are labeled with
stable isotopes (13C, 15N). After a sufficient number of cell divisions, the cells
are cultured in a heavy medium. Proteins derived from cells grown in heavy
media are now in a heavy state. The number of cell divisions required for the
complete labeling of proteins depends upon the rate of protein synthesis,
metabolism, degradation and turnover. Prior to quantification, the labeling
efficiency of the proteins should be tested. Labeled and unlabeled protein
extracts are mixed in a ratio of 1:1. The samples are then digested with the
help of trypsin to small peptides and they are analyzed with the LC-MS/MS
technique. The intensity of the signals from light and heavy samples allows for
a quantitative assessment of their relative abundance in the mixture (Fig.
17.6). Leucine, lysine and methionine are the essential amino acids that have
been used in SILAC. Though arginine is not an essential amino acid still
arginine has also been used in SILAC because it is essential for the growth of
some cells.
Fig. 17.6: Principle of stable isotope labeling by amino acids in cell culture
SILAC workflow
SILAC workflow consists of two phases namely the adaptation phase and the
experimental phase. In the adaptation phase, the cells are subjected to growth
in labeled and unlabeled media until the heavy amino acids have been
completely incorporated into the cellular proteins. The level of SILAC amino
acid incorporation is then determined by LC-MS/MS. The area under the curve
(AUC) of the MS peaks for the remaining light and heavy peptide pairings is
used to assess the degree of labeling (Fig. 17.7a). 163
Block 4 Applications of Genomics and Proteomics
During the experimental phase, after the full incorporation of heavy amino
acids has been confirmed, the two cell populations are subjected to different
treatments based on the experiment and then combined equally prior to
optional subcellular organelle purification, cell lysis, protein extraction and
protein digestion. The samples are then examined using LC-MS/MS to identify
and quantify the heavy peptide to light peptide ratios (Fig. 17.7b).
• Expression proteomics
• Protein-protein interactions
• Protein turnover
Unit 17 Proteome Analysis and Application of Proteomics
6$4
6$4
Fill in the blanks:
a) The …….………. and …….………. are the two distinct steps to separate
proteins in the Two-dimensional gel electrophoresis technique.
e) The Fluorescent cyanine dyes (CyDye) are used to label protein from
different samples in the …….………. technique.
Block 4 Applications of Genomics and Proteomics
Table 17.2: Drug targets that are identified by proteomics.
Protein expression study has a crucial role in drug development and research.
Proteomics allows for the study of the impacts on protein expression and
pattern, as well as the mechanism of action of drugs, their toxicological and
therapeutic effects, and abnormal protein expression in specific conditions. In
proteomics, both covalent modification and processing can be studied at the
protein level and this plays an important role in disease biology and gives
important information about disease-specific conditions for the use of
diagnostic markers or therapeutic agents against that disease. The
identification of the protein of interest, confirmation and purity checking can be
done with the help of proteomics.
Production
The discovery of the diagnostic markers and vaccine can be done with the
help of proteomics and it is a promising tool for disease-associated biomarker
detection. For the development and production of biomarkers and vaccines,
different tools of proteomics can be used such as 2D-PAGE, MALDI-TOF,
surface-enhanced laser desorption, ionization (SELDI) and protein chip
techniques. For the study of proteins on a large scale, proteomics is used and
with the help of different methods such as ICAT, 2D gel electrophoresis,
SILAC etc. sample isolation, identification and quantification is done which is
crucial for its production. Proteomics not only assists with producing
chemically stable, highly specific products, but it also ensures and predicts the
quality of the final product.
Unit 17 Proteome Analysis and Application of Proteomics
Safety
Target identification
2-Dimensional Gel Electrophoresis along with mass spectroscopy helps in
identifying the protein expression changes within a particular system. Using
protein sequence tags (PST), each protein is terminally tagged, isolated and
sequenced which helps in rapid identification of any set of proteins produced
by a cell. Apart from this, multidimensional protein identification technology
(MudPIT) uses strong cation exchange and reverse-phase adsorbent
separation columns for the identification of protein targets through LC/MS
analysis whereas isotope-coded affinity tagging (ICAT) uses an ICAT reagent
that binds to a particular amino acid usually a cysteine, light or heavy isotope
and an affinity tag biotin are incubated with each sample. The samples of
different groups or disease states are incubated with different isotopes and are
mixed with equal proportions and lysed after lysing the labeled peptides
identified by LC/MS technique. Control and drug-treated samples are
subjected to proteomics for target identification. Proteins whose expression is
upregulated or downregulated are expected to be the target for that drug. 167
Block 4 Applications of Genomics and Proteomics
Proteins make up the majority of therapeutic targets that are used to initiate
drug design processes. Proteomics is a powerful tool for identifying targets by
thoroughly analyzing changes in protein expression and protein-protein
interactions that take place over the course of a disease or after therapeutic
treatment. Analyzing the proteome profiles of cells treated with a drug is a
typical target discovery method. Compared to untreated cells, the changed
proteins in the signaling pathway or gene network regulator are studied..
Proteomics based research has been carried out to look into the cellular
pathways that drugs work on as well as the molecular basis of
pharmacological activity. Proteomics has facilitated the study of the
mechanisms by which small-molecule medicines interact with the proteome
via two sophisticated procedures: thermal proteome profiling (TPP) and
multiplexed proteome dynamics profiling (mPDP). The mPDP technique
permits the finding of regulated protein synthesis and degradation processes
brought on by small molecules, while TPP evaluates changes in protein
thermal stability in response to drug treatment and so provides information on
direct targets and downstream regulation events.
Target validation
Once the target is identified, the next step in drug discovery is the validation of
the identified target. The validations are done by overexpression or knockout
of the gene of interest by homologous recombination in the organism. In some
cases, RNAi approach is also used. Phenotypic changes are observed to
evaluate its essentiality. Samples are put through proteomics after
overexpression or knockout to observe how these events affect various
proteins. The upregulation or downregulation of proteins and their identity are
revealed by proteomics which highlights the class of proteins (metabolic
protein, transporter, cell cycle protein, etc.) whose expression is altered. This
provides an understanding of the metabolic pathways/signaling pathways/cell
cycle proteins that may have been affected due to overexpression and
knockout studies and may predict the essentiality of the knockout gene.
Unit 17 Proteome Analysis and Application of Proteomics
the lead compounds is done by screening the structure-activity relationships
among the drugs. Apart from the virtual screening, activity-based probes
(ABP) also help in identifying the potential drug compounds with particular
proteins.
Biomarkers should have high specificity for disease and proteomics offers
powerful techniques for biomarker identification, characterization and
validation. The process of confirming the assay, its performance
characteristics, and the necessary ideal conditions to produce reproducibility
and accuracy is known as analytical technique verification. Clinical or
biological validation is related to how a particular marker performs in a 169
Block 4 Applications of Genomics and Proteomics
population and between populations. The incorporation of validated proteomic
biomarkers into clinical drug development programs will improve the decision-
making process by adding critical information about the pharmacological and
pharmacodynamic mechanism of drug targets.
Drugs are metabolized and eliminated in the liver and it is often the most
targeted organ for studying toxicology. Hepatotoxicity is dose dependent and it
is studied in 28 days in vivo. Hepatotoxicity can be observed in later stages of
drug development and may cause hazards, so early detection by proteomics
study will help us in managing the hazard. Proteomics studies of functional
molecules will give insight into possible mechanisms of action. In vitro models
for hepatotoxicity are based on cell lines like HepG2, HepaRG and
hepatocytes. Hepatocytes are most widely used to study drug metabolism and
toxicity as they are capable of biotransforming drugs. After administration of a
drug, the expression of liver-specific proteins is checked. Also changes in the
level of CYPs 2B, 1A is checked by 2D Gel electrophoresis. In vivo analysis is
based on the use of test animals like rodents. About 28-90 days repeated
dose toxicity tests are carried out to observe the chronic effect, organ toxicity,
differential protein expression, etc. Proteomics investigations can be
performed on tissue, cellular fractions, plasma proteome, etc. An example of
hepatotoxicity studies using proteomic endpoint in human is the administration
of test compound acetaminophen, amiodarone and cyclosporine A. Protein
expression changes in HepG2 was studied using DIGE and mass
spectrometry. A total of 254 differentially expressed proteins were identified
and analyzed. High differential expression of secreted proteins such as serum
albumin, ApoA1, serotransferrin, and ER-Golgi transport network was
observed.
Lung tissues are collected and analyzed using iTRAQ technique. Changes in
170 the level of oxidative stress proteins and inflammatory mediators are studied.
Unit 17 Proteome Analysis and Application of Proteomics
Proteomics studies for heavy metal toxicity
The toxic effect of heavy metals on protein expression can be studied by
proteomics. Mechanisms of metal toxicity can be studied by evaluating
changes in protein after interaction with heavy metal. Biomarker identification
can help in designing diagnostic tests for detecting protein toxicity. In vitro
toxicity assay can be performed on cell lines and in cultured cells. In vivo
assay can be performed in zebrafish, insect, or rat models. Heavy metals tend
to accumulate in the brain and liver so they are the most focused organs. 2-DE
is the most widely used technique as it can simultaneously resolve many
proteins. Heavy metals have shown differential expression of proteins related
to antioxidant defense mechanisms. Many proteomics studies have shown
that enzymes involved in glutathione (GSH) are differentially regulated in case
of heavy metal poisoning. Heavy metal poisoning affects the heat shock
proteins (HSP) which are generally involved in protein folding, aggregation and
stability, so it can be assumed that HSP has a role in cellular defense against
heavy metal-induced stress. The upregulation of proteins associated with
energy production may be related to the higher energy required for
detoxification. Utilizing proteomic techniques, particularly quantitative
proteomics, will enable the generation of more precise and reliable results,
which will undoubtedly advance this developing field and lead to the
identification of novel biomarkers and new insights into the mechanism of
metal toxicity.
Block 4 Applications of Genomics and Proteomics
Step 2: Target Exposure: The library is then exposed to an immobilized
target such as a receptor, enzyme or ligand.
Step 5: Amplification: Eluted phages with specificity and affinity for binding to
the target are then replicated in bacteria. Amplification produces a phage
mixture that is enriched with binding to a specific target. The repeated cycling
of these steps is called biopanning (Fig. 17.9).
Unit 17 Proteome Analysis and Application of Proteomics
2D gel is used widely for the study of cellular proteins. However, it is labor-
intensive procedure and the variability is too high. So, affinity agents like
monoclonal antibodies are used to identify a given protein. Monoclonal
antibodies bind specially to a single epitope so they can be used to establish
the identity of a given protein. The Hybridoma technique was used earlier to
produce antibodies but it is difficult to produce a large number of monoclonal
antibodies needed for proteomics studies. The phage display technique
provides an alternative to produce a large number of monoclonal antibodies.
Cellular proteins are first separated by 2D Gel electrophoresis and blotted on
polyvinylidene fluoride membrane. Other sites on the PVDF membrane are
blocked so the only target available will be blotted antigen. The membrane is
incubated with a phage antibody library and washing is done to remove non-
bounded antibodies. Two to three rounds of selection are done to increase the
frequency of positive clones. Selected antibodies are studied by western
blotting.
The region of antigen to which the antibody binds is called epitope. Locating
the targeted antigen's binding sites where an antibody binds is known as
epitope mapping. The phage display library is used to display a number of
peptides. Antibodies have the capability to select peptides with high affinity for
their paratopes from these libraries. The phage display library is used to define
peptide structures and is recognized by major histocompatibility (MHC)
molecules. MHC molecules bind to peptide fragments derived from pathogens
and display them on the cell surface, which are then recognized by T cells.
Epitope mapping is significantly used in vaccine development and allows the
construction of peptide vaccines based on epitope specificity.
6$4
6$4
Fill in the blanks:
173
Block 4 Applications of Genomics and Proteomics
17.6 SUMMARY
• Proteomics is the large-scale study of proteomes. A proteome is a set of
proteins produced in an organism, system, or biological context. The
proteome is not constant; it differs from cell to cell and changes over
time.
Unit 17 Proteome Analysis and Application of Proteomics
• Phage display technology is used in protein-ligand interactions, protein-
protein interaction producing monoclonal antibodies and improving their
affinity and epitope mapping.
• The use of the phage display technique along with other techniques like
yeast two hybrid system can be very beneficial in understanding protein-
protein interaction.
17. 8 ANSWERS
Self-Assessment Questions
1. a) i) False, ii) True, iii) True, iv) True
b) isoelectric point
c) anionic
Block 4 Applications of Genomics and Proteomics
Terminal Questions
1. Refer to Sections 17.1 and 17.2.
176
GLOSSARY
Acrocentric : A chromosome where the centromere is not
chromosomes central and is instead located near the end of the
chromosome.
Albinism : It is derived from the Latin albus, meaning
"white," is a group of heritable conditions
associated with decreased or absence of melanin
in ectoderm-derived tissues (most notably the
skin, hair and eyes), yielding a characteristic
pallor.
Volume 2 Proteome Analysis and Applications
across two axes is provided by a basic heat map,
which enables users to rapidly identify the most
significant or pertinent data points. Complex data
sets can be understood by the viewer with more
intricate heat maps.
Volume 2 Proteome Analysis and Applications
PepNovo : It is a high throughput de novo peptide
sequencing tool for tandem mass spectrometry
data.
Volume 2 Proteome Analysis and Applications
Recombinant DNA : Recombinant DNA technology involves using
technology enzymes and various laboratory techniques to
manipulate and isolate DNA segments of interest.
This method can be used to combine (or splice)
DNA from different species or to create genes
with new functions. The resulting copies are often
referred to as recombinant DNA.
Reporter gene : A reporter gene is a nonendogenous gene
encoding an enzyme or fluorescent protein
whose expression is controlled by a promoter for
a separate gene of interest. It allows for
identifying, quantifying, visualizing and tracking
gene expression and protein distribution in cells.
RITS complex : RNA-induced transcriptional silencing (RITS)
complex, consisting of Ago1, Tas3 and
Chp1, binds to nascent transcripts from
centromeric chromatin (cenRNA). RITS leads to
transcriptional silencing by placing chromatin
marks and recruiting the RNA-dependent RNA
polymerase complex (RDRC).
RNA-induced : One strand of a small interfering RNA (siRNA) or
silencing complex micro RNA (miRNA) is incorporated into the
multiprotein complex known as the RNA-induced
silencing complex (RISC). The siRNA or miRNA
serves as a template for complementary mRNA
recognition in RISC. It initiates RNase activity and
cleaves the RNA when it comes across a
complementary strand. This process is crucial for
both defence against viral infections, which
frequently use double-stranded RNA as an
infectious vector, and for the control of genes by
microRNAs.
Scott syndrome : It is a rare autosomal recessive congenital
bleeding disorder caused by a defect in
blood coagulation.
SHERENGA : It is an algorithm for de novo interpretation of
MS/MS spectra.
Sodium Dodecyl : Technique used for the separation of proteins,
Sulphate- based on their molecular weight.
Polyacrylamide Gel
Electrophoresis
Southwestern Blotting : Technique used to study DNA-protein
interactions. This method detects specific DNA-
binding proteins by incubating radiolabeled DNA
with a gel blot, washing and visualizing through
autoradiography.
Telomeres : Telomeres are structures made from DNA
sequences and proteins found at the ends of
chromosomes. 181
Volume 2 Proteome Analysis and Applications
The International : The aim of this project is to determine the
HAPMAP Project common patterns of DNA sequence variation in
the human genome and to make this information
freely available in the public domain. There is an
international consortium involved in developing a
map of these patterns across the genome. This is
possible through determining the genotypes of
sequence variants, their frequencies and the
degree of association between them, in DNA
samples from populations with ancestry from
parts of Africa, Asia and Europe. This will lead to
the discovery of sequence variants that affect
common disease, thereby facilitating the
development of diagnostic tools.
Time of Flight Mass : Technology that utilizes an electric field to
Analyzer accelerate generated ions through the same
electrical potential, and then measures the time
each ion takes to reach the detector.
Transcription factors : These are proteins involved in the process of
converting or transcribing DNA into RNA.
Transfection : It refers to the introduction of foreign DNA
(genetic material other than host genomes) into
the cell. The main purpose of transfection is to
alter the host genome to express or block the
expression of the protein, associated with the
gene.
Transformation : It is a process by which foreign genetic material is
taken up by a cell. The process results in a stable
genetic change within the transformed cell.
Two-Dimensional Gel : This technique separates proteins, depending on
Electrophoresis two different steps: the first one is called
isoelectric focusing which separates proteins
according to isoelectric points (pI); the second
step is SDS-PAGE which separates proteins,
based on the molecular weights.
Western Blotting : Procedure for the immunodetection of proteins,
particularly proteins that are of low abundance.
This process involves the transfer of protein,
patterns from gel to microporous membrane.
Yeast One-Hybrid : Important technique for detecting physical
Assay interactions between sequence-specific
regulatory transcription factor proteins and their
DNA target sites. It involves two components: (1)
a reporter construct with DNA of interest cloned
upstream of a gene encoding a reporter protein
that can be easily detected; and (2) an
expression construct that generates a fusion (or
“hybrid”) between a transcription factor of interest
and a yeast transcription activation domain.
182