0% found this document useful (0 votes)
13 views103 pages

Genome and Proteome Analysis Overview

Block 4 of the course MZO-005 on Genomics and Proteomics covers applications of genomic and proteomic analysis, detailing techniques for genome manipulation, expression analysis, and proteome applications. The units discuss the sequencing of DNA, gene manipulation methods, and the role of proteomics in pharmaceuticals and drug development. The block aims to equip students with knowledge on genome analysis, gene expression, and the implications of these studies in various fields including medicine and agriculture.

Uploaded by

jobs843420
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
13 views103 pages

Genome and Proteome Analysis Overview

Block 4 of the course MZO-005 on Genomics and Proteomics covers applications of genomic and proteomic analysis, detailing techniques for genome manipulation, expression analysis, and proteome applications. The units discuss the sequencing of DNA, gene manipulation methods, and the role of proteomics in pharmaceuticals and drug development. The block aims to equip students with knowledge on genome analysis, gene expression, and the implications of these studies in various fields including medicine and agriculture.

Uploaded by

jobs843420
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

MZO-005

GENOMICS AND
Indira Gandhi PROTEOMICS
National Open University
School of Sciences

Block

4
APPLICATIONS OF GENOMICS AND
PROTEOMICS
UNIT 14
Analysis of the Genome 85
UNIT 15
Manipulation of the Genome 102
UNIT 16
Expression Analysis of Genome 134
UNIT 17
Proteome Analysis and Application of Proteomics 154



COURSE NAME: GENOMICS AND PROTEOMICS COURSE CODE: MZO-005

Course Design Committee


Prof. Sujatha Varma Prof. Bharati Chauhan Prof. Amrita Nigam (Retd.)
Former Director, School of Department of Zoology School of Sciences, IGNOU
Sciences, IGNOU, Maidan Garhi University of Rajasthan Maidan Garhi, New Delhi-110068
New Delhi-110068 Jaipur-302004
Prof. Neera Kapoor
Prof. Neeta Sehgal Prof. Devinder Kaur Kochar School of Sciences, IGNOU
Department of Zoology Department of Zoology, Punjab Maidan Garhi, New Delhi-110068
University of Delhi, Delhi-110007 Agricultural University
Dr. Siya Ram
Ludhiana-141004
Prof. Sarita Kumar School of Sciences, IGNOU
Department of Zoology, Acharya Prof. Varsha Wankhade Maidan Garhi, New Delhi-110068
Narender Dev College, University Department of Zoology
Dr. Ravi Rajwanshi
of Delhi, Delhi-110007 Savitribai Phule University
School of Sciences, IGNOU
Pune-411007
Prof. S. Dinakaran Maidan Garhi, New Delhi-110068
Department of Zoology Prof. Sarita Sachdeva
Madurai Kamaraj University Department of Biotechnology
Tamil Nadu-625021 Manav Rachna International
University, Sector-43, Faridabad
Haryana-121004

Block Preparation Team


Dr. Mulaka Maruthi Dr. Deepak Kumar School of Sciences
Department of Biochemistry, Department of Botany, Center of
Dr. Ravi Rajwanshi
Central University of Haryana Advanced Study in Botany,
Jant Pali, Mahendragarh Institute of Science, Banaras School of Sciences, IGNOU
Maidan Garhi, New Delhi-110068
Haryana, India (Unit 14) Hindu University
(Units 14, 15, 16 and 17)
Varanasi-221005 (Unit 16)
Dr. Gautam Kumar
Department of Life Sciences, Prof. Sushma Singh
Central University of South Bihar, Department of Biotechnology
Panchanpur, Gaya-824236 NIPER, Sector 67 S.A.S. Nagar
(Unit 15) Mohali, Punjab-160062 (Unit 17)

Programme Coordinators : Prof. Neera Kapoor and Dr. Siya Ram

Course Coordinators : Prof. Aryadeep Roy Choudhury and Dr. Ravi Rajwanshi
Course Editor : Prof. Satheeshkumar P.K.
Center of Advanced Study in Botany, Institute of
Science, Banaras Hindu University, Varanasi
UP-221005, India

Production : Mr. Tilak Raj


AR, MPDD, IGNOU
Acknowledgement:
• Dr. Ravi Rajwanshi and Mr. Ajit Kumar, suggestions for figures and Cover Design.
• Mr. Vikas Kumar, JAT for word processing and CRC preparation.
June, 2024
©Indira Gandhi National Open University, 2024

All rights reserved. No part of this work may be reproduced in any form, by mimeograph or any other
means, without permission in writing from Indira Gandhi National Open University.
Further information on Indira Gandhi National Open University courses may be obtained from the
University’s office at Maidan Garhi, New Delhi-110 068 or IGNOU website [Link].
Printed and published on behalf of Indira Gandhi National Open University, New Delhi by the Registrar,
MPDD, IGNOU.



BLOCK 4: APPLICATIONS OF GENOMICS AND
PROTEOMICS
Block 4, “Applications of Genomics and Proteomics”, consists of the four Units that
focus on the Genome-proteome analysis and their applications. The genome and proteome
are essential threads that tell the story of an organism's existence in the complex tapestry of
life. The proteome is the dynamic expression of genes that regulates the functions and
behaviors’ of living organisms, whereas the genome is similar to a blueprint that contains the
genetic information passed down through generations. By providing insights into evolution,
disease causes, and possible therapeutic approaches, the analysis and applications of these
genomic landscapes have revolutionised a variety of fields, including agriculture and
medicine. The present block will discuss the topics related to the analysis of the genome of
few important organisms including humans, various techniques for the manipulation of the
genome, expression analysis of the genome followed by the proteome analysis and its
applications for the benefit of mankind.

In Unit 14, you will learn about sequencing of DNA and the steps involved in the analysis of
the sequenced genome. The genome of few important organisms such as Plasmodium
falciparum, Mycobacterium tuberculosis along with humans has been discussed in detail.
Manipulation of genomes via PCR based Site-directed mutagenesis is also explained along
with its protocol.

The Unit 15 of this block focuses on the manipulation of the genome that involves the
techniques like cloning of genes as well as the role of reporter gene in the identification of the
locations of regulatory sequences like enhancer elements that drive a specific pattern of
gene expression. The Unit also discusses the Gene knockout and knockin methods in
transgenics that can alter genes in a selected model system, providing valuable insights into
the functioning of individual genes. In this Unit you will learn about the important genome
sequencing techniques and the role of Restriction Enzymes along with the cloning vectors in
the manipulation of the genome.

Unit 16 titled as “Expression analysis of genome” highlights the concept and function of
reporter genes. In this Unit, you will also learn about the basics of temporal and site-specific
gene expression. Also you will understand the detailed information provided in this Unit about
the Gene Silencing with special reference to the mechanism, biological functions and
applications of RNA interference.

Unit 17 is the last Unit of this block as well as course in which you will learn about the
Proteome analysis and applications of proteomics. This Unit will provide you the information
about proteomics and its various forms. The techniques of proteomics such as Two
dimensional gel electrophoresis, Two dimensional difference gel electrophoresis, Isotope
Coded Affinity Tag, Stable isotope labeling by amino acids in cell culture are discussed in
detail. The Unit also describes the different applications of proteomics in the field of
Pharmaceuticals, drug development and toxicology. Phage antibody as tool and application
of phage display in proteomics is also explained in this unit. You will also learn about the
various high throughput techniques used in proteomics based studies and application of
proteomics for the welfare of society. 83



Objectives
After studying this block, you would be able to:

• describe the steps involved in genome analysis,

• comprehend the results obtained from genomes analysis of different organisms


including the human genome,

• explain the methods of gene manipulation and techniques of genome sequencing,

• describe the different types of vectors constructed for cloning DNA,

• comprehend the mechanism of action of the restriction enzymes to cleave DNA and
steps of DNA cloning,

• describe the concept and function of reporter genes,

• explain the basics of temporal and site-specific gene expression,

• appreciate the concept of gene silencing,

• explain the mechanism, biological functions and applications of RNA interference,

• describe proteomics and its forms,

• discuss the techniques used to study proteomics, and

• explain the role of proteomics in drug development and toxicology and enumerate its
application in drug discovery in humans and pharmaceutical industry.

84





UNIT 14
$1$/<6,62)7+(*(120(
$1$/<6,62)7+(*(120(

6WUXFWXUH
6WUXFWXUH
14.1 Introduction Human Genome Project:
Implications for Medical
Objectives
Science
14.2 Genome
14.7 Genome Analysis of
DNA Sequencing Plasmodium Falciparum
Genome Analysis 14.8 Genome Analysis of
The Bulky Genomes in Plants Mycobacterium
Tuberculosis
14.3 Analysis of Genomes
Sequence Analysis
14.4 Steps of Genomic Data
Analysis Genes Encoding Proteins

Data Collection 14.9 Manipulation of Genomes:


Site-directed Mutagenesis
Quality Check and Cleaning
of Data Protocol: Site-directed
Mutagenesis using PCR
Data Processing
14.10 Summary
Exploratory Analysis of Data
and Modeling 14.11 Terminal Questions

Visualization and Reporting 14.12 Answers


14.5 Genomics-specific Data
Analysis
14.6 Analysis of the Human
Genome (Human Genome
Project)
Recent advancements in the
Human Genome Studies

14.1 INTRODUCTION
Genomic analyses allow researchers and clinicians to learn about differences
and changes in an organism’s or individual’s (be it an animal, a plant, a
bacterium, an archaean, a protist, a fungus, or a virus) genetic makeup,
leading to understand and discover the mechanism of disease progression
Block 4 Applications of Genomics and Proteomics
and help to design the treatment strategies. Genome projects ultimately aim to
determine the complete genome sequence of an organism and annotate
protein-coding genes, non-coding genes, and other important genome-
encoded features. The genome of an organism includes the complete DNA
sequences of each chromosome in the organism.

Mutagenesis is the process by which an organism's DNA sequence changes,


resulting in a gene mutation. Naturally, DNA mutagenesis occurs
spontaneously as a result of DNA exposure to endogenous and exogenous
mutagens, and by errors during DNA replication. Furthermore, in laboratories,
molecular biology techniques such as polymerase chain reaction (PCR), are
extensively used to create mutations in the DNA sequences and to analyse
the mutations, the modified DNA sequences are used to produce desired traits
in microbes, plants, and animals, for developing genetically modified
organisms (GMOs), GM food and therapeutic products.

The present Unit provides a brief description of the steps involved in genome
analysis, features of the whole genome analysis of Plasmodium,
Mycobacterium tuberculosis, and the human genome project. This Unit also
descries the genetic modification methods used to manipulate the gene
sequences, especially by site-directed mutagenesis.

2EMHFWLYHV
2EMHFWLYHV
After studying this Unit, you would be able to:

™ describe the steps involved in genome analysis,

™ comprehend the results obtained from genomes analysis of different


organisms including the human genome, and

™ explain the methods of gene manipulation in laboratory.

14.2 GENOME
The complete set of DNA content in an organism is called its genome.
Virtually, every single cell in the human body contains a complete copy of 3
billion DNA base pairs approximately, that make up the human genome. DNA
contains the information needed to build and function the entire organism. A
gene refers to the unit of DNA that carries the information for making a specific
protein or set of proteins. For example, each gene in the human genome
codes for an average of three proteins.

14.2.1 DNA Sequencing

Sequencing is considered the “gold standard” technique for the identification of


known as well as unknown/unspecified variants in genomic DNA. Sequencing
literally means determining the exact order of the nucleotide bases in a strand
of DNA. The most common type of sequencing (the Sanger or Next-
Generation Sequencing (NGS) techniques) used in recent times, is called
sequencing by synthesis, in which DNA polymerase generates a new strand of
86 DNA using the template (sample DNA). During this reaction, polymerase adds


Unit 14 Analysis of the Genome
fluorescently labeled deoxyribonucleoside 5’-triphosphates (dNTPs) into the
new DNA strand. The nucleotide is excited by a light source, resulting in the
emission of a fluorescent signal and detected by a detector.

Researchers use DNA sequencing to search for specific genetic variations


and/or mutations, that may play a role in the growth and development or
progression of a disease. The disease-causing source may be a small
substitution, addition, or deletion, of a single base pair or a large portion of
DNA.

14.2.2 Genome Analysis


Genomic analysis is the identification, comparison or measurement of specific
genomic features such as DNA sequence, gene expression, structural
variation, or regulatory and functional elements annotation at a genomic scale.
Methods for genomic analysis typically involve high-throughput sequencing or
microarray hybridization and computational analysis (bioinformatics).
Computational Gene analysis is relatively simple for the prokaryotes, in which
all the genes are transcribed into mRNA and then translated into proteins. The
analysis process is more complex for eukaryotic cells, in which the coding
DNA sequence is not continuous and is interrupted by random sequences
called introns.

Some of the questions which biologists want to answer using genome analysis
are:

• Classification of a newly sequenced genome into, the genes (coding)


and the non-coding regions.

• Given an organism’s DNA sequence, what regions of the sequence code


for a protein, and what regions are junk DNA?

• Classification the junk DNA as intron, dead genes, transposons,


untranslated region, regulatory elements, etc.

The importance of genome analysis can be understood by comparing the


genomes of humans and chimpanzees. The chimpanzee and human genomes
vary by only an average of 2% i.e. only about 160 enzymes. The complete
genome analysis would give a strong insight into the various mechanisms
responsible for the physiological and morphological differences between
different organisms (Table 14.1).

Table 14.1: List of estimated sizes of different organisms’ genome and


the number of genes.

Species Genome size (Mb) Number of genes

Mycoplasma genitalium 0.58 500

Streptococcus pneumoniae 2.2 2300

Escherichia coli 4.6 4400

Saccharomyces cerevisiae 12 5800

Caenorhabditi elegans 97 19,000


87


Block 4 Applications of Genomics and Proteomics

Arabidopsis thaiana 125 25,500

Drosophila melanogaster 180 13,700

Oryza sativa 466 45-55,000

Mus musculus 2500 29,000

Homo sapiens 3300 27,000

14.2.3 The Bulky Genomes in Plants


The bulky nature of plant genomes is due to two important factors, the ability of
the plants to duplicate their genomes in order to reproduce (polyploidization),
and the susceptibility of plants to mobile genetic elements. Polyploidization
allows plants to form hybrids easily when pollen and ova from different species
fertilize. Hybridization events produce plants with genomes that are the sum of
the two parent genome sizes (different from half of each parent’s genome in
normal sexual reproduction). It is common to observe the insertion of
transposable elements in intergenic regions of the plant’s genome. This also
explains the difference in the sizes of the genomes in plants among
themselves as well.

14.3 ANALYSIS OF GENOMES


In order to understand the genome structure, evolution, or function, it is not
enough to obtain only the DNA sequence, there is also a requirement for deep
and precise analysis using bioinformatics applications. Genome analysis is the
prediction of uncharacterized genomic sequences, identification, measurement,
or comparison of features such as DNA sequence, variations, or functional and
regulatory elements at a genomic scale. The genome analysis methods
typically require high-throughput sequencing, microarray hybridization, and
bioinformatic tools (Fig. 14.1).

Fig. 14.1: Major steps of genome analysis. To analyze any organism’s


genome, the genetic material (DNA/RNA) has to be extracted, which
further gets enriched by enrichment protocols. The enriched genetic
material is subjected to high-throughput sequencing to get the
sequence, and microarray techniques are used to quantify the gene
88 expression


Unit 14 Analysis of the Genome
The key to successful sequence analysis is the alignment of the sequence of
interest with another sequence whose function is known (reference genome).
This will reveal the function of the unknown genes, and also the evolutionary
relation between the sequences/organisms. Furthermore, the sequences can
be analyzed to find the significant matches between the domains of a
sequence that have been previously described which may have a huge impact
on the genes function.

14.4 STEPS OF GENOMIC DATA ANALYSIS


Irrespective of the analysis type, every data analysis follows a common
pattern, which typically includes steps of data collection, quality checking
and cleaning, processing, modeling, visualization, and reporting. In practice,
data analysis requires going through the same steps in a repeated fashion
in order to answer other related questions, which deal with data quality
issues that are later realized and include new data sets in the analysis.

14.4.1 Data Collection

Data collection refers to any experiment, source, or survey that provides


data for the data analysis. In genomics, data collection is done by high-
throughput assays. One can also use publicly available data sets and
specialized databases. The type of data collected depends on the type of
study and the technical and biological variability of the samples.

14.4.2 Quality Check and Cleaning of Data

Mostly, data analysis deals with imperfect data such as missing values or
measurements that are noisy. Data quality check and cleaning aims to
identify any data quality issues and rectify them by cleaning them from the
dataset. Identifying low-quality or missing bases and removing them from
the dataset will improve the read-mapping step.

14.4.3 Data Processing

This step involves the processing of data into a suitable format for
exploratory analysis and modeling. Sometimes, the data needs to be
converted into other formats by transforming data points (such as log
transformation, normalization, etc.), or fragment the data into subsets with
some arbitrary or pre-defined conditions. In terms of genomics, processing
includes aligning the sequence reads to the genome and quantification over
genes or regions of interest.

14.4.4 Exploratory Analysis of Data and Modeling


This phase is usually required for the processed or semi-processed data
and involves the application of machine learning or statistical methods to
explore the data. Typically, the relationship between variables and the
relationship between samples based on the variables are measured.
Additionally, cleanup or re-processing is performed to minimize the
anomalies. 89


Block 4 Applications of Genomics and Proteomics
In the context of genomics, modeling is predicting the disease status of the
patients from the gene expression values measured from their tissue
samples, if the variable of interest is disease status. This kind of approach is
generally called “predictive modeling”, and involves regression-based
machine learning methods.

14.4.5 Visualization and Reporting


Visualization is an important part of computational genomics data analysis.
In the final phase, data needs to be represented in the form of figures,
tables, graphs, and text that explains the outcome of the analysis. In
genomics, common data visualization methods, as well as specific
visualization methods such as R programming are developed for genomic
data analysis, to represent the data in basic plots (Histograms, scatter plots,
box plots, bar plots, heatmaps), ideograms, circus plots, meta profiles of
genomic features and visualization of quantitative assays for gene locus.

6$4
6$4
a) Choose the correct option in the multiple choices given below each
question.

i) Genomics is the study of:

A. Individual genes

B. Chromosomes

C. Whole genomes

D. Genetic disorders

ii) The technique used to determine the sequence of DNA is:

A. PCR (Polymerase Chain Reaction)

B. Microarray analysis

C. Sanger sequencing

D. Gel electrophoresis

b) Fill in the blanks:

i) The ……………… is the process by which an organism's DNA


sequence changes, resulting in a gene mutation.

ii) A ……………… refers to the unit of DNA that carries the


information for making a specific protein or set of proteins.

iii) The ……………… allows plants to form hybrids easily when pollen
and ova from different species fertilize.

iv) The key to successful sequence analysis is the alignment of the


sequence of interest with another sequence known as
……………… genome.

v) Any experiment, source, or survey that yields data for data


analysis is referred to as ……………… .
90


Unit 14 Analysis of the Genome

14.5 GENOMICS-SPECIFIC DATA ANALYSIS


Analysis and visualization tools like R and Bioconductor give access to a
multitude of other bioinformatics-specific algorithms, which can perform
sequence analysis such as TF binding motifs, GC content and CpG counts
of a given DNA sequence, Differential gene expression analysis, enriched
gene set/pathway analysis, identification of genomic intervals like CpG
islands and transcription start sites, single nucleotide polymorphisms (SNPs),
variants, and analysis of overlapping aligned reads can also be done.

To compare the data obtained from genome analysis, it is necessary to search


for information in different databases of biomedical information. The sources of
biomedical and genomic information are the NCBI (National Center for
Biotechnology Information), EMBL (European Molecular Biology Laboratories),
and DDBJ (DNA Data Bank of Japan). NCBI provides access to other
databases such as PubMed, OMIM, Entrez Gene, dbSNP, Variation Viewer,
and others.

14.6 ANALYSIS OF THE HUMAN GENOME


(HUMAN GENOME PROJECT)
Model organisms have been sequenced from both the plant and animal
kingdoms. The 21st century has witnessed the announcement of the draft
version of the human genome sequence. The Human Genome Project, which
was led by the National Human Genome Research Institute at the National
Institutes of Health (NIH), produced a very high-quality version of the freely
available human genome sequence. The sequence is a composite derived
from several individuals, thus referred to as a "representative" or generic
sequence. To ensure the privacy of the DNA donors, nearly 100 blood
samples were collected from volunteers for DNA sequencing, and no names of
the volunteers were attached to the samples that were analyzed.

The Human Genome Project was designed with the aim of generating a
resource that could be used for a wide range of biomedical studies such as the
study of genetic variations that increase the risk of certain diseases (like
cancer, and cardiac diseases), or to look for specific genetic mutations
frequently observed in cancerous cells. In its initial release in 2003 (Table.
14.2), the human genome project has covered only the euchromatic regions of
the genome, and not covered the important heterochromatic regions. After
multiple updates, the GRC in 2017 released GRCh38, which is the first
coordinate-changing assembly update since 2009; GRCh38 reflects the
resolution of roughly 1000 issues and covers modifications ranging from
thousands of single base changes to megabase-scale range reorganizations,
localization of previously orphaned sequences, and gap closures.

14.6.1 Recent Advancements in the Human


Genome Studies
In 2022, researchers of the Telomere-to-Telomere (T2T) Consortium
presented a complete 3.055 billion base pair (bp) sequence of a human
genome (T2T-CHM13), that includes corrected errors in the prior references,
gapless assemblies for all chromosomes (except Y), and introduced nearly 91


Block 4 Applications of Genomics and Proteomics
200 million bp of sequence containing 1,956 gene predictions, 99 of which are
predicted to be protein-coding (Fig. 14.2). The completed regions include
recent segmental duplications, all centromeric satellite arrays, and the short
arms of all five acrocentric chromosomes, uncovering these complex regions
of the genome and making them available for variational and functional
studies. Twenty years after the initial draft was published, a truly complete
sequence of a human genome reveals what has been missing.

Table 14.2: Recent human genome assemblies released

Release name Date of release

T2T-CHM13 January 2022

GRCh38 Dec 2013

GRCh37 Feb 2009

NCBI Build 36.1 Mar 2006

NCBI Build 35 May 2004

NCBI Build 34 Jul 2003

The T2T consortium used Long-read shotgun sequencing (PacBio’s multi-


kilobase, single-molecule reads) and >100 kbp “ultra-long” reads (Oxford
Nanopore’s) which are proved to be capable of resolving complex structural
variation and gaps from GRCh38, and enabled complete assemblies of a
human centromere (ChrY), and entire chromosome (ChrX). The resulting
sequence (T2T-CHM13 reference assembly) removed a 20-year-old barrier
that has hidden a significant portion (8%) of the genome from sequence-based
analysis.

14.6.2 Human Genome Project: Implications for


Medical Science
Since the human genome sequence is completely available, the major issue is
how to understand and decode the information of the DNA sequence. Despite
numerous genome-wide studies that have been performed already, the main
challenge is to determine the function of individual genes, gene products, and
their interaction. Any changes in the human genome are more likely to cause
pathological conditions or disease, functional analysis is highly important to
draw implications for human health.

With the vast amount of data about the human genome (generated by the
Human Genome Project and other genomics research consortiums), scientists
and clinicians have more powerful high-end tools, to investigate the role that
multiple genetic factors, which act together with the environmental factors and
play major role in much more complex diseases. The diseases, such as
diabetes, cancer, and cardiovascular disease constitute the majority of health
92 problems in global relevance. Genome-based research is already enabling


Unit 14 Analysis of the Genome
scientists and clinicians to develop improved diagnostics, evidence-based
approaches for improving clinical efficacy, highly effective therapeutic
strategies, and better decision-making technologies for patients and
healthcare providers. Ultimately, it appears inevitable in the future that,
treatments will be tailored to a patient's particular genomic makeup
(personalized medicine/treatment). Thus, the role of genetics in healthcare is
changing profoundly and the human genome sequence has accelerated the
era of genomic medicine.

Fig. 14.2: Summary of the complete T2T-CHM13 human genome assembly.


A. Pictogram of human genome assembly (T2T-CHM13) B. Non-
syntenic (additional) bases in the CHM13 human genome assembly
with acrocentrics

14.7 GENOME ANALYSIS OF PLASMODIUM


FALCIPARUM
Plasmodium falciparum (Pf) is the most dangerous parasite that causes
malaria in humans. It is associated with high levels of mortality, and it has
widespread resistance to antimalarial drugs. In P. falciparum, little is known 93


Block 4 Applications of Genomics and Proteomics
about the genome organization and its’ effect on gene expression. Using a
whole chromosome shotgun sequencing strategy, researchers have
determined the genome sequence of P. falciparum (3D7 clone).

The chromosomes were separated on pulsed-field gel electrophoresis, and


chromosomal DNA was extracted, and the extracted DNA was used to
construct shotgun libraries of 1–3-kilobase (kb) fragments of sheared DNA.
The shotgun sequences were assembled into contigs (contiguous DNA
sequences), shotgun sequences of yeast artificial chromosome (YAC) clones,
Sequence tagged sites (STSs), microsatellite markers, and HAPPY
mapping are used to assist the ordering, orientation of contigs, and gap
closure. The predicted restriction maps of the chromosome sequences were
compared with optical restriction maps to confirm the correct chromosomal
assembly. The P. falciparum 3D7 (nuclear genome) is composed of 22.8
megabases (Mb), distributed into 14 chromosomes, with sizes ranging from ~
0.643 to 3.29 Mb (Fig. 14.3). The P. falciparum genome is A+T rich (overall
80.6%, and rises to ~90% in introns and intergenic regions).

P. falciparum genome codes for approximately 5,300 proteins, suggesting an


average gene density of 1 gene per 4,338 base pairs (bp). Introns were found
in 54% of P. falciparum genes, excluding introns. The average length of P.
falciparum genes is 2.3 kb, comparatively larger than any other organism’s
average gene lengths which range from 1.3 to 1.6 kb. The reason behind the
increased gene length in P. falciparum is not clearly understood. Most of these
large-length genes encode uncharacterized proteins (hypothetical) that may
be cytosolic localized proteins since they do not possess recognizable signal
peptides. In P. flaciparum genome no retrotransposons or transposable
elements were identified.

Fig. 14.3: The haploid genome of Plasmodium falciparum contains 14


chromosomes. The size of the Plasmodium falciparum genome is 23
94 Mb and contains ~5500 protein-coding genes


Unit 14 Analysis of the Genome

14.8 GENOME ANALYSIS OF MYCOBACTERIUM


TUBERCULOSIS
Tuberculosis is a chronic infectious disease caused by tubercle bacillus
(Mycobacterium tuberculosis). It is the biggest killer of humans, countless
people have died from tuberculosis. Mycobacterium tuberculosis (MTB) is
considered one of the deadliest pathogens, with a toll of 1 billion people
succumbed to tuberculosis (TB)-related deaths in the last two centuries. The
World Health Organization (WHO), 2019 Annual Report estimated about 10
million new clinical cases and 1.7 million deaths per year, making MTB the
deadliest single infectious agent before the SARS-CoV-2 caused COVID-19
pandemic.

Whole Genome Sequencing (WGS) offers an opportunity for the molecular


epidemiology surveillance of TB in selected regions around the globe. The
development of local capacities (laboratories), bioinformatic analysis tools, and
the reduction of overall genome sequencing costs make WGS increasingly
accessible as an alternative for molecular epidemiologic studies. Drug
resistance profiles can be predicted for the majority of the existing antibiotic
treatments using the genetic sequencing information gained by WGS in
comparison to other analytical procedures, allowing for widespread drug
resistance monitoring.

The complete genome sequence of Mycobacterium tuberculosis (H37Rv


strain), has been determined and analyzed to improve the understanding of
the pathogens’ biology and to help the development of new prophylactic and
therapeutic interventions. The Mycobacterium tuberculosis genome comprises
4,411,529 base pairs, contains ~4,000 genes, and has a very high guanine-
cytosine (GC) content (65%).

14.8.1 Sequence Analysis


A combined approach was used to obtain the contiguous genome sequence,
that involved the systematic sequence analysis of selected large-insert clones
(BACs and cosmids), and random small-insert clones from a whole-genome
shotgun sequence library. This resulted in a composite sequence of 4,411,529
base pairs (Fig. 14.4), with a G + C content of 65.6%. Mycobacterium
tuberculosis genome represents the second-largest available bacterial
genome sequence after that of Escherichia coli. The Mycobacterium
tuberculosis genome is highly rich in repetitive DNA, particularly the insertion
sequences, and duplicated housekeeping genes. The G + C content of the
Mycobacterium tuberculosis genome is relatively constant throughout.

14.8.2 Genes Encoding Proteins


In the Mycobacterium tuberculosis genome, 3,924 open reading frames were
identified accounting for ‫׽‬91% of the potential coding capacity. Few of the
Mycobacterium tuberculosis genes appear to have in-frame stop codons or
use frameshifting during translation. The high G + C content of the genome,
resulted in GTG initiation codons (35%) more frequently than in E. coli (14%),
and Bacillus subtilis (9%), although ATG (61%) is the most commonly used
translational start codon. In the Mycobacterium tuberculosis genome, there is 95


Block 4 Applications of Genomics and Proteomics
a slight bias in the orientation of the genes with respect to the direction of
replication, as ‫׽‬59% are transcribed with the same polarity as replication,
compared with 75% in B. subtilis. The even distribution of gene polarity
in Mycobacterium tuberculosis may reflect the slower growth and infrequent
replication cycles of this pathogen.

Fig. 14.4: Chromosome map of Mycobacterium tuberculosis (H37Rv strain)

In the circular chromosome map, the outer circle shows the size scale in Mb,
with “0” representing the origin of replication. The first circle from the exterior
represents the positions of stable RNA genes (tRNAs- blue, other RNAs-pink)
and the direct repeat region (pink cube); the second circle inwards represents
the coding sequence by strand (clockwise-dark green; anticlockwise-light
green); the third circle shows repetitive DNA (insertion sequences-orange;
13E12 REP family, dark pink; prophage-blue); the fourth circle depicts the
positions of the PPE family members (green); the fifth circle shows the PE
family members (purple-excluding PGRS). The histogram (centre) shows G +
C content, with <65% G + C in yellow, and >65% G + C in red.

Mycobacterium tuberculosis (MTB) genome sequencing by Whole genome


sequencing (WGS) has shown to be useful for clinically predicting drug-
resistance, tracing the transmission rates, and defining outbreaks. WGS can
predict resistance to the full range of anti-tuberculosis drugs accurately, does
not require biological safety infrastructure, and has the potential to predict drug
resistance more efficiently and quickly than traditional testing systems
[phenotypic susceptibility testing (pDST)]. Molecular epidemiology using WGS
has a higher resolution capacity than strain-typing techniques (IS6110-RFLP,
MIRU-VNTR, or sp oligo typing) due to its specificity that can identify single
nucleotide polymorphisms (SNPs) in strains that are identical. Few countries
including the UK and the Netherlands, have implemented WGS-guided,
individualized (personal) treatment, through WGS-based monitoring and
96 treatment of all tuberculosis patients.


Unit 14 Analysis of the Genome

14.9 MANIPULATION OF GENOMES: SITE-


DIRECTED MUTAGENESIS
Mutagenesis is a modification of DNA sequence in an organism, to understand
the functioning of regulatory regions of genes, and the relationship between
the protein structure and its function. Site-directed mutagenesis is a method to
generate cloned DNAs with modified sequences (with insertions, deletions,
and substitutions) for examining the importance of specific residues/regions in
protein structure and function. Site-directed mutagenesis (SDM) is the most
widely used technique in molecular biology to investigate the structures and
functions of nucleic acids and proteins, the molecular mechanisms of
diseases, and to study the effect of genome modification. Using Site-directed
mutagenesis researchers can create specific, targeted changes in double-
stranded DNA.

Their main aims to make specific DNA alterations (insertions, deletions, and
substitutions) are:

• To study the changes in protein activity that occur as a result of the DNA
manipulation (which further changes protein sequence).

• To screen or selection of mutations (at the DNA, RNA, or protein level)


that have a desired property or activity.

• To introduce or remove restriction endonuclease sites or tags in


plasmids for diverse applications.

Although gene mutations (insertions, deletions, and substitutions) can be


achieved by a variety of methods; Site-directed mutagenesis mediated by
PCR (polymerase chain reaction) is the most powerful approach to generate
mutations in genes in vitro. One of the simplest PCR-mediated methods for
site-directed mutagenesis is PCR using a pair of complementary primers with
desired mutations (a double-stranded DNA fragment), where the DNA
template (usually a plasmid) after PCR is required to be digested by the DpnI
restriction enzyme before transformation (Fig. 14.5).

Depending on the number of sites in the DNA to be mutated, site-directed


mutagenesis can be divided into two types: simple or multiple mutations. For
simple or single mutations, SDM methods that are PCR-based are used which
involves amplification of double-stranded DNA from plasmids using
complementary (inverse) oligonucleotides (primers) carrying the mutation of
interest. Due to its simple protocol, less time is spent, and have high
efficiency. This is one of the most common and highly used strategies to
introduce mutations in DNA fragments.

14.9.1 Protocol: Site-directed Mutagenesis using PCR


To avoid the use of both forward and reverse primers in the same PCR
reaction, scientists have developed an SDM protocol that involves five steps
(Fig. 14.5) as given below:

1. Amplify the parental plasmid containing cDNA insert independently in


two separate PCR reactions using either forward or reverse single
primers. 97


Block 4 Applications of Genomics and Proteomics
2. Combine the two independent-primer PCR products from each reaction
in one test tube and denature at 95°C to separate the newly synthesized
DNA from the template DNA (plasmid).

3. Cool the tube gradually to 37°C to allow annealing of the newly


synthesized complementary strands.

4. Digest methylated non-mutated DNA (parental plasmid) with DpnI


restriction enzyme.

5. Perform transformation of the reannealed plasmids into competent E.


coli bacterial cells.

98 Fig. 14.5: Method of Site Directed Mutagenesis (PCR based method)


Unit 14 Analysis of the Genome

6$4
6$4
Choose the correct answer from the multiple options given below each
question.

a) The technique used to determine the sequence of DNA is

i) PCR (Polymerase Chain Reaction)

ii) Microarray analysis

iii) Sanger sequencing

iv) Gel electrophoresis

b) Mapping of the human genome involves

i) Sequencing multiple variations of each gene

ii) To study the entire DNA found in human cells

iii) To study the heterochromatin areas

iv) All of the above

c) Initial version of human genome was released in the year

i) 2003

ii) 1990

iii) 1998

iv) 2000

d) Commonly used vectors for human genome sequencing are

i) T-DNA

ii) BAC and YAC

iii) Expression vectors

iv) T/A cloning vectors

e) The sources of biomedical and genomic information are

i) NCBI (National Center for Biotechnology Information)

ii) EMBL (European Molecular Biology Laboratories)

iii) DDBJ (DNA Data Bank of Japan)

iv) All of the above

14.10 SUMMARY
• Genome analysis of an organism’s whole genome can provide many
predictions about diagnosis, or susceptibilities to conditions, and offers
the possibility of identifying all putative protein-encoding genes of a
given organism. 99


Block 4 Applications of Genomics and Proteomics
• The recent advances in gene sequencing and annotation technology
have allowed high throughput genomic sequencing to be done quickly
and relatively cheaply which propelled the work of genome analysis
forward.

• Precision medicine (PM) or personalized medicine is a concept where


disease treatment is based on knowing an individual's genome
abnormalities, and the causative pathogen’s acquired qualities, which
became relevant as whole-genome sequencing became more
accessible.

• Genome analysis has opened new avenues for developing new


therapeutic strategies, drug development, and gene-oriented treatment.

• Site-directed mutagenesis (SDM) is an invaluable tool to modify genes


and study the structural and functional properties of proteins, based on
their structure, function, and catalytic mechanism.

• SDM is the most widely used technique in molecular biology to


investigate the structures and functions of nucleic acids and proteins, the
molecular mechanisms of diseases, and to study the effect of genome
modification.

• Mutagenesis has multiple implications in agriculture, industry, and


clinical medicine.

14.11 TERMINAL QUESTIONS


1. What is genome and genome analysis? What are the major steps
involved in genome analysis?

2. What are the implications of the human genome project on medical


sciences?

3. Write a note on the genome analysis of Plasmodium falciparum.

4. How sequencing a pathogenic organism's DNA can help in controlling


the disease?

5. Describe the steps involved in Site-directed mutagenesis by PCR.

14.12 ANSWERS
Self-Assessment Questions
1. a) i) C) Whole genomes

ii) C) Sanger sequencing

b) i) mutagenesis, ii) gene, iii) polyploidization, iv) reference,


v) data collection.

2. a) iii) Sanger sequencing

100 b) i) Sequencing multiple variations of each gene


Unit 14 Analysis of the Genome
c) i) initial version published in 2003

d) ii) BAC (Bacterial artificial chromosome) and YAC (Yeast


artificial chromosome)

e) iv) All of the above

Terminal Questions
1. Refer to Sections 14.1, 14.2.

2. Refer to Subsection 14.6.2.

3. Refer to Section 14.7.

4. Refer to Section 14.8.

5. Refer to Section 14.9.

Acknowledgement of Figures
Fig 14.2: S. Nurk et al., SCIENCE, Vol 376, Issue 6588, pp. 44-53, 2022,
The complete sequence of a human genome, DOI:
10.1126/science.abj6987

101




UNIT 15
0$1,38/$7,212)7+(
*(120(
*(120(

6WUXFWXUH
6WUXFWXUH
15.1 Introduction Frequency of Occurrence of
Restriction Sites in DNA
Objectives
Restriction Sites and Creation
15.2 Converting Genomes into
of Recombinant DNA
Clones, and Clones into
Molecules
Genomes
15.5 Cloning Vectors and DNA
DNA Cloning
Cloning
Reporter Gene
Plasmid Cloning Vectors
15.3 Gene Knockout Method in
Artificial Chromosomes
Transgenics
15.6 Genomic Libraries
Gene Knockouts in Yeast
15.7 Chromosome Libraries
Gene Knockouts in the Mouse
15.8 DNA Sequencing and
Gene Knockin Method in
Analysis of DNA Sequences
Transgenics
15.9 Summary
Gene Knockin in the Mouse
15.10 Terminal Questions
Knockin Mice Mutations
15.11 Answers
15.4 Restriction Enzymes
General Properties of
Restriction Enzymes

15.1 INTRODUCTION
Genomics is the science of obtaining and analyzing the sequences of
complete genomes. At the core of genomics is recombinant DNA technology,
the ability to construct and clone individual fragments of a genome, and to
manipulate the cloned DNA in various ways, including expressing it in a
foreign cell. The development of molecular techniques for analyzing genes
and gene expression has revolutionized experimental biology. The present
Unit describes the techniques used for manipulation of genome.
Unit 15 Manipulation of the Genome

2EMHFWLYHV
After studying this unit you would be able to:

™ describe different steps of DNA cloning,

™ define the restriction enzymes,

™ comprehend the mechanism of action of the restriction enzymes to


cleave DNA,

™ understand the different types of vectors constructed for cloning DNA,


and

™ explain the Pyrosequencing.

15.2 CONVERTING GENOMES INTO CLONES,


AND CLONES INTO GENOMES
To study a genome, it must be broken into much smaller fragments that can
be worked within the lab, and you need to use an easily cultured host cell,
such as the easy-to-handle and manipulate microorganisms such as
Escherichia coli or yeast, to take up and maintain these small fragments so
that you can isolate many thousands of identical copies of each fragment. A
physical map of the genome is needed to be made; that is, a map of the
chromosomes showing the positions of important landmarks like genes and
promoters, as well as specific DNA base pairs, sequences, and regions that
vary between individuals. In a physical map, distances are measured in base
pairs. To make a physical map, you must determine where these landmarks
come from in the intact genome. This means taking the small fragments and
then reassembling a “virtual chromosome” from them. The first step is to
construct a genomic library, a collection of clones that contains at least one
copy of every DNA sequence in the genome of an organism. Since most
genomes contain millions or billions of base pairs, and a clone contains a
relatively small piece of DNA, genomic libraries must have many clones
(thousands to millions), with each clone containing a random small fragment of
genomic DNA carried by a cloning vector which is an artificially constructed
DNA molecule capable of replication in a host organism such as a bacterium.
A cloning vector allows us to make many copies of the small fragment of
genomic DNA.

15.2.1 DNA Cloning


In brief, DNA cloning can be explained by the following steps:

1. Isolate DNA from an organism.

2. Cut the DNA into pieces with a restriction enzyme, (an enzyme that
recognizes and cuts within a specific DNA sequence) and insert (ligate)
each piece individually into a cloning vector, that is, cut with the same
restriction enzyme to construct a recombinant DNA molecule, (a DNA
molecule constructed in vitro containing sequences from two or more
distinct DNA molecules). 103


Block 4 Applications of Genomics and Proteomics

3. Introduce the recombinant DNA molecules into a host such as E. coli


through transformation. Replication of the recombinant DNA molecule
(the process of molecular cloning) occurs in the host cell, producing
many identical copies called clones. As the host organism reproduces,
the recombinant DNA molecules are passed on to all the progeny, giving
rise to a population of cells carrying the cloned sequences.

15.2.2 Reporter Gene

An enzyme or fluorescent protein encoded by a nonendogenous gene whose


expression is regulated by a promoter for an unrelated gene of interest is
known as a reporter gene. In more complex organisms, numerous DNA
regulatory sequences function as enhancers to regulate transcription. For
example, in model organisms like Drosophila, researchers can develop
screens to search for these enhancers. The regulatory elements boost the
transcription of nearby genes, so a transgenic reporter construct can be
inserted that can be activated by any nearby enhancer. The construct contains
a transcription start site and a "reporter" gene, such as green fluorescent
protein (GFP) or blue dye-producing galactosidase, and is carried on a
transposon. Through crosses that mobilize this reporter construct, it
transposes itself into various sites in the genome, enabling the observation of
the distribution of the reporter protein product. This approach helps identify the
locations of enhancer elements that drive a specific pattern of gene
expression. The resulting reporter-transgene insertions, known as enhancer,
traps and reveal insights such as the expression of a particular reporter
insertion in developing Drosophila eye tissue only. This suggests that a gene
expressed in the eye likely resides nearby, making neighboring genes
potential candidates for involvement in eye development and worthy of
isolation and study.

15.3 GENE KNOCKOUT METHOD IN


TRANSGENICS
Gene-targeting techniques are used to create knock-in and knock-out
organisms. In knock-in animals novel externally supplied genes are
expressed, whereas in knockouts native genes are deleted or suppressed in
order to experimentally determine its function and the resulting phenotypic
changes. Major projects have been focused on systematically eliminating the
function of each gene in various organisms, including yeast, mouse, fruit fly,
Mycoplasma genitalium, and the nematode worm Caenorhabditis elegans.
There are various methods to disrupt the functions of protein-coding genes,
with gene knockouts and RNA interference (RNAi) being the most commonly
used techniques. Gene knockouts involve disrupting the gene on the
chromosome, and you will explore strategies for knocking out chromosomal
genes in yeast, mouse, and M. genitalium. On the other hand, RNA
interference (RNAi), also known as RNA silencing, utilizes small regulatory
104 RNAs to silence gene expression in eukaryotes.


Unit 15 Manipulation of the Genome
15.3.1 Gene Knockouts in Yeast
In yeast, gene function can be deactivated using a polymerase chain reaction
(PCR)-based approach, where PCR primers are designed based on the
known genome sequence to create an artificial linear DNA deletion module,
containing the gene sequence flanked by a selectable marker such as the
kanR (kanamycin) marker for resistance to a specific chemical G418.
Essentially, the kanR marker replaces a major portion of the gene of interest's
coding region, rendering the gene unable to produce its protein. When this
linear DNA is introduced into yeast, colonies resistant to G418 are chosen.
Unlike the previously mentioned plasmids, this linear DNA fragment cannot
replicate in the host cell due to the absence of an origin of replication. Through
a process known as homologous recombination, the linear plasmid integrates
into the yeast chromosome, effectively knocking out the chromosomal copy of
the gene of interest because the selectable marker replaces most of the
coding region.

15.3.2 Gene Knockouts in the Mouse


The mouse is a valuable organism for genetic research due to its genetic
similarities to humans and its suitability for lab studies involving genetic
techniques. For instance, gene knockouts in mice are used to investigate the
functions of mouse genes that are similar to unknown human genes, and to
explore fundamental questions about mammalian biology. The procedure for
knocking out the target gene in mouse cells is somewhat similar to that used
for yeast, but more complex. First, a cloned copy of the target gene is modified
to replace a central portion with a selectable marker, neoR, which enables
mouse cells to grow on neomycin. Then, a segment of DNA containing a
second selectable marker, tk, is added to the modified gene. When the
chemical ganciclovir is introduced to mouse cells expressing the tk gene, their
growth is inhibited, as the thymidine kinase phosphorylates ganciclovir, turning
it into an inhibitory chemical for DNA replication. The final product is the target
vector, which contains the disrupted target gene and the two selectable
markers.

15.3.3 Gene Knockin Method in Transgenics


Gene-targeting techniques are used to create knockin organisms, which are
organisms with a specific gene mutation introduced (i.e., knocked in). The
main goal of creating a knock-in model is to produce an animal with an allele
that has either enhanced or modified function. The creation of a knock-in
animal model typically involves inserting a specific DNA sequence into the
genome at a particular location, which can either replace an existing DNA
region or introduce a new DNA segment.

15.3.4 Gene Knockin in the Mouse


This approach is frequently employed to enable the expression of a specific
gene under the direct regulation of natural regulatory components. In this
instance, the first coding exon of the mouse ortholog has the insertion of a
human cDNA, followed by the mouse 3' untranslated region, and a
polyadenylation (pA) signal. The targeting strategy is intended to facilitate
human protein expression using mouse regulatory elements and to prevent
mouse protein expression. 105


Block 4 Applications of Genomics and Proteomics
15.3.5 Knockin Mice Mutations
Constitutive knockin mice: This particular model has been designed to
transport a cDNA sequence that encodes a protein to a specific location.

Point mutation knockin mice: This model involves substituting one or a few
DNA bases within the sequences of a designated gene.

Humanized mice: This model includes the replacement of murine gene by its
human counterpart.

15.4 RESTRICTION ENZYMES


To analyze genomic DNA, you must first cut it into smaller, more manageable
pieces. The tools for this are restriction enzymes. A restriction enzyme (or
restriction endonuclease) recognizes a specific nucleotide-pair sequence in
DNA called a restriction site and cleaves the DNA (hydrolyzes the
phosphodiester backbones) within or near that sequence. All restriction
enzymes cut DNA between the 3ƍ carbon and the phosphate moiety of the
phosphodiester bond so that fragments produced by restriction enzyme
digestion have 5ƍ phosphates and 3ƍ hydroxyls. Most restriction enzymes
function optimally at 37oC. Figure 15.1 shows the working principle of
restriction enzyme.

Restriction enzymes are used to produce a pool of DNA fragments to be


cloned. Restriction enzymes are also used to analyze the positions of
restriction sites in a piece of cloned DNA or in a segment of DNA in the
genome. Laboratories use of restriction enzyme digestions in an attempt to
“cut to completion” meaning that the enzyme is allowed to cut at each of its
restriction sites in the DNA. Such a digest will cut each genome copy of the
same organism into the same large set of pieces. As you will see, in certain
genomics applications it is desirable, instead, to do a “partial digest” in which
the enzyme does not have enough time to complete its task. As a result, only
some of the restriction sites are cut, and many are left uncut. As you are
cutting millions of identical DNA molecules, in a partial digest each will be cut
at a unique subset of the available restriction sites.

15.4.1 General Properties of Restriction Enzymes


Most restriction enzymes are found naturally in bacteria, although a handful
have been found in eukaryotes. In bacteria, restriction enzymes protect the
host organism against viruses by cutting up i.e. restricting the invading viral
DNA. The bacterium modifies its own restriction sites (by methylation) so that
its own DNA is protected from the action of the restriction enzyme(s) it makes.
Werner Arber, Daniel Nathans, and Hamilton O. Smith received the 1978
Nobel Prize in Physiology or Medicine “for their discovery of restriction
enzymes and their application in solving the problems of molecular genetics”.

More than 400 different restriction enzymes have been isolated, and at least
2,000 more have been characterized partially. They are named for the
organisms from which they are isolated. Conventionally, a three-letter system
106 is used. Commonly the first letter is that of the genus, and the second and


Unit 15 Manipulation of the Genome
third letters are from the species name. The letters are italicized or underlined,
followed by roman numerals that signify a specific restriction enzyme from that
organism. Additional letters sometimes are added just before the number to
signify a particular bacterial strain from which the enzymes were obtained. For
example, EcoRI and EcoRV are both from Escherichia coli strain RY13, but
recognize different restriction sites; HindIII is from Haemophilus influenzae
strain Rd. The Roman numerals indicate the order in which the restriction
enzymes from that strain were identified. Hence, EcoRI and EcoRV are the
first and fifth restriction enzymes identified for E. coli strain RY13. The names
are pronounced in ways that follow no set pattern. For example, BamHI is
“bam-H-one,” BglII is “bagel-two,” EcoRI is “echo-R-one” or “eeko-R-one,”
HindIII is “hin-D-three,” HhaI is “ha-ha-one,” and HpaII is “hepa-two.”

Fig. 15.1: Working principle of restriction enzymes

Many restriction sites have an axis of symmetry through the midpoint. Figure
15.2 shows this symmetry for the EcoRI restriction site: the nucleotide
sequence from 5ƍ to 3ƍ on one DNA strand is the same as the nucleotide
sequence from 5ƍ to 3ƍ on the complementary DNA strand. Thus, the
sequences are said to have two fold rotational symmetry. A number of
restriction sites are shown in Table 15.1. The most commonly used restriction
enzymes recognize four nucleotide pairs (for example, HhaI) or six nucleotide
pairs (for example, BamHI, EcoRI). Some enzymes recognize eight-nucleotide
pair sequences (for example, NotI [“not-one”]). Other classes of enzymes do
not fit our model because the restriction site is not symmetrical about the
center. HinfI (“hin-f-one”), for example, recognizes a five-nucleotide pair
sequence in which there is symmetry in the two nucleotide pairs on either side
of the central nucleotide pair, but the central nucleotide pair is obviously
asymmetrical within the sequence. BstXI (“b-s-t-x-one”) is representative of a
number of restriction enzymes with a nonspecific spacer region between
symmetrical sequences. 107


Block 4 Applications of Genomics and Proteomics
Table 15.1: Characteristics of some restriction enzymes

Fig. 15.2: Restriction site in DNA, showing the two fold rotational symmetry of
the sequence. The sequence reads the same from left to right (5ƍ to 3ƍ)
on the top strand (GAATTC, here) as it does from right to left (5ƍ to 3ƍ)
108 on the bottom strand. Shown is the restriction site for EcoRI


Unit 15 Manipulation of the Genome
15.4.2 Frequency of Occurrence of Restriction
Sites in DNA
Since each restriction enzyme cuts DNA at an enzyme-specific sequence, the
number of cuts the enzyme makes in a particular DNA molecule depends on
the number of times that particular restriction site occurs. When you cut a
number of copies of the same genome with a particular restriction enzyme, the
DNA is cleaved at the specific restriction sites by that enzyme, which are
distributed throughout the genome. Although this produces millions of
fragments of different sizes from one genome copy, all copies of the same
genome will be cut at identical places.

Based on probability principles, the frequency of a short nucleotide pair


sequence in the genome theoretically will be greater than the frequency of a
long nucleotide pair sequence, so an enzyme that recognizes a four-
nucleotide pair sequence will cut a DNA molecule more frequently than one
that recognizes a six-nucleotide pair sequence, and both enzymes will cut
more frequently than one that recognizes an eight-nucleotide pair sequence.

15.4.3 Restriction Sites and Creation of Recombinant


DNA Molecules
One major class of restriction enzymes recognizes a specific DNA sequence
and then cuts within that sequence. Another class of restriction enzymes
recognize a specific nucleotide-pair sequence, and then cut the two strands of
DNA outside of that sequence. This latter class of restriction enzymes is not
useful for creating recombinant DNA molecules and will not be considered
further.

Restriction enzymes in the first class cut DNA in different ways. As Table 15.1
indicates, some enzymes, such as SmaI (“sma-one”), cut both strands of DNA
between the same two nucleotide pairs to produce DNA fragments with blunt
ends (Fig. 15.3a). Other enzymes, such as BamHI, make staggered cuts in the
symmetrical nucleotide-pair sequence to produce DNA fragments with sticky
or staggered ends, either 5ƍ overhanging ends, as in the case of cleavage with
BamHI (Fig. 15.3b) or EcoRI, or 3ƍ overhanging ends, as in the case of
cleavage with PstI (“P-S-T-one”; Fig. 15.3c).Restriction enzymes that produce
sticky ends are of particular value in cloning DNA because every DNA
fragment generated by cutting a piece of DNA with the same restriction
enzyme has the same single-stranded nucleotide sequence at the two
overhanging ends. If the ends of two pieces of DNA produced by the action of
the same restriction enzyme (such as EcoRI)—a cloning vector and a
chromosomal DNA fragment, for example, come together in solution, base
pairing occurs between the overhanging ends; the two single-stranded DNA
ends are said to anneal (Fig. 15.4). Using DNA ligase, the two DNAs can be
covalently linked (ligated) to produce a longer DNA molecule with the
restriction sites reconstituted at the junction of the two fragments. Even DNA
fragments with blunt ends can be ligated together by DNA ligase at high
concentrations of the enzyme. The ligation of two DNA fragments is the
principle behind the formation of recombinant DNA molecules. Paul Berg
received part of the 1980 Nobel Prize in Chemistry “for his fundamental
studies of the biochemistry of nucleic acids, with particular regard to
recombinant-DNA.” 109


Block 4 Applications of Genomics and Proteomics

Fig. 15.3: Examples of how restriction enzymes cleave DNA. (a) SmaI results in
blunt ends, (b) BamHI results in 5ƍ overhanging (“sticky”) ends, (c) PstI
results in 3ƍ overhanging (“sticky”) ends.

Fig. 15.4: Cleavage of DNA by the restriction enzyme EcoRI. EcoRI makes
staggered, symmetrical cuts in DNA, leaving “sticky” ends. A DNA
fragment with a sticky end produced by EcoRI digestion can bind by
complementary base pairing (anneal) to any other DNA fragment with
a sticky end produced by EcoRI cleavage. The nicks can then be
110 sealed with DNA ligase


Unit 15 Manipulation of the Genome

6$4
6$4
a) Answer in one word only:

i) Restriction enzymes EcoRI and EcoRV are obtained from which


bacterial strain?

ii) How does restriction enzyme SmaI cleave DNA?

iii) In which technique specific gene is deleted in order to


experimentally determine its function and the resulting phenotypic
changes?

iv) The creation of which model typically involves inserting a specific


DNA sequence into the genome at a particular location, which can
either replace an existing DNA region or introduce a new DNA
segment?

b) Fill in the blanks with appropriate words:

i) An enzyme or fluorescent protein encoded by a nonendogenous


gene whose expression is regulated by a promoter for an
unrelated gene of interest is known as a ……………... .

ii) An enzyme that recognizes and cuts within a specific DNA


sequence is known as ……………... .

iii) Restriction enzyme HindIII is obtained from ……………... .

iv) Most restriction enzymes function optimally at ……………... .

v) The ……………... technique also known as RNA silencing, utilizes


small regulatory RNAs to silence gene expression in eukaryotes.

15.5 CLONING VECTORS AND DNA CLONING


To determine the sequence of a genome, you need to break the genome into
fragments and clone each fragment to produce multiple copies to be used for
DNA sequencing. Several types of vectors have been constructed specially for
cloning DNA. They include plasmids, bacteriophages (e.g., Ȝ and certain
single-stranded DNA species), cosmids (vectors with features of both plasmid
and bacteriophage vectors), and artificial chromosomes. The vector types
differ in their molecular properties and in the maximum amount of inserted
DNA they can hold. Each type of vector has been specially constructed in the
laboratory. In this section, we will focus on plasmid and artificial chromosome
vectors, as they have been work horses in genomics.

15.5.1 Plasmid Cloning Vectors


Bacterial plasmids are extra chromosomal elements that replicate
autonomously within cells. Plasmid DNA is double-stranded and often circular
and contains an origin sequence (ori) required for plasmid replication and 111


Block 4 Applications of Genomics and Proteomics
genes for the other functions of the plasmid. Plasmid cloning vectors are
derivatives of circular natural plasmids “engineered” to have features useful for
cloning DNA. We focus here on features of E. coli plasmid cloning vectors.

An E. coli plasmid cloning vector must have three features:

i) An ori (origin of DNA replication) sequence, is needed for the plasmid to


replicate in E. coli.

ii) A selectable marker, so that E. coli cells with the plasmid can be
distinguished easily from cells that lack the plasmid. A selectable marker
is a gene that allows us to determine easily if a cell does or does not
contain the cloning vector. For bacterial plasmid cloning vectors, typically
the selectable marker is a gene for resistance to an antibiotic, such as
the ampR gene for ampicillin resistance or the tetR gene for tetracycline
resistance. When plasmids carrying antibiotic-resistance genes are
added to a population of plasmid-free and therefore antibiotic sensitive
E. coli, the cells that take up the plasmid can be selected for by culturing
the cells on a solid medium containing the appropriate antibiotic. Only
bacteria with the plasmid will grow on the medium.

iii) One or more unique restriction enzyme cleavage sites, that is, the sites
present just once in the vector for the insertion of the DNA fragments to
be cloned. Typically, a number of sites are present in the vector, and
these sites tend to be engineered as a multiple cloning site or polylinker.
A multiple cloning site is a region of DNA containing several unique
restriction sites where a fragment of foreign DNA (not originally part of
the vector) can be inserted into the vector. With a number of different
sites available in the multiple cloning site of a vector, an investigator can
use the same vector in different cloning experiments by choosing
different restriction sites for the cloning purpose.

As an example, Fig. 15.5 shows the plasmid cloning vector pBluescript II. This
2,961 bp vector has the following features that make it useful for cloning DNA
in E. coli:

i) It has a high copy number, approaching 300-500 copies per cell


because it has a very active ori. As a result, many copies of a cloned
piece of DNA can be generated readily in a small number of host cells.

ii) It has the ampR selectable marker for ampicillin resistance.

iii) It has a multiple cloning site containing 18 restriction sites.

iv) The multiple cloning site is embedded in part of the E. coli ȕ-


galactosidase (lacZ+) gene (Fig. 15.5). pBluescript II, like other plasmids
constructed with such a lacZ gene fragment, is usually introduced into an
E. coli strain with a mutated lacZ gene. When the (unmodified) plasmid
is present in the cell, functional ȕ-galactosidase is produced. However,
when a piece of DNA is cloned into the multiple cloning site, the lacZ
fragment on the plasmid is disrupted and no functional ȕ-galactosidase
can be produced. Therefore, the presence or absence of ȕ-
galactosidase activity indicates whether the plasmid was introduced into
112 E. coli is the empty pBluescript II vector (no inserted DNA fragment:


Unit 15 Manipulation of the Genome
functional enzyme present) or pBluescript II with an inserted DNA
fragment (functional enzyme absent).The chemical X-gal, a colorless
artificial substrate for ȕ-galactosidase is included in the medium on
which the cells containing plasmids are plated as an indicator for ȕ-
galactosidase activity in cells of a colony. Cleavage of X-gal by ȕ-
galactosidase leads to the production of a blue dye. Thus, if a functional
enzyme is present (vector with no insert), the colony turns blue, whereas
if nonfunctional ȕ-galactosidase is made (vector with inserted DNA), the
colony is white. This protocol is called blue-white colony screening.

Fig. 15.5: The plasmid cloning vector pBluescript II. This plasmid cloning vector
has an origin of replication (ori), a selectable marker, and a multiple
cloning site located within part of the ȕ-galactosidase gene lacZ+

Figure 15.6 illustrates how a piece of DNA can be inserted into a plasmid
cloning vector such as pBluescript II. In the first step, pBluescript II is cut with
a restriction enzyme that has a site in the multiple cloning site. Next, the piece
of DNA to be cloned is generated by cutting high-molecular-weight DNA with
the same restriction enzyme. Since restriction sites are non uniformly arranged
in DNA, fragments of various sizes are produced. The DNA fragments are
mixed with the cut vector in the presence of DNA ligase; in some cases, the
DNA fragment becomes inserted between the two cut ends of the plasmid and
DNA ligase joins the two molecules covalently. The resulting recombinant
DNA plasmid is introduced into an E. coli host by transformation. This is done
either by incubating the recombinant DNA plasmids with E. coli cells treated
chemically (such as with CaCl2) to take up DNA, or by electroporation, a 113


Block 4 Applications of Genomics and Proteomics
method in which an electric shock is delivered to the cells, causing temporary
disruptions of the cell membrane to let the DNA enter. Transformed cells are
plated onto media containing ampicillin and X-gal. Cells that can grow and
divide on this medium, forming a colony, must have been transformed by a
plasmid. Colonies containing plasmids with an insert can be identified by the
blue–white colony screening method.

In a ligation reaction, the restriction enzyme-digested vector alone can


recircularize. Such recircularization is quite common because it is a reaction
involving only one DNA molecule, and thus more likely than ligation of two
DNA molecules, such as vector and insert. This can make it more difficult to
find the desired recombinant plasmids from amongst all the plasmids.
Fortunately, vector recircularization can be minimized by treating the digested
vector with the enzyme alkaline phosphatase to remove the 5ƍ phosphates,
leaving a 5ƍ–OH group at the two ends of the DNA. DNA ligase can only join a
3ƍ–OH to a 5ƍ–phosphate, so if you remove both phosphates from a vector, it
cannot recircularize. DNA to be inserted into the vector, that is, the insert DNA,
is not treated with phosphatase, so the insert DNA retains 5ƍ phosphate
groups and the 5ƍ ends of the insert DNA can be ligated to the 3ƍ ends of the
vector DNA. This ligation reaction creates a circular molecule with two nicks
where the phosphodiester backbone is broken but, since these nicks are far
apart, the complex holds together as a single molecule. If the digested vector
is treated with alkaline phosphatase before the ligation reaction, then, the
proportion of blue colonies among transformants is reduced drastically. In
other words, the alkaline phosphatase treatment makes the identification of
the desired clones more efficient.

Fig. 15.6: Insertion of a piece of DNA into the plasmid cloning vector pBluescript
II to produce a recombinant DNA molecule. The vector pBluescript II
contains several unique restriction enzyme sites localized in a multiple
cloning site that are convenient for constructing recombinant DNA
molecules. The insertion of a DNA fragment into the multiple cloning
site disrupts part of the ȕ-galactosidase (lacZ+) gene, leading to
nonfunctional ȕ-galactosidase in E. coli. The blue-white colony
screening method described in the text can be used to identify vectors
with or without inserts

DNA fragments of up to 15 kb may be cloned efficiently in E. coli plasmid


cloning vectors. Plasmids carrying larger DNA fragments often are unstable in
vivo and tend to lose most of the insert DNA. This size limitation means that
114 the plasmid vectors are of limited use in genomic analysis, since millions of


Unit 15 Manipulation of the Genome
clones would be needed to contain a single genome of a complex multicellular
organism such as a human. To clone larger DNA inserts, different vectors are
used such as cosmids and artificial chromosomes. A cosmid can
accommodate DNA inserts in the range of 40-45 kb for genomics uses. A
cosmid cloning vector is similar to a plasmid cloning vector, with an origin, a
drug resistance marker, and a multiple cloning site, but it is introduced into
host cells differently. Cosmids are frequently used as vectors when libraries
are made because they are able to hold larger inserts.

15.5.2 Artificial Chromosomes


Artificial chromosomes are cloning vectors that can accommodate very large
pieces of DNA, producing recombinant DNA molecules resembling small
chromosomes. Artificial chromosomes are useful in genomics applications
because you can use them to study large segments of chromosomes, and
they can contain an entire genome in a manageable number of clones. We
consider two examples here, bacterial artificial chromosomes and yeast
artificial chromosomes.

Bacterial Artificial Chromosomes

Bacterial artificial chromosomes (BACs) are cloning vectors containing the


origin of replication from a natural plasmid found in E. coli called the F factor, a
multiple cloning site, and one or more selectable markers. One BAC vector,
pBeloBAC11, is shown in Figure 15.7a. This particular vector can be used with
the blue-white colony screening method, just like a plasmid. The selectable
marker for this BAC is camR. This gene encodes an enzyme that degrades the
antibiotic chloramphenicol, and thus, cells carrying this vector (with or without
an insert) can grow in the presence of chloramphenicol while cells lacking this
vector are unable to grow if chloramphenicol is present. BACs accept inserts
up to 300 kb and have the advantage that they can be manipulated like giant
bacterial plasmids. One major difference between BACs and the plasmids you
have already learned about is that once transformed into E. coli, the F factor
origin of replication keeps the copy number of the BAC at one per cell, while
the origins of typical plasmid cloning vectors drive multiple rounds of DNA
replication to generate many copies of the plasmid in each cell. Unlike yeast
artificial chromosomes (YACs) that will be described next, BACs do not
undergo rearrangements in the host. Therefore, they have become the
preferred vector for making large clones in physical mapping studies of
genomes. Two disadvantages of BACs (and with other cloning vectors for E.
coli) are that AT-rich DNA fragments (DNA fragments with a high proportion of
A and T nucleotides) typically do not clone well, and some DNA sequences
are toxic to E. coli and, hence, are unclonable in that organism.

Yeast Artificial Chromosomes

Yeast artificial chromosomes (YACs) are cloning vectors that enable artificial
chromosomes to be made and replicated in yeast cells. YAC vectors can
accommodate DNA fragments that are several hundred kilobase pairs long,
much longer than the fragments that can be cloned in the plasmid, cosmid, or
BAC vectors. Therefore, YAC vectors have been used to clone very large DNA
fragments (between 0.2 and 2.0 Mb), for example, in creating physical maps of 115


Block 4 Applications of Genomics and Proteomics
large genomes such as the human genome. A YAC (shown in its linear form)
has the following features (Fig. 5.7b):

1. A yeast telomere (TEL) at each end. (Recall that all eukaryotic


chromosomes need a telomere at each end.)

2. A yeast centromere sequence (CEN) allowing regulated segregation


during mitosis.

3. A selectable marker on each arm for detecting and maintaining the YAC
in yeast (for example, TRP1 and URA3 to enable transformed trp1
[tryptophan requiring] ura3 [uracil requiring] mutant yeast to grow on a
medium lacking tryptophan and uracil).

4. An origin of replication sequence—ARS (autonomously replicating


sequence) that allows the vector to replicate in a yeast cell.

5. An origin of replication (ori) that allows a circular version of the empty


vector to replicate in E. coli, and a selectable marker such as ampR that
functions in E. coli.

6. A cloning region that contains one or more restriction sites; the


restriction enzymes cutting in this region should not have any other sites
in the YAC. This region is used for inserting foreign DNA.

There are two disadvantages associated with these very large YAC-based
clones. First, during the cloning process, a fraction of the YAC vectors accept
two or more inserts, rather than one, creating a chimeric YAC. A second
problem is that portions of the insert DNA are frequently deleted or otherwise
modified by the host cell, or undergo recombination with other DNA in the host
cell. The altered inserts in chimeric and rearranged YACs will confound the
assembly of the genome, because assembly requires that we compare how
different inserts in our library overlap. The alterations in these inserts will
cause us to misinterpret how they overlap with other clones, because a
chimeric clone might contain, for instance, DNA from chromosome 5 ligated to
DNA from chromosome 18. Determining which YACs are modified is often a
very slow and labor-intensive process, making the assembly of a genome
sequence more difficult.

Empty YAC vectors, ones that have yet to contain a DNA insert are
propagated in E. coli as circular plasmids; in this form the two telomeres are
end-to-end. This propagation step makes use of the bacterial origin of
replication and the bacterial selectable marker. Bacterial and eukaryotic
origins of replication are not functionally similar, which means that the yeast
ARS sequence will not work in a bacterial cell, just as the bacterial ori
sequence will not function in a yeast cell. In addition, bacterial and eukaryotic
promoters are different, meaning that the bacterial RNA polymerase cannot
transcribe the yeast TRP1 and URA3 genes, so those selectable markers will
function only in yeast, not in bacteria. Likewise, yeast RNA polymerase II is
unable to transcribe the ampR gene.

For cloning experiments, a circular YAC is cut with one restriction enzyme that
cuts in the multiple cloning site and with another restriction enzyme that cuts
116 between the two TELs. In this way, the left and right arms are produced. High-


Unit 15 Manipulation of the Genome
molecular-weight DNA, cut with the same restriction enzyme used to cut the
YAC multiple cloning site, is ligated to the two arms and the recombinant
molecules are transformed into yeast. By selecting for both TRP1 and URA3, it
can be ensured that the transformants have both the left and right arms.

Fig. 15.7: Examples of artificial chromosome cloning vectors. (a) A BAC


(bacterial artificial chromosome) vector, such as pBeloBAC11, is
similar to a plasmid vector, with one or more selectable markers (here,
camR for chloramphenicol resistance), a multiple cloning site present
in part of the lacZ+ gene, uses an origin derived from the F factor,
which limits the copy number of the BAC to one per E. coli cell. (b) A
YAC (yeast artificial chromosome) vector contains a yeast telomere
(TEL) at each end, a yeast centromere sequence (CEN), a yeast
selectable marker for each arm (here, TRP1 and URA3), a sequence
that allows autonomous replication in yeast (ARS), and restriction
sites for cloning

15.6 GENOMIC LIBRARIES


A genomic library is a collection of clones that, when successfully made,
theoretically contains at least one copy of every DNA sequence in the
genome. Genomic libraries have many uses in molecular biology and in
genomics. Remember that a key step in analysis of a genome is breaking the
genomic DNA into smaller, more easily manipulated fragments. A genomic
library will contain these smaller fragments, which are used in many types of 117


Block 4 Applications of Genomics and Proteomics
genetic analysis. A genomic library can also be used to isolate and study a
particular clone, such as that for a gene of interest. In this section, we will
focus on the construction of genomic libraries of eukaryotic DNA.

Genomic libraries are made using the basic cloning procedures already
described. A restriction enzyme is used to cut the genomic DNA, and a vector
is chosen so that the entire genome is represented in a manageable number
of clones. You might assume that it is as simple as digesting the genomic DNA
completely with a restriction enzyme and cloning the resulting DNA fragments
in a cloning vector. This will create a genomic library, but this library will have
serious functional limitations for four important reasons:

1. If the researcher wants to study a specific gene that contains one or


more restriction sites for the enzyme used to create the library, the gene
will be split into two or more fragments when genomic DNA is digested
completely by the restriction enzyme. As a result, the gene would then
be cloned in two or more pieces.

2. The average size of the fragment produced by the digestion of


eukaryotic DNA with restriction enzymes is small (about 4 kb for
restriction enzymes that have 6-bp recognition sequences). Not only
many genes are larger than 4 kb (especially those in mammals), but also
an entire genomic library would have to contain a very large number of
recombinant DNA molecules, and screening for a specific gene would be
very laborious.

3. The number of base pairs between adjacent restriction sites can vary
significantly; so, for instance, cutting a 10 kb fragment of DNA with
BamHI might yield fragments of 500, 2,500, and 7,000 base pairs. When
genomic DNA is digested, the resultant fragments will fall in a range of
sizes. Some of these fragments will be too large to clone. As a result,
part of the genome would be unclonable in this type of library.

4. The most troublesome aspect of this sort of library is the loss of


information. If you have a library made, say, of the BamHI-generated
fragments of the 10 kb fragment described above, it would contain three
clones. You would have no idea how the individual fragments were
positioned in the original fragment, and you could never determine that
order from the library itself. Extrapolating this issue to the thousands of
clones in a genomic library made using complete digestion of genomic
DNA, you would not be able to reassemble the cloned fragments into
their arrangement in the genome.

To deal with these functional limitations, you need to break the genomic DNA
differently. Specifically, you need to break the genomic DNA into fragments
that are of the correct size for your cloning vector and that overlap each other.
To generate these overlapping fragments, you can either mechanically break
(shear) the genomic DNA, or you can use a restriction enzyme under
conditions such that the genomic DNA is digested partially.

DNA is sheared by passing it through a syringe needle to produce a


population of overlapping DNA fragments of a particular size. However,
118 because the ends of the resulting DNA fragments have been generated by


Unit 15 Manipulation of the Genome
physical means and not by cutting with restriction enzymes, additional
enzymatic manipulations are necessary to add appropriate ends to the
molecules for their insertion into a restriction site of a cloning vector.

Large, overlapping DNA fragments of appropriate size for constructing a


genomic library can also be generated by using a partial digestion of the
genomic DNA with a restriction enzyme that recognizes a frequently occurring
6- or 4-bp recognition sequence (Fig. 5.8a). Partial digestion means that only a
random portion of the available restriction sites is cut by the enzyme. This is
achieved by limiting the amount of the enzyme used and/or the time of
incubation with the DNA.

DNA fragments generated by partial digestion with a restriction enzyme can be


cloned directly (Fig. 15.8b). Regardless of how you broke the DNA into
overlapping fragments, there will be a broad distribution of fragment sizes.
Now it is necessary to select the fragments that are of the right size for cloning
in the vector being used, and to eliminate those that are either too small or too
large. Consider a population of overlapping fragments generated by partial
digestion with a restriction enzyme (Fig. 15.9a). One common way to sort
fragments of the desired size for cloning is to use agarose gel electrophoresis
(Fig. 15.9a). In agarose gel electrophoresis, an electric field is used to move
the negatively charged DNA fragments through a gel matrix of agarose from
the negative pole to the positive pole. The gel, a horizontal slab of agarose
and a liquid buffer, is made by pouring a hot, liquid agarose/buffer mix into a
mold. A toothed comb is added, which creates “wells” in the gel. As the
agarose mixture cools, the agarose itself forms a “sieve” through which the
DNA transits. The DNA fragments (produced by shearing or restriction
digestion) are placed in a well in the gel. Other wells may contain a DNA
ladder (also called DNA size markers), a set of DNA molecules of known size.
For example, complete digestion of the phage lambda chromosome with
HindIII, which yields fragments of 23.1 kb,9.4 kb, 6.6 kb, 4.4 kb, 2.3 kb, 2.0 kb,
and 0.56 kb, is frequently used as a DNA ladder and is often called a lambda
ladder. An electric field is then applied to the gel and the DNA migrates toward
the positive pole. Smaller molecules are able to move through the gel more
rapidly, and larger molecules move more slowly (Fig. 15.9a).

The separated DNA fragments are invisible to the eye. They are made visible
by adding either ethidium bromide or SYBR® Green to stain the DNA. Both
chemicals bind tightly to DNA and emit visible light when excited with the
correct wavelength of light. Ethidium bromide emits visible light after being
excited with ultraviolet light, and SYBR® Green, when bound to DNA, emits
green light after being excited with blue light. The emission of visible light
makes the position of the DNA in the gel obvious. Since the wells are
rectangular, the DNA fragments form “bands” on the gel. In Figure 15.9b, an
agarose gel electrophoresis analysis shows partial digestion of genomic DNA.
The vertical “lanes” of the gel show how the DNA fragments in the samples
loaded into the wells at the top separated during the electrophoresis. Lane 1
contains the DNA ladder, in this case the lambda ladder. Note the discrete set
of bands of known sizes in the lane. Lane 2 shows a sample of genomic DNA
not treated with a restriction enzyme. There is not a highly discrete band, but a
concentrated mass of DNA in a region of the lane corresponding to the large
DNA fragments of the lambda ladder, and a smear of DNA going down the 119


Block 4 Applications of Genomics and Proteomics
lane from that point. The mass of DNA is the large DNA fragments of genomic
DNA that came out of the cell. It is unavoidable to break the genomic DNA
mechanically during isolation, so the size of the large DNA is much smaller
than the sizes of chromosomes. The mechanical shearing during isolation is
also responsible for the many bands of various sizes of DNA fragments that
are seen as a smear down the lane. Lane 3 shows genomic DNA digested
completely with a restriction enzyme. There are no discrete bands of DNA
fragments here either. Instead, a smear of fragments is seen, most of which
are smaller than the smallest visible lambda ladder fragment at 2.0 kb. Lanes
4 and 5 show the results of digesting the genomic DNA partially using the
same restriction enzyme. In both cases the DNA is of much larger size than
that seen in the complete digest lane, this being the expected outcome of
partial digestion. The partial digestion conditions were different for the samples
loaded in the two lanes, with more digestion carried out for the DNA in lane 4
than for the DNA in lane 5. The difference in partial digestion conditions is
reflected in the range of DNA fragment sizes on the gel; that is, larger DNA
fragments are seen in lane 5 than in lane 4. As for the complete digestion of
genomic DNA, partial digestion does not result in discrete bands when the
digested DNA is analyzed by agarose gel electrophoresis. Rather, there is a
smear of DNA fragments of different sizes. Since there is a DNA ladder in the
gel showing where DNA fragments of particular sizes migrated, researchers
can use that information and isolate DNA fragments of the desired size for
cloning from the partial digest lanes. The isolation is done simply by cutting out
a block of agarose containing the DNA fragments of the desired size and then
extracting the DNA from the gel piece.

Agarose gel electrophoresis is an important technique used commonly in the


lab to separate and visualize DNA fragments. It is useful for analyzing partial
digests of genomic DNA as we have discussed here as well as for analyzing
complete restriction digests of a variety of DNA molecules, including specific
clones, virus genomes, and organelle genomes.

While the aim of the methods is to produce a library of recombinant molecules


that contains all of the sequences in the genome but that is practically is not
possible. Some sequences are very difficult to clone and, as a result, will
either be absent or underrepresented in our library. For example, some
regions of eukaryotic chromosomes may contain sequences that affect the
ability of vectors containing them to replicate in E. coli; these sequences are
lost from the library.

How many clones are needed to contain all sequences in the genome? The
number of clones needed to include all sequences in the genome depends on
the size of the genome being cloned and the average size of the DNA
fragments inserted into the vector. The probability of having at least one copy
of any DNA sequence in the genomic library can be calculated from the
following formula:

where N is the necessary number of recombinant DNA molecules, P is the


probability desired, f is the fractional proportion of the genome in a single
recombinant DNA molecule (that is, f is the average size, in kilobase pairs, of
the fragments used to make the library divided by the size of the genome, in
120 kilobase pairs), and ln is the natural logarithm.


Unit 15 Manipulation of the Genome

Fig. 15.8: Partial digestion with a restriction enzyme to produce overlapping


DNA fragments of appropriate size for constructing a genomic library

Fig. 15.9: Separation of DNA fragments by agarose gel electrophoresis. (a)


Partial digestion of genomic DNA with a restriction enzyme, and
separation of the DNA fragments by agarose gel electrophoresis.(b)
Agarose gel electrophoresis analysis of genomic DNA partially
digested with a restriction enzyme. Lane 1: Lambda ladder (a type of
DNA ladder). The sizes for the DNA bands of the ladder are indicated
on the left side of the gel. Lane 2: Genomic DNA undigested by a
restriction enzyme. Lane 3: Genomic DNA digested completely with a
restriction enzyme. Lanes 4 and 5: Genomic DNA digested partially
with a restriction enzyme. Enzyme reaction conditions allowed for less
DNA digestion for the DNA in lane 5 than the DNA in lane 4 121


Block 4 Applications of Genomics and Proteomics

15.7 CHROMOSOME LIBRARIES


A genomic library must contain a very large number of clones to achieve
nearly complete representation of the genome. This is particularly a major
problem for larger genomes, like the human genome. One solution to this
problem is to simplify the library by making several smaller libraries, each from
an individual chromosome. A library consisting of a collection of cloned DNA
fragments derived from one chromosome is called a chromosome library. In
humans, this means 24 different libraries, one each for the 22 autosomes, the
X, and the Y. Since each chromosome is far smaller than the total genome,
the resulting libraries can also be smaller. If you wish to clone a specific gene
but do not have genomic sequences, libraries (either genomic or
chromosome) will be important tools for finding and cloning that gene.

Individual chromosomes can be separated if their morphologies and sizes are


distinct enough, as is the case for human chromosomes. In one separation
procedure, flow cytometry, chromosomes from cells in mitosis are stained with
a fluorescent dye and passed through a laser beam connected to a light
detector. This system sorts the chromosomes based on the differences in dye
intensity that result from subtle differences in the abilities of the various
chromosomes to bind the dye. Once the chromosomes have been sorted and
collected from a number of cells, a library of each chromosome type can be
made.

15.8 DNA SEQUENCING AND ANALYSIS OF DNA


SEQUENCES
A clone from a genomic library, or any other clone, can be analyzed to
determine the nucleotide sequence of the DNA insert, as well as to determine
the distribution and location of restriction sites. Its nucleotide sequence is the
most detailed information one can obtain about a DNA fragment. The
information is useful, for example, in computer database analyses for
comparing sequences from different genomes, which can tell us how closely
related two organisms are, or for identifying gene sequences and the
regulatory sequences like promoters, silencers, and enhancers that control
gene expression. Furthermore, the DNA sequence of a protein-coding gene
can be translated by computer to provide information about the properties of
the protein for which it codes. Such information can be helpful for an
investigator who wants to isolate and study a protein product of a gene for
which a clone is available. Walter Gilbert and Frederick Sanger shared one-
half of the 1980 Nobel Prize in Chemistry for their “contributions concerning
the determination of base sequences in nucleic acids.” The DNA sequence of
protein-coding genes is also useful for comparing the sequences of
homologous genes from different organisms. These analyses can compare
either the DNA sequences from the organisms or the predicted protein
sequences. Comparative genomics is a field that is growing as more and more
122 genomic sequences become available.


Unit 15 Manipulation of the Genome
Dideoxy sequencing

The most commonly used method of DNA sequencing called dideoxy


sequencing (developed by Fred Sanger in the 1970s), is based on DNA
replication. Using a sequence of interest already cloned into a vector as a
template, DNA polymerase adds nucleotides to a short primer, until the
extension of the new DNA strand is stopped by the inclusion of a modified
nucleotide. This generates an array of short fragments, which can be
interpreted by gel electrophoresis either in an automated DNA sequencer or in
a standard gel apparatus. Both linear DNA and circular DNA can be
sequenced using the dideoxy DNA sequencing method. Linear DNA fragments
can be generated, for example, by cutting plasmid DNA with a restriction
enzyme or enzymes, or by using the polymerase chain reaction (PCR).

In dideoxy DNA sequencing, the template DNA is first denatured to single


strands by heat treatment. Next, an oligonucleotide (short DNA strand) primer
is annealed to one of the two DNA strands (Fig. 15.10a). Typically the primer
is 10-20 nucleotides long. For simplicity, the primers shown in the DNA
sequencing figure are 3 nucleotides long. The oligonucleotide primer is
designed so that its end is next to the DNA sequence the investigator wishes
to determine. The oligonucleotide acts as a primer for DNA synthesis
catalyzed by a DNA polymerase enzyme and its orientation ensures that the
DNA made is a complementary copy of the DNA sequence of interest (Fig.
15.10a).

Commonly, the DNA sequence a researcher wishes to determine is that of the


insert in a cloning vector. This is the case for the inserts in a genomic library
when a complete genome sequence is the goal. Consider as an example, a
DNA fragment cloned into the plasmid cloning vector pBluescript II (Fig. 15.5).
For this discussion, the fragment cloned had a KpnI sticky end at one end and
a SacI sticky end at the other and was cloned into pBluescript II that had been
cut in the multiple cloning site with both KpnI and SacI (Fig. 15.10b).With an
oligonucleotide primer complementary to a DNA sequence adjacent to the
multiple cloning site, you can sequence into the DNA insert. In fact, most
plasmid cloning vectors have the same sequences flanking their multiple
cloning sites, so that with only two universal sequencing primers you can
sequence into any cloned insert in those vectors. Two such primers are the
SP6 and T7 universal sequencing primers and sites to which they anneal are
at the ends of the multiple cloning site in pBluescript II (Fig. 15.10b). Both
universal sequencing primers are ultimately useful in sequencing. For
instance, after a pBluescript II based clone is denatured with heat, the SP6
universal sequencing primer will anneal to one of the two strands, in this case
to a DNA region at the left end of the multiple cloning site (Fig. 5.10b). Using
this primer, you can sequence the DNA insert from this side. With a second
reaction that uses the T7 universal sequencing primer, which is
complementary to a short segment of DNA on the other side of the multiple
cloning site, you can sequence the DNA insert from that side. If the DNA insert
is small, the two sequencing reactions will cover much of the same DNA
sequence but will give the sequence of the two complementary strands. 123


Block 4 Applications of Genomics and Proteomics

Fig. 15.10: Primers for DNA sequencing. (a) In a DNA sequencing reaction,
double-stranded DNA is denatured to single strands, and the
sequencing primer anneals to a specific region of one of the two
strands. Extension of the primer by DNA polymerase produces new
DNA that is complementary to DNA to which the primer annealed; this
is the sequencing reaction. The other DNA strand plays no role in the
sequencing reaction. (b) Most commonly used vectors allow the use
of universal sequencing primers. For pBluescript II, the T7 universal
sequencing primer anneals near the KpnI site of the multiple cloning
site, and the SP6 universal sequencing primer anneals near the SacI
site at the other end of the multiple cloning site. The binding sites for
the primers are positioned so that, when a sequencing primer
anneals, extension of the primer by DNA polymerase produces a DNA
strand complementary to that of the DNA insert

Dideoxy sequencing is done using an automated DNA sequencer, that permits


rapid sequencing of DNA and computerized analysis of the results. For an
experiment using an automatic DNA sequencer, a single dideoxy sequencing
reaction is set up. Each reaction includes the template DNA to be sequenced
and a sequencing primer that, as you have just learned, sets the point from
which the DNA sequence will be determined. When the template DNA is
denatured to single strands by heat treatment, the primer anneals to one of the
124 two strands (Fig. 15.10b). DNA polymerase, the four normal deoxynucleotide


Unit 15 Manipulation of the Genome
precursors (dNTPs, that is dATP, dTTP,dCTP, and dGTP; Fig. 15.11a), and a
small amount of modified nucleotide precursors called dideoxynucleotides
(ddNTPs, that is ddATP, ddTTP, ddCTP, and ddGTP; Fig. 15.11b) are then
added. A dideoxynucleotide differs from a normal deoxynucleotide in that it
has a 3ƍ-H rather than a 3ƍ-OH on the deoxyribose sugar. Furthermore,
different fluorescent dye molecules are linked covalently to each of the four
dideoxynucleotides. These dyes absorb certain wavelengths of light, causing
them to emit very specific wavelengths of light. For instance, the ddGTP
appears blue-green because a dye is bound to it that emits light with a
wavelength of 520 nm (blue-green), while the ddATP appears green, the
ddCTP appears a different shade of green, and the ddTTP appears greenish
yellow. Generally, the dideoxynucleotide (ddNTP) precursors are present in
the reaction mixture at about one by one-hundredth the amount of the normal
deoxynucleotide (dNTP) precursors so that some DNA synthesis occurs in the
dideoxy sequencing reactions. When the dideoxysequencing reaction starts,
DNA polymerase adds a nucleotide to the 3ƍ-OH at the end of the primer. In
the example shown in Figure 15.12a, the template has an A nucleotide, so the
primer is extended by a T nucleotide. Since most of the DNA precursors in the
reaction are dNTPs, the probability is great that a dTTP will be used for this
extension step. However, there is a small chance that DNA polymerase will
use the ddTTP precursor for this extension step. If the normal dTTP precursor
is used, the extended DNA chain has a 3ƍ-OH at its end and, therefore,
another nucleotide can be added by DNA polymerase. However, if the dideoxy
ddTTP precursor is used, the extended DNA chain has a 3ƍ-H at its end and,
therefore, another nucleotide cannot be added by DNA polymerase. In other
words, the addition of a dideoxy nucleotide to a DNA chain being synthesized
terminates the DNA synthesis reaction. Therefore, in the example in Figure
15.12a, the addition of the normal T nucleotide leads to the next extension
step, during which again there is a choice of nucleotide precursor types, in this
case between dATP and ddATP.

Fig. 15.11: Deoxynucleotide (dNTP) and dideoxynucleotide (ddNTP) DNA


precursors 125


Block 4 Applications of Genomics and Proteomics

In a dideoxy sequencing reaction, there are millions of identical starting


template/primer pairs, all undergoing the same extension reaction. Therefore,
some reactions will stop at nucleotide 1 of the template DNA after
incorporating a dideoxy T nucleotide, others will stop at nucleotide 2 after
incorporating a dideoxy A nucleotide, yet others will stop at nucleotide 3 after
incorporating a dideoxy G nucleotide, and so on. Overall, a population of
newly synthesized DNA is produced with large numbers of new DNA
fragments ending at every position (Fig. 15.12b). Recall that each newly
synthesized fragment is color labeled by the dye attached to the dideoxy
nucleotide that is at the 3ƍ end of the fragment. In the reaction, the many
different-sized chains produced that end with ddT are all greenish yellow, all
chains ending with ddG are blue-green, and so on. In short, each DNA chain
synthesized starts from the same point and ends at the base determined by
the dideoxy nucleotide incorporated. The dye attached to the dideoxy
nucleotide color-codes the newly synthesized fragments, so you can identify
the last nucleotide added to that fragment.

The DNA chains in each reaction mixture are separated by a special, very
sensitive type of electrophoresis in a very small capillary, and a laser eye at
the end of the capillary detects the colored fragments as they exit the capillary.
While the dyes emit similar colors, the computer converts the minor color
differences into a far more obvious difference by assigning “false colors” to
each dye, such as using green for A, black for G, red for T, and blue for C. The
output is a series of colored peaks corresponding to each nucleotide position
in the sequence (Fig. 15.12c). The graphic representation is converted to a
sequence of nucleotides by a computer with the oversight of the researcher.
Automated sequencing is of great utility to research teams in determining the
complete sequences of various genomes because a single machine can
analyze 100 or more samples per day.

The DNA sequence of the newly synthesized strand is determined by the


computer associated with the laser by reading up the sequencing ladder from
the first colored fragment to exit the capillary (the smallest fragment with a
dye-labeled dideoxy nucleotide) to the last readable fragment to exit
(corresponding to the largest fragment with a dye-labeled dideoxy nucleotide)
to give the sequence in 5ƍ-to-3ƍ orientation. Generally, several hundred
nucleotides can be read by the laser before a “traffic jam” of fragments makes
it impossible to determine the exact order in which fragments exit the capillary.
In Figure 15.12b, the smallest DNA fragment ended with ddA, the second
smallest DNA fragment ended with ddT, and so on. “Reading” the sequence
from smallest fragment to largest gives 5ƍ-TACTGGTACAA-3ƍ. This sequence
is complementary to the sequence of the template sequence.

To sequence more nucleotides than can be read for a single reaction, the first
sequence obtained is used to design a custom primer that will anneal to the
DNA insert near the 3ƍ end of that sequence. The sequencing reaction using
the new primer generates a DNA sequence that partially overlaps the first
sequence. In this way, a researcher can step down a long DNA insert and
126 obtain its complete sequence.


Unit 15 Manipulation of the Genome

Fig. 15.12: Dideoxy sequencing. (a) A dideoxy sequencing reaction consists of


the template DNA, a sequencing primer, DNA polymerase, and a
mixture containing deoxy nucleotide (dNTP) DNA precursors and a
small amount of dideoxynucleotide (ddNTP) DNA precursors. When
DNA polymerase uses a (normal) dNTP precursor to extend the DNA
chain, a -OH on the incorporated nucleotide permits the addition of
another nucleotide. When DNA polymerase uses a ddNTP precursor
to extend the DNA chain, a -H on the incorporated nucleotide
prevents the addition of another nucleotide. (b )In a sequencing
reaction, a large number of template/primer pairs are present, which
leads to the synthesis of DNA fragments stopped at all possible
positions along the DNA template strand by the incorporation of a
dideoxynucleotide. The automated sequencer generates the curves
from the fluorescing bands on a gel. The colors are generated by the
machine and indicate the four bases: A is green, G is black, C is blue,
and T is red. Where bands cannot be distinguished clearly, an N is
listed 127


Block 4 Applications of Genomics and Proteomics
Pyrosequencing

A new automated technique, pyrosequencing, starts in a similar manner to


dideoxy sequencing, with a single-stranded DNA template and a sequencing
primer but the pyrosequencer machine detects the incorporation of nucleotides
into the growing strand without chain termination. Pyrosequencing is named
for the pyrophosphate molecule (two phosphate groups connected by a
covalent bond) that is released when a dNTP is used by DNA polymerase to
extend a new DNA strand. As you will see, the enzymatically based detection
of the released pyrophosphate by the pyrosequencer provides information
about the template sequence. Figure 15.13 illustrates the principles of the
pyrosequencing technique. The DNA to be sequenced is denatured to form
single-stranded DNA. The single-stranded DNA is attached to a solid,
microscopic bead that is placed in a microscopic well in the pyrosequencer.
The sequencing reaction mixture, consisting of a primer, DNA polymerase,
and three other enzymes, is added. The four dNTPs are not present in the
initial mix, but are added sequentially to and removed from the
pyrosequencing reaction, such that only one dNTP is present in the reaction at
any one time. This cycle of addition and removal of each dNTP in turn repeats
over and over. You will start with a reaction just as dCTP is added to the bead
(Fig. 5.13a). Since the first unpaired base in the template strand is a G, the
dCTP can be added to the end of the primer by DNA polymerase, and a
molecule of pyrophosphate (PPi) is released. Another enzyme in the mix uses
this pyrophosphate in a reaction that produces ATP, and a third enzyme uses
the energy stored in the newly produced ATP to produce light. The
pyrosequencer detects and quantifies the amount of light released and
correlates it to which dNTP was present in the reaction. Thus, for this
example, since light was emitted when dCTP was present, you know that C
was incorporated into the growing strand. Excess dCTP is destroyed by
another enzyme in the reaction. Now another dNTP is added, for example,
dTTP. In the example, no light is emitted when dTTP is added, because a
dTTP will not base-pair with the C on the template. The excess dTTP is
degraded enzymatically, and the pyrosequencer will next add dATP. Once
again, this cannot be added to the growing strand, so the dATP is destroyed
without powering the creation of light. The next addition is dGTP. Since the
next two bases on the template strand are both C, DNA polymerase adds two
molecules of dGTP to the growing strand after the C. This means that new
DNA with the sequence 5ƍCGG-3ƍ has been synthesized. You can tell that two
G residues were incorporated, since adding two G residues to the growing
strand releases two molecules of pyrophosphate, which are in turn used to
create two molecules of ATP, and twice as much light is produced as is the
case when one nucleotide is added to the strand. The pyrosequencer
measures exactly how much light is made as a particular dNTP is added, and,
based on the output of light, you can determine the exact sequence of the
DNA that has been synthesized based on the pyrogram. The pyrosequencer
continues this cyclical process, adding dCTP, then returning to dTTP, dATP,
dGTP, and so on. As for dideoxy sequencing, the DNA sequence obtained is
the complement of the sequence of the DNA template. The pyrosequencing
reaction has been described with one bead. The pyrosequencer has about
200,000 microscopic wells, in each of which a different pyrosequencing
128 reaction with a different single-stranded template DNA attached to a bead is


Unit 15 Manipulation of the Genome
carried out. Thus, the sequencing of many DNA templates is done
simultaneously, making it possible to obtain about 20 million nucleotides of
genome sequence in about 6 hours. The pyrosequencing technique is still
quite new and expensive, but it should become an important technique as the
equipment becomes refined and more affordable.

Fig. 15.13: In a pyrosequencing reaction, a single-stranded DNA template is


attached to a bead. A sequencing primer and several enzymes,
including DNA polymerase, are added. dNTPs are added to this mix
one at a time. In this example, dCTP has just been added to the
reaction. DNA polymerase can add a deoxy C nucleotide to the 3ƍ end
of the growing strand. This reaction releases pyrophosphate (PPi),
which is converted to ATP by a second enzyme in the mixture and
then a third enzyme in the mix breaks this ATP to release light. The
pyrosequencer quantifies the amount of light released. Excess dCTP
is consumed by yet another enzyme in the mixture and then another
dNTP is added. If the next dNTP is dTTP or dATP, no reaction occurs,
since neither can be added to the growing strand. Only when dGTP is
added can the new DNA strand be extended. In this case two units of
light will be created since the template has two adjacent C
nucleotides, so two deoxy G nucleotides can be added

6$4
6$4
a). Answer in one word only:

i) What is the most commonly used method of DNA sequencing?

ii) How the separated DNA fragments in the agarose gel are made
visible?

iii) How much insert size is accepted by BACs?

iv) Name the cloning vector that can carry the maximum size of DNA
insert.

v) What is produced due to the cleavage of X-gal by ȕ-


galactosidase? 129


Block 4 Applications of Genomics and Proteomics
b). Fill in the blanks with appropriate words:

i) DNA fragments of up to ……………. may be cloned efficiently in E.


coli plasmid cloning vectors.

ii) The selectable marker for BAC is camR for ……………. resistance.

iii) Complete digestion of the phage lambda chromosome with HindIII


is frequently used as a ……………. and is often called a lambda
ladder.

iv) A ……………. is a gene that allows us to determine easily if a cell


does or does not contain the cloning vector.

v) Modified nucleotide precursors called ……………. are used in the


DNA sequencing method called dideoxy sequencing.

vi) In ……………. reaction technique, a single-stranded DNA template


is attached to a bead.

c). Match the following:

Column A Column B

i) Werner Arber a) Haemophilus parainfluenzae

ii) SalI b) Haemophilus haemolyticus

iii) SmaI c) Dideoxy DNA sequencing

iv) Paul Berg d) Streptomyces albus

v) HpaII e) Sticky ends

vi) BamHI f) Restriction enzymes

vii) HhaI g) Blunt ends

viii) Frederick Sanger h) Recombinant-DNA

15.9 SUMMARY
• Genomics is the study of an organism's complete DNA sequence. The
process begins with cloning the organism's DNA into various types of
vectors. The exact nucleotide sequence of these clones is then
determined. These sequence data can be utilized in numerous analyses,
such as identifying regions that encode genes.

• DNA cloning involves introducing foreign DNA sequences into a specific


type of vector, an artificially constructed DNA molecule that allows the
foreign DNA to be replicated when inserted into a host cell, usually a
bacterium or yeast. Cloning entire chromosomes is generally impractical,
so the genomic DNA of an organism must be broken down into smaller
fragments before cloning. One method to achieve this is by using
130 restriction enzymes to cut the DNA.


Unit 15 Manipulation of the Genome
• Reporter genes are a type of protein-coding gene that are often tagged
to a gene of interest. A reporter gene is an exogenous coding region that
is joined to an expression vector containing a promoter sequence or
element in order to enable the measurement of promoter activity in cells.

• In knock-in animals novel externally supplied genes are expressed,


whereas in knockouts native genes are deleted or suppressed in order
to experimentally determine its function and the resulting phenotypic
changes. Through the use of knockout and knock-in technologies,
scientists can alter genes in a selected model system, providing valuable
insights into the functioning of individual genes.

• Various types of cloning vectors have been developed, with plasmids


being the most commonly used. Cloning vectors typically replicate within
one or more host organisms, have restriction sites for the insertion of
foreign DNA, and contain one or more selectable markers to identify
cells that harbor the vectors. Bacterial artificial chromosomes (BACs)
and yeast artificial chromosomes (YACs) enable the cloning of DNA
fragments that are several hundred kilobase pairs long in E. coli and
yeast, respectively.

• Restriction enzymes cut DNA at specific locations called restriction sites.


Each restriction enzyme recognizes a unique sequence of nucleotides
within the DNA and cleaves both strands, often producing a small
overhang known as a "sticky end". Complementary sticky ends can
reanneal with each other, joining two different pieces of DNA to form a
recombinant DNA molecule, provided they have been cut by the same
restriction enzyme or by enzymes that generate compatible ends. Some
restriction enzymes produce blunt ends when they cleave DNA. Blunt-
ended molecules can also be joined to form recombinant DNA
molecules.

• After DNA is cleaved by a restriction enzyme, it can be cloned into a


vector that is also cut by the same restriction enzyme. The genomic DNA
and vector DNA are mixed, allowing the sticky ends to anneal the
genomic DNA to the vector. The enzyme DNA ligase then restores the
phosphodiester backbone of the two DNA strands, covalently attaching
the two pieces together. The resulting recombinant DNA vector can now
be transformed into a host cell.

• Cloning vectors comprise several essential features: a multiple cloning


site (MCS), which is a collection of various restriction sites; an
appropriate origin of replication, ensuring the plasmid can replicate in the
chosen host cell; and a selectable marker, allowing the rare, transformed
cells to preferentially survive under certain conditions compared to their
untransformed neighbors. Common vectors include plasmids, cosmids,
yeast artificial chromosomes (YACs), and bacterial artificial
chromosomes (BACs), each with their own advantages and
disadvantages.

• To sequence a complete genome, it must be broken into fragments, and


each fragment must then be cloned and sequenced. A collection of 131


Block 4 Applications of Genomics and Proteomics

clones containing at least one copy of every DNA sequence in an


organism’s genome is known as a genomic library. The size of the library
depends on the size of the DNA inserts in the clones and the genome
size. For large genomes, a library may contain many thousands to
millions of clones. Vectors like BACs and YACs can hold larger DNA
fragments, so fewer clones are needed to construct a complete library
when using these vectors. A chromosome library is smaller than a
genomic library because it contains only the DNA from one specific
chromosome.

• Once a genomic library is completed, the DNA within that library can be
sequenced. One popular method of DNA sequencing involves the use of
dideoxynucleotides (ddNTPs) to terminate chain extension in a modified
version of DNA replication. These terminated fragments are detectable
because the individual ddNTPs are linked to colored dyes. The dyes
allow the fragments to be visualized and provide information on which
ddNTP terminated the fragment.

• A newer sequencing technique, called pyrosequencing, directly detects


the identity of each nucleotide as it is incorporated into the growing DNA
strand, eliminating the need for chain termination. This method relies on
the detection of pyrophosphate release during nucleotide incorporation,
which triggers a series of reactions resulting in a light signal proportional
to the number of nucleotides added.

15.10 TERMINAL QUESTIONS


1. How is a gene isolated and amplified by cloning?

2. Describe the importance of reporter gene in the study of DNA regulatory


sequences.

3. Write a note on restriction enzymes citing examples.

4. Describe the role of plasmid cloning vectors in cloning.

5. Discuss various types of artificial chromosomes.

6. Describe Genomic and Chromosome libraries.

7. Describe different methods of DNA sequencing.

15.11 ANSWERS
Self-Assessment Questions
1. a) i) Escherichia coli strain RY13, ii) Blunt ends, iii) Gene
knockout, iv) knock-in animal

b) i) reporter gene, ii) restriction enzyme, iii) Haemophilus


o
influenzae strain Rd, iv) 37 C, v) RNA interference
132 (RNAi)


Unit 15 Manipulation of the Genome
®
2. a) i) Dideoxy sequencing, ii) Ethidium bromide or SYBR Green,
iii) up to 300 kb, iv) Yeast artificial chromosomes,
v) blue dye

b) i) 15 kb, ii) chloramphenicol, iii) DNA ladder, iv) selectable


marker, v) dideoxynucleotides, vi) pyrosequencing

c) i) f, ii) d, iii) g, iv) h, v) a, vi) e, vii) b, viii) c

Terminal Questions
1. Refer to Section 15.2.

2. Refer to Subsection 15.2.2.

3. Refer to Section 15.4.

4. Refer to Subsection 15.5.1.

5. Refer to Subsection 15.5.2.

6. Refer to Section 15.6 and 15.7.

7. Refer to Section 15.8.

133




UNIT 16
(;35(66,21$1$/<6,62)
*(120(
*(120(

6WUXFWXUH
6WUXFWXUH
16.1 Introduction 16.5 RNA Interference

Objectives Mechanism of RNA


Interference
16.2 Reporter Genes
Cellular Mechanism
Transformation and
Transfection Assays Transcriptional Silencing

Gene Expression Assays Variation Among Organisms

Promoter Assays Biological Functions of RNAi

16.3 Temporal and Site-specific Applications of RNAi


Gene Expression
16.6 Summary
Temporal Analysis
16.7 Terminal Questions
Site-specific Analyses
16.8 Answers
16.4 Gene Silencing

Unravelling RNA Silencing

Important Features of RNA


Silencing

Components of Gene
Silencing

16.1 INTRODUCTION
You learned about the analysis and modification of the genome in Unit 15. In
the present Unit, you will learn about the principles of genomic expression
analysis. The simplest way to define gene expression analysis is the study of
how genes are transcribed to produce functional gene products, such as
functional RNA species or proteins. Understanding gene control helps
distinguish between abnormal or unhealthy cellular processes and normal
ones such as differentiation.
Unit 16 Expression Analysis of Genome

2EMHFWLYHV 

After studying this Unit, you would be able to:

™ describe the concept and function of reporter genes,

™ comprehend the basics of temporal and site-specific gene expression,

™ explain the concept of gene silencing,

™ explain the basics and mechanism of RNA interference,

™ comprehend the biological functions of RNA interference, and

™ discuss the applications of RNA interference.

16.2 REPORTER GENES


A reporter gene (commonly called as reporter) in molecular biology is a gene
that is attached to a regulatory sequence of another gene of interest in plants,
animals, cell cultures, or bacteria that researchers pay attention to. These
genes are referred to as reporters because the expressed traits in an
organism can be easily detected and quantified, or they are the selectable
markers. Reporter genes are frequently used to study whether a specific gene
has been taken up or expressed in a targeted cell or population. Figure 16.1
shows the mechanism of the expression of a reporter gene.

The gene of interest and the reporter gene are cloned in a DNA construct
before being transferred into the cell or organism. This construct typically
takes the shape of a plasmid, a circular DNA molecule that is present in
prokaryotic or bacterial cells in a culture. The expression of the reporter gene
is considered as a signal for successful uptake of the gene of interest by the
cell or organism and hence its study is essential.

Fluorescent and luminous proteins are frequently utilized as the reporter


genes that produce visually recognizable properties. Examples include the
enzyme luciferase, which catalyzes a reaction with luciferin producing
luminescence, green fluorescent protein (GFP) gene of the jellyfish expresses
a protein that exhibits green fluorescence when exposed to light in the blue to
ultraviolet range, and the red fluorescent protein from the gene dsRed [fr].
Research related to the genetic transformation of plants has traditionally used
the GUS gene, although luciferase and GFP are also increasingly popular.

One of the most commonly utilised reporter genes in bacteria is the lacZ gene
of E. coli, which codes for the protein beta-galactosidase. The enzyme
expressed by this gene gives bacteria a blue colour when they are grown on
medium containing the substrate analogue X-gal. Another example of a
bacterium-selectable marker that also serves as a reporter is the
chloramphenicol acetyltransferase (CAT) gene, which confers resistance to
the antibiotic chloramphenicol. 135


Block 4 Applications of Genomics and Proteomics

Fig. 16.1: The mechanism of the expression of a reporter gene

16.2.1 Transformation and Transfection Assays

Only a small percentage of the population responds well to several


transformation and transfection techniques, which are used to express a
modified or foreign gene in an organism. Therefore, a technique for locating
those rare instances of effective gene uptake is required. For instance, the
galactosidase system can be injected with IPTG to produce constitutive or
inducible reporter gene expression. When utilized in this manner, reporter
genes generally express themselves under the regulation of their own
promoters, irrespective of the inserted gene of interest's promoter (DNA region
that triggers gene transcription). As a result, the expression of the reporter
gene is unrelated to the expression of the gene of interest, which is
advantageous when the gene of interest is expressed only under specific
conditions circumstances or in difficult-to-access tissues.

16.2.2 Gene Expression Assays

Reporter genes can result in the production of a protein that has minimal direct
effects on the organism or cell culture. Reporter genes can be activated, that
is, they can be expressed constitutively by cloning the gene of interest
together with the reporter gene where both the genes are regulated by the
same promoter. This results in the transcription of a messenger RNA that code
for two protein-coding sequences (fusion protein). Despite being fused, it is
crucial that both proteins are able to correctly fold into their active
conformations and engage in interactions with their substrates. In order to
ensure that the reporter and gene product do not restrict each others function,
a piece of DNA encoding a flexible polypeptide linker region is typically
inserted while creating the DNA construct. Additionally, reporter genes may be
induced to express throughout growth. In these situations, the reporter gene is
expressed by the use of trans-acting elements, such as transcription factors.

Researchers need to find pathways, small molecule inhibitors, and activators


of protein targets to develop new drugs. Also, reporter gene assays are
increasingly being used in high throughput screening (HTS). As the reporter
enzymes, such as firefly luciferase, can be direct targets of minute molecules
and obstruct the interpretation of high-throughput sequencing data, novel
coincidence reporter designs featuring artefact suppression have been
136 developed.


Unit 16 Expression Analysis of Genome
16.2.3 Promoter Assays

Reporter genes can be employed to test for a specific promoter's activity in an


organism or cell. There isn't a distinct "gene of interest" in this instance; the
reporter gene is simply put under the target promoter's control, and the
reporter gene product's activity is measured. Results are typically expressed in
terms of activity under "consensus" promoters, which are known to strongly
drive gene expression.

6$4
6$4
Fill in the blanks:

a) These genes are referred to as reporters because the expressed traits in


an organism can be easily detected and quantified and hence can be
used as ……………….. .

b) The lacZ gene of E. coli codes for the protein ……………….. .

c) A bacterium selectable marker that also serves as a reporter is the


……………….. gene, which confers resistance to the antibiotic
chloramphenicol.

16.3 TEMPORAL AND SITE-SPECIFIC GENE


EXPRESSION
The action that converts a gene's informational content into the creation of a
functional product, often a protein, is referred to as the gene expression
process. Although certain genes, such as those that code for transfer RNAs,
ribosomal RNAs, and a few other short RNAs, have RNA as their functional
output but majority of genes in cells code for proteins. The mechanism of gene
expression through the processes of genetic transcription and translation is
represented in Figure 16.2.

Fig. 16.2: Gene expression: A gene's or genes' phenotypic expression through


the processes of genetic transcription and translation 137


Block 4 Applications of Genomics and Proteomics
16.3.1 Temporal Analysis

The triggering of genes inside an organism’s particular tissues at particular


developmental junctures is known as spatiotemporal gene expression. The
intricacy of gene activation patterns varies greatly. Some are simple and
unchanging, like the pattern of tubulin, which is always expressed in living
cells. The expression of some, on the other hand, varies drastically within
seconds or from cell to cell, making them extremely complicated and
challenging to anticipate and model. Since a cell's identity is determined by the
collection of genes it actively expresses, geographical and temporal variation
is essential for the diversity of cell types present in mature organisms. There
could only be one kind of cell if gene expression was constant over time and
space.

For example, the wingless gene, a member of the wnt gene family, is
expressed in alternating stripes separated by three cells in the fruit fly
Drosophila melanogaster during the early stages of embryonic development.
This pattern disappears by the time the embryo becomes a larva, but the
wingless gene is still present in certain tissues, such as the imaginal discs of
the wings, which are patches of tissue that eventually develop into the adult
wings. The spatiotemporal pattern of wingless gene expression is determined
by a network of regulatory interactions, which includes the effects of multiple
unique genes such as even-skipped and Krüppel.

What causes variations in a gene's expression throughout time and space? A


retrogressive challenge exists in determining what resulted in the initial
changes in gene expression since present expression patterns only depend on
past expression patterns. Symmetry breakdown is the process resulting in
temporal and geographical variation in uniform gene expression. For instance,
the messenger RNA (mRNA) for the genes nanos and bicoid in Drosophila
embryonic development is deposited in the poles of the egg by maternal cells
prior to implantation, resulting in asymmetric expression of these genes in the
oocyte.

16.3.2 Site-specific Analyses

Analyses of gene expression at specific sites are crucial for comprehending


tissue functions. Although DNA-related technologies have advanced quickly, it
is still difficult to analyze a tissue's full genome for expression at a given
space. Information on cell types and site-specific gene expression is
necessary for a complete knowledge of tissue functions. Imaging techniques
like in situ hybridization have proved effective for analyzing the site-specific
gene expression of tissue. However, with these techniques, only a few marker
genes may be evaluated concurrently. It is necessary to employ a next-
generation DNA sequencer together with an adequate site-specific sampling
technique in order to analyze many genes at once. Furthermore, site-specific
multi-omics (such as genomes, transcriptomics, and proteomics) studies will
be required to fully comprehend tissue functions. The tiny sample sizes in
these procedures necessitate the employment of a sensitive analyzer.
However, a number of extremely sensitive analytical methods have been
138 created recently. For instance, utilizing next-generation DNA sequencers,


Unit 16 Expression Analysis of Genome
specific gene expression, or genome analysis approaches for single cells have
been described. More information on the corresponding functions is now
available than was previously possible, due to the combination of tissue
microdissections and high throughput examination of single cells. To fully use
analytical technologies, a technique for site-specific cell or microdissection
collection from a tissue should be established.

6$4
6$4
Fill in the blanks:

a) The action that converts a gene's informational content into the creation
of a functional product, often a protein, is referred to as the ……………..
process.

b) The triggering of genes inside an organism’s particular tissues at


particular developmental junctures is known as …………….. gene
expression.

c) Imaging techniques like …………….. have proved effective for analyzing


the site-specific gene expression of tissue.

16.4 GENE SILENCING


A unique gene regulation mechanism called RNA silencing controls the
number of transcripts either by inhibiting transcription via transcriptional gene
silencing (TGS), or by causing sequence-specific RNA to degrade via
posttranscriptional gene silencing (PTGS), or by RNA interference (RNAi).
Silencing of particular genes has been linked to regulatory processes such as
antiviral defence mechanisms, chromosomal remodeling, transposon
silencing, and gene regulation. It has also been linked to two primitive
processes, quelling in fungi and co-suppression in plants.

The first PTGS/RNAi studies were made in plants, but afterward, nearly all
eukaryotic species, including parasites, insects, protozoa, nematodes, flies,
mouse, and human cell lines, were shown to be experiencing RNAi-related
events. Quelling in fungi, co-suppression or PTGS in plants, and RNAi in the
animal kingdom are the phenotypically diverse but mechanistically related
forms of RNA interference. Recently, it has been discovered that other aspects
of eukaryotic cells' naturally occurring RNAi processes, such as microRNA
production and heterochromatinization, also exist.

Gene silencing caused by RNAi is a two-step mechanism, as per the


extensive genomic and pharmacological study. In the first phase, an enzyme
that mimics RNase III breaks down dsRNA into small interfering RNAs
(siRNAs), which are 21 to 23 nucleotides long. The siRNAs join RISC (RNA-
induced silencing complex), an RNase complex that acts on the corresponding
mRNA and destroys it, in the second phase. In various organisms, helicases,
dsRNA endonucleases, Dicer, and RNA-dependent RNA polymerase have all
been found to play significant roles in RNAi. A few of these elements also 139


Block 4 Applications of Genomics and Proteomics
manage the maturation of multiple species by processing a significant number
of non-coding RNAs, or microRNAs. MicroRNA biosynthesis and function
share common characteristics with RNAi processes. Because of its
exceptional effectiveness and specificity, RNAi is being viewed as a crucial
tool for gene-specific therapeutic activities that affect the mRNAs of disease-
related genes, as well as for functional genomics.

16.4.1 Unravelling RNA Silencing


It would be wise to give a general summary of the homology-dependent RNA
silencing mechanism and highlight its salient characteristics in order to
comprehend it.

Post-Transcriptional Gene Silencing in Plants

RNA silencing in plants came to light accidentally while looking for transgenic
petunia blooms that were supposed to be more purple. The goal of R.
Jorgensen's lab in 1990 was to increase the activity of the chalcone synthase
(chsA) gene, an enzyme essential for the synthesis of anthocyanin.
Unexpectedly, few transgenic petunia plants carrying the chsA coding area
controlled by a 35S promoter lost both transgene and endogenous chalcone
synthase activity, leading to the development of white or variegated sectors in
many of the flowers. Run-on transcription assays in isolated nuclei showed
that the decrease in cytosolic chsA mRNA was not associated with decreased
transcription. The term "co-suppression" was created by Jorgensen to
characterize the loss of mRNAs from both the transgene and the endogene.
Both sense and antisense transgenes have the potential to cause PTGS, and
biochemical data show that comparable processes may be at play in both
situations. It is important to note that the co-suppression phenomena have
been proven in metazoans and mammals in addition to plants.

RNA Interference and Quelling

With increasing reports of PTGS in plants, independent findings of homology-


dependent gene silencing in fungal systems were also made. They were
referred to as quelling. Quelling was discovered while attempting to alleviate
the output of an orange pigment produced by the Neurospora crassa gene al1.
A plasmid containing a 1,500-bp segment of the al1 gene's coding sequence
was used to transform a strain of N. crassa with a wild-type al1+ gene (orange
phenotype). A few transformants had their phenotypes stabilized to albinism.
While the native al1 mRNA was significantly decreased in the al1-quelled
strain, the quantity of unspliced al1 mRNA was comparable to that of the wild-
type strain, demonstrating that quelling and not the transcription rate changed
the level of mature mRNA in a homology-dependent manner.

Owing to the discovery of Fire et al., who firmly showed the biological nature of
inducers in gene silencing by administering pure dsRNA directly into the body
of Caenorhabditis worms, the phenomena of RNAi initially gained attention.

16.4.2 Important Features of RNA Silencing


Investigations on numerous organisms with labels such as quelling in fungi,
140 PTGS in plants, RNAi in mammals, and virus-induced gene silencing (VIGS)


Unit 16 Expression Analysis of Genome
have independently led to the discovery of a universal paradigm for gene
control. The dsRNA serves as the inducer and the target RNA is destroyed in
a homology-dependent manner. Additionally, the degradative machinery
needs a set of proteins that are common to most species in terms of both
structure and function. siRNA production and systemic transmission of
silencing from its site of initiation are two characteristics that are present in the
majority of these activities.

16.4.3 Components of Gene Silencing

Understanding the silencing mechanism has been attempted using genetic


and molecular methods. In order to look for mutants deficient in RNA
interference, quelling, or PTGS, genetic screens were performed on the algae
Chlamydomonas reinhardtii, the nematode Caenorhabditis elegans, the
fungus Neurospora crassa, and the plant Arabidopsis thaliana.

Dicer

Members of the RNase III family are one of the few nucleases which
specifically target dsRNAs and cleave them with 3ƍ-hydroxyl and 5ƍ-phosphate
termini and 3ƍ overhangs of 2 to 3 nucleotides. This enzyme was given the
name Dicer because it can convert dsRNA into evenly sized short RNAs
(siRNA). The nucleases in question have been preserved throughout evolution
in flies, worms, fungi, animals, and plants. Dicer has four diverse domains,
including an amino-terminal helicase, a dsRNA binding domain, two RNase III
motifs, and a PAZ domain (a 110-amino-acid domain found in proteins like
Argo, Piwi, and Pinhead/Zwille). It also shares this domain with the
QDE2/RDE1/Argonaute family of proteins, which has been genetically linked
to RNAi by separate studies. Dicer's tandem RNase III domains are suggested
to catalyze cleavage.

RNase III enzyme fully degrades dsRNA to produce fragments of 12 to 15 bp,


which are about half as large as siRNAs. The RNase III enzyme functions as a
dimer, whereas the Dicer enzyme contains two catalytic domains in each
monomer, one of which differs from the consensus catalytic sequences. The
Dicer enzyme breaks down dsRNA with the help of two compound catalytic
centers. The model for the production of 23- to 28-mer diced siRNA products
were developed as a result of the recent discovery of the crystal structure of
the RNase III catalytic domain. According to this model, the two internal
domains with selective sequence identity forfeit their functional relevance as
the dimeric Dicer folds on the dsRNA substrate to yield four compound
catalyst sites, while the two terminal sites with the best levels of homology to
the consensus RNase III catalytic sequence continue to function. As a result,
the diced products, which are twice as large as the typical 12- to 15-mer
pieces, appear as digests of the RNase III enzymes. A similar model suggests
that specific adjustments to the Dicer structure may alter the distance joining
the two active terminal sites, leading to the production of siRNAs with different
sizes and species-specific imprints. Clearly, this model has to be verified using
the Dicer crystal structure. 141


Block 4 Applications of Genomics and Proteomics
RNA-Induced Silencing (RISC) Complex and the Guide RNAs
The inability of cellular isolates exposed to a Ca2+ dependent nuclease
(micrococcal nuclease, that can digest both RNA and DNA) to degrade the
homologous mRNAs and the lack of this impact with DNase I treatment served
as evidence that RNA was a critical element of the nuclease activity. The
RNA-induced silencing complex (RISC) was named after the sequence-
specific nuclease activity that was seen in the cellular extracts and was
responsible for abasing target mRNAs.

6$4
6$4
Fill in the blanks:

a) Quelling in fungi, co-suppression or PTGS in plants, and RNAi in the


animal kingdom are the phenotypically diverse but mechanistically
related forms of …………………… .

b) The …………………… biosynthesis and function share common


characteristics with RNAi processes.
c) The term …………………… was created by Jorgensen to characterize
the loss of mRNAs from both the transgene and the endogene.

d) Independent findings of homology-dependent gene silencing in fungal


systems were referred to as …………………… .

e) The enzyme …………………… convert dsRNA into evenly sized short


RNAs (siRNA).
f) The …………………… was named after the sequence-specific nuclease
activity that was seen in the cellular extracts and was responsible for
abasing target mRNAs.

16.5 RNA INTERFERENCE


RNA interference (RNAi) is the procedure by which the targeted mRNA
molecules are neutralized by RNA molecules to block the targeted gene
expression. RNAi involves the targeting of RNA molecules for the sequence-
specific inhibition of gene expression by double-stranded RNA via
transcriptional or translational repression. Andrew Fire and Craig C. Mello split
the 2006 Nobel Prize in Physiology or Medicine for their 1998 publication of
RNAi in the nematode Caenorhabditis elegans. Since, RNAi and its regulatory
capabilities were discovered; it has become clear that RNAi has enormous
prospects in suppressing desired genes. It is currently acknowledged that
RNAi is precise, stable, and effective. Small interfering RNA (siRNA) and
microRNA (miRNA) are two classes of tiny ribonucleic acid (RNA) molecules
that are essential to the RNAi pathway. Post-transcriptional silencing happens
as a result of mRNA degradation because protein translation is halted. The
pre-transcriptional silencing process of RNAi can limit transcription by
catalyzing DNA methylation at genomic sites corresponding to complex
miRNA or siRNA. RNAi plays a crucial part in protecting cells against parasitic
nucleotide sequences (such as transposons and viruses) and also affects the
142 organisms development.


Unit 16 Expression Analysis of Genome
Numerous eukaryotic and animal cells naturally include the RNAi pathway. It is
started by the enzyme Dicer, which breaks lengthy dsRNA molecules into
short double-stranded pieces called small interfering RNAs (siRNAs), which
have 21 to 23 nucleotides. The sense, that is, passenger strand and the
antisense, that is, guide strand of each siRNA are unwound to form two single-
stranded RNAs (ssRNAs), respectively. The protein Argonaute 2 then cleaves
the passenger strand (Ago2). The guide strand is integrated into the RISC
while the passenger strand is destroyed. The target mRNA is then bound and
degraded by the RISC assembly. The guide strand interacts with a
complementary strand in an mRNA molecule, which activates Ago2, a catalytic
subunit of the RISC, to initiate cleavage.

16.5.1 Mechanism of RNA Interference


It is a method for gene regulation that restricts the amount of transcript in two
ways:

• Reducing transcription (transcriptional gene silencing);

• Decreasing the quality of the RNA generated (post-transcriptional


gene silencing)

The following steps can be used to explain RNA interference's mechanism:

• With the help of an enzyme known as Dicer, long, double-stranded RNA


is cut into manageable pieces. Small interfering RNA (siRNA), is the
name given to these tiny bits.

• The siRNAs are transported through the complex of RNA-induced


silencing. The RNA is triggered as the duplex unwinds.

• One of the strand of the double-stranded RNA is deleted when the


siRNA binds to the Argonaute protein, which promotes RNA breakdown
and prevents translation. The mRNA target sequences are bound by the
remaining strand.

• To control the target sequence, the Argonaute protein either cleaves the
mRNA or enlists the help of other agents.

16.5.2 Cellular Mechanism


Short double-stranded RNA molecules engage with the catalytic RISC
component argonaute in the cytoplasm of a cell to begin the RNA-dependent
gene silencing process known as RNA interference. RNAi is regulated by
RISC and is started by these interactions (Fig. 16.3). If the dsRNA is
exogenous (either from laboratory manipulations or infection with an RNA-
based virus), the RNA is transported straight into the cytoplasm and cuts into
brief fragments by Dicer. Similar to pre-microRNAs generated from RNA-
coding genes inside the genome, endogenous dsRNA can likewise originate
within the cell. The main transcripts from these genes are first processed in
the nucleus to create pre-distinctive miRNA's stem-loop structure before being
released to the cytoplasm. Thus, the foreign and endogenous dsRNA routes
meet at the RISC. 143


Block 4 Applications of Genomics and Proteomics
By making the ribonuclease Dicer more active, which binds to and cleaves
human short hairpin RNAs (shRNAs) or exogenous dsRNAs into double-
stranded fragments with 20 –25 base pairs and a 2-nucleotide overhang at the
3ƍ end, exogenous dsRNA triggers RNAi. The RISC-Loading Complex
subsequently divides these siRNAs into single strands and incorporates them
into an active RISC (RLC). Dicer-2 and R2D2 are parts of RLC, which is
necessary to connect RISC and Ago2. TATA-binding protein-associated factor
11 (TAF11) induces Dcr-2-R2D2 tetramerization, which increases the binding
affinity to siRNA and facilitates RLC formation tenfold. The R2-D2-Initiator
(RDI) complex would change into the RLC by association with TAF11. R2D2
possesses tandem double-stranded RNA-binding domains that enable it to
recognise the thermodynamically stable end of siRNA duplexes, while Dicer-2
recognises the opposite, less stable extremity. Asymmetric loading is brought
about by Ago2's MID domain, which locates the thermodynamically stable end
of the siRNA. The "passenger" (sense) strand, whose 5ƍ end is abandoned by
MID, is emitted, and the "guide" (antisense) strand is preserved and
collaborates with AGO to build the RISC. After joining the RISC, siRNAs base-
pair to their target mRNA and cleave it to prevent it from being used as a
translation template. In contrast to siRNA, a miRNA-loaded RISC looks for
possible complementarity among cytoplasmic mRNAs. miRNAs bind to
mRNAs in the 3ƍ untranslated region (UTR) where they normally show weak
complementarity, obstructing ribosome access and preventing translation.
Exogenous dsRNA is recognised and bound by an effector protein called
R2D2 in Drosophila and RDE-4 in C. elegans, which increases dicer activity. It
is uncertain what mechanism results in this length selectivity in this protein,
which only binds to long dsRNAs.

siRNA

A class of double-stranded RNA (20–24 bp) known as small interfering RNA


(siRNA) functions in the RNA interference pathway by preventing the
production of genes that have complementary nucleotide sequences to siRNA,
hence causing mRNA degradation. A short RNA species (siRNA) with both
sense and antisense polarity was found as the by-product of RNA degradation
by PTGS, providing a critical insight into the process. The siRNA molecules
with specific chemical structures are produced and assembled as dsRNA
molecules. In contrast to non-silenced control plants, siRNAs were initially
identified in plants that were either undergoing virus-induced gene silencing or
co-suppression. Later on, siRNAs were discovered in both Drosophila embryo
extracts which were showing RNAi in vitro, and Drosophila embryos which had
received dsRNA injections. These siRNAs were found in Drosophila tissue
culture cells where RNAi was triggered by adding >500-nucleotide-long
exogenous dsRNA. The production of siRNA was shown to be the
distinguishing characteristic of any homology-dependent RNA-silencing
method.

miRNA

MicroRNAs are short noncoding RNAs with a length of about 22 nucleotides


that are involved in posttranscriptional gene silencing, regulation of gene
expression and have important biological functions, particularly during
144 development. The phenomena of RNAi comprises of both the silencing


Unit 16 Expression Analysis of Genome
mechanisms caused by the external dsRNA and the endogenously produced
gene silencing effects of miRNAs. Although mature miRNAs share structural
similarities with siRNAs made from external dsRNA, miRNAs must first go
through a significant amount of post-transcriptional modification. A miRNA is
produced in the cell nucleus from a much longer RNA-coding gene as a
primary transcript called a pri-miRNA, which is then processed by the
microprocessor complex into a 70-nucleotide stem-loop structure called a pre-
miRNA. This complex is made up of the dsRNA-binding protein DGCR8 and
the RNase III enzyme Drosha. Since Dicer binds to and cleaves the dsRNA
part of this pre-miRNA to create the mature miRNA molecule which can be
incorporated into the RISC, siRNA and miRNA, all have similar downstream
biological mechanisms. Epstein-Barr virus (EBV) was the first human virus
shown to express miRNAs. Since then, several microRNAs in viruses have
been identified.

Since miRNAs, particularly those found in animals, frequently have imperfect


base pairing to a target and impede the translation of several other mRNAs
with identical sequences, siRNAs produced from lengthy dsRNA precursors
vary from miRNAs in this regard. Comparatively, siRNAs often base-pair
flawlessly and cause mRNA cleavage in only one particular target. Different
argonaute proteins and dicer enzymes are used in C. elegans and Drosophila
to handle siRNA and miRNA.

Three prime untranslated regions and microRNAs


Regulating sequences in the 3’ untranslated regions (3ƍUTRs) of mRNAs are
often involved in RNAi post-transcriptionally. Such 3ƍ-UTRs frequently have
binding sites for regulatory proteins and miRNAs. miRNAs can reduce the
number of certain mRNAs that are expressed by binding to particular locations
in the 3ƍ-UTR and either preventing translation or result in the degeneration of
the transcript. Additionally, the 3ƍ-UTR may contain silencer sequences that
bind repressor proteins that prevent mRNA production. Frequently, microRNA
response elements (MREs) are found in the 3ƍ-UTR. The miRNAs attach to the
MRE sequences. These 3ƍ-UTRs frequently include these motifs. MREs
account for nearly half of all regulatory motifs found in the 3ƍ-UTRs.

RISC activation and catalysis


During RISC activation, the other anti-guide strand, also known as the
passenger strand, is weakened. Although strand selection is unaltered by the
direction in which dicer cleaves the dsRNA prior to RISC integration, the guide
strand is often the one whose 5ƍ end is less firmly linked to its counterpart.
Instead, the more stable 5ƍ end of the passenger strand may be bound by the
R2D2 protein, which would then operate as the distinguishing factor.

It is unclear how the active RISC finds mRNAs that are compatible in a cell.
Translation of the mRNA target is not necessary for RNAi-mediated
degradation, despite the fact that it has been suggested that the cleavage
process is tied to translation. P-bodies, also known as GW bodies or
cytoplasmic bodies, are areas of the cytoplasm where argonaute proteins are
localized and miRNA activity is likewise concentrated. P-bodies have high
rates of mRNA degradation. P-bodies are thought to be a crucial location in
the RNAi process since their disruption reduces RNAi's effectiveness. 145


Block 4 Applications of Genomics and Proteomics
16.5.3 Transcriptional Silencing
Many eukaryotes employ RNAi pathway components to maintain the structure
and organization of their genomes. Pre-transcriptional downregulation of
genes is achieved by modification of histones and the ensuing induction of
heterochromatin formation; this procedure is known as RNA-induced
transcriptional silencing (RITS), and it is carried out by a protein complex
known as the RITS complex. It is unclear how the RITS complex influence the
development and organization of heterochromatin. To sustain the existing
heterochromatin regions, RITS assembles a complex of siRNAs that are
homologous to the localized genes and firmly attach to the methylated
histones. This complex also acts co-transcriptionally to damage any nascent
pre-mRNA transcripts that are started by RNA polymerase.

Dicer is necessary to produce the first complement of siRNAs that target future
transcripts; therefore, it makes sense that while its maintenance is not dicer-
dependent, the development of such a heterochromatin area is. It has been
proposed that heterochromatin maintenance works as a self-reinforcing
feedback loop, with new siRNAs being produced from sporadic nascent
transcripts by RdRP and incorporated into regional RITS complexes.

16.5.4 Variation Among Organisms


The capacity of various organisms to absorb foreign dsRNA and utilize it in the
RNAi pathway varies. RNAi effects can be systemic and heritable in C.
elegans and plants in contrast to Drosophila or humans. In plants, siRNAs are
assumed to be transferred between cells by plasmodesmata to allow RNAi to
spread. Heritability results from the methylation of promoters that RNAi has
targeted; the altered pattern of methylation is replicated in each subsequent
cell generation.

Fig. 16.3: The mechanism of RNA Interference

The targeting of endogenously produced miRNAs can be used to draw broad


general distinctions between plants and animals. While in plants, miRNAs
146 usually cause direct mRNA fragmentation by RISC because they are perfectly


Unit 16 Expression Analysis of Genome
or almost perfectly complementary with their target genes. In animals, miRNAs
typically have more divergent sequences and suppress translation. It is
possible to prevent translation initiation factors from interacting with the
polyadenine tail of the mRNA in order to produce this translational effect.

Leishmania major and Trypanosoma cruzi are two examples of eukaryotic


protozoa that completely lack the RNAi pathway. Several fungi, including the
model organism, Saccharomyces cerevisiae, lack any or all of the components

16.5.5 Biological Functions of RNAi


Immunity

RNAi is a crucial component of the immune system's defence against viruses


and other foreign genetic material, particularly in plants where it may
additionally stop transposons from self-replicating. Many dicer homologs are
expressed by plants like Arabidopsis thaliana, which are adapted to respond
differently to various viruses. It became apparent that induced gene silencing
in plants affected the whole plant and could be transferred by grafting from
stock to scion plants before the RNAi pathway was well understood. Since
then, it has been determined that this process is an attribute of the plant
immune system that allows the entire plant to respond to a virus after the first
localised contact. As a result, several plant viruses have evolved highly
effective defences against the RNAi response. These include viral proteins
that bind dicer-produced small dsRNA fragments with single-stranded
overhang ends. In some plant genomes, endogenous siRNAs are also
produced in response to bacterial infection. These effects might be a
component of a broader response to pathogens that inhibits any host
metabolic activity that promotes infection.

RNAi can cause an antiviral response in some animals, despite the fact that
plants typically express more variations of the dicer enzyme than mammals.
RNAi is crucial for antiviral innate immunity in juvenile and adult Drosophila
and is effective against viruses like the Drosophila X virus. Worms that
overexpress components of the RNAi process are immune to viral infection
and produce higher levels of argonaute proteins in response to viruses.

There is a paucity of information on the function of RNAi in mammalian innate


immunity. There may be evidence for an RNAi-dependent mammalian immune
response in the presence of viruses that can disrupt the RNAi response in
mammalian cells, regardless of the theory's criticism that it lacks sufficient
proof. It has been demonstrated that mammalian cells have a functioning
antiviral RNAi pathway.

Downregulation of genes

The two main functions of endogenously expressed miRNAs are


morphogenesis timing and maintaining undifferentiated or partially
differentiated cell types, such as stem cells, which include both intronic and
intergenic miRNAs. This involves the regulation of development and
translational repression. Since the bulk of the genes controlled by miRNAs in
plants are transcription factors, their activity is extremely broad and controls
the expression of important regulatory genes, such as transcription factors and 147


Block 4 Applications of Genomics and Proteomics
F-box proteins, to regulate whole gene networks during development. miRNAs
are connected to the development of cancers and the disruption of the cell
cycle in many species, including humans. Here, miRNAs can act as both
tumor suppressors and oncogenes.

Evolution
Parsimony-based phylogenetic analysis suggests that an early RNAi pathway
was likely already present in the most recent common ancestor of all
eukaryotes; the absence of the process in some eukaryotes is regarded to be
a derived trait. This primitive RNAi system contained at least one dicer-like
protein, one argonaute, one PIWI protein, and one RNA-dependent RNA
polymerase, possibly with other biological functions. These elements were
most likely present in the eukaryotic crown group and may have closer
functional connections with RNA degradation mechanisms like the exosome.
This is supported by a large-scale comparative genomics investigation.
Additionally, this research reveals that the RNA-binding argonaute protein
family is homologous to and originally descended from components of the
translation initiation machinery and is found in eukaryotes, the majority of
archaea, and some bacteria (including Aquifex aeolicus).

16.5.6 Applications of RNAi


RNAi pathway for gene knockdown
Gene knockdown is a technique for reducing the expression of particular
genes in an organism. This is accomplished by the naturally occurring RNAi
mechanism. A double-stranded siRNA molecule that is generated with a
complementary sequence to the gene of interest is used in the gene
knockdown method. Once the Dicer enzyme begins to break down siRNA, the
RNAi cascade commences. The procedure results in the degradation of the
mRNA and the elimination of any instructions needed to produce specific
proteins. Using this technique, researchers can partially, but not completely,
suppress the expression of a particular gene.

Medications
The strategy of using RNAi treatments to shut down genes has proven to be
effective, as demonstrated by randomised controlled clinical trials. The
treatments in this class, which are expanding, work by reducing the expression
of the proteins that particular genes are able to encode using siRNA. To date,
regulatory agencies in the US and Europe have authorized four RNAi drugs:
patisiran (2018), givosiran (2019), lumasiran (2020), and inclisiran (2020 in
Europe with anticipated US approval in 2021).

While all of the RNAi therapeutics that have been currently approved by
regulatory bodies targeting liver-related illnesses; other drugs that are still in
the research phase focus on a variety of other conditions, such as cystic
fibrosis, cardiovascular problems, carcinoma, bleeding issues, gout, alcohol
use disorders, and eye problems.

Delivery mechanisms
For RNA interference to achieve its therapeutic potential, siRNA must be
148 efficiently delivered to the cells of the target tissues. However, a number of


Unit 16 Expression Analysis of Genome
obstacles need to be removed before it may be applied therapeutically. For
instance, "naked" siRNA is prone to a number of challenges that lower its
therapeutic efficiency. Additionally, bare RNA can activate the innate immune
system and be destroyed by serum nucleases once siRNA has reached the
circulation. Unmodified siRNA molecules cannot easily cross the cell
membrane because of their size and extremely polyanionic (carrying negative
charges at several locations) nature. Therefore, siRNA needs to be
synthetic or enclosed in nanoparticles. If therapeutic dosages are not adjusted,
siRNA transport across the cell membrane may result in unexpected toxicities,
and siRNAs may have off-target effects (e.g., unexpected suppression of
genes with partial sequence complementarity). Since their effects are
diminished with each cell division, frequent treatment is necessary even after
they have entered the cells. Lipid nanoparticles and conjugates are two
strategies that aid in siRNA distribution to target cells in response to these
possible problems and hurdles.
Lipid nanoparticles
The core of lipid nanoparticles (LNPs) is modelled following liposomes, which
are lipid shell-encased aqueous cores. Large unilamellar vesicles (LUVs),
which can be 100 nm in size, are the resting place for a subset of liposomal
structures utilized to carry medications to the target tissues. Plasmids,
CRISPR, and mRNA are examples of LNP delivery systems that may be used
to encase nucleic acids.
Conjugates
Targeted delivery for RNAi therapies using siRNA conjugates is an alternative
to LNPs (e.g., aptamers, carbohydrates, GalNAc, peptides, antibodies). In
addition to other cardiometabolic disorders including hypertension and non-
alcoholic steatohepatitis, therapeutics utilizing siRNA conjugates that are
developed for uncommon or inherited diseases such as hemophilia, acute
hepatic porphyria (AHP), hereditary ATTR amyloidosis (NASH), and primary
hyperoxaluria (PH).

Biotechnology
There have been numerous documented other applications for RNAi, such as
the manufacturing of insecticides, crops, and food. The RNAi pathway has
produced a wide range of products, including nutrient-fortified plants, arctic
apples, decaffeinated coffee, nicotine-free tobacco, and hypoallergenic crops.
A variety of new products could be made with the help of RNAi based
technology.
Viral infection
The creation of two unique antiviral therapies was one of the first uses of RNA
interference in medicine. The first type targets viral RNAs. Targeting viral
RNAs has been shown in multiple studies to reduce the replication of multiple
viruses, including adenovirus, hepatitis A, HIV, HPV, hepatitis B, SARS
coronavirus respiratory syncytial virus (RSV), SARS-CoV, influenza virus, and
measles virus. Targeting the host cell's genes is the second strategy used to
stop early viral invasions. For example, blocking the chemokine receptors
(CXCR4 and CCR5) can stop HIV entry. 149


Block 4 Applications of Genomics and Proteomics
Cancer

Conventional chemotherapy can kill cancer cells with effectiveness, but since
it lacks the ability to differentiate between normal and malignant cells, it
frequently results in serious side effects. Various studies have shown that
RNAi can give a more targeted method of preventing tumor growth by
targeting genes relevant to cancer (i.e., oncogene). Additionally, it has been
proposed that RNA interference (RNAi) may increase the susceptibility of
cancer cells to chemotherapeutic agents, providing a complementary
therapeutic approach to chemotherapy. Inhibiting cell invasion and migration is
yet another possible RNAi-based therapy. RNA interference therapies treat
cancer by suppressing particular genes that promote malignancy. By
complementing the cancer genes with RNA interference (RNAi), for example,
by keeping the mRNA sequences consistent with the RNAi drug, this is
achieved. RNA interference (RNAi) sequences should ideally be chemically
altered to enhance their ability to bind to cancer cells. RNAi uptake is
regulated and monitored by the kidneys.

Neurological diseases

Neurodegenerative illnesses may potentially be treated using RNAi


techniques. The amount of A peptide, which is connected with the origin of
Alzheimer's disease, may be considerably lowered by selectively targeting
Amyloid beta-producing genes (such as BACE1 and APP) via RNA
interference, according to studies conducted in cells and mice. Additionally,
these silencing-based methods for treating Parkinson's illness and
polyglutamine disorder show encouraging outcomes.

Transgenic plants

Transgenic crops express dsRNA that has been carefully chosen to silence
important genes in insect targets. These dsRNAs are exclusively meant to
affect insects that express specific gene sequences. As a proof-of-concept, a
2009 study showed that the dsRNAs could kill any one of four species of fruit
flies while harming none of the others.

Insecticides

RNAi is being developed as a pesticide using a range of methods, including


genetic modification and topical treatment. Certain insects' midgut cells absorb
the dsRNA molecules as part of the environmental RNAi mechanism. Some
insects experience systemic effects as a result of the signal reaching every
part of their body (referred to as systemic RNAi).

There are no harmful effects in animals given RNAi at dosages millions of


times higher than what is anticipated for human exposure. RNAi affects
various Lepidoptera species (butterflies and moths) in different ways. To
understand how RNAi functions in certain taxa of insects, models such as
Drosophila spp., Spodoptera spp., Locusta spp., Tribolium castaneum,
Helicoverpa armigera, Nilaparvata lugens, Bombyx mori, and Apis mellifera
have been extensively used. While Glossina morsitans contains three Ago2
150 genes, Musca domestica only has two.


Unit 16 Expression Analysis of Genome
Food

RNAi has been used to genetically modify plants such that they produce less
natural plant toxins. These methods make use of the RNAi phenotype that is
persistent and heritable in plant populations. Cotton seeds are a good source
of dietary protein; however, they should not be consumed by humans since
they naturally contain the hazardous terpenoid gossypol. A critical enzyme in
the production of gossypol, delta-cadinene synthase, has been reduced in
cotton stocks using RNA interference (RNAi), without impacting the production
of the enzyme in other parts of the plant, where gossypol is crucial for
protecting plant against pest damage. The amounts of allergens in tomatoes
have been successfully reduced through development efforts, and plants have
been fortified with nutrient-rich antioxidants.

Stimulation of immune response

The innate immune system, which may be further separated into acute
inflammatory responses and antiviral responses, is in charge of controlling
siRNA. Small signaling molecules known as cytokines provide messages that
trigger the inflammatory response. Tumor necrosis factor (TNF-), interleukin-6
(IL-6), interleukin-1 (IL-1), and interleukin-12 (IL-12) are a few of them.
Inflammation and antiviral responses produced by the innate immune system
result in the release of pattern recognition receptors (PRRs). These receptors
aid in classifying infections as bacterial, fungal, or viral. More PRRs should be
included in siRNA and the innate immune system in order to assist it in
identifying various RNA structures. In the event of an infection, the siRNA is
therefore more likely to trigger an immunostimulant response.

6$4
6$4
Fill in the blanks:

a) The ………..………. is the procedure by which the targeted mRNA


molecules are neutralized by RNA molecules to block the targeted gene
expression.

b) The ………..………. and ………..……. are two classes of tiny ribonucleic


acid (RNA) molecules that are essential to the RNAi pathway.

c) A miRNA is produced in the cell nucleus from a much longer RNA-


coding gene as a primary transcript called a ………..………. .

d) The ………..……… was the first human virus shown to express miRNAs.

16.7 SUMMARY
• A test gene with quantifiable expression is known as a reporter gene. It
can be present on plasmids that have their T-DNA integrated into the
genome of a cell. Examining how those genes are expressed after a cell
has undergone a transformation is important. 151


Block 4 Applications of Genomics and Proteomics
• Reporter genes are those sequences that may be examined to ascertain
how altered genes are expressed. It is possible to do a reporter gene
test by calculating the total amount of protein synthesised. They often
have luminous properties and provide visual clues for precise estimates.
Examples include, Green fluorescent proteins, luciferase, octopine
synthase, etc.

• The method through which a gene in a cell is activated to produce RNA


and proteins is called gene expression. The RNA or the protein
generated from the RNA, as well as the function of the protein in a cell
can be used to quantify gene expression.

• RNA interference has been called post-transcriptional gene silencing,


transgenic silencing, and quelling in the past.

• RNA interference, also known as "RNA-mediated interference," is a


method for RNA-guided control of gene expression. It involves double-
stranded RNA suppressing the gene expression with similar nucleotide
sequences.

• The enzyme dicer, which breaks down double-stranded RNA (dsRNA)


into 20–25 base pair-long short double-stranded fragments, starts the
RNAi pathway. The RNA-induced silencing complex (RISC) is
subsequently formed by incorporating one of each fragment's two
strands, referred to as the guide strand, which base pairs with
complementary sequences.

• Post-transcriptional gene silencing occurs when a messenger RNA


(mRNA) molecule pairs with the guide strand base, causing the mRNA
to be cut down by argonaute, which is the RISC's catalytic component.

• When created by RNA-coding genes in the cell's own genome, short


RNA fragments are referred to as microRNAs (miRNA) and small
interfering RNAs (siRNA) accordingly.

• Many model species, including Drosophila melanogaster, Arabidopsis


thaliana and Caenorhabditis elegans have been used to study the RNAi
process.

• Both in cell culture and in live animals, the selective and powerful impact
of RNAi on gene expression makes it an invaluable research tool.
Synthetic dsRNA put into cells can cause the suppression of the target
genes of interest.

16.8 TERMINAL QUESTIONS


1. What are reporter genes? Describe their role in transformation and
transfection assays.

2. Decribe the role of reporter genes in gene expression and promoter


assays.

3. Briefly explain about the type of RNAs that play role in RNA
152 interference?


Unit 16 Expression Analysis of Genome
4. Explain the concept of temporal and site-specific gene expression and
their analysis.

5. Briefly explain the mechanism of gene silencing.

6. Define RNA interference. Describe the mechanism and biological


functions of RNAi.

7. Describe few applications of RNAi.

16.9 ANSWERS
Self-Assessment Questions
1. a) selectable markers, b) beta-galactosidase, c) chloramphenicol
acetyltransferase (CAT)

2. a) gene expression, b) spatiotemporal, c) in situ hybridization

3. a) RNA interference, b) microRNA, c) co-suppression,


d) quelling, e) Dicer, f) RNA-induced silencing complex
(RISC)

4. a) RNA interference (RNAi), b) Small interfering RNA (siRNA);


microRNA (miRNA), c) pri-miRNA, d) Epstein-Barr virus (EBV)

Terminal Questions
1. Refer to Section 16.2 and Subsection 16.2.1.

2. Refer Subsections 16.2.2 and 16.2.3.

3. Refer Subsection 16.5.2.

4. Refer to Section 16.3.

5. Refer to Section 16.4.

6. Refer to Section 16.5, Subsections 16.5.1 and 16.5.5.

7. Refer Subsection 16.5.6.

153




UNIT 17
3527(20($1$/<6,6$1'
$33/,&$7,212)3527(20,&6
$33/,&$7,212)3527(20,&6

6WUXFWXUH
6WUXFWXUH
17.1 Introduction Stable Isotope Labeling by
Amino Acids in Cell Culture
Objectives
(SILAC)
17.2 Origin of Proteomics
17.5 Application of Proteomics
17.3 Classes of Proteomics
Pharmaceutical Field
Profiling Proteomics
Drug Discovery
Functional Proteomics
Drug Development and
Chemoproteomics Toxicology
Phosphoproteomics Phage Antibody as Tool
17.4 Techniques of Proteomics 17.6 Summary
Two Dimensional Gel 17.7 Terminal Questions
Electrophoresis (2DE)
17.8 Answers
Two Dimensional Difference
Gel Electrophoresis (2D-DIGE)

Isotope Coded Affinity Tag


(ICAT)

17.1 INTRODUCTION
Proteins are the molecules that play various roles in the biological system and
proteomics is the study of complete set of proteins at a time. With the advent
of technologies in the early seventies, genome sequencing was gaining
attention among scientists. The gene sequence cannot provide information
related to protein function, localization, post-translational modifications,
relative expression in different cell organelles, protein-protein interaction, etc.
The human genome has approximately 31,000 protein-encoding genes but the
protein products generated are estimated to be close to 1 million. This
indicates that the functional information in genes is actually located in the
proteome. Therefore, understanding of ‘proteome’ means the complete set of
proteins within the cell is important. Hence, the term ‘proteomics’ which was
Unit 17 Proteome Analysis and Application of Proteomics
first coined in 1995 refers to the study and characterization of complete set of
proteins in a cell, tissue or organism. This study plays an important role in
biomarker identification, drug discovery, disease pathogenesis, identification of
drug targets for various diseases, and so on.

2EMHFWLYHV
2EMHFWLYHV
After studying this Unit you would be able to:

™ describe proteomics,

™ explain various forms of proteomics,

™ comprehend the techniques used to study proteomics,

™ discuss the role of proteomics in drug development and toxicology, and

™ enumerate its application in drug discovery in humans and


pharmaceutical industry.

17.2 ORIGIN OF PROTEOMICS


It started way back in 1975 when two-dimensional gel electrophoresis was
introduced by O'Farrell, Klose and Scheele. Proteins from mouse, guinea pig
and Escherichia coli were studied. Proteins were successfully separated and
visualized using silver stain, however, they could not be identified and this was
a major drawback. However, the development of techniques for protein
identification like sequencing of proteins by Edman degradation followed by
mass spectrometry was a major breakthrough.

17.3 CLASSES OF PROTEOMICS


Based on the applications; proteomics can be divided into the following types:

17.3.1 Profiling Proteomics


Proteomics which helps to identify proteins expressed in any biological sample
eg. blood, other body fluids, tissues, at some specific time or the proteins that
are differentially expressed proteins (DEPs) between two samples: healthy
and diseased is termed profiling proteomics. This is also called protein
expression profile and protein signature. Specific proteins may be upregulated
(overexpressed) or downregulated (underexpressed) which may be the reason
for disease pathogenesis. This protein could be a cell cycle protein, signaling
protein, metabolic protein, etc. Profiling proteomics is done using 2-
dimensional electrophoresis (2DE) /mass spectrometry technique. One
example is the identification of breast cancer biomarkers.

17.3.2 Functional Proteomics


Understanding protein function is crucial in biology. This requires the
understanding the underlying molecular mechanisms and identifying the
biological role of the unidentified proteins. The interaction of an unknown
protein with partners belonging to a specific protein complex could indicate 155


Block 4 Applications of Genomics and Proteomics
crucial biological functions. Studying protein-protein interaction by a two-hybrid
approach is common in functional proteomics. For example, X protein is
interacting with A/B/C/D proteins within the cell. We need to identify which
protein is A/B/C/D. For this immunoprecipitation technique can be used. An
expression construct is generated with gene X along with a tag such as FLAG,
GFP, or c-myc, etc. When it is transfected into the cell (bacteria/fungi/yeast
/mammalian) it expresses fusion protein. This is used as bait or ligand to catch
the prey which is A/B/C/D proteins in this case. Following fusion protein
expression in the host, this protein interacts with A/B/C/D proteins; in order to
decipher the process in which this protein is involved, you must identify this
complex. The cell extract is immunoprecipitated with anti-tag antibodies (anti-
FLAG/anti-GFP/anti-cmyc etc.). The protein components are eluted and
separated by Sodium dodecyl-sulfate polyacrylamide gel electrophoresis
(SDS-PAGE). The protein bands are subjected to in-gel digestion or are
extracted and trypsinized. The interacting proteins A/B/C/D are identified after
the peptide mixtures are separated using the liquid chromatography with
tandem mass spectrometry (LC-MS-MS) approach. Thus, the integration of
molecular biology, protein tagging, immunoprecipitation, and mass
spectrometry facilitates the high-throughput analysis of protein complexes that
are produced within cells.

17.3.3 Chemoproteomics
Understanding the mechanism of action of drugs and small molecules remains
one of the biggest challenges in chemistry and biology sciences. Chemical
proteomics is a new field in chemical biology that aims to design small
molecules to understand protein function. Chemical proteomics is used to
identify the protein binding partners or targets of small molecules in live cells.
In this approach instead of using proteins as bait, small molecules/drugs are
used as bait to look out for interacting proteins.

17.3.4 Phosphoproteomics
Post-translational modifications of proteins for example, phosphorylation,
acetylation, ubiquitination and SUMOylation takes place within the cell after
the translation process. These modifications are essential to regulate protein
activation/inactivation and protein-protein interaction in cells. Phosphorylation
of serine, threonine and tyrosine residues of proteins play a very crucial role
as it can switch on or switch off the function of protein. This is seen in
transcription factors where phosphorylation of some factors activates the
protein and dephosphorylation deactivates it. Understanding the
phosphoproteome is crucial in biology as it provides better insights into protein
function and regulation. It's important to comprehend which proteins within the
cell are phosphorylated as well as how a particular protein site influences
the protein interactions with other proteins. Phosphorylation mapping by mass
spectrometry (MS) has helped in the understanding of phosphoproteome.

6$4
6$4
a) State whether these statements are “True” or “False”:

156 i) Proteomics is the study of genes.


Unit 17 Proteome Analysis and Application of Proteomics
ii) 2-dimensional electrophoresis is a technique to study proteomics.

iii) In Chemiproteomics, small molecules/drugs are used to study


interacting partners.

iv) Phosphoproteomics involves the study of phosphorylated proteins.

b) Fill in the blanks:

i) The term ……………….. which was first coined in 1995 refers to


the study and characterization of complete set of proteins in a cell,
tissue or organism.

ii) The ……………. includes the description of the whole proteome of


a cell, tissue, organ or organism and comprises organelle mapping
and differential measurement of expression levels between cells or
conditions.

iii) The …………….. aims to characterize protein activity, by


determining protein interactions and the presence
of posttranslational modifications.

iv) The field of …………… seeks to characterize interactions between


small molecules and their protein targets.

v) The ……………… deals with the large-scale analysis of protein


phosphorylation sites to define signaling network regulation and
dysregulation in normal and pathological conditions.

17.4 TECHNIQUES OF PROTEOMICS


There are several techniques which are used to study proteomics. A few of
them are mentioned below:

17.4.1 Two Dimensional Gel Electrophoresis (2DE)


It is the most widely used technique in proteomics and can be used for the
separation of complex protein mixtures from biological samples, cells, tissue,
etc. Two dimensional gel electrophoresis involves two distinct steps to
separate proteins (Fig. 17.1). The first step is isoelectric focusing (IEF), in
which proteins are separated on the basis of isoelectric point (pI) and the
second step involves Sodium dodecyl-sulfate polyacrylamide gel
electrophoresis (SDS-PAGE) in which proteins are separated on the basis of
molecular weight.

Two protein molecules with the same molecular weight cannot be


distinguished using SDS-PAGE, so with the use of 2DE, we can resolve two
proteins having similar molecular weights. 2DE is based on two properties of
proteins: isoelectric point and molecular weight. Proteins can have positive,
negative and zero charge based on the pH of the environment. A pH at which
protein has zero net charge is called the isoelectric point. Below the isoelectric
point, the protein has a net positive charge, while above the isoelectric point,
the protein has a negative charge. Isoelectric focusing is done on strips that 157


Block 4 Applications of Genomics and Proteomics
have a pH gradient. Under the influence of an applied electric field, a protein
migrates towards the electrode with an opposite charge to that of the protein. It
will migrate until it reaches a point on the strip where the pH of the strip is
equal to the isoelectric point of the protein, as the protein shows no migration
after reaching its isoelectric point.

Fig. 17.1: Steps in 2-Dimensional Gel electrophoresis (2DE)

The second step is SDS-PAGE, which is based upon separation of proteins


and peptides on the basis of molecular weight. Polyacrylamide gel is used to
separate proteins. Sodium dodecyl sulfate (SDS) is an anionic detergent that
binds to protein and denatures it to produce a linear polypeptide chain. The
protein-SDS complex has a negative charge so under the influence of the
applied field, the complex migrates toward the anode. The protein gets
resolved due to the sieving effect of the resolving gel (Fig. 17.2). The resolved
protein can be stained with commassie brilliant blue or with silver stain. Silver
staining is highly sensitive and can detect up to 0.1-1 ng of protein. The silver
staining method is based upon the reduction of silver at the site close to the
protein molecule, to insoluble metallic silver producing a brown-black color.
Coomassie dye binds to the basic and hydrophobic region of the protein and
changes its color from reddish brown to blue. This method can detect up to 8-
10 ng of protein.

Fig. 17.2: Resolving protein in 2D Gel electrophoresis

Evaluation of 2D Gel image


Identification of proteins separated in gel is carried out by various techniques
like mass spectrometry (MS), Matrix-assisted laser desorption/ionization
158 (MALDI), Matrix-assisted laser desorption/ionization-time of flight (MALDI-


Unit 17 Proteome Analysis and Application of Proteomics
TOF) mass spectrometry (MS), tandem time-of-flight (TOF/TOF) mass
spectrometer (TOF/TOF MS), Electrospray ionization mass spectrometry (ESI-
MS/MS). Software like Melanie, PDQuest, Proteomweaver, Decyder 2D,
Progenesis, REDEFIN etc. are able to identify differentially expressed proteins
in 2DE gel. The differentially expressed proteins are marked on gel and
excised manually. Excised gel with spots are washed and destained. In-Gel
digestion of protein is carried out and samples are further spotted on MALDI
target plate. Sample is analyzed by using mass spectrometry and mass
spectra are obtained. The spectra are submitted to software like MASCOT 1.9
for database search against the National Center for Biotechnology Information
(NCBI) Database and the protein whose expression is changed is identified.
Application and utilities of 2D Gel electrophoresis
It is a powerful technique for proteome analysis and has capability to resolve
thousands of protein at once. Various applications of 2DE are:

• Whole proteome analysis

• Detection of biomarkers
• Drug discovery and cancer research

• Protein characterization
• Study of post-translational modification
• Protein-protein interaction

17.4.2 Two Dimensional Difference Gel


Electrophoresis (2D-DIGE)
Derivatization of protein with fluorophore in complex protein mixture before IEF
and SDS-PAGE, allows differential analysis of proteins in samples. The 2D-
DIGE technique is based on separation of multiple protein samples on same
2D Gel. Fluorescent cyanine dyes (CyDye) are used to label protein from
different samples (Table 17.1). CyDye DIGE fluors are available in two forms:
minimal and saturation dye. Minimal dyes are used where sufficient amount of
sample is available. Saturation dye is used where small amounts of sample
are available. The fluorescent labeled samples are mixed in equal ratio and
separated on the same gel. Same protein from different samples will migrate
to the same position. These proteins are differentiated on the basis of different
flurophore labeled dyes. The fluorophore used to tag two samples has
different fluorescent emission spectra but has identical mass and
electrophoretic mobility. The intensity of these fluors is measured to check for
the differential expression of a specific protein in two different samples viz.
sample from healthy and diseased individual.
Table 17.1: Dyes used in 2D DIGE technique.

Name of Dye Dyes available Amino acids Sensitivity


commercially labeled

CyDye DIGE Fluor Cy2, Cy3, Cy5 Lysine residues Similar to silver
minimal dye staining

CyDye DIGE Fluor Cy3, Cy5 Cysteine residues 100 times to silver
staining
Saturation dye
159


Block 4 Applications of Genomics and Proteomics
Use of internal standard decreases the gel-to-gel variation and is used to
match and normalize the protein patterns across different gels. Two samples
mixed in equal amounts and labeled with a third dye other than the dye used
for labeling samples, serve as an internal standard. The internal standard is
mixed with protein samples and separated on gel. A fluorescence image is
captured on a multiwavelength scanner and image analysis is carried out. The
relative intensity of the labeled sample protein spot of two test samples are
compared to the intensity of the corresponding spot in standard. DeCyder
software allows the identification of spots, co-detection of spots, spot volume
ratio, etc. The steps involved in 2D DIGE are shown in Figure 17.3.

Fig. 17.3: Steps involved in two-dimensional difference gel electrophoresis


(2D-DIGE)

Advantages of 2D-DIGE

• Lower experimental variation

• Image produced in less than one hour

• Saturation dye is about 100 times more sensitive than silver staining

• Use of internal standard eliminates experimental gel-to-gel variation and


allows accurate quantification

17.4.3 Isotope Coded Affinity Tag (ICAT)


Isotope-coded affinity tag (ICAT) technique was developed by Gygi in 1999 for
differential expression proteomics. ICAT is a labeling method for proteins
containing cysteine residues. This is an in-vitro method for quantitative
proteomics that uses chemical labeling agents. With the help of this technique,
high coverage, high throughput and high accuracy can be achieved. ICAT
technique can quantify the protein from a complex mixture which the gel-
based technique could not do. ICAT is a useful technique for the identification
and accurate quantification of proteins in complex mixtures. This technique
can also be used to study protein expression changes and protein
modification. There are two types of reagents available: the light form and the
heavy form. The structure of the ICAT reagent structure is depicted in Figure
160 17.4.


Unit 17 Proteome Analysis and Application of Proteomics

Fig. 17.4: ICAT Reagent structure

The light form is also known as the normal form and the heavy form is known
as the deuterated form. In the heavy form, a hydrogen atom is replaced with a
deuterium atom. The deuterium atom is absent in light form and this results in
mass difference between the light form and heavy form. A mass difference of
at least 5 Da is desirable to allow tagged-peptide ion separation. The isotope-
coded affinity tag reagent consists of three elements, that is, an affinity tag,
linker, and thiol reactive group. The affinity tag (biotin) is used for the isolation
of ICAT-labelled peptides with the help of avidin affinity chromatography. The
linker part helps to form a stable isotope and generates the mass difference
and cysteine residue which is present in protein, covalently forms a complex
with thiol-reactive group and they can be recovered from the mixture of
protein. The ICAT reagent does not change the property of the protein after
labeling.

ICAT Workflow

ICAT consists of 4 main steps such as 1) lysis and labelling 2) proteolysis 3)


isolation of peptides by affinity chromatography 4) identification and
quantification. In the ICAT analysis technique, only two samples of protein can
be quantified because of the availability of two reagents (Fig. 17.5).

• Lysis and labeling: Protein samples that contain cysteine chains are
isolated from cells by various methods like cell lysis by freeze-thaw,
sonication, etc. and are labelled. In tagging or labeling, one protein
sample is tagged with the isotopically light form of the ICAT reagent.
Here the cysteinyl residue forms a complex with the thiol reactive group.
Another protein sample is tagged with isotopically heavy reagent.

• Proteolysis: Then both the samples are combined in the ratio of 1:1 and
the proteolysis is done in the presence of proteolytic enzymes like
trypsin etc. for the formation of peptide fragments.

• Isolation of peptides by avidin affinity chromatography: In this step


with the help of avidin affinity chromatography, the tagged cysteine-
containing peptides are isolated. Here the interaction between
immobilized avidin on the column and biotin occurs.

• Identification and quantification by mass spectrometry: This is the


most crucial step. The tagged peptides are quantified and identified with 161


Block 4 Applications of Genomics and Proteomics
the help of micro-capillary high-performance liquid-mass spectrometry.
Protein quantification is accomplished by comparison of integrated peak
intensities. The ratios of the maxima of lower and upper mass
components provide a precise estimate of the relative abundance of
peptides.

Fig. 17.5: ICAT Workflow

Applications

Applications of ICAT are listed below:

• Identification and quantification of protein

• Analysis of protein expression changes

• Analysis of protein dynamics of complex protein

17.4.4 Stable Isotope Labeling by Amino Acids


in Cell Culture (SILAC)
Stable isotope labeling by amino acids in cell culture is abbreviated as SILAC.
SILAC was developed by Ong and colleagues in 2002 that detects proteome
changes under differential treatment. SILAC is a mass spectrometry-based
approach that identifies changes in protein abundance across samples by
employing non-radioactive isotopic labeling. SILAC plays an important role in
characterizing the changes in the proteome within the different biological
samples, determines that the changes in protein occur due to the post-
translational modifications and compares specific interactions with the other
proteins inside the cell. Stable isotope-labeled amino acids are directly added
to the cell culture which helps in comparing the set of proteins produced within
the cell in different cellular states.

SILAC is a method that involves the incorporation of stable isotopes 13C and
15
N. Two different cell populations are grown in two different culture media
162 which are referred to as light medium and heavy medium respectively. The


Unit 17 Proteome Analysis and Application of Proteomics
light medium contains amino acids that are labeled with natural isotopes (12C,
14
N) whereas the heavy medium contains amino acids that are labeled with
stable isotopes (13C, 15N). After a sufficient number of cell divisions, the cells
are cultured in a heavy medium. Proteins derived from cells grown in heavy
media are now in a heavy state. The number of cell divisions required for the
complete labeling of proteins depends upon the rate of protein synthesis,
metabolism, degradation and turnover. Prior to quantification, the labeling
efficiency of the proteins should be tested. Labeled and unlabeled protein
extracts are mixed in a ratio of 1:1. The samples are then digested with the
help of trypsin to small peptides and they are analyzed with the LC-MS/MS
technique. The intensity of the signals from light and heavy samples allows for
a quantitative assessment of their relative abundance in the mixture (Fig.
17.6). Leucine, lysine and methionine are the essential amino acids that have
been used in SILAC. Though arginine is not an essential amino acid still
arginine has also been used in SILAC because it is essential for the growth of
some cells.

Fig. 17.6: Principle of stable isotope labeling by amino acids in cell culture

SILAC workflow

SILAC workflow consists of two phases namely the adaptation phase and the
experimental phase. In the adaptation phase, the cells are subjected to growth
in labeled and unlabeled media until the heavy amino acids have been
completely incorporated into the cellular proteins. The level of SILAC amino
acid incorporation is then determined by LC-MS/MS. The area under the curve
(AUC) of the MS peaks for the remaining light and heavy peptide pairings is
used to assess the degree of labeling (Fig. 17.7a). 163


Block 4 Applications of Genomics and Proteomics

Fig. 17.7: a) Adaptation phase of SILAC

During the experimental phase, after the full incorporation of heavy amino
acids has been confirmed, the two cell populations are subjected to different
treatments based on the experiment and then combined equally prior to
optional subcellular organelle purification, cell lysis, protein extraction and
protein digestion. The samples are then examined using LC-MS/MS to identify
and quantify the heavy peptide to light peptide ratios (Fig. 17.7b).

Fig. 17.7: b) Experimental phase of SILAC

SILAC has wide applications in the analysis of proteomes as listed below:

• Expression proteomics

• Protein-protein interactions

• Dynamic changes of protein post-translational modifications

• Protein turnover

164 • Characterization of changes observed in proteome


Unit 17 Proteome Analysis and Application of Proteomics

6$4
6$4
Fill in the blanks:

a) The …….………. and …….………. are the two distinct steps to separate
proteins in the Two-dimensional gel electrophoresis technique.

b) A pH at which protein has zero net charge is called the …….………. .

c) Sodium dodecyl sulfate (SDS) is an …….………. detergent that binds to


protein and denatures it to produce a linear polypeptide chain.

d) In SDS-PAGE, the resolved protein can be stained with …….………. or


with …….………. .

e) The Fluorescent cyanine dyes (CyDye) are used to label protein from
different samples in the …….………. technique.

f) The …….………. is a labeling method for proteins containing cysteine


residues.

g) The …….………. is a mass spectrometry-based approach that identifies


changes in protein abundance across samples by employing non-
radioactive isotopic labeling.

17.5 APPLICATION OF PROTEOMICS


17.5.1 Pharmaceutical Field
The term proteomics was first introduced by Marc Wilkins in 1996. With the
help of this term, he represented the protein complement of the genome. The
term proteome means the categorization of the total protein composition of a
cell by its location, post-translational alteration and turnover at a specified
period. Due to the accuracy, sensitivity and clear resolution, the use of
proteomics in the drug discovery field is increased. Proteomics is used in
various parts of the pharmaceutical field such as chemical profiling, identifying
drug targets, drug ADME (absorption, distribution, metabolism and excretion),
safety and toxicity studies. Proteins are the main targets in drug discovery and
proteomics plays an important role in the identification of drug targets in the
drug development process. Studies related to changes in protein expression,
post-translational modification and protein-protein interaction in disease vs.
healthy individuals can be done using proteomic analysis. Proteomics also
provides an important tool for studying protein expression profiles following
drug treatment.

Research and development

In research and development, proteomics plays an important role and serves


as a powerful tool to examine biochemical processes at the basic molecular
level. Protein identification, characterization, quantification of whole cell
proteome and target identification can be easily done (Table 17.2). 165


Block 4 Applications of Genomics and Proteomics
Table 17.2: Drug targets that are identified by proteomics.

Disease Drug Target

Cancer Pyrido[2,3-d]pyrimidine Src,FGFR, RICK

Bisindolyl-Maleimide-III Adenosine kinase, CDK2, Prohibitin,


heme binding protein

Chronic myeloid Imatinib BCR-AB L, ABL,PDGFR,NQO2


leukemia

Bosutinib ABL, STE, SRC family kinases, TEC


family kinase, CAMK2G

Protein expression study has a crucial role in drug development and research.
Proteomics allows for the study of the impacts on protein expression and
pattern, as well as the mechanism of action of drugs, their toxicological and
therapeutic effects, and abnormal protein expression in specific conditions. In
proteomics, both covalent modification and processing can be studied at the
protein level and this plays an important role in disease biology and gives
important information about disease-specific conditions for the use of
diagnostic markers or therapeutic agents against that disease. The
identification of the protein of interest, confirmation and purity checking can be
done with the help of proteomics.

Production
The discovery of the diagnostic markers and vaccine can be done with the
help of proteomics and it is a promising tool for disease-associated biomarker
detection. For the development and production of biomarkers and vaccines,
different tools of proteomics can be used such as 2D-PAGE, MALDI-TOF,
surface-enhanced laser desorption, ionization (SELDI) and protein chip
techniques. For the study of proteins on a large scale, proteomics is used and
with the help of different methods such as ICAT, 2D gel electrophoresis,
SILAC etc. sample isolation, identification and quantification is done which is
crucial for its production. Proteomics not only assists with producing
chemically stable, highly specific products, but it also ensures and predicts the
quality of the final product.

Quality control and validation


Parameters like quality control (QC) and quality assurance (QA) are able to
detect inaccuracy, ensure consistency and avoid error. Adverse impacts of
missing QC in early proteomics research can be avoided if the effective quality
control operation like proteomics quality control (PTXQC) is used in various
processes such as downstream applications. To produce peptides, sample is
digested and analyzed by HPLC which separates the peptides based on
physicochemical quality. The eluate is subsequently ionized and the
mass/charge ratio of the peptides is determined. All these stages have an
impact on the quality of final spectra. In order to generate output for
proteomics quality control evaluation, the raw data is submitted to the
MaxQuant programme. If the output meets the quality criteria, the downstream
analysis is approved. MaxQuant parameters are adjusted to remove the error
166 indicated by PTXQC.


Unit 17 Proteome Analysis and Application of Proteomics
Safety

Early identification of side effects of new drugs as well as understanding of


their mechanism may improve drug safety and efficacy. The identification of
optimal targets and effective doses for patients can be analyzed with the help
of proteomics to reduce the unwanted side effects of the drugs. Sometimes
drug toxicity is a major cause of drug discovery and development failure so
during in-process control factors like protein structure changes, contamination,
protein-protein interaction, microbial control, detection of resistance and
protein-microbe interaction can be studied. Evaluation of the patient’s distinct
characteristics at the molecular level for the development of personalized drug
therapy to reduce the side effects and increase safety is also done.

Regulation related to proteomics

Proteomics data is expected to be submitted to the U.S Food and Drug


Administration (USFDA) more frequently in order to validate or in connection
with requests for the approval of medication. However, there is no set
standard for the creation, evaluation and submission of proteomics data that
will be examined by the regulatory bodies.

17.5.2 Drug Discovery


Proteomics along with computational methods play an important role in
discovering disease biomarkers, identifying and validating drug targets,
designing effective drugs and assessing the efficacy of drugs and patient
response. Proteomics along with the disciplines of biology, chemistry and
computational biology plays a significant role in drug discovery. Figure 17.8
shows important stages in drug discovery as depicted below:

Fig. 17.8: Important stages in drug discovery

Target identification
2-Dimensional Gel Electrophoresis along with mass spectroscopy helps in
identifying the protein expression changes within a particular system. Using
protein sequence tags (PST), each protein is terminally tagged, isolated and
sequenced which helps in rapid identification of any set of proteins produced
by a cell. Apart from this, multidimensional protein identification technology
(MudPIT) uses strong cation exchange and reverse-phase adsorbent
separation columns for the identification of protein targets through LC/MS
analysis whereas isotope-coded affinity tagging (ICAT) uses an ICAT reagent
that binds to a particular amino acid usually a cysteine, light or heavy isotope
and an affinity tag biotin are incubated with each sample. The samples of
different groups or disease states are incubated with different isotopes and are
mixed with equal proportions and lysed after lysing the labeled peptides
identified by LC/MS technique. Control and drug-treated samples are
subjected to proteomics for target identification. Proteins whose expression is
upregulated or downregulated are expected to be the target for that drug. 167


Block 4 Applications of Genomics and Proteomics
Proteins make up the majority of therapeutic targets that are used to initiate
drug design processes. Proteomics is a powerful tool for identifying targets by
thoroughly analyzing changes in protein expression and protein-protein
interactions that take place over the course of a disease or after therapeutic
treatment. Analyzing the proteome profiles of cells treated with a drug is a
typical target discovery method. Compared to untreated cells, the changed
proteins in the signaling pathway or gene network regulator are studied..
Proteomics based research has been carried out to look into the cellular
pathways that drugs work on as well as the molecular basis of
pharmacological activity. Proteomics has facilitated the study of the
mechanisms by which small-molecule medicines interact with the proteome
via two sophisticated procedures: thermal proteome profiling (TPP) and
multiplexed proteome dynamics profiling (mPDP). The mPDP technique
permits the finding of regulated protein synthesis and degradation processes
brought on by small molecules, while TPP evaluates changes in protein
thermal stability in response to drug treatment and so provides information on
direct targets and downstream regulation events.

Target validation

Once the target is identified, the next step in drug discovery is the validation of
the identified target. The validations are done by overexpression or knockout
of the gene of interest by homologous recombination in the organism. In some
cases, RNAi approach is also used. Phenotypic changes are observed to
evaluate its essentiality. Samples are put through proteomics after
overexpression or knockout to observe how these events affect various
proteins. The upregulation or downregulation of proteins and their identity are
revealed by proteomics which highlights the class of proteins (metabolic
protein, transporter, cell cycle protein, etc.) whose expression is altered. This
provides an understanding of the metabolic pathways/signaling pathways/cell
cycle proteins that may have been affected due to overexpression and
knockout studies and may predict the essentiality of the knockout gene.

Identifying protein modifications

Determination of post-translational modifications of protein is important in


investigating the processes of events within the cell such as cell division,
growth and differentiation. Some of the examples include the use of 2-DE to
assess the lymphocytes in a healthy individual and the lymphocyte in patients
with Scott syndrome. This showed changes in the tyrosine phosphorylation of
immunoglobulin chain precursor, fascin and actin-associated proteins.

Drug design and lead optimization

Structural proteomics plays an important role in drug design and lead


optimization. Once the target is validated, crystallization of the protein is
carried out. The structure of the protein is developed by multiple-wavelength
anomalous diffraction (MAD) phasing which has adjustable high-energy X-rays
with increased computational abilities. The protein's crystal structure can be
used to directly synthesise appropriate medications with the use of
bioinformatics techniques that will identify inhibitors that bind to the specific
protein target. Virtual screening software not only selects the drugs that bind to
168 the target protein but also optimizes the lead compound. The optimization of


Unit 17 Proteome Analysis and Application of Proteomics
the lead compounds is done by screening the structure-activity relationships
among the drugs. Apart from the virtual screening, activity-based probes
(ABP) also help in identifying the potential drug compounds with particular
proteins.

Clinical stages of drug discovery

Drug-Centric chemoproteomic profiling is an approach where the bioactive


molecule of interest is chemically conjugated to a suitable affinity moiety (e.g.,
biotin) or directly immobilised on a resin such as a sepharose bead. Chemical
synthesis of a suitable functionalized analogue of the compound is required in
both cases (typically bearing an amine, carboxyl, hydroxyl or sulfhydryl group).
To verify that the functionalized molecule maintains similar target binding and
biological activity properties, detailed information on the structure-activity
relationship (SAR) of a compound is required. After incubating the affinity
probe with cell extracts, the bound proteins are identified using mass
spectrometry.

17.5.3 Drug Development and Toxicology


Drug development is a complex process and involves various disciplines like
metabolomics, structural biology, bioinformatics and proteomics. Recent
development in the field of proteomics has brought many advances in the drug
development process. Proteomics helps in identifying new drug targets, mode
of action of drugs and their possible toxicity. The study of proteomics is crucial
for developing new drugs. Proteomics and bioinformatics can address the
needs of the pharmaceutical industry in identifying novel targets to understand
insight into drug action because most pharmaceuticals act by targeting
proteins or they are proteins themselves.

A biomarker is defined as “a characteristic feature that is objectively measured


and evaluated as an indicator of normal biological processes, pathogenic
processes or pharmacological responses to a therapeutic intervention”.
Biomarkers are studied and evaluated in many stages of drug development
and drug discovery. Various biological pathways are involved in disease
development and numerous protein expression pathways may get altered.
These alterations can be studied by evaluating physical properties like heart
rate and blood pressure, molecular biomarkers like gene expression and
protein level, and imaging techniques like magnetic resonance imaging (MRI).
Many biomarkers of diseases like cancer are identified by proteomics studies.
There are two types of biomarkers: prognostic and predictive biomarkers.
Prognostic biomarkers are used to identify the likelihood of disease
occurrence, progression, or recurrence while predictive is used to identify
people who are more prone or respond to a particular medical agent or
environmental agent.

Biomarkers should have high specificity for disease and proteomics offers
powerful techniques for biomarker identification, characterization and
validation. The process of confirming the assay, its performance
characteristics, and the necessary ideal conditions to produce reproducibility
and accuracy is known as analytical technique verification. Clinical or
biological validation is related to how a particular marker performs in a 169


Block 4 Applications of Genomics and Proteomics
population and between populations. The incorporation of validated proteomic
biomarkers into clinical drug development programs will improve the decision-
making process by adding critical information about the pharmacological and
pharmacodynamic mechanism of drug targets.

Proteomics in toxicology study

Toxicogenomics is the study of genome response to environmental stress or


toxins by genome-wide mRNA expression profiling (transcriptomics) and
protein expression profiling (proteomics). The interaction between gene
dysfunction and disease development is studied by toxicogenomics analysis.
Proteomics analysis can be done to study a) the composition of sub-cellular
structures like mitochondria and nuclei to understand differential protein
expression in disease vs. normal state b) Biomarker identification c) Analysis
of post-translational modification like glycosylation and phosphorylation etc.
Safety assessment of drugs is one of the most important steps in drug
development. It can be done in live animals, but the main problem associated
with in vivo studies is the generation of false negative results. Proteomics
provides an alternative approach to studying the toxic effects of drugs.
Analysis of proteins in secreted fluids like blood, urine, CSF etc. are carried
out to find specific biomarkers for detection of toxicity.

Proteomics studies for hepatotoxicity

Drugs are metabolized and eliminated in the liver and it is often the most
targeted organ for studying toxicology. Hepatotoxicity is dose dependent and it
is studied in 28 days in vivo. Hepatotoxicity can be observed in later stages of
drug development and may cause hazards, so early detection by proteomics
study will help us in managing the hazard. Proteomics studies of functional
molecules will give insight into possible mechanisms of action. In vitro models
for hepatotoxicity are based on cell lines like HepG2, HepaRG and
hepatocytes. Hepatocytes are most widely used to study drug metabolism and
toxicity as they are capable of biotransforming drugs. After administration of a
drug, the expression of liver-specific proteins is checked. Also changes in the
level of CYPs 2B, 1A is checked by 2D Gel electrophoresis. In vivo analysis is
based on the use of test animals like rodents. About 28-90 days repeated
dose toxicity tests are carried out to observe the chronic effect, organ toxicity,
differential protein expression, etc. Proteomics investigations can be
performed on tissue, cellular fractions, plasma proteome, etc. An example of
hepatotoxicity studies using proteomic endpoint in human is the administration
of test compound acetaminophen, amiodarone and cyclosporine A. Protein
expression changes in HepG2 was studied using DIGE and mass
spectrometry. A total of 254 differentially expressed proteins were identified
and analyzed. High differential expression of secreted proteins such as serum
albumin, ApoA1, serotransferrin, and ER-Golgi transport network was
observed.

Proteomics studies for pulmonary toxicity

Lung tissues are collected and analyzed using iTRAQ technique. Changes in
170 the level of oxidative stress proteins and inflammatory mediators are studied.


Unit 17 Proteome Analysis and Application of Proteomics
Proteomics studies for heavy metal toxicity
The toxic effect of heavy metals on protein expression can be studied by
proteomics. Mechanisms of metal toxicity can be studied by evaluating
changes in protein after interaction with heavy metal. Biomarker identification
can help in designing diagnostic tests for detecting protein toxicity. In vitro
toxicity assay can be performed on cell lines and in cultured cells. In vivo
assay can be performed in zebrafish, insect, or rat models. Heavy metals tend
to accumulate in the brain and liver so they are the most focused organs. 2-DE
is the most widely used technique as it can simultaneously resolve many
proteins. Heavy metals have shown differential expression of proteins related
to antioxidant defense mechanisms. Many proteomics studies have shown
that enzymes involved in glutathione (GSH) are differentially regulated in case
of heavy metal poisoning. Heavy metal poisoning affects the heat shock
proteins (HSP) which are generally involved in protein folding, aggregation and
stability, so it can be assumed that HSP has a role in cellular defense against
heavy metal-induced stress. The upregulation of proteins associated with
energy production may be related to the higher energy required for
detoxification. Utilizing proteomic techniques, particularly quantitative
proteomics, will enable the generation of more precise and reliable results,
which will undoubtedly advance this developing field and lead to the
identification of novel biomarkers and new insights into the mechanism of
metal toxicity.

17.5.4 Phage Antibody as Tool


Phage display technology was invented by George Smith in 1985. It is a
technique in which a gene encoding a protein of interest is inserted into a
phage coat protein gene, which leads to the phage expressing the protein as a
fusion product with one of the phage coat proteins. Phage display allows the
presentation of large protein libraries on the surface of filamentous phage,
which leads to the selection of proteins and antibodies with high affinity and
specificity to almost any target. This technology has been used in various
fields of immunology, cell biology, pharmacology and drug discovery. It is used
in protein-ligand interactions, protein-protein interaction producing monoclonal
antibodies and improving their affinity and epitope mapping.

Bacteriophages used in the phage display technique are single-stranded DNA


viruses that infect a number of bacteria. Most popular choice of phage in
phage display is M13. It is non-lytic bacteriophage and it doesn't destroy host
bacteria during infection, also it has proteins that can be displayed on the
surface. It consists of different proteins like pIII, pVI, pVII, pIX, pVIII. Major
protein, pVIII forms an envelope. About 2700 copies of pVIII are present in
filamentous M13.
Steps in phage display technique
Step 1: Production of phage display library: Phage display begins by
inserting a diverse set of genes into the phagemid vector or recombinant
phage. Mostly a segment of foreign DNA is inserted into gene III or gene VIII.
The modified gene III contains an added segment that expresses an antibody,
small protein or peptide on the surface of the phage. A collection of phages
displaying related but diverse proteins or peptide is called a library. 171


Block 4 Applications of Genomics and Proteomics
Step 2: Target Exposure: The library is then exposed to an immobilized
target such as a receptor, enzyme or ligand.

Step 3: Binding: When the library is exposed to a target, some members of


the library will bind to the target through specific interactions between the
displayed molecule and the target. According to the affinity of phages, some
have higher binding affinity and some have lower binding affinity towards the
target protein.

Step 4: Washing: After binding of phage to the target, washing is done to


remove phage that did not bind to the target and only those which show affinity
for the receptor remains bound to the target. Bound phage is eluted by
changing pH, addition of salt, detergent, etc.

Step 5: Amplification: Eluted phages with specificity and affinity for binding to
the target are then replicated in bacteria. Amplification produces a phage
mixture that is enriched with binding to a specific target. The repeated cycling
of these steps is called biopanning (Fig. 17.9).

Fig.17.9: Steps in phage display technique

Application of phage display in proteomics

Protein-protein interaction studies are very important in understanding cellular


function and dysfunction. The proteins for phage display are derived from
cDNA, open reading frame, genomic DNA, etc. Highly diverse library of
phages are created to identify ligands with high affinity. The use of the phage
display technique along with other techniques like yeast two hybrid system can
172 be very beneficial in understanding protein-protein interaction.


Unit 17 Proteome Analysis and Application of Proteomics
2D gel is used widely for the study of cellular proteins. However, it is labor-
intensive procedure and the variability is too high. So, affinity agents like
monoclonal antibodies are used to identify a given protein. Monoclonal
antibodies bind specially to a single epitope so they can be used to establish
the identity of a given protein. The Hybridoma technique was used earlier to
produce antibodies but it is difficult to produce a large number of monoclonal
antibodies needed for proteomics studies. The phage display technique
provides an alternative to produce a large number of monoclonal antibodies.
Cellular proteins are first separated by 2D Gel electrophoresis and blotted on
polyvinylidene fluoride membrane. Other sites on the PVDF membrane are
blocked so the only target available will be blotted antigen. The membrane is
incubated with a phage antibody library and washing is done to remove non-
bounded antibodies. Two to three rounds of selection are done to increase the
frequency of positive clones. Selected antibodies are studied by western
blotting.

The region of antigen to which the antibody binds is called epitope. Locating
the targeted antigen's binding sites where an antibody binds is known as
epitope mapping. The phage display library is used to display a number of
peptides. Antibodies have the capability to select peptides with high affinity for
their paratopes from these libraries. The phage display library is used to define
peptide structures and is recognized by major histocompatibility (MHC)
molecules. MHC molecules bind to peptide fragments derived from pathogens
and display them on the cell surface, which are then recognized by T cells.
Epitope mapping is significantly used in vaccine development and allows the
construction of peptide vaccines based on epitope specificity.

6$4
6$4
Fill in the blanks:

a) The term proteomics was first introduced by ……………….. in 1996.

b) To generate output for proteomics quality control evaluation, the raw


data is submitted to the ……………….. programme.

c) The protein's ……………….. structure can be used to directly synthesise


appropriate medications with the use of bioinformatics techniques that
will identify inhibitors that bind to the specific protein target.

d) The ……………….. profiling is an approach where the bioactive


molecule of interest is chemically conjugated to a suitable affinity moiety

e) Heavy metal poisoning affects the ……………….. which are generally


involved in protein folding, aggregation and stability.

f) The ……………….. is non-lytic bacteriophage and it doesn't destroy host


bacteria during infection, also it has proteins that can be displayed on
the surface.

g) The region of antigen to which the antibody binds is called …………… .

173


Block 4 Applications of Genomics and Proteomics

17.6 SUMMARY
• Proteomics is the large-scale study of proteomes. A proteome is a set of
proteins produced in an organism, system, or biological context. The
proteome is not constant; it differs from cell to cell and changes over
time.

• The entire proteome of a cell, tissue, organ, or organism is described as


part of the profiling proteomics, which also involves organelle mapping
and differential expression level measurement between cells or
conditions.

• By identifying protein interactions and the existence of posttranslational


modifications, functional proteomics seeks to characterise the activity
specific proteins.

• Characterising the interactions between small molecules and their


protein targets is the aim of the field of chemoproteomics.

• Phosphoproteomics involves the extensive analysis of protein


phosphorylation sites to identify the regulation and dysregulation of
signalling networks in healthy as well as diseased states.

• Proteins can be separated employing two different methods in two-


dimensional gel electrophoresis: isoelectric focusing (IEF) and sodium
dodecyl-sulfate polyacrylamide gel electrophoresis (SDS-PAGE).

• The 2D-DIGE technique is based on separation of multiple protein


samples on same 2D gel in which Fluorescent cyanine dyes (CyDye) are
used to label protein from different samples.

• The technology of isotope-coded affinity tag is employed in proteomics


of differential expression. Using isotope-coded affinity tag (ICAT),
proteins with cysteine residues can be labelled. This is a chemical
labelling agent-based in vitro approach for quantitative proteomics.

• Stable isotope labelling by amino acids (SILAC) in cell culture, is a mass


spectrometry-based technique that uses non-radioactive isotopic
labelling to detect variations in protein abundance between samples.

• Proteomics is employed in many pharmaceutical applications, including


safety and toxicity research, medication ADME (absorption, distribution,
metabolism, and excretion), and chemical profiling.

• The protein's crystal structure can be used to directly synthesise


appropriate medications with the use of bioinformatics techniques that
will identify inhibitors that bind to the specific protein target.

• Drug development is a complex process and involves various disciplines


like metabolomics, structural biology, bioinformatics and proteomics.

• A characteristic feature that is objectively tested and assessed as an


indicator of pathogenic processes, normal biological processes, or
pharmacological responses to a therapeutic intervention is called a
174 biomarker.


Unit 17 Proteome Analysis and Application of Proteomics
• Phage display technology is used in protein-ligand interactions, protein-
protein interaction producing monoclonal antibodies and improving their
affinity and epitope mapping.

• The use of the phage display technique along with other techniques like
yeast two hybrid system can be very beneficial in understanding protein-
protein interaction.

• Epitope mapping is significantly used in vaccine development and allows


the construction of peptide vaccines based on epitope specificity.

17.7 TERMINAL QUESTIONS


1. Describe Proteomics and explain its origin.

2. Describe the different classes of proteomics?

3. Write a short note on following techniques of Proteomics:

a) Two dimensional gel electrophoresis (2DE)

b) Two dimensional difference gel electrophoresis (2D-DIGE)

c) Isotope Coded Affinity Tag (ICAT)

d) Stable isotope labeling by amino acids in cell culture (SILAC)

4. Discuss the application of Proteomics in the field of pharmaceuticals,


drug discovery, toxicology and phage antibody as tool.

17. 8 ANSWERS
Self-Assessment Questions
1. a) i) False, ii) True, iii) True, iv) True

b) i) proteomics, ii) profiling proteomics, iii) functional


proteomics, iv) chemoproteomics, v) phosphoproteomics

2. a) isoelectric focusing (IEF); Sodium dodecyl-sulfate polyacrylamide


gel electrophoresis (SDS-PAGE)

b) isoelectric point

c) anionic

d) commassie brilliant blue; silver stain

e) Two-dimensional difference gel electrophoresis (2D-DIGE)

f) Isotope-coded affinity tag (ICAT)

g) stable isotope labeling by amino acids in cell culture (SILAC)

3. a) Marc Wilkins, b) MaxQuant, c) crystal, d) drug-centric


chemoproteomic, e) heat shock proteins (HSP), g) M13,
h) epitope 175


Block 4 Applications of Genomics and Proteomics
Terminal Questions
1. Refer to Sections 17.1 and 17.2.

2. Refer to Section 17.3.

3. i) Refer to Subsection 17.4.1.

ii) Refer to Subsection 17.4.2.

iii) Refer to Subsection 17.4.3.

iv) Refer to Subsection 17.4.4.

4. Refer to Section 17.5.

176


GLOSSARY
Acrocentric : A chromosome where the centromere is not
chromosomes central and is instead located near the end of the
chromosome.
Albinism : It is derived from the Latin albus, meaning
"white," is a group of heritable conditions
associated with decreased or absence of melanin
in ectoderm-derived tissues (most notably the
skin, hair and eyes), yielding a characteristic
pallor.

Argonaute proteins : An evolutionarily highly conserved family of


proteins associated with the silencing of gene
expression in pathways such as RNA
interference (RNAi) by interaction with small,
single-stranded, noncoding RNA, leading to the
formation of RNA-Induced Silencing Complex
(RISC).

Biomarker : A biomarker (short for biological marker) is an


objective measure that captures what is
happening in a cell or an organism at a given
moment.
Box plot : It visually show the distribution of numerical data
and skewness by displaying the data quartiles (or
percentiles) and averages. Box plots show the
five-number summary of a set of data: including
the minimum score, first (lower) quartile, median,
third (upper) quartile and maximum score.
Centromere : The region of the chromosome to which the
spindle fiber is attached during cell division (both
mitosis and meiosis). The centromere is the
constricted point at which the two chromatids
forming the chromosome are joined together.

ChIP-on-chip : Chromatin immunoprecipitation (ChIP) followed


Microarray by microarray hybridization (ChIP-chip) or high-
throughput sequencing (ChIP-seq) allows
genome-wide discovery of protein-DNA
interactions such as transcription factor bindings
and histone modifications.

Comparative Genomic : This technique permits the detection of


Hybridization chromosomal copy number changes without the
need for cell culturing. It provides a global
overview of chromosomal gains and losses
throughout the whole genome of a tumor.

Dicer : It is a general name for a family of enzymes that


generate short pieces of RNA, about 21–23
nucleotides in length.
Volume 2 Proteome Analysis and Applications
DNAse Footprinting : Fragments of a 5' end-labelled, double-stranded
Assay DNA segment, partially degraded by DNAase in
the presence and absence of the binding protein,
are visualized by electrophoresis and
autoradiography, utilizing the base-specific
reaction products of the Maxam-Gilbert
sequencing method. It is possible to see the
protective "footprint" of the binding protein on the
DNA sequence.

DROSHA : It is a nuclear RNase III enzyme, responsible for


cleaving primary microRNAs (miRNAs) into
precursor miRNAs and is thus essential for the
biogenesis of canonical miRNAs.

Edman Degradation : Method of sequencing amino acids in a peptide


by labeling the amino terminal residue and
cleaving from the peptide without disrupting the
peptide bonds between other amino acid
residues.

Electrophoretic : Technique used to detect protein complexes with


Mobility Shift Assay nucleic acids. Solutions of protein and nucleic
acid are combined and the resulting mixtures are
subjected to electrophoresis under native
conditions through polyacrylamide gel. After
electrophoresis, the distribution of species
containing nucleic acid is determined, usually by
autoradiography of 32P-labeled nucleic acid.

Electrospray : Technique used in mass spectrometry to produce


Ionization ions using an electrospray in which a high voltage
is applied to a liquid to create an aerosol.

Epidemiology : It is the study of the determinants, occurrence


and distribution of health and disease in a defined
population.

Epitope mapping : It is the process of identifying and characterizing


the specific regions on an antigen to which
antibodies bind.

F-box protein : The F-box is a protein motif of approximately 50


amino acids that functions as a site of protein-
protein interaction.

Gene knockin : A gene knockin refers to a genetic engineering


method that involves the insertion of a protein,
coding a cDNA sequence at a particular locus, in
an organism's chromosome.

Heat map : A heat map is a two-dimensional representation


of data in which various values are represented
178 by colors. An instant visual summary of data




Volume 2 Proteome Analysis and Applications
across two axes is provided by a basic heat map,
which enables users to rapidly identify the most
significant or pertinent data points. Complex data
sets can be understood by the viewer with more
intricate heat maps.

Hepatotoxicity : It is the medical term for damage to the liver


caused by a medicine, chemical and herbal or
dietary supplement.

Ideograms : A picture or symbol used in a system of writing to


represent an idea, but not a particular word or
phrase for it.

Isoelectric Focusing : A technique that separates charged molecules,


usually proteins or peptides, on the basis of their
isoelectric point (pI) which is the pH at which a
molecule has no overall charge.

Isoelectric point : The isoelectric point (pI) is the pH of a solution at


which the net charge of a protein becomes zero.

Knockout : A knockout, as related to genomics, refers to the


use of genetic engineering to inactivate or
remove one or more specific genes from an
organism. Scientists create knockout organisms
to study the impact of removing a gene from an
organism, which often allows them to then learn
about that gene's function.

Mass Spectrometry : This technique is based on the ionization of


sample molecules in the gas phase, followed by
separation and detection of the resulting ions
according to mass-to-charge ratio (m/z). The
ionized sample molecules can be efficiently
fragmented to yield product ions. Results are
displayed in the form of a mass spectrum, a
graphical representation of ion abundance
versus m/z.

Matrix Assisted Laser : Technique for soft ionization of mass


Desorption Ionization spectrometry that uses a laser energy absorbing
matrix to create ions from large molecules with
minimal fragmentation.

Mutagen : A mutagen is a chemical or physical agent


capable of inducing changes in DNA called
mutations.

P-bodies : Processing bodies (P-bodies) are cytoplasmic


ribonucleoprotein (RNP) granules comprised
primarily of mRNAs in complex with proteins,
associated with translation repression and 5ƍ-to-3ƍ
mRNA decay. 179



Volume 2 Proteome Analysis and Applications
PepNovo : It is a high throughput de novo peptide
sequencing tool for tandem mass spectrometry
data.

Peptide Mass : In this technique, after separation of proteins by


Fingerprinting gel electrophoresis or liquid chromatography and
cleaving with a proteolytic enzyme, experimental
peptide mass is derived through mass
spectrometry.

Phage Display : Technology utilized in studying protein-ligand


interactions, receptor binding sites and in
improving or modifying the affinity of proteins for
their binding partners. Generating monoclonal
antibodies and improving their affinity, cloning
antibodies from unstable hybridoma cells and
identifying epitopes, mimotopes and functional or
accessible sites from antigens are important
advantages of this technology.

Phage display : It is a molecular biology technique by which


phage genomes are modified in such a way that
the coat proteins of assembled virions are fused
to other proteins or peptides of interest (of any
origin), displaying them to the external milieu.

Polylinker : A short DNA sequence containing several re-


striction enzyme recognition sites that are
contained in cloning vectors.

Protein Microarray : An emerging technique that provides a versatile


platform for the characterization of hundreds of
thousands of proteins in a highly parallel and
highǦthroughput manner. Protein microarrays are
composed of two major classes, viz., analytical
and functional.

Protein turnover : It is the net result of continuous synthesis and


breakdown of body proteins and ensures
maintenance of optimally functioning proteins.

Proteolysis : It is a hydrolysis reaction of peptide bonds in


which proteins breakdown into smaller peptides
and/or into individual amino acid residues.

Pyrosequencing : It is a replication-based sequencing method in


which addition of the correct nucleotide
to immobilized template DNA is signaled by a
photometrically detectable reaction.

Quelling : It is related to co-suppression, observed in plants,


and RNA interference in animals; it requires an
Argonaute protein and acts by generating small
RNA molecules (about 25 nt long), which in turn
180 target mRNAs to be silenced.




Volume 2 Proteome Analysis and Applications
Recombinant DNA : Recombinant DNA technology involves using
technology enzymes and various laboratory techniques to
manipulate and isolate DNA segments of interest.
This method can be used to combine (or splice)
DNA from different species or to create genes
with new functions. The resulting copies are often
referred to as recombinant DNA.
Reporter gene : A reporter gene is a nonendogenous gene
encoding an enzyme or fluorescent protein
whose expression is controlled by a promoter for
a separate gene of interest. It allows for
identifying, quantifying, visualizing and tracking
gene expression and protein distribution in cells.
RITS complex : RNA-induced transcriptional silencing (RITS)
complex, consisting of Ago1, Tas3 and
Chp1, binds to nascent transcripts from
centromeric chromatin (cenRNA). RITS leads to
transcriptional silencing by placing chromatin
marks and recruiting the RNA-dependent RNA
polymerase complex (RDRC).
RNA-induced : One strand of a small interfering RNA (siRNA) or
silencing complex micro RNA (miRNA) is incorporated into the
multiprotein complex known as the RNA-induced
silencing complex (RISC). The siRNA or miRNA
serves as a template for complementary mRNA
recognition in RISC. It initiates RNase activity and
cleaves the RNA when it comes across a
complementary strand. This process is crucial for
both defence against viral infections, which
frequently use double-stranded RNA as an
infectious vector, and for the control of genes by
microRNAs.
Scott syndrome : It is a rare autosomal recessive congenital
bleeding disorder caused by a defect in
blood coagulation.
SHERENGA : It is an algorithm for de novo interpretation of
MS/MS spectra.
Sodium Dodecyl : Technique used for the separation of proteins,
Sulphate- based on their molecular weight.
Polyacrylamide Gel
Electrophoresis
Southwestern Blotting : Technique used to study DNA-protein
interactions. This method detects specific DNA-
binding proteins by incubating radiolabeled DNA
with a gel blot, washing and visualizing through
autoradiography.
Telomeres : Telomeres are structures made from DNA
sequences and proteins found at the ends of
chromosomes. 181



Volume 2 Proteome Analysis and Applications
The International : The aim of this project is to determine the
HAPMAP Project common patterns of DNA sequence variation in
the human genome and to make this information
freely available in the public domain. There is an
international consortium involved in developing a
map of these patterns across the genome. This is
possible through determining the genotypes of
sequence variants, their frequencies and the
degree of association between them, in DNA
samples from populations with ancestry from
parts of Africa, Asia and Europe. This will lead to
the discovery of sequence variants that affect
common disease, thereby facilitating the
development of diagnostic tools.
Time of Flight Mass : Technology that utilizes an electric field to
Analyzer accelerate generated ions through the same
electrical potential, and then measures the time
each ion takes to reach the detector.
Transcription factors : These are proteins involved in the process of
converting or transcribing DNA into RNA.
Transfection : It refers to the introduction of foreign DNA
(genetic material other than host genomes) into
the cell. The main purpose of transfection is to
alter the host genome to express or block the
expression of the protein, associated with the
gene.
Transformation : It is a process by which foreign genetic material is
taken up by a cell. The process results in a stable
genetic change within the transformed cell.
Two-Dimensional Gel : This technique separates proteins, depending on
Electrophoresis two different steps: the first one is called
isoelectric focusing which separates proteins
according to isoelectric points (pI); the second
step is SDS-PAGE which separates proteins,
based on the molecular weights.
Western Blotting : Procedure for the immunodetection of proteins,
particularly proteins that are of low abundance.
This process involves the transfer of protein,
patterns from gel to microporous membrane.
Yeast One-Hybrid : Important technique for detecting physical
Assay interactions between sequence-specific
regulatory transcription factor proteins and their
DNA target sites. It involves two components: (1)
a reporter construct with DNA of interest cloned
upstream of a gene encoding a reporter protein
that can be easily detected; and (2) an
expression construct that generates a fusion (or
“hybrid”) between a transcription factor of interest
and a yeast transcription activation domain.
182

You might also like