MODULE 1
BIOINFORMATICS
● Research, development, or application of computational tools and approaches for
expanding the use of biological, medical, behavioral or health data, including those to
acquire, store, organize, archive, analyze, or visualize such data.
● Bioinformatics is an interdisciplinary research area at the interface between computer
science and biological science.
● Bioinformatics involves the technology that uses computers for storage, retrieval,
manipulation, and distribution of information related to biological macromolecules such
as DNA, RNA, and proteins
● The field of science in which biology, computer science and information technology
merge into a single discipline
Computational Biology
The development and application of data-analytical and theoretical methods, mathematical
modeling and computational simulation techniques to the study of biological, behavioral, and
social systems.
GOALS
1. The ultimate goal of bioinformatics is to better understand a living cell and how it
functions at the molecular level.
2. Cellular functions are mainly performed by proteins whose capabilities are ultimately
determined by their sequences. Therefore, solving functional problems using sequence
and sometimes structural approaches has proved to be a fruitful endeavor
SCOPE
● Bioinformatics consists of two subfields: the development of computational tools and
databases and the application of these tools and databases in generating biological
knowledge to better understand living systems.
● These tools are used in three areas of genomic and molecular biological research:
○ Molecular sequence analysis
○ Molecular structural analysis
○
TRACE KTU
Molecular functional analysis
APPLICATIONS
1. Knowledge-based drug design.
2. Forensic DNA analysis
3. Agricultural biotechnology
Central Dogma Of Molecular Biology
The central dogma of molecular biology is an explanation of the flow of genetic information within a
biological system. It is often stated as "DNA makes RNA and RNA makes protein,"[1]although this is
not its original meaning. It was first stated by Francis Crick in 1958:[2]
The Central Dogma. This states that once 'information' has passed into protein it cannot
“
get out again. In more detail, the transfer of information from nucleic acid to nucleic
acid, or from nucleic acid to protein may be possible, but transfer from protein to
protein, or from protein to nucleic acid is impossible. Information means here the
precise determination of sequence, either of bases in the nucleic acid or of amino acid
”
residues in the protein.
“ TRACE KTU
The central dogma of molecular biology deals with the detailed residue-by-residue transfer
of sequential information. It states that such information cannot be transferred back from
protein to either protein or nucleic acid.
Information flow in biological systems
Biological sequence information
The biopolymers that comprise DNA, RNA and (poly)peptides are linear polymers (i.e.: each
monomer is connected to at most two other monomers). The sequence of their monomers
effectively encodes information. The transfers of information described by the central dogma
ideally are faithful, deterministic transfers, wherein one biopolymers sequence is used as a
template for the construction of another biopolymer with a sequence that is entirely dependent
on the original biopolymers sequence.
TRACE KTU
DNA contains the complete genetic information that defines the structure and function of an
organism. Proteins are formed using the genetic code of the DNA. Three different processes are
responsible for the inheritance of genetic information and for its conversion from one form to
another:
1. Replication : a double stranded nucleic acid is duplicated to give identical copies. This
process perpetuates the genetic information.
2. Transcription : a DNA segment that constitutes a gene is read and transcribed into a
Vierstraete Andy (version 1.01) 1/02/2000 -Page 2 - single stranded sequence of RNA. The RNA
moves from the nucleus into the cytoplasm.
3. Translation : the RNA sequence is translated into a sequence of amino acids as the protein is
formed. During translation, the ribosome reads three bases (a codon) at a time from the RNA and
translates them into one amino acid In eukaryotic cells, the second step (transcription) is
necessary because the genetic material in the nucleus is physically separated from the site of
protein synthesis in the cytoplasm in the cell. Therefore, it is not possible to translate DNA
directly into protein, but an intermediary must be made to carry the information from one
compartment to another
Major Types of RNA
TRACE KTU
There are three main types of RNA – messenger RNA or mRNA, ribosomal or rRNA, and
transfer RNA or tRNA. These 3 types of RNA are discussed below.
Messenger RNA (mRNA)
mRNA accounts for just 5% of the total RNA in the cell. mRNA is the most heterogeneous of
the 3 types of RNA in terms of both base sequence and size. It carries the genetic code copied
from the DNA during transcription in the form of triplets of nucleotides called codons to the
ribosome.
Ribosomal RNA (rRNA)
rRNAs are found in the ribosomes and account for 80% of the total RNA present in the cell.
Different rRNAs present in the ribosomes include small rRNAs and large rRNAs, which
denote their presence in the small and large subunits of the ribosome.
rRNAs combine with proteins in the cytoplasm to form ribosomes, which act as the site of
protein synthesis and has the enzymes needed for the process. These complex structures
travel along the mRNA molecule during translation and facilitate the assembly of amino
acids to form a polypeptide chain. They bind to tRNAs and other molecules that are crucial
for protein synthesis.
In bacteria, the small and large rRNAs contain about 1500 and 3000 nucleotides,
respectively, whereas in humans, they have about 1800 and 5000 nucleotides, respectively.
However, the structure and function of ribosomes is largely similar across all species.
Transfer RNA (tRNA)
tRNA is the smallest of the 3 types of RNA having about 75-95 nucleotides. tRNAs are an
essential component of translation, where their main function is the transfer of amino acids
during protein synthesis. Therefore they are called transfer RNAs.
Each of the 20 amino acids has a specific tRNA that binds with it and transfers it to the
growing polypeptide chain. tRNAs also act as adapters in the translation of the genetic
TRACE KTU
sequence of mRNA into proteins. Therefore they are also called adapter molecules.
tRNAs have a clover leaf structure which is stabilized by strong hydrogen bonds between the
nucleotides.
Coding RNA and Non Coding RNA
RNA can be readily classified into either protein-coding or non-protein–coding categories
Coding RNA
● Coding RNA is a RNA molecule that can be translated into a protein.
● Messenger RNA (mRNA) is a large family of RNA molecules that convey genetic
information from DNA to the ribosome, where they specify the amino acid sequence of
the protein products of gene expression.
Non Coding RNA
● A non-coding RNA (ncRNA) is an RNA molecule that is not translated into a protein.
● The DNA sequence from which a functional non-coding RNA is transcribed is often called an
RNA gene.
● Abundant and functionally important types of non-coding RNAs include transfer RNAs
(tRNAs) and ribosomal RNAs (rRNAs), as well as small RNAs such as microRNAs, siRNAs,
piRNAs, snoRNAs, snRNAs, exRNAs, scaRNAs and the long ncRNAs such as Xist and
HOTAIR.
● The number of non-coding RNAs within the human genome is unknown; however, recent
transcriptomic and bioinformatic studies suggest that there are thousands of them. Many of
the newly identified ncRNAs have not been validated for their function.
● Non-coding RNAs contribute to diseases including cancer and Alzheimer's.
miRNA
microRNA (abbreviated miRNA) is a small non-coding RNA molecule (containing about 22
nucleotides) found in plants, animals and some viruses, that functions in RNA silencing and
post-transcriptional regulation of gene expression
Encoded by nuclear DNA in plants and animals and by viral DNA in certain viruses whose
genome is based on DNA, miRNAs function via base-pairing with complementary sequences
TRACE KTU
within mRNA molecules
miRNAs are abundant in many mammalian cell types[7][8] and appear to target about 60% of the
genes of humans and other mammals.
RNAi
The term RNA interference ( RNAi) was coined to describe a cellular mechanism that use the
gene's own DNA sequence of gene to turn it off, a process that researchers call silencing. In
a wide variety of organisms, including animals, plants, and fungi, RNAi is triggered by
double-stranded RNA (dsRNA).
Two types of small ribonucleic acid (RNA) molecules – microRNA (miRNA) and small interfering
RNA (siRNA) – are central to RNA interference.
Nucleic Acid Structure - DNA and RNA Structure
Nucleic acids are molecules that allow organisms to transfer genetic information from one
generation to the next. These macromolecules store the genetic information that determines
TRACE KTU
traits and makes protein synthesis possible. Two examples of nucleic acids include:
deoxyribonucleic acid (better known as DNA) and ribonucleic acid (better known as
RNA). These molecules are composed of long strands of nucleotides held together by
covalent bonds. Nucleic acids can be found within the nucleus and cytoplasm of our cells.
The Building Blocks
Three types of chemicals make up the building blocks for nucleic acids: an aromatic base,
a sugar ring and a phosphate group. The sugar and the phosphate constitute the
non-specific backbone of the DNA. The base constitutes the specific part which actually holds
the information and accounts for most of the interactions with other molecules.
DNA nucleotides are made up of a sugar (deoxyribose), base (adenine, guanine,
cytosine, thymine) and phosphate.
RNA nucleotides are made up of a sugar (ribose), base (adenine, guanine,
cytosine, uracil) and phosphate.
Base
The group that gives each nucleic acid unit its specificity is the organic base. DNA contains
two purine bases (adenine and guanine) and two pyrimidine bases (cytosine and thymine).
In RNA the thymine base is replaced by uracil. Polar atoms in the ring or attached to the ring
TRACE KTU
are capable of creating hydrogen bonds with polar atoms of other bases.
Sugar
The sugars in DNA and RNA are pentoses (5 carbons). The ring of the sugar is flexible and
non-planar (in contrast to the rings of the base), and can adopt several conformations.
TRACE KTU
Phosphate
The inorganic acid H3PO4 (phosphoric acid) gives the nucleic acids an overall net negative
charge. Most of the interactions between the DNA and proteins which are not specific to the
DNA sequence are with the phosphate groups.
TRACE KTU
DNA Structure
DNA is the cellular molecule that contains instructions for the performance of
all cell functions. When a cell divides, its DNA is copied and passed from one
cell generation to the next generation. DNA is organized into chromosomes
and found within the nucleus of our cells. It contains the "programmatic
instructions" for cellular activities. When organisms produce offspring, these
instructions in are passed down through DNA.
DNA commonly exists as a double stranded molecule with a twisted double
helixshape. DNA is composed of a phosphate-deoxyribose sugar backbone and
the four nitrogenous bases: adenine (A), guanine (G), cytosine (C), and
thymine (T). In double stranded DNA, adenine pairs with thymine (A-T)
and guanine pairs with cytosine (G-C).
TRACE KTU
RNA STRUCTURE
RNA is essential for the synthesis of proteins. Information contained within
the genetic code is typically passed from DNA to RNA to the resulting proteins.
There are several different types of RNA.
Messenger RNA (mRNA) is the RNA transcript or RNA copy of the
DNA message produced during DNA transcription. Messenger RNA is
translated to form proteins.
Transfer RNA (tRNA) has a three dimensional shape and is
necessary for the translation of mRNA in protein synthesis.
Ribosomal RNA (rRNA) is a component of ribosomes and is also
involved in protein synthesis.
MicroRNAs (miRNAs) are small RNAs that help to regulate gene
expression.
TRACE KTU
RNA most commonly exists as a single stranded molecule composed of a
phosphate-ribose sugar backbone and the nitrogenous bases adenine,
guanine, cytosine and uracil (U). When DNA is transcribed into an RNA
transcript during DNA transcription, guanine pairs with cytosine (G-C) and
adenine pairs with uracil (A-U).
DNA versus RNA
The nucleic acids DNA and RNA differ in composition and structure. The
differences are listed as follows:
DNA
● Nitrogenous Bases: Adenine, Guanine, Cytosine, and Thymine
● Five-Carbon Sugar: Deoxyribose
● Structure: Double-stranded
DNA is commonly found in its three dimensional, double helix shape. This
twisted structure makes it possible for DNA to unwind for DNA replication
and protein synthesis.
RNA
● Nitrogenous Bases: Adenine, Guanine, Cytosine, and Uracil
● Five-Carbon Sugar: Ribose
● Structure: Single-stranded
TRACE KTU
While RNA does not take on a double helix shape like DNA, this molecule is
able to form complex three dimensional shapes. This is possible because RNA
bases form complementary pairs with other bases on the same RNA strand.
The base pairing causes RNA to fold forming various shapes.
TRACE KTU
Functions of Nucleic Acids
● The main functions is store and transfer genetic information.
● To use the genetic information to direct the synthesis of new protein.
● The deoxyribonucleic acid is the storage for place for genetic information in the cell.
● DNA controls the synthesis of RNA in the cell.
● The genetic information is transmitted from DNA to the protein synthesizers in the
cell.
● RNA also directs the production of new protein by transmitting genetic information
to the protein building structures.
● The function of the nitrogenous base sequences in the DNA backbone determines the
proteins being synthesized.
● The function of the double helix of the DNA is that no disorders occur in the genetic
information if it is lost or damaged.
● RNA directs synthesis of proteins.
● m-RNA takes genetic message from RNA.
● t-RNA transfers activated amino acid, to the site of protein synthesis.
● r-RNA are mostly present in the ribosomes, and responsible for stability of m-RNA.
TRACE KTU
Genetic code and its characteristics
The pathway of protein synthesis is called Translation because the language of nucleotide
sequence on mRNA is translated in to the language of an amino acid sequence. The process
of Translation requires a Genetic code, through which the information contained in nucleic
acid sequence is expressed to produce a specific sequence of amino acids.
● The letters A, G, T and C correspond to the nucleotides found in DNA. They are
organized into codons.
● The collection of codons is called Genetic code.
● For 20 amino acids there should be 20 codons.
● Each codon should have 3 nucleotides to impart specificity to each of the amino acid
for a specific codon
● 1 Nucleotide- 4 combinations
● 2 Nucleotides- 16 combinations
● 3 Nucleotides- 64 combinations ( Most suited for 20 amino acids).
Genetic code
● Genetic code is a dictionary that corresponds with sequence of nucleotides and
sequence of Amino Acids.
● Words in dictionary are in the form of codons
● Each codon is a triplet of nucleotides
● 64 codons in total and three out of these are Non Sense codons (Figure)
● 61 codons for 20 amino acids.
TRACE KTU
Figure- Genetic code is a dictionary that corresponds with sequence of nucleotides and
sequence of Amino Acids.
Genetic Code-Characteristics (Table)
1. Specificity
a. Genetic code is specific (Unambiguous)
b. A specific codon always codes for the same amino acid.
c. e.g. UUU codes for Phenyl Alanine, it cannot code for any other amino acid.
2. Universal
a. In all living organism Genetic code is the same.
b. The exception to universality is found in mitochondria! Codons -
c. where AUA codes for Methionine and UGA for tryptophan, instead of
termination codon respectively of cytoplasmic protein synthesizing
machinery.
d. AGA and AGG code for Arginine in cytoplasm but in mitochondria they are
termination codons.
3. Redundant
a. Genetic code is Redundant, also called Degenerate.
b. Although each codon corresponds to a single amino acid but a single amino
acid can have multiple codons.
TRACE KTU
c. Except Tryptophan and Methionine each amino acid has multiple codons.
4. Non Overlapping and Non Punctuated
a. All codons are independent sets of 3 bases.
b. There is no overlapping , Codon is read from a fixed starting point as a
continuous sequence of bases, taken three at a time.
c. The starting point is extremely important and this is called Reading frame.
5. Non Sense Codons
a. There are 3 codons out of 64 in genetic code which do not encode for any
Amino Acid.
b. These are called termination codons or stop codons or nonsense codons.
c. The stop codons are UAA, UAG, and UGA. They encode no amino acid.
d. The ribosome pauses and falls off the mRNA.
6. Initiator codon
a. AUG is the initiator codon in majority of proteins
b. In a few cases GUG may be the initiator codon
c. Methionine is the only amino acid specified by just one codon, AUG.
Table- Characteristics of Genetic code (Summary)
Sl. No Feature Details
1 Specific/ Unambiguous Given a specific codon, only a single amino
acid is indicated.
2 Universal In all living organism Genetic code is the
same (Except mitochondria) codons)
3 Redundant/ Degenerate Multiple codons can decode the same
amino acid
4
TRACE KTU
Non Overlapping The reading of the genetic code during the
process of protein synthesis does not
involve any overlap of codons
5 Non Punctuated Once the reading is commenced at a
specific codon, there is no punctuation
between codons, and the message is read in
a continuing sequence of nucleotide
triplets until a translation stop codon is
reached.