0% found this document useful (0 votes)
3 views33 pages

Introduction to Bioinformatics Concepts

Introduction to Bioinformatics

Uploaded by

gorabbi07
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views33 pages

Introduction to Bioinformatics Concepts

Introduction to Bioinformatics

Uploaded by

gorabbi07
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Introduction to Bioinformatics

CSE 4463
Lecture – 1

Wahida Ferdose Urmi


Lecturer, CSE
Introduction

• The term was introduced in 1970s by


Paulien Hogeweg and Ben Hesper

• Study biotic systems

• Information processing in various


forms
Introduction
• Bioinformatics is an interdisciplinary field that develop
methods and software tools for understanding biological
data.

• Bioinformatics is about searching, managing and analyzing


large amount of biological data using different
computational approaches.

• Union of biology and informatics: bioinformatics involves


the technology that uses computers for storage, retrieval,
manipulation, and distribution of information related to
biological macromolecules such as DNA, RNA, & proteins.

• As an interdisciplinary field of science, bioinformatics


combines computer science, statistics, and mathematics to
analyze and interpret biological data.
Outlines
• Cell
• Chromosomes
• DNA replication, transcription, translation
• Genes
• Human Genome
• Gene Expression Datasets
• Gene Regulatory Network
• Application Areas of Bioinformatics
Cell

• Basic unit of life


• Different types of cell:
• Skin, brain, red/white blood
• Different biological function
• Cells produced by cells
• Cell division (mitosis)
• 2 daughter cells
Types of Cells
Chromosomes
● Each cell has nucleus
● Rod-shaped particles inside
– Are chromosomes
– Which we think of in pairs
● Different number for species
– Human(46),tobacco(48)
– Goldfish(94),chimp(48)
– Usually paired up
● X & Y Chromosomes
– Humans: Male(xy), Female(xx)
– Birds: Male(xx), Female(xy)
DNA
All Life depends on 3 critical molecules
• DNAs
– Hold information on how cell works

• RNAs
– Act to transfer short pieces of information to different parts of cell

– Provide templates to synthesize into protein

• Proteins
– Form enzymes that send signals to other cells and regulate gene activity

– Form body’s major components (e.g. hair, skin, etc.)

– Are life’s laborers!


DeoxyriboNucleic Acid (DNA) : Structure
and Role
• Double Helix Structure (Watson and Crick, Nature 1953)
• Carrier of genetic instructions
• Two complementary antiparallel strands, one runs from 5’ to 3’
end and another runs from 3’ to 5’ end
• 3 major parts – Nitrogenous Base, 5-Carbon Deoxyribose Sugar
and Phosphate Group
• Four nitrogenous bases – Adenine (A), Cytosine (C), Guanine
(G), Thymine (T)
• A-T is Double Hydrogen Bond and G-C is Triple Hydrogen
Bond
• DNA is more stable than RNA due to its Deoxyribose Sugar
Structure
RNA
• Unlike DNA, RNA is single-stranded and much
shorter in length.
• It carries the genetic instructions for a single gene
rather than an entire genome.
• Another key difference is the sugar in RNA—ribose,
rather than the deoxyribose found in DNA.
• Additionally, RNA uses uracil (U) instead of
thymine (T). These differences are crucial for RNA’s
role in translating genetic information into action."
The Central Dogma
DNA Replication
The parent molecule unwinds, and two new daughter strands are built based on
base-pairing rules
DNA Replication is Semi-Conservative, because, in new sets of DNA,
one strand is newly created but the other strand comes from the
ancestor.
Principle of DNA Replication
• The Basic Principle: Base Pairing to a Template
Strand The relationship between structure and
function is manifest in the double helix

• Since the two strands of DNA are complementary


each strand acts as a template for building a new
strand in replication
DNA Replication
▹Initiation
- Helicase enzyme unwinds DNA strands
- Replication fork is created
- RNA Primer is created by Primase enzyme
- Primer is starting point of elongation
▹Elongation
- New DNA Strand grows 1 base at a time as complimentary of leading strand (5’ to 3’)
- DNA Polymerase enzyme controls it
- Complimentary strand of lagging strand is created in small fragments called Okazaki
Fragments (3’ to 5’)
▹Termination
- Exonuclease enzyme removes all the primer sequences from new strands
- Again, DNA Polymerase fills the gaps
- DNA Ligase enzyme seals all the gaps
DNA Transcription
DNA transcription is the process where DNA is converted into mRNA
(messenger RNA) to be used for protein synthesis.
RNA Splicing
DNA translation
DNA translation is the process where the genetic code in mRNA is
used to synthesize proteins.
▪ mRNA (Messenger RNA) – Carries genetic instructions from DNA.
▪ tRNA (Transfer RNA) – Brings amino acids to the ribosome.
▪ Ribosome – Assembles proteins by linking amino acids.
▪ Codons – Three-nucleotide sequences in mRNA that specify amino acids.
▪ Start & Stop Codons – Initiate and terminate protein synthesis.
Codons
Proteins
• proteins are molecules composed of one or more
polypeptides

• a polypeptide is a polymer composed of amino acids

• cells build their proteins from 20 different amino acids

• a polypeptide can be thought of as a string composed from


a 20-character alphabet
Protein Functions
• structural support
• storage of amino acids
• transport of other substances
• coordination of an organism’s activities
• response of cell to chemical stimuli
• movement
• protection against disease
• selective acceleration of chemical reactions
Amino Acids
Alanine Ala A
Arginine Arg R
Aspartic Acid Asp D
Asparagine Asn N
Cysteine Cys C
Glutamic Acid Glu E
Glutamine Gln Q
Glycine Gly G
Histidine His H
Isoleucine Ile I
Leucine Leu L
Lysine Lys K
Methionine Met M
Phenylalanine Phe F
Proline Pro P
Serine Ser S
Threonine Thr T
Tryptophan Trp W
Tyrosine Tyr Y
Valine Val V
Amino Acid Sequence: Hexokinase
5 10 15 20 25 30
1 A A S X D X S L V E V H X X V F I V P P X I L Q A V V S I A
31 T T R X D D X D S A A A S I P M V P G W V L K Q V X G S Q A
61 G S F L A I V M G G G D L E V I L I X L A G Y Q E S S I X A
91 S R S L A A S M X T T A I P S D L W G N X A X S N A A F S S
121 X E F S S X A G S V P L G F T F X E A G A K E X V I K G Q I
151 T X Q A X A F S L A X L X K L I S A M X N A X F P A G D X X
181 X X V A D I X D S H G I L X X V N Y T D A X I K M G I I F G
211 S G V N A A Y W C D S T X I A D A A D A G X X G G A G X M X
241 V C C X Q D S F R K A F P S L P Q I X Y X X T L N X X S P X
271 A X K T F E K N S X A K N X G Q S L R D V L M X Y K X X G Q
301 X H X X X A X D F X A A N V E N S S Y P A K I Q K L P H F D
331 L R X X X D L F X G D Q G I A X K T X M K X V V R R X L F L
361 I A A Y A F R L V V C X I X A I C Q K K G Y S S G H I A A X
391 G S X R D Y S G F S X N S A T X N X N I Y G W P Q S A X X S
421 K P I X I T P A I D G E G A A X X V I X S I A S S Q X X X A
451 X X S A X X A

• enzyme involved in glycolysis


• in every organism known from bacteria to humans
Genes
• a gene is a sequence of bases that carries the information required for
constructing a particular protein (more accurately, polypeptide)
• such a gene is said to encode a protein
• the human genome comprises ~ 25,000 protein-coding genes
• Often the active part of a gene is spit into exons - Separated by introns

• not all of the DNA in a genome encodes protein:

• bacteria ~90% coding gene/kb


• human ~1.5% coding gene/35kb
Genes
❑ Bioinformatics helps identify genes within a
long DNA sequence.
❑ This technique locates a gene simply by
analyzing sequence data using a computer

Why is gene prediction important?


• Helps scientists to distinguish between coding
and non-coding regions of a genome,
• Explain genes in terms of their function,
• Conduct research related to detection,
treatment, and prevention of genetic disorder
diseases, etc.
Human Genome Project
• The Human Genome Project (HGP) was
first proposed by the U.S. National
Research Council in 1990.
• scientists wanted to understand all the genetic
information in humans.

• Goals:
Create maps of the human genome: This means
figuring out where genes are located and how
they work.

Study other organisms' genomes: While


mapping the human genome, scientists also
studied the genomes of other creatures to help
understand ours better.
Gene Expression Data
• Gene Expression Data is the biological data to
extract meaningful hidden information from the
gene dataset.
• This gene information is used for disease
diagnosis especially in cancer treatment based on
the variations in gene expression levels.

Data generation methods:


• Gene expression levels are typically measured
using techniques like microarray or RNA
sequencing (RNA-Seq), which provide a
snapshot of the RNA transcripts produced by a
gene, indicating its activity level.
Examples of publicly available gene expression
datasets:
Gene Expression Omnibus (GEO):
• A large repository maintained by NCBI (National Center for Biotechnology
Information ) containing gene expression data from various studies across diverse
organisms.
The Cancer Genome Atlas (TCGA):
• A comprehensive collection of genomic data including gene expression profiles
for various cancer types.
ArrayExpress:

• A European repository for microarray data including gene expression datasets.


Gene Regulatory Network
• A set of genes, proteins, small molecules which interact mutually to control rate of
transcription
• In unicellular organisms regulatory networks respond to the external environment,
to make the cell survival (Yeast)
• In multicellular organisms regulatory networks control transcription, cell signaling
and development
Function:
• GRNs play a critical role in regulating cellular functions by controlling which
genes are expressed at a given time, depending on environmental conditions or
developmental stages.
Gene Regulatory Networks (GRNs): Components
and Function
• GRNs are made up of thousands of DNA
sequences in a cell Inputs are signaling
pathways and regulatory proteins known as
transcription factors
• Signaling pathways respond to signals and
activate the transcription factor proteins
Transcription factors bind to genes and make
mRNA The mRNA synthesizes the required
proteins
Application Areas of Bioinformatics
1. Medical and Healthcare
✔ Drug Discovery
✔ Personalized Medicine
✔ Genetic Disease Diagnosis
✔ Gene Therapy Development
2. Genomics and Proteomics
✔ Genome Sequencing and Assembly
✔ Protein Structure Prediction
✔ Protein-Protein Interaction Analysis
3. Evolutionary Biology
✔ Comparative Genomics
✔ Phylogenetic Analysis
Application Areas of Bioinformatics
4. Microbiology and Immunology
✔ Microbial Genome Analysis
✔ Vaccine Development
5. Agriculture and Biotechnology
✔ Crop Improvement
✔ Animal Breeding
6. Environmental Science
✔ Metagenomics
✔ Bioremediation
Thank You

You might also like