0% found this document useful (0 votes)
2 views19 pages

Human Genome Project Bioinfo

The Human Genome Project (HGP) was an international initiative aimed at mapping and sequencing the entire human genome, identifying approximately 30,000 genes and revealing that only about 2% of DNA codes for proteins. It has significant implications for molecular medicine, agriculture, and biotechnology, and laid the groundwork for modern genomics and bioinformatics. Bioinformatics itself is a multidisciplinary field that manages and analyzes biological data, essential for handling the vast amounts of information generated by genomic studies.

Uploaded by

manasvisuchindra
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views19 pages

Human Genome Project Bioinfo

The Human Genome Project (HGP) was an international initiative aimed at mapping and sequencing the entire human genome, identifying approximately 30,000 genes and revealing that only about 2% of DNA codes for proteins. It has significant implications for molecular medicine, agriculture, and biotechnology, and laid the groundwork for modern genomics and bioinformatics. Bioinformatics itself is a multidisciplinary field that manages and analyzes biological data, essential for handling the vast amounts of information generated by genomic studies.

Uploaded by

manasvisuchindra
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Human Genome Project (HGP)

The Human Genome Project (HGP) was an international scientific research program aimed at understanding
the complete genetic makeup of humans. It focused on identifying, mapping, and sequencing all the DNA
present in the human genome.

What is the Human Genome?

The human genome is the entire DNA content present in the nucleus of a human cell. It contains genes that
provide instructions for making proteins required for growth, development, and maintenance of the body. It
also includes non-coding DNA and genes for rRNA and tRNA involved in protein synthesis.

Goals of the Human Genome Project

The major objectives of the HGP were:

1. To identify all the genes present in human DNA

2. To determine the sequence of all 3 billion DNA base pairs (A, T, G, C)

3. To store this genetic information in public databases

4. To develop faster and more accurate DNA sequencing technologies

5. To create tools for analyzing large genomic data (bioinformatics)

6. To address ethical, legal, and social issues (ELSI) arising from genome research

Timeline and Milestones

 1990: Project initiated jointly by the U.S. Department of Energy (DOE) and National Institutes of
Health (NIH)

 June 2000: A working draft covering over 90% of the genome was completed

 February 2001: Analysis of the draft genome published

 April 2003: Final genome sequence completed — two years ahead of schedule

Key Findings of the HGP

 The human genome contains about 3 billion DNA base pairs

 Humans have approximately 30,000 genes, much fewer than earlier estimates

 Only less than 2% of DNA codes for proteins

 Around 50–75% of DNA is non-coding or repetitive (“junk DNA”)

 99.9% of DNA is identical among all humans

 Chromosome 1 has the most genes, while the Y chromosome has the fewest
 About 3 million SNPs (single nucleotide polymorphisms) were identified, explaining individual
differences

Importance and Applications

1. Molecular Medicine

o Improved disease diagnosis

o Detection of genetic disease risk

o Development of targeted drugs

o Gene therapy

o Personalized medicine (pharmacogenomics)

2. Microbial Genomics

o Faster detection of pathogens

o Development of biofuels

o Environmental monitoring

o Bioremediation of toxic waste

o Protection against biological warfare

3. Agriculture and Biotechnology

o Disease-resistant crops

o Improved livestock breeding

o Nutrient-rich food production

Conclusion

The Human Genome Project revolutionized biology and medicine by providing a complete reference of human
DNA. It laid the foundation for modern genomics, personalized medicine, bioinformatics, and future research
into complex diseases. The HGP remains one of the most important scientific achievements in human history.

Bioinformatics – 10 Marks Answer

Bioinformatics is a multidisciplinary field that combines biology, computer science, mathematics, and
information technology to store, manage, analyze, and interpret biological data. It mainly deals with large
amounts of data generated from DNA, RNA, and protein studies.

Bioinformatics is often described as a “marriage between biology and computers” and is considered the
electronic infrastructure of modern molecular biology.

Definition of Bioinformatics
Bioinformatics is the science of collecting, storing, analyzing, and interpreting biological data, especially
genetic and protein data, using computers and computational tools.
Its ultimate goal is to discover new biological insights and understand life processes at a molecular level.

Why Bioinformatics is Needed

Modern biology produces huge amounts of data, especially from projects like the Human Genome Project.
Manual analysis is impossible, so bioinformatics helps to:

 Handle massive datasets

 Compare new sequences with existing databases

 Predict structure and function of genes and proteins

What is Done in Bioinformatics

Bioinformatics involves:

 Analysis of DNA and RNA sequences

 Analysis of protein sequences and structures

 Identification of genes in genomic sequences

 Sequence alignment and comparison

 Studying evolutionary relationships using phylogenetic trees

Major Areas (Branches) of Bioinformatics

1. Genomics – gene finding, genome mapping, sequence alignment

2. Proteomics – protein structure prediction and function analysis

3. Computer-Aided Drug Design – identifying drug targets and designing drugs

4. Databases and Data Mining – storing and retrieving biological data

5. Molecular Phylogenetics – studying evolutionary relationships

Bioinformatics Tools and Technologies

 Biological databases (DNA, RNA, protein databases)

 Algorithms like BLAST, FASTA, and Smith-Waterman

 Programming and scripting languages such as Perl and Python

 Use of Linux operating systems and internet-based tools


Applications of Bioinformatics

1. Medical Applications

o Understanding genetic diseases

o Drug discovery and development

o Personalized medicine and pharmacogenomics

2. Agricultural Applications

o Disease-resistant and high-yield crops

o Drought-resistant plants

3. Pharmaceutical and Biotechnology Industry

o Structure-based and gene-based drug design

Role of Computers in Bioinformatics

Computers help in:

 Locating genes in DNA sequences

 Predicting RNA and protein structures

 Studying protein-protein interactions

 Managing and integrating biological databases

Challenges in Bioinformatics

 Data generation is faster than data analysis

 Need for experts skilled in both biology and computers

 Bioinformatics predictions require experimental validation

Conclusion

Bioinformatics plays a crucial role in modern biological research by enabling efficient analysis of large biological
datasets. It supports genomics, proteomics, drug discovery, and disease research, making it an essential field in
biotechnology and life sciences.

Biological Databases – Simple Explanation

Biological databases are like digital libraries that store information related to living organisms. They collect
and organize biological data so that scientists can easily store, search, and analyze it using computers.

Sources of Biological Databases

The information in biological databases comes from:


 Scientific experiments

 Research papers and published literature

 High-throughput technologies (experiments that generate large amounts of data quickly)

 Computer-based analyses (computational analysis)

Types of Information Stored

Biological databases contain data from many areas of life sciences, such as:

 Genomics – information about DNA and genes

 Proteomics – information about proteins

 Metabolomics – information about metabolites and biochemical pathways

 Microarray gene expression – data showing which genes are active

 Phylogenetics – evolutionary relationships between organisms

Details Stored in Biological Databases

These databases store important biological information including:

 Gene function – what a gene does

 Gene and protein structure – shape and arrangement

 Localization – where a gene or protein is found

o inside the cell

o on a chromosome

 Clinical effects of mutations – how DNA changes cause diseases

 Sequence and structure similarities – similarities between genes and proteins of different organisms

Primary Databases – Simple Explanation

Primary databases are biological databases that store original data collected directly from experiments.
The data is submitted by scientists (authors/experimentalists) who performed the research.

Key Features of Primary Databases

 Data comes from direct experimental results, not from analysis or interpretation

 Authors submit their own data to the database

 The submitted data is checked and curated (reviewed for correctness) by the database team

 They store different levels of sequence information, such as:

o Raw DNA sequences

o Processed or annotated sequences

 The content is controlled by the submitter, meaning the original researcher is responsible for
accuracy
Example: Nucleotide Sequence Databases

These databases store DNA and RNA sequences.

1. GenBank

 Hosted by the National Center for Biotechnology Information (NCBI)

 One of the largest nucleotide sequence databases in the world

2. EMBL

 European Molecular Biology Laboratory

 Maintains nucleotide sequence data for Europe

3. DDBJ

 DNA Data Bank of Japan

 Maintains nucleotide sequences for Asia

🔹 These three databases exchange data with each other, so the same sequence can be found in all of them.

Secondary Databases – Simple Explanation

Secondary databases are databases that are created using information taken from primary databases.
They do not store raw experimental data. Instead, they store analyzed, curated, and organized information.

Key Features of Secondary Databases

 Built from primary data (data taken from primary databases like GenBank)

 Information is curated by the database experts, not by the original authors

 They are specialized databases that focus on specific types of biological information

 The content is controlled by third-party organizations, such as NCBI

Examples of Secondary Databases

1. RefSeq (Reference Sequence)

 A curated, non-redundant collection of DNA, RNA, and protein sequences

 Provides standard reference sequences for genes

2. TPA (Third Party Annotation)

 Contains sequence annotations submitted by researchers who are not the original data submitters

3. RefSNP

 Stores information about single nucleotide polymorphisms (SNPs)

 Helps in studying genetic variations and diseases

4. UniGene

 Organizes gene sequences into clusters representing a single gene

5. NCBI Protein
 Contains curated protein sequences derived from nucleotide data

6. Structure Database

 Stores 3D structures of proteins and nucleic acids

Metadatabases – Simple Explanation

A metadatabase is a database of databases.


Instead of storing raw data, it collects data from multiple databases and presents it in a more organized and
user-friendly way.

Metadatabases often focus on:

 A specific organism

 A specific disease

 A specific biological function

Examples of Metadatabases

1. Entrez

 Developed by NCBI (National Center for Biotechnology Information)

 Integrates data from many databases such as genes, proteins, structures, and literature

 Provides a single search interface for multiple biological databases

2. euGenes

 Integrates gene-related data from different sources

 Helps in comparative genomics studies

Genome Databases – Simple Explanation

Genome databases are databases that collect, store, analyze, and annotate complete genome sequences of
organisms.
They make genome information publicly available so researchers around the world can study genes, proteins,
and diseases.

Some genome databases:

 Store genomes of many species

 Others focus on one model organism

Many of these databases also curate information from scientific literature to improve the accuracy of gene
annotations.

Examples of Genome Databases

1. CAMERA

 Cyberinfrastructure for Advanced Microbial Ecology Research and Analysis

 Focuses on microbial genomics and metagenomics


 Used for studying environmental microbes

2. Corn (Maize) Genetics and Genomics Database

 Contains genome information of maize (corn)

 Useful for plant genetics and agricultural research

Protein Databases – Simple Explanation

Protein databases store information related to protein sequences, structures, domains, and models.
They help scientists understand protein function, structure, and evolution.

Protein databases are mainly of three types:

1. Protein sequence databases

2. Protein structure databases

3. Protein model databases

1. Protein Sequence Databases

These databases store amino-acid sequences of proteins, along with functional and biological information.

Examples

UniProt – Universal Protein Resource

 One of the most comprehensive protein databases

 Contains protein sequences and functional information

Swiss-Prot – Protein Knowledgebase

 Maintained by the Swiss Institute of Bioinformatics (SIB)

 Manually curated and highly reliable

 Contains well-annotated protein sequences

2. Protein Structure Databases

These databases store 3-dimensional (3D) structures of proteins.

Protein Data Bank (PDB)

 Maintained by the Research Collaboratory for Structural Bioinformatics (RCSB)

 Stores experimentally determined 3D protein and nucleic acid structures

Specialized Biological Databases – Simple Explanation

These databases store specific types of biological information such as carbohydrates, protein interactions,
signaling pathways, and metabolic pathways.

1. Carbohydrate Structure Databases

These databases store information about carbohydrate (sugar) structures and sequences, which are important
for cell recognition and signaling.
2. Protein–Protein Interaction Databases

These databases store information about how proteins interact with each other inside the cell.

3. Signaling Pathway Databases

These databases store information about cell signaling pathways, which control cell communication and
responses.

4. Metabolic Pathway Databases

These databases store information about biochemical and metabolic pathways in organisms.

Clinical and Mutation Databases – Simple Explanation

Clinical and mutation databases store information about genes, mutations, and diseases.
They help researchers and doctors understand genetic disorders, disease causes, and inheritance patterns.

1. OMIM – Online Mendelian Inheritance in Man

OMIM is a comprehensive database that focuses on human genes and genetic diseases.

Key Features of OMIM

 Contains information on disease-linked genes

 Describes associated phenotypes (observable characteristics of diseases)

 Focuses mainly on Mendelian (inherited) disorders

 Provides detailed summaries of genetic conditions

Database Links

OMIM is linked to:

 Entrez

 GDB (Genome Database)

 Other related biological databases

This allows easy access to gene sequences, literature, and additional data.

3. Disease-Specific Mutation Databases

These databases focus on mutations related to a specific disease or gene.

Key Features

 Provide detailed mutation information for individual diseases

 Useful for:

o Clinical diagnosis

o Genetic counseling

o Disease-specific research
Feature GenBank RefSeq

Type of database Primary nucleotide sequence database Secondary, curated reference database

Full form GenBank (no expanded full form) RefSeq – Reference Sequence

NCBI (National Center for Biotechnology NCBI (National Center for Biotechnology
Maintained by
Information) Information)

Derived from GenBank and other primary


Source of data Direct submissions from researchers
databases

Curated, reviewed, and non-redundant


Nature of data Original experimental sequence data
sequences

Highly redundant (same gene may Non-redundant (one standard reference


Redundancy
appear many times) sequence per gene)

Data control Controlled by the original submitter Controlled and curated by NCBI experts

Annotation
Variable quality; depends on submitter High-quality, standardized annotation
quality

To act as a global archive of all submitted To provide a standard reference for genes,
Purpose
sequences transcripts, and proteins

Sequence
Updated when submitters revise data Regularly updated and corrected by curators
updates

Examples of Raw DNA/RNA sequences from Reference genomic, mRNA, and protein
entries experiments sequences

Comparative genomics, annotation, clinical


Usage Data submission and archival
research

NCBI – National Center for Biotechnology Information (10 Marks)

Introduction

NCBI stands for National Center for Biotechnology Information.


It is a major public bioinformatics organization that provides free access to biological data, tools, and
databases for researchers, students, doctors, and scientists all over the world.

NCBI was established in 1988 and is a part of the National Institutes of Health (NIH), USA.

Main Objectives of NCBI

The main goals of NCBI are:

 To collect and store biological data

 To organize and analyze genetic and molecular information

 To provide free public access to biological databases

 To support research in genomics, proteomics, medicine, and biotechnology


Major Databases Maintained by NCBI

1. GenBank

o A primary database for DNA and RNA sequences

o Accepts direct submissions from researchers worldwide

2. RefSeq (Reference Sequence)

o A curated, non-redundant database

o Provides standard reference sequences for genes, transcripts, and proteins

3. Protein Database

o Stores protein sequences and functional annotations

4. Structure Database

o Stores 3D structures of proteins and nucleic acids

5. OMIM

o Database of human genes and genetic diseases

6. SNP / RefSNP

o Contains information on genetic variations (SNPs)

Entrez System

 Entrez is an integrated search system developed by NCBI

 It links:

o DNA sequences

o Protein sequences

o Structures

o PubMed / MEDLINE (biomedical literature)

 Allows easy navigation between related biological data

Tools Provided by NCBI

1. BLAST (Basic Local Alignment Search Tool)

o Used to compare DNA or protein sequences

o Helps find similarities and identify unknown sequences

2. Genome Browsers
o Used to view and analyze genome data

3. Analysis Tools

o Used for gene prediction, alignment, and annotation

Importance of NCBI

 Acts as a central hub for bioinformatics

 Essential for:

o Genome research

o Disease studies

o Drug discovery

o Evolutionary biology

 Supports projects like the Human Genome Project

Users of NCBI

 Researchers

 Students

 Clinicians

 Biotechnologists

 Pharmaceutical scientists

OMIM – Online Mendelian Inheritance in Man (10 Marks Answer)

Introduction

OMIM (Online Mendelian Inheritance in Man) is a comprehensive clinical genetics database that provides
detailed information about human genes, genetic disorders, and inherited traits. It mainly focuses on
Mendelian (single-gene) diseases and their relationship with specific genes.

Full Form

OMIM – Online Mendelian Inheritance in Man

Maintained By

 Johns Hopkins University, USA

 Developed originally by Dr. Victor McKusick, known as the father of medical genetics
Type of Database

 Clinical and mutation database

 Secondary database (information curated from literature and primary databases)

What OMIM Consists Of

OMIM contains detailed, text-based and structured information such as:

 Human gene descriptions

 Genetic mutations

 Inherited genetic disorders

 Phenotypes (observable traits or symptoms)

 Mode of inheritance (autosomal dominant, recessive, X-linked, etc.)

 Clinical features and disease mechanisms

 References from scientific and medical literature

OMIM Entry Structure

Each OMIM entry includes:

 A unique OMIM identification number

 Gene name and symbol

 Disease name (if applicable)

 Clinical summary

 Molecular genetics information

 References and external links

Key Features

 Focuses only on humans

 Regularly updated with new research findings

 Strongly linked with NCBI databases, including:

o Entrez Gene

o GenBank

o RefSeq

o PubMed
 Allows easy navigation from gene → mutation → disease

Importance of OMIM

 Helps in understanding the genetic basis of human diseases

 Useful for diagnosis and classification of genetic disorders

 Supports medical genetics research and education

 Plays an important role in genetic counseling

Applications

 Identification of disease-causing genes

 Study of inherited disorders

 Clinical diagnosis and prognosis

 Precision medicine and personalized treatment

 Teaching and research in medical genetics

Relation with NCBI

 OMIM is not directly maintained by NCBI

 However, it is integrated with NCBI through the Entrez system

 OMIM entries are linked to:

o Gene sequences

o Protein sequences

o Scientific publications

. PDB – Protein Data Bank (10 Marks Answer)

Introduction

The Protein Data Bank (PDB) is a biological database that stores three-dimensional (3D) structures of
proteins and nucleic acids, which are essential for understanding molecular function.

Full Form

PDB – Protein Data Bank

Maintained by

 RCSB (Research Collaboratory for Structural Bioinformatics)

What PDB Consists Of

 3D structures of proteins
 3D structures of DNA and RNA

 Structural coordinates and models

Sources of Data

 X-ray crystallography

 Nuclear Magnetic Resonance (NMR)

 Cryo-Electron Microscopy (Cryo-EM)

Key Features

 Freely accessible worldwide

 High-quality experimentally determined structures

 Linked to NCBI databases via Entrez

Type of Database

 Protein structure database

Importance

 Helps understand structure–function relationship of proteins

 Essential in drug design and molecular modeling

Applications

 Structure-based drug discovery

 Protein engineering

 Structural biology research

Conclusion

PDB is a vital resource that provides detailed structural information of biomolecules, helping scientists
understand biological processes at the molecular level.

Why are match and mismatch penalties important in a scoring scheme?

In sequence alignment, a scoring scheme is used to decide how similar two biological sequences (DNA, RNA,
or protein) are. Match scores and mismatch penalties are important because they help distinguish real
biological similarity from random similarity.

1. Reflect biological reality

 A match means the same nucleotide or amino acid occurs in both sequences.

 This often indicates conservation, meaning the position is important for structure or function.

 Therefore, matches are given positive scores.

 A mismatch indicates a difference, which may be due to mutation, so it is given a penalty (negative
score).
2. Helps find the best alignment

 Without match rewards and mismatch penalties, many alignments would look equally good.

 Penalties force the algorithm to choose alignments with more biologically meaningful similarities.

3. Reduces random alignments

 Random sequences can accidentally match at some positions.

 Penalizing mismatches ensures that random similarities do not get high scores.

4. Differentiates conserved and variable regions

 Highly conserved regions will have many matches and few mismatches, giving high scores.

 Variable regions will accumulate penalties, lowering the score.

 This helps identify functional or evolutionary important regions.

5. Essential for evolutionary interpretation

 Match/mismatch scoring models mutations over time.

 Fewer mismatches suggest closer evolutionary relationship.

 More mismatches suggest distant relationship.

6. Improves accuracy of sequence comparison

 Correct scoring leads to accurate homology detection.

 This is crucial for:

o Gene prediction

o Protein function annotation

o Phylogenetic analysis

Conclusion

Match scores and mismatch penalties are essential components of a scoring scheme because they ensure
that sequence alignments are biologically meaningful, accurate, and reliable, rather than random or
misleading.

Difference between Global and Local Sequence Alignment


Global Sequence Alignment Local Sequence Alignment

Tries to align the entire length of both sequences from Finds and aligns only the most similar regions of the
start to end. sequences.

All letters from both sequences are included in the


Only substrings (parts) of the sequences are aligned.
alignment.

Best when sequences are of similar length. Can be used for sequences of different lengths.

Suitable for distantly related or unrelated


Suitable when sequences are closely related.
sequences.

Ignores poorly matching regions and focuses on high


Forces alignment even in regions with low similarity.
similarity areas.

Used to compare homologous genes or proteins with


Used to find conserved regions, domains, or motifs.
the same function.

Gives an overall similarity score for the whole


Gives a score only for the best matching region.
sequence.

More sensitive to gaps and mismatches over the whole


Less affected by gaps outside the matching region.
sequence.

Needleman–Wunsch algorithm is commonly used. Smith–Waterman algorithm is commonly used.

Examples of tools: EMBOSS Needle, Needleman–


Examples of tools: BLAST, EMBOSS Water, LALIGN
Wunsch

Sequence Alignment (10 Marks Answer)

Introduction

Sequence alignment is a fundamental concept in bioinformatics used to compare biological sequences such
as DNA, RNA, or proteins. It arranges sequences in a way that similar characters (nucleotides or amino acids)
are placed in the same column to identify similarities, differences, and evolutionary relationships.

Definition

Sequence alignment is the process of arranging two or more biological sequences to maximize similarity by
introducing gaps, so that conserved regions, mutations, insertions, and deletions can be identified.

Types of Sequences Aligned

 DNA sequences

 RNA sequences

 Protein sequences
Purpose of Sequence Alignment

 To identify homologous sequences

 To study evolutionary relationships

 To predict structure and function of genes and proteins

 To find conserved regions, motifs, and domains

 To detect mutations and variations

Types of Sequence Alignment

1. Pairwise Alignment

o Aligns two sequences

o Can be:

 Global alignment (entire sequence)

 Local alignment (best matching regions)

2. Multiple Sequence Alignment (MSA)

o Aligns three or more sequences

o Helps find highly conserved regions

Scoring Scheme in Sequence Alignment

A scoring system is used to evaluate the quality of an alignment:

 Match score – given when characters are the same

 Mismatch penalty – given when characters differ

 Gap penalty – given for insertions or deletions

This ensures biologically meaningful alignments.

Common Algorithms

 Needleman–Wunsch – Global alignment

 Smith–Waterman – Local alignment

 BLAST – Fast local alignment tool

 ClustalW – Multiple sequence alignment

Applications of Sequence Alignment


 Gene and protein annotation

 Comparative genomics

 Phylogenetic analysis

 Disease mutation identification

 Drug discovery and protein modeling

Importance of Sequence Alignment

 Helps understand biological function

 Reveals conserved functional regions

 Forms the basis of many bioinformatics tools and analyses

You might also like