0% found this document useful (0 votes)
13 views8 pages

Bioinformatics: Data, History, and Omics

This document describes bioinformatics as a discipline that uses computer, mathematical, and statistical tools to analyze biological data. It explains the concepts of genomics, transcriptomics, and proteomics, and describes the role of biological databases in storing genetic information.

Translated by

ScribdTranslations
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
13 views8 pages

Bioinformatics: Data, History, and Omics

This document describes bioinformatics as a discipline that uses computer, mathematical, and statistical tools to analyze biological data. It explains the concepts of genomics, transcriptomics, and proteomics, and describes the role of biological databases in storing genetic information.

Translated by

ScribdTranslations
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

1erChapter

Databases in bioinformatics

I. Bioinformatics: definition, description, approach and history


[Link]
Bioinformatics is a discipline of life sciences that relies on tools
computing, mathematics, and statistics for storing, analyzing, and visualizing data
biologicals such as DNA sequences (genomes), proteins,
sugars or results of experiments.

Bioinformatics is the information related to biological molecules: their sequence, their


name, their structure(s), their function(s), their 'kinship' links, their interactions and
their integration into the cell ...

This bioinformation comes from various disciplines: biochemistry, genetics,


structural genomics, functional genomics, transcriptomics, proteomics...

Definition of bioinformatics according to NCBI (2001): "Bioinformatics is the field of


science in which biology, computer science, and information technology merge into a
single discipline.

Description

recent discipline (a few decades).


hybrid discipline: it is based on concepts (general ideas) and
formalisms derived from biology, computer science, mathematics and
physics, chemistry (sequencing techniques, ...).
discipline that uses the full processing potential of computing: models
theoretical,algorithms and programs, databases, computers, Internet network,
communication protocolslanguages,...

3. Approach

Compilation and organization of biological data in databases

2. Systematic data processing: One of the objectives of bioinformatics is to


identify and characterize an important function and/or biological structure.
results of these treatments constitute new biological data obtained 'in
silico1.

1
in silicoa research or an essay conducted using complex computerized calculations or computer models.
Silico is widely used in bioinformatics, for example for gene research that can be done in silico through programs.
gene detection, then situpour to experimentally validate the predictions made by computer.

1
3. Development of strategies

bring additional biological knowledge by combining the data


initial biological data and the biological data obtained 'in silico'.
This knowledge, in turn, allows for the development of new concepts in
biology, which, to be validated, may require the development of new
theories and tools in mathematics and computer science.

4. History
Key steps in molecular biology, computing, and bioinformatics

Margaret Dayhoff et al.: First compilation of proteins


1965
Atlas of Protein Sequences

Algorithm for global sequence alignment: Saul


1970
Needleman & Christian Wunsch

Cloning of DNA fragments in a virus, the DNA


1972
recombined

1973 Discovery of restriction enzymes

Secondary structure prediction program of


proteins: "Prediction of Protein Conformation"Chou &
Fasman.
1974

Development of the concept of networks connecting computers


within an 'internet'

Development of microcomputers accessible to everyone.


1977
DNA sequencing techniques: Frederick Sanger

Directed mutagenesis; Sequencing of the first DNA genome,


1978 - 1980 bacteriophage phiX174 (Frederick Sanger
First databases: EMBL, GenBank, PIR
1981: 370,000 nucleotides
270 sequences Local sequence alignment program

DNA amplification: polymerization reaction in


1984
chain (PCR)

1985 Local sequence alignment program

Taq polymerase, thermostable enzyme for PCR.


1988 Creation of the 'National Centre for Biotechnology'
Information (NCBI).

2
1990 Local sequence alignment program

1992 Complete sequencing of yeast chromosome III

1996 Complete sequencing of yeast (European consortium).

11 genomes bacteria sequenced


1997
Evolutions of BLAST

1998 Sequencing of 2 million nucleotides per day.

2000 Sequencing of the first plant genome: Arabidopsis thaliana

Access to scientific journals and periodicals: development


2000s
of 'open access'.

2003 complete sequencing of the human genome

Emergence of new high-throughput sequencing technologies


broadband, referred to as second generation and now thirdeme
generation.
2007 - 2008
Awareness of the 'big data' phenomenon (not
only in biology) which gradually becomes a discipline
scientific.

II. The Omic Family


Reminder
There are two types of molecules that support bioinformation: nucleic acids
(DNA or RNA) and proteins. The sequence is the ordered and oriented chain of
nucleotides (nucleic acids: DNA and RNA) or amino acids (proteins). The sequence
constitutes the 'base material' of genomics, transcriptomics, and proteomics.

There are many scientific fields whose name has been created with the suffix "omic."
"Omics" is an Anglo-Saxon neologism.

3
5. Genomics
The genome is the set of chromosomes of an organism (coding sequences +
non-coding sequences

The size of genomes varies from one individual to another:

Prokaryotes: from 500,000 bp to 13 Mb


Eucaryotes: certains champignons (8Mb); Homme (3.2 Gb); Blé: (16 Gb);
amibe (686 Gb)

Genomics is a discipline that allows for the comprehensive study and analysis.
multidisciplinary of genomes. It aims to inventory all the genes of a
organism to localize on the chromosomes and to characterize their sequences as well as
study their functions

Genomics began with the first large sequencing projects that used the
Frederick Sanger method

Haemophilus influenzae 1995 Arabidopsis thaliana2000


Saccharomyces cerevisiae 1996 Drosophila melanogaster 2000
Escherichia coli K-121997 Man2001
Caenorhabditis elegans 1998 House mouse 2002

6. Proteomics
The proteome is the entire set of proteins expressed in a cell, a part of a
cell (membranes, organelles) or a group of cells (organ, organism) in
conditions given at a given time.

Proteomics groups research on detection, separation, and identification.


sequencing of the entire set of proteins in a proteome, to determine their activities, their
functions to analyze their interactions and their modifications over time.

The causes of the variability and complexity of the proteome:

alternative splicing of primary transcripts (multiple mRNAs for a gene),


post-translational modifications of proteins

for each condition environmental (condition physiological


normal stress conditions) a cell is characterized by a proteome adapted to
this condition even though it always has the same genome.

For example: plants adapt to variations in light and biotic stress.

4
In addition to post-translational modifications, proteins undergo
transformations once synthesized: cleavage of the signaling address peptide,
activation of the native form from a precursor (zymogen), assembly into
oligomeric complexes, association with cofactors.

III. Information storage: databases

In computing, a database is a collection of objects exhibiting properties.


and/or common characters that can be reused in a processing procedure.

Biological sequences (nucleic or protein) are collected in databases.


biological data. The greatest contribution of databases to the community of
biologists aim to make sequences accessible.

1. General-purpose databases

They correspond to a data collection that is as comprehensive as possible and provides a


rather heterogeneous set of information (viruses, bacteria, fungi, plants,
animals, .....)

Generalist databases are essential to the scientific community because they


regroup essential data and results. They mainly contain
experimental results, but which are neither verified nor analyzed.

There are a large number of generalist databases of biological interest. These include:

A. Bases of nucleic sequences:

American GenBank base 216 million sequences (October 2019) managed by the
National Center for Biotechnology Information (NCBI)
The provided text is a URL and contains no translatable [Link]
EMBLbase European maintained by the European Bioinformatics Institute (EBI)
DDBJ (DNA Database of Japan) Japanese base

These three databases manage all nucleic acid sequences and their annotations: they
cooperate and exchange their data daily to ensure consistency
maximum in the provision of sequences to the scientific community.

GenBank data format

Example: Consult the GenBank database to search for the sequence XM_015777817.2

5
Each entry corresponds to a primary nucleic sequence associated with
annotations2The sequence is available in a plain text file format.3where the lines
corresponding to keyword/value associations in a format specific to the GenBank database
called GenBank format.

The entry is structured into four parts:

1erapart: The header containing general information about the sequence: identifier
unique, accession number, definition, keyword, taxonomy of the organism whose
sequence comes from

2thPart: Describe the bibliographic references associated with the sequence.

3part: essential, describes the biological annotations associated with the sequence below
standardized form: we talk about features the characteristics of annotations.

4emepart: contains the nucleic sequence itself in GenBank format. The


The format used in bioinformatics is the FASTA format.

A. Protein sequence databases

Origin of the sequences:

Automatic translation of DNA sequences (mostly)


protein sequencing (rare because it is long and expensive)
Proteins whose 3D structure is known

Origin of the annotations

Mass spectrometry: regulation and localization of protein expression; but


also identification and post-transcriptional modification
Study of interactions: how proteins assemble with each other or with others
molecules to form molecular complexes
Crystallography and nuclear magnetic resonance: to determine the 3D shape
end of the protein

The protein databases are as follows:

. PIR Protein Information Resource: American bank


. SWISSPROT: European bank

2
Annotate: to accompany a text (for example) with notes or comments. Genome annotation consists of predicting and
localize the entire set of coding sequences (genes) of the genome and determine and identify their structure (annotation)
syntax), their function (functional annotation) as well as the relationships between the biological entities related to the genome
(relational annotation). The resulting information enriches biological databases.

3
A flat file is an unencrypted file, usually in text form, whose content can be interpreted.
independently of software.

6
. TrEMBL: automatic translation of contained coding sequences
in EMBL

From 2002 onwards, these three banks joined forces to give rise to UniProt.
Universal Protein Resource. In 2019, UniProt contains 559,000 sequences, with a
precise, coherent, and rich annotation.

Example: Consult the UniProt database to search for the sequence P02769

B. Bases of protein structures

In the field of protein structures, the Protein Databank (PDB)


([Link] and disseminate all available data on the structures
crystallographic data of proteins as well as some nucleotide structures. The PDB
contains the three-dimensional atomic coordinates of proteins, nucleic acids,
nucleoprotein complexes (it contains more than 159,000 structures as of January 2020).

Example: Consult the PDB database to determine the structure of the protein 4F5S

C. Objectives of general-purpose databases

Make the sequences and any other type of information, such as the broadcast, public.
of the results of thesequencing of the human genome.

Searching for similarities between sequences recorded in the same database


data with a new sequence.

Analysis of evolutionary type thanks to the great diversity of organisms represented in the
database.

Presence of information accompanying the sequences: the annotations and the


bibliography.

Presence of links to other databases

[Link] databases
They correspond to more homogeneous data established around a theme:

biological theme: database of receptors coupled to proteins


bacillus subtilis, drosophila melanogaster...
Technology: NMR spectrum, mass spectrometry map, electrophoresis gel
two-dimensional
Data types: sequences, structures, image, spectrum, interaction

Specialized databases have the advantage of being maintained by experts in the field.
domain.

7
A. Resources for prokaryotes

The two most commonly used databases of complete prokaryotic genomes


are:

The "Microbial Genomes" section of the NCBI RefSeq database


Unable to translate URL content.
The 'Ensembl Bacteria' section of the 'Ensembl Genomes' database
[Link]

The Ensembl project provides an integrated environment of databases and interfaces.


graphics to annotate and compare large chromosomal sequences from
the entire set of available data.

D. Resources for animals

Ensembl is a key data resource for eukaryotic genomes.

EnsemblMetazoa:[Link]
Vertebrate:[Link]

E. Resources for plants

EnsemblPlants:[Link]
TAIR: The Arabidopsis Information Resource. This database centralizes the
most of the information available on Arabidopsis thaliana
-FLAGdb++: this database integrates genomic data from Arabidopsis,
rice, poplar, grapevine, tomato, and melon.
Gramene: international reference for cereals

F. Resources for mushrooms

EnsemblFungi[Link] is the mushroom part of the Ensembl project


SGD: Saccharomyces Genome Database, a database focused on molecular biology
and the genetics of the baker's yeast S. cerevisiae.

You might also like