Understanding Bioinformatics Basics
Understanding Bioinformatics Basics
1.1 Introduction
Bioinformatics is a newly emerged scientific discipline for the computational analysis and storage of biological
data. The word bioinformatics has been derived from two words ‘Bio’ means biology and ‘Informatique’ (a French
word) meaning ‘data processing’.
Bioinformatics is the combination of biology and information technology. The discipline encompasses any
computational tools and methods used to manage, analyse and manipulate large sets of biological data. Essentially,
bioinformatics has three components:
• The creation of databases allowing the storage and management of large biological data sets.
• The development of algorithms and statistics to determine relationships among members of large data sets.
• The use of these tools for the analysis and interpretation of various types of biological data, including DNA,
RNA and protein sequences, protein structures, gene expression profiles, and biochemical pathways.
The term bioinformatics first came into use in the 1990s and was originally synonymous with the management and
analysis of DNA, RNA and protein sequence data. Computational tools for sequence analysis had been available since
the 1960s, but this was a minority interest until advances in sequencing technology, which led to a rapid expansion
in the number of stored sequences in databases such as GenBank. Now, the term has expanded to incorporate many
other types of biological data, for example protein structures, gene expression profiles and protein interactions. Each
of these areas requires its own set of databases, algorithms and statistical methods.
Bioinformatics is largely, although not exclusively, a computer-based discipline. Computers are important in
bioinformatics for two reasons:
Firstly, many bioinformatics problems require the same task to be repeated millions of times. For example, comparing
a new sequence to every other sequence stored in a database or comparing a group of sequences systematically
to determine evolutionary relationships. In such cases, the ability of computers to process information and test
alternative solutions rapidly is indispensable.
Secondly, computers are required for their problem-solving power. Typical problems that might be addressed using
bioinformatics could include solving the folding pathways of protein given its amino acid sequence, or deducing a
biochemical pathway given a collection of RNA expression profiles. Computers can help with such problems, but
it is important to note that expert input and robust original data are also required.
Bioinformatics is the field in which biology, computer science and information technology merge into single discipline
for managing and analysing biological data using advanced computing techniques. Bioinformatics has emerged
as a full-fledged interdisciplinary subject that interfaces the developments of computer science and information
technology with biological sciences. The knowledge of computer science and information technology is applied for
creation as well as management of databases, data warehousing, data mining and overall communication networking
throughout the world.
Bioinformatics is the application of computer technology to the management and analysis of biological
data. The result is that computers are being used to gather, store, analyse and merge biological data.
The ultimate goal of bioinformatics is to uncover the wealth of biological information hidden in the mass of data
and obtain a clearer insight into the fundamental biology of organisms. This new knowledge could have profound
impacts on fields as varied as human health, agriculture, the environment, energy and biotechnology.
Bioinformatics is conceptualising biology in terms of molecules (in the sense of physical-chemistry) and then
applying “informatics” techniques (derived from disciplines such as applied math, CS, and statistics) to understand
and organise the information associated with these molecules, on a large-scale.
2
The three terms: bioinformatics, computational biology and bioinformation infrastructure are very similar and most
of the time used interchangeably. However;
• Bioinformatics refers to database like activities, involving persistent sets of data that are maintained in a consistent
state over essentially indefinite periods of time.
Computational biology encompasses the use of algorithmic tools to facilitate biological analyses.
Bioinformation infrastructure comprises the entire collection of information management systems, analysis
tools and communication networks supporting biology. Thus, the latter may be viewed as a computational
scaffold of the former two.
The future of bioinformatics is integration. For example, integration of a wide variety of data sources such as clinical
and genomic data will allow us to use disease symptoms to predict genetic mutations and vice versa. The integration
of GIS data, such as maps, weather systems, with crop health and genotype data, will allow us to predict successful
outcomes of agriculture experiments.
Another future area of research in bioinformatics is large-scale comparative genomics. For example, the development
of tools that can do 10-way comparisons of genomes will push forward the discovery rate in this field of bioinformatics.
Along these lines, the modelling and visualisation of full networks of complex systems could be used in the future
to predict how the system (or cell) reacts to a drug for example. A technical set of challenges faces bioinformatics
and is being addressed by faster computers, technological advances in disk storage space, and increased bandwidth.
Finally, a key research question for the future of bioinformatics will be how to computationally compare complex
biological observations, such as gene expression patterns and protein networks. Bioinformatics is about converting
biological observations to a model that a computer will understand. This is a very challenging task since biology
can be very complex. This problem of how to digitise phenotypic data such as behavior, electrocardiograms, and
crop health into a computer readable form offers exciting challenges for future bioinformaticians.
A bioinformaticist is an expert who not only knows how to use bioinformatics tools, but also knows how to write
interfaces for effective use of the tools. A bioinformatician, on the other hand, is a trained individual who only knows
to use bioinformatics tools without a deeper understanding.
3
Bioinformatics
Equally exciting is the potential for uncovering evolutionary relationships and patterns between different forms
of life. With the aid of nucleotide and protein sequences, it should be possible to find the ancestral ties between
different organisms. Thus far, experience has taught us that closely related organisms have similar sequences and
that more distantly related organisms have more dissimilar sequences. Proteins that show significant sequence
conservation, indicating a clear evolutionary relationship, are said to be from the same protein family. By studying
protein folds (distinct protein building blocks) and families, scientists are able to reconstruct the evolutionary
relationship between two species and to estimate the time of divergence between two organisms since they last
shared a common ancestor.
Since Mendel, bioinformatics and genetic record keeping have come a long way. The understanding of genetics has
advanced remarkably in the last thirty years. In 1972, Paul berg made the first recombinant DNA molecule using
ligase. In that same year, Stanley Cohen, Annie Chang and Herbert Boyer produced the first recombinant DNA
organism. In 1973, two important things happened in the field of genomics:
• Joseph Sambrook led a team that refined DNA electrophoresis technique using agarose gel, and
• Herbert Boyer and Stanely Cohen invented DNA cloning. By 1977, a method for sequencing DNA was discovered
and the first genetic engineering company, Genetech was founded.
By 1981, 579 human genes had been mapped and mapping by in situ hybridisation had become a standard method.
Marvin Carruthers and Leory Hood made a huge leap in bioinformatics when they invented a method for automated
DNA sequencing. In 1988, the Human Genome Organisation (HUGO) was founded. This is an international
organisation of scientists involved in Human Genome Project. In 1989, the first complete genome map was published
of the bacteria Haemophilus influenza. The following year, the Human Genome Project was started. By 1991, a total
of 1879 human genes had been mapped. In 1993, Genethon, a human genome research centre in France produced a
physical map of the human genome. Three years later, Genethon published the final version of the Human Genetic
Map. This concluded the end of the first phase of the Human Genome Project.
In the mid-1970s, it would take a laboratory at least two months to sequence 150 nucleotides. Ten years ago, the
only way to track genes was to scour large, well documented family trees of relatively inbred populations, such as
the Ashkenzai Jews from Europe. These types of genealogical searches 11 million nucleotides a day for its corpo-
rate clients and company research.
Bioinformatics was fuelled by the need to create huge databases, such as GenBank and EMBL and DNA Database of
Japan to store and compare the DNA sequence data erupting from the human genome and other genome sequencing
projects. Today, bioinformatics embraces protein structure analysis, gene and protein functional information, data
from patients, pre-clinical and clinical trials, and the metabolic pathways of numerous species.
Origin of internet
The management and, more importantly, accessibility of this data is directly attributable to the development of the
Internet, particularly the World Wide Web (WWW). Originally developed for military purposes in the 60’s and
expanded by the National Science Foundation in the 80’s, scientific use of the Internet grew dramatically following
the release of the WWW by CERN in 1992.
4
HTML
The WWW is a graphical interface based on hypertext by which text and graphics can be displayed and highlighted.
Each highlighted element is a pointer to another document or an element in another document, which can reside on
any internet host computer. Page display, hypertext links and other features are coded using a simple, cross-platform
HyperText Markup Language (HTML) and viewed on UNIX workstations, PCs and Apple Macs as WWW pages
using a browser.
Java
The first graphical WWW browser - Mosaic for X and the first molecular biology WWW server - ExPASy were made
available in 1993. In 1995, Sun Microsystems released Java, an object-oriented, portable programming language
based on C++. In addition to being a standalone programming language in the classic sense, Java provides a highly
interactive, dynamic content to the Internet and offers a uniform operational level for all types of computers, provided
they implement the ‘Java Virtual Machine’ (JVM). Thus, programs can be written, transmitted over the internet and
executed on any other type of remote machine running a JVM. Java is also integrated into Netscape and Microsoft
browsers, providing both the common interface and programming capability, which is vital in sorting through and
interpreting the gigabytes of bioinformatics data now available and increasing at an exponential rate.
XML
The new XML standard 8 is a project of the World-Wide Web Consortium (W3C), which extends the power of the
WWW to deliver not only HTML documents but an unlimited range of document types using customised markup.
This will enable the bioinformatics community to exchange data objects such as sequence alignments, chemical
structures, spectra and so on together with appropriate tools to display them, just as easily as they exchange HTML
documents today. Both Microsoft and Netscape support this new technology in their latest browsers.
CORBA
Another new technology, called CORBA, provides a way of bringing together many existing or ‘legacy’ tools
and databases with a common interface that can be used to drive them and access data. CORBA frameworks for
bioinformatics tools and databases have been developed by, for example, NetGenics and the European Bioinformatics
Institute (EBI).
Representatives from industry and the public sector under the umbrella of the Object Management Group are
working on open CORBA-based standards for biological information representation The Internet offers scientists a
universal platform on which to share and search for data and the tools to ease data searching, processing, integration
and interpretation. The same hardware and software tools are also used by companies and organisations in more
private yet still global Intranet networks. One such company, Oxford GlycoSciences in the UK, has developed a
bioinformatics system as a key part of its proteomics activity.
ROSETTA
ROSETTA focuses on protein expression data and sets out to identify the specific proteins, which are up or down-
regulated in a particular disease; characterise these proteins with respect to their primary structure, post-translational
modifications and biological function; evaluate them as drug targets and markers of disease; and develop novel drug
candidates. OGS uses a technique called fluorescent IPG-PAGE to separate and measure different protein types in
a biological sample such as a body fluid or purified cell extract. After separation, each protein is collected and then
broken up into many different fragments using controlled techniques. The mass and sequence of these fragments is
determined with great accuracy using a technique called mass spectrometry. The sequence of the original protein
can then be theoretically reconstructed by fitting these fragments back together in a kind of jigsaw. This reassembly
of the protein sequence is a task well-suited to signal processing and statistical methods.
ROSETTA is built on an object-relational database system, which stores demographic and clinical data on sample
donors and tracks the processing of samples and analytical results. It also interprets protein sequence data and matches
this data with that held in public, client and proprietary protein and gene databases. ROSETTA comprises a suite
of linked HTML pages, which allow data to be entered, modified and searched and allows the user easy access to
5
Bioinformatics
other databases. A high level of intelligence is provided through a sophisticated suite of proprietary search, analytical
and computational algorithms. These algorithms facilitate searching through the gigabytes of data generated by the
Company’s proteome projects, matching sequence data, carrying out de novo peptide sequencing and correlating
results with clinical data. These processing tools are mostly written in C, C++ or Java to run on a variety of computer
platforms and use the networking protocol of the internet, TCP/IP, to co-ordinate the activities of a wide range of
laboratory instrument computers, reliably identifying samples and collecting data for analysis.
The need to analyse ever increasing numbers of biological samples using increasingly complex analytical techniques
is insatiable. Searching for signals and trends in noisy data continues to be a challenging task, requiring great
computing power. Fortunately, this power is available with today’s computers, but of key importance is the integration
of analytical data, functional data and biostatistics. The protein expression data in ROSETTA forms only part of an
elaborate network of the type of data, which can now be brought to bear in biology. The need to integrate different
information systems into a collaborative network with a friendly face is bringing together an exciting mixture of
talents in the software world and has brought the new science of bioinformatics to life.
Data Bank followed in 1972 with a collection of ten X-ray crystallographic protein structures, and the SWISSPROT
protein sequence database began in 1987.A huge variety of divergent data resources of different type sand sizes
are now available either in the public domain or more recently from commercial third parties. All of the original
databases were organised in a very simple way with data entries being stored in flat files, either one per entry, or
as a single large text file. Re-write - Later on lookup indexes were added to allow convenient keyword searching
of header information.
Origin of tools
After the formation of the databases, tools became available to search sequence databases - at first in a very simple
way, looking for keyword matches and short sequence words, and then more sophisticated pattern matching and
alignment based methods. The rapid but less rigorous BLAST algorithm has been the mainstay of sequence database
searching since its introduction a decade ago, complemented by the more rigorous and slower FASTA and Smith
Waterman algorithms. Suites of analysis algorithms, written by leading academic researchers at Stanford, CA,
Cambridge, UK and Madison, WI for their in-house projects, began to become more widely available for basic
sequence analysis. These algorithms were typically single function black boxes that took input and produced output
in the form of formatted files. UNIX style commands were used to operate the algorithms, with some suites having
hundreds of possible commands, each taking different command options and input formats. Since these early efforts,
significant advances have been made in automating the collection of sequence information.
Rapid innovation in biochemistry and instrumentation has brought us to the point where the entire genomic
sequence of at least 20 organisms, mainly microbial pathogens, are known and projects to elucidate at least 100
more prokaryotic and eukaryotic genomes are currently under way. Groups are now even competing to finish the
sequence of the entire human genome. With new technologies we can directly examine the changes in expression
levels of both mRNA and proteins in living cells, both in a disease state or following an external challenge. We can
go on to identify patterns of response in cells that lead us to an understanding of the mechanism of action of an agent
on a tissue. The volume of data arising from projects of this nature is unprecedented in the pharmaceutical industry,
and will have a profound effect on the ways in which data are used and experiments performed in drug discovery
and development projects. This is true not least because, with much of the available interesting data being in the
hands of commercial genomics companies, pharmacies are unable to get exclusive access to many gene sequences
or their expression profiles.
6
The competition between co-licensees of a genomic database is effectively a race to establish a mechanistic role
or other utility for a gene in a disease state in order to secure a patent position on that gene. Much of this work is
carried out by informatics tools. Despite the huge progress in sequencing and expression analysis technologies, and
the corresponding magnitude of more data that is held in the public, private and commercial databases, the tools
used for storage, retrieval, analysis and dissemination of data in bioinformatics are still very similar to the original
systems gathered together by researchers 15-20 years ago.
Many are simple extensions of the original academic systems, which have served the needs of both academic and
commercial users for many years. These systems are now beginning to fall behind as they struggle to keep up
with the pace of change in the pharmaceutical industry. Databases are still gathered, organised, disseminated and
searched using flat files. Relational databases are still few and far between, and object-relational or fully object
oriented systems are rarer still in mainstream applications. Interfaces still rely on command lines, fat client interfaces,
which must be installed on every desktop, or HTML/CGI forms. Whilst they were in the hands of bioinformatics
specialists, pharmacies have been relatively undemanding of their tools. Now the problems have expanded to cover
the mainstream discovery process, much more flexible and scalable solutions are needed to serve pharmaceutical
R&D informatics requirements.
There are different views of origin of Bioinformatics- From T K Attwood and D J ParrySmith’s “Introduction to
Bioinformatics”, Prentice-Hall 1999 [Longman Higher Education; ISBN 0582327881]: “The term bioinformatics
is used to encompass almost all computer applications in biological sciences, but was originally coined in the
mid-1980s for the analysis of biological sequence data.” From Mark S. Boguski’s article in the “Trends Guide to
Bioinformatics” Elsevier, Trends Supplement 1998 p1: “The term “bioinformatics” is a relatively recent invention,
not appearing in the literature until 1991 and then only in the context of the emergence of electronic publishing. The
National Center for Biotechnology Information (NCBI), is celebrating its 10th anniversary this year, having been
written into existence by US Congressman Claude Pepper and President Ronald Reagan in 1988. So, bioinformatics
has, in fact, been in existence for more than 30 years and is now middle-aged.
Sequence generation, and its subsequent storage, interpretation and analysis are entirely computer dependent tasks.
However, the molecular biology of an organism is a very complex issue with research being carried out at different levels
including the genome, proteome, transcriptome and metabalome levels. Following on from the explosion in volume of
genomic data, similar increase in data have been observed in the fields of proteomics, transcriptomics and metabalomics.
The first challenge facing the bioinformatics community today is the intelligent and efficient storage of this mass of
data. It is then their responsibility to provide easy and reliable access to this data. The data itself is meaningless before
analysis and the sheer volume present makes it impossible for even a trained biologist to begin to interpret it manually.
Therefore, incisive computer tools must be developed to allow the extraction of meaningful biological information.
There are three central biological processes around, which bioinformatics tools must be developed:
• DNA sequence determines protein sequence
• Protein sequence determines protein structure
• Protein structure determines protein function
The integration of information learned about these key biological processes should allow us to achieve the long
term goal of the complete understanding of the biology of organisms.
7
Bioinformatics
Due to the spectacular growth of biotechnology and molecular biology tremendous amount of data of nucleotide
sequences or protein sequences are being produced. Here, comes the role of bioinformatics:
• To uncover the wealth of biological information hidden in the mass of nucleotide sequence.
• Knowing the amino acid sequence on the basis of nucleotide sequences.
• Knowing structure of proteins on the basis of amino acid sequences.
• Prediction of functional aspects of proteins on the basis of its structure.
Therefore, it is clear that the knowledge of bioinformatics not merely limited to the computation of data, but in
reality it can be used to solve many biological problems and can be applied how living things work.
The major applications of bioinformatics being to access, search, visualise and retrieve the information of databases
of the sequences as well as to understand structural information of biomolecules proteome analysis and so on. Other
applications include cell metabolism, biodiversity, downstream processing in chemical engineering, drug and vaccine
design. These are the areas in which bioinformatics is an integral component. Current efforts in molecular biology
(example, genome projects) are producing a large quantity of data that is not only providing exciting opportunities
for knowledge discovery, but also increasing problem of information overload. Bioinformatics also concerns the
development of new tools for the analysis of genomic and molecular biological data. This can be applied to all fields
of biological science as agricultural science, environmental science, pharmaceutical science, chemical science and
medical science.
8
Fig. 1.1 Genes encode the recipes for proteins
(Source: [Link]
9
1.10 Bioinformatics Applications
Molecular medicine
The human genome will have profound effects on the fields of biomedical research and clinical medicine. Every
disease has a genetic component. This may be inherited (as is the case with an estimated 3000-4000 hereditary
disease including Cystic Fibrosis and Huntingtons disease) or a result of the body’s response to an environmental
stress which causes alterations in the genome (example, cancers, heart disease, diabetes.).
The completion of the human genome means that we can search for the genes directly associated with different
diseases and begin to understand the molecular basis of these diseases more clearly. This new knowledge of the
molecular mechanisms of disease will enable better treatments, cures and even preventative tests to be developed.
Personalised medicine
Clinical medicine will become more personalised with the development of the field of pharmacogenomics. This is
the study of how an individual’s genetic inheritance affects the body’s response to drugs. At present, some drugs fail
to make it to the market because a small percentage of the clinical patient population show adverse affects to a drug
due to sequence variants in their DNA. As a result, potentially life saving drugs never makes it to the marketplace.
Today, doctors have to use trial and error to find the best drug to treat a particular patient as those with the same
clinical symptoms can show a wide range of responses to the same treatment. In the future, doctors will be able to
analyse a patient’s genetic profile and prescribe the best available drug therapy and dosage from the beginning.
Preventative medicine
With the specific details of the genetic mechanisms of diseases being unraveled, the development of diagnostic tests
to measure a person’s susceptibility to different diseases may become a distinct reality. Preventative actions such
as change of lifestyle or having treatment at the earliest possible stages when they are more likely to be successful,
could result in huge advances in our struggle to conquer disease.
Gene therapy
In the not too distant future, the potential for using genes themselves to treat disease may become a reality. Gene
therapy is the approach used to treat, cure or even prevent disease by changing the expression of a person’s genes.
Currently, this field is in its infantile stage with clinical trials for many different types of cancer and other diseases
ongoing.
Drug development
At present all drugs on the market target only about 500 proteins. With an improved understanding of disease
mechanisms and using computational tools to identify and validate new drug targets, more specific medicines that
act on the cause, not merely the symptoms, of the disease can be developed. These highly specific drugs promise
to have fewer side effects than many of today’s medicines.
Department of Energy (DOE) initiated the MGP (Microbial Genome Project) to sequence genomes of bacteria
useful in energy production, environmental cleanup, industrial processing and toxic waste reduction. By studying
the genetic material of these organisms, scientists can begin to understand these microbes at a very fundamental
level and isolate the genes that give them their unique abilities to survive under extreme conditions.
13
Bioinformatics
Waste cleanup
The world’s toughest bacterium is the most radiation resistant organism known. Scientists are interested in this
organism because of its potential usefulness in cleaning up waste sites that contain radiation and toxic chemicals.
(Department of Energy, USA) launched a program to decrease atmospheric carbon dioxide levels. One method of
doing so is to study the genomes of microbes that use carbon dioxide as their sole carbon source.
Biotechnology
Some archaeon and the bacterium have potential for practical applications in industry and government-funded
environmental remediation. These microorganisms thrive in water temperatures above the boiling point and therefore,
may provide the DOE, the Department of Defence, and private companies with heat-stable enzymes suitable for
use in industrial processes.
Other industrially useful microbes are of high industrial interest as a research object because it is used by the chemical
industry for the biotechnological production of the amino acid lysine. The substance is employed as a source of protein
in animal nutrition. Lysine is one of the essential amino acids in animal nutrition. Biotechnologically produced lysine
is added to feed concentrates as a source of protein, and is an alternative to soybeans or meat and bonemeal.
Micro-organisms are useful in the dairy industry, for manufacturing dairy products like buttermilk, yogurt and
cheese. They are also used to prepare pickled vegetables, beer, wine, some bread and sausages and other fermented
foods. Researchers anticipate that understanding the physiology and genetic make-up of this bacterium will prove
invaluable for food manufacturers as well as the pharmaceutical industry as a vehicle for delivering drugs.
Antibiotic resistance
Scientists have been examining the genome of a bacterium. They have discovered a virulence region made up of a
number of antibiotic-resistant genes that may contribute to the bacterium’s transformation from harmless gut bacteria
to a menacing invader. The discovery of the region, known as a pathogenicity island, could provide useful markers
for detecting pathogenic strains and help to establish controls to prevent the spread of infection in wards.
Evolutionary studies
The sequencing of genomes from all three domains of life, eukaryota, bacteria and archaea means that evolutionary
studies can be performed in a quest to determine the tree of life and the last universal common ancestor.
14
Crop improvement
Comparative genetics of the plant genomes has shown that the organisation of their genes has remained more
conserved over evolutionary time than was previously believed. These findings suggest that information obtained
from the model crop systems can be used to suggest improvements to other food crops. At present the complete
genomes of water cress and rice are available.
Insect resistance
Genes from Bacillus thuringiensis that can control a number of serious pests have been successfully transferred to
cotton, maize and potatoes. This new ability of the plants to resist insect attack means that the amount of insecticides
being used can be reduced and hence the nutritional quality of the crops is increased.
Vetinary science
Sequencing projects of many farm animals including cows, pigs and sheep are now well under way in the hope that
a better understanding of the biology of these organisms will have huge impacts for improving the production and
health of livestock and ultimately have benefits for human nutrition.
Comparative studies
Analysing and comparing the genetic material of different species is an important method for studying the functions
of genes, the mechanisms of inherited diseases and species evolution. Bioinformatics tools can be used to make
comparisons between the numbers, locations and biochemical functions of genes in different organisms. Organisms
that are suitable for use in experimental research are termed model organisms.
They have a number of properties that make them ideal for research purposes including short life spans, rapid
reproduction, being easy to handle, inexpensive and they can be manipulated at the genetic level.
An example of a human model organism is the mouse. Mouse and human are very closely related (>98%) and for
the most part we see a one to one correspondence between genes in the two species. Manipulation of the mouse at
the molecular level and genome comparisons between the two species can and is revealing detailed information on
the functions of human genes, the evolutionary relationship between the two species and the molecular mechanisms
of many human diseases.
15
Bioinformatics
Chapter II
Biological Databases
Aim
The aim of this chapter is to:
Objectives
The objectives of this chapter are to:
Learning outcome
At the end of this chapter, you will be able to:
20
2.1 Introduction
The modern genomic research leads to the generation of huge amounts of raw sequence data. Sophisticated
computational methodologies are required to manage the mass of data as the volume of genomic data grows. The
challenge in the genomics era is to store and handle the volume of information through the establishment and use
of computer databases. Thus, the development of databases to handle the vast amount of molecular biological data
is a fundamental task of bioinformatics.
A biological database is a large, organised body of persistent data, usually associated with computerised software
designed to update, query and retrieve components of the data stored within system. A simple database might
be a single file containing many records, each including the same set of information. A record associated with a
nucleotide sequence database contains information such as contact name, input sequence with a description of the
type of molecule, scientific name of the source organism from which it was isolated and literature citations associated
with sequence.
For researchers to benefit from the data stored in a database, two additional requirements must be met:
• Easy access to the information
• A method for extracting only that information needed to answer a specific biological question
Currently, a lot of bioinformatics work is concerned with the technology of databases. These databases include both
‘public’ repositories of gene data such as GenBank or Protein.
DataBank (PDB), and private databases such as those used by research groups involved in gene mapping projects or
those held by biotech companies. Making such databases accessible through open standards (such as Web) is very
important since consumers of bioinformatics data use a range of computer platforms, from the more powerful and
forbidding UNIX boxes favoured by the developers and curators to the far friendlier Macs often found populating
the labs of computer-wary biologists.
RNA and DNA are the proteins that store hereditary information about an organism. These macromolecules have a
fixed structure analysed by biologists with the help of bioinformatic tools and databases. A few popular databases
are GenBank from NCBI (National Center for Biotechnology Information), SwissProt from Swiss Institute of
Bioinformatics and PIR from Protein Information Resource.
21
Bioinformatics
Depending on the research project, biological data comes in many different flavours. Most researches work with
a number of different formats even though they may not realise this at first hand. Some of the data, which can be
found when researching any biological question, is briefly listed below. Some of these data in the databases are
partly overlapping and referring to each other.
• Text: Examples of text databases are PubMed and OMIM containing textual information and references related
to biological data.
• Sequence data: GenBank and UniProt exemplify biological databases containing DNA and protein sequences,
respectively.
• Protein structure: You can also find databases specifically related to protein structure files (for example, the
PDB, SCOP and CATH databases).
• Links: Most databases contain information on sequence data within a specific field or subject. A different type
of database is for example, the InterPro database consisting of a collection of links from protein domains and
families to other databases providing related resources.
• Images: In the field of 2D gel and microscopic images you can also find various databases containing data, for
example, identified on reference gel images.
• Numerical data: Gene expression data as well as other microarray data are also accessible from a number of
databases. An example is the ArrayExpress database of the European Bioinformatics Institute, EBI.
• Biological matter: Frozen bacterial strains, vectors and so on are also to be found in databases collecting
information on each of these specific biological matters, for example, UniVec database hosted by NCBI.
EMBL and GenBank are the two major nucleotide databases. EMBL is the European version and GenBank is the
American. EMBL and GenBank collaborate and synchronise their databases so that the databases will contain same
information. The rate of growth of DNA databases has been following an exponential trend, with a doubling time
now estimated to be 9 to12 months. In January 1998, EMBL contained more than a million entries, representing more
than 15,500 species, although most data is from model organisms. These databases are updated on a daily basis.
22
The principal requirements on the public data services are:
• Data quality: Data quality has to be of the highest priority. However, because the data services in most cases
lack access to supporting data, the quality of the data must remain the primary responsibility of the submitter.
• Supporting data: Database users will need to examine the primary experimental data, either in the database
itself, or by following cross-references back to network accessible laboratory databases.
• Deep annotation: Deep, consistent annotation comprising supporting and ancillary information should be attached
to each basic data object in the database.
• Timeliness: The basic data should be available on an Internet-accessible server within days (or hours) of
publication or submission.
• Integration: Each data object in the database should be cross-referenced to representation of the same or related
biological entities in other databases. Data services should provide capabilities for following these links from
one database or data service to another.
Biological databases can be broadly classified into sequence and structure databases:
Sequence databases
With the current speed of sequencing projects, to store and organise sequence data a lot of work is needed. Most
sequence databases store additional information along with the sequence. This could be references to the original
research papers stored in PubMed, information about annotated regions, regions were conflicting residues have been
published, information on species and much more. So far, a common standard for handling all of this information
has not been created. Thus, every database has its own standard on how to store the data. However, most data is
stored in a plain text format (flat file) and can thus, be opened in standard software such as Word, Notepad and so
on. However, large amounts of plain text may not be easy comprehensible. Another problem by storing in a flat
file format is the size of the database. Databases with long sequence entries may become too large to handle on a
normal PC for most users.
Sequence databases are applicable to both nucleic acid sequences and protein sequences, whereas structure database
is applicable to only proteins. The first database was created within a short period after the insulin protein sequence
was made available in 1956. Incidentally, Insulin is the first protein to be sequenced. The sequence of Insulin
consisted of just 51 residues, which characterise the sequence. An alternative approach used by most websites with
large databases is to store all the information in a relational database.
Relational databases have connections or pointers to additional data in other databases or tables. Thus, one can
easily and very fast retrieve a large amount of information on one particular sequence.
One of the characteristics of these databases is that they are maintained and kept up to date on a regular basis. The
four major sequence databases are:
• GenBank: A US-based comprehensive collection of various biological data.
• EMBL: The main European Resource of Nucleotide Sequence Data.
• DDBJ: The DNA Data Bank of Japan.
• UniProt: The universal protein resource.
23
Bioinformatics
GenBank at NCBI
GenBank (Genetic Sequence Databank) is one of the fastest growing repositories of known genetic sequences. It
has a flat file structure, which is an ASCII text file, readable by both humans and computers. In addition to sequence
data, GenBank files contain information like accession numbers and gene names, phylogenetic classification and
references to published literature. There are approximately 191,400,000 bases and 183,000 sequences as of June
1994.
The National Institute of Health hosted at [Link] has achieved a strong position in collecting
biological data of almost any kind. In addition to storing sequence data, NCBI stores almost all kinds of biological
sequence related data. PubMed is probably the mostly used service that NCBI offers to theirs users together with
BLAST, an option for searching for homologous sequences in the entire database. The NCBI staff provides software
tools for handling sequence data.
Bases in GenBank
90
Billions
80
70
60
50
40
30
20
10
0
Jan- Jan- Jan- Jan- Jan- Jan- Jan- Jan- Jan- Jan- Jan- Jan- Jan-
83 85 87 89 91 93 95 97 99 01 03 05 07
24
Fig. 2.2 GenBank file format
(Source: [Link]/teaching/c40/ppt/intro_databases.ppt)
25
Bioinformatics
EMBL
EMBL Nucleotide Sequence Database is a comprehensive database of DNA and RNA sequences collected from the
scientific literature and patent applications and directly submitted from researchers and sequencing groups. Data
collection is done in collaboration with GenBank (USA) and the DNA Database of Japan (DDBJ). The database
currently doubles in size every 18 months and currently (June 1994) contains nearly 2 million bases from 182,615
sequence entries.
The EMBL Nucleotide Sequence Database is hosted by EBI - the European Bioinformatics Institute, at the European
Molecular Biology Laboratory (EMBL), hosted at [Link] DNA and RNA sequences are directly
submitted to the EMBL nucleotide sequence database by individual researchers wand also by genome sequencing
projects and patent applications, and the database is produced and maintained collaborating with both GenBank and
the DNA Data Bank of Japan (DDBJ). The international collection of sequence data is exchanged between EMBL,
GenBank and DDBJ on a daily basis and knowledge of global sequence information can be retrieved from any of
the three entries.
UniProt
UniProt is the universal protein resource, and as stated on its website the database intends to be both comprehensive and
of high quality. At the UniProt website, [Link] data has been divided into three classifications:
• Core data
• Supporting data
• Information
You can search information of your sequence in these three categories. You can also do BLAST searches and
create alignments; a couple of other services are also provided. The UniProt Knowledgebase, UniProtKB, contains
translations of the coding sequences submitted to EMBL, GenBank and DDBJ and the UniProtKB contains all
publicly available protein sequences.
26
EMBL GenBank
Europe USA
Japan
NIG
CIB
SwissProt
This is a protein sequence database that provides a high level of integration with other databases and also has a very
low level of redundancy (means less identical sequences are present in the database).
PubMed
PubMed gives biological data in text format and this service provided by the U.S. National Library of Medicine
links to more than 17 million resources from different journals within the field of life science. A relatively new
functionality at NCBI website is possibility to sign up for an account at My NCBI, which is a service, offering a
customised and automated PubMed update. After registration at My NCBI, you can save your searches and set up
automated searches alerting you by e-mail. You can also customise, for example, filtering options on the searches.
PubMed can be accessed at [Link]
Ensembl
Ensembl is a project developing software for automatic annotation of eukaryotic genomes.
EMBL - EBI and the Sanger Institute are behind the project and at the website, http:
//[Link]/[Link], you can search within all the data from the Ensemble project divided into
species.
EBI
One of the larger European bioinformatic centers, European Bioinformatics Institute (EBI), hosts a number of
databases and a lot of methods to help analyse all this data. The EBI website [Link] also stores the
EMBL nucleotide sequence database ([Link]
27
Bioinformatics
InterPro
Most of the databases mentioned above also provide links to related information in other databases. InterPro database
try to a larger extent to link from one protein domain or family to a number of different databases, which individually
contains a lot of relevant information. InterPro database does not contain any sequence information but is largely a
mesh of hyperlinks to various other resources. The link to InterPro is [Link]
Pfam
A very useful database for finding protein domains is the Pfam database. Pfam currently stores information on more
than 9000 protein families. When working on an unknown protein, it is often very valuable to retrieve information of
the actual protein family as identification of functional domains within a protein sequence can benefit your knowledge
about the role and function of the protein. You can access the Pfam database at [Link]
Structure databases
Information about protein structure is not developing as fast as sequence data information due to slower pace in
solving 3 D structures of proteins. The RCSB Protein Data Bank hosted at [Link] holds slightly more
than 48000 structures. At the website, you can download structure files and you are provided a number of tools for
structure studies.
• SCOP: Structural Classification of Proteins is accessible at [Link] The SCOP database
describes structural and evolutionary relationships between all known protein structures and also provides a
number of links to other on-line resources related to protein structure and to sequence databases in general.
• CATH Protein Structure Classification. The CATH database hosted at [Link] classifies
protein structures from the PDB according to a four-level hierarchy.
Species-specific databases
A number of species-specific databases usually hold very detailed information about only one particular species.
During this period, 3D structure of proteins were studied and well known PDB was developed as the first protein
structure database with only 10 entries in 1972, which has now grown into a large database with over 10,000 entries.
While the initial databases of protein sequences were maintained at the individual laboratories, the development
of a consolidated formal database known as SWISS-PROT protein sequence database was initiated in 1986, which
recently has about 70,000 protein sequences from more than 5000 model organisms, a small fraction of all known
organisms. These huge varieties of data resources are now available for study and research by both academic
institutions and industries. These are made available as public domain information in the larger interest of research
community through Internet ([Link]) and CD-ROMs (on request from [Link]).
Databases can be classified into:
Primary databases
A primary database contains information of the sequence or structure alone. Examples of these include Swiss-Prot and
PIR for protein sequences, GenBank and DDBJ for Genome sequences and Protein Databank for protein structures.
Biological databases are archives of consistent data stored in an efficient manner. These databases contain data from
a wide spectrum of molecular biology areas. Primary or archived databases contain information and annotation of
DNA and protein sequences, DNA and protein structures and DNA and protein expression profiles.
Secondary databases
A secondary database contains derived information from the primary database. A secondary sequence database
contains information such as the conserved sequence, signature sequence and active site residues of the protein
families arrived by multiple sequence alignment of a set of related proteins.
28
A secondary structure database contains entries of the PDB in an organised way. These contain entries classified
according to their structure such as all alpha proteins, all beta proteins, and so on. These also contain information
on conserved secondary structure motifs of a particular protein. Some of the secondary database created and hosted
by various researchers at their individual laboratories includes SCOP, developed at Cambridge University; CATH
developed at University College of London, PROSITE of Swiss Institute of Bioinformatics, eMOTIF at Stanford.
Secondary or derived databases contain the results of analysis on the primary resources including information on
sequence patterns or motifs, variants and mutations and evolutionary relationships. Information from the literature
is contained in bibliographic databases (Medline). These databases are easily accessible and that an intuitive query
system is provided to allow researchers to obtain very specific information on a particular biological subject. The
data should be provided in a clear, consistent manner with some visualisation tools for biological interpretation.
Specialist databases for particular subjects have been set-up, for example EMBL database for nucleotide sequence
data, UniProtKB/Swiss-Prot protein database and PDB (a 3D protein structure database).
Scientists also need to be able to integrate the information obtained from the underlying heterogeneous databases in
a sensible manner for having an overview of their biological subject. Sequence Retrieval System (SRS) is a power-
ful, querying tool provided by EBI that links information from more than 150 heterogeneous resources.
Composite databases
Composite database amalgamates different primary database sources, which obviates the need to search multiple
resources. Different composite database use different primary database and different criteria in their search algorithm.
Various options for search have also been incorporated in the composite database. NCBI hosts these nucleotide and
protein databases in their large high available redundant array of computer servers. NCBI provides free access to
various persons involved in research.
This also has link to OMIM (Online Mendelian Inheritance in Man), which contains information about the proteins
involved in genetic diseases. The growth of the primary databases gave rise to questions on the format of sequences,
reliability and comprehensiveness of databases. To address the format issues, in-house software solutions have been
developed to convert format of one database to another. A public domain software (FORCON) can also be used.
The newer software tools are used for analysis to accept data in multiple formats.
The problem in the data reliability is the possibility of misannotations. The misannotations are some time introduced
due to the process of automation of annotation process carried out with computers. A misannotation (if introduced)
multiplies in subsequent additions and may accumulate to an unbelievable extent and create confusion. A possible
solution to prevent this from happening is to flag the protein sequence, which has been annotated by sequence
comparison but whose function has not been validated by experimental methods.
While most biological databases contain nucleotide and protein sequence information, there are also databases, which
include taxonomic information such as the structural and biochemical characteristics of organisms. However, the
power and ease of using sequence information has made it the method of choice in modern analysis.
Contributions from the fields of biology and chemistry have facilitated an increase in the speed of sequencing
genes and proteins. The advent of cloning technology allowed foreign DNA sequences to be easily introduced into
bacteria. In this way, rapid mass production of particular DNA sequences became possible. Oligonucleotide synthesis
provided researchers with the ability to construct short fragments of DNA with sequences of their own choosing.
These oligonucleotides could then be used in probing vast libraries of DNA to extract genes containing that sequence.
29
Bioinformatics
Alternatively, these DNA fragments could also be used in polymerase chain reactions to amplify existing DNA
sequences or to modify these sequences. With these techniques in place, progress in biological research increased
exponentially. However, for researchers to benefit from all this information, two additional things were required:
• Ready access to the collected pool of sequence information and
• A way to extract from this pool only those sequences of interest to a given researcher
Collecting all necessary sequence information of interest to a given project from published journal articles quickly
became a formidable task. After collection, the organisation and analysis of this data still remained. It could take
weeks to months for a researcher to search sequences by hand in order to find related genes or proteins. Computer
technology has provided the obvious solution to this problem. Not only can computers be used to store and organise
sequence information into databases, but they can also be used to analyse sequence data rapidly.
The evolution of computing power and storage capacity has, so far, been able to outpace the increase in sequence
information being created. Theoretical scientists have derived new and sophisticated algorithms, which allow
sequences to be readily compared using probability theories. These comparisons become the basis for determining
gene function, developing phylogenetic relationships and simulating protein models.
The physical linking of a vast array of computers in the 1970’s provided a few biologists with ready access to the
expanding pool of sequence information. This web of connections, now known as the Internet, has evolved and
expanded so that nearly everyone has access to this information and the tools necessary to analyse it. Databases
of existing sequencing data can be used to identify homologues of new molecules that have been amplified and
sequenced in the lab. The property of sharing a common ancestor, homology, can be a very powerful indicator in
bioinformatics.
Analysis of data
Both types of sequence can then be analysed in many ways with bioinformatics tools. They can be assembled. Note
that this is one of the occasions when the meaning of a biological term differs markedly from a computational one.
Computer scientists, banish from your mind any thought of assembly language. Sequencing can only be performed
for relatively short stretches of a biomolecule and finished sequences are therefore, prepared by arranging overlapping
‘reads’ of monomers (single beads on a molecular chain) into a single continuous passage of ‘code’.
This is the bioinformatic sense of assembly. They can be mapped (that is, their sequences can be parsed to find sites
where so-called ‘restriction enzymes’ will cut them). They can be compared, usually by aligning corresponding
segments and looking for matching and mismatching letters in their sequences. Genes or proteins, which are
sufficiently similar are likely to be related and are therefore, said to be ‘homologous’ to each other the whole truth
is rather more complicated than this. Such cousins are called ‘homologues’. If a homologue (a related molecule)
exists then a newly discovered protein may be modelled that is the three dimensional structure of the gene product
can be predicted without doing laboratory experiments.
Bioinformatics is used in primer design. Primers are short sequences needed to make many copies of (amplify) a
piece of DNA as used in PCR (the Polymerase Chain Reaction). Bioinformatics is used to attempt to predict the
function of actual gene products. Information about the similarity, and, by implication, the relatedness of proteins
is used to trace the family trees’ of different molecules through evolutionary time.
30
There are various other applications of computer analysis to sequence data, but, with so much raw data being
generated by the Human Genome Project and other initiatives in biology, computers are presently essential for
many biologists just to manage their day-to-day results Molecular modelling/structural biology is a growing field,
which can be considered part of bioinformatics. There are, for example, tools which allow (often via the Net) to
make pretty good predictions of the secondary structure of proteins arising from a given amino acid sequence, often
based on known ‘solved’ structures and other sequenced molecules acquired by structural biologists. Structural
biologists use ‘bioinformatics’ to handle the vast and complex data from X-ray crystallography, nuclear magnetic
resonance (NMR) and electron microscopy investigations and create the 3-D models of molecules that seem to be
everywhere in the media.
Factors that must be taken into consideration when designing these tools are:
• The end user (the biologist) may not be a frequent user of computer technology
• These software tools must be made available over the internet given the global distribution of the scientific
research community
Structural analysis
These sets of tools allow you to compare structures with the known structure databases. The function of a protein
is more directly a consequence of its structure rather than its sequence with structural homologs tending to share
functions. The determination of a protein’s 2D/3D structure is crucial in the study of its function.
31
Bioinformatics
BLAST
Sequence data are compared with one another using the Basic Local Alignment Search Tool or BLAST (Altschul et
al., 1990). This algorithm attempts to find ‘‘high-scoring segment pairs’’ (HSPs), which are pairs of sequences that
can be aligned with one another and, when aligned, meet certain scoring and statistical criteria.
BLAST is a set of similarity search programs designed to explore all of the available sequence databases regardless
of whether the query is protein or DNA. The BLAST programs have been designed for speed, with a minimal
sacrifice of sensitivity. The scores assigned in a BLAST search have a well-defined statistical interpretation, making
real matches easier to distinguish from random background hits. BLAST uses a heuristic algorithm, which seeks
local as opposed to global alignments and is therefore able to detect relationships among sequences which share
only isolated regions of similarity. This is a primary criterion in sequence analysis. Other tool available includes
CLUSTALW for multiple sequence alignment.
BLAST (Basic Local Alignment Search Tool) comes under the category of homology and similarity tools. It is a set
of search programs designed for the Windows platform and is used to perform fast similarity searches regardless of
whether the query is for protein or DNA. Comparison of nucleotide sequences in a database can be performed. Also
a protein database can be searched to find a match against the queried protein sequence. NCBI has also introduced
the new queuing system to BLAST (Q BLAST) that allows users to retrieve results at their convenience and format
their results multiple times with different formatting options.
32
Nucleotide sequence databases
• nr: All GenBank+EMBL+DDBJ+PDB sequences (but no EST, STS, GSS, or phase 0, 1 or 2 HTGS sequences).
No longer ‘non-redundant’.
• month : All new or revised GenBank+EMBL+DDBJ+PDB sequences released in the last 30 days.
• Drosophila genome : Drosophila genome provided by Celera and Berkeley Drosophila Genome Project )
• dbest: Database of GenBank+EMBL+DDBJ sequences from EST Divisions
• dbsts: Database of GenBank+EMBL+DDBJ sequences from STS Divisions
• htgs: Unfinished High Throughput Genomic Sequences: phases 0, 1 and 2
• gss: Genome Survey Sequence, includes single-pass genomic data, exon-trapped sequences, and Alu PCR
sequences.
• Yeast: Yeast (Saccharomyces cerevisiae) genomic nucleotide sequences
• E. coli : Escherichia coli genomic nucleotide sequences
• pdb :Sequences derived from the 3-dimensional structure from Brookhaven Protein Data Bank
• kabat: Kabat’s database of sequences of immunological interest
• vector : Vector subset of GenBank(R), NCBI, in [Link]
• mito : Database of mitochondrial sequences
• alu: Select Alu repeats from REPBASE, suitable for masking Alu repeats from query sequences. It is available
by anonymous FTP from [Link] (under the /pub/jmc/alu directory).
• Epd: Eukaryotic Promotor Database found on the web at [Link]
33
Bioinformatics
FASTA
FASTA is an alignment program for protein sequences created by Pearsin and Lipman in 1988. The program is one of
the many heuristic algorithms proposed to speed up sequence comparison. The basic idea is to add a fast pre screen
step to locate the highly matching segments between two sequences, and then extend these matching segments to
local alignments using more rigorous algorithms such as Smith-Waterman.
EMBOSS
EMBOSS (European Molecular Biology Open Software Suite) is a software-analysis package. It can work with
data in a range of formats and also retrieve sequence data transparently from the Web. Extensive libraries are also
provided with this package, allowing other scientists to release their software as open source. It provides a set of
sequence-analysis programs, and also supports all UNIX platforms.
Clustalw
It is a fully automated sequence alignment tool for DNA and protein sequences. It returns the best match over a total
length of input sequences, be it a protein or a nucleic acid.
34
RasMol
It is a powerful research tool to display the structure of DNA, proteins, and smaller molecules. Protein Explorer, a
derivative of RasMol, is an easier to use program.
PROSPECT
PROSPECT (Protein Structure Prediction and Evaluation Computer ToolKit) is a protein structure prediction system
that employs a computational technique called protein threading to construct a protein’s 3-D model.
PatternHunter
PatternHunter, based on Java, can identify all approximate repeats in a complete genome in a short time using little
memory on a desktop computer. Its features are its advanced patented algorithm and data structures, and the java
language used to create it. The Java language version of PatternHunter is just 40 KB, only 1% the size of Blast,
while offering a large portion of its functionality.
COPIA
COPIA (Consensus Pattern Identification and Analysis) is a protein structure analysis tool for discovering motifs
(conserved regions) in a family of protein sequences. Such motifs can be then used to determine membership to
the family for new protein sequences, predict secondary and tertiary structure and function of proteins and study
evolution history of the sequences.
JAVA in bioinformatics
Since research centers are scattered all around the globe ranging from private to academic settings, and a range of
hardware and OSs are being used, Java is emerging as a key player in bioinformatics. Physiome Sciences’ computer-
based biological simulation technologies and Bioinformatics Solutions’ PatternHunter are two examples of the
growing adoption of Java in bioinformatics.
Perl in bioinformatics
String manipulation, regular expression matching, file parsing, data format inter-conversion and so on are the common
text-processing tasks performed in bioinformatics. Perl excels in such tasks and is being used by many developers.
Yet, there are no standard modules designed in Perl specifically for the field of bioinformatics. However, developers
have designed several of their own individual modules for the purpose, which have become quite popular and are
coordinated by the BioPerl project.
35
Chapter IV
Sequence Alignment
Aim
The aim of this chapter is to:
Objectives
The objectives of this chapter are to:
Learning outcome
At the end of this chapter, you will be able to:
• identity matrix
51
Bioinformatics
4.1 Introduction
Once a genome is completely sequenced, there are sorts of analyses performed on it. Some of the goals of sequence
analysis are the following:
• Identify the genes.
• Determine the function of each gene. One way to hypothesise the function is to find another gene (possibly
from another organism) whose function is known and to which the new gene has high sequence similarity. This
assumes that sequence similarity implies functional similarity, which may or may not be true.
• Identify the proteins involved in the regulation of gene expression.
• Identify sequence repeats.
• Identify other functional regions.
Many of these tasks are computational in nature. Given the incredible rate at which sequence data is being produced,
the integration of computer science, mathematics, and biology will be integral to analysing those sequences.
Sequence alignment in bioinformatics is a field of research focused on developing tools for comparing and finding
similar sequences of amino acids or DNA base pairs with the aid of computers. The sequence similarity is used to
assess gene and protein homology, classify genes and proteins, predict biological function, secondary and tertiary
protein structure, detect point mutations, construct evolutionary trees, and so on. There are two main areas of
sequence alignment: pairwise sequence alignment and multiple sequence alignment.
Sequence alignment is an arrangement of two or more sequences, highlighting their similarity. The sequences are
padded with gaps (dashes) so that wherever possible, columns contain identical characters from the sequences
involved.
tcctctgcctctgccatcat---caaccccaaagt
|||| ||| ||||| ||||| ||||||||||||
tcctgtgcatctgcaatcatgggcaaccccaaagt
The problem has tractable solutions by means of dynamic programming and Hidden Markov Models and is the basis
of popular heuristic search methods such as FASTA or BLAST. Needleman and Wunsch (1970), were the first to
present a dynamic programming algorithm that could find the global alignment between two amino acid sequences.
Smith and Waterman (1981), introduced a new algorithm with a different method of scoring similarity aimed at
finding optimum local alignment sub-sequences, at the expense of the global score. Global algorithms are generally
not sensitive for highly diverged sequences with some localised similarities within them.
A particular application of pairwise sequence alignment is quickly searching large DNA and protein databases for
matches to a query sequence. Popular heuristic algorithms, such as those from the FASTA (Pearson and Lipman
1985, 1988) or BLAST (Altschul et al 1990, 1997) families are much faster than algorithms based on dynamic
programming.
Pairwise sequence alignment methods are concerned with finding the best-matching piecewise local or global
alignments of protein (amino acid) or DNA (nucleic acid) sequences. Typically, the purpose of this is to find
homologues (relatives) of a gene or gene-product in a database of known examples.
52
This information is useful for answering a variety of biological questions:
• The identification of sequences of unknown structure or function.
• The study of molecular evolution.
Global alignment
A global alignment between two sequences is an alignment in which all the characters in both sequences participate
in the alignment. Global alignments are useful mostly for finding closely-related sequences. As these sequences
are also easily identified by local alignment methods global alignment is now somewhat deprecated as a technique.
Further, there are several complications to molecular evolution (such as domain shuffling), which prevent these
methods from being useful. Find the global best fit between two sequences.
Alignment cost
The cost of the entire alignment:
c
M = ∑ s ( xi , yi )
i =1
Scoring function
The cost for aligning the two sequences s = VIVALASVEGAS and t = VIVADAVIS
A(s,t) = V I V A L A S V E G A S
| | | | | | |
V I V A D A - V - - I S
53
Bioinformatics
is:
M(A) = 7 matches + 2 mismatches + 3 gaps
=7 –2 –3 =2
Local alignment
Local alignment methods find related regions within sequences they can consist of a subset of the characters within
each sequence. For example, positions 20-40 of sequence A might be aligned with positions 50-70 of sequence B.
This is a more flexible technique than global alignment and has the advantage that related regions, which appear in
a different order in the two proteins (which is known as domain shuffling) can be identified as being related. This
is not possible with global alignment methods.
As a result, it has largely been replaced in practical use by the BLAST algorithm; although not guaranteed to find
optimal alignments, BLAST is much more efficient.
Identity matrix
For biological sequences it is known how one sequence can mutate into another one. First there are point mutation
that is one nucleotide or amino acid is changed into another one. Secondly, there is deletion that is one element
(nucleotide or amino acid) or a whole subsequence of element is deleted from the sequence. Thirdly, there are
insertions such as one element or a subsequence is inserted into the sequence. First approach the similarity of two
biological sequences that can be expressed through the minimal number of mutations to transform one sequence
into another one. All mutations are not equally likely. Point mutations are more likely because an amino acid can be
replaced by an amino acid with similar chemical properties without changing the function. Deletions and insertions
are more prone to destroying the function of the protein, where the length of deletions and insertions must be taken
into account. For simplicity we can count the length of insertions and deletions. Finally, we are left with simply
counting the number of amino acids, which match in the two sequences (it is the length of both sequences added
together and insertions, deletions and two times the mismatches subtracted, finally divided by two).
54
Here is an example:
BIOINFORMATICS BIOIN-FORMATICS
!
BOILING FOR MANICS B-OILINGFORMANICS
The hit count gives 12 identical letters out of the 14 letters of BIOINFORMATICS. The mutations would be:
• delete I BOINFORMATICS
• insert LI BOILINFORMATICS
• insert G BOILINGFORMATICS
• change T into N BOILINGFORMANICS
These two texts seem to be very similar. Note that insertions or deletions cannot be distinguished if two sequences
are presented (is I deleted form the first string or inserted in the second?). Therefore, both are denoted by a “-” (note,
two “-” are not matched to one another). The task for bioinformatics algorithms is to find from the two strings (left
hand side in above example) the optimal alignment (right hand side in above example). The optimal alignment is
the arrangement of the two strings in a way that the number of mutations is minimal. The optimality criterion scores
matches (the same amino acid) with 1 and mismatches (different amino acids) with 0. If these scores for pairs of
amino acids are written in matrix form, then the identity matrix is obtained. The number of mutation is one criterion
for optimality but there exists more. In general, an alignment algorithm searches for the arrangement of two sequences
such that a criterion is optimised. The sequences can be arranged by inserting “-” into the strings and moving them
horizontally against each other. For long sequences the search for an optimal alignment can be very difficult.
One tool for representing alignments is the dot matrix, where one sequence is written horizontally on the top and
the other one vertically on the left. This gives a matrix where each letter of the first sequence is paired with each
letter of the second sequence. For each matching of letters, a dot is written in the according position in the matrix.
Which pairs appear in the optimal alignment? We will see later, that each path through the dot matrix corresponds
to an alignment. The dots on diagonals correspond to matching regions.
B I O I N F O R M A T I C S
B
O
I
L
I
N
G
F
O
R
M
A
N
I
C
S
55
Bioinformatics
B I O I N F O R M A T I C S
B
O
I
L
I
N
G
F
O
R
M
A
N
I
C
S
Molecular distances of evolution between species can be calculated using various metrics based on DNA or protein
sequence difference. The smaller the number of differences in the DNA and/or protein sequences of similar genes
from two related organisms, the less they have evolutionarily diverged from each other.
56
Bioinformatics
Chapter V
Phylogenetic Analysis
Aim
The aim of this chapter is to:
• define phylogenetics
Objectives
The objectives of this chapter are to:
• define bootstrapping
Learning outcome
At the end of this chapter, you will be able to:
60
5.1 Introduction
Phylogenetic analysis is the process you use to determine the evolutionary relationships between organisms. The
results of an analysis can be drawn in a hierarchical diagram called a cladogram or phylogram (phylogenetic tree).
The branches in a tree are based on the hypothesised evolutionary relationships (phylogeny) between organisms. Each
member in a branch, also known as a monophyletic group, is assumed to be descended from a common ancestor.
Originally, phylogenetic trees were created using morphology, but now, determining evolutionary relationships
includes matching patterns in nucleic acid and protein sequences.
Phylogenetics is the study of evolutionary relationships. Phylogenetic analysis is the means of inferring or estimating
these relationships. The evolutionary history inferred from phylogenetic analysis is usually depicted as branching,
treelike diagrams that represent an estimated pedigree of the inherited relationships among molecules (‘gene trees”),
organisms, or both. Phylogenetics is sometimes called cladistics because the word ‘clade,’ a set of descendants
from a single ancestor, is derived from the Greek word for branch. However, cladistics is a particular method of
hypothesising about evolutionary relationships.
The basic tenet behind cladistics is that members of a group or clade share a common evolutionary history and
are more related to each other than to members of another group. A given group is recognised by sharing unique
features that were not present in distant ancestors. These shared, derived characteristics can be anything that can be
observed and described from two organisms having developed a spine to two sequences having developed a mutation
at a certain base pair of a gene. Usually, cladistic analysis is performed by comparing multiple characteristics or
‘characters’ at once, either multiple phenotypic characters or multiple base pairs or amino acids in a sequence.
• There are three basic assumptions in cladistics. Any group of organisms is related by descent from a common
ancestor (fundamental tenet of evolutionary theory).
• There is a bifurcating pattern of cladogenesis. This assumption is controversial.
• Change in characteristics occurs in lineages over time. This is a necessary condition for cladistics to work.
The resulting relationships from cladistic analysis are most commonly represented by a phylogenetic tree:
A node
Human
A clade
Mouse
Fly
Even with this simple tree, a number of terms that are used frequently in phylogenetic analysis can be introduced:
• A clade is a monophyletic taxon. Clades are groups of organisms or genes that include the most recent common
ancestor of all of its members and all of the descendants of that most recent common ancestor. Clade is derived
from the Greek word ‘klados,’ meaning branch or twig.
• A taxon is any named group of organisms but not necessarily a clade.
• In some analyses, branch lengths correspond to divergence (example, in the above example, mouse is slightly
more related to fly than human is to fly).
• A node is a bifurcating branch point.
61
Bioinformatics
Macromolecules, especially sequences, have surpassed morphological and other organism characters as the most
popular form of data for phylogenetic or cladistic analysis. Although numerous phylogenetic algorithms, procedures,
and computer programs have been devised, their reliability and practicality are, in all cases, dependent on the
structure and size of the data.
The danger of generating incorrect results is inherently greater in computational phylogenetics than in many other
fields of science. The events yielding a phylogeny happened in the past and can only be inferred or estimated.
Despite the well-documented limitations of available phylogenetic procedures, current biological literature is repleted
with examples of conclusions derived from the results of analyses in which data had been simply run through
one or another phylogeny program. Occasionally, the limiting factor in phylogenetic analysis is not so much the
computational method used; more often than not, the limiting factor is the users’ understanding of what the method
is actually doing with the data.
62
Fig. 5.2 A phylogenetic tree
(Source: [Link]
Example of a phylogenetic tree based on genes that do not match organismal phylogeny, suggesting horizontal gene
transfer has occurred. The ancestor of protozoan eukaryote 1 (underlined and marked with an arrow) appears to
have obtained the gene from the ancestor of Bacteria 1, 2, and 3, as this is the simplest explanation for the results.
This unexpected result is not without precedent: there have been a number of reported phylogenetic analyses that
suggest that protozoa have taken up genes from bacteria, most likely from bacteria that they have ingested.
There are additional assumptions that are defaults in some methods but can be at least partially corrected for in
others:
• The sequences in the sample evolved according to a single stochastic process.
• All positions in the sequence evolved according to the same stochastic process.
• Each position in the sequence evolved independently.
Errors in published phylogenetic analyses can often be attributed to violations of one or more of the foregoing
assumptions. Every sequence data set must be evaluated against these assumptions, with other possible explanations
for the observed results considered.
Studies of protein and gene evolution involve the comparison of homologs sequences that have common origins but
may or may not have common activity. Sequences that share an arbitrary, threshold level of similarity determined
by alignment of matching bases are termed homologous. They are inherited from a common ancestor that possessed
similar structure, although the structure of the ancestor may be difficult to determine because it has been modified
through descent.
63
Bioinformatics
Occasionally, the %(G _ C) content may be so vastly different from the average gene in the current host that a
conclusion of external origin is nearly inescapable, however often it is unclear whether a gene has horizontal origins.
Function of xenologs can be variable depending on how significant the change in context was for the horizontally
moving gene; however, in general, the function tends to be similar.
Each step is critical for the analysis and should be handled accordingly. For example, trees are only as good as the
alignment they are based on. When performing a phylogenetic analysis, it is often insightful to build trees based on
different modifications of the alignment to see how the alignment proposed influences the resulting tree.
Aligned sequence positions subjected to phylogenetic analysis represent a priori phylogenetic conclusions because
the sites themselves (not the actual bases) are effectively assumed to be genealogically related, or homologous. Sites
at which one is confident of homology and that contain changes in character states useful for the given phylogenetic
analysis are often referred to as ‘informative sites.’
Steps in building the alignment include selection of the alignment procedure(s) and extraction of a phylogenetic
data set from the alignment. The latter procedure requires determination of how ambiguously aligned regions and
insertion/deletions (referred to as indels, or gaps) will be treated in the tree-building procedure.
A typical alignment procedure involves the application of a program such as CLUSTAL W, followed by manual
alignment editing and submission to a tree building program. This procedure should be performed with the following
questions and considerations in mind.
64
5.4.3 Tree-Building Methods
Tree building methods can be sorted into distance-based vs. character-based methods. Much of the discussion in
molecular phylogenetics dwells on the utility of distance and character-based methods (example, Saitou, 1996; Li,
1997). Distance methods compute pairwise distances according to some measure and then discard the actual data,
using only the fixed distances to derive trees. Character-based methods derive trees that optimise the distribution
of the actual data patterns for each character. Pairwise distances are, therefore, not fixed, as they are determined
by the tree topology. The most commonly applied distance-based methods include neighbor-joining and the most
common character-based methods include maximum parsimony and maximum likelihood.
• Distance-based: Transform the data into pairwise distances (dissimilarities), and then use a matrix during tree
building.
• Character-based: Use the aligned characters, such as DNA or protein sequences, directly during tree inference
– based on substitutions.
Bootstrap
Bootstrapping is a resampling tree evaluation method that works with distance, parsimony, likelihood, and just about
any other tree derivation method. It was invented in 1979 (Efron, 1979) and introduced as a tree evaluation method
in phylogenetic analysis by Felsenstein (1985). The result of bootstrap analysis is typically a number associated
with a particular branch in the phylogenetic tree that gives the proportion of bootstrap replicates that supports the
monophyly of the clade.
Bootstrapping can be considered a two-step process comprising the generation of (many) new data sets from the
original set and the computation of a number that gives the proportion of times that a particular branch (example, a
taxon) appeared in the tree. That number is commonly referred to as the bootstrap value. New data sets are created
from the original data set by sampling columns of characters at random from the original data set with replacement.
‘With replacement’ means that each site can be sampled again with the same probability as any of the other sites.
As a consequence, each of the newly created data sets has the same number of total positions as the original data
set, but some positions are duplicated or triplicated and others are missing. It is therefore possible that some of the
newly created data sets are completely identical to the original set—or, on the other extreme, that only one of the
sites is replicated, say, 500 times, whereas the remaining 499 positions in the original data set are dropped.
Although it has become common practice to include bootstrapping as part of a thorough phylogenetic analysis,
there is some discussion on what exactly is measured by this method. It was originally suggested that the bootstrap
value is a measure of repeatability (Felsenstein, 1985). In more recent interpretations, it has been considered to be
a measure of accuracy a biologically more relevant parameter that gives the probability that the true phylogeny has
been recovered. On the basis of simulation studies, it has been suggested that, under favourable conditions (roughly
equal rates of change, symmetric branches), bootstrap values greater than 70% correspond to a probability of greater
than 95% that the true phylogeny has been found (Hillis and Bull, 1993). By the same token, under less favourable
conditions, bootstrap values greater than 50% will be overestimates of accuracy (Hillis and Bull, 1993). Simply put,
under certain conditions, high bootstrap values can make the wrong phylogeny look good; therefore, the conditions
of the analysis must be considered. Bootstrapping can be used in experiments in which trees are recomputed after
internal branches are deleted one at a time. The results provide information on branching orders that are ambiguous
in the full data set (cf. Leipe et al., 1994).
65
The widespread availability of diverse biological databases amplifies researchers' ability to conduct comprehensive, cross-disciplinary studies by integrating genomic, proteomic, and other biological datasets . This diversity facilitates advancements in systems biology and other fields but also introduces challenges in ensuring compatibility, data consistency, and retrieval efficiency .
Biological databases serve as a large, organized collection of persistent data which is crucial for storing and handling the vast volumes of molecular biological data generated by modern genomic research . They must provide easy access to information and methods for extracting specific data to address particular biological questions . Key features of biological databases include autonomy, heterogeneous data formats, dynamic data content, broad domain knowledge, and workflow orientation, with a requirement for information integration .
Historical developments have shaped biological databases by establishing standardized formats and structuring data organization, such as the common format agreed by EMBL/GenBank/DDBJ in 1988 . Advances such as PCR greatly increased nucleotide sequence data, necessitating more sophisticated databases and forming the basis for databases like GenBank and Swiss-Prot that continue to influence current data repository structures and functions .
Dynamic programming in sequence alignment involves breaking down the alignment problem into simpler subproblems, allowing for systematic exploration of alignment possibilities to maximize the alignment score . The Needleman-Wunsch algorithm uses this approach for global alignment by scoring all possible alignments to find the optimal one, while Smith-Waterman focuses on local alignments .
Tools like BLAST play a critical role in sequence alignment by providing fast heuristic search methods to match query sequences with database entries, which is crucial given the immense volume of sequence data . Fundamentally, BLAST uses scoring systems and algorithms to identify high-similarity regions, optimizing the search process by comparing subsequences rather than entire sequences .
Bioinformatics databases face challenges including the need for high computational power to manage data volume growth, execute complex queries efficiently, and integrate heterogeneous datasets . Additionally, dynamic changes in data content and the broad domain knowledge required pose significant demands on database frameworks and computational strategies .
Sequence databases such as GenBank provide essential nucleotide data that supports genomic research needs by offering easily accessible, well-annotated sequence information . These databases assist with identifying genes, understanding genetic functions, and facilitating the connectivity across academic and industrial research efforts .
Information integration is crucial because biological data are diverse, collected from various experiments, and stored in different database formats . Integration allows researchers to have a comprehensive overview of data related to their inquiries, achieved by linking or aggregating data from multiple databases using tools like the Sequence Retrieval System (SRS) that connects over 150 heterogeneous resources .
Primary biological databases contain core data such as sequences or structures directly obtained from experiments, for example, GenBank for genome sequences and Swiss-Prot for protein sequences . Secondary databases, on the other hand, contain derived information, including conserved sequences and protein family motifs obtained through the analysis of primary data .
Java and Perl have significantly impacted bioinformatics tool development. Java's robustness and portability support complex applications and computational tasks. Perl excels in text processing tasks necessary for bioinformatics data format conversion and regular expression matching . These languages have facilitated the creation of versatile and efficient bioinformatics tools, fostering innovation and ease of use in database management and querying .