Bioinformatics – concepts and methods
A dvanced level, 7.5 ECTS
Course Introduction
Björn Olsson
Associate professor in computer science and bioinformatics
[Link]@[Link]
Bild 1
Welcome to the course! The first lecture contains some administrative information,
followed by a general introduction to the subject. The following topics will be covered:
• Aims and learning goals for the course
• Overview of the course structure (lectures, exercises and assignments)
• Examination
• Course literature
• The history of bioinformatics
• The biological foundations of bioinformatics
• Examples of analysis tasks performed by bioinformaticians
1
Course contents
The course will address the following questions:
• What is bioinformatics?
• What kinds of biological questions are commonly addressed using
bioinformatic methods?
• What types of biological data do we analyze with bioinformatic tools?
• What are the most important biological databases?
You will develop skills in the following bioinformatic tasks:
• Comparing biological sequences (i.e. RNA and DNA sequences)
• Searching for homologs in sequence databases
• Modelling evolutionary relationships
• Primer design
• Analysing parts of a genome sequence, finding and characterising the
genes within it
Bild 2
Ths slides gives an overview of what is covered in the course. The course has an emphasis
on practical skills. The list shows the specific analysis tasks that you should be able to
perform, using bioinformatic tools, after completing the course.
Learning goals
At the end of the course, the student should be able to:
Extensively define concepts and questions that are central to
bioinformatics, and critically evaluate practical applications of
bioinformatics.
In an extensive way define the practical applications of bioinformatics.
Describe how bioinformatics has developed.
Describe and evaluate the public information sources that are central
to bioinformatics, and how these are structured.
Independently apply and extensively describe methods and tools used
for analysis of large-scale data in molecular biology and biomedicine.
Independently draw conclusions regarding strategies to solve a given
bioinformatic problem and critically analyze the results
Official course curriculum document on the course site:
Kursplan_BI760A.pdf (swe), Courseplan_BI760A.pdf (eng)
Bild 3
The learning goals shown above are taken from the official course plan, which regulates
what the course should include. You can find the complete document in the course site
(English version “Courseplan_BI760A”, Swedish version “Kursplan_BI760A”). As indicated by
the learning goals, this is an overview course that introduces the subject of bioinformatics.
By the end of the course you should have a good idea of the applications of bioinformatics,
i.e. for what biological analysis tasks we use bioinformatic tools. You should also gain some
understanding of how bioinformatics has developed over the years, which is an important
issue since this subject changes very rapidly. Another very important topic is to be
knowledgeable about the public information sources, i.e. biological databases, that are
available. You should understand how they are structured and how to access the data in
them. The major part of the course is to learn about different methods and tools which are
used to analyse large datasets. You will learn about these tools in the lectures and also
practice using them during the assignments. All software that we use in this course is
available publically and free, either through web page interfaces or for download to your
computer. The final learning point shows how the knowledge is integrated and put into
practical use, in other words how to work as a bioinformatician in practice. You should
develop this skill during the assignments, where you will be given specific problems to
solve. You will have to select strategies for solving the problem and use your ability of
critical thinking when analysing the results that are produced by the bioinformatic tools.
For each assignment it is required that you write a report that shows how you solved the
problem and how you analyzed the results. Examination and grading is based on the quality
of your reports, so you should put a lot of effort into writing.
Course literature
All required material will be available on the
course site, and consists of:
Lecture slides + Lecture notes
Exercises
Assignments
It is also recommended to read:
Suggested overview articles that will be uploaded
on the course site
And/or a text book
Bild 4
The main literature for the course consists of the lecture slides and the accompanying
lecture notes. These will be uploaded to the course site and it is obligatory to read and
study all lecture materials. In addition, there will be exercises and assignments. The
exercises and assignments cover the obligatory knowledge covered by the course.
However, examination is based only on the assignments. In other words, you will do the
exercises as practice, as part of your learning process, whereas you do the assignments
mostly in order to demonstrate your knowledge and skills. In addition to the obligatory
readings, it is suggested that you also read articles and/or a text book covering the topics in
more detail. Suggested articles will be added on the course site as the course progresses. It
can be very good to also read a text book. When doing so, you should keep in mind that
bioinformatics is a quickly developing area, which means that tools and databases
described in a text book may have changed since the book was published. So when reading
books your focus should be on understanding the general principles.
Course contents
• Assignment reports can only be submitted through Canvas
• Submission closes at 23:59 on the deadline date
Description of how the course is organized:
Bild 5
[Link]
The course covers a number topics in bioinformatics. For each topic you will get lecture
material consisting of slides and text to read. The first and second modules cover the topics
“Biological databases” and “Sequence alignments”. These two topics constitute a very
central part of “traditional” bioinformatics and lays a foundation of knowledge that anyone
should have when working with bioinformatic tools. You will become familiar with the most
important databases and how to use the search and alignment tools that any
bioinformatician should have expertise in. For the first module ("Databases"), the
examination is done with a quiz that you solve online. The quiz has a time limit of two
hours and you can only make one attempt. It will be available online during the days shown
in the table. For the second module, the examination is through an assignment, where you
will get a few tasks to solve and your should document your solutions in a written report.
Assignments must be solved individually, meaning that each student must submit a unique
solution and an individual report. (And we do check if different reports contain material
that has been “borrowed” from other students). The third module has the topic
“Bioinformatics literature survey”. For that topic, you will get articles to read and
summarize in a written report. The fourth assignment covers “Primer design”, which is an
important step for successful polymerase chain reaction (PCR) experiments. PCR is an
essential technique for amplification of genetic material and is frequently used in genetics
and molecular biology. The fifth and final study module covers the topics “Genomics” and
“Gene prediction”. In the first part, you learn how to use genome databases, how to
identify novel genes in genome sequences and how to make cross-genome comparisons
using bioinformatic tools. In the second part, you learn about methods identifying genes
within genome sequences. This final module is examined with a written exam that you
need to do on campus on a specific date, as shown in the table.
Course site in Canvas, organized by Modules
When a module has opened, you will there find:
• Recorded lecture(s) to watch
• Lecture notes to read
• A recommended article or other additional reading
• A set of exercises to work with for practice
• A link to a Discussion where you can ask questions and get help
• A link to the Assignment, which is the examination to pass for the Module
Bild 6
All course material will be distributed through the course site in Canvas. The first thing you
should do there is to read the Study Guide. After that, proceed to look at the list of
modules. Each module will open on a specific date. First look at the dates and plan your
studies accordingly. All of this should be self-explanatory if you work through the material,
one document at the time, from the first module onwards.
6
Teacher contact information
Bild 7
You can find the contact information to each teacher through the links on the course page.
7
Communication with teachers
• Note:
• We do not provide help in solving the assignments!
• But we do provide help in solving the exercises
• Therefore, if anything is unclear or difficult when you
work with an exercise, make sure to ask the teacher
who is responsible for that exercise
• For each Lecture+Exercise we open a new Discussion
on the course site, where you can ask questions on that
theme (and see what questions other students have
already asked)
Bild 8
Please note carefully the following: We can not provide any help in solving assignments.
This is because the assignments are only for examination and testing your knowledge. By
the time you solve an assignment, you should already have studied the lecture material
and the exercise on the same topic. If you feel that anything is unclear in the lecture or in
the exercise, then that is the right time to ask your questions! When solving the
assignment, it is already too late to ask questions, since by that time you are being
examined.
Examination
5 examinations
1.5 ECTS credits each
Module 1: Quiz
Module 2, 3 and 4: Assignments
Module 5: Written exam
Assignments must be solved individually
If you fail in an examination, you will get ONE chance to re-take it
Grading
For each module you get a grade, A, B, C, D, E or F
Pass: A-E
You get an overall grade for the whole course when you have
passed all five modules
Overall grade = average of the 5 module grades
Swedish version of grades: A = Utmärkt, B = Mycket bra, C = Bra,
D = Tillfredställande, E = Tillräcklig, F = Underkänd.
Bild 9
Each module is worth 1.5 ECTS credits, which will be registered as soon as you have passed
the examination for the module. You then get your overall grade for the course when you
have passed all five modules. The assignments mainly consist of practical use of
bioinformatics tools, but you also have to write a report where you describe and discuss
what you have done. It is important that you analyse the results, since your knowledge of
the subject will be evaluated based on the report. Each student must solve the task
individually and write a report him-/herself. If you fail in an examination, you will get ONE
chance to re-take it. Your overall grade in the course is based on the five module grades.
Bioinformatics
>LOC_Os01g01010.1|12001.m06748|1000bp
Acgacgcagctatggcctccccgcccaccaggccgcc
10
Ggcttcctaggtagggatcccatcccttcgattccct
Tttgatttgatttaattcgattgcctgcttttcaggt
ggcgagggagATG
>LOC_Os01g01030.1|12001.m06750|1000bp
5
gtaggtggcgcggtccgcgcggatatggaggaggctg
Log fold change
acgcgtggagcggaacagcatcgcctgcagctttgt
cccaccttgacctccgcttcgaattcgaattcgaat
ccgATG
0
CBF2/AT4G25470
>LOC_Os01g01040.1|12001.m06751|1000bp
POLASIG3 ABRE 1.98e-63
CBF3/AT4G25480
CBF1/AT4G25490
CBF4/AT5G51990
Attagccgttgtgtgatatctgaagatcaacttacta AT4G32800
DREB3/AT5G11590
Cccaaacctagactcacagaaaccaattgattgatat AT2G44940
ROOTMOTIFTAPOX1 ABRE
ggattaagacctgaagcttcttgggttttccta 1.98e-63 AT2G35700
-5
0 5 10 15 20 25
GT-1 binding ABRE 1.98e-63 Time
ROOTMOTIFTAPOX1 ABRE 1.98e-63
GT-1 binding ABRE 1.98e-63
Biological data Analysis
Biological experiment
Bio+informatics = using software tools (informatics) to collect, organize, store,
analyze and interpret data obtained from biological experiments
Bild 10
Bioinformatics is about the use of information technology (computers and computer
programs) to handle molecular biological data. The data sets are generated in experiments
carried out by biologists in the wet lab, which means that anyone working with
bioinformatics must have knowledge about both biology and computer science. The role of
bioinformatics is to use computers and software to collect, organize, and store the data in
databases. Bioinformatics is also about the development and use of computer programs for
analyzing biological data. In some cases, this analysis also includes development of
computer models of biological systems, but this approach is more common in systems
biology than in bioinformatics. As you can see from this overview, bioinformatics is a cross-
disciplinary field. Most bioinformaticians have a first degree in one subject, either in
biology or computer science, and on top of that a second degree specialising in
bioinformatics.
Bioinformatics – a wide and
diverse field
Applied bioinformatics
Use of existing tools and software
Focus on biological findings
Suitable background: Molecular biology, genetics,
biomedicine, and related fields in biology
Computational bioinformatics
Development of new algorithms and tools
Development and maintenance of biological databases
Suitable background: Computer science, programming,
mathematics, statistics, and related fields in data sciences
Bild 11
Applied bioinformatics is focused on using existing software tools to analyze data and
deliver new biological findings. This is often done by bioinformaticians in close
collaboration with biologists, or by the biologists themselves, if they have sufficient
knowledge of bioinformatic tools. For applied bioinformatics it is necessary to be very
knowledgeable about the advantages and disadvantages of different bioinformatic tools, so
that you can choose the most suitable software for a particular task. It is also necessary to
be very knowledgeable about the biological system being studied, since the focus is on the
biological findings.
Computational bioinformatics is focused on developing new bioinformatic algorithms and
tools, or improving the existing methods. To specialize in computational bioinformatics, you
normally need a degree in computer science, since the main task is to design of algorithms
and implementation of software. The course you are taking now has an orientation
towards applied bioinformatics. The goal is that after completion of the course you should
be an advanced user of some common bioinformatic tools. By “advanced user” we mean
that you not only know how to run the tools, but you are also able to adjust their
parameters in a way that suits the data set you are analyzing. You should also be able to
choose between different tools for a given task, and know the advantages and
disadvantages of a particular tool compared to another.
A brief history of bioinformatics
Bild 12
How did bioinformatics develop? Here is a brief selection of events in the history of the field, including advances both in
experimental techniques and bioinformatics analysis methods.
Until early 1970s biologists had to compare DNA and protein sequences manually, which was very time-consuming.
Needleman and Wunsch in 1970 published the first computer algorithm for sequence comparison and alignment. In 1971
the Brookhaven Protein Data Bank was founded, which holds data about the 3D structure of proteins. The1970s saw the
first attempts to use computers to predict protein secondary structure prediction. In 1978, the term “bioinformatics” was
used for the first time, and defined as “studies of informatics processes in biotic systems”. In 1980 Smith and Waterman
published an improved, and very influential, algorithm for pairwise sequence comparison and alignment. In 1983
polymerase chain reaction (PCR) was invented. In 1986 the SWISS-PROT database was founded, holding annotated
protein sequences (such as description of function, domain structure, post-translational modifications etc..). In 1986 the
FASTA algorithm was published, which made it possible to speed up sequence comparison and handle the rapidly
growing amounts of sequence data. In 1990 the BLAST (Basic Local Alignment Search Tool) program was released, which
allows similarity search in databases containing several million sequences. BLAST is more time-efficient than FASTA, due
to searching only for the most significant patterns in the sequences. In 1991 the use of Expressed Sequence Tags was
described, opening the possibility of sequencing all expressed genes in an organism (the transcriptome) in a single
experiment. In 1995 the first microarrays for gene expression profiling became available, which made it necessary to
develop bioinformatic methods for comparing expression profiles of genes. In 1996 the PROSITE database of protein
motifs was released. In 1997 the complete genome sequence of C. elegans was published. In 2001 the first draft of the
human genome sequence was published. Since 2002 the rat, mouse, malaria and chimpanzee genomes have been
published, as well as many other important eukaryotic genomes. In the last few years, massively parallel sequencing
(“next generation sequencing”, or NGS) has produced new datasets and new challenges for bioinformaticians to develop
analysis tools for such data. Currently the NGS technique is developing rapidly not only for genome (DNA) sequencing,
but also RNA sequencing, and the so called RNAseq approach can be used to measure gene expression in cases where
previously the EST or microarray approaches were used. A current development (2010 onwards) is that the distinction
between bioinformatics and systems biology is gradually disappearing and the two subjects are becoming integrated with
each other.
12
Databases
Bioinformatics is
largely concerned
with the functions of
genes & proteins at
cellular and
sub-cellular level
Example database entry
Human BRCA2
Describes:
- Molecular functions
- Biological processes
- Subcellular locations
Roles of bioinformatics:
- Derive
- Store
- Organize
- Systematize
- Make available
- Utilize
.. such data and
information Bild 13
The focus of bioinformatics on molecular biology, genetics, cell biology and medicine, is
reflected in the kinds of data being stored in the databases that are developed and
maintained by bioinformaticians. This slide shows as example the UniProt database, which
has a focus on proteins. For each protein this database stores information on the protein’s
molecular function in the cell, which biological processes the protein is involved in, and in
which subcellular locations it is present. As we will see in this course, one of the core roles
of bioinformatics is to derive such information from experimental data, store it in
databases, and organize these databases in a systematic way so that the information can
be made available to the research community. Bioinformatics also provides the tools for
utilizing the data and information for further research.
13
Bioinformatics
Organize, Process, and Make Sense of Complex Biological Data Sets
DNA RNA Protein Pathways
Genome Transcriptome Proteome Interactome
Examples of bioinformatic tasks
Sequence analysis Differential expression Structure prediction Network analysis
Alignment Clustering Docking Network reconstruction
Database search Marker identification Protein motifs Pathway alignment
Phylogeny Family classification
Bild 14
This is an overview of how bioinformatics is used for various analysis tasks and how these
relate to the different levels along the path from the genome to the phenotype. Different
experimental methods are used for studying biological phenomena at the genome,
transcriptome, proteome and interactome levels.
Studies at DNA level primarily produce sequence data which is analysed for example by
aligning different sequences to identify their similarities and differences. Sequences are
also used in database searches, to find homologous (evolutionarily related) sequences,
which is an efficient way of characterizing new sequences. Sequence data are also used as
input for phylogeny analysis, which results in phylogenetic trees that depict the
evolutionary relationships among a set of related sequences.
Experiments at the transcriptome level produce numerical data showing the expression
levels of genes. The analysis tasks include identification of differential expression, i.e. to
find genes that “react” by increased/decreased expression in different circumstances. This
is for example used in biomedical research to find genes that can serve as markers for
disease. Very recently, it has also become increasingly common to use massively parallel
sequencing techniques to sequence all RNA found in a sample, and currently many new
bioinformatic tools are being developed for analysis of such data sets.
At the proteome level experiments may focus on individual proteins, to determine their
three-dimensional structure, or at the interaction between proteins and ligands (molecules
that they bind to). The bioinformatic tasks here include trying to predict protein structure
from sequence data, using simulation software to study docking, etc. In addition, sequence
data from groups of proteins are used to identify conserved sequence regions, i.e. motifs,
14
which can be used for the purpose of family classification.
At the interactome level experiments are done on a larger scale, sometimes covering all
proteins in an organism, with the goal of understanding the whole network of interaction.
The analysis tasks for this are in the border-area between bioinformatics and systems
biology, where systems biology has the aim to understand complex biological systems by
studying the interactions between all components of the system.
14
In this course you will learn to use the following methods (and a few more):
• Search algorithms for finding homologs
• Given an "new" unknown gene and a database of >100 million genes,
find the new gene’s homologs and assign it a tentative function
• Alignment algorithms for sequence comparison
• Given two genes or proteins, which parts of their sequences differ,
which parts are the same, and what does this tell us?
• Given a set of related genes or proteins, which regions of their
sequences are conserved and which regions are more variable? What
can we learn about their function and/or structure from this?
• Methods for classification of proteins
• Given an unknown protein sequence, can we predict what class or
family of proteins it belongs to?
• Gene prediction algorithms
• Given a complete (and newly sequenced) genome, can we predict the
locations of all its genes?
• Methods to structure and analyze gene annotation data
• How can information about the molecular function or biological role of a
gene be standardized and written in such a form that the information is
comparable between different databases?
Bild 15
This list includes the majority of methods covered in the course, but a few additional ones
will also be introduced.
15
>unknown protein
Search tools for finding
MVHLTPDEKNAVCALWGKVNVEEVGGEALG
homologs
RLLVVYPWTQRFFDSFGDLSSPSAVMGNPK
VKAHGKKVLSAFSEGLNHLDNLKGTFAKLS
ELHCDKLHVDPENFRLLGNVLVVVLAHHFG
KDFTPQVQAAYQKVVAGVATALAHKYH
Homology search
Protein sequence database
• Similar sequences
• Same protein from related species
• Analysis of set of homologs
• Conclusion: The unknown
protein is a beta-globin from
the otolemur
Bild 16
Homology search is one of the concepts introduced in Lecture 3, “Pairwise alignment”,
because homology search algorithms are based on alignment algorithms. This is one of the
most central and important concepts of bioinformatics. The term “homology” means
shared ancestry. Genes with shared ancestry tend to have sequences that are similar to a
degree that reflects their evolutionary distance. This fact is taken advantage of in many
different applications of bioinformatics. Thanks to the huge sequence databases available
today, it is often easy to determine to origin of an unknown sequence just by comparing it
to all similar sequences in the databases. In the example shown on the slide, we perform
homology search in a protein sequence database and by analyzing the identified set of
homologs we are able to conclude what type of protein the sequence represents, as well as
the organism it comes from.
16
Alignment tools for
sequence comparison
Using alignment tools we can:
• Infer how proteins or organisms have evolved
• Derive functional and structural feature of proteins,
by analysing conserved regions
• Pinpoint individual mutations of biological interest
• Identify sub-families within a large protein family
Alignment of beta globins from five species
Bild 17
Alignment tools can be used for many different purposes, including those listed on the
slide. The figure shows how a tool for multiple alignment (i.e. alignment of more than two
sequences) is used to compare sequences of beta globin proteins from five organisms.
Using the alignment, we can quantify the similarity between each pair of sequences by
counting the percentage of positions at which they differ. The similarity measures can then
be used to trace the evolutionary history of the five species by building a so called
phylogenetic tree, i.e. a graph indicating the relatedness of the species. As expected, the
proteins from the three primate species are very closely related, wheras the proteins from
the reptile and bird species are distantly related to the primate sequences.
17
Using clustering and classification Analysis of transcriptome
algorithms we can: data
• Identify genes that are co-expressed
• Identify markers for diseases or other important conditions
• Create classifiers for diagnosis and prognosis
• 70 gene signature for
outcome prediction
Metastasis in <5 years
Over-expressed Metastasis free >5 years Bild 18
Under-expressed Excluded (BRCA1 germline)
Experimental techniques in transcriptomics are used to measure the expression levels of
genes, i.e. to what degree each gene is actively producing transcripts (messenger RNAs,
mRNAs). Bioinformatic tools for analysis of such data are used to perform clustering or
classification. Clustering groups together genes that show similar expression patterns,
which is indicates that they are co-expressed and/or co-regulated. This, in turn, can be
indicative of the genes having similar function or being involved in similar biological
processes. With tools based classification algorithms we can identify those genes that are
associated with a particular biological condition or outcome. The example on the slide
shows clusterings of genes whose expression levels were measured in breast cancer
tumors. Patterns are visible (in red and green) of groups of genes which are over-expressed
in tumors from patients with good outcome (metastasis free for >5 years) and under-
expressed in tumors from patients with poor outcome. This type of analysis can lead to
discovery of marker genes to be used for outome prediction, and classification algorithms
can be used to optimize the selection of marker genes, to get as robust predictions as
possible.
18
Bild 19
19
• End of lecture 1
• Topic of next lecture:
Biological databases
Bild 20
In connection with Lecture 2, Biological databases, you will also get exercises to work with,
so that you can start practicing how to access biological data from the public databases.
20