Introduction to Bioinformatics online course: IBT_2025
Module topic: Multiple Sequence Alignment
Contact session title: Building a Multiple Sequence Alignment
Trainer: Ahmed M. Alzohairy
Assignment – Building a Multiple Sequence Alignment
Introduction
Building multiple sequence alignments is far from an exact science. It’s
more art than science, requiring that you use everything you know in
bioinformatics and biology. The main idea behind building a multiple
sequence alignment is to put similar amino acids or nucleotides in the
same column if they contain the same criterion. There are four major
criteria most scientists use to build a multiple alignment of sequences
that all have different properties.
NB: Hand-in information - please upload your completed assignment to
the Vula ‘Assignments’ tab. Take note of the final hand-in date for each
assignment indicated on Vula
Tools used in this session
For building an MSA, the most popular programs. We will see the
differences between ClustalW, MUSCLE, and Tcoffee. And for analysis,
we will use other online tools.
Please note
Hand-in information: Please upload your completed assignment
to the Vula assignments tab. Take note of the final hand-in date,
which will be indicated on Vula.
Task 1: Reading the slides
Task 1: instructions
Read and understand the contents of the PPT slides
Introduction to Bioinformatics online course: IBT_2025
Task 1: Sequence Retrieval and FASTA Formatting
Task:
Retrieve the HSF1 (Heat Shock Factor 1) protein sequence for Human,
Mouse, and Chicken from UniProt. Save them in a single file in FASTA
format.
Task 2: Aligning Sequences and Interpreting Conserved Regions
Task:
Align the retrieved sequences using Clustal Omega, then identify one
conserved block using the sequence logo from WebLogo.
Task 3: Extracting annotation (identifying conserved domain)
Task 3: instructions
Task 3: Memo answer MSA Software Matching Task: Match each MSA
tool to its optimal use case: COBALT, ClustalW, T-Coffee, MUSCLE,
MAFFT.
Best Use Case
- Domain-based alignment with constraints
- Widely used standard alignment method
- Small-size, highly accurate alignments
- Fast alignment for medium-sized datasets
- Fast and scalable for large datasets
Task 4: Fast alignment for medium-sized datasets
Task 4: List four biological applications of multiple sequence
alignments (MSA) and briefly explain their purpose.
Task 5: Troubleshooting a Poor MSA
Task 5: You attempted an MSA and obtained an alignment with many
gaps and low conservation. What steps can you take to improve it?