0% found this document useful (0 votes)
3 views8 pages

Mastering Multiple Sequence Alignment

The lecture notes provide structured explanations of core management and marketing concepts. They cover management functions, theories, strategic decision-making, marketing fundamentals, consumer behavior, and the marketing mix.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views8 pages

Mastering Multiple Sequence Alignment

The lecture notes provide structured explanations of core management and marketing concepts. They cover management functions, theories, strategic decision-making, marketing fundamentals, consumer behavior, and the marketing mix.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Introduction to Bioinformatics online course: IBT_2025

Module topic: Multiple Sequence Alignment


Contact session title: Building a Multiple Sequence Alignment
Trainer: Ahmed M. Alzohairy
Participant: Moyosore Emmanuel
Date: 15/06/2025

Assignment – Building a Multiple Sequence Alignment

Introduction

Building multiple sequence alignments is far from an exact science. It’s


more art than science, requiring that you use everything you know in
bioinformatics and biology. The main idea behind building a multiple
sequence alignment is to put similar amino acids or nucleotides in the
same column if they contain the same criterion. There are four major
criteria most scientists use to build a multiple alignment of sequences
that all have different properties.

NB: Hand-in information - please upload your completed assignment to


the Vula ‘Assignments’ tab. Take note of the final hand-in date for each
assignment indicated on Vula

Tools used in this session

For building an MSA, the most popular programs. We will see the
differences between ClustalW, MUSCLE, and Tcoffee. And for analysis,
we will use other online tools.

Please note
 Hand-in information: Please upload your completed assignment
to the Vula assignments tab. Take note of the final hand-in date,
which will be indicated on Vula.
Task 1: Reading the slides
Introduction to Bioinformatics online course: IBT_2025

Task 1: instructions
Read and understand the contents of the PPT slides

Task 1: Sequence Retrieval and FASTA Formatting


Task:
Retrieve the HSF1 (Heat Shock Factor 1) protein sequence for Human,
Mouse, and Chicken from UniProt. Save them in a single file in FASTA
format.

Task 2: Aligning Sequences and Interpreting Conserved Regions


Task:
Align the retrieved sequences using Clustal Omega, then identify one
conserved block using the sequence logo from WebLogo.

Task 3: Extracting annotation (identifying conserved domain)


Task 3: instructions
Task 3: Memo answer MSA Software Matching Task: Match each MSA
tool to its optimal use case: COBALT, ClustalW, T-Coffee, MUSCLE,
MAFFT.
Best Use Case
- Domain-based alignment with constraints
- Widely used standard alignment method
- Small-size, highly accurate alignments
- Fast alignment for medium-sized datasets
- Fast and scalable for large datasets

Task 4: Fast alignment for medium-sized datasets


Task 4: List four biological applications of multiple sequence
alignments (MSA) and briefly explain their purpose.

Task 5: Troubleshooting a Poor MSA


Task 5: You attempted an MSA and obtained an alignment with many
gaps and low conservation. What steps can you take to improve it?

Task 1
Introduction to Bioinformatics online course: IBT_2025

>sp|P38529|HSF1_CHICK Heat shock factor protein 1 OS=Gallus gallus


OX=9031 GN=HSF1 PE=1 SV=1
MEGPGAAAAAVGAGPGGSNVSAFLTKLWTLVEDPETDPLICWSPSGNSFHVFD
QGQFAKEVLPKYFKHNNMASFVRQLNMYGFRKVVHIEQGGLVKPEKDDTEFQH
PYFIRGQEHLLENIKRKVTSVSSIKNEDIKVRQDNVTKLLTDIQVMKGKQESMDSK
LIAMKHENEALWREVASLRQKHAQQQKVVNKLIQFLISLVQSNRILGVKRKIPLM
LNDSSSAHSMPKYSRQYSLEHVHGSSPYAASSPAYSGSNIYSPDSSTNSGPIISDVT
ELAQSSPSASPSGSLDERSSPVVRIKEEPPSPSRSPKENEPSTTTAAAGNSTEQPQP
QEKCLSVACLDKNELNDHLDTIDSNLDNLQTMLSTHGFSVDTTALLDLFSPSMTV
TDMNLPDLDSSLASIQDLLSSQEQQKPSEADAAAADTGKQLVHYTAQPLFLVDSS
AVDVGSGDLPIFFELGEGSYFTDGDEYNEDPTISLLSGTEQPKPKDPTVS

>sp|P38532|HSF1_MOUSE Heat shock factor protein 1 OS=Mus


Introduction to Bioinformatics online course: IBT_2025

musculus OX=10090 GN=Hsf1 PE=1 SV=2


MDLAVGPGAAGPSNVPAFLTKLWTLVSDPDTDALICWSPSGNSFHVFDQGQFA
KEVLPKYFKHNNMASFVRQLNMYGFRKVVHIEQGGLVKPERDDTEFQHPCFLR
GQEQLLENIKRKVTSVSTLKSEDIKIRQDSVTRLLTDVQLMKGKQECMDSKLLAM
KHENEALWREVASLRQKHAQQQKVVNKLIQFLISLVQSNRILGVKRKIPLMLSDS
NSAHSVPKYGRQYSLEHVHPGPYSAPSPAYSSSSLYSSDAVTSSGPIISDITELAPTS
PLASPGRSIDERPLSSSTLVRVKQEPPSPPHSPRVLEASPGRPSSMDTPLSPTAFIDS
ILRESEPTPAASNTAPMDTTGAQAPALPTPSTPEKCLSVACLDKNELSDHLDAMD
SNLDNLQTMLTSHGFSVDTSALLDLFSPSVTMPDMSLPDLDSSLASIQELLSPQEP
PRPIEAENSNPDSGKQLVHYTAQPLFLLDPDAVDTGSSELPVLFELGESSYFSEGD
DYTDDPTISLLTGTEPHKAKDPTVS

>sp|Q00613|HSF1_HUMAN Heat shock factor protein 1 OS=Homo


sapiens OX=9606 GN=HSF1 PE=1 SV=1
MDLPVGPGAAGPSNVPAFLTKLWTLVSDPDTDALICWSPSGNSFHVFDQGQFA
KEVLPKYFKHNNMASFVRQLNMYGFRKVVHIEQGGLVKPERDDTEFQHPCFLR
GQEQLLENIKRKVTSVSTLKSEDIKIRQDSVTKLLTDVQLMKGKQECMDSKLLAM
KHENEALWREVASLRQKHAQQQKVVNKLIQFLISLVQSNRILGVKRKIPLMLNDS
GSAHSMPKYSRQFSLEHVHGSGPYSAPSPAYSSSSLYAPDAVASSGPIISDITELAP
ASPMASPGGSIDERPLSSSPLVRVKEEPPSPPQSPRVEEASPGRPSSVDTLLSPTAL
IDSILRESEPAPASVTALTDARGHTDTEGRPPSPPPTSTPEKCLSVACLDKNELSDH
LDAMDSNLDNLQTMLSSHGFSVDTSALLDLFSPSVTVPDMSLPDLDSSLASIQELL
SPQEPPRPPEAENSSPDSGKQLVHYTAQPLFLLDPGSVDTGSNDLPVLFELGEGSY
FSEGDGFAEDPTISLLTGSEPPKAKDPTVS.

Task 2
Introduction to Bioinformatics online course: IBT_2025
Introduction to Bioinformatics online course: IBT_2025

One conserved block using the sequence logo from WebLogo is


Tryptophan (W)
Introduction to Bioinformatics online course: IBT_2025

Task 3
MSA Tool
COBALT - Domain-based alignment with constraints
ClustalW - Widely used standard alignment method
T-Coffee - Small-size, highly accurate alignments
MUSCLE - Fast alignment for medium-sized datasets
MAFFT - Fast and scalable for large datasets

Task 4
Four biological applications of multiple sequence alignments (MSA)
i. Extrapolation: A good multiple alignment can help convince you
that an uncharacterized sequence is really a member of a protein
family. Alignments that include Swiss-Prot sequences are the most
informative.
ii. Structure prediction: A good multiple sequence alignment can give
an almost perfect prediction of protein secondary structure for
both proteins and RNA. Sometimes it can also help in the building
of a 3-D model”.
iii. Phylogenetic Analysis: Multiple Sequence Alignment (MSA) can
help infer evolutionary relationships between organisms based on
sequence similarities. By carefully choosing the sequences to
include in a multiple sequence alignment analysis, it is possible to
reconstruct the history of these proteins.
iv. Pattern identification: Multiple Sequence Alignment identifies
functionally important conserved regions such as active sites. By
discovering very conserved positions, one can identify a region
that is characteristic of a function (in proteins or nucleic-acid
sequences).
Introduction to Bioinformatics online course: IBT_2025

Task 5

i. Remove gaps
ii. Remove extremities
iii. Keep informative blocks

You might also like