0% found this document useful (0 votes)
46 views41 pages

Motif Descriptors in Protein Sequences

This is about the Profile Pattern
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
46 views41 pages

Motif Descriptors in Protein Sequences

This is about the Profile Pattern
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Using Patterns and

Profiles as Motif Descriptors


Random Sequence
atgaccgggatactgataccgtatttggcctaggcgtacacattagataaacgtatgaagtacgttagactcggcgccgccg

acccctattttttgagcagatttagtgacctggaaaaaaaatttgagtacaaaacttttccgaatactgggcataaggtaca

tgagtatccctgggatgacttttgggaacactatagtgctctcccgatttttgaatatgtaggatcattcgccagggtccga

gctgagaattggatgaccttgtaagtgttttccacgcaatcgcgaaccaacgcggacccaaaggcaagaccgataaaggaga

tcccttttgcggtaatgtgccgggaggctggttacgtagggaagccctaacggacttaatggcccacttagtccacttatag

gtcaatcatgttcttgtgaatggatttttaactgagggcatagaccgcttggcgcacccaaattcagtgtgggcgagcgcaa

cggttttggcccttgttagaggcccccgtactgatggaaactttcaattatgagagagctaatctatcgcgtgcgtgttcat

aacttgagttggtttcgaaaatgctctggggcacatacaagaggagtcttccttatcagttaatgctgtatgacactatgta

ttggcccattggctaaaagcccaacttgacaaatggaagatagaatccttgcatttcaacgtatgccgaaccgaaagggaag

ctggtgagcaacgacagattcttacgtgcattagctcgcttccggggatctaatagcacgaagcttctgggtactgatagca
Sequences like (AAAAAAAGGGGGGG)
repeats
atgaccgggatactgataaaaaaaagggggggggcgtacacattagataaacgtatgaagtacgttagactcggcgccgccg

acccctattttttgagcagatttagtgacctggaaaaaaaatttgagtacaaaacttttccgaataaaaaaaaaggggggga

tgagtatccctgggatgacttaaaaaaaagggggggtgctctcccgatttttgaatatgtaggatcattcgccagggtccga

gctgagaattggatgaaaaaaaagggggggtccacgcaatcgcgaaccaacgcggacccaaaggcaagaccgataaaggaga

tcccttttgcggtaatgtgccgggaggctggttacgtagggaagccctaacggacttaataaaaaaaagggggggcttatag

gtcaatcatgttcttgtgaatggatttaaaaaaaaggggggggaccgcttggcgcacccaaattcagtgtgggcgagcgcaa

cggttttggcccttgttagaggcccccgtaaaaaaaagggggggcaattatgagagagctaatctatcgcgtgcgtgttcat

aacttgagttaaaaaaaagggggggctggggcacatacaagaggagtcttccttatcagttaatgctgtatgacactatgta

ttggcccattggctaaaagcccaacttgacaaatggaagatagaatccttgcataaaaaaaagggggggaccgaaagggaag

ctggtgagcaacgacagattcttacgtgcattagctcgcttccggggatctaatagcacgaagcttaaaaaaaaggggggga
Where is the Implanted ?
atgaccgggatactgatAAAAAAAAGGGGGGGggcgtacacattagataaacgtatgaagtacgttagactcggcgccgccg

acccctattttttgagcagatttagtgacctggaaaaaaaatttgagtacaaaacttttccgaataAAAAAAAAGGGGGGGa

tgagtatccctgggatgacttAAAAAAAAGGGGGGGtgctctcccgatttttgaatatgtaggatcattcgccagggtccga

gctgagaattggatgAAAAAAAAGGGGGGGtccacgcaatcgcgaaccaacgcggacccaaaggcaagaccgataaaggaga

tcccttttgcggtaatgtgccgggaggctggttacgtagggaagccctaacggacttaatAAAAAAAAGGGGGGGcttatag

gtcaatcatgttcttgtgaatggatttAAAAAAAAGGGGGGGgaccgcttggcgcacccaaattcagtgtgggcgagcgcaa

cggttttggcccttgttagaggcccccgtAAAAAAAAGGGGGGGcaattatgagagagctaatctatcgcgtgcgtgttcat

aacttgagttAAAAAAAAGGGGGGGctggggcacatacaagaggagtcttccttatcagttaatgctgtatgacactatgta

ttggcccattggctaaaagcccaacttgacaaatggaagatagaatccttgcatAAAAAAAAGGGGGGGaccgaaagggaag

ctggtgagcaacgacagattcttacgtgcattagctcgcttccggggatctaatagcacgaagcttAAAAAAAAGGGGGGGa
Introduction
• Among the various databases dedicated to the identification of protein families and
domains, PROSITE is the first one created.

• The close relationship of PROSITE with the SWISS-PROT protein database allows the
evaluation of the sensitivity and specificity of the PROSITE motifs and their periodic
reviewing.

• The motif descriptors used in PROSITE are either patterns or profiles, which are
derived from multiple alignments of homologous sequences.
PROSITE PATTERNS
• In some cases the sequence of an unknown protein is too distantly related to any protein
of known structure to detect its resemblance by pairwise sequence alignment.

• However, relationships can be revealed by the occurrence in its sequence of a particular


cluster of residue types, which is variously known as a pattern, motif, signature or
fingerprint.

• These motifs, typically around 10 to 20 amino acids in length.

• Because of specific residues and regions (position[s]) proved to be important to the


biological function of a group of proteins.

• Conserved in both structure and sequence during evolution.


These biologically significant regions or residues are generally:

• Enzyme catalytic sites (see figure).

• Prostethic group attachment sites (heme,


pyridoxal-phosphate, biotin, etc.).

• Amino acids involved in binding a metal ion.

• Cysteines involved in disulphide bonds.

• Regions involved in binding a molecule (ADP/ATP,


GDP/GTP, calcium, DNA,etc.) or another protein.
• As the sequence of biologically meaningful motifs is evolutionarily conserved, a multiple
alignment of them can be reduced to a consensus expression called a regular expression or
pattern.

• Each position of such a pattern can be occupied by any residue from a specified set of
acceptable residues, and in addition can be repeated a variable number of times within a
specified range.

• At strictly conserved positions only one particular amino acid is accepted, whereas at other
positions several amino acids with similar physicochemical properties can be accepted.
• A regular expression is qualitative; it either does match or does not. There is no
threshold above which it consider the match as statistically significant.

• It should be noticed that some families or domains are defined not just by one pattern
but by the co-occurrence of two or more patterns of low specificity.

• The presence of just one of these patterns is not sufficient to assign a protein to a
particular family and/or domain. However, the simultaneous occurrence of linked
patterns gives good confidence that the matched protein belongs to the set being
considered.
PROSITE PROFILES
• Profiles are quantitative motif descriptors providing numerical weights for each
possible match or mismatch between a sequence residue and a profile position.

• A mismatch at a highly conserved position can thus be accepted provided that


the rest of the sequence displays a sufficiently high level of similarity.
Rules for prosite pattern (pattern syntax)
• Pattern syntax
• The standard IUPAC one letter code for the amino acids is used in PROSITE.
• The symbol 'x' is used for a position where any amino acid is accepted.
• Ambiguities are indicated by listing the acceptable amino acids for a given position, between square
brackets '[ ]'. For example: [ALT] stands for Ala or Leu or Thr.
• Ambiguities are also indicated by listing between a pair of curly brackets '{ }' the amino acids that are not
accepted at a given position. For example: {AM} stands for all any amino acid except Ala and Met.
• Each element in a pattern is separated from its neighbor by a '-'.
• Repetition of an element of the pattern can be indicated by following that element with a numerical value
or, if it is a gap ('x'), by a numerical range between parentheses.
Examples:
• x(3) corresponds to x-x-x
• x(2,4) corresponds to x-x or x-x-x or x-x-x-x
• A(3) corresponds to A-A-A
• When a pattern is restricted to either the N- or C-terminal of a sequence, that pattern respectively starts
with a '<' symbol or ends with a '>' symbol.
In some rare cases (e.g. PS00267 or PS00539), '>' can also occur inside square brackets for the C-terminal
element. 'F-[GSTV]-P-R-L-[G>]' is equivalent to 'F-[GSTV]-P-R-L-G' or 'F-[GSTV]-P-R-L>'.

[Link]
Where is the Motif???
atgaccgggatactgatagaagaaaggttgggggcgtacacattagataaacgtatgaagtacgttagactcggcgccgccg

acccctattttttgagcagatttagtgacctggaaaaaaaatttgagtacaaaacttttccgaatacaataaaacggcggga

tgagtatccctgggatgacttaaaataatggagtggtgctctcccgatttttgaatatgtaggatcattcgccagggtccga

gctgagaattggatgcaaaaaaagggattgtccacgcaatcgcgaaccaacgcggacccaaaggcaagaccgataaaggaga

tcccttttgcggtaatgtgccgggaggctggttacgtagggaagccctaacggacttaatataataaaggaagggcttatag

gtcaatcatgttcttgtgaatggatttaacaataagggctgggaccgcttggcgcacccaaattcagtgtgggcgagcgcaa

cggttttggcccttgttagaggcccccgtataaacaaggagggccaattatgagagagctaatctatcgcgtgcgtgttcat

aacttgagttaaaaaatagggagccctggggcacatacaagaggagtcttccttatcagttaatgctgtatgacactatgta

ttggcccattggctaaaagcccaacttgacaaatggaagatagaatccttgcatactaaaaaggagcggaccgaaagggaag

ctggtgagcaacgacagattcttacgtgcattagctcgcttccggggatctaatagcacgaagcttactaaaaaggagcgga

Sequence motifs are short, recurring patterns in DNA that are presumed to have a biological function
Finding (15,4) Motif is Difficult?

atgaccgggatactgatAgAAgAAAGGttGGGggcgtacacattagataaacgtatgaagtacgttagactcggcgccgccg

acccctattttttgagcagatttagtgacctggaaaaaaaatttgagtacaaaacttttccgaatacAAtAAAAcGGcGGGa

tgagtatccctgggatgacttAAAAtAAtGGaGtGGtgctctcccgatttttgaatatgtaggatcattcgccagggtccga

gctgagaattggatgcAAAAAAAGGGattGtccacgcaatcgcgaaccaacgcggacccaaaggcaagaccgataaaggaga

tcccttttgcggtaatgtgccgggaggctggttacgtagggaagccctaacggacttaatAtAAtAAAGGaaGGGcttatag

gtcaatcatgttcttgtgaatggatttAAcAAtAAGGGctGGgaccgcttggcgcacccaaattcagtgtgggcgagcgcaa

cggttttggcccttgttagaggcccccgtAtAAAcAAGGaGGGccaattatgagagagctaatctatcgcgtgcgtgttcat

aacttgagttAAAAAAtAGGGaGccctggggcacatacaagaggagtcttccttatcagttaatgctgtatgacactatgta

ttggcccattggctaaaagcccaacttgacaaatggaagatagaatccttgcatActAAAAAGGaGcGGaccgaaagggaag

ctggtgagcaacgacagattcttacgtgcattagctcgcttccggggatctaatagcacgaagcttActAAAAAGGaGcGGa

AgAAgAAAGGttGGG
..|..|||.|..|||
cAAtAAAAcGGcGGG
Challenge Problem

• Find a motif in a sample of

- 20 “random” sequences (e.g. 600 nt long)

- each sequence containing an implanted pattern of length 15,

- each pattern appearing with 4 mismatches as (15,4)-motif.


Regulatory Regions
• Every gene contains a regulatory region (RR) typically stretching 100-1000 bp
upstream of the transcriptional start site

• Located within the RR are the Transcription Factor Binding Sites (TFBS), also
known as motifs, specific for a given transcription factor

• TFs influence gene expression by binding to a specific location in the respective


gene’s regulatory region - TFBS
Transcription Factor Binding Sites

• A TFBS can be located anywhere within the


Regulatory Region.

• TFBS may vary slightly across different regulatory regions since non-
essential bases could mutate
Motifs and Transcriptional Start Sites

ATCCCG gene

TTCCGG gene

ATCCCG gene

ATGCCG gene

ATGCCC gene
Transcription Factors and Motifs
Motif Logo
TGGGGGA
• Motifs can mutate on non important TGAGAGA
bases TGGGGGA
• The five motifs in five different genes TGAGAGA
have mutations in position 3 and 5 TGAGGGA
• Representations called motif logos
illustrate the conserved and variable
regions of a motif
Sequence logos
• In a ‘sequence logo,’ developed by Schneider and Stephens, each
stack is scaled with the information content of the base frequencies
at that position:

where fb,I indicates the frequency of base, b at position i.


Motif Logos: An Example

([Link]
Regulatory Proteins
• Gene X encodes regulatory protein, a.k.a. a transcription
factor (TF)

• The 20 unexpressed genes rely on gene X’s TF to induce


transcription

• A single TF may regulate multiple genes


Identifying Motifs

• Genes are turned on or off by regulatory proteins

• These proteins bind to upstream regulatory regions of genes to either


attract or block an RNA polymerase
• Regulatory protein (TF) binds to a short DNA sequence called a motif
(TFBS)

• So finding the same motif in multiple genes’ regulatory regions


suggests a regulatory relationship amongst those genes
Identifying Motifs: Complications

• We do not know the motif sequence

• We do not know where it is located relative to the genes start

• Motifs can differ slightly from one gene to the next

• How to discern it from “random” motifs?


ROX1 binding sites and sequence motif.
(a) Eight known genomic binding sites in three S. cerevisiae genes.
(b) Degenerate consensus sequence.
(c) (C and D) Frequencies of nucleotides at each position.
(e) Sequence logo showing the frequencies scaled relative to the information
content (measure of conservation) at each position.
(f) Energy normalized logo using relative entropy to adjust for low GC content in S.
cerevisiae.
The Matrix
• A position weight matrix (PWM)
• also called position-specific weight matrix (PSWM)
• also called position-frequency matrix (PFM)
• also called position-specific scoring matrix (PSSM)
• or just matrix
• Alternative to the consensus.
• There is a matrix element for all possible bases at
every position
The Matrix Formats
Information Theory

• Information theory is a branch of applied mathematics involved with the


quantification of information.

• It has been applied to DNA motifs in order to determine the amount of


uncertainty at each position in a site.

• Uncertainty is measured in bits of information, which is on a log2 scale.

• Information is a decrease in uncertainty


Algorithm for greedy profile motif search
• Use P-most probable l-mers to adjust new start positions until we reach the “best”

profile; this will be pronounced as the motif.

• Select random starting positions, then:

1. Create a profile P from the l-mers at these starting positions.

2. Find the P-most probable l-mer a in each sequence and change the starting positions to the starting

positions of a’s.

3. Go to step 1 and re-iterate until we cannot increase the score anymore.


Summary of greedy motif discovery
• Since we choose starting positions randomly, there is little chance that our guess will be

close to an optimal motif, meaning it will take a very long time to find the optimal motif.

• In practice, this algorithm is run many times with the hope that random starting positions

will be close to the optimum solution simply by chance.

• The algorithm may be improved by heuristic knowledge, were approximately we should

start or by more sophisticated statistical techniques, like Gibbs sampling that estimates

the most probable start positions for motifs, where we ought to start our iterative process

of motif discovery.

You might also like