Using Patterns and
Profiles as Motif Descriptors
Random Sequence
atgaccgggatactgataccgtatttggcctaggcgtacacattagataaacgtatgaagtacgttagactcggcgccgccg
acccctattttttgagcagatttagtgacctggaaaaaaaatttgagtacaaaacttttccgaatactgggcataaggtaca
tgagtatccctgggatgacttttgggaacactatagtgctctcccgatttttgaatatgtaggatcattcgccagggtccga
gctgagaattggatgaccttgtaagtgttttccacgcaatcgcgaaccaacgcggacccaaaggcaagaccgataaaggaga
tcccttttgcggtaatgtgccgggaggctggttacgtagggaagccctaacggacttaatggcccacttagtccacttatag
gtcaatcatgttcttgtgaatggatttttaactgagggcatagaccgcttggcgcacccaaattcagtgtgggcgagcgcaa
cggttttggcccttgttagaggcccccgtactgatggaaactttcaattatgagagagctaatctatcgcgtgcgtgttcat
aacttgagttggtttcgaaaatgctctggggcacatacaagaggagtcttccttatcagttaatgctgtatgacactatgta
ttggcccattggctaaaagcccaacttgacaaatggaagatagaatccttgcatttcaacgtatgccgaaccgaaagggaag
ctggtgagcaacgacagattcttacgtgcattagctcgcttccggggatctaatagcacgaagcttctgggtactgatagca
Sequences like (AAAAAAAGGGGGGG)
repeats
atgaccgggatactgataaaaaaaagggggggggcgtacacattagataaacgtatgaagtacgttagactcggcgccgccg
acccctattttttgagcagatttagtgacctggaaaaaaaatttgagtacaaaacttttccgaataaaaaaaaaggggggga
tgagtatccctgggatgacttaaaaaaaagggggggtgctctcccgatttttgaatatgtaggatcattcgccagggtccga
gctgagaattggatgaaaaaaaagggggggtccacgcaatcgcgaaccaacgcggacccaaaggcaagaccgataaaggaga
tcccttttgcggtaatgtgccgggaggctggttacgtagggaagccctaacggacttaataaaaaaaagggggggcttatag
gtcaatcatgttcttgtgaatggatttaaaaaaaaggggggggaccgcttggcgcacccaaattcagtgtgggcgagcgcaa
cggttttggcccttgttagaggcccccgtaaaaaaaagggggggcaattatgagagagctaatctatcgcgtgcgtgttcat
aacttgagttaaaaaaaagggggggctggggcacatacaagaggagtcttccttatcagttaatgctgtatgacactatgta
ttggcccattggctaaaagcccaacttgacaaatggaagatagaatccttgcataaaaaaaagggggggaccgaaagggaag
ctggtgagcaacgacagattcttacgtgcattagctcgcttccggggatctaatagcacgaagcttaaaaaaaaggggggga
Where is the Implanted ?
atgaccgggatactgatAAAAAAAAGGGGGGGggcgtacacattagataaacgtatgaagtacgttagactcggcgccgccg
acccctattttttgagcagatttagtgacctggaaaaaaaatttgagtacaaaacttttccgaataAAAAAAAAGGGGGGGa
tgagtatccctgggatgacttAAAAAAAAGGGGGGGtgctctcccgatttttgaatatgtaggatcattcgccagggtccga
gctgagaattggatgAAAAAAAAGGGGGGGtccacgcaatcgcgaaccaacgcggacccaaaggcaagaccgataaaggaga
tcccttttgcggtaatgtgccgggaggctggttacgtagggaagccctaacggacttaatAAAAAAAAGGGGGGGcttatag
gtcaatcatgttcttgtgaatggatttAAAAAAAAGGGGGGGgaccgcttggcgcacccaaattcagtgtgggcgagcgcaa
cggttttggcccttgttagaggcccccgtAAAAAAAAGGGGGGGcaattatgagagagctaatctatcgcgtgcgtgttcat
aacttgagttAAAAAAAAGGGGGGGctggggcacatacaagaggagtcttccttatcagttaatgctgtatgacactatgta
ttggcccattggctaaaagcccaacttgacaaatggaagatagaatccttgcatAAAAAAAAGGGGGGGaccgaaagggaag
ctggtgagcaacgacagattcttacgtgcattagctcgcttccggggatctaatagcacgaagcttAAAAAAAAGGGGGGGa
Introduction
• Among the various databases dedicated to the identification of protein families and
domains, PROSITE is the first one created.
• The close relationship of PROSITE with the SWISS-PROT protein database allows the
evaluation of the sensitivity and specificity of the PROSITE motifs and their periodic
reviewing.
• The motif descriptors used in PROSITE are either patterns or profiles, which are
derived from multiple alignments of homologous sequences.
PROSITE PATTERNS
• In some cases the sequence of an unknown protein is too distantly related to any protein
of known structure to detect its resemblance by pairwise sequence alignment.
• However, relationships can be revealed by the occurrence in its sequence of a particular
cluster of residue types, which is variously known as a pattern, motif, signature or
fingerprint.
• These motifs, typically around 10 to 20 amino acids in length.
• Because of specific residues and regions (position[s]) proved to be important to the
biological function of a group of proteins.
• Conserved in both structure and sequence during evolution.
These biologically significant regions or residues are generally:
• Enzyme catalytic sites (see figure).
• Prostethic group attachment sites (heme,
pyridoxal-phosphate, biotin, etc.).
• Amino acids involved in binding a metal ion.
• Cysteines involved in disulphide bonds.
• Regions involved in binding a molecule (ADP/ATP,
GDP/GTP, calcium, DNA,etc.) or another protein.
• As the sequence of biologically meaningful motifs is evolutionarily conserved, a multiple
alignment of them can be reduced to a consensus expression called a regular expression or
pattern.
• Each position of such a pattern can be occupied by any residue from a specified set of
acceptable residues, and in addition can be repeated a variable number of times within a
specified range.
• At strictly conserved positions only one particular amino acid is accepted, whereas at other
positions several amino acids with similar physicochemical properties can be accepted.
• A regular expression is qualitative; it either does match or does not. There is no
threshold above which it consider the match as statistically significant.
• It should be noticed that some families or domains are defined not just by one pattern
but by the co-occurrence of two or more patterns of low specificity.
• The presence of just one of these patterns is not sufficient to assign a protein to a
particular family and/or domain. However, the simultaneous occurrence of linked
patterns gives good confidence that the matched protein belongs to the set being
considered.
PROSITE PROFILES
• Profiles are quantitative motif descriptors providing numerical weights for each
possible match or mismatch between a sequence residue and a profile position.
• A mismatch at a highly conserved position can thus be accepted provided that
the rest of the sequence displays a sufficiently high level of similarity.
Rules for prosite pattern (pattern syntax)
• Pattern syntax
• The standard IUPAC one letter code for the amino acids is used in PROSITE.
• The symbol 'x' is used for a position where any amino acid is accepted.
• Ambiguities are indicated by listing the acceptable amino acids for a given position, between square
brackets '[ ]'. For example: [ALT] stands for Ala or Leu or Thr.
• Ambiguities are also indicated by listing between a pair of curly brackets '{ }' the amino acids that are not
accepted at a given position. For example: {AM} stands for all any amino acid except Ala and Met.
• Each element in a pattern is separated from its neighbor by a '-'.
• Repetition of an element of the pattern can be indicated by following that element with a numerical value
or, if it is a gap ('x'), by a numerical range between parentheses.
Examples:
• x(3) corresponds to x-x-x
• x(2,4) corresponds to x-x or x-x-x or x-x-x-x
• A(3) corresponds to A-A-A
• When a pattern is restricted to either the N- or C-terminal of a sequence, that pattern respectively starts
with a '<' symbol or ends with a '>' symbol.
In some rare cases (e.g. PS00267 or PS00539), '>' can also occur inside square brackets for the C-terminal
element. 'F-[GSTV]-P-R-L-[G>]' is equivalent to 'F-[GSTV]-P-R-L-G' or 'F-[GSTV]-P-R-L>'.
[Link]
Where is the Motif???
atgaccgggatactgatagaagaaaggttgggggcgtacacattagataaacgtatgaagtacgttagactcggcgccgccg
acccctattttttgagcagatttagtgacctggaaaaaaaatttgagtacaaaacttttccgaatacaataaaacggcggga
tgagtatccctgggatgacttaaaataatggagtggtgctctcccgatttttgaatatgtaggatcattcgccagggtccga
gctgagaattggatgcaaaaaaagggattgtccacgcaatcgcgaaccaacgcggacccaaaggcaagaccgataaaggaga
tcccttttgcggtaatgtgccgggaggctggttacgtagggaagccctaacggacttaatataataaaggaagggcttatag
gtcaatcatgttcttgtgaatggatttaacaataagggctgggaccgcttggcgcacccaaattcagtgtgggcgagcgcaa
cggttttggcccttgttagaggcccccgtataaacaaggagggccaattatgagagagctaatctatcgcgtgcgtgttcat
aacttgagttaaaaaatagggagccctggggcacatacaagaggagtcttccttatcagttaatgctgtatgacactatgta
ttggcccattggctaaaagcccaacttgacaaatggaagatagaatccttgcatactaaaaaggagcggaccgaaagggaag
ctggtgagcaacgacagattcttacgtgcattagctcgcttccggggatctaatagcacgaagcttactaaaaaggagcgga
Sequence motifs are short, recurring patterns in DNA that are presumed to have a biological function
Finding (15,4) Motif is Difficult?
atgaccgggatactgatAgAAgAAAGGttGGGggcgtacacattagataaacgtatgaagtacgttagactcggcgccgccg
acccctattttttgagcagatttagtgacctggaaaaaaaatttgagtacaaaacttttccgaatacAAtAAAAcGGcGGGa
tgagtatccctgggatgacttAAAAtAAtGGaGtGGtgctctcccgatttttgaatatgtaggatcattcgccagggtccga
gctgagaattggatgcAAAAAAAGGGattGtccacgcaatcgcgaaccaacgcggacccaaaggcaagaccgataaaggaga
tcccttttgcggtaatgtgccgggaggctggttacgtagggaagccctaacggacttaatAtAAtAAAGGaaGGGcttatag
gtcaatcatgttcttgtgaatggatttAAcAAtAAGGGctGGgaccgcttggcgcacccaaattcagtgtgggcgagcgcaa
cggttttggcccttgttagaggcccccgtAtAAAcAAGGaGGGccaattatgagagagctaatctatcgcgtgcgtgttcat
aacttgagttAAAAAAtAGGGaGccctggggcacatacaagaggagtcttccttatcagttaatgctgtatgacactatgta
ttggcccattggctaaaagcccaacttgacaaatggaagatagaatccttgcatActAAAAAGGaGcGGaccgaaagggaag
ctggtgagcaacgacagattcttacgtgcattagctcgcttccggggatctaatagcacgaagcttActAAAAAGGaGcGGa
AgAAgAAAGGttGGG
..|..|||.|..|||
cAAtAAAAcGGcGGG
Challenge Problem
• Find a motif in a sample of
- 20 “random” sequences (e.g. 600 nt long)
- each sequence containing an implanted pattern of length 15,
- each pattern appearing with 4 mismatches as (15,4)-motif.
Regulatory Regions
• Every gene contains a regulatory region (RR) typically stretching 100-1000 bp
upstream of the transcriptional start site
• Located within the RR are the Transcription Factor Binding Sites (TFBS), also
known as motifs, specific for a given transcription factor
• TFs influence gene expression by binding to a specific location in the respective
gene’s regulatory region - TFBS
Transcription Factor Binding Sites
• A TFBS can be located anywhere within the
Regulatory Region.
• TFBS may vary slightly across different regulatory regions since non-
essential bases could mutate
Motifs and Transcriptional Start Sites
ATCCCG gene
TTCCGG gene
ATCCCG gene
ATGCCG gene
ATGCCC gene
Transcription Factors and Motifs
Motif Logo
TGGGGGA
• Motifs can mutate on non important TGAGAGA
bases TGGGGGA
• The five motifs in five different genes TGAGAGA
have mutations in position 3 and 5 TGAGGGA
• Representations called motif logos
illustrate the conserved and variable
regions of a motif
Sequence logos
• In a ‘sequence logo,’ developed by Schneider and Stephens, each
stack is scaled with the information content of the base frequencies
at that position:
where fb,I indicates the frequency of base, b at position i.
Motif Logos: An Example
([Link]
Regulatory Proteins
• Gene X encodes regulatory protein, a.k.a. a transcription
factor (TF)
• The 20 unexpressed genes rely on gene X’s TF to induce
transcription
• A single TF may regulate multiple genes
Identifying Motifs
• Genes are turned on or off by regulatory proteins
• These proteins bind to upstream regulatory regions of genes to either
attract or block an RNA polymerase
• Regulatory protein (TF) binds to a short DNA sequence called a motif
(TFBS)
• So finding the same motif in multiple genes’ regulatory regions
suggests a regulatory relationship amongst those genes
Identifying Motifs: Complications
• We do not know the motif sequence
• We do not know where it is located relative to the genes start
• Motifs can differ slightly from one gene to the next
• How to discern it from “random” motifs?
ROX1 binding sites and sequence motif.
(a) Eight known genomic binding sites in three S. cerevisiae genes.
(b) Degenerate consensus sequence.
(c) (C and D) Frequencies of nucleotides at each position.
(e) Sequence logo showing the frequencies scaled relative to the information
content (measure of conservation) at each position.
(f) Energy normalized logo using relative entropy to adjust for low GC content in S.
cerevisiae.
The Matrix
• A position weight matrix (PWM)
• also called position-specific weight matrix (PSWM)
• also called position-frequency matrix (PFM)
• also called position-specific scoring matrix (PSSM)
• or just matrix
• Alternative to the consensus.
• There is a matrix element for all possible bases at
every position
The Matrix Formats
Information Theory
• Information theory is a branch of applied mathematics involved with the
quantification of information.
• It has been applied to DNA motifs in order to determine the amount of
uncertainty at each position in a site.
• Uncertainty is measured in bits of information, which is on a log2 scale.
• Information is a decrease in uncertainty
Algorithm for greedy profile motif search
• Use P-most probable l-mers to adjust new start positions until we reach the “best”
profile; this will be pronounced as the motif.
• Select random starting positions, then:
1. Create a profile P from the l-mers at these starting positions.
2. Find the P-most probable l-mer a in each sequence and change the starting positions to the starting
positions of a’s.
3. Go to step 1 and re-iterate until we cannot increase the score anymore.
Summary of greedy motif discovery
• Since we choose starting positions randomly, there is little chance that our guess will be
close to an optimal motif, meaning it will take a very long time to find the optimal motif.
• In practice, this algorithm is run many times with the hope that random starting positions
will be close to the optimum solution simply by chance.
• The algorithm may be improved by heuristic knowledge, were approximately we should
start or by more sophisticated statistical techniques, like Gibbs sampling that estimates
the most probable start positions for motifs, where we ought to start our iterative process
of motif discovery.