Proteomics
Genome sequencing projects increase in the number of protein
sequences.
GAP
Experimental methods for structure determination are complex, time-
consuming and limited in their application 1- 3 years
Structure determination X-ray or NMR techniques.
There are many important proteins
Knowledge of protein structure
Protein function prediction Ligand binding Enzyme catalysis
What is the Solution ???????
There are two main areas in protein structure prediction
secondary structure prediction and
tertiary structure prediction
Secondary structure prediction
Structural differences between
globular proteins
&
transmembrane
proteins
Different approaches to predict respective
secondary structural elements
Secondary structure prediction for globular proteins
1. ab initio based method
make use of single sequence
information only
Prediction is based on statistical calculations of the residues of a
single query sequence.
It measures the relative propensity of each amino acid belonging
to a certain secondary structure element (SSE)
The propensity scores are derived from known crystal structures.
Eg: Chou-Fasman and GOR methods.
2. homology based method make use of multiple sequence
alignment information.
evolutionary informations are considered.
ab intio method + information form homologous multiple
sequence alignments.
Incorporation of multiple sequence alignment provides validity to
the prediction as the close protein homologs adopt same
secondary and tertiary structure.
Eg: PHD, PSIPRED, PROF, Jpred
Secondary structure prediction for transmembrane proteins
Functions
signal transduction
cross-membrane transport 30 % of proteins
energy conversion
drug target HMMTOP,
TMHMM, Coils,
Structure is very difficult to resolve.
Multicoil, 2ZIP
Prediction is very important.
Tertiary structure prediction
1. Ab initio methods
2. Comparative or homology modeling
3. Threading
1. Ab initio methods
Calculating co-ordinates for sequences from first principles.
There is some information in the
sequence that provides the
instruction for the proteins to
find their native structures
Little success
More theory
No structure is used as a template. Rosetta
2. Homology Modeling
Predicting the structure of a protein sequence based on its
sequence homology with known structures.
If two proteins share a high
enough sequence similarity,
they are likely to have very
similar three dimensional
structures.
If one of the protein
sequences has a known
structure, then the structure
can be copied to the unknown
Comparative modeling protein with a high degree of
confidence.
Modeller
3. Threading
Falling somewhere between comparative modeling and
ab initio prediction.
There are only small number of
It predicts the structural fold of an unknown protein folds available (<1000),
protein sequence by fitting the sequence into compared to millions of protein
structural database and selecting the best sequences.
fitting fold.
Protein structures tend to be
more conserved than sequences.
Therefore, many proteins can
share similar fold even in the
absence of sequence similarities.
This allowed the development of
computational methods to predict
protein structures beyond
sequence similarities.
<25% pair-wise sequence identity
Remote homology modeling.
Fold recognition 3D-PSSM, GenTHREADER
Homology Modeling Process
Template recognition
Alignment
Determining structurally conserved regions
Backbone generation
Building loops or variable regions
Conformational search for side chains
Refinement of structure
Validating structures
Template Recognition
• First we search the related proteins sequence(templates) to the target
sequence in any structural database of proteins
• The accuracy of model depends on the selection of proper template
• FASTA and BLAST from EMBL-EBI and NCBI can be used
• This gives a probable set of templates but the final one is not yet decided
• After intial aligments and finding structurally conserved regions among
templates, we choose the final template
Determining Structurally Conserved Regions (SCRs)
• When two or more reference protein structures are available
• Establish structural guidelines for the family of proteins under consideration
• First step in building a model protein by homology is determining what regions
are structurally conserved or constant among all the reference proteins
• Target protein is supposed to assume the same conformation in conserved
regions
Structurally Conserved Regions
SCRs are region in all proteins of a particular family that are nearly
identical in structures.
Tend to be at inner cores of the proteins
Usually contains alpha-helices and beta sheets
No SCR can span more than one secondary structure
Assignment of coordinates within conserved
region
Once the correspondence between amino acids in the reference and test
sequences has been made, the coordinates for an SCR can be assigned
The reference proteins' coordinates are used as a basis for this assignment
Where the side chains of the reference and model proteins are the same at
corresponding locations along the sequence, all the coordinates for the amino
acid are transferred
Where they differ, the backbone coordinates are transferred , but the side
chain atoms are automatically replaced to preserve the model protein's
residue types
Assignment of coordinates in loop or variable
region
Two main methods
Finding similar peptide segments in other proteins
Generating a segment de-novo
Assignment of coordinates in
loop or variable region
Finding similar peptide segment in other proteins
Advantage: all loops found are guaranteed to have reasonable
internal geometries and conformations
Disadvantage: may not fit properly into the given model protein’s
framework
In this case, de-novo method is advisable
Refinement of model using Molecular
Mechanics
Many structural artifacts can be introduced while the model protein is
being built
Substitution of large side chains for small ones
Strained peptide bonds between segments taken from difference
reference proteins
Non optimum conformation of loops
Optimisation Approaches
Energy Minimisation is used to produce a chemically and
conformationally reasonable model protein structure
Two mainly used optimisation algorithms are
Steepest Descent
Conjugate Gradients
Molecular Dynamics is used to explore the conformational space a
molecule could visit
Model Validation
Every homology model contains [Link] main reasons
% sequence identity between reference and model
The number of errors in templates
Hence it is essential to check the correctness of overall fold/ structure,
errors of localized regions and stereochemical parameters: bond
lengths, angles, geometries
Model Evaluation
WHAT IF [Link]
SOV [Link]
PROVE [Link]
ANOLEA [Link]
ERRAT [Link]
VERIFY3D
[Link]
BIOTECH [Link]
ProsaII [Link]
WHATCHECK [Link]
[Link]/whatcheck/
Challenges
To model proteins with lower similarities( eg < 30% sequence identity)
To increase accuracy of models and to make it fully automated
Improvements may include simulataneous optimization techniques in
side chain modeling and loop modeling
Developing better optimizers and potential function, which can lead
the model structure away from template towards the correct
structure
Although comparative modelling needs significant
improvement, it is already a mature technique that can
be used to address many practical problems
Automated Web-Based Homology Modeling
SWISS Model : [Link]
WHAT IF : [Link]
The CPHModels Server : [Link]
3D Jigsaw : [Link]
SDSC1 : [Link]
EsyPred3D : [Link]
Comparative Modeling Server & Program
COMPOSER
[Link]
ml
MODELER [Link]
InsightII [Link]
SYBYL [Link]
SCR
Loop region