0% found this document useful (0 votes)
21 views27 pages

Protein Structure Prediction Methods

Proteomics projects have increased the number of protein sequences, but experimental structure determination methods are complex, time-consuming, and limited. Protein structure prediction, including secondary and tertiary structure prediction, helps address this gap. Homology or comparative modeling is commonly used for tertiary structure prediction and involves identifying a template, aligning sequences, determining conserved regions, building models, and refining structures. However, modeling accuracy depends on template selection and quality.

Uploaded by

Biju Thomas
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
21 views27 pages

Protein Structure Prediction Methods

Proteomics projects have increased the number of protein sequences, but experimental structure determination methods are complex, time-consuming, and limited. Protein structure prediction, including secondary and tertiary structure prediction, helps address this gap. Homology or comparative modeling is commonly used for tertiary structure prediction and involves identifying a template, aligning sequences, determining conserved regions, building models, and refining structures. However, modeling accuracy depends on template selection and quality.

Uploaded by

Biju Thomas
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd

Proteomics

Genome sequencing projects increase in the number of protein


sequences.

GAP

Experimental methods for structure determination are complex, time-


consuming and limited in their application 1- 3 years
Structure determination X-ray or NMR techniques.

There are many important proteins


Knowledge of protein structure

Protein function prediction Ligand binding Enzyme catalysis


What is the Solution ???????

There are two main areas in protein structure prediction


secondary structure prediction and
tertiary structure prediction
Secondary structure prediction
Structural differences between

globular proteins

&

transmembrane
proteins
Different approaches to predict respective
secondary structural elements
Secondary structure prediction for globular proteins

1. ab initio based method


make use of single sequence
information only

Prediction is based on statistical calculations of the residues of a


single query sequence.

It measures the relative propensity of each amino acid belonging


to a certain secondary structure element (SSE)

The propensity scores are derived from known crystal structures.

Eg: Chou-Fasman and GOR methods.


2. homology based method make use of multiple sequence
alignment information.
evolutionary informations are considered.

ab intio method + information form homologous multiple


sequence alignments.

Incorporation of multiple sequence alignment provides validity to


the prediction as the close protein homologs adopt same
secondary and tertiary structure.

Eg: PHD, PSIPRED, PROF, Jpred


Secondary structure prediction for transmembrane proteins

Functions
signal transduction
cross-membrane transport 30 % of proteins
energy conversion
drug target HMMTOP,
TMHMM, Coils,
Structure is very difficult to resolve.
Multicoil, 2ZIP
Prediction is very important.
Tertiary structure prediction
1. Ab initio methods
2. Comparative or homology modeling
3. Threading
1. Ab initio methods
Calculating co-ordinates for sequences from first principles.

There is some information in the


sequence that provides the
instruction for the proteins to
find their native structures

Little success
More theory
No structure is used as a template. Rosetta
2. Homology Modeling
Predicting the structure of a protein sequence based on its
sequence homology with known structures.

If two proteins share a high


enough sequence similarity,
they are likely to have very
similar three dimensional
structures.

If one of the protein


sequences has a known
structure, then the structure
can be copied to the unknown
Comparative modeling protein with a high degree of
confidence.

Modeller
3. Threading
Falling somewhere between comparative modeling and
ab initio prediction.
There are only small number of
It predicts the structural fold of an unknown protein folds available (<1000),
protein sequence by fitting the sequence into compared to millions of protein
structural database and selecting the best sequences.
fitting fold.
Protein structures tend to be
more conserved than sequences.
Therefore, many proteins can
share similar fold even in the
absence of sequence similarities.

This allowed the development of


computational methods to predict
protein structures beyond
sequence similarities.
<25% pair-wise sequence identity 
Remote homology modeling.

Fold recognition 3D-PSSM, GenTHREADER


Homology Modeling Process
 Template recognition
 Alignment
 Determining structurally conserved regions
 Backbone generation
 Building loops or variable regions
 Conformational search for side chains
 Refinement of structure
 Validating structures
Template Recognition
• First we search the related proteins sequence(templates) to the target
sequence in any structural database of proteins

• The accuracy of model depends on the selection of proper template

• FASTA and BLAST from EMBL-EBI and NCBI can be used

• This gives a probable set of templates but the final one is not yet decided

• After intial aligments and finding structurally conserved regions among


templates, we choose the final template
Determining Structurally Conserved Regions (SCRs)

• When two or more reference protein structures are available

• Establish structural guidelines for the family of proteins under consideration

• First step in building a model protein by homology is determining what regions


are structurally conserved or constant among all the reference proteins

• Target protein is supposed to assume the same conformation in conserved


regions
Structurally Conserved Regions
 SCRs are region in all proteins of a particular family that are nearly
identical in structures.

 Tend to be at inner cores of the proteins

 Usually contains alpha-helices and beta sheets

 No SCR can span more than one secondary structure


Assignment of coordinates within conserved
region

 Once the correspondence between amino acids in the reference and test
sequences has been made, the coordinates for an SCR can be assigned
 The reference proteins' coordinates are used as a basis for this assignment

 Where the side chains of the reference and model proteins are the same at
corresponding locations along the sequence, all the coordinates for the amino
acid are transferred

 Where they differ, the backbone coordinates are transferred , but the side
chain atoms are automatically replaced to preserve the model protein's
residue types
Assignment of coordinates in loop or variable
region

Two main methods

 Finding similar peptide segments in other proteins

 Generating a segment de-novo


Assignment of coordinates in
loop or variable region
Finding similar peptide segment in other proteins

 Advantage: all loops found are guaranteed to have reasonable


internal geometries and conformations

 Disadvantage: may not fit properly into the given model protein’s
framework

In this case, de-novo method is advisable


Refinement of model using Molecular
Mechanics

Many structural artifacts can be introduced while the model protein is


being built

 Substitution of large side chains for small ones

 Strained peptide bonds between segments taken from difference


reference proteins

 Non optimum conformation of loops


Optimisation Approaches

 Energy Minimisation is used to produce a chemically and


conformationally reasonable model protein structure
Two mainly used optimisation algorithms are
 Steepest Descent

 Conjugate Gradients

 Molecular Dynamics is used to explore the conformational space a


molecule could visit
Model Validation

 Every homology model contains [Link] main reasons

 % sequence identity between reference and model

 The number of errors in templates

 Hence it is essential to check the correctness of overall fold/ structure,


errors of localized regions and stereochemical parameters: bond
lengths, angles, geometries
Model Evaluation

 WHAT IF [Link]
 SOV [Link]
 PROVE [Link]
 ANOLEA [Link]
 ERRAT [Link]
 VERIFY3D
[Link]
 BIOTECH [Link]
 ProsaII [Link]
 WHATCHECK [Link]
[Link]/whatcheck/
Challenges
 To model proteins with lower similarities( eg < 30% sequence identity)

 To increase accuracy of models and to make it fully automated

 Improvements may include simulataneous optimization techniques in


side chain modeling and loop modeling

 Developing better optimizers and potential function, which can lead


the model structure away from template towards the correct
structure

 Although comparative modelling needs significant


improvement, it is already a mature technique that can
be used to address many practical problems
Automated Web-Based Homology Modeling

 SWISS Model : [Link]

 WHAT IF : [Link]

 The CPHModels Server : [Link]

 3D Jigsaw : [Link]

 SDSC1 : [Link]

 EsyPred3D : [Link]
Comparative Modeling Server & Program

 COMPOSER
[Link]
ml

 MODELER [Link]

 InsightII [Link]

 SYBYL [Link]
SCR

Loop region

You might also like