Unit 2
Unit 2
Protein tertiary structure prediction is a computational process aimed at determining the three-
dimensional (3D) conformation of a protein based on its amino acid sequence. Understanding
the 3D structure of proteins is essential because their biological functions are inherently tied to
their shape. The tertiary structure involves the folding of the protein into a unique three-
dimensional arrangement, stabilized by various non-covalent interactions, such as hydrogen
bonds, hydrophobic interactions, electrostatic forces, and disulfide bridges.
Predicting the tertiary structure of a protein from its sequence is a complex challenge due to the
vast number of possible configurations the polypeptide chain can adopt. The process involves
various techniques, such as homology modeling, ab initio prediction, and threading.
Recent advancements in artificial intelligence (AI) and machine learning, particularly the
development of models like AlphaFold by DeepMind, have made significant strides in improving
the accuracy of tertiary structure predictions. These innovations have opened up new
possibilities in drug discovery, biotechnology, and understanding protein functions, advancing
our overall understanding of molecular biology.
1. Sequence Similarity: The basic assumption is that proteins with high sequence similarity
in the conserved regions are likely to share structural similarity. Identifying a
homologous protein with a solved 3D structure (template) is crucial for the accuracy of
the model.
2. Conservation of Fold: Protein sequences that have similar functions typically adopt
similar folds. Homologous proteins, even with some variations in their sequences, will
often preserve critical elements of their 3D structure. Thus, structural information can
be inferred from the homologous template.
3. Template Selection: The quality of the model depends on the choice of the template. A
good template is one that has a high sequence identity (typically >30%) with the target
protein. The higher the sequence identity, the more accurate the homology model is
likely to be.
4. Alignment Accuracy: The correct alignment of the target sequence to the template
sequence is key. Misalignments can lead to errors in predicting the 3D structure.
5. Modeling Regions of Uncertainty: Areas of the protein sequence that do not align well
with the template may need to be modeled using specialized techniques, and these
regions may have less structural accuracy in the final model.
o Identify homologous templates: The first step is to search for a template protein
with a known 3D structure. This is done using sequence similarity search tools
such as BLAST (Basic Local Alignment Search Tool) or HMMER (Hidden Markov
Models) to find proteins in structural databases like the Protein Data Bank (PDB).
o Assess template quality: Once templates are identified, their quality and
relevance to the target protein must be assessed. This includes considering the
resolution of the template structure and the sequence identity between the
target and template proteins.
3. Model Building:
Tools: MODELLER, SWISS MODEL, Phyre 2
o Construct the initial model: Using the sequence alignment, a 3D model of the
target protein is built by transferring the coordinates of the conserved regions
from the template. This process involves mapping the backbone of the target
protein onto that of the template and adding side chains to the appropriate
positions. Software tools like MODELLER, SWISS-MODEL, or Phyre2 can perform
this step automatically.
o Insert loop regions: In cases where the target protein sequence deviates
significantly from the template (such as insertions or deletions), the
corresponding regions (loops) need to be modeled separately. Loops are typically
modeled based on the local structural environment and flexibility. Specialized
algorithms or libraries of loop conformations are used for this purpose. eg, rosetta, i-tasser
o Refinement tools: Tools like ROSETTA or MODELLER can refine the model by
iterating through different conformations and selecting the one with the lowest
energy.
o Validation: The quality of the model is assessed using various validation tools to
check for steric clashes, bond angles, and the overall geometry. Tools like
PROCHECK, MolProbity, or Verify3D can help assess the model's structural
quality.
o Comparison with other templates: The model can also be compared with other
homologous proteins to ensure that the predicted structure is reasonable.
o Loop and side-chain validation: Areas with low confidence, such as loops or
surface-exposed side chains, should be further scrutinized for proper geometry
and conformation.
6. Model Usage:
o Functional interpretation: Once validated, the homology model can be used for
functional interpretation, such as understanding protein-ligand interactions, drug
design, or mutation studies.
o Further refinement (if necessary): If the model is to be used for drug design or
structural simulation, additional refinement and docking studies might be
required.
• Useful for unknown proteins: It can provide structural insights for proteins with
unknown 3D structures but homologous to well-studied proteins.
• Handling of novel folds: Homology modeling cannot be used for proteins with no known
homologous structures (novel folds), which requires alternative approaches like ab initio
prediction.
• Geometrical Optimization: The bond angles, bond lengths, and torsion angles may
deviate from their ideal values, resulting in a less stable structure.
• Improving Accuracy: The model might not completely represent the native
conformation, and refinement helps bring the structure closer to the biologically
relevant state.
1. Energy Minimization:
Tools: ROSETTA, MODELLER, GROMACS, Chimera
o What it does: This is the first step of refinement, where the model undergoes a
process of energy optimization to remove steric clashes and improve the overall
structure. The goal is to minimize the total potential energy of the protein model
by adjusting atomic positions.
o Tools: Energy minimization can be done using software tools like MODELLER,
ROSETTA, GROMACS, or Chimera.
o Outcomes: The result is a model with reduced steric clashes and more favorable
bond geometries, making the structure more stable.
o What it does: Loops (regions in the protein that do not align well with the
template structure) are often less accurate in the initial homology model.
Refining these loops to adopt biologically relevant conformations is crucial for
improving the overall structure.
o Tools: Tools like MODELLER or ROSETTA can be used for loop refinement by
sampling possible loop conformations and selecting the most favorable one
based on energy calculations.
o Outcomes: Improved loop regions that better fit the rest of the protein structure,
reducing steric clashes and improving the model’s overall accuracy.
4. Side-Chain Refinement:
Tools: SCWRL, MODELLER, ROSETTA, PyRosetta
o What it does: Side-chain conformations may not always be accurate in the initial
model, especially if the target protein has residues that differ significantly from
the template. Refining side chains improves the accuracy of interactions at the
protein surface and within the protein core.
o Tools: Tools like SCWRL, MODELLER, ROSETTA, or PyRosetta are often used for
side-chain refinement.
o Tools: Tools like Chimera, PyMOL, MODELLER, and ROSETTA offer features for
fine-tuning bond angles and performing additional energy minimization steps.
o What it does: After refinement, it’s essential to validate the model to ensure it is
biologically plausible and of high quality.
o Methods: Validation tools assess the stereochemistry and overall quality of the
model. The primary validation parameters include:
• Purpose: The Ramachandran plot evaluates the backbone dihedral angles (φ and ψ) of a
protein, which determine the protein's secondary structure. It helps identify regions of
the model that might have steric clashes or unusual conformations.
• What it does: This plot visualizes the distribution of φ and ψ angles for each residue in a
protein structure. The plot is divided into favored, allowed, and disallowed regions.
o Favored regions: Most of the residues should ideally lie in these regions,
indicating favorable conformations.
o Allowed regions: Some residues may fall here but still adopt slightly less
favorable conformations.
• Outcomes: A good model should have a majority of its residues in the favored regions,
with minimal residues in the disallowed regions.
• What it does: This process adjusts atomic positions to reduce unfavorable interactions,
like steric clashes between atoms that are too close or incorrect bond angles. The model
is then evaluated based on its energy score, which reflects the stability of the structure.
• Outcomes: After energy minimization, the model should exhibit lower energy states,
with minimized steric clashes and better overall stability.
3. Clashscore:
• Purpose: Clashscore assesses the number and severity of steric clashes in a protein
model. It measures how close atoms are to each other, with the aim of minimizing
atomic overlaps.
• What it does: It quantifies the steric clashes between atoms in the model. A high
clashscore indicates severe atomic collisions, while a low clashscore suggests a more
stable and plausible structure.
• Outcomes: A low clashscore (usually <10) indicates a well-refined model with minimal
steric clashes.
4. MolProbity:
o Rotamer outliers: Detects side-chain positions that deviate from the expected
conformations.
• Tools: MolProbity.
• Outcomes: It provides a detailed report, including scores for steric clashes, geometry,
and other quality parameters. Models with a high MolProbity score are considered more
reliable.
5. Verify3D:
• What it does: It calculates a score that reflects the quality of the 3D structure based on
how well it matches the sequence. A higher score indicates that the model is consistent
with the physical constraints of the protein’s amino acid sequence.
• Tools: Verify3D.
• Outcomes: A Verify3D score close to 1.0 suggests that the model is of high quality, with
minimal deviation from expected structural norms.
• Purpose: ProSA is used to evaluate the overall quality of a protein model by computing
an energy score that reflects how well the structure fits with known protein structures.
• What it does: ProSA generates a Z-score based on the overall energy of the structure. A
Z-score close to 0 indicates that the model is similar to experimentally determined
protein structures, while a significantly negative or positive Z-score suggests potential
problems.
• Tools: ProSA.
• Outcomes: A ProSA Z-score within the range observed for native structures suggests that
the model is reliable. Z-scores that are far from 0 may indicate problems with the model.
• Purpose: RMSD is a measure of the average distance between atoms (usually backbone
atoms) in two superimposed protein structures. It is a useful metric for comparing a
model to the template or other models.
• Outcomes: An RMSD value of less than 2 Å is generally considered acceptable for high-
quality models, while values above 3-4 Å indicate significant differences.
• Purpose: This strategy checks the consistency of torsion angles (φ and ψ) throughout
the protein structure. It helps to validate whether the predicted torsion angles are
consistent with known structural data.
• What it does: Torsion angle validation looks for unusual or incorrect angle distributions,
which may indicate errors in the model.
• Tools: PROCHECK, MolProbity.
• Outcomes: A model with appropriate torsion angles, especially in the favorable regions
of the Ramachandran plot, is more likely to be accurate.
• Purpose: If experimental data (such as NMR or cryo-EM maps) are available, the
homology model can be validated by comparing it to these data.
• What it does: The model can be aligned with experimental electron density maps or
NMR chemical shift data to determine how well it matches the experimental
observations.
• Outcomes: A model that closely fits the experimental data is considered more reliable.
• Purpose: If there are multiple potential templates, comparing models generated from
different templates helps to assess the consistency and reliability of the model.
• What it does: By building models from different templates, it is possible to evaluate how
consistent the models are. Significant discrepancies suggest potential weaknesses in the
model.
1. Template Selection:
• Purpose: Identify the best template (a homologous protein structure) for the target
sequence to align with.
• Methods:
o BLAST: Basic Local Alignment Search Tool, used to search sequence databases for
similar sequences.
o FASTA: Another sequence alignment tool, often used for identifying homologous
templates.
• Advantages:
• Limitations:
2. Sequence Alignment:
• Purpose: Align the target sequence with the template sequence to map equivalent
residues for structure prediction.
• Methods:
o Global Alignment: Aligns the entire target sequence with the template (e.g.,
ClustalW, T-Coffee).
o Local Alignment: Focuses on aligning highly conserved regions, helpful when the
target and template have divergent regions (e.g., Smith-Waterman, BLAST).
• Advantages:
• Limitations:
o Difficulties arise when there are gaps, insertions, or deletions in the alignment.
3. Model Building:
• Purpose: Generate an initial 3D structure based on the aligned sequences of the target
and template.
• Methods:
• Advantages:
• Limitations:
o Structural regions not present in the template (e.g., loops) can be modeled with
lower accuracy.
• Purpose: Optimize the generated model to remove steric clashes and ensure the
structure is geometrically stable.
• Methods:
• Advantages:
• Limitations:
5. Loop Modeling:
• Purpose: Refine or rebuild flexible loop regions that may not be well-modeled in the
initial homology model.
• Methods:
• Advantages:
o Crucial for accurately modeling regions with high flexibility or structural variation.
• Limitations:
o May be less accurate if there are no suitable template loop structures available
for comparison.
6. Side-Chain Refinement:
• Purpose: Improve the orientation of the side chains to reduce steric clashes and
optimize interactions.
• Methods:
• Advantages:
• Limitations:
7. Model Validation:
• Purpose: Assess the quality and accuracy of the generated homology model to ensure it
is reliable for further analysis.
• Methods:
o Verify3D and ProSA: Assess how well the model's 3D structure aligns with its
sequence and known protein data.
o RMSD (Root Mean Square Deviation): Measures the deviation between the
predicted model and the template or experimental structure.
• Advantages:
• Limitations:
• Methods:
• Advantages:
• Limitations:
o If templates are of low quality or significantly divergent, this strategy may lead to
conflicting models.
Concept of Threading
In fold recognition, a template is selected by "threading" a target protein sequence through a library of known
protein structures, evaluating the resulting alignments with scoring functions to identify the template that best
matches the sequence while maintaining structural compatibility, essentially finding the structure that best
"fits" the target sequence based on its amino acid properties and potential interactions; the highest scoring
template is chosen as the predicted fold for the target protein.
Threading methods are designed to predict a protein's tertiary structure by aligning its amino
acid sequence to a set of known 3D protein structures (templates or folds) that may not
necessarily share high sequence similarity with the target. The primary goal is to identify the
best-fitting structural template for the target sequence, even when there is no obvious
sequence homology.
3. Energy Scoring: Each alignment is evaluated based on how well the sequence fits into
the template structure, considering both sequence-structure compatibility and the
physical principles of protein folding.
4. Ranking and Selection: The alignments are ranked according to their scores, and the
best-fitting template is selected. This provides the initial structure of the target protein.
Threading is especially useful for proteins with unknown folds or when sequence-based
methods (e.g., homology modeling) cannot provide reliable results due to low sequence
identity.
Threading methods employ various algorithms to search for the best template, align the target
sequence with it, and evaluate the resulting structure. These algorithms can be broadly
categorized into scoring functions, search strategies, and alignment techniques. Below are
some of the commonly used algorithms and approaches:
1. Scoring Functions:
Scoring functions are used to evaluate how well a given target sequence fits a potential
template. They aim to assess the sequence-template compatibility and the overall stability of
the resulting structure.
• Contact Potential: Measures how well the sequence fits into the template based on the
interactions between pairs of amino acids (or side chains) in the structure. Contact
potentials are derived from known protein structures, and they estimate the likelihood
of two amino acids being in close proximity in a stable fold.
o Examples: Karplus potential, Cohen potential.
• Threading Energy Score: Combines multiple energy terms such as van der Waals
interactions, electrostatic interactions, and hydrophobic interactions to compute a score
for a particular alignment. The score reflects how favorable the sequence-template
combination is based on physical chemistry principles.
• Pairwise Potentials: These capture amino acid residue pair preferences in 3D space.
They allow a scoring function to account for residue-residue interactions in the
structure.
2. Search Strategies:
Threading methods involve searching for the best alignment of the target sequence onto a given
template. This search can be optimized using various strategies, including:
• Exhaustive Search: This method evaluates all possible alignments of the target sequence
onto the template. While it guarantees finding the best possible alignment, it is
computationally expensive and impractical for large proteins or large template
databases.
• Greedy Search: Greedy algorithms try to find a good alignment by making locally
optimal choices at each step, without considering the global optimum. While faster than
exhaustive search, they can sometimes miss the best possible alignment.
• Monte Carlo Simulation: A stochastic search method that uses random sampling to
explore the possible configurations of a target sequence on a template. It evaluates the
quality of each alignment based on the scoring function and iterates toward an optimal
solution by accepting or rejecting new configurations probabilistically.
3. Alignment Techniques:
Threading methods use various techniques to align a sequence with a template structure. These
techniques are central to determining how well a sequence "fits" into a given template.
• Dynamic Programming (DP): A popular method for sequence alignment, especially for
pairwise sequence-template alignment. It optimizes alignment by breaking down the
problem into smaller subproblems and solving them recursively.
pairwise alignment
o Needleman-Wunsch algorithm (global alignment): Finds the optimal global
alignment between two sequences.
4. Template Database:
Threading algorithms often rely on pre-built databases of protein structures to search for the
best templates. These databases contain known protein folds, structural motifs, or entire
protein families and are continuously updated as more experimental structures are solved.
Some of the commonly used template databases are:
5. Hybrid Approaches:
Some threading methods combine threading with other computational approaches to improve
accuracy and overcome some of the inherent limitations of threading alone.
• Goal: To identify potential interacting partners and understand the interface between
two proteins.
• Example: Analyzing the structural interface between a receptor and its ligand to design
inhibitors or modulators of the interaction.
• Goal: To classify proteins into families or folds based on structural similarities and to
understand evolutionary relationships.
• Example: The SCOP (Structural Classification of Proteins) and CATH (Class, Architecture,
Topology, Homologous superfamily) databases use structural comparison to organize
proteins by their folds and structural families, contributing to evolutionary studies.
• Goal: To build accurate protein models when experimental structures are not available.
• Example: Using a known template (e.g., from PDB) to model a protein with unknown
structure and then comparing the resulting model to other similar structures for
validation.
• Goal: To design or screen small molecules (drugs) that interact with specific protein
targets.
• Example: Comparing the open and closed forms of enzymes or receptors to study how
conformational changes facilitate their function.
• Goal: To investigate the arrangement and stability of protein complexes and large multi-
subunit assemblies.
• Goal: To identify the effects of mutations or structural variations on protein stability and
function.
• Example: Comparing the structure of a wild-type protein with a mutant that causes a
genetic disorder (e.g., cystic fibrosis) to understand how the mutation disrupts protein
folding or function.
• Goal: To engineer proteins with desired properties, such as increased stability, altered
specificity, or new functionality.
• Goal: To explore protein folding pathways and understand the causes of protein
misfolding diseases.
• Example: Comparing the amyloid fibril structure of a misfolded protein like prion to its
native state to understand how misfolding contributes to neurodegenerative diseases.
• Goal: To identify biologically relevant sites on a protein, such as active sites, binding
pockets, or regions involved in protein-protein interactions.
• Example: Identifying the active site of a kinase by comparing its structure to other
known kinases.