Bioinformatics
— Unit V: Protein Structure Classification—
Dr. Chandra Mohan D
Assistant Professor
Computer Science and Engineering Group
Indian Institute of Information Technology, Sri City
If you know your own DNA sequence than you know every thing about your self
February 10, 2025 Bioinformatics 1
Outline
Introduction (11-03-2023)
Secondary Structure elements
Ramachandran plot
Propensity
Secondary and 3D structure prediction
Visualization tools
Structural classification
February 10, 2025 Bioinformatics 2
Introduction
Proteins perform most essential biological and chemical functions in
a cell
They play important roles in structural, enzymatic, transport, and
regulatory functions
The protein functions are strictly determined by their structures
Therefore, protein structural analysis is an essential element of
bioinformatics
The building blocks of proteins are 20 naturally occurring amino
acids, small molecules
Amino acids contain a free amino group (NH2) and a free carboxyl
group (COOH)
Both of these groups are linked to a
central carbon (Cα), which is attached
to a hydrogen and a side chain group (R)
February 10, 2025 Bioinformatics 3
Introduction
Amino acids differ only by the side chain R group
The chemical reactivities of the R groups determine the specific
properties of the amino acids
Amino acids can be grouped into several categories based on the
chemical and physical properties of the side chains (R), such as
size and affinity for water
According to these properties, the side chain groups can be divided
into small, large, hydrophobic, and hydrophilic categories
Within the hydrophobic set of amino acids, they can be further
divided into Aliphatic and Aromatic
Aliphatic side chains are linear
hydrocarbon chains and
Aromatic side chains are cyclic rings
February 10, 2025 Bioinformatics 4
Introduction
Within the hydrophilic set, amino
acids can be subdivided into polar
and charged
Charged amino acids can be either
positively charged (basic) or
negatively charged (acidic)
The R side chains also play an
important structural role
Of particular interest within the twenty
amino acids are glycine and proline
Glycine, the smallest amino acid, has
a hydrogen atom as the R group
February 10, 2025 Bioinformatics 5
Introduction
It can therefore adopt more flexible conformations that are not
possible for other amino acids
Glycine increase local flexibility in structures
Proline is on the other extreme of flexibility
Its side chain forms a bond with its own backbone amino group,
causing it to be cyclic
The cyclic conformation makes it very rigid,
Unable to occupy many of the main chain conformations adopted
by other amino acids
Cysteine, which can react with another cysteine to form a cross-link
that can stabilize the protein structure
February 10, 2025 Bioinformatics 6
Protein structure Terminology
Certain amino acids are subject to modifications after a protein is
translated in a cell
This is called posttranslational modification
The peptide formation involves two amino acids covalently joined
together between
the carboxyl group of one amino acid and the amino group of
another
This reaction is a condensation reaction involving removal of
elements of water from the two molecules
The resulting product is called a dipeptide
February 10, 2025 Bioinformatics 7
Protein structure Terminology
The newly formed covalent bond connecting the two amino acids is
called a peptide bond
Once an amino acid is incorporated into a peptide, it becomes an amino
acid residue
Multiple amino acids can be joined together to form a longer chain of
amino acid polymer
A linear polymer of more than fifty amino acid residues is referred to as
a polypeptide
A polypeptide, also called a protein, has a well-defined 3D arrangement
On the other hand, a polymer with fewer than fifty residues is usually
called a peptide without a well-defined 3D structure
February 10, 2025 Bioinformatics 8
Protein structure Terminology
The residues in a peptide or polypeptide are numbered
beginning with the residue containing the amino group, referred to
as the N-terminus, and
ending with the residue containing the carboxyl group, known as the
C-terminus
The actual sequence of amino acid residues in a polypeptide determines
its ultimate structure and function
The atoms involved in forming the peptide bond are referred to as the
backbone atoms
February 10, 2025 Bioinformatics 9
Protein structure Terminology
They are the nitrogen of the amino group, the α carbon to which the
side chain is attached and carbon of the carbonyl group
The polypeptide chain is first assembled on the ribosome using the
codon sequence on mRNA as a template
The resulting linear chain forms secondary structures through
The formation of hydrogen bonds between amino acids in the chain
Through further interactions among amino acid side groups, these 2D
structures then fold into a 3D structure
February 10, 2025 Bioinformatics 10
Protein structure Terminology
Proteins are chains of amino acids joined by peptide bonds
Many conformations of the chain
are possible due to the rotation of
the chain about each Cα atom
These conformational variations
are responsible for differences in
the 3D structures of proteins
Each amino acid in the chain is polar,
It has separated positive and negatively charged regions with a
chemically free C=O group, which can act as a hydrogen bond
acceptor, and an NH group, which can act as a hydrogen bond
donor
These groups interact in protein structures
February 10, 2025 Bioinformatics 11
Protein structure Terminology
Dihedral Angels:
A peptide bond is actually a partial double bond owing to shared
electrons between O=C–N atoms
The rigid double bond structure forces atoms associated with the
peptide bond to lie in the same plane, called the peptide plane
Because of the planar nature of the peptide bond and the size of the R
groups,
There are considerable restrictions on the rotational freedom
The angle of rotation about the bond is referred to as the dihedral angle
(also called the Torsional angle)
For a peptide unit, the atoms linked to the peptide bond can be moved
to a certain extent by the rotation of two bonds flanking the peptide
bond
This is measured by two dihedral angles
February 10, 2025 Bioinformatics 12
Protein structure Terminology
Dihedral Angels:
One is the dihedral angle along the N–Cα bond, which is defined as phi
(φ); and
The other is the angle along the Cα–C bond, which is called psi (ψ)
Various combinations of φ and ψ angles allow the proteins to fold in
many different ways
February 10, 2025 Bioinformatics 13
Protein Secondary Structures
Secondary Structures:
Regular recurring arrangements in space of adjacent amino acid
residues in a polypeptide chain
In these 2D structures, regular patterns of H bonds are formed between
neighboring amino acids, and the amino acids have similar ϕ and ψ
angles
The formation of these
structures neutralizes
the polar groups on
each amino acid
The secondary structures
are tightly packed in the
protein core in a hydrophobic environment
February 10, 2025 Bioinformatics 14
Protein Secondary Structures
Types of Secondary Structures:
There are four commonly occurring 2D structures
1. α-helices
2. β-strands (sheets)
3. Turns (bends)
4. Coil (irregular)
1. α Helix:
The α helix is the most abundant type of secondary
structure in proteins
The helix has 3.6 amino acids per turn with an H bond formed between
every fourth residue
The helix has 5.4Ӑ per turn.
The average length is 10 amino acids (3 turns)
February 10, 2025 Bioinformatics 15
Protein Secondary Structures
The backbone of the chain is shown in red, the Cα atoms, the C=O
and NH groups are shown in blue, yellow and green
respectively
The alignment of the H bonds creates a dipole moment
for the helix with a resulting partial positive charge at
the amino end of the helix
Because this region has free NH2 groups, it will
interact with negatively charged groups such as
phosphates
In the α helix, note that each C=O group at amino acid
position n in the sequence is hydrogen-bonded with the
NH group at position n + 4
February 10, 2025 Bioinformatics 16
Protein Secondary Structures
α Helix:
The helix is usually right-handed, but short sections of
3–5 amino acids of left-handed helices occur occasionally
The average ϕ and ψ angles of the amino acids in
The right-handed helix are approximately 60° and 40°,
respectively
The R side chains of the amino acids are on the outside
of the helix
February 10, 2025 Bioinformatics 17
Protein Secondary Structures
β Sheets:
β Sheets are formed by H bonds between an average of 5–10
consecutive amino acids in one portion of the chain with another
5–10 farther down the chain
The interacting regions may be adjacent,
with a short loop in between, or far apart,
with other structures in between
Every chain may run in the same direction
to form a parallel sheet, every other chain
may run in the reverse chemical direction
to form an antiparallel sheet, or the chains
may be parallel and antiparallel to form a
mixed sheet
February 10, 2025 Bioinformatics 18
Protein Secondary Structures
β Sheets:
Each amino acid in the interior strands of the sheet forms two H bonds
with neighboring amino acids, whereas
Each amino acid on the outside strands forms only one bond with an
interior strand
Coils and Loops:
There are also local structures that do not belong to regular secondary
structures (α-helices and β-strands)
The irregular structures are coils or loops
The loops are often characterized by sharp turns or hairpin-like
structures
If the connecting regions are completely irregular, they belong to
random coils
February 10, 2025 Bioinformatics 19
Protein Secondary Structures
Coils and Loops:
Residues in the loop or coil regions tend to be charged, polar and
located on the surface of the protein structure
They are often the evolutionarily variable regions where mutations,
deletions, and insertions frequently occur
They can be functionally significant because these locations are often
the active sites of proteins
Coiled Coils
Coiled coils are a special type of super secondary structure
characterized by a bundle of two or more α-helices wrapping
around each other
The helices forming coiled coils have a unique pattern of
hydrophobicity, which repeats every seven residues (five hydrophobic
and two hydrophilic)
February 10, 2025 Bioinformatics 20
Protein Secondary Structures
Turns:
One-third of all residues in globular proteins contained in turns that
serve to reverse the direction of the polypeptide chain
Turn contains the hydrogen bond between the carbonyl oxygen of
residue i and the amide nitrogen of i+3
There are three types of turns, named I, II, III
Type I occur more frequently (2-3 times faster than type-II)
The backbone dihedral angels of the residue are (-60,-30) and (-90,0)
of residues of i+1 and i+2 of type-I turn
Prolin is often found in position i+1 in type-I
turns as its phi angle is restricted to -60
February 10, 2025 Bioinformatics 21
Protein Secondary Structures
Turns:
The backbone dihedral angels of the residue are (-60,120) and (80,0) of
residues of i+1 and i+2 of type-II turn
Glycine is favoured in this position i+1 in type-II as it is requires a
positive phi value
The backbone dihedral angels of the residue
are (-60,-30) and (-60,-30) of residues of i+1
and i+2 of classical type-III turn
February 10, 2025 Bioinformatics 22
Ramachandran Plot
G.N Ramachandran used computer models of small polypeptides to
systematically vary phi and psi with the objective of finding stable
conformations
He plots the phi value on the X-axis and the psi value on the Y-axis
Plotting the torsional angles in this way graphically shows which
combination of angles are possible
February 10, 2025 Bioinformatics 23
Ramachandran Plot
In the diagram the white areas correspond to conformations where
atoms in polypeptide come closer than the sum of their van der waals
radii
These regions are sterically
disallowed for all amino acids
except glycine which is unique
in that it lacks a side chain
The red regions correspond to
conformations where there are no
steric clashes i.e these are the
allowed regions namely the
α- helical and β-sheet
conformations
February 10, 2025 Bioinformatics 24
Ramachandran Plot
The yellow area shows the allowed regions if slightly shorter van der
waals radi are used in the calculation, i.e the atoms are allowed to
come a little closer together
This bring out an additional region which corresponds to the left
handed α-helix
Glycine has no side chain and therefore can adapt phi and psi angels
in all four quadrants of the Ramachandran plot
This is the Ramachandran plot for
pyruvate kinase (glycolysis) with
all amino acids accounted for except
Glycine
February 10, 2025 Bioinformatics 25
Propensity Value
February 10, 2025 Bioinformatics 26
Propensity Value
February 10, 2025 Bioinformatics 27
Propensity Value
February 10, 2025 Bioinformatics 28
Protein Structure Prediction
Classes of Protein Structure
Four principal classes of protein structure were recognized based on
the types and arrangements of secondary structural elements
Class α comprises a bundle of α helices connected by loops on the
surface of the proteins
Structure of α class proteins. (A) Diagram showing a-helical pattern of
this class. a helices are red cylinders, and black lines are loops.
(B) Example of the class, hemoglobin, PDB file 3hhb displayed
using Rasmol, using ribbons display
February 10, 2025
and group color
Bioinformatics 29
Protein Structure Prediction
Classes of Protein Structure
Class β comprises antiparallel β sheets, usually two sheets in close
contact forming a sandwich
Alternatively, a sheet can twist into a barrel with the first and last
strands touching
Examples are enzymes, transport proteins, antibodies, and virus coat
proteins such as neuraminidase
February 10, 2025 Bioinformatics 30
Protein Structure Prediction
Classes of Protein Structure
Class α/β comprises mainly parallel β sheets with intervening α
helices, but may also have mixed β sheets
In addition to forming a sheet in some proteins in this class, as
illustrated below, in others parallel β strands may form into a barrel
structure that is surrounded by α
helices
This class of proteins includes
many metabolic enzymes
February 10, 2025 Bioinformatics 31
Protein Structure Prediction
Classes of Protein Structure
Class α+β comprises mainly segregated α helices and antiparallel β
sheets
February 10, 2025 Bioinformatics 32
Protein Structure Prediction
Classes of Protein Structure
Multidomain (α and β) proteins comprise domains representing more than one of the
above four classes
Membrane and cell-surface proteins and peptides excluding proteins of the immune
system comprise this class
Diagram showing typical
arrangement of membrane
-traversing, hydrophobic
a helices (red)
Membrane bilayer shown
as green
February lines
10, 2025 Bioinformatics 33
Protein Tertiary Structure
The overall packing and arrangement of secondary structures form
the tertiary structure of a protein ([Link]
v=piXHivrTT-E)
The tertiary structure can come in various forms but is generally
classified as either globular or membrane proteins
The former exists in solvents through hydrophilic interactions with
solvent molecules;
The latter exists in membrane lipids and is stabilized through
hydrophobic interactions with the lipid molecules
Membrane proteins are long, strand-like proteins that are insoluble in
water, weak acids, and weak bases
Globular proteins have a spherical shape and are soluble in water,
acids, and bases
February 10, 2025 Bioinformatics 34
Protein Tertiary Structure
Globular Proteins
These are usually soluble and surrounded by water molecules
They tend to have an overall compact structure of spherical shape with
polar or hydrophilic residues on the surface and hydrophobic residues in
the core
Such an arrangement is energetically
favorable because
It minimizes contacts with water by
hydrophobic residues in the core and
maximizes interactions with water by
surface polar and charged residues
Examples: enzymes, cytokines, and protein hormones
February 10, 2025 Bioinformatics 35
Protein Tertiary Structure
Integral Membrane Proteins
Membrane proteins exist in lipid bilayers of cell
membranes
Because they are surrounded by lipids, the
exterior of the proteins spanning the membrane
must be very hydrophobic to be stable
Most typical transmembrane segments are
α-helices
Occasionally, for some bacterial periplasmic
membrane proteins, they are composed of
β-strands
Examples of membrane proteins are rhodopsins,
cytochrome c oxidase, and ion channel proteins
February 10, 2025 Bioinformatics 36
Protein Structure Visualization, Comparison,
Classification
Protein Structure Visualization:
With the development of computer hardware and software technology,
Sophisticated computer graphics programs have been developed
for visualizing and manipulating complicated 3D structures
The computer graphics help to analyze and compare protein structures
to gain insight to functions of the proteins
PDB data file for a protein structure contains only x, y, and z
coordinates of atoms
The most basic requirement for a visualization program is to build
connectivity between atoms to make a view of a molecule
The visualization program should also be able to produce molecular
structures in different styles, which include
Wire frames, balls and sticks, space-filling spheres, and ribbons
February 10, 2025 Bioinformatics 37
Protein Structure Visualization, Comparison,
Classification
Protein Structure Visualization:
A. Wire frames,
B. Balls and sticks,
C. Space-filling spheres
D. Ribbons
February 10, 2025 Bioinformatics 38
Protein Structure Visualization, Comparison,
Classification
Protein Structure Visualization:
Different representation styles can be used in combination to highlight
a certain feature of a structure while deemphasizing the structures
surrounding it
For example, a cofactor of an enzyme can be shown as space-filling
spheres while the rest of the protein structure is shown as wire frames or
ribbons
Some of the widely used and freely available software programs for
molecular graphics are
RasMol
Molscript
Ribbons
Grasp
February 10, 2025 Bioinformatics 39
Protein Structure Visualization, Comparison,
Classification
Protein Structure Visualization:
RasMol is a command-line–based viewing program that calculates
connectivity of a coordinate file and
Displays wireframe, cylinder, stick bonds, α-carbon trace, space-
filling (CPK) spheres, and ribbons
It reads both PDB and mmCIF formats and can display a whole
molecule or specific parts of it
Molscript is a UNIX program capable of generating wire-frame, space-
filling, or ball-and-stick styles
In particular, secondary structure elements can be drawn with solid
spirals and arrows representing α-helices and β-strands, respectively
The drawback is that the program is command-line–based and not very
user friendly
February 10, 2025 Bioinformatics 40
Protein Structure Visualization, Comparison,
Classification
Protein Structure Visualization:
Ribbons another UNIX program similar to Molscript, generates ribbon
diagrams depicting protein secondary structures
Aesthetically appealing images can be produced that are of publication
quality
However, the program, which is also command-line-based, is extremely
difficult to use
Grasp is a UNIX program that generates solid molecular surface images
and uses a gradated coloring scheme to display electrostatic charges on
the surface
There are also a number of web-based visualization tools that use Java
applets
These programs tend to have limited molecular display features and
low-quality images
February 10, 2025 Bioinformatics 41
Protein Structure Visualization, Comparison,
Classification
Protein Structure Visualization:
Molecular graphics generated by
A. RasMol,
B. Molscript,
C. Ribbons
D. Grasp
February 10, 2025 Bioinformatics 42
Protein Structure Visualization, Comparison,
Classification
Protein Structure Comparison:
With the visualization and computer graphics tools available, it
becomes easy to observe and compare protein structures
To compare protein structures is to analyze two or more protein
structures for similarity
The comparative analysis often, but not always, involves the direct
alignment and superimposition of structures in a 3D space to reveal
Which part of structure is conserved and
Which part is different at the 3D level
This structure comparison is one of the fundamental techniques in
protein structure analysis
The comparative approach is important in finding remote protein
homologs
February 10, 2025 Bioinformatics 43
Protein Structure Visualization, Comparison,
Classification
Protein Structure Comparison:
Because protein structures have a much higher degree of conservation
than the sequences,
Proteins can share common structures even without sequence
similarity
Thus, structure comparison can often reveal distant evolutionary
relationships between proteins,
Which is not feasible using the sequence-based alignment
approach alone
In addition, protein structure comparison is a prerequisite for protein
structural classification into different fold classes
It is also useful in evaluating protein prediction methods by comparing
theoretically predicted structures with experimentally determined ones
February 10, 2025 Bioinformatics 44
Protein Structure Visualization, Comparison,
Classification
Protein Structure Comparison:
The algorithmic approaches to compare protein geometric properties
can be divided into three categories:
The first superposes protein structures by minimizing inter-
molecular distances
The second relies on measuring intra-molecular distances of a
structure and
The third includes algorithms that combine both inter-molecular
and intra-molecular approaches
Inter-molecular Method
This approach is normally applied to relatively similar structures
To compare and superpose two protein structures, one of the structures
has to be moved with respect to the other in such a way that the two
structures have a maximum overlap in a 3D space
February 10, 2025 Bioinformatics 45
Protein Structure Visualization, Comparison,
Classification
Inter-molecular Method
This procedure starts with identifying equivalent residues or atoms
After residue–residue correspondence is established,
One of the structures is moved laterally and vertically toward the
other structure, a process known as translation, to allow the two
structures to be in the same location
The structures are further rotated relative to each other around the 3D
axes, during which process the distances between equivalent positions
are constantly measured
The rotation continues until the shortest intermolecular distance is
reached
At this point, an optimal superimposition of the two structures is
reached
February 10, 2025 Bioinformatics 46
Protein Structure Visualization, Comparison,
Classification
Inter-molecular Method
After superimposition, equivalent residue pairs can be identified, which
helps to quantitate the fitting between the two structures
February 10, 2025 Bioinformatics 47
Protein Structure Visualization, Comparison,
Classification
Inter-molecular Method
An important measurement of the structure fit during superposition is
the distance between equivalent positions on the protein structures
This requires using a least square-fitting function called root mean
square deviation (RMSD), which is the square root of the averaged sum
of the squared differences of the atomic distances
where D is the distance between coordinate data points and N is the
total number of corresponding residue pairs
In practice, only the distances between Cα carbons of corresponding
residues are measured
The goal of structural comparison is to achieve a minimum RMSD
February 10, 2025 Bioinformatics 48
Protein Structure Visualization, Comparison,
Classification
Inter-molecular Method
However, the problem with RMSD is that it depends on the size of the
proteins being compared
For the same degree of sequence identity, large proteins tend to have
higher RMSD values than small proteins when an optimal alignment is
reached
Recently, a logarithmic factor has been proposed to correct this size-
dependency problem
This new measure is called RMSD100 and is determined by the
following formula:
where N is the total number of corresponding atoms
February 10, 2025 Bioinformatics 49
Protein Structure Visualization, Comparison,
Classification
Intra-molecular Method
This approach relies on structural internal statistics and therefore does
not depend on sequence similarity between the proteins to be compared
In addition, this method does not generate a physical superposition of
structures, but instead
Provides a quantitative evaluation of the structural similarity
between corresponding residue pairs
The method works by generating a distance matrix between residues of
the same protein
In comparing two protein structures, the distance matrices from the two
structures are moved relative to each other to achieve maximum overlaps
February 10, 2025 Bioinformatics 50
Protein Structure Visualization, Comparison,
Classification
Intra-molecular Method
By overlaying two distance matrices, similar intra-molecular distance
patterns representing similar structure folding regions can be identified
For the ease of comparison, each matrix is decomposed into smaller
submatrices consisting of hexapeptide fragments
To maximize the similarity regions between two structures, a Monte
Carlo procedure is used
By reducing 3D information into 2D information,
This strategy identifies overall structural resemblances and
common structure cores
February 10, 2025 Bioinformatics 51
Protein Structure Visualization, Comparison,
Classification
Combined Method
A recent development in structure comparison involves combining both
inter- and intra-molecular approaches
In the hybrid approach, corresponding residues can be identified using
the intra-molecular method
Subsequent structure superposition can be performed based on residue
equivalent relationships
In addition to using RMSD as a measure during alignment, additional
structural properties such as
Secondary structure types, torsion angles, accessibility, and
Local hydrogen bonding environment can be used
Dynamic programming is often employed to maximize overlaps in both
inter- and intra-molecular comparisons
February 10, 2025 Bioinformatics 52
Protein Structure Visualization, Comparison,
Classification
Protein Structure Classification:
One of the applications of protein structure comparison is structural
classification
The ability to compare protein structures allows classification of the
structure data and identification of relationships among structures
The reason to develop a protein structure classification system is
To establish hierarchical relationships among protein structures
and
To provide a comprehensive and evolutionary view of known
structures
Once a hierarchical classification system is established, a newly
obtained protein structure can find its place in a proper category
As a result, its functions can be better understood based on association
with other proteins
February 10, 2025 Bioinformatics 53
Protein Structure Visualization, Comparison,
Classification
Protein Structure Classification:
Proteins may be classified according to
both structural and sequence similarity
For structural classification, the sizes
and spatial arrangements of secondary
structures are compared in known 3D
structures
For classification by sequence similarity, alignments of protein
sequences are made using different alignment methods
Protein Structure Classification Systems:
The two most popular classification systems are
Structural Classification of Proteins (SCOP) and
Class, Architecture, Topology and Homologous (CATH)
February 10, 2025 Bioinformatics 54
Protein Structure Visualization, Comparison,
Classification
SCOP
SCOP is a database for comparing and classifying protein structures
It is constructed almost entirely based on manual examination of
protein structures
The proteins are grouped into hierarchies of classes, folds, super-
families, and families
In the SCOP v1.65 there are 7 classes, 800 folds, 1,294 super-families,
and 2,327 families
The SCOP families consist of proteins having high sequence identity
(>30%)
Thus, the proteins within a family clearly share close evolutionary
relationships and normally have the same functionality
The protein structures at this level are also extremely similar
February 10, 2025 Bioinformatics 55
Protein Structure Visualization, Comparison,
Classification
Super-families consist of families with similar structures, but weak
sequence similarity
It is believed that members of the same super-family share a common
ancestral origin, although
The relationships between families are considered distant
Folds consist of super-families with a common core structure, which
is determined manually
This level describes similar overall secondary structures with similar
orientation and connectivity between them
Members within the same fold do not always have evolutionary
relationships
Classes consist of folds with similar core structures
February 10, 2025 Bioinformatics 56
Protein Structure Visualization, Comparison,
Classification
CATH
CATH classifies proteins based on the automatic structural alignment
program SSAP as well as manual comparison
Structural domain separation is carried out also as a combined effort of
a human expert and computer programs
Individual domain structures are classified at five major levels:
Class, architecture, fold/topology, homologous superfamily, and
homologous family
In the CATH v2.5.1 there are
4 classes, 37 architectures, 813 topologies, 1,467 homologous
super-families, and 4,036 homologous families
The definition for class in CATH is similar to that in SCOP, and is
based on secondary structure content
February 10, 2025 Bioinformatics 57
Protein Structure Visualization, Comparison,
Classification
CATH
Architecture is a unique level in CATH, intermediate between fold and
class
This level describes the overall packing and arrangement of secondary
structures independent of connectivity between the elements
The topology level is equivalent to the fold level in SCOP, which
describes
Overall orientation of secondary structures and takes into account
the sequence connectivity between the secondary structure
elements
The homologous superfamily and homologous family levels are
equivalent to
The superfamily and family levels in SCOP with similar
evolutionary definitions, respectively
February 10, 2025 Bioinformatics 58
Protein Structure Visualization, Comparison,
Classification
Comparison of SCOP and CATH
SCOP is almost entirely based on manual comparison of structures by
human experts with no quantitative criteria to group proteins
It is argued that this approach offers some flexibility in recognizing
distant structural relatives, because
Human brains may be more adept at recognizing slightly dissimilar
structures that essentially have the same architecture
However, this reliance on human expertise also renders the method
subjective
The exact boundaries between levels and groups are sometimes arbitrary
CATH is a combination of manual curation and automated procedure,
which makes the process less subjective
February 10, 2025 Bioinformatics 59
Protein Structure Visualization, Comparison,
Classification
Comparison of SCOP and CATH
For example, in defining domains, CATH first relies on the consensus of
three different algorithms to recognize domains
When the computer programs disagree, human intervention will take
place
In addition, the extra Architecture level in CATH makes the structure
classification more continuous
The drawback of the systems is that the fixed thresholds in structural
comparison may make assignment less accurate
Due to the differences in classification criteria, one might expect that
there would be huge differences in classification results
In fact, the classification results from both systems are quite similar
Exhaustive analysis has shown that the results from the two systems
converge at about 80% of the time
February 10, 2025 Bioinformatics 60
Protein 3D Structure Prediction
There are many important proteins for which the sequence information
is available, but their 3Dstructures remain unknown
The full understanding of the biological roles of these proteins
requires knowledge of their structures
Hence, the lack of such information hinders many aspects of the
analysis, ranging from
protein function and ligand binding
Therefore, it is often necessary to obtain approximate protein
structures through computer modeling
There are three computational approaches to protein 3D structural
modeling and prediction
Homology modeling,
Threading, and
Ab initio prediction
February 10, 2025 Bioinformatics 61
Protein 3D Structure Prediction
The first two are knowledge-based methods
They predict protein structures based on knowledge of existing protein
structural information in databases
Homology modeling builds an atomic model based on an
experimentally determined structure that is closely related at the
sequence level
Threading identifies proteins that are structurally similar, with or
without detectable sequence similarities
The Ab initio approach is simulation based and predicts structures
based on physicochemical principles governing protein folding
without the use of structural templates
February 10, 2025 Bioinformatics 62
Protein 3D Structure Prediction
Homology Modeling:
It predicts protein structures based on sequence homology with known
structures
It is also known as comparative modeling
The principle behind it is that
if two proteins share a high enough sequence similarity, they are
likely to have very similar 3D structures
If one of the protein sequences has a known structure, then
The structure can be copied to the unknown protein with a high
degree of confidence
Homology modeling produces an all-atom model based on alignment
with template proteins
The overall homology modeling procedure consists of six steps
February 10, 2025 Bioinformatics 63
Protein 3D Structure Prediction
Homology Modeling:
Step1: Template selection, which involves identification of homologous
sequences in the protein structure database to be used as templates for
modeling
Step2: Alignment of the target and template sequences
Step3: Build a framework structure for the target protein consisting of
main chain atoms
Step4: Model building includes the addition and optimization of side
chain atoms and loops
Step5: To refine and optimize the entire model according to energy
criteria
Step6: Evaluating of the overall quality of the model obtained
If necessary, alignment and model building are repeated until a
satisfactory result is obtained
February 10, 2025 Bioinformatics 64
Protein 3D Structure Prediction
Homology Modeling:
February 10, 2025 Bioinformatics 65
Protein 3D Structure Prediction
Homology Modeling:
Step1:Template Selection
The first step in protein structural modeling is to select appropriate
structural templates
This forms the foundation for rest of the modeling process
It involves searching the PDB for homologous proteins with determined
structures
The search can be performed using a heuristic pairwise alignment search
program such as
BLAST or FASTA
However, the use of dynamic programming based search programs such
as
SSEARCH or ScanPS can result in more sensitive search results
February 10, 2025 Bioinformatics 66
Protein 3D Structure Prediction
Step1:Template Selection
The relatively small size of the structural database means that
The search time using the exhaustive method is still within
reasonable limits, while giving a more sensitive result to ensure the
best possible similarity hits
A database protein should have at least 30% sequence identity with the
query sequence to be selected as template
Multiple database structures with significant similarity can be found as a
result of the search
In that case, it is recommended that the structure(s) with
The highest percentage identity,
Highest resolution, and
The most appropriate cofactors is selected as a template
February 10, 2025 Bioinformatics 67
Protein 3D Structure Prediction
Step2:Sequence Alignment
Once the structure with the highest sequence similarity is identified as a
template,
The full-length sequences of the template and target proteins need
to be realigned using refined alignment algorithms to obtain
optimal alignment
This realignment is the most critical step in homology modeling, which
directly affects the quality of the final model
This is because incorrect alignment at this stage leads to
Incorrect designation of homologous residues and therefore to
incorrect structural models
Errors made in the alignment step cannot be corrected in the following
modeling steps
February 10, 2025 Bioinformatics 68
Protein 3D Structure Prediction
Step2:Sequence Alignment
Therefore, the best possible multiple alignment algorithms, such as
Praline and T-Coffee should be used for this purpose
Even alignment using the best alignment program may not be error free
and should be visually inspected to ensure that conserved key residues
are correctly aligned
If necessary, manual refinement of the alignment should be carried out
to improve alignment quality
Step3:Backbone Model Building
In optimal alignment, residues in the aligned regions of the target
protein can assume a similar structure as the template proteins, meaning
that
The coordinates of the corresponding residues of the template
proteins can be simply copied onto the target protein
February 10, 2025 Bioinformatics 69
Protein 3D Structure Prediction
Step3:Backbone Model Building
If the two aligned residues are identical, coordinates of the side chain
atoms are copied along with the main chain atoms
If the two residues differ, only the backbone atoms can be copied
The side chain atoms are rebuilt in a subsequent procedure
In backbone modeling, it is simplest to use only one template structure
The structure with the best quality and highest resolution is normally
chosen if multiple options are available
This structure tends to carry the fewest errors
Occasionally, multiple template structures are available for modeling
In this situation, the template structures have to be optimally aligned and
superimposed before being used as templates in model building
One can either choose to use average coordinate values of the templates
or the best parts from each of the templates to model
February 10, 2025 Bioinformatics 70
Protein 3D Structure Prediction
Step4a:Loop Modeling
In the sequence alignment for modeling, there are often regions caused
by insertions and deletions producing gaps in sequence alignment
The gaps cannot be directly modeled, creating “holes” in the model
Closing the gaps requires loop modeling, which is a very difficult
problem in homology modeling and is also a major source of error
Loop modeling can be considered a mini–protein modeling problem by
itself
Unfortunately, there are no mature methods available that can model
loops reliably
Currently, there are two main techniques used to approach the problem:
The database searching method and
The Ab initio method
February 10, 2025 Bioinformatics 71
Protein 3D Structure Prediction
Step4a:Loop Modeling
The database method involves finding “spare parts” from known protein
structures in a database that
Fit onto the two stem regions of the target protein
The stems are defined as the main chain atoms that precede and follow
the loop to be modeled
The procedure begins by measuring the orientation and distance of the
anchor regions in the stems and searching PDB for
Segments of the same length that also match the above endpoint
conformation
Usually, many different alternative segments that fit the endpoints of the
stems are available
The best loop can be selected based on sequence similarity as well as
minimal steric clashes with the neighboring parts of the structure
February 10, 2025 Bioinformatics 72
Protein 3D Structure Prediction
Step4a:Loop Modeling
The conformation of the best matching fragments is then copied onto the
anchoring points of the stems
The Ab initio method generates many random loops and searches for the
one that does not clash with nearby side chains and
Also has reasonably low energy and φ and ψ angles in the allowable
regions in the Ramachandran plot
The specialized programs for loop modeling: FREAD, PETRA, CODA
February 10, 2025 Bioinformatics 73
Protein 3D Structure Prediction
Step4b:Side Chain Refinement
Once main chain atoms are built, the positions of side chains that are not
modeled must be determined
Modeling side chain geometry is very important in evaluating protein–
ligand interactions at active sites and protein–protein interactions at the
contact interface
A side chain can be built by searching every possible conformation at
every torsion angle of the side chain
To select the one that has the lowest interaction energy with
neighboring atoms
However, this approach is computationally prohibitive in most cases
In fact, most current side chain prediction programs use the concept of
Rotamers, which are favored side chain torsion angles extracted
from known protein crystal structures
February 10, 2025 Bioinformatics 74
Protein 3D Structure Prediction
Step4b:Side Chain Refinement
A collection of preferred side chain conformations is a
Rotamer library in which the rotamers are ranked by their
frequency of occurrence
Having a rotamer library reduces the computational time significantly
because
Only a small number of favored torsion angles are examined
In prediction of side chain conformation, only the possible rotamers
with the lowest interaction energy with nearby atoms are selected
Most modeling packages incorporate the side chain refinement function
A specialized side chain modeling program that has reasonably good
performance is
SCWRL (sidechain placement with a rotamer library)
February 10, 2025 Bioinformatics 75
Protein 3D Structure Prediction
Step5: Model Refinement Using Energy Function
In these loop modeling and side chain modeling steps, potential energy
calculations are applied to improve the model
However, this does not guarantee that
The entire raw homology model is free of structural irregularities
such as unfavorable bond angles, bond lengths, or close atomic
contacts
These kinds of structural irregularities can be corrected by
Applying the energy minimization procedure on the entire model,
which moves the atoms in such a way that the overall
conformation has the lowest energy potential
The goal of energy minimization is
To relieve steric collisions and strains without significantly
altering the overall structure
February 10, 2025 Bioinformatics 76
Protein 3D Structure Prediction
Step5: Model Refinement Using Energy Function
However, energy minimization has to be used with caution because
Excessive energy minimization often moves residues away from
their correct positions
Therefore, only limited energy minimization is recommended (a few
hundred iterations)
To remove major errors, such as short bond distances and close
atomic clashes
Key conserved residues and those involved in cofactor binding have
To be restrained if necessary during the process
Step6: Model Evaluation
The final homology model has to be evaluated to make sure that
The structural features of the model are consistent with the
physicochemical rules
February 10, 2025 Bioinformatics 77
Protein 3D Structure Prediction
Step6: Model Evaluation
This involves checking anomalies in φ–ψ angles, bond lengths, close
contacts, and so on
Another way of checking the quality of a protein model is to implicitly
take these stereo-chemical properties into account
This is a method that detects errors by
Compiling statistical profiles of spatial features and interaction
energy from experimentally determined structures
By comparing the statistical parameters with the constructed model,
The method reveals which regions of a sequence appear to be
folded normally and which regions do not
If structural irregularities are found, the region is considered to have
errors and has to be further refined
February 10, 2025 Bioinformatics 78
Thank You
February 10, 2025 Bioinformatics 79