0% found this document useful (0 votes)
12 views79 pages

Protein Structure Classification in Bioinformatics

The document outlines the classification and analysis of protein structures, emphasizing the importance of protein structure in determining biological functions. It discusses various aspects of protein structure, including secondary structures, dihedral angles, and the Ramachandran plot, which illustrates allowed conformations of polypeptides. The document also touches on the roles of amino acids in protein structure and the significance of posttranslational modifications.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPT, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
12 views79 pages

Protein Structure Classification in Bioinformatics

The document outlines the classification and analysis of protein structures, emphasizing the importance of protein structure in determining biological functions. It discusses various aspects of protein structure, including secondary structures, dihedral angles, and the Ramachandran plot, which illustrates allowed conformations of polypeptides. The document also touches on the roles of amino acids in protein structure and the significance of posttranslational modifications.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPT, PDF, TXT or read online on Scribd

Bioinformatics

— Unit V: Protein Structure Classification—

Dr. Chandra Mohan D


Assistant Professor
Computer Science and Engineering Group
Indian Institute of Information Technology, Sri City

If you know your own DNA sequence than you know every thing about your self

February 10, 2025 Bioinformatics 1


Outline
 Introduction (11-03-2023)

 Secondary Structure elements

 Ramachandran plot

 Propensity

 Secondary and 3D structure prediction

 Visualization tools

 Structural classification

February 10, 2025 Bioinformatics 2


Introduction
 Proteins perform most essential biological and chemical functions in
a cell
 They play important roles in structural, enzymatic, transport, and
regulatory functions
 The protein functions are strictly determined by their structures
 Therefore, protein structural analysis is an essential element of
bioinformatics
 The building blocks of proteins are 20 naturally occurring amino
acids, small molecules
 Amino acids contain a free amino group (NH2) and a free carboxyl
group (COOH)
 Both of these groups are linked to a
central carbon (Cα), which is attached
to a hydrogen and a side chain group (R)
February 10, 2025 Bioinformatics 3
Introduction
 Amino acids differ only by the side chain R group
 The chemical reactivities of the R groups determine the specific
properties of the amino acids
 Amino acids can be grouped into several categories based on the
chemical and physical properties of the side chains (R), such as
 size and affinity for water

 According to these properties, the side chain groups can be divided


into small, large, hydrophobic, and hydrophilic categories
 Within the hydrophobic set of amino acids, they can be further
divided into Aliphatic and Aromatic
 Aliphatic side chains are linear
hydrocarbon chains and
Aromatic side chains are cyclic rings
February 10, 2025 Bioinformatics 4
Introduction
 Within the hydrophilic set, amino
acids can be subdivided into polar
and charged
 Charged amino acids can be either
positively charged (basic) or
negatively charged (acidic)
 The R side chains also play an
important structural role
 Of particular interest within the twenty
amino acids are glycine and proline
 Glycine, the smallest amino acid, has
a hydrogen atom as the R group

February 10, 2025 Bioinformatics 5


Introduction
 It can therefore adopt more flexible conformations that are not
possible for other amino acids
 Glycine increase local flexibility in structures

 Proline is on the other extreme of flexibility


 Its side chain forms a bond with its own backbone amino group,
causing it to be cyclic

 The cyclic conformation makes it very rigid,


 Unable to occupy many of the main chain conformations adopted

by other amino acids


 Cysteine, which can react with another cysteine to form a cross-link
that can stabilize the protein structure
February 10, 2025 Bioinformatics 6
Protein structure Terminology
 Certain amino acids are subject to modifications after a protein is
translated in a cell
 This is called posttranslational modification
 The peptide formation involves two amino acids covalently joined
together between
 the carboxyl group of one amino acid and the amino group of

another
 This reaction is a condensation reaction involving removal of
elements of water from the two molecules
 The resulting product is called a dipeptide

February 10, 2025 Bioinformatics 7


Protein structure Terminology
 The newly formed covalent bond connecting the two amino acids is
called a peptide bond
 Once an amino acid is incorporated into a peptide, it becomes an amino
acid residue
 Multiple amino acids can be joined together to form a longer chain of
amino acid polymer
 A linear polymer of more than fifty amino acid residues is referred to as
a polypeptide
 A polypeptide, also called a protein, has a well-defined 3D arrangement
 On the other hand, a polymer with fewer than fifty residues is usually
called a peptide without a well-defined 3D structure

February 10, 2025 Bioinformatics 8


Protein structure Terminology
 The residues in a peptide or polypeptide are numbered
 beginning with the residue containing the amino group, referred to

as the N-terminus, and


 ending with the residue containing the carboxyl group, known as the

C-terminus

 The actual sequence of amino acid residues in a polypeptide determines


its ultimate structure and function
 The atoms involved in forming the peptide bond are referred to as the
backbone atoms
February 10, 2025 Bioinformatics 9
Protein structure Terminology
 They are the nitrogen of the amino group, the α carbon to which the
side chain is attached and carbon of the carbonyl group
 The polypeptide chain is first assembled on the ribosome using the
codon sequence on mRNA as a template
 The resulting linear chain forms secondary structures through
 The formation of hydrogen bonds between amino acids in the chain

 Through further interactions among amino acid side groups, these 2D


structures then fold into a 3D structure

February 10, 2025 Bioinformatics 10


Protein structure Terminology
 Proteins are chains of amino acids joined by peptide bonds
 Many conformations of the chain
are possible due to the rotation of
the chain about each Cα atom
 These conformational variations
are responsible for differences in
the 3D structures of proteins
 Each amino acid in the chain is polar,
 It has separated positive and negatively charged regions with a

chemically free C=O group, which can act as a hydrogen bond


acceptor, and an NH group, which can act as a hydrogen bond
donor
 These groups interact in protein structures
February 10, 2025 Bioinformatics 11
Protein structure Terminology
Dihedral Angels:
A peptide bond is actually a partial double bond owing to shared

electrons between O=C–N atoms


The rigid double bond structure forces atoms associated with the

peptide bond to lie in the same plane, called the peptide plane
Because of the planar nature of the peptide bond and the size of the R

groups,
 There are considerable restrictions on the rotational freedom

The angle of rotation about the bond is referred to as the dihedral angle

(also called the Torsional angle)


For a peptide unit, the atoms linked to the peptide bond can be moved

to a certain extent by the rotation of two bonds flanking the peptide


bond
This is measured by two dihedral angles
February 10, 2025 Bioinformatics 12
Protein structure Terminology
Dihedral Angels:
One is the dihedral angle along the N–Cα bond, which is defined as phi

(φ); and
The other is the angle along the Cα–C bond, which is called psi (ψ)

Various combinations of φ and ψ angles allow the proteins to fold in

many different ways

February 10, 2025 Bioinformatics 13


Protein Secondary Structures
Secondary Structures:
Regular recurring arrangements in space of adjacent amino acid

residues in a polypeptide chain


In these 2D structures, regular patterns of H bonds are formed between

neighboring amino acids, and the amino acids have similar ϕ and ψ
angles
The formation of these

structures neutralizes
the polar groups on
each amino acid
The secondary structures

are tightly packed in the


protein core in a hydrophobic environment
February 10, 2025 Bioinformatics 14
Protein Secondary Structures
Types of Secondary Structures:
There are four commonly occurring 2D structures

1. α-helices
2. β-strands (sheets)
3. Turns (bends)
4. Coil (irregular)
1. α Helix:
The α helix is the most abundant type of secondary

structure in proteins
The helix has 3.6 amino acids per turn with an H bond formed between

every fourth residue


The helix has 5.4Ӑ per turn.

The average length is 10 amino acids (3 turns)

February 10, 2025 Bioinformatics 15


Protein Secondary Structures
 The backbone of the chain is shown in red, the Cα atoms, the C=O
and NH groups are shown in blue, yellow and green
respectively
 The alignment of the H bonds creates a dipole moment
for the helix with a resulting partial positive charge at
the amino end of the helix
 Because this region has free NH2 groups, it will
interact with negatively charged groups such as
phosphates
 In the α helix, note that each C=O group at amino acid
position n in the sequence is hydrogen-bonded with the
NH group at position n + 4

February 10, 2025 Bioinformatics 16


Protein Secondary Structures
α Helix:
The helix is usually right-handed, but short sections of

3–5 amino acids of left-handed helices occur occasionally


The average ϕ and ψ angles of the amino acids in

 The right-handed helix are approximately 60° and 40°,

respectively
The R side chains of the amino acids are on the outside

of the helix

February 10, 2025 Bioinformatics 17


Protein Secondary Structures
β Sheets:
β Sheets are formed by H bonds between an average of 5–10

consecutive amino acids in one portion of the chain with another


5–10 farther down the chain
The interacting regions may be adjacent,

with a short loop in between, or far apart,


with other structures in between
Every chain may run in the same direction

to form a parallel sheet, every other chain


may run in the reverse chemical direction
to form an antiparallel sheet, or the chains
may be parallel and antiparallel to form a
mixed sheet
February 10, 2025 Bioinformatics 18
Protein Secondary Structures
β Sheets:
Each amino acid in the interior strands of the sheet forms two H bonds

with neighboring amino acids, whereas


Each amino acid on the outside strands forms only one bond with an

interior strand
Coils and Loops:
There are also local structures that do not belong to regular secondary

structures (α-helices and β-strands)


The irregular structures are coils or loops

The loops are often characterized by sharp turns or hairpin-like

structures
If the connecting regions are completely irregular, they belong to

random coils

February 10, 2025 Bioinformatics 19


Protein Secondary Structures
Coils and Loops:
Residues in the loop or coil regions tend to be charged, polar and

located on the surface of the protein structure


They are often the evolutionarily variable regions where mutations,

deletions, and insertions frequently occur


They can be functionally significant because these locations are often

the active sites of proteins


Coiled Coils
Coiled coils are a special type of super secondary structure

characterized by a bundle of two or more α-helices wrapping


around each other
The helices forming coiled coils have a unique pattern of

hydrophobicity, which repeats every seven residues (five hydrophobic


and two hydrophilic)
February 10, 2025 Bioinformatics 20
Protein Secondary Structures
Turns:
One-third of all residues in globular proteins contained in turns that

serve to reverse the direction of the polypeptide chain


Turn contains the hydrogen bond between the carbonyl oxygen of

residue i and the amide nitrogen of i+3


 There are three types of turns, named I, II, III

Type I occur more frequently (2-3 times faster than type-II)

The backbone dihedral angels of the residue are (-60,-30) and (-90,0)

of residues of i+1 and i+2 of type-I turn


Prolin is often found in position i+1 in type-I

turns as its phi angle is restricted to -60

February 10, 2025 Bioinformatics 21


Protein Secondary Structures
Turns:
The backbone dihedral angels of the residue are (-60,120) and (80,0) of

residues of i+1 and i+2 of type-II turn


Glycine is favoured in this position i+1 in type-II as it is requires a

positive phi value

 The backbone dihedral angels of the residue


are (-60,-30) and (-60,-30) of residues of i+1
and i+2 of classical type-III turn

February 10, 2025 Bioinformatics 22


Ramachandran Plot
 G.N Ramachandran used computer models of small polypeptides to
systematically vary phi and psi with the objective of finding stable
conformations
 He plots the phi value on the X-axis and the psi value on the Y-axis
 Plotting the torsional angles in this way graphically shows which
combination of angles are possible

February 10, 2025 Bioinformatics 23


Ramachandran Plot
 In the diagram the white areas correspond to conformations where
atoms in polypeptide come closer than the sum of their van der waals
radii
 These regions are sterically
disallowed for all amino acids
except glycine which is unique
in that it lacks a side chain
 The red regions correspond to
conformations where there are no
steric clashes i.e these are the
allowed regions namely the
α- helical and β-sheet
conformations
February 10, 2025 Bioinformatics 24
Ramachandran Plot
 The yellow area shows the allowed regions if slightly shorter van der
waals radi are used in the calculation, i.e the atoms are allowed to
come a little closer together
 This bring out an additional region which corresponds to the left
handed α-helix
 Glycine has no side chain and therefore can adapt phi and psi angels
in all four quadrants of the Ramachandran plot
 This is the Ramachandran plot for
pyruvate kinase (glycolysis) with
all amino acids accounted for except
Glycine

February 10, 2025 Bioinformatics 25


Propensity Value

February 10, 2025 Bioinformatics 26


Propensity Value

February 10, 2025 Bioinformatics 27


Propensity Value

February 10, 2025 Bioinformatics 28


Protein Structure Prediction
Classes of Protein Structure
Four principal classes of protein structure were recognized based on

the types and arrangements of secondary structural elements


Class α comprises a bundle of α helices connected by loops on the

surface of the proteins

 Structure of α class proteins. (A) Diagram showing a-helical pattern of


this class. a helices are red cylinders, and black lines are loops.
(B) Example of the class, hemoglobin, PDB file 3hhb displayed
using Rasmol, using ribbons display
February 10, 2025
and group color
Bioinformatics 29
Protein Structure Prediction
Classes of Protein Structure
Class β comprises antiparallel β sheets, usually two sheets in close

contact forming a sandwich


Alternatively, a sheet can twist into a barrel with the first and last

strands touching
Examples are enzymes, transport proteins, antibodies, and virus coat

proteins such as neuraminidase

February 10, 2025 Bioinformatics 30


Protein Structure Prediction
Classes of Protein Structure
Class α/β comprises mainly parallel β sheets with intervening α

helices, but may also have mixed β sheets


In addition to forming a sheet in some proteins in this class, as

illustrated below, in others parallel β strands may form into a barrel


structure that is surrounded by α
helices
This class of proteins includes

many metabolic enzymes

February 10, 2025 Bioinformatics 31


Protein Structure Prediction
Classes of Protein Structure
Class α+β comprises mainly segregated α helices and antiparallel β

sheets

February 10, 2025 Bioinformatics 32


Protein Structure Prediction
Classes of Protein Structure
Multidomain (α and β) proteins comprise domains representing more than one of the

above four classes


Membrane and cell-surface proteins and peptides excluding proteins of the immune

system comprise this class

Diagram showing typical


arrangement of membrane
-traversing, hydrophobic
a helices (red)
Membrane bilayer shown

as green
February lines
10, 2025 Bioinformatics 33
Protein Tertiary Structure
 The overall packing and arrangement of secondary structures form
the tertiary structure of a protein ([Link]
v=piXHivrTT-E)
 The tertiary structure can come in various forms but is generally
classified as either globular or membrane proteins
 The former exists in solvents through hydrophilic interactions with
solvent molecules;
 The latter exists in membrane lipids and is stabilized through
hydrophobic interactions with the lipid molecules
 Membrane proteins are long, strand-like proteins that are insoluble in
water, weak acids, and weak bases
 Globular proteins have a spherical shape and are soluble in water,
acids, and bases

February 10, 2025 Bioinformatics 34


Protein Tertiary Structure
Globular Proteins
These are usually soluble and surrounded by water molecules

They tend to have an overall compact structure of spherical shape with

polar or hydrophilic residues on the surface and hydrophobic residues in


the core
Such an arrangement is energetically

favorable because
 It minimizes contacts with water by

hydrophobic residues in the core and


maximizes interactions with water by
surface polar and charged residues
Examples: enzymes, cytokines, and protein hormones

February 10, 2025 Bioinformatics 35


Protein Tertiary Structure
Integral Membrane Proteins
Membrane proteins exist in lipid bilayers of cell

membranes
Because they are surrounded by lipids, the

exterior of the proteins spanning the membrane


must be very hydrophobic to be stable
Most typical transmembrane segments are

α-helices
Occasionally, for some bacterial periplasmic

membrane proteins, they are composed of


β-strands
Examples of membrane proteins are rhodopsins,

cytochrome c oxidase, and ion channel proteins


February 10, 2025 Bioinformatics 36
Protein Structure Visualization, Comparison,
Classification

Protein Structure Visualization:


With the development of computer hardware and software technology,
 Sophisticated computer graphics programs have been developed

for visualizing and manipulating complicated 3D structures


The computer graphics help to analyze and compare protein structures

to gain insight to functions of the proteins


PDB data file for a protein structure contains only x, y, and z

coordinates of atoms
The most basic requirement for a visualization program is to build

connectivity between atoms to make a view of a molecule


The visualization program should also be able to produce molecular

structures in different styles, which include


 Wire frames, balls and sticks, space-filling spheres, and ribbons

February 10, 2025 Bioinformatics 37


Protein Structure Visualization, Comparison,
Classification

Protein Structure Visualization:

A. Wire frames,
B. Balls and sticks,
C. Space-filling spheres
D. Ribbons

February 10, 2025 Bioinformatics 38


Protein Structure Visualization, Comparison,
Classification

Protein Structure Visualization:


Different representation styles can be used in combination to highlight
a certain feature of a structure while deemphasizing the structures
surrounding it
For example, a cofactor of an enzyme can be shown as space-filling

spheres while the rest of the protein structure is shown as wire frames or
ribbons
Some of the widely used and freely available software programs for

molecular graphics are


 RasMol

 Molscript

 Ribbons

 Grasp

February 10, 2025 Bioinformatics 39


Protein Structure Visualization, Comparison,
Classification

Protein Structure Visualization:


RasMol is a command-line–based viewing program that calculates
connectivity of a coordinate file and
 Displays wireframe, cylinder, stick bonds, α-carbon trace, space-

filling (CPK) spheres, and ribbons


It reads both PDB and mmCIF formats and can display a whole

molecule or specific parts of it


Molscript is a UNIX program capable of generating wire-frame, space-

filling, or ball-and-stick styles


In particular, secondary structure elements can be drawn with solid

spirals and arrows representing α-helices and β-strands, respectively


The drawback is that the program is command-line–based and not very

user friendly

February 10, 2025 Bioinformatics 40


Protein Structure Visualization, Comparison,
Classification

Protein Structure Visualization:


Ribbons another UNIX program similar to Molscript, generates ribbon
diagrams depicting protein secondary structures
Aesthetically appealing images can be produced that are of publication

quality
However, the program, which is also command-line-based, is extremely

difficult to use
Grasp is a UNIX program that generates solid molecular surface images

and uses a gradated coloring scheme to display electrostatic charges on


the surface
There are also a number of web-based visualization tools that use Java

applets
These programs tend to have limited molecular display features and

low-quality images
February 10, 2025 Bioinformatics 41
Protein Structure Visualization, Comparison,
Classification

Protein Structure Visualization:


Molecular graphics generated by
A. RasMol,
B. Molscript,
C. Ribbons
D. Grasp

February 10, 2025 Bioinformatics 42


Protein Structure Visualization, Comparison,
Classification

Protein Structure Comparison:


With the visualization and computer graphics tools available, it
becomes easy to observe and compare protein structures
To compare protein structures is to analyze two or more protein

structures for similarity


The comparative analysis often, but not always, involves the direct

alignment and superimposition of structures in a 3D space to reveal


 Which part of structure is conserved and

 Which part is different at the 3D level

This structure comparison is one of the fundamental techniques in

protein structure analysis


The comparative approach is important in finding remote protein

homologs
February 10, 2025 Bioinformatics 43
Protein Structure Visualization, Comparison,
Classification

Protein Structure Comparison:


Because protein structures have a much higher degree of conservation
than the sequences,
 Proteins can share common structures even without sequence

similarity
Thus, structure comparison can often reveal distant evolutionary

relationships between proteins,


 Which is not feasible using the sequence-based alignment

approach alone
In addition, protein structure comparison is a prerequisite for protein

structural classification into different fold classes


It is also useful in evaluating protein prediction methods by comparing

theoretically predicted structures with experimentally determined ones

February 10, 2025 Bioinformatics 44


Protein Structure Visualization, Comparison,
Classification

Protein Structure Comparison:


The algorithmic approaches to compare protein geometric properties
can be divided into three categories:
 The first superposes protein structures by minimizing inter-

molecular distances
 The second relies on measuring intra-molecular distances of a

structure and
 The third includes algorithms that combine both inter-molecular

and intra-molecular approaches


Inter-molecular Method
This approach is normally applied to relatively similar structures

To compare and superpose two protein structures, one of the structures

has to be moved with respect to the other in such a way that the two
structures have a maximum overlap in a 3D space
February 10, 2025 Bioinformatics 45
Protein Structure Visualization, Comparison,
Classification
Inter-molecular Method
This procedure starts with identifying equivalent residues or atoms

After residue–residue correspondence is established,

 One of the structures is moved laterally and vertically toward the

other structure, a process known as translation, to allow the two


structures to be in the same location
The structures are further rotated relative to each other around the 3D

axes, during which process the distances between equivalent positions


are constantly measured
The rotation continues until the shortest intermolecular distance is

reached
At this point, an optimal superimposition of the two structures is

reached

February 10, 2025 Bioinformatics 46


Protein Structure Visualization, Comparison,
Classification
Inter-molecular Method
After superimposition, equivalent residue pairs can be identified, which

helps to quantitate the fitting between the two structures

February 10, 2025 Bioinformatics 47


Protein Structure Visualization, Comparison,
Classification
Inter-molecular Method
An important measurement of the structure fit during superposition is

the distance between equivalent positions on the protein structures


This requires using a least square-fitting function called root mean

square deviation (RMSD), which is the square root of the averaged sum
of the squared differences of the atomic distances

where D is the distance between coordinate data points and N is the


total number of corresponding residue pairs
In practice, only the distances between Cα carbons of corresponding

residues are measured


The goal of structural comparison is to achieve a minimum RMSD

February 10, 2025 Bioinformatics 48


Protein Structure Visualization, Comparison,
Classification
Inter-molecular Method
However, the problem with RMSD is that it depends on the size of the

proteins being compared


For the same degree of sequence identity, large proteins tend to have

higher RMSD values than small proteins when an optimal alignment is


reached
Recently, a logarithmic factor has been proposed to correct this size-

dependency problem
This new measure is called RMSD100 and is determined by the

following formula:

 where N is the total number of corresponding atoms

February 10, 2025 Bioinformatics 49


Protein Structure Visualization, Comparison,
Classification

Intra-molecular Method
This approach relies on structural internal statistics and therefore does
not depend on sequence similarity between the proteins to be compared
In addition, this method does not generate a physical superposition of

structures, but instead


 Provides a quantitative evaluation of the structural similarity

between corresponding residue pairs


The method works by generating a distance matrix between residues of

the same protein


In comparing two protein structures, the distance matrices from the two

structures are moved relative to each other to achieve maximum overlaps

February 10, 2025 Bioinformatics 50


Protein Structure Visualization, Comparison,
Classification

Intra-molecular Method
By overlaying two distance matrices, similar intra-molecular distance
patterns representing similar structure folding regions can be identified
For the ease of comparison, each matrix is decomposed into smaller

submatrices consisting of hexapeptide fragments


To maximize the similarity regions between two structures, a Monte

Carlo procedure is used


By reducing 3D information into 2D information,

 This strategy identifies overall structural resemblances and

common structure cores

February 10, 2025 Bioinformatics 51


Protein Structure Visualization, Comparison,
Classification

Combined Method
A recent development in structure comparison involves combining both
inter- and intra-molecular approaches
In the hybrid approach, corresponding residues can be identified using

the intra-molecular method


Subsequent structure superposition can be performed based on residue

equivalent relationships
In addition to using RMSD as a measure during alignment, additional

structural properties such as


 Secondary structure types, torsion angles, accessibility, and

 Local hydrogen bonding environment can be used

Dynamic programming is often employed to maximize overlaps in both

inter- and intra-molecular comparisons


February 10, 2025 Bioinformatics 52
Protein Structure Visualization, Comparison,
Classification
Protein Structure Classification:
One of the applications of protein structure comparison is structural
classification
The ability to compare protein structures allows classification of the

structure data and identification of relationships among structures


The reason to develop a protein structure classification system is

 To establish hierarchical relationships among protein structures

and
 To provide a comprehensive and evolutionary view of known

structures
Once a hierarchical classification system is established, a newly

obtained protein structure can find its place in a proper category


As a result, its functions can be better understood based on association

with other proteins


February 10, 2025 Bioinformatics 53
Protein Structure Visualization, Comparison,
Classification
Protein Structure Classification:
Proteins may be classified according to
both structural and sequence similarity
For structural classification, the sizes

and spatial arrangements of secondary


structures are compared in known 3D
structures
For classification by sequence similarity, alignments of protein

sequences are made using different alignment methods


Protein Structure Classification Systems:
 The two most popular classification systems are
 Structural Classification of Proteins (SCOP) and

 Class, Architecture, Topology and Homologous (CATH)

February 10, 2025 Bioinformatics 54


Protein Structure Visualization, Comparison,
Classification
SCOP
SCOP is a database for comparing and classifying protein structures

It is constructed almost entirely based on manual examination of

protein structures
The proteins are grouped into hierarchies of classes, folds, super-

families, and families


In the SCOP v1.65 there are 7 classes, 800 folds, 1,294 super-families,

and 2,327 families


The SCOP families consist of proteins having high sequence identity

(>30%)
Thus, the proteins within a family clearly share close evolutionary

relationships and normally have the same functionality


The protein structures at this level are also extremely similar

February 10, 2025 Bioinformatics 55


Protein Structure Visualization, Comparison,
Classification
 Super-families consist of families with similar structures, but weak
sequence similarity
 It is believed that members of the same super-family share a common
ancestral origin, although
 The relationships between families are considered distant

 Folds consist of super-families with a common core structure, which


is determined manually
 This level describes similar overall secondary structures with similar
orientation and connectivity between them
 Members within the same fold do not always have evolutionary
relationships
 Classes consist of folds with similar core structures

February 10, 2025 Bioinformatics 56


Protein Structure Visualization, Comparison,
Classification
CATH
CATH classifies proteins based on the automatic structural alignment

program SSAP as well as manual comparison


Structural domain separation is carried out also as a combined effort of

a human expert and computer programs


Individual domain structures are classified at five major levels:

 Class, architecture, fold/topology, homologous superfamily, and

homologous family
In the CATH v2.5.1 there are

 4 classes, 37 architectures, 813 topologies, 1,467 homologous

super-families, and 4,036 homologous families


The definition for class in CATH is similar to that in SCOP, and is

based on secondary structure content

February 10, 2025 Bioinformatics 57


Protein Structure Visualization, Comparison,
Classification
CATH
Architecture is a unique level in CATH, intermediate between fold and

class
This level describes the overall packing and arrangement of secondary

structures independent of connectivity between the elements


The topology level is equivalent to the fold level in SCOP, which

describes
 Overall orientation of secondary structures and takes into account

the sequence connectivity between the secondary structure


elements
 The homologous superfamily and homologous family levels are

equivalent to
 The superfamily and family levels in SCOP with similar

evolutionary definitions, respectively


February 10, 2025 Bioinformatics 58
Protein Structure Visualization, Comparison,
Classification
Comparison of SCOP and CATH
SCOP is almost entirely based on manual comparison of structures by

human experts with no quantitative criteria to group proteins


It is argued that this approach offers some flexibility in recognizing

distant structural relatives, because


 Human brains may be more adept at recognizing slightly dissimilar

structures that essentially have the same architecture


However, this reliance on human expertise also renders the method

subjective
The exact boundaries between levels and groups are sometimes arbitrary

CATH is a combination of manual curation and automated procedure,

which makes the process less subjective

February 10, 2025 Bioinformatics 59


Protein Structure Visualization, Comparison,
Classification
Comparison of SCOP and CATH
For example, in defining domains, CATH first relies on the consensus of

three different algorithms to recognize domains


When the computer programs disagree, human intervention will take

place
In addition, the extra Architecture level in CATH makes the structure

classification more continuous


The drawback of the systems is that the fixed thresholds in structural

comparison may make assignment less accurate


Due to the differences in classification criteria, one might expect that

there would be huge differences in classification results


In fact, the classification results from both systems are quite similar

Exhaustive analysis has shown that the results from the two systems

converge at about 80% of the time


February 10, 2025 Bioinformatics 60
Protein 3D Structure Prediction
 There are many important proteins for which the sequence information
is available, but their 3Dstructures remain unknown
 The full understanding of the biological roles of these proteins
requires knowledge of their structures
 Hence, the lack of such information hinders many aspects of the
analysis, ranging from
 protein function and ligand binding

 Therefore, it is often necessary to obtain approximate protein


structures through computer modeling
 There are three computational approaches to protein 3D structural
modeling and prediction
 Homology modeling,

 Threading, and

 Ab initio prediction

February 10, 2025 Bioinformatics 61


Protein 3D Structure Prediction
 The first two are knowledge-based methods
 They predict protein structures based on knowledge of existing protein
structural information in databases
 Homology modeling builds an atomic model based on an
experimentally determined structure that is closely related at the
sequence level
 Threading identifies proteins that are structurally similar, with or
without detectable sequence similarities
 The Ab initio approach is simulation based and predicts structures
based on physicochemical principles governing protein folding
without the use of structural templates

February 10, 2025 Bioinformatics 62


Protein 3D Structure Prediction
Homology Modeling:
It predicts protein structures based on sequence homology with known

structures
It is also known as comparative modeling

The principle behind it is that

 if two proteins share a high enough sequence similarity, they are

likely to have very similar 3D structures


If one of the protein sequences has a known structure, then

 The structure can be copied to the unknown protein with a high

degree of confidence
Homology modeling produces an all-atom model based on alignment

with template proteins


The overall homology modeling procedure consists of six steps

February 10, 2025 Bioinformatics 63


Protein 3D Structure Prediction
Homology Modeling:
Step1: Template selection, which involves identification of homologous

sequences in the protein structure database to be used as templates for


modeling
Step2: Alignment of the target and template sequences

Step3: Build a framework structure for the target protein consisting of

main chain atoms


Step4: Model building includes the addition and optimization of side

chain atoms and loops


Step5: To refine and optimize the entire model according to energy

criteria
Step6: Evaluating of the overall quality of the model obtained

If necessary, alignment and model building are repeated until a

satisfactory result is obtained


February 10, 2025 Bioinformatics 64
Protein 3D Structure Prediction
Homology Modeling:

February 10, 2025 Bioinformatics 65


Protein 3D Structure Prediction
Homology Modeling:
Step1:Template Selection
The first step in protein structural modeling is to select appropriate

structural templates
This forms the foundation for rest of the modeling process

It involves searching the PDB for homologous proteins with determined

structures
The search can be performed using a heuristic pairwise alignment search

program such as
 BLAST or FASTA

However, the use of dynamic programming based search programs such

as
 SSEARCH or ScanPS can result in more sensitive search results

February 10, 2025 Bioinformatics 66


Protein 3D Structure Prediction
Step1:Template Selection
The relatively small size of the structural database means that

 The search time using the exhaustive method is still within

reasonable limits, while giving a more sensitive result to ensure the


best possible similarity hits
A database protein should have at least 30% sequence identity with the

query sequence to be selected as template


Multiple database structures with significant similarity can be found as a

result of the search


In that case, it is recommended that the structure(s) with

 The highest percentage identity,

 Highest resolution, and

 The most appropriate cofactors is selected as a template

February 10, 2025 Bioinformatics 67


Protein 3D Structure Prediction
Step2:Sequence Alignment
Once the structure with the highest sequence similarity is identified as a

template,
 The full-length sequences of the template and target proteins need

to be realigned using refined alignment algorithms to obtain


optimal alignment
This realignment is the most critical step in homology modeling, which

directly affects the quality of the final model


This is because incorrect alignment at this stage leads to

 Incorrect designation of homologous residues and therefore to

incorrect structural models


Errors made in the alignment step cannot be corrected in the following

modeling steps

February 10, 2025 Bioinformatics 68


Protein 3D Structure Prediction
Step2:Sequence Alignment
Therefore, the best possible multiple alignment algorithms, such as

 Praline and T-Coffee should be used for this purpose

Even alignment using the best alignment program may not be error free

and should be visually inspected to ensure that conserved key residues


are correctly aligned
If necessary, manual refinement of the alignment should be carried out

to improve alignment quality


Step3:Backbone Model Building
In optimal alignment, residues in the aligned regions of the target

protein can assume a similar structure as the template proteins, meaning


that
 The coordinates of the corresponding residues of the template

proteins can be simply copied onto the target protein


February 10, 2025 Bioinformatics 69
Protein 3D Structure Prediction
Step3:Backbone Model Building
If the two aligned residues are identical, coordinates of the side chain

atoms are copied along with the main chain atoms


If the two residues differ, only the backbone atoms can be copied

The side chain atoms are rebuilt in a subsequent procedure

In backbone modeling, it is simplest to use only one template structure

The structure with the best quality and highest resolution is normally

chosen if multiple options are available


This structure tends to carry the fewest errors

Occasionally, multiple template structures are available for modeling

In this situation, the template structures have to be optimally aligned and

superimposed before being used as templates in model building


One can either choose to use average coordinate values of the templates

or the best parts from each of the templates to model


February 10, 2025 Bioinformatics 70
Protein 3D Structure Prediction
Step4a:Loop Modeling
In the sequence alignment for modeling, there are often regions caused

by insertions and deletions producing gaps in sequence alignment


The gaps cannot be directly modeled, creating “holes” in the model

Closing the gaps requires loop modeling, which is a very difficult

problem in homology modeling and is also a major source of error


Loop modeling can be considered a mini–protein modeling problem by

itself
Unfortunately, there are no mature methods available that can model

loops reliably
Currently, there are two main techniques used to approach the problem:

 The database searching method and

 The Ab initio method

February 10, 2025 Bioinformatics 71


Protein 3D Structure Prediction
Step4a:Loop Modeling
The database method involves finding “spare parts” from known protein

structures in a database that


 Fit onto the two stem regions of the target protein

The stems are defined as the main chain atoms that precede and follow

the loop to be modeled


The procedure begins by measuring the orientation and distance of the

anchor regions in the stems and searching PDB for


 Segments of the same length that also match the above endpoint

conformation
Usually, many different alternative segments that fit the endpoints of the

stems are available


The best loop can be selected based on sequence similarity as well as

minimal steric clashes with the neighboring parts of the structure


February 10, 2025 Bioinformatics 72
Protein 3D Structure Prediction
Step4a:Loop Modeling
The conformation of the best matching fragments is then copied onto the

anchoring points of the stems

The Ab initio method generates many random loops and searches for the
one that does not clash with nearby side chains and
Also has reasonably low energy and φ and ψ angles in the allowable

regions in the Ramachandran plot


The specialized programs for loop modeling: FREAD, PETRA, CODA

February 10, 2025 Bioinformatics 73


Protein 3D Structure Prediction
Step4b:Side Chain Refinement
Once main chain atoms are built, the positions of side chains that are not

modeled must be determined


Modeling side chain geometry is very important in evaluating protein–

ligand interactions at active sites and protein–protein interactions at the


contact interface
A side chain can be built by searching every possible conformation at

every torsion angle of the side chain


 To select the one that has the lowest interaction energy with

neighboring atoms
However, this approach is computationally prohibitive in most cases

In fact, most current side chain prediction programs use the concept of

 Rotamers, which are favored side chain torsion angles extracted

from known protein crystal structures


February 10, 2025 Bioinformatics 74
Protein 3D Structure Prediction
Step4b:Side Chain Refinement
A collection of preferred side chain conformations is a

 Rotamer library in which the rotamers are ranked by their

frequency of occurrence
Having a rotamer library reduces the computational time significantly

because
 Only a small number of favored torsion angles are examined

In prediction of side chain conformation, only the possible rotamers

with the lowest interaction energy with nearby atoms are selected
Most modeling packages incorporate the side chain refinement function

A specialized side chain modeling program that has reasonably good

performance is
 SCWRL (sidechain placement with a rotamer library)

February 10, 2025 Bioinformatics 75


Protein 3D Structure Prediction
Step5: Model Refinement Using Energy Function
In these loop modeling and side chain modeling steps, potential energy

calculations are applied to improve the model


However, this does not guarantee that

 The entire raw homology model is free of structural irregularities

such as unfavorable bond angles, bond lengths, or close atomic


contacts
These kinds of structural irregularities can be corrected by

 Applying the energy minimization procedure on the entire model,

which moves the atoms in such a way that the overall


conformation has the lowest energy potential
The goal of energy minimization is

 To relieve steric collisions and strains without significantly

altering the overall structure


February 10, 2025 Bioinformatics 76
Protein 3D Structure Prediction
Step5: Model Refinement Using Energy Function
However, energy minimization has to be used with caution because

 Excessive energy minimization often moves residues away from

their correct positions


Therefore, only limited energy minimization is recommended (a few

hundred iterations)
 To remove major errors, such as short bond distances and close

atomic clashes
Key conserved residues and those involved in cofactor binding have

 To be restrained if necessary during the process

Step6: Model Evaluation


The final homology model has to be evaluated to make sure that

 The structural features of the model are consistent with the

physicochemical rules
February 10, 2025 Bioinformatics 77
Protein 3D Structure Prediction
Step6: Model Evaluation
This involves checking anomalies in φ–ψ angles, bond lengths, close

contacts, and so on
Another way of checking the quality of a protein model is to implicitly

take these stereo-chemical properties into account


This is a method that detects errors by

 Compiling statistical profiles of spatial features and interaction

energy from experimentally determined structures


By comparing the statistical parameters with the constructed model,

 The method reveals which regions of a sequence appear to be

folded normally and which regions do not


If structural irregularities are found, the region is considered to have

errors and has to be further refined

February 10, 2025 Bioinformatics 78


Thank You

February 10, 2025 Bioinformatics 79

You might also like