0% found this document useful (0 votes)
4 views20 pages

Protein Structure Overview

Structural genomics focuses on the high-throughput determination of protein structures, emphasizing the relationship between amino acid sequences and their three-dimensional configurations. Proteins, made from 20 different amino acids, exhibit a hierarchical structure comprising primary, secondary, tertiary, and quaternary levels, each defined by unique interactions and arrangements. Understanding these structures is crucial for predicting protein function and interactions within biological systems.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views20 pages

Protein Structure Overview

Structural genomics focuses on the high-throughput determination of protein structures, emphasizing the relationship between amino acid sequences and their three-dimensional configurations. Proteins, made from 20 different amino acids, exhibit a hierarchical structure comprising primary, secondary, tertiary, and quaternary levels, each defined by unique interactions and arrangements. Understanding these structures is crucial for predicting protein function and interactions within biological systems.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

PROTEIN STRUCTURE
18 OVERVIEW

18.1 INTRODUCTION
Structural genomics refers to high-throughput three-dimensional structure determination and
analysis of biological macromolecules. Structural genomics concerns primarily with individual
protein domains. It is one of the major goals of bioinformatics to understand the relationship
between amino acid sequence and three-dimensional structure in proteins. If this relationship
is known, it can be used to predict the protein structure from the amino acid sequence. In this
chapter, as the first step, we will discuss the basics of protein structure.

18.2 OVERVIEW OF THE PROTEIN STRUCTURE


Proteins are macromolecules (heteropolymers) made up from 20 different L-a-amino acids,
also referred to as residues. A certain number of residues are necessary to perform a particular
biochemical function, and around 40–50 residues appear to be the lower limit for a functional
domain size. Protein sizes range from this lower limit to several hundred residues in multi-
functional proteins. Very large aggregates can be formed from protein subunits, e.g. many
thousand actin molecules assemble into an actin filament. Large protein complexes with RNA
are found in the ribosome particles, which are, in fact, ribozymes.
Proteins usually contain thousands of atoms precisely arranged in a three-dimensional
structure that is unique for each protein type. As a protein is made, it “folds” itself into a
complex, three-dimensional shape like a piece of ribbon that has been crumpled up. Each
protein has one folded shape, and consistently folds into it, usually in less than a second. This
complicated folded shape of a protein dictates how it works, and also, how it interacts with
other biomolecules.
A gene in the DNA of living cells codes the specific sequence of amino acids that make
up each protein. A protein cannot be synthesized without its mRNA being present, but a protein
can persist in the cell when its mRNA is no longer present. However, mRNA may be present
in abundance, but the message is not translated into proteins unless genetically required. There
is, thus, no good correlation between mRNA and protein in a cell at any given time.
287

ISTUDY
288  Protein Structure Overview

Protein structure is hierarchical and can be discussed in terms of four levels of


organization. Primary structure is the amino acid sequence of its polypeptide chain(s). Every
protein has a unique amino acid sequence. Secondary structure is the spatial arrangement of
the polypeptide backbone, ignoring the conformation of the side chains. Tertiary structure is
the three-dimensional structure of the entire polypeptide. Quaternary structure refers to the
three-dimensional structure of proteins that are composed of two or more polypeptide chains,
called sub-units.
Whereas the primary sequence of a protein is held together by covalent peptide bonds,
the three-dimensional structure of a protein may be held together by a variety of bonds—
hydrogen bonds, ionic bonds, covalent bonds, and hydrophobic interactions.

Structure of Amino Acids


The basic structure of an a-amino acid is quite simple. R denotes any one of the 20 possible
side chains (see Table 18.1 and Figure 18.1). The a-C atom has four different ligands
(the H is omitted in the drawing) and is thus chiral. An easy way to remember the correct
L-form is the CORN-rule—when the Ca-atom is viewed with the H in front, the residues read
“CO-R-N” in a clockwise direction.

FIGURE 18.1 Amino acid structure.

The different side chains R determine the chemical properties of the amino acid or
residue (the residue is the amino acid side chain plus the peptide backbone). There is a wide
diversity in the chemical properties of amino acid side chains, but they can be grouped into
six classes (Table 18.1).

TABLE 18.1 Classification of Amino Acids Based on Side Chains

Side Chain Amino Acids


Aliphatic (non-polar R groups), Hydroxyl- or glycine, alanine, valine, leucine, isoleucineserine,
sulphur containing (polar, uncharged R groups), cysteine, threonine (hydroxyl), methionine
aromatic (sulphur) phenylalanine, tyrosine, tryptophan
Basic (positively charged R groups), acidic histidine, lysine, arginine aspartic acid, glutamic
(negatively charged R groups) and their amides, acid (acidic) and aspargine, glutamine (amides,
cyclic having polar, uncharged R groups) proline

ISTUDY
Overview of the Protein Structure  289
Proline does not fit into any class because it is cyclic though it shares many properties with
the aliphatic group. The side chain sulphydryl groups of cysteine can undergo a reversible
oxidation reaction to form cystine, which contains a disulphide bond. The presence of
disulphide bonds in proteins is often a crucial structural feature as it plays a very important
role in biogenesis and structure of many proteins that are present in the inter-membrane space
of mitochondria.
Proteins contain several classes of weak acid groups. The ionization state of the side chain
weak acid groups controls the charge on the protein. The aromatic amino acids, tryptophan,
tyrosine and phenylalanine absorb light in the ultraviolet region (250–300 nm). Tryptophan
has the highest molar absorptivity, followed by tyrosine, with phenylalanine making only a
small contribution.
A summary of the physical properties of standard amino acids is given in Table 18.2.

TABLE 18.2 Physical Properties of Standard Amino Acids

Name (Residue) Single Relative MW pK VdW Volume Charged,


Code Abundance (Å) Polar,
(%) E.C. Hydrophobic
Alanine A 13.0 71 67 H
Arginine R 5.3 157 12.5 148 C+
Asparagine N 9.9 114 96 P
Aspartate D 9.9 114 3.9 91 C–
Cysteine C 1.8 103 86 P
Glutamate E 10.8 128 4.3 109 C-
Glutamine Q 10.8 128 114 P
Glycine G 7.8 57 48 –
Histidine H 0.7 137 6.0 118 P, C+
Isoleucine I 4.4 113 124 H
Leucine L 7.8 113 124 H
Lysine K 7.0 129 10.5 135 C+
Methionine M 3.8 131 124 H
Phenylalanine F 3.3 147 135 H
Proline P 4.6 97 90 H
Serine S 6.0 87 73 P
Threonine T 4.6 101 93 P
Tryptophan W 1.0 186 163 P
Tyrosine Y 2.2 163 10.1 141 P
Valine V 6.0 99 105 H

ISTUDY
290  Protein Structure Overview

18.3 PROTEIN STRUCTURE HIERARCHY


Protein structure can be understood at different levels—primary, secondary, tertiary and
quarternary (Figure 18.2). Each of the levels has a different way of representing the structure.
The following sections discuss each of these levels.

(a) –Arg–Val–Glu–Lys–Met–Gly–Val–Ala–
Metal ion

COOH
C H C C H


N N
C H O
N H C
C C
O
N C O
C C
H O H C
H2 N
N
C N
H
(c)

-helix
(b) (d)

FIGURE 18.2 Four levels of protein organization (a) primary structure, (b) secondary structure—
a-helix, (c) tertiary structure, and (d) quarternary structure.

Primary Structure
The primary structure of proteins comprises the sequence of amino acids linked via peptide
bonds. The amino acid sequence refers to the sequential order of a polypeptide (AA1-AA2-
AA3...AAn). The relative amount of each of the 20 AAs in a peptide (e.g., 10% glycine, 6%
tryptophan, etc.) is referred to as the composition of the protein.
Amino acids are joined together via the peptide bond that is formed by the reaction
of the a-carboxyl group of one amino acid with the a-amino group of another amino acid
(Figure 18.3). If this process is repeated many times, then a long linear chain of amino acids, a
polypeptide, is produced. As a convention, the sequence of a polypeptide is written beginning
with the residue containing the free a-amino group (the N-terminal or amino terminal) and
ending with the residue containing the free a-carboxyl group (the C-terminal or carboxyl
terminal), e.g. NH2-Glu-Gly-Ala-Lys-COOH. Due to the partial double bond character of

ISTUDY
Protein Structure Hierarchy  291
the peptide bond, the O, C, N and H atoms are nearly planar and there is no rotation about
the peptide bonds. The peptide bond is planar due to the delocalization of the electrons. The
planarity of these elements has important consequences for the three-dimensional structure of
proteins. Deviation from planarity is rare and when there are highly non-planar peptides, they
are integral components of protein structure related to local and tertiary structural features that
tend to be conserved among homologs.

R H O
R R
R
C N C
C O—H C O— H 

H—N C C O— H
H—N C H — N C

H O H O H O R

FIGURE 18.3 Primary structure of proteins.

Two amino acids are combined in a condensation reaction. The sequence of the different
amino acids is considered as the primary structure of the peptide or protein. Counting of
residues always starts at the N-terminal end (NH2-group).
In contrast to the rather rigid peptide bond angle w (always close to 180º), the bond
angles phi j and psi f can have a certain range of possible values. They are restrained by
geometry to allow ranges typical for particular secondary structure elements, and represented
in a Ramachandran plot (discussed later in the chapter). A few important bond lengths are
given in Table 18.3. Each amino acid in a protein is unique and determination of the amino
acid sequence is an important part of characterizing proteins.

TABLE 18.3 Peptide Bonds

Peptide Bond Average Length Single Bond Average Length Hydrogen Bond Average (±0.3)
C–C 1.53 (Å) C–C 1.54 (Å) O–H … O–H 2.8 (Å)
C–N 1.33 (Å) C–N 1.48 (Å) N–H … O=C 2.9 (Å)
N–C 1.46 (Å) C–O 1.43 (Å) O–H … O=C 2.8 (Å)

You can use PROWL ([Link] for getting


information about the protein primary structures (Figure 18.4). This is a site that provides
information on properties of amino acids, bond geometry they allow, probability for being
interior versus exterior residues and which amino acid substitutions are likely to maintain
structure and function. PROWL uses data from mass spectrometry. PROWL provides other
information that is used to make predictions about protein structure based on primary amino
acid sequence.

ISTUDY
292  Protein Structure Overview

FIGURE 18.4 PROWL can provide a lot of information on amino acids.

PROWL is a very useful source for obtaining information on amino acids. Table 18.4
gives the information about solubility, density and pI for various amino acids from PROWL.

TABLE 18.4 Solubility, Density and pI from PROWL

Name Solubility (g/100 g, 25ºC) Crystal Density (g/ml) pI at 25ºC


Alanine 16.65 1.401 6.107
Arginine 15 1.1 10.76
Aspartic Acid 0.778 1.66 2.98
Asparagine 3.53 1.54 –
Cysteine very – 5.02
Glutamic Acid 0.864 1.460 3.08
Glutamine 2.5 – –
(Cont.)

ISTUDY
Protein Structure Hierarchy  293

Name Solubility (g/100 g, 25ºC) Crystal Density (g/ml) pI at 25ºC


Glycine 24.99 1.607 6.064
Histidine 4.19 – 7.64
Isoleucine 4.117 – 6.038
Leucine 2.426 1.191 6.036
Lysine very – 9.47
Methionine 3.381 1.340 5.74
Phenylalanine 2.965 – 5.91
Proline 162.3 – 6.3
Serine 5.023 1.537 5.68
Threonine very – –
Tryptophan 1.136 – 5.88
Tyrosine 0.0453 1.456 5.63
Valine 8.85 1.230 6.002

Secondary Structure
Secondary structure in proteins is formed due to the hydrogen-bonding properties of the
peptide backbone, the interaction of the side-chain atoms with the backbone and the
chiral nature of amino acids. The two most common secondary structure arrangements are the
right-handed a-helix and the b-sheet, which can be connected into a larger tertiary structure
(or fold) by turns and loops of a variety of types. These two secondary structural elements
satisfy a strong hydrogen bond network within the geometric constraints of the bond angles w,
j and f. The b-sheets can be formed by parallel or, most common, antiparallel arrangement
of individual b-strands.
The peptide bond has partial double bond character that forces the OCNH atoms of the
polypeptide backbone to be planar. Thus,
the only degrees of freedom for rotation The Ramachandran Plot
in the polypeptide backbone are around 180
Beta-sheet
the bonds to the Ca carbon—phi (f) or
psi (y). However, there are significant +psi Left
limitations as to which angles of f and handed
y can be used due to steric clashes alpha-helix
between atoms. The Ramachandran plot 0
shows those regions of f and y where
there are no steric conflicts (Figure
–psi
18.5). Ramachandran plot is an integral Right handed
tool for protein structure research and alpha-helix
education.
Using PDB, we can get the –180
–180 –psi 0 +psi –180
Ramachandran plot. For example, the
Ramachandran plot for the protein FIGURE 18.5 Schematic of Ramachandran plot.
1NXB is given in Figure 18.6.

ISTUDY
294  Protein Structure Overview

FIGURE 18.6 Ramachandran plot for the protein 1NXB from PDB.

MolProbity is a structure-validation web service that provides broad-spectrum solidly based


evaluation of model quality at both the global and local levels for both proteins and nucleic
acids ([Link] Please see Figure 18.7.

FIGURE 18.7 MolProbity home page.

ISTUDY
Protein Structure Hierarchy  295
You can use PROCHECK ([Link]
.html) to get the summary of above Ramchandran plot. The PROCHECK output for 1NXB
is given as Table 18.5.

TABLE 18.5 Ramachandran Plot Statistics

No. of Residues Percentage


Most favored regions [A, B, L] 32 62.7%**
Additional allowed regions [a, b, l, p] 14 27.5%
Generously allowed regions [~a, ~b, ~l, ~p] 4 7.8%
Disallowed regions [XX] 1 2.0%*
Non-glycine and non-proline residues 51 100.0%
End-residues (excl. Gly and Pro) 2
Glycine residues 5
Proline residues 4
Total number of residues 62

Based on an analysis of 118 structures of resolution of at least 2.0 Angstroms and


R-factor no greater than 20.0 a good quality model would be expected to have over 90% in
the most favoured regions [A, B, L].
Proteins are organized with the hydrophobic side chains in the interior and the hydrophilic
side chains on the surface. Because the main chain peptide bond elements, C=O and N-H, are
polar, placing them in the hydrophobic interior of a protein could create a major problem.
This problem is solved by the formation of hydrogen bonds between the amide protons and
carbonyl oxygens of the peptide bonds in the main chain. These hydrogen bonds cause the
main chain to adopt a-helix and b-sheet conformations.
In the a-helix, the protein twists to form a tightly wound spiral with the R groups facing
out from the central axis. The helical structure is held together by extensive formation of
hydrogen bonds that form between the carbonyl oxygens and amino hydrogens from residues
four places further along the chain. In the b-conformation (or b-pleated sheet), the b-pleated
sheet consists of peptide chains running side by side with hydrogen bonds holding the chains
together. The sheet is most stable with the chains in an anti-parallel array, although a parallel
arrangement may be seen as well. In the a-helix, intra-chain hydrogen bonds are formed
between the peptide bond elements four residues apart in the primary sequence. The peptide
bond has a dipole moment. Because the peptide bond has a dipole moment arising from the
polarity of the NH and C=O groups, the helix itself has a dipole moment (that runs the length
of the helix with the amino terminal of the helix carrying a partial positive charge and the
carboxyl terminal a partial negative charge). The helix dipole moment plays an important role
in binding charged ligands to proteins.
Some of the common protein secondary structures are discussed here.

ISTUDY
296  Protein Structure Overview

Alpha-helix
Alpha-helix, the right-handed helix is the most common arrangement. Alpha-helix results from
the conformation of peptide bonds and R-groups seeking the most stable arrangement. There
are 3.6 residues/turn in right-handed helix. The periodicity or distance that the helix rises
per turn is 0.54 nm. The rise per amino acid is 0.15 nm. H-bonding between every fourth
peptide bond (amide and carboxyl) stabilizes alpha-helix. In helical region of a protein, every
peptide bond, therefore, participates in H-bonding, which leads to a very stable structure. The
consequence of the structure of the alpha-helix are those amino acids which are three to four
amino acids apart in primary structure are spatially close to each other in the helix. Hydrogen
bonding is the strongest when groups are in straight line which is the configuration in alpha-
helix. Tight packing within the core of the helix also stabilizes the helix.
One method of visualization is by using Kinemages ([Link]). Kinemage 1 is
amphipathic. The other method is using Chemscape Chime Molecular Visualization.
Agadir ([Link] predicts the helical behaviour of monomeric peptides.
It only considers short range interactions. Conditions such as pH, temperature and ionic
strength are used in the calculation. Modifications of the termini are also allowed.

Beta-sheet
Beta-strands are typically 5–8 amino acid residues that form an extended, chain of silk/spider
webs. They form a fully extended structure. Adjacent amino acids are 3.5 Å away (compared
to only 1.5 for alpha-helix). These strands can form sheets when two strands are grouped
and hydrogen bonded. Kinemage anti-parallel beta-sheet can be used to visualize beta-sheets.
Arrangements of b-sheets can be:
1. Anti-parallel, which have amino-terminal end at one end of the polypeptide
oriented toward carboxy-terminus of polypeptide. Anti-parallel is the most common
arrangement.
2. Parallel, which have amino and carboxy-termini of peptide chains oriented in the same
direction. They are often grouped into structural units of 2–5 sheets. Beta-sheets can
have a mix of parallel and anti-parallel components (Figure 18.8).

(a) (b) (c)


FIGURE 18.8 (a) Parallel, (b) anti-parallel, and (c) mixed arrangements of beta-sheets.

ISTUDY
Protein Structure Hierarchy  297

Loops and Turns


Most proteins contain combinations of a-helices and b-sheets, which are connected by loops.
Loops have irregular lengths and shapes, unlike the a-helices and b-sheets and are on the
surface of the protein. The elements of the loop do not usually form hydrogen bonds to each
other, but form hydrogen bonds to water. Loops are a diverse class of secondary structures
and are comprised of turns, coils and strands, which connect the main secondary structures.
Loop regions that connect two anti-parallel b-strands are called hairpin loops or turns.
Turns are common secondary structure elements in proteins and are responsible for changing
the direction of the polypeptide chain within a globular protein. Turns are classified into Types
I and II according to the (phi, psi) angles of the two central residues. Since a beta-turn has
several unsaturated backbone hydrogen bond donors and acceptors, it is polar, and is usually
found near the surface of the protein. Proline is very common in beta-turns, as it always has
the correct phi angle (–60º) and it has one less unsaturated hydrogen donor.
SuperLooper ([Link]
show=loop_db) searches a comprehensive compilation of protein segments, which integrates
two databases—loops in proteins (compiled from PDB) and loops in membrane proteins
(LIMP).
NetTurnP ([Link] predicts if an amino acid is
located in a beta-turn or not.

Motifs
A motif is a recurring thematic element, i.e., it is found in many molecules not uniquely in just
one. Motifs can be seen in both molecular sequences and structures and provide a means to
bridge these views of molecules. An example of a DNA motif is a repressor binding sequence
that occurs in many locations as part of its regulatory function, and it is centrally important
because it mediates the interaction between the repressor molecule and the gene it regulates.
Examples of protein motifs range from sub-structures such as the calcium binding EF-hand, to
whole domains or “folds” such as the globin fold.
A basic helix-loop-helix (bHLH) is a protein structural motif that characterizes one of
the largest families of dimerizing transcription. The motif is characterized by two a-helices
connected by a loop. 1MDY myogenic protein is shown to have HLH DNA binding motif
(Figure 18.9).
Another motif is the Greek key motif. This motif consists of four adjacent anti-parallel
strands and their linking loops. It consists of three anti-parallel strands connected by hairpins,
while the fourth is adjacent to the first and linked to the third by a longer loop. This type of
structure forms easily during the protein folding process. An example is 3HZ2 protein which
has a beta/gamma Greek key (Figure 18.10).
While the profile method can be used with either nucleic acids or proteins, the intimate
coupling between sequence and structure in proteins makes the profile method especially well-
suited to describe protein motifs. When one looks at sequence alignments of protein families, a
distinctive pattern of conserved (slowly evolving) regions and unconserved (rapidly evolving)
regions can be easily seen.

ISTUDY
298  Protein Structure Overview

FIGURE 18.9 HLH DNA binding motif in 1MDY.

FIGURE 18.10 Beta/Gamma key motif in 3HZ2.

Profiles provide a means to describe the sequence features that are important to a motif.
This information is encoded in a two-dimensional table in which the rows correspond to the
positions in the motif, and the columns to all of the possible nucleotide bases or amino acid
residues that could be found at each position. A profile is similar to a multiple sequence
alignment where the rows are the aligned positions and the columns tabulate the frequencies
of the residues or bases at each position.
Despite that there are about 100,000 different proteins expressed in eukaryotic systems,
there are much fewer different structural motifs and folds, partly as a consequence of evolved
pathways and mechanisms. Motif refers to a small specific combination of secondary structure
elements (such as helix-turn-helix), and not to the contents of the asymmetric unit cell as used
as a crystallographic term.

ISTUDY
Protein Structure Hierarchy  299
A typical small motif is the calcium binding EF-hand in calmodulin, a ubiquitous
molecule undergoing Ca-dependent conformational changes. It contains four Ca++ ions that
are coordinated in a typical fashion in a helix-turn-helix motif called the EF-hand.
The positively charged calcium atom Ca++ is coordinated through hydrogen bonds with
acidic (negatively charged) aspartate and glutamate residues as well as with backbone oxygen
atoms. Other typical motif in the alpha-domain structures is the four-helix bundle. Ferritin,
Cytochrome b562 and apo-E are typical examples. The helices are amphipathic—hydrophobic
residues on one side, charged ones on the other and pack anti-parallel with the hydrophobic
sides towards each other forming a hydrophobic core.
You can get more information on the EF-hand calcium binding proteins from the EF-
hand calcium-binding proteins data library ([Link]
[Link]).
The use of motifs in protein function prediction is discussed in Chapter 19.

Random Coils
Coils comprise amino acid residues in chains that do not lead to a consistent secondary structure
or disrupt alpha-helix or beta-sheet. Coils can form due to electrostatic repulsion or bulkiness
of R-groups that disrupt structure. This can happen due to amino acids like proline, where
N is the part of a rigid ring that prevents rotation of the ring carbons and leads to a “kink”.

Coiled Coils
A random coil is not truly a secondary structure and has no specific shape. However, a coiled
coil is a motif in which 2–7 alpha helices are coiled together. Coiled coils are formed by about
3–5% of all amino acids in proteins. A coiled coil consists of two to five helices wrapped
around each other into a left-handed helix which forms a supercoil. Many coiled coils are
important from biological functions perspective, e.g. regulation of gene expression.
Coils ([Link] is a program that
compares a sequence to a database of known parallel two-stranded coiled coils and derives a
similarity score. After comparing this score to the distribution of scores in globular and coiled
coil proteins, the program then calculates the probability that the sequence will adopt a coiled
coil conformation.
Paircoil ([Link] predicts the location of coiled
coil regions in amino acid sequences.
MultiCoil program ([Link] predicts the location
of coiled coil regions in amino acid sequences and classifies the predictions as dimeric or
trimeric. The method is based on the PairCoil algorithm and extends the two-stranded coiled
coil PairCoil to the identification of three-stranded coiled coils.

Folds
A fold is an arrangement of the secondary structure elements. Fold assignment follows after
assignment of the secondary structure. Proteins are defined as having a common fold if they
have the same secondary structures in the same arrangement with a similar topology.

ISTUDY
300  Protein Structure Overview

Superfamilies and families are defined as having a common fold if their proteins have
same major secondary structures in same arrangement with the same topological connections.
Different proteins, but having same fold, usually have peripheral elements of secondary
structure. When proteins are placed together in the same fold category, there are similarities
in structure which probably are a result of the physics and chemistry of proteins.
Fold analysis can reveal evolutionary relationships. A hierarchical clustering can provide
a dendrogram of fold neighbors at various hierarchical levels of structural similarity.
Fold analysis can also help in better understanding of protein function and activity.
Many cell functions depend on the ability of proteins to fold correctly and misstep in folding
can lead to disease. For example, the proteins catalyze disulphide bond formation, which is a
step in the oxidative folding pathway and takes place in endoplasmic reticulum in eukaryotes
and in the bacterial periplasm in the prokaryotes. The structures in prokaryotes and eukaryotes
have several similarities and perform a similar function.
Finally, fold analysis can help in the design of new proteins with pre-defined structure
and activity. This can happen in many cases where a common evolutionary pattern gets
obscured by the extent of the divergence. In such cases, it is possible that the discovery of
new structures, with folds between those of the previously known structures, will make clear
their common evolutionary relationship.
There is considerable interest in assigning a sequence to a folding class. There are two
approaches used for this purpose. One approach uses threading algorithms which, given a
group of structures and a sequence, identify the structure that is closest to this sequence. The
second approach is taxonometric and assumes that the number of folds is restricted, and hence,
the objective is to predict in the context of a particular classification of three-dimensional
folds.
CATH ([Link] provides evolutionary relationships for protein
domains—it is a hierarchical classification of protein domains based on their folding patterns.
It consists of class, architecture, topology (folds) and homologous superfamily (Figures 18.11
and 18.12).

FIGURE 18.11 CATH hierarchy.

ISTUDY
Protein Structure Hierarchy  301

FIGURE 18.12 CATH statistics.

More than such 1000 unique folds have been identified.


Structural classification of proteins or SCOP—([Link]
provides the details of proteins possessing the same fold in these classifications. SCOP does
the classification based on the protein domains. Small proteins and most the proteins of
medium size have a single domain and are treated as a whole. Domains in large proteins are
classified individually.
SCOP’s hierarchy is family, superfamily, common fold and class. Class is based on
grouping of different folds. Most of the folds are assigned to one of the five structural classes
based on their secondary structures. These four classes are given in Table 18.6. The fifth class
is multi-domain—for those with domains of different fold and for which no homologs are
known at present.

TABLE 18.6 Protein Structure Classification

Type Description Sub-types/Examples


Alpha or Predominantly alpha-helices Bundle and non-bundle
Beta or Predominantly beta-sheets Single sheet
Roll Barrel
Sandwich
Predominantly alternating motifs and combinations of this
Alpha/Beta or a/b alpha-helix and beta-strand motif (usually parallel b-sheets).
Typically, b-sheets are found in the
interior of proteins while alpha-helixes are
found on the exterior of proteins.
Alpha and beta Alpha-helices and beta-strand Anti-parallel -sheets, segregated and
or + regions as separate groupings -regions.
Helices typically on one side of the sheet.

ISTUDY
302  Protein Structure Overview

Steric properties of amino acids can greatly affect the local structures that a protein
adopts. The best example of this is perhaps proline. Proline exhibits reduced torsional freedom
with the main chain phi angle fixed. This leads to the conformation of the peptide backbone
being locked in a turn, and with the loss of a hydrogen bonding in the N, this leads to the
residue often appearing on the surface loops of a protein.
Other particular types of secondary structure that have reasonable probabilities of having
particular types of amino acid that lend themselves to those structural configurations are given
in Table 18.7. However, one of the difficulties in protein structure prediction is that the context
in which amino acids find themselves in a protein has a large effect on their conformation
(phi and psi rotation), and thus, on their propensity to form particular secondary structures.

TABLE 18.7 Structural Configurations for Some Amino Acids

Structural Configuration Amino Acid Comments


Propensity for alpha-helix Ala, Leu, Glu Amino acids without much bulk, or polar
atoms, close to the alpha-carbon
Disrupt alpha-helix Ser, Asp, Asn
Propensity for -sheets and Val and Ile Side chains project out of plane and thus
disrupt alpha-helix conducive to sheets
Propensity for turns Gly, Pro, Asp

Families of structurally similar proteins (FSSP) database provides an automated


classification scheme that employs exhaustive structure-to-structure alignment of proteins
using the DALI alignment methods ([Link]
ProFold ([Link] is a fold classifier combining the protein
structural and functional information.

Tertiary Structure
The tertiary structure of protein is a three-dimensional representation of proteins. This
representation consists of a single polypeptide chain, forming the backbone of the structure
with one or more protein secondary structures, called the protein domains. The interactions
of side chains within a particular protein determine its tertiary structure. A number of tertiary
structures may fold into a quaternary structure. Generally, the information for protein structure
is contained within the amino acid sequence of the protein itself.
SWISS-MODEL ([Link] is an automated system for modeling
the three-dimensional structure of a protein from its amino acid sequence using homology
modeling techniques.
Tertiary structure represents spatial arrangement of amino acids that are far apart in
the linear sequence. It results from interactions between R-groups via van der Waals, ionic,
hydrophobic and hydrogen bonding. Groupings or arrangements of secondary structures into
combinations that are present in a variety of proteins are often described as motifs or domains.
Examples include:

ISTUDY
Protein Structure Hierarchy  303
™ Parallel and anti-parallel -sheets (also known as the hairpin beta-motif)
™ Greek key motif—four adjacent anti-parallel b-strands
™ Beta–alpha–beta motif (can be either two parallel or anti-parallel b)

Domains
Polypeptide chains of >200 amino acids that fold into two or more compact globular clusters
are called domains. There are three main types of domains:
™ a–domains are composed of a-helices.
™ b–domains contain anti-parallel b-sheets, and usually, contain two b-sheets packed
against each other.
™ a/b-domains contain the b-a-b motif of parallel b-sheets.
Adjacent domains are connected by one or two segments of the polypeptide chain.
All globular proteins have a defined structure. In the tertiary structure of proteins,
almost all the hydrophobic side chains are found in the interior of the protein and
almost all hydrophilic side chains are found on the outside of the protein, interacting with
water. Globular proteins are compact and there is no space inside. Hence, water is effectively
excluded from the hydrophobic interior. Nearly all hydrogen bond donors, e.g. Ser, form
hydrogen bonds with hydrogen bond acceptors like Gln. Hydrogen bond formation neutralizes
the polarity of the hydrogen-bonding group. b-sheets are usually twisted or wrapped into
barrel structures and loops and turns are on the outside of the protein. Mutations, which
place a hydrophobic side chain on the surface, can cause significant changes in the folding of
the protein.

Quaternary Structure
Many large proteins contain more than one polypeptide chain. The spatial arrangement of
these subunits is the quaternary structure of the protein (e.g., hemoglobin). The forces that
hold sub-units together are the same weak bonds as those that stabilize the tertiary structure
of proteins—van der Waals and London dispersion forces, salt bridges and hydrogen bonds.
The contact region between sub-units resembles the interior of a protein. The sub-units may
be identical or non-identical, but there is always a defined stoichiometry.
Several proteins function by forming transient complexes with other proteins. For example,
electron transfer from cytochrome c to cytochrome c peroxidase involves the formation of
such a complex. Because such complexes are transient, the interactions are weaker than
those that stabilize the quaternary structure of hemoglobin. In the cytochrome c-cytochrome
c peroxidase complex, there are specific electrostatic interactions that guide the two proteins
together in the correct orientation. The quaternary structure is not required for all proteins to
be functional; many proteins may have only 2º or 3º structure.

ISTUDY
304  Protein Structure Overview

18.4 DOMAIN ARCHITECTURE DATABASES


Conserved domain database or CDD ([Link]
Structure/cdd/[Link])
Conserved domain database (CDD) on NCBI is a collection of sequence alignments
and profiles, representing protein domains conserved in molecular evolution. Proteins often
contain several modules or domains, each with a distinct evolutionary origin and function.
The CD-search service may be used to identify the conserved domains present in a protein
sequence.
It uses the following domain definition:
™ A compact sub-structure of a protein based on the three-dimensional fold
™ Defined without regard to the sequence conservation shared with other members of
protein families
™ Defined without regard to sequence continuity (i.e., topology)
Conserved domains are defined based on recurring sequence patterns or motifs. CDD
contains domains derived from Simple Modular Architecture Research Tool (Smart —http://
[Link]/), Pfam ([Link] and COG ([Link]
[Link]/COG/). The source databases also provide descriptions and links to citations.
Since conserved domains correspond to compact structural units, CDs contain links to three-
dimensional structure via Cn3-D whenever possible.
To identify conserved domains in a protein sequence, the CD-search service employs the
reverse position-specific BLAST algorithm. The query sequence is compared to a position-
specific score matrix prepared from the underlying conserved domain alignment. Hits are
displayed as a pair-wise alignment of the query sequence with a representative domain
sequence, or as a multiple alignment. CD-search is then run by default in parallel with protein
BLAST searches.

CDART ([Link]
.cgi?cmd=rps)
CDART is the conserved domain architecture retrieval tool that can be used to search for
proteins with similar domain architectures. CDART uses pre-computed CD-search results to
quickly identify proteins with a set of domains similar to that of the query.

18.5 PROTEIN CLASSIFICATION APPROACHES


Several databases attempt to use structural similarities of proteins for their classification.
Proteins are classified to reflect both structural and evolutionary relatedness. Many levels
exist in the hierarchy, but the principal levels are family, superfamily and fold. These are
described below.

ISTUDY
Protein Classification Approaches  305
The demarcation of boundaries between these levels is to some degree subjective. The
evolutionary classification is generally conservative—where any doubt about relatedness exists,
we made new divisions at the family and superfamily levels. Thus, some researchers may
prefer to focus on the higher levels of the classification tree, where proteins with structural
similarity are clustered.
The different major levels in the hierarchy are:

Family: Clear Evolutionarily Relationship


Proteins clustered together into families are clearly evolutionarily related. Generally, this
means that pair-wise residue identities between the proteins are 30% and greater. However, in
some cases, similar functions and structures provide definitive evidence of common descent in
the absence of high sequence identity, e.g. many globins form a family though some members
have sequence identities of only 15%.

Superfamily: Probable Common Evolutionary Origin


Proteins that have low sequence identities, but whose structural and functional features
suggest that a common evolutionary origin is probable, are placed together in superfamilies.
For example, actin, the ATPase domain of the heat shock protein, and hexakinase together
form a superfamily.

Fold: Major Structural Similarity


As discussed earlier, proteins are defined as having a common fold if they have the same
major secondary structures in the same arrangement and with the same topological connections.
Different proteins with the same fold often have peripheral elements of secondary structure
and turn regions that differ in size and conformation. In some cases, these differing peripheral
regions may comprise half the structure. Proteins placed together in the same fold category
may not have a common evolutionary origin—the structural similarities could arise just from
the physics and chemistry of proteins favouring certain packing arrangements and chain
topologies.
SCOP and CATH are two main database approaches used in protein classification,
briefly discussed below.

SCOP
SCOP provides a broad survey of all known protein folds, detailed information about the close
relatives of any particular protein, and a framework for future research and classification. The
SCOP database aims to provide a detailed and comprehensive description of the structural and
evolutionary relationships between all proteins whose structure is known, including all entries
in the protein data bank. It is available as a set of tightly linked hypertext documents, which
make the large database comprehensible and accessible.

ISTUDY
306  Protein Structure Overview

CATH
CATH provides a hierarchical classification of protein domain structures, which clusters
proteins at four major levels. The homologous superfamilies cluster proteins with highly similar
structures and functions. The assignments of structures to topology families and homologous
superfamilies are made by sequence and structure comparisons.

SUMMARY

In this chapter, you have learned the fundamentals of protein structure. You have also learned
how to analyze the physical properties of proteins. The visualization of protein structure and
classification tools for proteins have also been discussed in detail. Using the discussion in this
chapter as the basis, you have learned how to predict protein structure and function from the
sequence information.

REVIEW QUESTIONS

1. Which tools are recommended for obtaining information about the primary structure of
protein?
2. How can you analyze physical properties of proteins?
3. Is it possible to predict the structure and function of a protein from sequence information?
4. Find out by sequence analysis whether the following protein contains multiple,
independently evolving, domains.
>SCYKLO40C * NFUI NfulpqSGD:SOOOI523 * GI:6322811 * chromosome 11
MFKSV AKLGK SPIFYLNSQR LIHIKTL TTP NENALKFLST DGEMLQTRGS
KSIVIKNillE NLINHSKLAQ QIFLQCPGVE SLMIGDDFL T INKDRMVHWN
SIKPEIIDLL TKQLAYGEDV ISKEFHAVQE EEGEGGYKIN MPKFELTEED
EEVSELIEEL IDTRIRP AIL EDGGDIDYRG WDPKTGTVYL RLQGACTSCS
SSEVTLKYGI ESMLKHYVDE VKEVIQIMDP EQEIALKEFD KLEKKLESSK
NTSHEK
5. Calmodulin is a protein that assists in modulation of other protein activity. Its structure
is variable and is modulated by calcium. Find the three-dimensional structure of
calmodulin in PDB. You can visualize these models using RasMol.

ISTUDY

You might also like