Protein Structure Overview
Protein Structure Overview
PROTEIN STRUCTURE
18 OVERVIEW
18.1 INTRODUCTION
Structural genomics refers to high-throughput three-dimensional structure determination and
analysis of biological macromolecules. Structural genomics concerns primarily with individual
protein domains. It is one of the major goals of bioinformatics to understand the relationship
between amino acid sequence and three-dimensional structure in proteins. If this relationship
is known, it can be used to predict the protein structure from the amino acid sequence. In this
chapter, as the first step, we will discuss the basics of protein structure.
ISTUDY
288 Protein Structure Overview
The different side chains R determine the chemical properties of the amino acid or
residue (the residue is the amino acid side chain plus the peptide backbone). There is a wide
diversity in the chemical properties of amino acid side chains, but they can be grouped into
six classes (Table 18.1).
ISTUDY
Overview of the Protein Structure 289
Proline does not fit into any class because it is cyclic though it shares many properties with
the aliphatic group. The side chain sulphydryl groups of cysteine can undergo a reversible
oxidation reaction to form cystine, which contains a disulphide bond. The presence of
disulphide bonds in proteins is often a crucial structural feature as it plays a very important
role in biogenesis and structure of many proteins that are present in the inter-membrane space
of mitochondria.
Proteins contain several classes of weak acid groups. The ionization state of the side chain
weak acid groups controls the charge on the protein. The aromatic amino acids, tryptophan,
tyrosine and phenylalanine absorb light in the ultraviolet region (250–300 nm). Tryptophan
has the highest molar absorptivity, followed by tyrosine, with phenylalanine making only a
small contribution.
A summary of the physical properties of standard amino acids is given in Table 18.2.
ISTUDY
290 Protein Structure Overview
(a) –Arg–Val–Glu–Lys–Met–Gly–Val–Ala–
Metal ion
COOH
C H C C H
–
N N
C H O
N H C
C C
O
N C O
C C
H O H C
H2 N
N
C N
H
(c)
-helix
(b) (d)
FIGURE 18.2 Four levels of protein organization (a) primary structure, (b) secondary structure—
a-helix, (c) tertiary structure, and (d) quarternary structure.
Primary Structure
The primary structure of proteins comprises the sequence of amino acids linked via peptide
bonds. The amino acid sequence refers to the sequential order of a polypeptide (AA1-AA2-
AA3...AAn). The relative amount of each of the 20 AAs in a peptide (e.g., 10% glycine, 6%
tryptophan, etc.) is referred to as the composition of the protein.
Amino acids are joined together via the peptide bond that is formed by the reaction
of the a-carboxyl group of one amino acid with the a-amino group of another amino acid
(Figure 18.3). If this process is repeated many times, then a long linear chain of amino acids, a
polypeptide, is produced. As a convention, the sequence of a polypeptide is written beginning
with the residue containing the free a-amino group (the N-terminal or amino terminal) and
ending with the residue containing the free a-carboxyl group (the C-terminal or carboxyl
terminal), e.g. NH2-Glu-Gly-Ala-Lys-COOH. Due to the partial double bond character of
ISTUDY
Protein Structure Hierarchy 291
the peptide bond, the O, C, N and H atoms are nearly planar and there is no rotation about
the peptide bonds. The peptide bond is planar due to the delocalization of the electrons. The
planarity of these elements has important consequences for the three-dimensional structure of
proteins. Deviation from planarity is rare and when there are highly non-planar peptides, they
are integral components of protein structure related to local and tertiary structural features that
tend to be conserved among homologs.
R H O
R R
R
C N C
C O—H C O— H
H—N C C O— H
H—N C H — N C
H O H O H O R
Two amino acids are combined in a condensation reaction. The sequence of the different
amino acids is considered as the primary structure of the peptide or protein. Counting of
residues always starts at the N-terminal end (NH2-group).
In contrast to the rather rigid peptide bond angle w (always close to 180º), the bond
angles phi j and psi f can have a certain range of possible values. They are restrained by
geometry to allow ranges typical for particular secondary structure elements, and represented
in a Ramachandran plot (discussed later in the chapter). A few important bond lengths are
given in Table 18.3. Each amino acid in a protein is unique and determination of the amino
acid sequence is an important part of characterizing proteins.
Peptide Bond Average Length Single Bond Average Length Hydrogen Bond Average (±0.3)
C–C 1.53 (Å) C–C 1.54 (Å) O–H … O–H 2.8 (Å)
C–N 1.33 (Å) C–N 1.48 (Å) N–H … O=C 2.9 (Å)
N–C 1.46 (Å) C–O 1.43 (Å) O–H … O=C 2.8 (Å)
ISTUDY
292 Protein Structure Overview
PROWL is a very useful source for obtaining information on amino acids. Table 18.4
gives the information about solubility, density and pI for various amino acids from PROWL.
ISTUDY
Protein Structure Hierarchy 293
Secondary Structure
Secondary structure in proteins is formed due to the hydrogen-bonding properties of the
peptide backbone, the interaction of the side-chain atoms with the backbone and the
chiral nature of amino acids. The two most common secondary structure arrangements are the
right-handed a-helix and the b-sheet, which can be connected into a larger tertiary structure
(or fold) by turns and loops of a variety of types. These two secondary structural elements
satisfy a strong hydrogen bond network within the geometric constraints of the bond angles w,
j and f. The b-sheets can be formed by parallel or, most common, antiparallel arrangement
of individual b-strands.
The peptide bond has partial double bond character that forces the OCNH atoms of the
polypeptide backbone to be planar. Thus,
the only degrees of freedom for rotation The Ramachandran Plot
in the polypeptide backbone are around 180
Beta-sheet
the bonds to the Ca carbon—phi (f) or
psi (y). However, there are significant +psi Left
limitations as to which angles of f and handed
y can be used due to steric clashes alpha-helix
between atoms. The Ramachandran plot 0
shows those regions of f and y where
there are no steric conflicts (Figure
–psi
18.5). Ramachandran plot is an integral Right handed
tool for protein structure research and alpha-helix
education.
Using PDB, we can get the –180
–180 –psi 0 +psi –180
Ramachandran plot. For example, the
Ramachandran plot for the protein FIGURE 18.5 Schematic of Ramachandran plot.
1NXB is given in Figure 18.6.
ISTUDY
294 Protein Structure Overview
FIGURE 18.6 Ramachandran plot for the protein 1NXB from PDB.
ISTUDY
Protein Structure Hierarchy 295
You can use PROCHECK ([Link]
.html) to get the summary of above Ramchandran plot. The PROCHECK output for 1NXB
is given as Table 18.5.
ISTUDY
296 Protein Structure Overview
Alpha-helix
Alpha-helix, the right-handed helix is the most common arrangement. Alpha-helix results from
the conformation of peptide bonds and R-groups seeking the most stable arrangement. There
are 3.6 residues/turn in right-handed helix. The periodicity or distance that the helix rises
per turn is 0.54 nm. The rise per amino acid is 0.15 nm. H-bonding between every fourth
peptide bond (amide and carboxyl) stabilizes alpha-helix. In helical region of a protein, every
peptide bond, therefore, participates in H-bonding, which leads to a very stable structure. The
consequence of the structure of the alpha-helix are those amino acids which are three to four
amino acids apart in primary structure are spatially close to each other in the helix. Hydrogen
bonding is the strongest when groups are in straight line which is the configuration in alpha-
helix. Tight packing within the core of the helix also stabilizes the helix.
One method of visualization is by using Kinemages ([Link]). Kinemage 1 is
amphipathic. The other method is using Chemscape Chime Molecular Visualization.
Agadir ([Link] predicts the helical behaviour of monomeric peptides.
It only considers short range interactions. Conditions such as pH, temperature and ionic
strength are used in the calculation. Modifications of the termini are also allowed.
Beta-sheet
Beta-strands are typically 5–8 amino acid residues that form an extended, chain of silk/spider
webs. They form a fully extended structure. Adjacent amino acids are 3.5 Å away (compared
to only 1.5 for alpha-helix). These strands can form sheets when two strands are grouped
and hydrogen bonded. Kinemage anti-parallel beta-sheet can be used to visualize beta-sheets.
Arrangements of b-sheets can be:
1. Anti-parallel, which have amino-terminal end at one end of the polypeptide
oriented toward carboxy-terminus of polypeptide. Anti-parallel is the most common
arrangement.
2. Parallel, which have amino and carboxy-termini of peptide chains oriented in the same
direction. They are often grouped into structural units of 2–5 sheets. Beta-sheets can
have a mix of parallel and anti-parallel components (Figure 18.8).
ISTUDY
Protein Structure Hierarchy 297
Motifs
A motif is a recurring thematic element, i.e., it is found in many molecules not uniquely in just
one. Motifs can be seen in both molecular sequences and structures and provide a means to
bridge these views of molecules. An example of a DNA motif is a repressor binding sequence
that occurs in many locations as part of its regulatory function, and it is centrally important
because it mediates the interaction between the repressor molecule and the gene it regulates.
Examples of protein motifs range from sub-structures such as the calcium binding EF-hand, to
whole domains or “folds” such as the globin fold.
A basic helix-loop-helix (bHLH) is a protein structural motif that characterizes one of
the largest families of dimerizing transcription. The motif is characterized by two a-helices
connected by a loop. 1MDY myogenic protein is shown to have HLH DNA binding motif
(Figure 18.9).
Another motif is the Greek key motif. This motif consists of four adjacent anti-parallel
strands and their linking loops. It consists of three anti-parallel strands connected by hairpins,
while the fourth is adjacent to the first and linked to the third by a longer loop. This type of
structure forms easily during the protein folding process. An example is 3HZ2 protein which
has a beta/gamma Greek key (Figure 18.10).
While the profile method can be used with either nucleic acids or proteins, the intimate
coupling between sequence and structure in proteins makes the profile method especially well-
suited to describe protein motifs. When one looks at sequence alignments of protein families, a
distinctive pattern of conserved (slowly evolving) regions and unconserved (rapidly evolving)
regions can be easily seen.
ISTUDY
298 Protein Structure Overview
Profiles provide a means to describe the sequence features that are important to a motif.
This information is encoded in a two-dimensional table in which the rows correspond to the
positions in the motif, and the columns to all of the possible nucleotide bases or amino acid
residues that could be found at each position. A profile is similar to a multiple sequence
alignment where the rows are the aligned positions and the columns tabulate the frequencies
of the residues or bases at each position.
Despite that there are about 100,000 different proteins expressed in eukaryotic systems,
there are much fewer different structural motifs and folds, partly as a consequence of evolved
pathways and mechanisms. Motif refers to a small specific combination of secondary structure
elements (such as helix-turn-helix), and not to the contents of the asymmetric unit cell as used
as a crystallographic term.
ISTUDY
Protein Structure Hierarchy 299
A typical small motif is the calcium binding EF-hand in calmodulin, a ubiquitous
molecule undergoing Ca-dependent conformational changes. It contains four Ca++ ions that
are coordinated in a typical fashion in a helix-turn-helix motif called the EF-hand.
The positively charged calcium atom Ca++ is coordinated through hydrogen bonds with
acidic (negatively charged) aspartate and glutamate residues as well as with backbone oxygen
atoms. Other typical motif in the alpha-domain structures is the four-helix bundle. Ferritin,
Cytochrome b562 and apo-E are typical examples. The helices are amphipathic—hydrophobic
residues on one side, charged ones on the other and pack anti-parallel with the hydrophobic
sides towards each other forming a hydrophobic core.
You can get more information on the EF-hand calcium binding proteins from the EF-
hand calcium-binding proteins data library ([Link]
[Link]).
The use of motifs in protein function prediction is discussed in Chapter 19.
Random Coils
Coils comprise amino acid residues in chains that do not lead to a consistent secondary structure
or disrupt alpha-helix or beta-sheet. Coils can form due to electrostatic repulsion or bulkiness
of R-groups that disrupt structure. This can happen due to amino acids like proline, where
N is the part of a rigid ring that prevents rotation of the ring carbons and leads to a “kink”.
Coiled Coils
A random coil is not truly a secondary structure and has no specific shape. However, a coiled
coil is a motif in which 2–7 alpha helices are coiled together. Coiled coils are formed by about
3–5% of all amino acids in proteins. A coiled coil consists of two to five helices wrapped
around each other into a left-handed helix which forms a supercoil. Many coiled coils are
important from biological functions perspective, e.g. regulation of gene expression.
Coils ([Link] is a program that
compares a sequence to a database of known parallel two-stranded coiled coils and derives a
similarity score. After comparing this score to the distribution of scores in globular and coiled
coil proteins, the program then calculates the probability that the sequence will adopt a coiled
coil conformation.
Paircoil ([Link] predicts the location of coiled
coil regions in amino acid sequences.
MultiCoil program ([Link] predicts the location
of coiled coil regions in amino acid sequences and classifies the predictions as dimeric or
trimeric. The method is based on the PairCoil algorithm and extends the two-stranded coiled
coil PairCoil to the identification of three-stranded coiled coils.
Folds
A fold is an arrangement of the secondary structure elements. Fold assignment follows after
assignment of the secondary structure. Proteins are defined as having a common fold if they
have the same secondary structures in the same arrangement with a similar topology.
ISTUDY
300 Protein Structure Overview
Superfamilies and families are defined as having a common fold if their proteins have
same major secondary structures in same arrangement with the same topological connections.
Different proteins, but having same fold, usually have peripheral elements of secondary
structure. When proteins are placed together in the same fold category, there are similarities
in structure which probably are a result of the physics and chemistry of proteins.
Fold analysis can reveal evolutionary relationships. A hierarchical clustering can provide
a dendrogram of fold neighbors at various hierarchical levels of structural similarity.
Fold analysis can also help in better understanding of protein function and activity.
Many cell functions depend on the ability of proteins to fold correctly and misstep in folding
can lead to disease. For example, the proteins catalyze disulphide bond formation, which is a
step in the oxidative folding pathway and takes place in endoplasmic reticulum in eukaryotes
and in the bacterial periplasm in the prokaryotes. The structures in prokaryotes and eukaryotes
have several similarities and perform a similar function.
Finally, fold analysis can help in the design of new proteins with pre-defined structure
and activity. This can happen in many cases where a common evolutionary pattern gets
obscured by the extent of the divergence. In such cases, it is possible that the discovery of
new structures, with folds between those of the previously known structures, will make clear
their common evolutionary relationship.
There is considerable interest in assigning a sequence to a folding class. There are two
approaches used for this purpose. One approach uses threading algorithms which, given a
group of structures and a sequence, identify the structure that is closest to this sequence. The
second approach is taxonometric and assumes that the number of folds is restricted, and hence,
the objective is to predict in the context of a particular classification of three-dimensional
folds.
CATH ([Link] provides evolutionary relationships for protein
domains—it is a hierarchical classification of protein domains based on their folding patterns.
It consists of class, architecture, topology (folds) and homologous superfamily (Figures 18.11
and 18.12).
ISTUDY
Protein Structure Hierarchy 301
ISTUDY
302 Protein Structure Overview
Steric properties of amino acids can greatly affect the local structures that a protein
adopts. The best example of this is perhaps proline. Proline exhibits reduced torsional freedom
with the main chain phi angle fixed. This leads to the conformation of the peptide backbone
being locked in a turn, and with the loss of a hydrogen bonding in the N, this leads to the
residue often appearing on the surface loops of a protein.
Other particular types of secondary structure that have reasonable probabilities of having
particular types of amino acid that lend themselves to those structural configurations are given
in Table 18.7. However, one of the difficulties in protein structure prediction is that the context
in which amino acids find themselves in a protein has a large effect on their conformation
(phi and psi rotation), and thus, on their propensity to form particular secondary structures.
Tertiary Structure
The tertiary structure of protein is a three-dimensional representation of proteins. This
representation consists of a single polypeptide chain, forming the backbone of the structure
with one or more protein secondary structures, called the protein domains. The interactions
of side chains within a particular protein determine its tertiary structure. A number of tertiary
structures may fold into a quaternary structure. Generally, the information for protein structure
is contained within the amino acid sequence of the protein itself.
SWISS-MODEL ([Link] is an automated system for modeling
the three-dimensional structure of a protein from its amino acid sequence using homology
modeling techniques.
Tertiary structure represents spatial arrangement of amino acids that are far apart in
the linear sequence. It results from interactions between R-groups via van der Waals, ionic,
hydrophobic and hydrogen bonding. Groupings or arrangements of secondary structures into
combinations that are present in a variety of proteins are often described as motifs or domains.
Examples include:
ISTUDY
Protein Structure Hierarchy 303
Parallel and anti-parallel -sheets (also known as the hairpin beta-motif)
Greek key motif—four adjacent anti-parallel b-strands
Beta–alpha–beta motif (can be either two parallel or anti-parallel b)
Domains
Polypeptide chains of >200 amino acids that fold into two or more compact globular clusters
are called domains. There are three main types of domains:
a–domains are composed of a-helices.
b–domains contain anti-parallel b-sheets, and usually, contain two b-sheets packed
against each other.
a/b-domains contain the b-a-b motif of parallel b-sheets.
Adjacent domains are connected by one or two segments of the polypeptide chain.
All globular proteins have a defined structure. In the tertiary structure of proteins,
almost all the hydrophobic side chains are found in the interior of the protein and
almost all hydrophilic side chains are found on the outside of the protein, interacting with
water. Globular proteins are compact and there is no space inside. Hence, water is effectively
excluded from the hydrophobic interior. Nearly all hydrogen bond donors, e.g. Ser, form
hydrogen bonds with hydrogen bond acceptors like Gln. Hydrogen bond formation neutralizes
the polarity of the hydrogen-bonding group. b-sheets are usually twisted or wrapped into
barrel structures and loops and turns are on the outside of the protein. Mutations, which
place a hydrophobic side chain on the surface, can cause significant changes in the folding of
the protein.
Quaternary Structure
Many large proteins contain more than one polypeptide chain. The spatial arrangement of
these subunits is the quaternary structure of the protein (e.g., hemoglobin). The forces that
hold sub-units together are the same weak bonds as those that stabilize the tertiary structure
of proteins—van der Waals and London dispersion forces, salt bridges and hydrogen bonds.
The contact region between sub-units resembles the interior of a protein. The sub-units may
be identical or non-identical, but there is always a defined stoichiometry.
Several proteins function by forming transient complexes with other proteins. For example,
electron transfer from cytochrome c to cytochrome c peroxidase involves the formation of
such a complex. Because such complexes are transient, the interactions are weaker than
those that stabilize the quaternary structure of hemoglobin. In the cytochrome c-cytochrome
c peroxidase complex, there are specific electrostatic interactions that guide the two proteins
together in the correct orientation. The quaternary structure is not required for all proteins to
be functional; many proteins may have only 2º or 3º structure.
ISTUDY
304 Protein Structure Overview
CDART ([Link]
.cgi?cmd=rps)
CDART is the conserved domain architecture retrieval tool that can be used to search for
proteins with similar domain architectures. CDART uses pre-computed CD-search results to
quickly identify proteins with a set of domains similar to that of the query.
ISTUDY
Protein Classification Approaches 305
The demarcation of boundaries between these levels is to some degree subjective. The
evolutionary classification is generally conservative—where any doubt about relatedness exists,
we made new divisions at the family and superfamily levels. Thus, some researchers may
prefer to focus on the higher levels of the classification tree, where proteins with structural
similarity are clustered.
The different major levels in the hierarchy are:
SCOP
SCOP provides a broad survey of all known protein folds, detailed information about the close
relatives of any particular protein, and a framework for future research and classification. The
SCOP database aims to provide a detailed and comprehensive description of the structural and
evolutionary relationships between all proteins whose structure is known, including all entries
in the protein data bank. It is available as a set of tightly linked hypertext documents, which
make the large database comprehensible and accessible.
ISTUDY
306 Protein Structure Overview
CATH
CATH provides a hierarchical classification of protein domain structures, which clusters
proteins at four major levels. The homologous superfamilies cluster proteins with highly similar
structures and functions. The assignments of structures to topology families and homologous
superfamilies are made by sequence and structure comparisons.
SUMMARY
In this chapter, you have learned the fundamentals of protein structure. You have also learned
how to analyze the physical properties of proteins. The visualization of protein structure and
classification tools for proteins have also been discussed in detail. Using the discussion in this
chapter as the basis, you have learned how to predict protein structure and function from the
sequence information.
REVIEW QUESTIONS
1. Which tools are recommended for obtaining information about the primary structure of
protein?
2. How can you analyze physical properties of proteins?
3. Is it possible to predict the structure and function of a protein from sequence information?
4. Find out by sequence analysis whether the following protein contains multiple,
independently evolving, domains.
>SCYKLO40C * NFUI NfulpqSGD:SOOOI523 * GI:6322811 * chromosome 11
MFKSV AKLGK SPIFYLNSQR LIHIKTL TTP NENALKFLST DGEMLQTRGS
KSIVIKNillE NLINHSKLAQ QIFLQCPGVE SLMIGDDFL T INKDRMVHWN
SIKPEIIDLL TKQLAYGEDV ISKEFHAVQE EEGEGGYKIN MPKFELTEED
EEVSELIEEL IDTRIRP AIL EDGGDIDYRG WDPKTGTVYL RLQGACTSCS
SSEVTLKYGI ESMLKHYVDE VKEVIQIMDP EQEIALKEFD KLEKKLESSK
NTSHEK
5. Calmodulin is a protein that assists in modulation of other protein activity. Its structure
is variable and is modulated by calcium. Find the three-dimensional structure of
calmodulin in PDB. You can visualize these models using RasMol.
ISTUDY